Non-natural base pair compositions and methods of use
Engineered host cells with modified nucleoside triphosphate transporters and DNA repair proteins enable the incorporation of non-natural nucleotides, addressing the limitations of the natural genetic alphabet and enhancing information storage and gene expression in semi-synthetic organisms.
Patent Information
- Application Number
- JP2023182921
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-29
- Filing Date
- 2023-10-25
- Publication Date
- 2026-05-28
- Estimated Expiration
- 2038-12-28
AI Technical Summary
The limited chemical and physical diversity of the natural genetic alphabet restricts the ability to sequence-specifically synthesize and amplify oligonucleotides, limiting the information storage capacity in cells and hindering the creation of semi-synthetic organisms with enhanced gene expression products.
Engineered host cells are developed with modified nucleoside triphosphate transporters, transposition-associated proteins, and DNA repair proteins to incorporate non-natural nucleotides into nucleic acid molecules, enhancing the stability and retention of non-natural base pairs.
The incorporation of non-natural nucleotides into nucleic acid molecules increases the information storage capacity and stability of gene expression products, facilitating the development of semi-synthetic organisms.
Smart Images

Figure 0007866730000026 
Figure 0007866730000027 
Figure 0007866730000028
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the interests of U.S. Provisional Patent Application No. 62 / 612,062, filed on 29 December 2017, which is incorporated herein by reference in its entirety.
[0002] Description of research funded by the federal government. The inventions disclosed herein were made, at least in part, with the support of the U.S. Government under National Institutes of Health (NIH) Grant No. R35 GM118178 / GM / NIGMS. Accordingly, the U.S. Government has certain rights in these inventions.
[0003] Sequence List This application includes a sequence listing, which was submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The above ASCII copy, prepared on December 20, 2018, is named 46085-712_601_SL.txt and has a size of 116,287 bytes. [Background technology]
[0004] The application of the ability to sequence-specifically synthesize / amplify oligonucleotides (DNA or RNA) with polymerases is limited by the limited chemical / physical diversity present in the natural genetic alphabet (four natural nucleotides A, C, G, and T in DNA, and four natural nucleotides A, C, G, and U in RNA). An extended genetic alphabet including non-natural nucleic acids increases the amount of information that can be stored in cells, facilitating the creation of semi-synthetic organisms (SSOs) that utilize this increased information to create new forms of gene expression products. [Overview of the project] [Problems that the invention aims to solve]
[0005] In certain embodiments, methods, cells, engineered microorganisms, plasmids, and kits for increasing the production of nucleic acid molecules containing non-natural nucleotides are described herein. In some embodiments, cells, engineered microorganisms, plasmids, and methods of use utilizing modified transfer-related proteins, modified DNA repair proteins, or combinations thereof for increasing the production of nucleic acid molecules containing non-natural nucleotides are also described herein. [Means for solving the problem]
[0006] Embodiments disclosed herein provide an engineered host cell comprising a first nucleic acid molecule containing a non-native nucleotide; and optionally, a second nucleic acid molecule encoding a modified translocation-associated protein or transposable element. In some embodiments, the engineered host cell further comprises a third nucleic acid molecule encoding a modified nucleoside triphosphate transporter, which is incorporated into the genome sequence of the engineered host cell or comprises a plasmid encoding the modified nucleoside triphosphate transporter. In some embodiments, the modified nucleoside triphosphate transporter exhibits higher stability of expression in the engineered host cell compared to expression in an equivalent engineered host cell that does not contain the second nucleic acid molecule encoding the modified translocation-associated protein. In some embodiments, the modified nucleoside triphosphate transporter comprises deletion of the entire nucleic acid molecule encoding the nucleoside triphosphate transporter, N-terminal cleavage, C-terminal cleavage, or cleavage of both ends. In some embodiments, the modified nucleoside triphosphate transporter is derived from Pheodactylum tricornuum. The modified nucleoside triphosphate transporter (PtNTT2) is derived from *Tricornutum*. In some embodiments, the modified nucleoside triphosphate transporter includes deletions. In some embodiments, the deletions are terminal deletions or internal deletions. In some embodiments, the deletions include N-terminal cleavage, C-terminal cleavage, or bi-terminal cleavage. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues at the N-terminus. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 66 amino acid residues at the N-terminus. In some embodiments, PtNTT2 is under the control of a promoter selected from a pSC plasmid or a promoter derived from a lac operon. In some embodiments, the engineered host cell further comprises a Cas9 polypeptide or a variant thereof, and a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold, and the combination of the Cas9 polypeptide or a variant thereof and the sgRNA regulates the replication of a first nucleic acid molecule encoding a non-native nucleotide. In some embodiments, the sgRNA contains a target motif that recognizes the modification at a non-native nucleotide site within the nucleic acid molecule. In some embodiments, the sgRNA further comprises a protospacer-adjacent motif (PAM) recognition factor. In some embodiments, the PAM factor is adjacent to the 3' end of the target motif. In some embodiments, the target motif is 15–30 nucleotides long. In some embodiments, the combination of the Cas9 polypeptide or a variant thereof and the sgRNA reduces the replication rate of the nucleic acid molecule containing the modification by about 80%, 85%, 95%, 99%, or more. In some embodiments, the Cas9 polypeptide is wild-type Cas9.In some embodiments, the second nucleic acid molecule comprises a gene containing catalase (cat), IS1 protein insB-4 (insB-4), IS1 protein insA-4 (insA-4), or a combination thereof. In some embodiments, the modified transposterior-related protein comprises insertion factor IS1 4 protein InsB, insertion factor IS1 4 protein InsA, or a combination thereof; the modified transposterior factor comprises IS1. In some embodiments, the gene comprises one or more deletions, one or more of which include N-terminal deletions, C-terminal deletions, cleavage at both ends, internal deletions, and / or deletions of the entire gene. In some embodiments, the engineered host cell further comprises a fifth nucleic acid molecule encoding a modified DNA repair response-related protein, the DNA repair response comprising recombination repair, SOS response, nucleotide excision repair, or methyl-directed mismatch repair, or a combination thereof. In some embodiments, the modified DNA repair response-related protein comprises RecA, Rad51, RadA, or LexA, or a combination thereof. In some embodiments, the manipulated host cells are prokaryotic cells, including Escherichia coli cells and Escherichia coli BL21(DE3) cells.In some embodiments, non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I ), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihy. Drocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azouracil, 6-azocytosine, azacytosine, 6-azothymine, 6-thioguanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-de Aza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenine and guanine; N4-ethylcytosine, N-2-substituted purines, N-6-substituted purines, O-6-substituted purines, enhancers of double-strand formation stability, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazinecytidine ([5,4-b [1,4]benzoxazine-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindole cytidine (H-pyrimido[3',2':4,5]pyrrolo[2,3 -d] Pyrimidine-2-one), 5-Fluorouracil, 5-Bromouracil, 5-Chlorouracil, 5-Iodouracil, Hypoxanthine, Xanthine, 4-Acetylcytosine, 5-(Carboxyhydroxylmethyl)uracil, 5-Carboxymethylaminomethyl-2-thiouridine, 5-Carboxymethylaminomethyluracil, Dihydrouracil, β-D-Galactosylqueosin, Inosine, N6-Isopentenyladenine, 1-Methylguanine, 1-Methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylcuosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weibtoxosin, pseudouracil, cue The non-natural bases are selected from the group consisting of osin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, as well as those in which a purine or pyrimidine group is replaced by a heterocyclic group. In some embodiments, the non-natural bases are, [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl L, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10Alkynyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, as selected from the group consisting of any combination thereof. In some embodiments, the engineered host cell further comprises polymerase. In some embodiments, polymerase is constitutively expressed. In some embodiments, polymerase is overexpressed. In some embodiments, the polymerase is DNA polymerase. In some embodiments, the DNA polymerase is DNA polymerase II. In some embodiments, the polymerase is encoded by the polB gene. In some embodiments, the polB gene is derepressed. In some embodiments, the polB gene is derepressed through integration across half a site of the operator. In some embodiments, the operator is the lexA operator. In some embodiments, the polymerase is DNA polymerase I. In some embodiments, the polymerase is encoded by the polA gene. In some embodiments, the polymerase is DNA polymerase III. In some embodiments, the polymerase is encoded by the dnaQ gene.
[0007] Embodiments disclosed herein provide a method for increasing the production of nucleic acid molecules, including non-natural nucleotides, comprising: incubating engineered host cells with a plurality of non-natural nucleotides, comprising a modified nucleoside triphosphate transporter and optionally a modified transposition-associated protein or transposable element; and incorporating the plurality of non-natural nucleotides into one or more newly synthesized DNA strands, thereby generating a non-natural nucleic acid molecule, wherein the modified transposition-associated protein or transposable element and the modified nucleoside triphosphate transporter enhance the retention of non-natural base pairs, including non-natural nucleotides, in one or more newly synthesized DNA strands. In some embodiments, the modified transposition-associated protein comprises insertion factor IS1 4 protein InsB, insertion factor IS1 4 protein InsA, or a combination thereof; the modified transposable element comprises IS1. In some embodiments, the modified nucleoside triphosphate transporter comprises a codon-optimized nucleoside triphosphate transporter (PtNTT2) derived from Pheodactylum tricornutum. In some embodiments, the modified nucleoside triphosphate transporter includes deletions. In some embodiments, the deletions are terminal deletions or internal deletions. In some embodiments, the deletions are N-terminal cleavage, C-terminal cleavage, or bi-terminal cleavage. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues at the N-terminus. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 66 amino acid residues at the N-terminus. In some embodiments, the engineered host cell is Cas9 polypeptide or its barrier The Cas9 polypeptide or a variant thereof, and the sgRNA combination further comprises a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold, and regulates the replication of a first nucleic acid molecule encoding a non-native nucleotide. In some embodiments, the sgRNA contains a target motif that recognizes the modification at a non-native nucleotide site within the nucleic acid molecule. In some embodiments, the sgRNA further comprises a protospacer-adjacent motif (PAM) recognition factor. In some embodiments, the PAM factor is adjacent to the 3' end of the target motif. In some embodiments, the target motif is 15–30 nucleotides long. In some embodiments, the Cas9 polypeptide or a variant thereof, and the sgRNA combination reduces the replication rate of the nucleic acid molecule containing the modification by about 80%, 85%, 95%, 99%, or more. In some embodiments, the Cas9 polypeptide is wild-type Cas9. In some embodiments, non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine;5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-Fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azo-uracil, 6-azocytosine, azacytosine, 6-azothymine, 6-thio-guanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenine and guanine;N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazinecytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one) (n), G-clamps, phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindole cytidine (H-pyrimido[3',2':4,5]pyrrolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5- Iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosyl quosine, inosine, N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methyl Ruguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylcuosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weibtoxosin, pseudouracil, cuosin, 2-thiositol; The non-natural bases are selected from the group consisting of syn, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, as well as those in which a purine or pyrimidine group is replaced by a heterocyclic group. In some embodiments, the non-natural base is [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10Alkynyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, as selected from the group consisting of any combination thereof. In some embodiments, the engineered host cell further comprises polymerase. In some embodiments, polymerase is constitutively expressed. In some embodiments, polymerase is overexpressed. In some embodiments, the polymerase is DNA polymerase. In some embodiments, the DNA polymerase is DNA polymerase II. In some embodiments, the polymerase is encoded by the polB gene. In some embodiments, the polB gene is derepressed. In some embodiments, the polB gene is derepressed through integration across half a site of the operator. In some embodiments, the operator is the lexA operator. In some embodiments, the polymerase is DNA polymerase I. In some embodiments, the polymerase is encoded by the polA gene. In some embodiments, the polymerase is DNA polymerase III. In some embodiments, the polymerase is encoded by the dnaQ gene.
[0008] Embodiments disclosed herein provide a method for preparing a modified polypeptide comprising non-natural amino acids, comprising incubating an engineered host cell comprising a modified nucleoside triphosphate transporter and optionally a modified transposition-associated protein or transposable molecule with a plurality of non-natural nucleotides; and incorporating the plurality of non-natural nucleotides into one or more newly synthesized DNA strands, thereby generating a non-natural nucleic acid molecule, wherein the modified transposition-associated protein or transposable molecule and the modified nucleoside triphosphate transporter facilitate the incorporation of the plurality of non-natural nucleotides into the newly synthesized polypeptide, thereby increasing the retention of non-natural base pairs that produce the modified polypeptide. In some embodiments, the modified transposition-associated protein comprises insertion factor IS1 4 protein InsB, insertion factor IS1 4 protein InsA, or a combination thereof; the modified transposable molecule comprises IS1. In some embodiments, the modified nucleoside triphosphate transporter comprises a codon-optimized nucleoside triphosphate transporter (PtNTT2) derived from Pheodactylum tricornutum. In some embodiments, the modified nucleoside triphosphate transporter includes deletions. In some embodiments, the deletions are terminal deletions or internal deletions. In some embodiments, the deletions include N-terminal cleavage, C-terminal cleavage, or bi-terminal cleavage. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues at the N-terminus. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 66 amino acid residues at the N-terminus. In some embodiments, the manipulated host cell further comprises a Cas9 polypeptide or a variant thereof, and a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold.A combination of the Cas9 polypeptide or a variant thereof, and an sgRNA, regulates the replication of a first nucleic acid molecule encoding a non-native nucleotide. In some embodiments, the sgRNA includes a target motif that recognizes the modification at a non-native nucleotide site within the nucleic acid molecule. In some embodiments, the sgRNA further includes a protospacer-adjacent motif (PAM) recognition factor. In some embodiments, the PAM factor is adjacent to the 3' end of the target motif. In some embodiments, the target motif is 15–30 nucleotides long. In some embodiments, the combination of the Cas9 polypeptide or a variant thereof, and an sgRNA, reduces the replication rate of the nucleic acid molecule containing the modification by about 80%, 85%, 95%, 99%, or more. In some embodiments, the Cas9 polypeptide is wild-type Cas9. In some embodiments, non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil Cyl and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azouracil,6-Azocytosine, Azacytosine, 6-Azothymine, 6-Thio-Guanine, 7-Methylguanine, 7-Methyladenine, 7-Deazaguanine, 7-Deazaguanosine, 7-Deaza-Adenine, 7-Deaza-8-Azaguanine, 8-Azaguanine, 8-Azaadenine, 8-Halo, 8-Amino, 8-Thiol, 8-Thioalkyl, and 8-Hydroxyl-substituted adenines and guanines; N4-Ethylcytosine, N-2-substituted purines, N-6-substituted purines, O-6-substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nuclei Acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindole cytidine (H-pyrimido[3' [2':4,5]pyrrolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueusin, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyl Adenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylquosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weibtoxosin, pseudouracil, quosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil,It contains unnatural bases selected from the group consisting of methyl uracil-5-oxyacetate, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, and those in which the purine or pyrimidine group is replaced by a heterocyclic ring. In some embodiments, the unnatural base is [Chemical formula] selected from the group consisting of. In some embodiments, the unnatural nucleotide has a modification at the 2'-position: OH; substituted lower alkyl, arylalkyl, aralkyl, O-aryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'- OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 alkyl, C2-C 10 alkenyl, C2-C 10The non-natural sugar moieties are selected from the group consisting of alkynyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, as well as any combination thereof. In some embodiments, the engineered host cell further comprises a polymerase. In some embodiments, the polymerase is constitutively expressed. In some embodiments, the polymerase is overexpressed. In some embodiments, the polymerase is DNA polymerase. In some embodiments, the DNA polymerase is DNA polymerase II. In some embodiments, the polymerase is encoded by the polB gene. In some embodiments, the polB gene is derepressed. In some embodiments, the polB gene is derepressed through integration across half a site of the operator. In some embodiments, the operator is the lexA operator. In some embodiments, the polymerase is DNA polymerase I. In some embodiments, the polymerase is encoded by the polA gene. In some embodiments, the polymerase is DNA polymerase III. In some embodiments, the polymerase is encoded by the dnaQ gene.
[0009] Embodiments disclosed herein provide engineered host cells for producing non-natural products comprising a modified DNA repair response-related protein. In some embodiments, the DNA repair response includes recombinant repair. In some embodiments, the DNA repair response includes an SOS response. In some embodiments, the engineered host cell is a prokaryotic cell, a eukaryotic cell, or a yeast cell. In some embodiments, the engineered host cell is a prokaryotic cell. In some embodiments, the prokaryotic cell is an E. coli cell. In some embodiments, the E. coli cell is an E. coli BL21(DE3) cell. In some embodiments, the modified DNA repair response-related protein is RecA. In some embodiments, the engineered host cell is engineered to express a gene encoding RecA. In some embodiments, the modified DNA repair response-related protein is Rad51. In some embodiments, the engineered host cell is engineered to express a gene encoding Rad51. In some embodiments, the modified DNA repair response-related protein is RadA. In some embodiments, the modified DNA repair response-related protein is LexA. In some embodiments, the gene encoding the modified DNA repair response-related protein contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, the gene contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA, rad51, and / or radA contain one or more mutations, one or more deletions, or a combination thereof. In some embodiments, recA, rad51, and radA each independently contain an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an internal deletion at residues 2-347. In some embodiments, lexA contains one or more mutations, one or more deletions, or a combination thereof.In some embodiments, lexA includes a mutation at amino acid position S119, possibly an S119A mutation. In some embodiments, the manipulated host cells are polymers. The polymerase further comprises [a specific gene]. In some embodiments, the polymerase is constitutively expressed. In some embodiments, the polymerase is overexpressed. In some embodiments, the polymerase is DNA polymerase. In some embodiments, the DNA polymerase is DNA polymerase II. In some embodiments, the polymerase is encoded by the polB gene. In some embodiments, the polB gene is derepressed. In some embodiments, the polB gene is derepressed through integration across half a site of the operator. In some embodiments, the operator is the lexA operator. In some embodiments, the polymerase is DNA polymerase I. In some embodiments, the polymerase is encoded by the polA gene. In some embodiments, the polymerase is DNA polymerase III. In some embodiments, the polymerase is encoded by the dnaQ gene.
[0010] Embodiments disclosed herein provide engineered host cells for producing non-natural products comprising modified DNA repair response-related proteins and polymerases, wherein the engineered host cells exhibit elevated expression of the polymerase compared to equivalent host cells comprising equivalent polymerases at basic expression levels. In some embodiments, the DNA repair response comprises recombinant repair. In some embodiments, the DNA repair response comprises an SOS response. In some embodiments, the polymerase is constitutively expressed. In some embodiments, the polymerase is DNA polymerase II. In some embodiments, the DNA repair response comprises recombinant repair, an SOS response, nucleotide excision repair, or methyl-directed mismatch repair. In some embodiments, the DNA repair response comprises recombinant repair. In some embodiments, the DNA repair response comprises an SOS response. In some embodiments, the engineered host cell is a prokaryotic cell, a eukaryotic cell, or a yeast cell. In some embodiments, the engineered host cell is a prokaryotic cell. In some embodiments, the prokaryotic cell is an E. coli cell. In some embodiments, the E. coli cell is an E. coli BL21(DE3) cell. In some embodiments, the modified DNA repair response-related protein is RecA. In some embodiments, the modified DNA repair response-related protein is Rad51. In some embodiments, the modified DNA repair response-related protein is RadA. In some embodiments, the modified DNA repair response-related protein is LexA. In some embodiments, the gene encoding the defect protein includes one or more mutations, one or more deletions, or a combination thereof. In some embodiments, the gene includes an N-terminal deletion, a C-terminal deletion, a break at both ends, or an internal deletion. In some embodiments, recA, rad51, and / or radA include one or more mutations, one or more deletions, or a combination thereof. In some embodiments, recA, rad51, and radA each independently include an N-terminal deletion, a C-terminal deletion, a break at both ends, or an internal deletion.In some embodiments, recA includes an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA includes an internal deletion at residues 2–347. In some embodiments, lexA includes one or more mutations, one or more deletions, or a combination thereof. In some embodiments, lexA includes a mutation at amino acid position S119, optionally an S119A mutation. In some embodiments, the engineered host cell further comprises a nucleoside triphosphate transporter (PtNTT2) derived from Pheodactylum tricornuate. In some embodiments, the nucleoside triphosphate transporter derived from PtNTT2 is modified. In some embodiments, the modified nucleoside triphosphate transporter is encoded by a nucleic acid molecule. In some embodiments, the nucleic acid molecule encoding the modified nucleoside triphosphate transporter is incorporated into the genome sequence of the engineered host cell. In some embodiments, the engineered host cell comprises a plasmid containing a nucleic acid molecule encoding the modified nucleoside triphosphate transporter. In some embodiments, the modified nucleoside triphosphate transporter is Pheodactylum trichophosphate. This is a codon-optimized nucleoside triphosphate transporter derived from Lunutum. In some embodiments, the modified nucleoside triphosphate transporter includes deletions. In some embodiments, the deletions are terminal deletions or internal deletions. In some embodiments, the deletions include N-terminal cleavage, C-terminal cleavage, or bi-terminal cleavage. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 5, 10, 15, 20, 22, 25, 30, 40, 44, 50, 60, 66, 70 or more amino acid residues at the N-terminus. In some embodiments, the modified nucleoside triphosphate transporter includes deletions of about 66 amino acid residues at the N-terminus. In some embodiments, the modified nucleoside triphosphate transporter is under the control of a promoter selected from a pSC plasmid or a promoter derived from a lac operon. In some embodiments, the lac operon is an E. coli lac operon. In some embodiments, the lac operon is a P bla , P lac , P lacUV5 , P H207 , P λ , P tac , or P N25 Selected from. In some embodiments, the modified nucleoside triphosphate transporter is promoter P lacUV5It is under the control of [unclear]. In some embodiments, the engineered host cell further comprises a Cas9 polypeptide or a variant thereof, and a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold, and the combination of the Cas9 polypeptide or a variant thereof and the sgRNA regulates the replication of nucleic acid molecules containing non-native nucleotides. In some embodiments, the sgRNA comprises a target motif that recognizes modifications at non-native nucleotide sites within the nucleic acid molecule. In some embodiments, the sgRNA further comprises a protospacer-adjacent motif (PAM) recognition factor. In some embodiments, the PAM factor is adjacent to the 3' end of the target motif. In some embodiments, the target motif is 15 to 30 nucleotides long. In some embodiments, the combination of the Cas9 polypeptide or a variant thereof and the sgRNA reduces the replication rate of nucleic acid molecules containing modifications by about 80%, 85%, 95%, 99%, or more. In some embodiments, the Cas9 polypeptide is wild-type Cas9. In some embodiments, the engineered host cell further comprises non-native nucleotides.In some embodiments, the non-natural nucleotides are 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil Cyl, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methyl-cytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5- Chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azouracil, 6-azocytosine, azacytosine N, 6-azo-thymine, 6-thio-guanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenines and guanines; N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, those that enhance the stability of double-chain formation, and uni. Versal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrim (4,5-b)indole-2-one), pyridoindolecytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylquosine Inosine, N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylquosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyl adenine, uracil-5-oxyacetic acid The non-natural bases are selected from the group consisting of , [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10 Alkynnyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to approximately 10; The group consists of and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position, 4'-S, heterocycloalkyl, heterocycloalkalyl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, and any combination thereof. In some embodiments, the non-natural base is selected from the group. [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural backbone. In some embodiments, the non-natural backbone comprises phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkyl phosphotryesters, C1-C 10The non-natural nucleotide is selected from the group consisting of phosphonates, 3'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, and boranophosphates. In some embodiments, the non-natural nucleotide is dNaMTP and / or dTPT3TP. In some embodiments, the non-natural nucleotide is incorporated into the manipulated host cell genome. In some embodiments, the non-natural nucleotide is incorporated into a chromosome. In some embodiments, the non-natural nucleotide is incorporated into the arsB locus. In some embodiments, the engineered host cells enable retention of approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more non-native base pairs compared to equivalent engineered host cells in the absence of the modified DNA repair response-related protein or in the absence of the modified DNA repair response-related protein combined with an overexpressed polymerase. In some embodiments, the engineered host cells enable retention of at least 50% non-native base pairs after more than 50, more than 100, more than 120, more than 130, more than 150, or more than 200 passages. In some embodiments, the engineered host cell allows for the retention of at least 55% of non-natural base pairs after more than 50, more than 100, more than 120, more than 130, more than 137, more than 150, or more than 200 passages. In some embodiments, the non-natural product is a nucleic acid molecule containing non-natural nucleotides. In some embodiments, the non-natural product is a polypeptide containing non-natural amino acids. In some embodiments, the engineered host cell is a semi-synthetic organism.
[0011] The embodiments disclosed herein provide nucleic acid molecules comprising non-natural nucleotides produced by the manipulated host cells described herein.
[0012] Embodiments disclosed herein provide polypeptides comprising one or more non-natural amino acids produced by the manipulated host cells described herein.
[0013] Aspects disclosed herein are methods for increasing the replication fidelity of nucleic acid molecules containing non-natural nucleotides, wherein (a) the manipulated host cells described herein are subjected to multiple non-natural nucleotides The method provides a way to increase the fidelity of replication of non-natural base pairs containing non-natural nucleotides in one or more newly synthesized DNA strands, comprising (b) incubation with ocide; and (b) incorporating multiple non-natural nucleotides into one or more newly synthesized DNA strands, thereby generating a non-natural nucleic acid molecule, wherein the modified DNA repair response-related protein and optionally overexpressed polymerase enhance the fidelity of replication of non-natural base pairs containing non-natural nucleotides in one or more newly synthesized DNA strands. In some embodiments, the DNA repair response includes recombination repair. In some embodiments, the DNA repair response includes an SOS response. In some embodiments, the increased production of nucleic acid molecules containing non-natural nucleotides is associated with the production of nucleic acid molecules in equivalent host cells in the absence of the modified DNA repair response-related protein and optionally overexpressed polymerase. In some embodiments, the increased production of nucleic acid molecules is at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 99% higher than the production of nucleic acid molecules in equivalent host cells in the absence of the modified DNA repair response-related protein and optionally overexpressed polymerase. In some embodiments, the increase in nucleic acid molecule production is more than 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 100 or more than the production of nucleic acid molecules in equivalent host cells in the absence of modified DNA repair response-related proteins and optionally overexpressed polymerases. In some embodiments, the increase in nucleic acid molecule production is 1 to 5 times, 5 to 10 times, 10 to 15 times, 15 to 20 times, 20 to 25 times, 25 to 30 times, 30 to 40 times, 40 to 50 times, 50 to 60 times, 60 to 70 times, 70 to 80 times, 80 to 90 times, 90 to 100 times, or 100 to 200 times higher than the production of nucleic acid molecules in equivalent host cells in the absence of modified DNA repair response-related proteins and optionally overexpressed polymerases. In some embodiments, non-natural nucleotides are 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-amino-adenine, 2-amino-propyl-adenine,2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thiothymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyri Midine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azouracil, 6-azocytosine, azacy Tosine, 6-azo-thymine, 6-thio-guanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenines and guanines; N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, Fluorinated nucleic acids, tricyclic pyrimidines, phenoxadine cytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxadine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindole cytidine (H-pyrimido[3',2':4,5]pyrrolo, [2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueucine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl- The non-natural bases are selected from the group consisting of 2-thiouracil, β-D-mannosylquosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weybutoxosin, pseudouracil, quosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, as well as those in which a purine or pyrimidine group is replaced by a heterocyclic group. In some embodiments, the non-natural bases are [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10 Alkynnyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, and any combination thereof are selected from the group. In some embodiments, non-natural bases are, [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural backbone. In some embodiments, the non-natural backbone comprises phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkyl phosphotryesters, C1-C 10The non-natural nucleotide is selected from the group consisting of phosphonates, 3'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, and boranophosphates. In some embodiments, the non-natural nucleotide is dNaMTP and / or dTPT3TP. In some embodiments, the non-natural nucleotide is incorporated into the engineered host cell genome. In some embodiments, the non-natural nucleotide is incorporated into the chromosome. In some embodiments, the non-natural nucleotide is incorporated into the arsB locus. In some embodiments, the modified DNA repair response-related protein is RecA. In some embodiments, the modified DNA repair response-related protein is Rad51. In some embodiments, the modified DNA repair response-related protein is RadA. In some embodiments, the modified DNA repair response-related protein is LexA. In some embodiments, the gene encoding the modified DNA repair response-related protein contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, the gene contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA, rad51, and / or radA contain one or more mutations, one or more deletions, or a combination thereof. In some embodiments, recA, rad51, and radA each independently contain an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an internal deletion at residues 2-347. In some embodiments, lexA contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, lexA contains a mutation at amino acid position S119, possibly an S119A mutation.
[0014] Embodiments disclosed herein provide a method for increasing the production of nucleic acid molecules containing non-natural nucleotides, comprising (a) incubating a manipulated host cell described herein with a plurality of non-natural nucleotides; and (b) incorporating the plurality of non-natural nucleotides into one or more newly synthesized DNA strands, thereby generating a non-natural nucleic acid molecule, wherein a modified DNA repair response-related protein and optionally overexpressed polymerase enhance the retention of non-natural base pairs containing non-natural nucleotides in one or more newly synthesized DNA strands. In some embodiments, the DNA repair response includes recombinant repair. In some embodiments, the DNA repair response includes an SOS response. In some embodiments, the increased production of nucleic acid molecules containing non-natural nucleotides is achieved by a modified DNA repair response-related protein and optionally overexpressed polymerase. This relates to the production of nucleic acid molecules in equivalent host cells in the absence of the polymerase. In some embodiments, the increase in nucleic acid molecule production is at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 99% higher than the production of nucleic acid molecules in equivalent host cells in the absence of the modified DNA repair response-related protein and optionally overexpressed polymerase. In some embodiments, the increase in nucleic acid molecule production is more than 1x, 2x, 3x, 4x, 5x, 10x, 15x, 20x, 25x, 30x, 40x, 50x, 100x or more than the production of nucleic acid molecules in equivalent host cells in the absence of the modified DNA repair response-related protein and optionally overexpressed polymerase. In some embodiments, the increase in nucleic acid molecule production is 1 to 5 times, 5 to 10 times, 10 to 15 times, 15 to 20 times, 20 to 25 times, 25 to 30 times, 30 to 40 times, 40 to 50 times, 50 to 60 times, 60 to 70 times, 70 to 80 times, 80 to 90 times, 90 to 100 times, or 100 to 200 times higher than the production of nucleic acid molecules in equivalent host cells in the absence of modified DNA repair response-related proteins and optionally overexpressed polymerases. In some embodiments, non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine;5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-Fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azo-uracil, 6-azocytosine, azacytosine, 6-azothymine, 6-thio-guanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenine and guanine;N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazinecytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazinecytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindolecytidine (H-pyrimido[3',2': 4,5]pyrrolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosyl quosine, inosine, N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thio; The non-natural bases are selected from the group consisting of uracil, β-D-mannosylquosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weybutoxosin, pseudouracil, quosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, as well as those in which a purine or pyrimidine group is replaced by a heterocyclic group. In some embodiments, the non-natural bases are [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10Alkynnyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, and any combination thereof are selected from the group. In some embodiments, non-natural bases are, [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural backbone. In some embodiments, the non-natural backbone comprises phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkyl phosphotryesters, C1-C 10The non-natural nucleotide is selected from the group consisting of phosphonates, 3'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, and boranophosphates. In some embodiments, the non-natural nucleotide is dNaMTP and / or dTPT3TP. In some embodiments, the non-natural nucleotide is incorporated into the engineered host cell genome. In some embodiments, the non-natural nucleotide is incorporated into the chromosome. In some embodiments, the non-natural nucleotide is incorporated into the arsB locus. In some embodiments, the modified DNA repair response-related protein is RecA. In some embodiments, the modified DNA repair response-related protein is Rad51. In some embodiments, the modified DNA repair response-related protein is RadA. In some embodiments, the modified DNA repair response-related protein is LexA. In some embodiments, the gene encoding the modified DNA repair response-related protein contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, the gene contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA, rad51, and / or radA contain one or more mutations, one or more deletions, or a combination thereof. In some embodiments, recA, rad51, and radA each independently contain an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an internal deletion at residues 2-347. In some embodiments, lexA contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, lexA contains a mutation at amino acid position S119, possibly an S119A mutation.
[0015] Embodiments disclosed herein provide a method for preparing a modified polypeptide containing non-natural amino acids, comprising (a) incubating a manipulated host cell described herein with a plurality of non-natural amino acids; and (b) incorporating the plurality of non-natural amino acids into a newly synthesized polypeptide to produce a modified polypeptide, wherein a modified DNA repair response-related protein and optionally an overexpressed polymerase facilitate the incorporation of the plurality of non-natural amino acids into the newly synthesized polypeptide, thereby increasing the retention of non-natural base pairs to produce the modified polypeptide. In some embodiments, the DNA repair response includes recombination repair. In some embodiments, the DNA repair response includes an SOS response. In some embodiments, the modified polypeptide is further coupled with a conjugated moiety to produce a modified polypeptide conjugate. In some embodiments, the conjugated moiety is a protein or its binding fragment, a polymer, a therapeutic agent, a contrast agent, or a combination thereof. In some embodiments, the modified polypeptide is further conjugated with a therapeutic agent. In some embodiments, the modified polypeptide is a contrast agent. In some embodiments, the modified polypeptide conjugate is further formulated with pharmaceutical additives to produce a pharmaceutical composition. In some embodiments, non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-Dihydrocytosine, 5-Iodocytosine, Hydroxyurea, Iodouracil, 5-Nitrocytosine, 5-Bromouracil, 5-Chlorouracil, 5-Fluorouracil, and 5-Iodouracil, 6-alkyl derivatives of adenine and guanine, 6-Azapyrimidine, 6-Azouracil, 6-Azocytosine, Azacytosine, 6-Azothymine, 6-Thio-Guanine, 7-Methylguanine, 7-Methyladenine, 7-Deazaguanine, 7-Deazaguanosine, 7-Deaza-A Denine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenine and guanine; N4-ethylcytosine, N-2-substituted purines, N-6-substituted purines, O-6-substituted purines, enhancers of double-strand formation stability, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1, 4)Benzoxazine-2(3H)-one), Phenothiazinecytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, Phenoxazinecytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), Carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), Pyridindolecytidine (H-pyrimido[3',2':4,5]pyrrolo[2,3-d ]Pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosyl quosine, inosine, N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylcuosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weibtoxosin, pseudouracil, cue The non-natural bases are selected from the group consisting of osin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, as well as those in which a purine or pyrimidine group is replaced by a heterocyclic group. In some embodiments, the non-natural bases are, [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10Alkynnyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3-O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, and any combination thereof are selected from the group. In some embodiments, non-natural bases are [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural backbone. In some embodiments, the non-natural backbone comprises phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkyl phosphotryesters, C1-C 10 From the group consisting of phosphonates, 3'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, 3'-aminophosphoamidates, aminoalkylphosphoamidates, thionophosphoamidates, thionoalkylphosphonates, thionoalkylphosphotryesters, and boranophosphates Selected. In some embodiments, the non-native nucleotide is dNaMTP and / or dTPT3TP. In some embodiments, the non-native nucleotide is incorporated into the engineered host cell genome. In some embodiments, the non-native nucleotide is incorporated into a chromosome. In some embodiments, the non-native nucleotide is incorporated into the arsB locus. In some embodiments, the modified DNA repair response-related protein is RecA. In some embodiments, the modified DNA repair response-related protein is Rad51. In some embodiments, the modified DNA repair response-related protein is RadA. In some embodiments, the modified DNA repair response-related protein is LexA. In some embodiments, the gene encoding the modified DNA repair response-related protein contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, the gene contains an N-terminal deletion, a C-terminal deletion, a break at both ends, or an internal deletion. In some embodiments, recA, rad51, and / or radA contain one or more mutations, one or more deletions, or a combination thereof. In some embodiments, recA, rad51, and radA each independently include an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA includes an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA includes an internal deletion at residues 2-347. In some embodiments, lexA includes one or more mutations, one or more deletions, or a combination thereof. In some embodiments, lexA includes a mutation at amino acid position S119, possibly an S119A mutation.
[0016] The embodiments disclosed herein provide a method for treating a disease or symptom, comprising administering to a subject in need a pharmaceutical composition comprising a modified polypeptide prepared by the method disclosed herein, thereby treating the disease or symptom.
[0017] Embodiments disclosed herein provide a kit comprising the manipulated host cells described herein.
[0018] Embodiments disclosed herein provide engineered host cells for producing non-natural products containing modified RecA. In some embodiments, the gene encoding modified RecA contains one or more mutations, one or more deletions, or a combination thereof. In some embodiments, the gene contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, or an internal deletion. In some embodiments, recA contains an internal deletion of residues 2-347.
[0019] Embodiments disclosed herein provide engineered host cells for producing non-natural products comprising modified RecA and overexpressed DNA polymerase II, wherein the expression level of overexpressed DNA polymerase II is relative to an equivalent host cell comprising DNA polymerase II at a basic expression level.
[0020] Embodiments disclosed herein are methods for increasing the production of nucleic acid molecules containing non-natural nucleotides, comprising: (a) incubating engineered host cells containing modified RecA and optionally overexpressed DNA polymerase II with a plurality of non-natural nucleotides, wherein the expression level of the overexpressed DNA polymerase II is relative to an equivalent host cell containing an equivalent level of DNA polymerase II; and (b) incorporating the plurality of non-natural nucleotides into one or more newly synthesized DNA strands, thereby generating a non-natural nucleic acid molecule. The expressed DNA repair response-related proteins and optionally overexpressed polymerases provide a method for enhancing the retention of non-native base pairs, including non-native nucleotides, in one or more newly synthesized DNA strands.
[0021] Embodiments disclosed herein provide a method for preparing a modified polypeptide comprising non-natural amino acids, comprising (a) incubating an engineered host cell comprising a modified RecA and optionally overexpressed DNA polymerase II with a plurality of non-natural amino acids, wherein the expression level of the overexpressed DNA polymerase II is related to an equivalent host cell comprising an equivalent basic expression level of DNA polymerase II; and (b) incorporating the plurality of non-natural amino acids into a newly synthesized DNA strand, thereby producing a modified polypeptide, wherein the modified DNA repair response-related protein and optionally overexpressed polymerase facilitate the incorporation of the plurality of non-natural amino acids into the newly synthesized polypeptide, thereby increasing the retention of non-natural base pairs that produce the modified polypeptide. In some embodiments, the DNA repair response comprises recombination repair. In some embodiments, the DNA repair response comprises an SOS response. In some embodiments, non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thio-thymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine;5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-Fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azo-uracil, 6-azocytosine, azacytosine, 6-azothymine, 6-thio-guanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenine and guanine;N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazinecytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazinecytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indo (H-pyridindolecytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosyl quosine, inosine, N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-; Methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylcuosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, weibtoxosin, pseudouracil, cuosin, 2-thiocytosine, 5-methyl This includes non-natural bases selected from the group consisting of thiouracil-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine, as well as those in which a purine or pyrimidine group is replaced by a heterocyclic group.
[0022] In some embodiments, the non-natural base is [ka] Selected from the group consisting of
[0023] In some embodiments, the non-natural nucleotide further comprises a non-natural sugar moiety. In some embodiments, the non-natural sugar moiety is modified at the 2' position:OH; substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2~C 10 Alkenyl, C2~C 10 Alkynnyl, -O[(CH2)nO]mCH3, -O(CH2)nOCH3, -O(CH2)nNH2, -O(CH2)nCH3, -O(CH2)n-ONH2, and -O(CH2)nON[(CH2)nCH3)]2, where n and m are 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides, and any combination thereof are selected from the group. In some embodiments, non-natural bases are, [ka] Selected from the group consisting of the following. In some embodiments, the non-natural nucleotide further comprises a non-natural backbone. In some embodiments, the non-natural backbone comprises phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkyl phosphotryesters, C1-C 10 The non-natural nucleotide is selected from the group consisting of phosphonates, 3'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, and boranophosphates. In some embodiments, the non-natural nucleotide is dNaMTP and / or dTPT3TP. In some embodiments, the non-natural nucleotide is incorporated into the manipulated host cell genome. In some embodiments, the non-natural nucleotide is incorporated into a chromosome. In some embodiments, the non-natural nucleotide is incorporated into the arsB locus.
[0024] Various aspects of the present invention are described in particular in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by referring to the detailed description below, which describes exemplary embodiments in which the principles of the present invention are utilized, and to the accompanying drawings. [Brief explanation of the drawing]
[0025] [Figure 1-1]Figures 1A to 1E show non-native base pairs (UBPs) and the contributions of DNA damage and tolerance pathways to their retention. Figure 1A shows dNaM-dTPT3 UBPs and native dG-dC base pairs. Figure 1B shows strains lacking NER(ΔuvrC), MMR(ΔmutH), or RER(ΔrecA). Figure 1C shows strains lacking RER and SOS(ΔrecA), as well as strains lacking only SOS(lexA(S119A)). Figure 1D shows strains lacking SOS regulatory polymerase PolII(ΔpolB) or PolIV and V(ΔdinBΔumuCD) or RER and SOS(ΔrecA). Figure 1E shows strains of PolIexo-(polA(D424A,K890R)) or PolIIIexo-(dnaQ(D12N)) with wild-type, ΔpolB, or ΔpolBΔrecA backgrounds. In each case, the shown strains were loaded with replicated plasmids in which UBP was embedded within the shown sequence (X=dNaM). For all data shown, n≧3; dots represent individual replicates; bars represent the sample mean; and error bars represent SD. [Figure 1-2] Continuation of Figure 1-1. [Figure 1-3] Continuation of Figure 1-2. [Figure 2-1]Figures 2A to 2C are diagrams showing replisome reprogramming that results in an optimized UBP retention rate. Figure 2A shows the retention rate of UBP in individual clones of WT-Opt (medium gray), ΔrecA-Opt (dark gray), and PolII+ΔrecA-Opt (light gray) after selection of the solid growth medium. Each strain was loaded with a replication pINF-derived UBP with an array configuration (GTAXAGA<TCCXCGT<TCCXGGT) having various difficulties. Each point represents an individual clone, and n≥12 for each distribution. Figure 2B shows the growth curves of chromosomal UBP integrants of WT-Opt (medium gray), ΔrecA-Opt (dark gray), and PolII+Δred-Opt (light gray) cells during logarithmic growth in media with (circles / solid lines) and without (squares / dashed lines) dNaMTP and dTPT3TP. The data are fitted to a theoretical logarithmic growth curve. n = 3; small points represent individual replicates; large points represent sample means; error bars with respect to time and OD600 represent S.D. Figure 2C shows that the retention rate of chromosomal dNaM-dTPT3 UBP in WT-Opt (medium gray), ΔrecA-Opt (dark gray), and PolII+ΔrecA-Opt (light gray) cells was measured over long-term growth. n = 3; small points represent individual replicates; large points represent sample means; error bars represent two S.D.s for both cell doubling and retention rate other than for PolII+ΔrecA-Opt data. After approximately 70 doublings, one replicate of the PolII+ΔrecA-Opt strain was contaminated with WT-Opt cells. Therefore, the data at and after the black arrow represent the average of only two independent experiments for PolII+ΔrecA-Opt. [Figure 2-2] Continuation of Figure 2-1. [Figure 3] A diagram showing the increasing pTNTT2 activity over long-term growth (10 passages) of a strain containing a knockout of IS1, compared to the YZ3 strain engineered to constitutively express the modified PtNTT2 nucleotide transporter gene from the chromosomal lacZYA locus. [Figure 4A]This figure illustrates the PtNTT2 expression constructs. Expression constructs for PtNTT2 (66-575) are shown. Figure 4A shows that all data presented in Figure 1, excluding the PolIIIExo- strain, were generated using pACS2. [Figure 4B] This figure illustrates the PtNTT2 expression constructs. Expression constructs for PtNTT2 (66-575) are shown. Figure 4B shows that data for the PolIIIexo- strain was generated using pACS2+dnaQ(D12N). [Figure 4C] This figure illustrates the PtNTT2 expression constructs. The expression constructs for PtNTT2 (66-575) are shown. Figure 4C shows that all the data in Figure 3 were generated using chromosome expression from the lacZYA locus. [Figure 5] This figure shows exonuclease-deficient polymerases that replicate TCAXAGT pINF replication data for exonuclease-deficient polymerase strains. The same strains from Figure 1E were also tested for their ability to replicate TCAXAGT(X=dNaM). For all data shown, N>≧3; error bars represent 95% empirical bootstrap confidence intervals. [Figure 6] Figures 6A and 6B show the polA(D424A,K890R) and P_polB designs. The constructive strategy for polA(D424A,K890R) and the unrepressed PpolB are shown. Figure 6A shows that polA was cleaved into its 5”---3” exonuclease domain (corresponding to PolA(1-341)). The desired D424A mutation was then introduced. The K890R mutation was generated by PCR and was predicted to have a limited effect on PolI function. Figure 6B shows that PpolB was unrepressed (PPolII+) through integration into one of the lexA operator half-sites (bold) located upstream of the -35 sequence of the promoter. [Figure 7-1]Figures 7A–7C illustrate UBS chromosome integration. Figure 7A shows the construction strategy for the arsB::UBP integration cassette. The integration cassette was constructed by superimposing a PCR of a short UBP containing DNA onto a neocassette of the pKD13 plasmid. Figure 7B shows that successful integration of the chromosomal UBP was confirmed by PCR and biotin-shift PCR. Confirmation of ΔrecA-Opt and PolII+ΔrecA-Opt SSO integraters (A2 and B3, respectively) is shown. Teal bands indicate overexposure. Figure 7C shows that reseeding of the initial integraters and isolation of individual clones rapidly identified 100% retention clones for ΔrecA-Opt and PolII+ΔrecA-Opt (A2.1 and B3.1, respectively). The same procedure for the WT-Opt integrater (C1) did not. A representative subset of the reseeded clones is shown. Red bands indicate overexposure. In panels B and C, details of the primer sets used to generate each gel are given above each gel. Molecular weight is presented in base pair counts following the size standard. Streptavidin-DNA and DNA species are indicated by black and red arrows, respectively, where the relevant shift percentage values are presented below the lane. [Figure 7-2] Continuation of Figure 7-1. [Figure 8] Figures 8A and 8B show the doubling time characteristics of reprogrammed strains and chromosome integrations. Figure 8A shows growth curves for reprogrammed strains lacking chromosomal UBP (WT-Opt (red), ΔrecA-Opt (blue), and PolII+ΔrecA-Opt (gold)) and wild-type BL21 (DE3) with chloramphenicol resistance (lacZYA::cat (black)). Circles / solid lines represent growth in media with dNaMTP and dTPT3TP. Squares / dotted lines represent growth in media without dNaMTP and dTPT3TP. Figure 8B shows that the average measured doubling time (n=3) is presented for all strains with and without chromosomal UBP, and with and without dXTP added. Figure 6B discloses sequence numbers 28 and 29, respectively, in order of appearance. [Figure 9] Figures 9A and 9B show contamination of the PolII+ΔrecA-Opt chromosome UBP integrater by WT-Opt cells. Replica 3 of the PolII+ΔrecA-Opt integrater was contaminated by WT-Opt cells at passage 13. Figure 9A shows that the PpolB locus was monitored by PCR of a gDNA sample from passage 3 of replica 3 concerning the PolII+ΔrecA-Opt integrater. Strains with the PPolII+ mutation produce larger amplicons than wild-type BL21(DE3)(a) with chloramphenicol resistance (lacZYA::cat), as can be seen from the analysis of PolII+ΔrecA-Opt before UBP integration (b). Figure 9B shows that the recA locus was monitored by PCR of a gDNA sample from passage 3 of replica 3 concerning the PolII+ΔrecA-Opt integrater. As can be seen from the analysis of PolII+ΔrecA-Opt before UBP integration (b), strains with the ΔrecA mutation produce smaller amplicons than the chloramphenicol-resistant (lacZYA::cat) wild-type BL21(DE3)(a). [Figure 10A] This figure shows mutations in the WT-Opt chromosome UBP embedding PtNTT2(66~575) during passage. The figure also shows the PtNTT2(66~575) mutations during passage of WT-Opt and their characterization. Figure 10A shows that the region between cat and IS1 (upper panel) is where the C-terminus of PtNTT2(66~575) is cut into IS1 (middle panel), indicating that the WT-Opt mutation occurs during passage. Sequencing confirmed this transposition (lower panel). Figure 10A reveals sequence numbers 30-32 in order of appearance. [Figure 10B]This figure shows mutations in the WT-Opt chromosome UBP embedding PtNTT2(66~575) during passage. Figure 10B shows the inactivation of PtNTT2(66~575) by the IS1 transposon, monitored by PCR of gDNA from passaged WT-Opt (see Table S1 for primers). The transposition event inactivates PtNTT2(66~575), with a size of approximately 3000~4000 bp. Inactivation occurs during the rapid phase of UBP loss. Additional amplicons (approximately 1500 bp in size) are also generated by these primers for chloramphenicol-resistant wild-type BL21(DE3)(lacZYA::cat)(a), pre-UBP-embedded WT-Opt(b), and wild-type BL21(DE3)(c). [Modes for carrying out the invention]
[0026] The development of non-natural base pairs (UBPs), which enable cells to store and retrieve increased information, will promote the production of proteins containing non-natural amino acids for development as therapeutic agents. This has a significant impact on practical applications, including applications to human health. However, the retention of UBP within a population of cells is sequence-dependent, and in some sequences, UBP is not adequately maintained for practical applications (e.g., protein expression) or is maintained at low levels, thereby limiting the number of usable codons.
[0027] While UBP loss during extended proliferation can be mitigated by applying UBP retention through Cas9 expression targeting selective pressure on triphosphate incorporation and cleaving and thus degrading UBP-depleted DNA sequences, retention remains a challenging issue in some sequence situations. Furthermore, this approach requires optimizing different guide RNAs for each sequence to be retained, which is a difficult challenge in many applications, such as those involving the proliferation of random DNA sequences. In addition, encoding UBP information in chromosomes, in contrast to plasmids, was expected to be unsuitable for applying this selective pressure due to unwanted cleavage of UBP-containing sequences and / or chromosomal disruption by cleavage, in contrast to minimal significant erasure in one of many copies of plasmids.
[0028] In some embodiments, modified DNA repair-related proteins, such as proteins involved in recombination repair, SOS response, nucleotide excision repair, or methyl-directed mismatch repair, and / or modified transposition-related proteins, such as insertion factor IS1. Methods, compositions, cells, engineered microorganisms, plasmids, and kits for increasing the retention of UBP using four proteins InsB, insertion factor IS1, and four proteins InsA are disclosed herein. In some embodiments, constitutive expression or overexpression of DNA repair-related proteins and / or deletion or reduction of expression of transfer-related proteins promote the high stability of the nucleoside triphosphate transporter, resulting in the creation of SSOs characterized by high UBP chromosome retention.
[0029] In certain embodiments, methods, compositions, cells, engineered microorganisms, plasmids, and kits for increasing the production of nucleic acid molecules containing non-natural nucleotides are disclosed herein. In some examples, engineered cells are disclosed herein comprising (a) a first nucleic acid molecule containing a non-natural nucleotide; and (b) a second nucleic acid molecule encoding a modified transfer-associated protein. In some embodiments, the engineered cell further comprises a third nucleic acid molecule, which is a third nucleic acid molecule encoding a modified nucleoside triphosphate transporter and is incorporated into the genome sequence of the engineered host cell, or comprises a plasmid encoding the modified nucleoside triphosphate transporter. In some embodiments, the engineered cell further comprises a Cas9 polypeptide or a variant thereof, and a single guide RNA (sgRNA) comprising a crRNA-tracrRNA scaffold, wherein the combination of the Cas9 polypeptide or a variant thereof and the sgRNA regulates the replication of the first nucleic acid molecule encoding a non-natural nucleotide. In certain embodiments, the manipulated cell further comprises (a) a fourth nucleic acid molecule encoding a Cas9 polypeptide or a variant thereof; and (b) a fifth nucleic acid molecule encoding a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold. In some examples, the first, second, third, fourth, and fifth nucleic acid molecules are encoded by one or more plasmids, and the sgRNA encoded by the fifth nucleic acid molecule contains a target motif that recognizes a modification at a non-native nucleotide site within the first nucleic acid molecule.
[0030] In some embodiments, as further provided herein, there are nucleic acid molecules containing non-natural nucleotides produced by a method comprising incubating the manipulated cells with: (a) a first nucleic acid molecule containing non-natural nucleotides; (b) a second nucleic acid molecule encoding a modified transfer-related protein; (c) a third nucleic acid molecule encoding a modified nucleoside triphosphate transporter; (d) Cas9 polypeptide (e) a fourth nucleic acid molecule encoding the Cas9 polypeptide or a variant thereof; and a fifth nucleic acid molecule encoding a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold. In some examples, modifications at non-native nucleotide sites within the first nucleic acid molecule produce a modified first nucleic acid molecule, and the combination of the Cas9 polypeptide or a variant thereof and the sgRNA regulates the replication of the modified first nucleic acid molecule, leading to the production of a nucleic acid molecule containing non-native nucleotides. In some examples, the expression of a modified translocation-associated protein in the engineered cell increases the stability of the triphosphate transporter. In some embodiments, the high stability of the triphosphate transporter contributes to (i) increased production of a modified polypeptide containing non-native amino acids encoded by non-native nucleotides, and / or (ii) high retention of non-native nucleotides in the genome of the engineered cell.
[0031] In some embodiments, as additionally provided herein, semisynthetic organisms (SSOs) produced by a method comprising incubating an organism with: (a) a first nucleic acid molecule containing a non-natural nucleotide; (b) a second nucleic acid molecule encoding a modified transfer-related protein; (c) a third nucleic acid molecule encoding a modified nucleoside triphosphate transporter; (d) a fourth nucleic acid molecule encoding a Cas9 polypeptide or a variant thereof; and (e) a fifth nucleic acid molecule encoding a single guide RNA (sgRNA) containing a crRNA-tracrRNA scaffold. In some examples, modification at a non-natural nucleotide site within the first nucleic acid molecule generates a modified first nucleic acid molecule, and the combination of the Cas9 polypeptide or a variant thereof and the sgRNA regulates the replication of the modified first nucleic acid molecule, leading to the production of a semisynthetic organism containing a nucleic acid molecule containing a non-natural nucleotide. In some examples, expression of the modified transfer-related protein in the engineered cells increases the stability of the triphosphate transporter. In some embodiments, the high stability of the triphosphate transporter contributes to (i) increased production of modified polypeptides containing non-natural amino acids encoded by non-natural nucleotides, and / or (ii) high retention of non-natural nucleotides in the SSO genome.
[0032] DNA repair mechanism DNA repair mechanisms include nucleotide excision repair (NER), ribonucleotide excision repair (RER), SOS response, methyl-directed mismatch repair (MMR), and recombination repair. NER, MMR, RER, and SOS response are signal-induced, which can be mimicked by introducing UBP into the host genome. Non-limiting examples of DNA repair-related proteins in prokaryotic cells involved in recombination repair and / or SOS response include RecA, Rad51, RadA, and LexA. Non-limiting examples of DNA repair-related proteins in prokaryotic cells involved in recombination repair include RecO, RecR, RecN, and RuvABC. Non-limiting examples of DNA repair-related proteins in prokaryotic cells involved in NER include UvrA and UvrB. Non-limiting examples of DNA repair-related proteins in prokaryotic cells involved in MMR include MutS, MutH, and MutL.
[0033] In some embodiments, modified DNA repair-related proteins are introduced into the manipulated cells or SSOs described herein to enhance chromosomal UBP retention. In some embodiments, the modified DNA repair-related proteins include deletions of RecA, Rad51, RadA, LexA, RecO, RecR, RecN, RuvABC, MutS, MutH, MutL, UvrA, and / or UvrB. In some embodiments, the deletions include N-terminal deletions, C-terminal deletions, cleavage at both ends, internal deletions, and / or deletions of the entire gene. In some embodiments, deletions or mutations in the nucleic acid molecule encoding the DNA repair-related protein are modified to result in deletions.
[0034] Translocation-related proteins In E. coli, there are replication and conservation (non-replication) mechanisms for translocation of transposable elements (e.g., ISIs) consisting of nucleic acid sequences. In the replication pathway, a new copy of the transposable element is generated during the translocation event. As a result of the translocation, one copy appears at a new site, and one copy remains at the old site. In the conservation pathway, replication does not occur. Instead, the element is cleaved from the chromosome or plasmid and incorporated into a new site. In these cases, DNA replication of the element does not occur, and the element is lost at its original chromosomal location. Deletion of a transposable element is responsible for a high incidence of deletions in its vicinity (e.g., deletion of the transposable element in addition to adjacent or surrounding DNA).
[0035] The insB-4 and insA-4 genes encode two proteins, InsB and InsA, necessary for the transposition of the IS1 transposon. IS1 transposition involves targeted replication of 9 to 8 base pairs. Deletion of insB-4 suppresses the abnormal transposition event mediated by InsB.
[0036] In some embodiments, the methods, manipulated cells, and semisynthetic organisms described herein include a modified nucleic acid molecule encoding a translocation-related protein. In some embodiments, the translocation-related protein includes insB and / or insA. In some embodiments, the modified nucleic acid molecule encoding the translocation-related protein includes a deletion or mutation. In some embodiments, the deletion includes an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, an internal deletion, and / or a deletion of the entire gene. In some embodiments, the mutation results in a decrease in the expression of insB and / or InsA. In some embodiments, the deletion or mutation of the modified nucleic acid molecule encoding the translocation-related protein is effective in stabilizing the expression and / or activity of the triphosphate nucleoside transporter, thereby increasing the retention of UBP.
[0037] In some embodiments, the methods, manipulated cells, and semisynthetic organisms described herein include a modified nucleic acid molecule encoding an IS1 transposer. In some embodiments, the modified nucleic acid molecule encoding an IS1 transposer includes a deletion or mutation. In some embodiments, the deletion includes the knockout or knockdown of all or part of the nucleic acid molecule encoding the IS1 transposon. In some embodiments, the mutation results in a decrease in the expression of the IS1 transposon. In some embodiments, the deletion or mutation of the modified nucleic acid molecule encoding the IS1 transposon is effective in stabilizing the expression and / or activity of the triphosphate nucleoside transporter, thereby increasing the retention of UBP. In some examples, the modified nucleic acid molecule encoding an IS1 transposer includes SEQ ID NO: 4.
[0038] CRISPR / CRISPR-related (CAS) editing system In some embodiments, the methods, cells, and manipulated microorganisms disclosed herein utilize the CRISPR / CRISPR-related (Cas) system for modification of nucleic acid molecules, including non-native nucleotides. In some examples, the CRISPR / Cas system modulates the retention of the modified nucleic acid molecule, including the modification at its non-native nucleotide site. In some examples, retention is a decrease in the replication of the modified nucleic acid molecule. In some examples, the CRISPR / Cas system generates double-strand breaks within the modified nucleic acid molecule that lead to degradation involving DNA repair proteins such as RecBCD and its associated nucleases.
[0039] In some embodiments, the CRISPR / Cas system (1) incorporates short regions of genetic material homologous to the nucleic acid molecule of interest, containing non-native nucleotides called "spacers," into a clustered array in the host genome, and (2) short guides from the spacers. The system includes (3) the expression of RNA (crRNA), (4) the binding of crRNA to a specific portion of the target nucleic acid molecule called a protospacer, and (5) the degradation of the protospacer by a CRISPR-related nuclease (Cas). In some cases, the type II CRISPR system has been described in the bacterium Streptococcus pyogenes, in which Cas9 and two non-coding small RNAs (pre-crRNA and tracrRNA (trans-activated CRISPR RNA)) act in a sequence-specific manner to target and degrade the target nucleic acid molecule (Jinek et al., "A Programmable Dual-RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity," Science 337(6096): pp. 816-821 (August 2012, electronically published June 28, 2012)).
[0040] In some examples, two non-coding RNAs are further fused into a single guide RNA (sgRNA). In some examples, the sgRNA contains a target motif that recognizes a modification at a non-native nucleotide site within the nucleic acid molecule of interest. In some embodiments, the modification is a substitution, insertion, or deletion. In some cases, the sgRNA contains a target motif that recognizes a substitution at a non-native nucleotide site within the nucleic acid molecule of interest. In some cases, the sgRNA contains a target motif that recognizes a deletion at a non-native nucleotide site within the nucleic acid molecule of interest. In some cases, the sgRNA contains a target motif that recognizes an insertion at a non-native nucleotide site within the nucleic acid molecule of interest.
[0041] In some cases, the target motif is 10–30 nucleotides long. In some examples, the target motif is 15–30 nucleotides long. In some cases, the target motif is approximately 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long. In some cases, the target motif is approximately 15, 16, 17, 18, 19, 20, 21, or 22 nucleotides long.
[0042] In some cases, the sgRNA further contains a protospacer-adjacent motif (PAM) recognition factor. In some examples, the PAM is located adjacent to the 3' end of the target motif. In some cases, the nucleotides in the target motif that form a Watson-Crick base pair with the modification at the non-native nucleotide site within the nucleic acid molecule of interest are located 3–22, 5–20, 5–18, 5–15, 5–12, or 5–10 nucleotides from the 5' end of the PAM. In some cases, the nucleotides in the target motif that form a Watson-Crick base pair with the modification at the non-native nucleotide site within the nucleic acid molecule of interest are located approximately 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides from the 5' end of the PAM.
[0043] In some cases, the CRISPR / Cas system utilizes the Cas9 polypeptide or a variant thereof. Cas9 is a double-stranded nuclease having two active cleavage sites, one for each strand of the double helix. In some cases, the Cas9 polypeptide or a variant thereof produces double-strand breaks. In some cases, the Cas9 polypeptide is wild-type Cas9. In some cases, the Cas9 polypeptide is Cas9 optimized for expression in the cells and / or engineered microorganisms described herein.
[0044] In some embodiments, the Cas9 / sgRNA complex binds to a portion of the target nucleic acid molecule (e.g., DNA) containing a sequence that matches, for example, 17-20 nucleotides upstream of the sgRNA of PAM. Once bound, the two independent nuclease domains in Cas9 then each cleave one of the DNA strands three bases upstream of PAM, creating a blunt-ended D NA double-strand breaks (DSBs) remain. Subsequently, in some cases, the presence of DSBs leads to the degradation of the target DNA by RecBCD and its associated nucleases.
[0045] In some cases, the Cas9 / sgRNA complex modulates the retention of modified nucleic acid molecules, including modifications at their non-native nucleotide sites. In some embodiments, retention refers to a decrease in the replication of the modified nucleic acid molecule. In some cases, Cas9 / sgRNA reduces the replication rate of the modified nucleic acid molecule by approximately 80%, 85%, 95%, 99%, or more.
[0046] In some cases, the production of nucleic acid molecules containing non-natural nucleotides increases by approximately 30%, 40%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0047] In some cases, the retention of nucleic acid molecules containing non-natural nucleotides increases by approximately 30%, 40%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0048] In some embodiments, the CRISPR / Cas system includes two or more sgRNAs. In some examples, each of the two or more sgRNAs independently includes a target motif that recognizes a modification at a non-native nucleotide site within the nucleic acid molecule of interest. In some embodiments, the modification is a substitution, insertion, or deletion. In some cases, each of the two or more sgRNAs includes a target motif that recognizes a substitution at a non-native nucleotide site within the nucleic acid molecule of interest. In some cases, each of the two or more sgRNAs includes a target motif that recognizes a deletion at a non-native nucleotide site within the nucleic acid molecule of interest. In some cases, each of the two or more sgRNAs includes a target motif that recognizes an insertion at a non-native nucleotide site within the nucleic acid molecule of interest.
[0049] In some embodiments, the specificity of CRISPR component binding to the target nucleic acid molecule is controlled by a non-repeatable spacer factor at the pre-crRNA portion of the sgRNA. When transcription occurs along the tracrRNA portion, the Cas9 nuclease is directed to the protospacer:crRNA heteroduplex, inducing double-strand break (DSB) formation. In some examples, the specificity of the sgRNA is approximately 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or higher. In some examples, the off-target binding rate of the sgRNA is approximately 20%, 15%, 10%, 5%, 3%, 1%, or less.
[0050] nucleic acid molecule In some embodiments, nucleic acids (for example, also referred to herein as the nucleic acid molecule of interest) are derived from any source or composition such as DNA, cDNA, gDNA (genomic DNA), RNA, siRNA (short inhibitory RNA), RNAi, tRNA, mRNA, or rRNA (ribosomal RNA), and are in any form (e.g., linear, cyclic, supercoiled, single-stranded, double-stranded, etc.). In some embodiments, nucleic acids include nucleotides, nucleosides, or polynucleotides. In some cases, nucleic acids include native and non-native nucleic acids. In some cases, nucleic acids also include non-native nucleic acids such as DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-native backbone, etc.). The term “nucleic acid” does not refer to or presume to a polynucleotide chain of a specific length, and therefore polynucleotides and oligonucleotides are also included in the definition. It is understood that, as an example of natural nucleotides, there are, but are not limited to, ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, GMP, dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Examples of natural deoxyribonucleotides include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Examples of natural ribonucleotides include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, and GMP. For RNA, the uracil base is uridine. Nucleic acids are sometimes vectors, plasmids, phagemids, autonomous replication sequences (ARS), centromeres, artificial chromosomes, yeast artificial chromosomes (e.g., YACs), or other nucleic acids that can or are replicated within a host cell. In some cases, non-natural nucleic acids are nucleic acid analogs. In additional cases, non-natural nucleic acids are derived from extracellular sources. In other cases, non-natural nucleic acids are available within the intracellular space of organisms provided herein, such as genetically modified organisms.
[0051] unnatural nucleic acid Nucleotide analogs, or non-natural nucleotides, include nucleotides containing several types of modifications to bases, sugars, or phosphate moieties. In some embodiments, the modifications include chemical modifications. In some cases, the modifications occur in the 3'OH or 5'OH group, the backbone, sugar components, or nucleotide bases. In some examples, the modifications may include linker molecules that do not exist naturally, and / or interchain or intrachain crosslinks. In one embodiment, the modified nucleic acid includes modifications to one or more of the 3'OH or 5'OH group, the backbone, sugar components, or nucleotide bases, and / or the addition of linker molecules that do not exist naturally. In one embodiment, the modified backbone includes a backbone other than a phosphodiester backbone. In one embodiment, the modified sugar includes a sugar other than deoxyribose (in the modified DNA) or a sugar other than ribose (in the modified RNA). In one embodiment, the modified bases include bases other than adenine, guanine, cytosine, or thymine (in the modified DNA), or bases other than adenine, guanine, cytosine, or uracil (in the modified RNA).
[0052] In some embodiments, the nucleic acid contains at least one modified base. In some examples, the nucleic acid contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more modified bases. In some cases, modifications to the base moiety include A, C, G, and T / U, as well as natural and synthetic modifications of different purine or pyrimidine bases. In some embodiments, the modification is a modified form of adenine, guanine, cytosine, or thymine (in the modified DNA), or a modified form of adenine, guanine, cytosine, or uracil (in the modified RNA).
[0053] Modified bases of non-natural nucleic acids include, but are not limited to, uracil-5-yl, hypoxanthin-9-yl(I), 2-aminoadenine-9-yl, 5-methylcytosine(5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, and 6-azo. Examples include uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. 2-aminopropyladenine, 5-p Certain non-natural nucleic acids, including ropinyluracil and 5-propynylcytosine, such as 5-substituted pyrimidines, 6-azapyrimidines and N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, 5-propynylcytosine, 5-methylcytosine, those that enhance the stability of double-strand formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, 5-substituted pyrimidines, 6-azapyrimidines, and certain non-natural nucleic acids such as N-2, N-6 and O-6 substituted purines. 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl, other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyl(-C≡C-CI1 / 4)uracil, 5-propynylcytosine, other alkynyl derivatives of pyrimidine nucleic acids, 6-azouracil, 6-azocytosine, 6-azocymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenine and guanine, 5-halo, especially 5-bromo, 5-trifluoromethyl, and other 5-substituted u Racil and cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidine, phenoxazinecytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimidine) Mido[5,4-b][1,4]benzothiazine-2(3H)-one), G-clamp, phenoxazinecytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindolecytidine (H-pyrimido[3',2':4,5]pyrrolo[2,3-d]pyrimidine-2-one), purine or pyrimidine group replaced by another heterocycle, 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, azacytosine, 5-bromocytosine, bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine Tosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, 2-aminoadenine, 6-thioguanine, 2-thiothymine, 4-thiothymine, 5-propynyluracil, 4-thiouracil, N4-ethylcytosine, 7-deazaguanine, 7-deaza-8-azaguanine, 5-hydroxycytosine, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, and U.S. Patent No. 3,687,808; U.S. Patent No. 4,845,205 US No. 4,910,300; US No. 4,948,882; US No. 5,093,232; US No. 5,130,302; US No. 5,134,066; US No. 5,175,273; US No. 5,367,066; US No. 5,432,272; US No. 5,457,187; US No. 5,459,255; US No. 5,484,908; US No. 5,502,177; US No. 5,525,711; US No. 5,552,540 U.S. Patent No. 5,587,469; U.S. Patent No. 5,594,121; U.S. Patent No. 5,596,091; U.S. Patent No. 5,614,617; U.S. Patent No. 5,645,985; U.S. Patent No. 5,681,941; U.S. Patent No. 5,750,692; U.S. Patent No. 5,763,588; U.S. Patent No. 5,830,653; and U.S. Patent No. 6,005,096; WO99 / 62923; Kandimalla et al., (2001) Bioorg. Med. Chem. 9: pp. 807-813; The Concise Encyclopedia of Polymer Science and Engineering, Kroschwitz,J.I., John Wiley & Sons, 1990, pp. 858-859; Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613; and S. This is described in Sanghvi, Chapter 15, Antisense Research and Applications, edited by Crookeand Lebleu, CRC Press, 1993, pp. 273-288. Additional base modifications can be found, for example, in U.S. Patent No. 3,687,808, Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613; and Sanghvi, Chapter 15, Antisense Research and Applications, pp. 289-302, edited by Crookeand Lebleu, CRC Press, 1993.
[0054] Non-natural nucleic acids, comprising various heterocyclic bases and various sugar moieties (and sugar analogs), are available in the art, and in some cases, the nucleic acids contain one or more heterocyclic bases other than the five major basic base components of naturally occurring nucleic acids. For example, heterocyclic bases in some cases include uracil-5-yl, cytosine-5-yl, adenine-7-yl, adenine-8-yl, guanine-7-yl, guanine-8-yl, 4-aminopyrrolo[2.3-d]pyrimidine-5-yl, 2-amino-4-oxopyrrolo[2.3-d]pyrimidine-5-yl, and 2-amino-4-oxopyrrolo[2.3-d]pyrimidine-3-yl groups, where purines are attached to the sugar moiety of the nucleic acid via position 9, pyrimidines via position 1, pyrrolopyrimidines via position 7, and pyrazolopyrimidines via position 1.
[0055] In some embodiments, modified bases of non-natural nucleic acids are shown below, where the wavy lines indicate the point of binding to (deoxy)ribose or ribose.
[0056] [ka] [ka] [ka] [ka] [ka] [ka]
[0057] In some embodiments, nucleotide analogs are also modified at the phosphate moiety. Modified phosphate moieties include, but are not limited to, those having modifications at the bond between two nucleotides, and include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates including 3'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates including 3'-aminophosphoamides and aminoalkylphosphoamides, thionophosphoamides, thionoalkyl phosphonates, thionoalkyl phosphotriesters, and boranophosphates. These phosphate or modified phosphate bonds between two nucleotides are through a 3'-5' bond or a 2'-5' bond, and it is understood that the bond contains inverted polarity such as 3'-5' to 5'-3' or 2'-5' to 5'-2'. Various salts, mixed salts, and free acid forms are also included. Many U.S. patents teach, but are not limited to, how to prepare and use nucleotides containing modified phosphates, including, but are limited to, patents 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; and 5,321,131. Nos. 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050 are examples.
[0058] In some embodiments, non-natural nucleic acids include 2',3'-dideoxy-2',3'-didehydro-nucleosides (PCT / US2002 / 006460), 5'-substituted DNA and RNA derivatives (PCT / US2011 / 033961; Saha et al., J. Org Chem., 1995, 60, pp. 788-789; Wang et al., Bioorganic & Medicinal Chemistry Letters, 1999, 9, pp. 885-890; and Mikhailov et al., Nucleosides & Nucleotides, 1991, 10(1-3), pp. 339-343; Leonid et al., 1995, 14(3-5), pp. 901-905; and Eppacher et al., Helvetica Chimica This includes Acta, 2004, 87, pp. 3004-3020; PCT / JP2000 / 004720; PCT / JP2003 / 002342; PCT / JP2004 / 013216; PCT / JP2005 / 020435; PCT / JP2006 / 315479; PCT / JP2006 / 324484; PCT / JP2009 / 056718; PCT / JP2010 / 067560), or 5'-substituted monomers prepared as monophosphates with modified bases (Wang et al., Nucleosides Nucleotides & Nucleic Acids, 2004, 23(1&2), pp. 317-337).
[0059] In some embodiments, the non-natural nucleic acids include modifications at the 5' and 2' positions of the sugar ring (PCT / US94 / 02993), e.g., 5'-CH2-substituted 2'-O-protected nucleosides (Wu et al., Helvetica Chimica Acta, 2000, 83, pp. 1127-1143, and Wu et al., Bioconjugate Chem. 1999, 10, pp. 921-924). In some cases, the non-natural nucleic acids include amide-linked nucleoside dimers prepared for incorporation into oligonucleotides, where the 3'-linked nucleosides in the dimer (5' to 3') include 2'-OCH3 and 5'-(S)-CH3 (Mesmaeker et al., Synlett, 1997, pp. 1287-1290). Non-natural nucleic acids may include 2'-substituted 5'-CH2 (or O) modified nucleosides (PCT / US92 / 01020). Non-natural nucleic acids may include 5'-methylenephosphonate DNA and RNA monomers, as well as dimers (Bohringer et al., Tet. Lett., 1993, 34, pp. 2723-2726; Collingwood et al., Synlett, 1995, 7, pp. 703-705; and Hutter et al., Helvetica Chimica Acta, 2002, 85, pp. 2777-2806). Non-natural nucleic acids may include 5'-phosphonate monomers with 2' substitution (US2006 / 0074035) and other modified 5'-phosphonate monomers (WO1997 / 35869). Non-natural nucleic acids may include 5'-modified methylenephosphonate monomers (EP614907 and EP629633). Non-natural nucleic acids may also include analogues of 5' or 6'-phosphonate ribbonucleosides containing hydroxyl groups at the 5' and / or 6' positions (Chen et al., Phosphorus, Sulfur and Silicon, 2002, pp. 777, 1783-1786; Jung et al., Bioorg. Med. Chem., 2000, 8, pp. 2501-2509; Gallier et al., Eur. J. Org. Chem., 2007, pp. 925-933; and Hampton et al., J. Med. Chem., 1976, 19(8), pp. 1029-1033).Non-natural nucleic acids can include 5'-phosphonate deoxyribonucleoside monomers and dimers having a 5'-phosphate group (Nawrot et al., Oligonucleotides, 2006, 16(1), 6). (pp. 8-82). Non-natural nucleic acids may contain nucleosides having a 6'-phosphonate group, where the 5' and / or 6' position is either not substituted with a thio-tert-butyl group (SC(CH3)3) (and its analogues); a methyleneamino group (CH2NH2) (and its analogues); or a cyano group (CN) (and its analogues) (Fairhurst et al., Synlett, 2001, 4, pp. 467-472; Kappler et al., J. Med. Chem., 1986, 29, pp. 1030-1030). (Page 38; Kappler et al., J. Med. Chem., 1982, 25, pp. 1179-1184; Vrudhula et al., J. Med. Chem., 1987, 30, pp. 888-894; Hampton et al., J. Med. Chem., 1976, 19, pp. 1371-1377; Geze et al., J. Am. Chem. Soc., 1983, 105(26), pp. 7638-7640; and Hampton et al., J. Am. Chem. Soc., 1973, 95(13), pp. 4404-4414).
[0060] In some embodiments, non-natural nucleic acids also involve modification of the sugar moiety. In some cases, the nucleic acid contains one or more nucleosides, and the sugar moiety is modified. Such modified sugar nucleosides can confer high nuclease stability, high binding affinity, or several other beneficial biological properties. In certain embodiments, the nucleic acid includes a chemically modified ribofuranose ring moiety. Examples of chemically modified ribofuranose rings, but not limited to, include the addition of substituents (including 5' and / or 2' substituents); bridging of two ring atoms to form a bicyclic nucleic acid (BNA); S, N(R) or C(Ri)(R2)(R=H, C1~C 12Examples include substitution of the ribosyl ring oxygen atom with alkyl or protecting groups; and combinations thereof. Examples of chemically modified sugars can be found in WO2008 / 101157, US2005 / 0130923, and WO2007 / 134181.
[0061] In some cases, modified nucleic acids contain modified sugars or sugar analogs. Thus, in addition to ribose and deoxyribose, the sugar moiety may be a pentose, deoxypentose, hexose, deoxyhexose, glucose, arabinose, xylose, lyxose, or sugar "analogous" cyclopentyl group. The sugar may be in pyranosyl or furanosyl form. The sugar moiety is a furanoside of ribose, deoxyribose, arabinose, or 2'-O-alkylribose, and the sugar can be attached to the respective heterocyclic base in either an [alpha] or [beta] anomeric configuration. Sugar modifications include, but are not limited to, 2'-alkoxy-RNA analogs, 2'-amino-RNA analogs, 2'-fluoro-DNA, and 2'-alkoxy- or amino-RNA / DNA chimeras. For example, sugar modifications may include 2'-O-methyluridine or 2'-O-methylcytidine. Sugar modifications include 2'-O-alkyl-substituted deoxyribonucleosides and 2'-O-ethylene glycol-like ribonucleosides. The preparation of these sugars or sugar analogs, and their respective "nucleosides," is known, and such sugars or analogs are bonded to heterocyclic bases (nucleic acid bases). Sugar modifications can also be prepared and combined with other modifications.
[0062] Modifications to the sugar moiety include natural and non-natural modifications of ribose and deoxyribose. Sugar modifications include, but are not limited to, the following modifications at the 2' position: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl or C2-C 10It can be an alkenyl or alkynyl. 2' sugar modification is also possible, but is not limited to -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n ONH2 and -O(CH2) n ON[(CH2) n CH3)2 is included, and n and m are between 1 and approximately 10.
[0063] Other modifications in position 2' are not limited to C1~C 10This includes lower alkyl groups, substituted lower alkyl groups, alkaryl groups, aralkyl groups, O-alkaryl groups, O-aralkyl groups, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl groups, heterocycloalkaryl groups, aminoalkylamino groups, polyalkylamino groups, substituted silyl groups, RNA cleavage groups, reporter groups, intercalators, groups for improving the pharmacokinetic properties of oligonucleotides, or groups for improving the pharmacodynamic properties of oligonucleotides, and other substituents having similar properties. Similar modifications can also be made at other positions of the sugar, particularly on the 3' terminal nucleotide or at the 3' position of the sugar in 2'-5' bonded oligonucleotides, and at the 5' position of the 5' terminal nucleotide. Modified sugars also include those containing modifications at cross-linked ring oxygen such as CH2 and S. Nucleotide sugar analogs may also have sugar mimes such as cyclobutyl moieties instead of pentofuranosyl sugars.There are many U.S. patents that teach the preparation of such modified sugar structures and describe in detail the scope of base modification, for example, U.S. Patents 4,981,957; 5,118,800; 5,319,080; 5,359,044; 5,393,878; 5,446,137; 5,466,786; 5,514,785; 5,519,134; 5,567,811; 5,576,427; 5,591,722; 5,597,909; 5,610,300; 5,627,053; 5,639,873; 5,646,265; 5,658,873 US Patent No. 5,670,633; US Patent No. 4,845,205; US Patent No. 5,130,302; US Patent No. 5,134,066; US Patent No. 5,175,273; US Patent No. 5,367,066; US Patent No. 5,432,272; US Patent No. 5,457,187; US Patent No. 5,459,255; US Patent No. 5,484,908; US Patent No. 5, U.S. Patent Nos. 502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121, 5,596,091; 5,614,617; 5,681,941; and 5,700,920, each of which is incorporated herein by reference in its entirety.
[0064] Examples of nucleic acids having modified sugar moieties include, but are not limited to, nucleic acids having 5'-vinyl, 5'-methyl(R or S), 4'-S, 2'-F, 2'-OCH3, and 2'-O(CH2)2OCH3 substituents. Substituents at the 2' position also include allyl, amino, azide, thio, O-allyl, and O-(C1~C 10 Alkyl), OCF3, O(CH2)2SCH3, O(CH2)2-ON(R m )(R n ), and O-CH2-C(=O)-N(R m )(R n) can be selected from, where each R m and R n These are independently H or substituted or unsubstituted C1~C 10 It is alkyl.
[0065] In certain embodiments, the nucleic acids described herein comprise one or more bicyclic nucleic acids. In certain embodiments, the bicyclic nucleic acid comprises a bridge between 4' and 2' ribosyl ring atoms. In certain embodiments, the nucleic acids provided herein comprise one or more bicyclic nucleic acids, wherein the bridge comprises a 4' to 2' bicyclic nucleic acid. Examples of such 4'-to-2' bicyclic nucleic acids include, but are not limited to, formulas: 4'-(CH2)-O-2'(LNA); 4'-(CH2)-S-2'; 4'-(CH2)2-O-2'(ENA); 4'-CH(CH3)-O-2' and 4'-CH(CH2OCH3)-O-2' and one of their analogues (see U.S. Patent No. 7,399,845); 4'-C(CH3)(CH3)-O-2' and its analogues (WO2009 / 006478, WO2008 / 150729, U.S.2004 / 0171570, U.S. Patent No. 7,427,672, Chattopadhyaya et al., J. Org. Chem., 2009, 74 (See pages 118-134 and WO2008 / 154401.) Also, for example, Singh et al., Chem.Commun., 1998, 4, pp. 455-456; Koshkin et al., Tetrahedron, 1998, 54, pp. 3607-3630; Wahlestedt et al., Proc.Natl.Acad.Sci.USA, 2000, 97, pp. 5633-5638; Kumar et al., Bioorg.Med.Chem.Lett., 1998, 8, pp. 2219-2222; Singh et al., J.Org.Chem., 1998, 63, pp. 10035-10039; Srivastava et al., J.Am.Chem.Soc., 2007, 129(26), pp. 8362-8379; Elayadi et al., Curr.Opinion Invens.Drugs, 2001, Vol. 2, pp. 558-561; Braasch et al., Chem.Biol, 2001, Vol. 8, pp. 1-7; Oram et al., Curr.Opinion Mol. Ther., 2001, 3, pp. 239-243; U.S. Patent Nos. 4,849,513; U.S. Patent Nos. 5,015,733; U.S. Patent Nos. 5,118,800; U.S. Patent Nos. 5,118,802; U.S. Patent Nos. 7,053,207; U.S. Patent Nos. 6,268,490; U.S. Patent Nos. 6,770,748; U.S. Patent Nos. 6,794,499; U.S. Patent Nos. 7,034,133; U.S. Patent Nos. 6,525,191; U.S. Patent Nos. 6,670,461; and U.S. Patent Nos. 7,399,845; International Publication Numbers WO2004 / 106356, WO1994 / 14226, WO2005 / 021570, WO2007 / 090071, etc. See also WO2007 / 134181; U.S. Patent Publication Nos. US2004 / 0171570, US2007 / 0287831, and US2008 / 0039618; U.S. Provisional Application Nos. 60 / 989,574, U.S. Provisional Application Nos. 61 / 026,995, U.S. Provisional Application Nos. 61 / 026,998, U.S. Provisional Application Nos. 61 / 056,564, U.S. Provisional Application Nos. 61 / 086,231, U.S. Provisional Application Nos. 61 / 097,787, and U.S. Provisional Application Nos. 61 / 099,844; and International Application Nos. PCT / US2008 / 064591, PCTUS2008 / 066154, PCTUS2008 / 068922, and PCT / DK98 / 00393.
[0066] In certain embodiments, nucleic acids include bound nucleic acids. Nucleic acids can be bound together using any inter-nucleic acid bonds. Two main classes of inter-nucleic acid linkages are defined by the presence or absence of a phosphorus atom. Representative phosphorus groups containing inter-nucleic acid bonds include, but are not limited to, phosphodiesters, phosphotriesters, methylphosphonates, phosphoramidates, and phosphorothioates (P=S). Representative non-phosphorus groups containing inter-nucleic acid bonds include, but are not limited to, methylenemethylimino (-CH2-N(CH3)-O-CH2-), thiodiesters (-OC(O)-S-), thionocarbamates (-OC(O)(NH)-S-); siloxanes (-O-Si(H)2-O-); and N,N * -Dimethylhydrazine (-CH2-N(CH3)-N(CH3)) is an example. In certain embodiments, internucleic acid bonds having chimeric atoms can be prepared as a racemic mixture, or as separate enantiomers, such as alkylphosphonates and phosphorothioates. Non-natural nucleic acids may contain a single modification. Non-natural nucleic acids may contain multiple modifications in one or between different parts.
[0067] Main chain phosphate modifications to nucleic acids include, but are not limited to, methylphosphonates, phosphorothioates, phosphoramidates (crosslinked or uncrosslinked), phosphotryesters, phosphorodithioates, phosphodithioates, and boranophosphates, and can be used in any combination. Other non-phosphate bindings can also be used.
[0068] In some embodiments, main chain modifications (e.g., internucleotide bonding of methylphosphonates, phosphorothioates, phosphoramidates, and phosphorodithioates) can confer immunomodulatory activity to modified nucleic acids and / or enhance their stability in vivo.
[0069] In some cases, the phosphorus derivative (or modified phosphate group) is attached to a sugar or sugar analog moiety and may be a monophosphate, diphosphate, triphosphate, alkylphosphonate, phosphorothioate, phosphorodithioate, phosphoramidate, etc. Exemplary polynucleotides containing modified phosphate or non-phosphate bonds are cited in Peyrottes et al., 1996, Nucleic Acids Res. 24: pp. 1841-1848; Chaturvedi et al., 1996, Nucleic Acids Res. 24: pp. 2318-2323; and Schultz et al., (1996) Nucleic Acids Res. 24: pp. 2966-2973; Matteucci, 1997, “Oligonucleotide Analogs: an Overview” in Oligonucleotides as Therapeutic Agents, (edited by Chadwick and Cardew), John Wiley and Sons, New York, NY; Zon, 1993, “Oligonucleoside Phosphorothioates” in Protocols for Oligonucleotides and Analogs, Synthesis and Properties, Humana This can be found in Press, pp. 165-190; Miller et al., 1971, JACS 93: pp. 6657-6665; Jager et al., 1988, Biochem. 27: pp. 7247-7246; Nelson et al., 1997, JOC 62: pp. 7278-7287; U.S. Patent No. 5,453,496; and Micklefield, 2001, Curr. Med. Chem. 8: pp. 1157-1179.
[0070] In some cases, backbone modifications involve replacing phosphodiester bonds with alternative moieties such as anionic, neutral, or cationic groups. Examples of such modifications include anionic nucleoside bonds; N3' to P5' phosphoramidate modifications; boranophosphate DNA; prooligonucleotides; neutral nucleoside bonds such as methylphosphonates; amide-linked DNA; methylene (methylimino) bonds; formacetal and thioformacetal bonds; backchains containing sulfonyl groups; morpholino oligonucleotides; peptide nucleic acids (PNA); and positively charged deoxyribonucleic acid guanidine (DNG) oligonucleotides (Micklefield, 2001, Current Medicinal Chemistry 8: pp. 1157-1179). Modified nucleic acids may include chimeric or mixed backchains containing one or more modifications, such as combinations of phosphate bonds, including combinations of phosphodiester and phosphorothioate bonds.
[0071] Substituents for the phosphate include, for example, short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatoms and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatoms or heterocyclic nucleoside bonds. These include those having morpholino bonds (partially formed from the sugar moiety of the nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2 component moieties. Many U.S. patents disclose, but are not limited to, how such types of phosphate substitutions are fabricated and used, including U.S. Patents 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; and 5,602,240. U.S. Patents 5,610,289, 5,602,240, 5,608,046, 5,610,289, 5,618,704, 5,623,070, 5,663,312, 5,633,360, 5,677,437, and 5,677,439 are examples. It is also understood in nucleotide substituents that both the sugar and phosphate portions of a nucleotide can be replaced by, for example, an amide-type bond (aminoethylglycine) (PNA). U.S. Patents 5,539,082, 5,714,331, and 5,719,262 teach how to prepare and use PNA molecules, each of which is incorporated herein by reference in its entirety. See also Nielsen et al., Science, 1991, 254, pp. 1497-1500. It is also possible to bind other types of molecules (conjugates) to nucleotides or nucleotide analogs to enhance, for example, cellular uptake. Conjugates can be chemically bonded to nucleotides or nucleotide analogs.Such conjugates are not limited to, but include the cholesterol portion (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, pp. 6553-6556), cholic acid (Manoharan et al., Bioorg. Med. Chem. Let., 1994, 4, pp. 1053-1060), thioethers, for example, hexyl-S-tritylthiol (Manoharan et al., Ann. KY. Acad. Sci., 1992, pp. 660, 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, pp. 2765-2770), and thiocholesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, pp. 533-538), aliphatic chains, e.g., dodecanediol or undecyl residues (Saison-Behmoaras et al., EM5OJ, 1991, 10, pp. 1111-1118; Kabanov et al., FEBS Lett., 1990, 259, pp. 327-330; Svinarchuk et al., Biochimie, 1993, 75, pp. 49-54), phospholipids, e.g., di-hexadecyl-glycerol or triethylammonium 1-di-O-hexadecyl-rac-glycero-SH-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, pp. 3651-3654; Shea et al., Nucl. Acids It includes lipid moieties such as Res., 1990, 18, pp. 3777-3783, polyamines or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, pp. 969-973), or adamantane acetate (Manoharan et al., Tetrahedron Lett., 1995, 36, pp. 3651-3654), palmityl moieties (Mishra et al., Biochem. Biophys. Acta, 1995, 1264, pp. 229-237), or octadecylamine or hexylamino-carbonyl-oxycholesterol moieties (Crooke et al., J. Pharmacol. Exp. Ther., 1996, pp. 277, pp. 923-937).Many U.S. patents teach the manufacture of such conjugates, but are not limited to, U.S. Patents 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414 ,077; US Patent No. 5,486,603; US Patent No. 5,512,439; US Patent No. 5,578,718; US Patent No. 5,608,046; US Patent No. 4,587,044; US Patent No. 4,605,735; US Patent No. 4,667,025; US Patent No. 4,762,779; US Patent No. 4,789,737; US Patent No. 4,824,941; US Patent No. 4,835,263; US Patent No. 4,876,335; US Patent No. 4,904,582; US Patent No. 4,958,013; US Patent No. 5,082,830; US Patent No. 5,112,963; US Patent No. 5,214. ,136; US Patent No. 5,082,830; US Patent No. 5,112,963; US Patent No. 5,214,136; US Patent No. 5,245,022; US Patent No. 5,254,469; US Patent No. 5,258,506; US Patent No. 5,262,536; US Patent No. 5,272,250; US Patent No. 5,292,873; US Patent No. 5,317,098; US Patent No. 5,371,241, US Patent No. 5,391,723; US Patent No. 5,416,203, US Patent No. 5,45 Examples include US Patent No. 1,463; US Patent No. 5,510,475; US Patent No. 5,512,667; US Patent No. 5,514,785; US Patent No. 5,565,552; US Patent No. 5,567,810; US Patent No. 5,574,142; US Patent No. 5,585,481; US Patent No. 5,587,371; US Patent No. 5,595,726; US Patent No. 5,597,696; US Patent No. 5,599,923; US Patent No. 5,599,928, and US Patent No. 5,688,941.
[0072] Nucleic acid base pairing characteristics In some embodiments, non-natural nucleic acids form base pairs with other nucleic acids. In some embodiments, stably incorporated non-natural nucleic acids are non-natural nucleic acids that can form base pairs with other nucleic acids, e.g., natural or non-natural nucleic acids. In some embodiments, stably incorporated non-natural nucleic acids are non-natural nucleic acids that can form base pairs with other non-natural nucleic acids (non-natural nucleic acid base pairs (UBPs)). For example, a first non-natural nucleic acid can form base pairs with a second non-natural nucleic acid. For example, a pair of non-natural nucleotide triphosphates that can form base pairs when incorporated into nucleic acids includes a triphot phosphate of d5SICS (d5SICSTP) and a triphot phosphate of dNaM (dNaMTP). Such non-natural nucleotides may have a ribose or deoxyribose sugar moiety. In some embodiments, non-natural nucleic acids substantially do not form base pairs with natural nucleic acids (A, T, G, C). In some embodiments, stably incorporated non-natural nucleic acids can form base pairs with natural nucleic acids.
[0073] In some embodiments, the stably incorporated non-natural nucleic acid is a non-natural nucleic acid that can form UBPs but substantially does not form base pairs with each of the four natural nucleic acids. In some embodiments, the stably incorporated non-natural nucleic acid is a non-natural nucleic acid that can form UBPs but substantially does not form base pairs with one or more natural nucleic acids. For example, the stably incorporated non-natural nucleic acid substantially does not form base pairs with A, T, and C but can form base pairs with G. For example, the stably incorporated non-natural nucleic acid substantially does not form base pairs with A, T, and G but can form base pairs with C. For example, the stably incorporated non-natural nucleic acid substantially does not form base pairs with C, G, and A but can form base pairs with T. For example, the stably incorporated non-natural nucleic acid substantially does not form base pairs with C, G, and T but can form base pairs with A. For example, the stably incorporated non-natural nucleic acid substantially does not form base pairs with A and T but can form base pairs with C and G. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with A and C, but can form base pairs with T and G. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with A and G, but can form base pairs with C and T. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with C and T, but can form base pairs with A and G. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with C and G, but can form base pairs with T and G. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with T and G, but can form base pairs with A and G. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with G, but can form base pairs with A, T, and C. For example, a stably incorporated non-natural nucleic acid substantially does not form base pairs with A, but can form base pairs with G, T, and C. For example, a stably incorporated non-natural nucleic acid may not substantially form base pairs with T, but may form base pairs with G, A, and C.
[0074] Exemplary examples of non-natural nucleotides capable of forming non-natural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, d5SICS, NAM, dNaM, dTPT3, and mixtures thereof. In some embodiments, the non-natural nucleotides are: [ka] Includes.
[0075] Manipulated organisms In some embodiments, the methods and plasmids disclosed herein are further used to generate manipulated organisms, such as organisms that incorporate and replicate non-natural nucleotides or non-natural nucleic acid base pairs (UBPs) with improved UBP retention, and further transcribe and translate nucleic acids containing non-natural nucleotides or non-natural nucleic acid base pairs into proteins containing non-natural amino acid residues. In some examples, the organism is a semi-synthetic organism (SSO). In some examples, the SSO is a cell.
[0076] In some cases, the cells being utilized are nucleoside triphosphate transporters, which are capable of transporting heterologous proteins, such as non-natural nucleotide triphosphates, into the cell. Cells are genetically transformed with an expression cassette encoding a modified translocation-related protein that enhances the stability of the nucleotide triphosphate transporter, a CRISPR / Cas9 system that removes modifications at non-natural nucleotide triphosphate sites, and / or a polymerase with high fidelity to non-natural nucleic acids, resulting in the non-natural nucleotides being incorporated into the cellular nucleic acid and forming non-natural base pairs, for example, under in vivo conditions. In some cases, the cells further contain high activity for the uptake of non-natural nucleic acids. In some cases, the cells further contain high activity for the translocation of non-natural nucleic acids. In some cases, the cells further contain high polymerase activity for non-natural nucleic acids.
[0077] In some embodiments, Cas9 and sgRNA are encoded on separate plasmids. In some examples, Cas9 and sgRNA are encoded on the same plasmid. In some cases, the nucleic acid molecule encoding Cas9, sgRNA, or a nucleic acid molecule containing non-natural nucleotides is located on one or more plasmids. In some examples, Cas9 is encoded on a first plasmid, and the nucleic acid molecule containing sgRNA and a nucleic acid molecule containing non-natural nucleotides is encoded on a second plasmid. In some examples, the nucleic acid molecule containing Cas9, sgRNA, and a non-natural nucleotide is encoded on the same plasmid. In some examples, the nucleic acid molecule contains two or more non-natural nucleotides.
[0078] In some cases, a first plasmid encoding Cas9 and sgRNA, as well as a second plasmid encoding a nucleic acid molecule containing non-natural nucleotides, are introduced into the engineered microorganism. In some cases, a first plasmid encoding Cas9, as well as a second plasmid encoding a nucleic acid molecule containing sgRNA and non-natural nucleotides, are introduced into the engineered microorganism. In some cases, a plasmid encoding a nucleic acid molecule containing Cas9, sgRNA, and non-natural nucleotides is introduced into the engineered microorganism. In some cases, the nucleic acid molecule contains two or more non-natural nucleotides.
[0079] In some embodiments, living cells are generated in which at least one non-natural nucleotide and / or at least one non-natural base pair (UBP) is incorporated into its nucleic acid. In some examples, if the non-natural mutual base-paired nucleotide is taken up into the cell by the action of a nucleotide triphosphate transporter as its respective triphodes, the non-natural base pair includes a pair of non-natural mutual base-paired nucleotides capable of forming a non-natural base pair under in vivo conditions. The cell can be genetically transformed with an expression cassette encoding a nucleotide triphosphate transporter, as a result of which the nucleotide triphosphate transporter is expressed and available for transporting non-natural nucleotides into the cell. The cell can be a prokaryotic or eukaryotic cell, and the pair of non-natural mutual base-paired nucleotides may be a triphodes of d5SICS (d5SICSTP) and a triphodes of dNaM (dNaMTP) as their respective triphodes.
[0080] In some embodiments, the cells are genetically transformed cells with nucleic acids, for example, cells that have been encoded with an expression cassette that encodes a nucleotide triphosphate transporter capable of transporting such non-natural nucleotides into the cell. The cells may contain a heterologous nucleotide triphosphate transporter that can transport natural or non-natural nucleotide triphosphates into the cell. The cells may contain a heterologous polymerase that is active against non-natural nucleic acids.
[0081] In some cases, the methods described herein also include potassium phosphate and / or phosphate. The process involves contacting genetically transformed cells with non-natural nucleotides in their respective triphot forms in the presence of an atase or nucleotidase inhibitor. During or after such contact, the cells can be placed in a life support medium suitable for cell growth and replication. The cells can be maintained in the life support medium so that each triphot form of the non-natural nucleotide is incorporated into the cellular nucleic acid through at least one replication cycle of the cell. The pairs of non-natural cross-base paired nucleotides may include the triphot form of d5SICS (d5SICSTP) and the triphot form of dNaM (dNaMTP) as their respective triphot forms, and the cells may be E. coli, where d5SICSTP and dNaMTP can be effectively transferred into E. coli by the transporter PtNTT2, and E. coli polymerases such as PolI can effectively use the non-natural triphot forms to replicate DNA, thereby incorporating the non-natural nucleotides and / or non-natural base pairs into the cellular nucleic acid in the cellular environment.
[0082] By the method of the present invention, those skilled in the art can obtain a population of living, growing cells having at least one non-natural nucleotide and / or at least one non-natural base pair (UBP) in at least one nucleic acid maintained in at least some of the individual cells, the at least one nucleic acid which grows stably within the cell, and the cells express a nucleotide triphosphate transporter suitable for cellular uptake of one or more non-natural nucleotides in triphot form when brought into contact with the non-natural nucleotide (e.g., by growing in its presence) in a life-sustaining medium suitable for the growth and replication of organisms.
[0083] Following transport into the cell by nucleotide triphosphate transporters, non-native base-paired nucleotides are incorporated into intracellular nucleic acids by cellular mechanisms, such as the cell's own DNA and / or RNA polymerase, heterologous polymerase, or polymerases that have evolved using directional evolution (Chen T, Romesberg FE, FEBS Lett. January 21, 2014; 588(2): pp. 219-29; Betz K et al., J Am Chem Soc. December 11, 2013; 135(49): 186 (pp. 37-43). Non-natural nucleotides can be incorporated into cellular nucleic acids such as genomic DNA, genomic RNA, mRNA, structural RNA, microRNA, and autonomously replicating nucleic acids (e.g., plasmids, viruses, or vectors).
[0084] In some cases, genetically engineered cells are produced by introducing nucleic acids, such as heterologous nucleic acids, into cells. Any cell described herein may be a host cell and may contain an expression vector. In one embodiment, the host cell is a prokaryotic cell. In another embodiment, the host cell is Escherichia coli. In some embodiments, the cell contains one or more heterologous polynucleotides. Nucleic acid reagents can be introduced into microorganisms using a variety of techniques. Non-limiting examples of methods used to introduce heterologous nucleic acids into various organisms include transformation, transfection, electroporation, sonication-mediated transformation, and particle shock. In some examples, the addition of carrier molecules (e.g., bis-benzimidazole compounds, see, e.g., U.S. Patent No. 5,595,899) can typically enhance DNA uptake into cells that are considered difficult to transform by conventional methods. Conventional methods of transformation are readily available to those skilled in the art, as seen in Maniatis, T., E.F. Fritsch and J. Sambrook (1982) Molecular Cloning:a Laboratory Manual;Cold Spring It can be found at Harbor Laboratory, Cold Spring Harbor, and in New York.
[0085] In some cases, genetic transformation is performed using, but is not limited to, direct introduction of expression cassettes in plasmids, viral vectors, viral nucleic acids, phage nucleic acids, phages, cosmids, and artificial chromosomes, or using intracellular genetic material or cationic liposomes. The result is obtained by introducing a transvestite carrier. Such methods are available in the art and are readily adaptable to use in the methods described herein. The transvestite vector may be any nucleotide construct (e.g., plasmid) used to deliver a gene into a cell and may be, for example, part of a recombinant retrovirus or adenovirus as part of a general strategy for gene delivery (Ram et al., Cancer Res. 53: pp. 83-88, (1993)). Suitable means for transfection, including viral vectors, chemical transvestites, or physicomechanical methods, such as electroporation and direct diffusion of DNA, are described, for example, in Wolff, JA et al., Science, 247, pp. 1465-1468, (1990); and Wolff, JA, Nature, 352, pp. 815-818, (1991).
[0086] For example, nucleotide triphosphate transporters or polymerase nucleic acid molecules, expression cassettes, and / or vectors can be introduced into cells by any method, including, but not limited to, calcium-mediated transfection, electroporation, microinjection, lipofection, particle impact, etc.
[0087] In some cases, cells contain non-natural nucleotide triphosphates incorporated into one or more nucleic acids within the cell. For example, a cell may be a living cell capable of incorporating at least one non-natural nucleotide into DNA or RNA maintained within the cell. A cell can also incorporate at least one non-natural base pair (UBP), containing a pair of non-natural mutual base-paired nucleotides, into nucleic acids within the cell under in vivo conditions, and the non-natural mutual base-paired nucleotides, e.g., each triphosphate, are taken up into the cell by the action of nucleotide triphosphate transporters, and the gene is present in the cell (e.g., introduced) by genetic transformation. For example, when incorporation into nucleic acids maintained within a cell occurs, d5SICS and dNaM can form stable non-natural base pairs that can grow stably by the DNA replication mechanism of the organism, for example, when grown in a life-sustaining medium containing d5SICS and dNaM.
[0088] In some cases, cells are capable of replicating non-natural nucleic acids. Such methods may include genetically transforming cells with an expression cassette encoding a nucleotide triphosphate transporter capable of transporting one or more non-natural nucleotides into the cell under in vivo conditions, each as a triphot phosphate. Alternatively, previously genetically transformed cells may be utilized with an expression cassette capable of expressing the encoded nucleotide triphosphate transporter. The method also includes contacting or exposing the genetically transformed cells to potassium phosphate and each triphot phosphate form of at least one non-natural nucleotide (e.g., two cross-base-paired nucleotides capable of forming a non-natural base pair (UBP)) in a life-sustaining medium suitable for cell proliferation and replication, and maintaining the transformed cells in the life-sustaining medium in the presence of each triphot phosphate form of at least one non-natural nucleotide (e.g., two cross-base-paired nucleotides capable of forming a non-natural base pair (UBP)) for at least one replication cycle of the cell under in vivo conditions.
[0089] In some embodiments, the cell contains stably incorporated non-natural nucleic acids. Some embodiments include a cell (e.g., E. coli) that stably incorporates nucleotides other than A, G, T, and C in nucleic acids maintained within the cell. For example, nucleotides other than A, G, T, and C may be d5SICS, dNaM, and dTPT3, which, upon incorporation into the cell's nucleic acid, can form stable non-natural base pairs within the nucleic acid. In one embodiment, non-natural nucleotides and non-natural base pairs are obtained by an organism genetically transformed against a triphosphate transporter, which then incorporates potassium phosphate and d5SICS, dNaM, and d When grown in a life-sustaining medium containing the triphot form of TPT3, it can be stably propagated by the organism's replication machinery.
[0090] In some cases, cells contain an extended genetic alphabet. Cells may contain stably incorporated non-natural nucleic acids. In some embodiments, cells having an extended genetic alphabet contain non-natural nucleic acids that can form base pairs (bp) with another nucleic acid, e.g., a natural or non-natural nucleic acid. In some embodiments, cells having an extended genetic alphabet contain non-natural nucleic acids that are hydrogen-bonded to another nucleic acid. In some embodiments, cells having an extended genetic alphabet contain non-natural nucleic acids that are not hydrogen-bonded to the other nucleic acid with which they are base-paired. In some embodiments, cells having an extended genetic alphabet contain non-natural nucleic acids that form base pairs with another nucleic acid via hydrophobic interactions. In some embodiments, cells having an extended genetic alphabet contain non-natural nucleic acids that form base pairs with another nucleic acid via non-hydrogen-bonding interactions. Cells having an extended genetic alphabet may be cells that can copy heterologous nucleic acids to form nucleic acids containing non-natural nucleic acids. Cells having an extended genetic alphabet may be cells that contain non-natural nucleic acids that are base-paired with another non-natural nucleic acid (non-natural nucleic acid base pairs (UBPs)).
[0091] In some embodiments, cells derived from non-natural DNA undergo base pairing (UBP) from introduced non-natural nucleotides under in vivo conditions. In some embodiments, potassium phosphate and / or phosphatase inhibitors and / or nucleotidase activity can facilitate the transport of non-natural nucleic acids. The method involves the use of cells expressing heterologous nucleotide triphosphate transporters. When such cells come into contact with one or more nucleotide triphosphates, the nucleotide triphosphates are transported into the cell. The cells may be in the presence of potassium phosphate and / or phosphatase and nucleotidase inhibitors. Non-natural nucleotide triphosphates can be incorporated into intracellular nucleic acids by the cell's natural mechanisms, for example, by mutual base pairing to form non-natural base pairs within the cell's nucleic acids.
[0092] In some embodiments, UBP can be incorporated into cells or populations of cells upon exposure to non-natural triphot phosphates. In some embodiments, UBP can be incorporated into cells or populations of cells substantially consistently upon exposure to non-natural triphot phosphates. In some embodiments, UBP replication does not result in a substantial reduction in growth rate. In some embodiments, the replication and expression of heterologous proteins, such as the transport of nucleotide triphosphates, does not result in a substantial reduction in growth rate.
[0093] In some embodiments, induction of the expression of heterologous genes, such as NTT, within cells can result in slower cell proliferation and higher uptake of non-native nucleic acids compared to cell proliferation and uptake without induction of heterologous gene expression. In some embodiments, induction of the expression of heterologous genes, such as NTT, within cells can result in increased cell proliferation and higher uptake of non-native nucleic acids compared to cell proliferation and uptake without induction of heterologous gene expression.
[0094] In some embodiments, UBP is incorporated during the logarithmic growth phase. In some embodiments, UBP is incorporated during the non-logarithmic growth phase. In some embodiments, UBP is incorporated during the substantially linear growth phase. In some embodiments, UBP is stably incorporated into cells or populations of cells after growing over a period of time. For example, UBP grows over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or 50 or more replications. After growth, it can be stably incorporated into cells or populations of cells. For example, UBP can be stably incorporated into cells or populations of cells after growth for at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours. For example, UBP can be stably incorporated into cells or populations of cells after growth for at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or 31 days. For example, UBP can be stably incorporated into cells or populations of cells after proliferation for at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months. For example, UBP can be stably incorporated into cells or populations of cells after proliferation over at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or 50 years.
[0095] In some embodiments, cells further utilize the polymerase disclosed herein to produce mutant mRNA containing a mutant codon comprising one or more non-natural nucleic acid bases. In some examples, cells further utilize the polymerase disclosed herein to produce mutant tRNA containing a mutant anticodon comprising one or more non-natural nucleic acid bases. In some examples, the mutant anticodon represents a non-natural nucleic acid. In some examples, the anticodon of the mutant tRNA pairs with the codon of the mutant mRNA during translation into synthesis, forming a protein containing a non-natural amino acid.
[0096] As used herein, an amino acid residue may refer to a molecule containing both an amino group and a carboxyl group. Suitable amino acids include, but are not limited to, both D- and L-isomers of naturally occurring amino acids, as well as non-naturally occurring amino acids prepared by organic synthesis or other metabolic pathways. The term amino acid, as used herein, includes, but is not limited to, α-amino acids, natural amino acids, non-natural amino acids, and amino acid analogs.
[0097] The term "α-amino acid" can refer to a molecule that contains both an amino group and a carboxyl group bonded to a carbon atom that represents an α-carbon.
[0098] The term "β-amino acid" can refer to a molecule that contains both an amino group and a carboxyl group in its β-stereoconfiguration.
[0099] "Naturally occurring amino acids" can refer to any one of the 12 amino acids commonly found in naturally synthesized peptides, and are known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V.
[0100] The table below provides an overview of the properties of natural amino acids.
[0101] [Table 1]
[0102] "Hydrophobic amino acids" include small hydrophobic amino acids and large hydrophobic amino acids. "Small hydrophobic amino acids" may be glycine, alanine, proline, and their analogues. "Large hydrophobic amino acids" may be valine, leucine, isoleucine, phenylalanine, methionine, tryptophan, and their analogues. "Polar amino acids" may be serine, threonine, asparagine, glutamine, cysteine, tyrosine, and their analogues. "Charged amino acids" may be lysine, arginine, histidine, aspartate, glutamate, and their analogues.
[0103] An "amino acid analog" is a molecule that is structurally similar to an amino acid and may be substituted for an amino acid in the formation of a peptide-mimicking macrocyclic molecule. Amino acid analogs include, but are not limited to, β-amino acids and amino acids, in which the amino or carboxyl group is substituted with a similar reactive group (e.g., substitution of a primary amine with a secondary or tertiary amine, or substitution of a carboxyl group with an ester).
[0104] "Non-natural amino acids" are the 20 amino acids that are typically found in naturally synthesized peptides. It can be an amino acid that is not one of the ano acids, and is known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V.
[0105] Amino acid analogs may include β-amino acid analogs. Examples of β-amino acid analogs, but are not limited to, cyclic β-amino acid analogs; β-alanine; (R)-β-phenylalanine; (R)-1,2,3,4-tetrahydroisoquinoline-3-acetic acid; (R)-3-amino-4-(1-naphthyl)-butyric acid; (R)-3-amino-4-(2,4-dichlorophenyl)butyric acid; (R)-3-amino-4-(2-chlorophenyl)-butyric acid; (R)-3-amino-4-(2-cyanophenyl)-butyric acid; (R)-3-amino-4-(2-fluorophenyl)-butyric acid; (R)-3-amino-4-(2-fri- (R)-butyric acid; (R)-3-amino-4-(2-methylphenyl)-butyric acid; (R)-3-amino-4-(2-naphthyl)-butyric acid; (R)-3-amino-4-(2-thienyl)-butyric acid; (R)-3-amino-4-(2-trifluoromethylphenyl)-butyric acid; (R)-3-amino-4-(3,4-dichlorophenyl)butyric acid; (R)-3-amino-4-(3,4-difluorophenyl)butyric acid; (R)-3-amino-4-(3-benzothienyl)-butyric acid; (R)-3-amino-4-(3-chlorophenyl)-butyric acid; (R)-3-amino-4-(3 (R)-3-amino-4-(3-fluorophenyl)-butyric acid; (R)-3-amino-4-(3-methylphenyl)-butyric acid; (R)-3-amino-4-(3-pyridyl)-butyric acid; (R)-3-amino-4-(3-thienyl)-butyric acid; (R)-3-amino-4-(3-trifluoromethylphenyl)-butyric acid; (R)-3-amino-4-(4-bromophenyl)-butyric acid; (R)-3-amino-4-(4-chlorophenyl)-butyric acid; (R)-3-amino-4-(4-cyanophenyl)-butyric acid; (R)-3-amino- 4-(4-fluorophenyl)-butyric acid; (R)-3-amino-4-(4-iodophenyl)-butyric acid; (R)-3-amino-4-(4-methylphenyl)-butyric acid; (R)-3-amino-4-(4-nitrophenyl)-butyric acid; (R)-3-amino-4-(4-pyridyl)-butyric acid; (R)-3-amino-4-(4-trifluoromethylphenyl)-butyric acid; (R)-3-amino-4-pentafluorophenylbutyric acid; (R)-3-amino-5-hexenoic acid; (R)-3-amino-5-hexic acid; (R)-3-amino-5-phenylpentanoic acid;(R)-3-amino-6-phenyl-5-hexenoic acid; (S)-1,2,3,4-tetrahydroisoquinoline-3-acetic acid; (S)-3-amino-4-(1-naphthyl)-butyric acid; (S)-3-amino-4-(2,4-dichlorophenyl)butyric acid; (S)-3-amino-4-(2-chlorophenyl)butyric acid; (S)-3-amino-4-(2-cyanophenyl)butyric acid; (S)-3-amino-4-(2-fluorophenyl)butyric acid; (S)-3-amino-4-(2-furyl)butyric acid; (S)-3-amino- No-4-(2-methylphenyl)-butyric acid; (S)-3-amino-4-(2-naphthyl)-butyric acid; (S)-3-amino-4-(2-thienyl)-butyric acid; (S)-3-amino-4-(2-trifluoromethylphenyl)-butyric acid; (S)-3-amino-4-(3,4-dichlorophenyl)butyric acid; (S)-3-amino-4-(3,4-difluorophenyl)butyric acid; (S)-3-amino-4-(3-benzothienyl)-butyric acid; (S)-3-amino-4-(3-chlorophenyl)-butyric acid; (S)-3-amino- 4-(3-cyanophenyl)-butyric acid; (S)-3-amino-4-(3-fluorophenyl)-butyric acid; (S)-3-amino-4-(3-methylphenyl)-butyric acid; (S)-3-amino-4-(3-pyridyl)-butyric acid; (S)-3-amino-4-(3-thienyl)-butyric acid; (S)-3-amino-4-(3-trifluoromethylphenyl)-butyric acid; (S)-3-amino-4-(4-bromophenyl)-butyric acid; (S)-3-amino-4-(4-chlorophenyl)butyric acid; (S)-3-amino-4-(4-cyanophenyl)butyric acid (S)-3-amino-4-(4-fluorophenyl)butyric acid; (S)-3-amino-4-(4-iodophenyl)butyric acid; (S)-3-amino-4-(4-methylphenyl)butyric acid; (S)-3-amino-4-(4-nitrophenyl)butyric acid; (S)-3-amino-4-(4-pyridyl)butyric acid; (S)-3-amino-4-(4-trifluoromethylphenyl)butyric acid; (S)-3-amino-4-pentafluorophenylbutyric acid; (S)-3-amino-5-hexenoic acid; (; (S)-3-amino-5-hexic acid; (S)-3-amino-5-phenylpentanoic acid; (S)-3-amino-6-phenyl-5-hexenoic acid; 1,2,5,6-tetrahydropyridine-3-carboxylic acid; 1,2,5,6-tetrahydropyridine-4-carboxylic acid; 3-amino-3-(2-chlorophenyl)-propionic acid; 3-amino-3-(2-thienyl)-propionic acid; 3-amino-3-(3-bromophenyl)-propionic acid; 3-amino-3-(4-chlorophenyl )-propionic acid; 3-amino-3-(4-methoxyphenyl)-propionic acid; 3-amino-4,4,4-trifluorobutyric acid; 3-aminoadipic acid; D-β-phenylalanine; β-leucine; L-β-homoalanine; L-β-homoaspartic acid γ-benzyl ester; L-β-homoglutamic acid δ-benzyl ester; L-β-homoisoleucine; L-β-homoleucine; L-β-homomethionine; L-β-homophenylalanine; L-β-homoproline; L-β- Homotryptophan; L-β-homovaline; L-Nω-benzyloxycarbonyl-β-homolysine; Nω-L-β-homoarginine; O-benzyl-L-β-homohydroxyproline; O-benzyl-L-β-homoserine; O-benzyl-L-β-homothreonine; O-benzyl-L-β-homotyrosine; γ-trityl-L-β-homoasparagine; (R)-β-phenylalanine; L-β-homoaspartic acid γ-t-butyl ester; L-β-homoglutamic acid δ-t- Examples include butyl esters; L-Nω-β-homolidine; Nδ-trityl-L-β-homoglutamine; Nω-2,2,4,6,7-pentamethyl-dihydrobenzofuran-5-sulfonyl-L-β-homoarginine; Ot-butyl-L-β-homohydroxy-proline; Ot-butyl-L-β-homoserine; Ot-butyl-L-β-homothreonine; Ot-butyl-L-β-homotyrosine; 2-aminocyclopentanecarboxylic acid; and 2-aminocyclohexanecarboxylic acid.
[0106] Amino acid analogs may include analogs of alanine, valine, glycine, or leucine. Examples of amino acid analogs of alanine, valine, glycine, and leucine, but are not limited to: α-methoxyglycine; α-allyl-L-alanine; α-aminoisobutyric acid; α-methylleucine; β-(1-naphthyl)-D-alanine; β-(1-naphthyl)-L-alanine; β-(2-naphthyl)-D-alanine; β-(2-naphthyl)-L-alanine; β-(2-pyridyl)-D-alanine; β-(2-pyridyl)-L-alanine; β-(2-thienyl )-D-alanine; β-(2-thienyl)-L-alanine; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; β-(3-pyridyl)-D-alanine; β-(3-pyridyl)-L-alanine; β-(4-pyridyl)-D-alanine; β-(4-pyridyl)-L-alanine; β-chloro-L-alanine; β-cyano-L-alanine; β-cyclohexyl-D-alanine; β-cyclohexyl-L-alanine; β-cyclo Penten-1-yl-alanine; β-cyclopentyl-alanine; β-cyclopropyl-L-Ala-OH.dicyclohexylammonium salt; β-t-butyl-D-alanine; β-t-butyl-L-alanine; γ-aminobutyric acid; L-α,β-diaminopropionic acid; 2,4-dinitrophenylglycine; 2,5-dihydro-D-phenylglycine; 2-amino-4,4,4-trifluorobutyric acid; 2-fluorophenylglycine; 3-amino-4,4 4-Trifluorobutyric acid; 3-Fluoro-valine; 4,4,4-Trifluoro-valine; 4,5-Dehydro-L-leu-OH.Dicyclohexylammonium salt; 4-Fluoro-D-phenylglycine; 4-Fluoro-L-phenylglycine; 4-Hydroxy-D-phenylglycine; 5,5,5-Trifluoro-leucine; 6-Aminohexanoic acid; Cyclopentyl-D-Gly-OH.Dicyclohexylammonium salt;Cyclopentyl-Gly-OH.Dicyclohexylammonium salt; D-α,β-diaminopropionic acid; D-α-aminobutyric acid; D-α-t-butylglycine; D-(2-thienyl)glycine; D-(3-thienyl)glycine; D-2-aminocaproic acid; D-2-indanylglycine; D-allylglycine-dicyclohexylammonium salt; D-cyclohexylglycine; D-norvaline; D-phenylglycine; β-aminobutyric acid; β-aminoisobutyric acid; (2-bromophenyl)glycine; (2-methoxyglycine) (Ciphenyl)glycine; (2-methylphenyl)glycine; (2-thiazoyl)glycine; (2-thienyl)glycine; 2-amino-3-(dimethylamino)-propionic acid; L-α,β-diaminopropionic acid; L-α-aminobutyric acid; L-α-t-butylglycine; L-(3-thienyl)glycine; L-2-amino-3-(dimethylamino)-propionic acid; L-2-aminocaproic acid dicyclohexylammonium salt; L-2-indanylglycine; L-allylglycine dicyclohexylammonium salt; L-cyclo Hexylglycine; L-phenylglycine; L-propargylglycine; L-norvaline; N-α-aminomethyl-L-alanine; D-α,γ-diaminobutyric acid; L-α,γ-diaminobutyric acid; β-cyclopropyl-L-alanine; (N-β-(2,4-dinitrophenyl))-L-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-D-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl) (N-β-4-methyltrityl)-L-α,β-diaminopropionic acid; (N-β-allyloxycarbonyl)-L-α,β-diaminopropionic acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-D-α,γ-diaminobutyric acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-L-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-D-α,γ-diaminobutyric acid; ( Examples include N-γ-4-methyltrityl)-L-α,γ-diaminobutyric acid; (N-γ-allyloxycarbonyl)-L-α,γ-diaminobutyric acid; D-α,γ-diaminobutyric acid; 4,5-dehydro-L-leucine; cyclopentyl-D-Gly-OH; cyclopentyl-Gly-OH; D-allylglycine; D-homocyclohexylalanine; L-1-pyrenylalanine; L-2-aminocaproic acid; L-allylglycine; L-homocyclohexylalanine; and N-(2-hydroxy-4-methoxy-Bzl)-Gly-OH.
[0107] Amino acid analogs may include analogs of arginine or lysine. Examples of amino acid analogs of arginine and lysine include, but are not limited to, citrulline; L-2-amino-3-guanidinopropionic acid; L-2-amino-3-ureidopropionic acid; L-citrulline; Lys(Me)2-OH; Lys(N3)-OH; Nδ-benzyloxycarbonyl-L-ornithine; Nω-nitro-D-arginine; Nω-nitro-L-arginine; α-methyl-ornithine; 2,6-diaminoheptanodic acid; L-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxocyclohexe-1-ylidene)ethyl)-D-ornithine Examples include (Nδ-1-(4,4-dimethyl-2,6-dioxocyclohexe-1-ylidene)ethyl)-L-ornithine; (Nδ-4-methyltrityl)-D-ornithine; (Nδ-4-methyltrityl)-L-ornithine; D-ornithine; L-ornithine; Arg(Me)(Pbf)-OH; Arg(Me)2-OH (asymmetric); Arg(Me)2-OH (symmetric); Lys(ivDde)-OH; Lys(Me)2-OH.HCl; Lys(Me3)-OH chloride; Nω-nitro-D-arginine; and Nω-nitro-L-arginine.
[0108] The amino acid analogs may include analogs of aspartic acid or glutamic acid. Examples of amino acid analogs of aspartic acid and glutamic acid include, but are not limited to, α-methyl-D-aspartic acid; α-methyl-glutamic acid; α-methyl-L-aspartic acid; γ-methylene-glutamic acid; (N-γ-ethyl)-L-glutamine; [N-α-(4-aminobenzoyl)]-L-glutamic acid; 2,6-diaminopimelic acid; L-α-aminosuberic acid; D-2-aminoadipic acid; D-α-aminosuberic acid; α-aminopimelic acid; iminodiacetic acid; L-2-aminoadipic acid; threo-β-methyl-aspartic acid; γ-carboxy-D-glutamic acid γ,γ-di-t-butyl ester; γ-carboxy-L-glutamic acid γ,γ-di-t-butyl ester; Glu(OAll)-OH; L-Asu(OtBu)-OH; and pyroglutamic acid.
[0109] Amino acid analogs may include analogs of cysteine and methionine. Examples of amino acid analogs of cysteine and methionine include, but are not limited to, Cys(farnesyl)-OH, Cys(farnesyl)-OMe, α-methylmethionine, Cys(2-hydroxyethyl)-OH, Cys(3-aminopropyl)-OH, 2-amino-4-(ethylthio)butyric acid, butionine, butionine sulfoximine, ethionine, methionine methylsulfonium chloride, selenomethionine, cysteic acid, [2-(4-pyridyl)ethyl]-DL-penicillamine, [2-(4-pyridyl)ethyl]-L-cysteine, 4-methoxybenzyl-D-penicillamine, 4-methoxybenzyl-L-penicillamine, 4-methylbenzyl-D-penicillamine Examples include mine, 4-methylbenzyl-L-penicillamine, benzyl-D-cysteine, benzyl-L-cysteine, benzyl-DL-homocysteine, carbamoyl-L-cysteine, carboxyethyl-L-cysteine, carboxymethyl-L-cysteine, diphenylmethyl-L-cysteine, ethyl-L-cysteine, methyl-L-cysteine, t-butyl-D-cysteine, trityl-L-homocysteine, trityl-D-penicillamine, cystathionine, homocystine, L-homocystine, (2-aminoethyl)-L-cysteine, seleno-L-cystine, cystathionine, Cys(StBu)-OH, and acetamidomethyl-D-penicillamine.
[0110] Amino acid analogs may include analogs of phenylalanine and tyrosine. Examples of amino acid analogs of phenylalanine and tyrosine include β-methylphenylalanine, β-hydroxyphenylalanine, α-methyl-3-methoxy-DL-phenylalanine, α-methyl-D-phenylalanine, α-methyl-L-phenylalanine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, 2,4-dichlorophenylalanine, 2-(trifluoromethyl)-D-phenylalanine, and 2-(trifluoromethyl)-L-phenylalanine. Nin, 2-bromo-D-phenylalanine, 2-bromo-L-phenylalanine, 2-chloro-D-phenylalanine, 2-chloro-L-phenylalanine, 2-cyano-D-phenylalanine, 2-cyano-L-phenylalanine, 2-fluoro-D-phenylalanine, 2-fluoro-L-phenylalanine, 2-methyl-D-phenylalanine, 2-methyl-L-phenylalanine, 2-nitro-D-phenylalanine, 2-nitro-L-phenylalanine, 2;4;5-Trihydroxyphenylalanine, 3,4,5-Trifluoro-D-phenylalanine, 3,4,5-Trifluoro-L-phenylalanine, 3,4-Dichloro-D-phenylalanine, 3,4-Dichloro-L-phenylalanine, 3,4-Difluoro-D-phenylalanine, 3,4-Difluoro-L-phenylalanine, 3,4-Dihydroxy-L-phenylalanine, 3,4-Dimethoxy-L-phenylalanine, 3,5,3'-Triiodine -L-thyronine, 3,5-diiodo-D-tyrosine, 3,5-diiodo-L-tyrosine, 3,5-diiodo-L-thyronine, 3-(trifluoromethyl)-D-phenylalanine, 3-(trifluoromethyl)-L-phenylalanine, 3-amino-L-tyrosine, 3-bromo-D-phenylalanine, 3-bromo-L-phenylalanine, 3-chloro-D-phenylalanine, 3-chloro-L-phenylalanine, 3-chloro-L-tyrosine, 3-shea No-D-phenylalanine, 3-cyano-L-phenylalanine, 3-fluoro-D-phenylalanine, 3-fluoro-L-phenylalanine, 3-fluoro-tyrosine, 3-iodo-D-phenylalanine, 3-iodo-L-phenylalanine, 3-iodo-L-tyrosine, 3-methoxy-L-tyrosine, 3-methyl-D-phenylalanine, 3-methyl-L-phenylalanine, 3-nitro-D-phenylalanine, 3-nitro-L-phenylalanine Nin, 3-nitro-L-tyrosine, 4-(trifluoromethyl)-D-phenylalanine, 4-(trifluoromethyl)-L-phenylalanine, 4-amino-D-phenylalanine, 4-amino-L-phenylalanine, 4-benzoyl-D-phenylalanine, 4-benzoyl-L-phenylalanine, 4-bis(2-chloroethyl)amino-L-phenylalanine, 4-bromo-D-phenylalanine, 4-bromo-L-phenylalanine, 4-; Examples include chloro-D-phenylalanine, 4-chloro-L-phenylalanine, 4-cyano-D-phenylalanine, 4-cyano-L-phenylalanine, 4-fluoro-D-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-D-phenylalanine, 4-iodo-L-phenylalanine, homophenylalanine, thyroxine, 3,3-diphenylalanine, thyronine, ethyl-tyrosine, and methyl-tyrosine.
[0111] Amino acid analogs may include proline analogs. Examples of proline amino acid analogs include, but are not limited to, 3,4-dehydro-proline, 4-fluoro-proline, cis-4-hydroxy-proline, thiazolidinedione-2-carboxylic acid, and trans-4-fluoro-proline.
[0112] Amino acid analogs may include serine and threonine analogs. Examples of serine and threonine amino acid analogs include, but are not limited to, 3-amino-2-hydroxy-5-methylhexanoic acid, 2-amino-3-hydroxy-4-methylpentanoic acid, 2-amino-3-ethoxybutanoic acid, 2-amino-3-methoxybutanoic acid, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-ethoxypropionic acid, 4-amino-3-hydroxybutanoic acid, and α-methylserine.
[0113] Amino acid analogs can include tryptophan analogs. Examples of tryptophan amino acid analogs, but are not limited to, α-methyltryptophan; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; 1-methyltryptophan; 4-methyltryptophan; 5-benzyloxytryptophan; 5-bromotryptophan; 5-chlorotryptophan; 5-fluorotryptophan; 5-hydroxytryptophan; 5-hydroxy-L-tryptophan; 5-methoxytryptophan; 5-methoxy-L-tryptophan; 5-methyltryptophan; 6-bromotryptophan Examples include 6-chloro-D-tryptophan, 6-chloro-tryptophan, 6-fluoro-tryptophan, 6-methyl-tryptophan, 7-benzyloxy-tryptophan, 7-bromo-tryptophan, 7-methyl-tryptophan, D-1,2,3,4-tetrahydro-norharmann-3-carboxylic acid, 6-methoxy-1,2,3,4-tetrahydronorharmann-1-carboxylic acid, 7-azatryptophan, L-1,2,3,4-tetrahydro-norharmann-3-carboxylic acid, 5-methoxy-2-methyl-tryptophan, and 6-chloro-L-tryptophan.
[0114] Amino acid analogs can be racemic. In some cases, the D isomer of the amino acid analog is used. In some cases, the L isomer of the amino acid analog is used. In some cases, the amino acid analog contains a chiral center in an R or S configuration. Sometimes, the amino group of the β-amino acid analog is substituted with a protecting group, such as tert-butyloxycarbonyl (BOC group), 9-fluorenylmethyloxycarbonyl (FMOC), tosyl, etc. Sometimes, the carboxylic acid functional group of the β-amino acid analog is protected, for example, as its ester derivative. In some cases, salts of the amino acid analog are used.
[0115] In some embodiments, the non-natural amino acid is the non-natural amino acid described in Liu CC, Schultz, PG, Annu. Rev. Biochem. 2010, pp. 79, 413. In some embodiments, the non-natural amino acid includes N6(2-azidoethoxy)-carbonyl-L-lysine.
[0116] cell type In some embodiments, many types of cells / microorganisms are transformed or genetically modified, for example. It is used to perform the following actions. In some embodiments, the cells are prokaryotic or eukaryotic cells. In some cases, the cells are microorganisms such as bacterial cells, fungal cells, yeast, or single-celled protozoa. In other cases, the cells are eukaryotic cells such as cultured animals, plants, or human cells. In further cases, the cells are present in living organisms such as plants or animals.
[0117] In some embodiments, the manipulated microorganism is a single-celled organism that is often capable of division and reproduction. Microorganisms may include one or more of the following forms: aerobic bacteria, anaerobic bacteria, filamentous fungi, non-filamentous fungi, haploid, diploid, trophoblasts, and / or non-trophoblasts. In certain embodiments, the manipulated microorganism is a prokaryotic microorganism (e.g., bacteria), and in certain embodiments, the manipulated microorganism is a non-prokaryotic microorganism. In some embodiments, the manipulated microorganism is a eukaryotic microorganism (e.g., yeast, fungi, amoeba). In some embodiments, the manipulated microorganism is a fungus. In some embodiments, the manipulated organism is yeast.
[0118] Any suitable yeast is selected as a host microorganism, engineered microorganism, genetically modified organism, or source of heterologous or modified polynucleotides. Yeasts include, but are not limited to, Yarouia yeast (e.g., Y. lipolytica (formerly Candida lipolytica)). Classified as lipolytica), Candida yeast (e.g., C. revkaufi, C. viswanathii, C. pulcherrima, C. tropicalis, C. utilis), Rhodotorula yeast (e.g., R. glutinus, R. graminis), Rhodosporidiium yeast (e.g., R. toruloides), Saccharomyces yeast (e.g., S. se This includes S. cerevisiae, S. bayanus, S. pastorianus, S. carlsbergensis, Cryptococcus yeasts, Trichosporon yeasts (e.g., T. pullans, T. cutaneum, Pichia yeasts (e.g., P. pastoris), and Lipomyces yeasts (e.g., L. starkeyii, L. lipoferus).In some embodiments, suitable yeasts include Arachniotus, Aspergillus, Aureobasidium, Auxarthron, Blastomyces, Candida, Chrysosporium, and Chrysosporium devariomyces. Debaryomyces, Coccidiodes, Cryptococcus, Gymnoascus, Hansenula, Histoplasma, Issatchenkia, Kluyveromyces, Lipomyces, Lssatchenkia, Microsporum, Myxotrichum, Myxozyma, Eudiodendron (O These are from the genera Idiodendron, Pachysolen, Penicillium, Pichia, Rhodosporidium, Rhodotorula, Rhodotorula, Saccharomyces, Schizosaccharomyces, Scopulariopsis, Sepedonium, Trichosporon, or Yarrowia. In some embodiments, appropriate yeast is used. My mother is Arachniotus flavoluteus, Aspergillus flavus, Aspergillus fumigatus, Aspergillus niger, Aureobasidium pullulans, Auxarthron thaxteri, Blastomyces dermatitidis, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida gilliermonzii Candida guilliermondii), Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobiii (Candida xestobii), Chrysosporuim keratinophilum, Coccidiodes immitis, Cryptococcus albidus var. difluensCryptococcus diffluens), Cryptococcus laurentii, Cryptococcus neofomans, Debaryomyces hansenii, Gymnoascus dugwayensis, Hansenula anomala, Histoplasma capsulatum, Issatchenkia occidentalis, Isstachenkia orientalis, Kluyveromyces lactis, Kluyveromyces marxianus marxianus), Kluyveromyces thermotolerans, Kluyveromyces warthii. waltii), Lipomyces lipoferus, Lipomyces starkeyii, Microsporum gypseum, Myxotrichum deflexum, Oidiodendron echinulatum, Pachysolen tannophilis, Penicillium notatum, Pichia anomala Anomala), Pichia pastoris, Pichia stipitis, Rhodosporidium toruloides, Rhodotorula glutinus, Rhodotorula graminis, Saccharomyces cerevisiae, Saccharomyces kluyveri, Schizosaccharomyces pombe These include species such as *Ccharomyces pombe*, *Scopulariopsis acremonium*, *Sepedonium chrysospermum*, *Trichosporon cutaneum*, *Trichosporon pullans*, *Yarrowia lipolytica*, or *Yarrowia lipolytica* (formerly classified as *Candida lipolytica*). In some embodiments, the yeast is a strain of Y. lipolytica, including but not limited to ATCC20362, ATCC8862, ATCC18944, ATCC20228, ATCC76982, and LGAM S(7)1 (Papanikolaou S. and Aggelis G., Bioresour. Technol. 82(1):43~9 (2002)). In certain embodiments, the yeast is a Candida species (i.e., Candida spp.) yeast. Any suitable Candida species can be used and / or genetically modified to produce fatty dicarboxylic acids (e.g., octanedioic acid, decandioic acid, dodecandioic acid, tetradecandioic acid, hexadecanedioic acid, octadecandioic acid, eicosanedioic acid).In some embodiments, suitable Candida species include, but are not limited to, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, and Candida levukaufi, as described herein. This includes Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, and any other Candida species yeast. Non-exclusive examples of Candida strains include, but are not limited to, sAA001 (ATCC20336), sAA002 (ATCC20913), sAA003 (ATCC20962), sAA496 (US2012 / 0077252), sAA106 (US2012 / 0077252), SU-2 (ura3- / ura3-), and H5343 (beta-oxidation blocked; US Patent No. 5648247). Any suitable strain from the Candida yeast species is used as the parent strain for genetic modification.
[0119] Yeast genera, species, and strains are often very closely related in genetic material, which can make them difficult to distinguish, classify, and / or name. In some cases, strains of C. lipolytica and Y. lipolytica can be difficult to distinguish, classify, and / or name, and in some cases, they may be considered the same organism. In some cases, various strains of C. tropicalis and C. viswanathii can be difficult to distinguish, classify, and / or name (see, e.g., Arie et al., J. Gen. Appl. Microbiol., 46, 257-262 (2000)). Several C. tropicalis and C. viswanathii strains obtained from ATCC and other commercial or academic suppliers can be considered equivalent and equally suitable for the embodiments described herein. In some embodiments, C. tropicalis Some parent plants of C. tropicalis and C. viswanathii are considered to differ only in name.
[0120] Any suitable fungus is selected as a host microorganism, an engineered microorganism, or a source of heterologous polynucleotides. Non-limiting examples of fungi include, but are not limited to, Asperchilus fungi (e.g., A. parasiticus, A. nidulans), Traustochtrium fungi, Schizochtrium fungi, and Rhizopus fungi (e.g., R. arrhizus, R. oryzae, R. nigricans). In some embodiments, the fungus is an A. parasiticus strain, including but not limited to strain ATCC24690, and in certain embodiments, the fungus is an A. nidulans strain, including but not limited to strain ATCC38163.
[0121] Any suitable prokaryotes are selected as the host microorganism, the manipulated microorganism, or the source of heterologous polynucleotides. Gram-negative or Gram-positive bacteria are selected. Examples of bacteria include, but are not limited to, Bacillus bacteria (e.g., B. subtilis, B. megaterium), Acinetobacter bacteria, Norcardia bacteria, Xanthobacter bacteria, Sescherichia bacteria (e.g., Escherichia coli (e.g., strains DH10B, Stbl2, DH5-alpha, DB3, DB3.1), DB4, DB5, JDP682), and ccdA-over (e.g., US Patent Application No. 09 / 518, 18) This includes bacteria such as Streptomyces, Erwinia, Klebsiella, Serratia (e.g., S. marcessans), Pseudomonas (e.g., P. aeruginosa), Salmonella (e.g., S. typhimurium, S. typhi), and Megasphaera (e.g., Megasphaera elsdenii). Bacteria include, but are not limited to, photosynthetic bacteria (e.g., green non-sulfur bacteria (e.g., Chloroflexus bacteria (e.g., C. aurantiacus), Chloronema bacteria (e.g., C. gigateum), green sulfur bacteria (e.g., Chlorobium bacteria (e.g., C. limicola), Perodictyon bacteria (e.g., P. luteolum)), This also includes violet sulfur bacteria (e.g., Chromatium bacteria (e.g., C. okenii)), and violet non-sulfur bacteria (e.g., Rhodospirillum bacteria (e.g., R. rubrum), Rhodobacter bacteria (e.g., R. sphaeroides, R. capsulatus), and Rhodomicrobium bacteria (e.g., R. vanellii).
[0122] Cells derived from non-microorganisms can be used as host microorganisms, engineered microorganisms, or sources of heterologous polynucleotides. Examples of such cells include, but are not limited to, insect cells (e.g., Drosophila melanogaster, Spodoptera (e.g., S. frugiperda Sf9 or Sf21 cells), and Trichopulsa (e.g., High-Five cells); nematode cells (e.g., C. elegans cells); bird cells; amphibian cells (e.g., Xenopus laevis cells); reptile cells; mammalian cells (e.g., NIH3T3, 293, CHO, COS, VERO, C127, BHK, Per-C6, Bowes melanoma, and HeLa cells); and plant cells (e.g., Arabidopsis thaliana, tobacco, Cuphea acinifolia, Cuphea aekipetala). aequipetala), Cuphea angustifolia a) Cuphea appendiculata, Cuphea avigera, Cuphea avigera var. pulcherrima, Cuphea axilliflora, Cuphea bahiensis, Cuphea baillonis, Cuphea brachypoda, Cuphea bustamanta, Cuphea calcarata, Cuphea calophylla, Cuphea calophylla subsp. mesostemon, Cuphea calsagenesis carthagenensis, Cuphea circaeoides Cuphea circaeoides), Cuphea confertiflora, Cuphea cordata, Cuphea crassiflora, Cuphea cyanea, Cuphea decandra, Cuphea denticulata, Cuphea disperma, Cuphea epilobiifolia, Cuphea ericoides, Cuphea flava, Cuphea flavisetula, Cuphea fuchsiifolia, Cuphea gaumeli (Cuphea Cuphea gaumeri), Cuphea glutinosa, Cuphea heterophylla, Cuphea hookeriana, Cuphea hyssopifolia (Mexican willow), Cuphea hyssopoides, Cuphea ignea, Cuphea ingrata, Cuphea jorullensis, Cuphea lanceolata, Cuphea linarioides, Cuphea llavea, Cuphea lophostoma, Cuphea lutea Cuphea lutea), Cuphea lutescens, Cuphea melanium, Cuphea melvilla, Cuphea micrantha, Cuphea micropetala, Cuphea mimuloides, Cuphea nitidula, Cuphea palustris, Cuphea parsonsia (Cuphea Cuphea parsonsia), Cuphea pascuorum, Cuphea paucipetala, Cuphea procumbens, Cuphea pseudosilene, Cuphea pseudovaccinium, Cuphea pulchra, Cuphea racemosa, Cuphea repens, Cuphea salicifolia, Cuphea salvadorensis, Cuphea schumannii, Cuphea sessiliflora, Cuphea sessilifolia (Cuphea Cuphea sessilifolia, Cuphea setosa, Cuphea spectabilis, Cuphea spermacoce, Khufu Cuphea splendida, Cuphea splendida var. viridiflava, Cuphea strigulosa, Cuphea subuligera, Cuphea teleandra, Cuphea thymoides, Cuphea tolucana, Cuphea urens, Cuphea utriculosa, Cuphea viscosissima, Cuphea watsoniana, Cuphea wrightii, Cuphea lanceorata It includes lanceolata.
[0123] Microorganisms or cells used as host organisms or sources of heterologous polynucleotides are commercially available. The microorganisms and cells described herein, and other suitable microorganisms and cells, are available, for example, from Invitrogen Corporation (Carlsbad, CA), American Type Culture Collection (Manassas, Virginia), and Agricultural Research Culture Collection (NRRL; Peoria, Illinois). Host microorganisms and engineered microorganisms are provided in any suitable form. For example, such microorganisms are provided as primary cultures or in liquid cultures or solid cultures (e.g., agar-based media), having been subcultured one or more times (e.g., diluted and cultured). Microorganisms are also provided in frozen or dried form (e.g., lyophilized). Microorganisms are provided in any suitable concentration.
[0124] polymerase A particularly useful function of polymerases is to catalyze the polymerization of nucleic acid chains using existing nucleic acids as templates. Other useful functions are described elsewhere in this specification. Examples of useful polymerases include DNA polymerase and RNA polymerase.
[0125] The ability of polymerases to improve the specificity, processability, or other characteristics of non-natural nucleic acids is considered highly desirable in a variety of contexts where the incorporation of non-natural nucleic acids is desired, including, for example, amplification, sequencing, labeling, detection, cloning, and many others. This invention provides polymerases whose properties are modified with respect to non-natural nucleic acids, methods for producing such polymerases, methods for using such polymerases, and many other features that will be revealed by a full examination below.
[0126] In some cases, those disclosed herein include polymerases for incorporating non-natural nucleic acids into a growth template copy, for example, during DNA amplification. In some embodiments, polymerases can be modified such that the active site of the polymerase is altered to reduce inhibition of steric entry of non-natural nucleic acids into the active site. In some embodiments, polymerases can be modified to provide complementarity to one or more non-natural features of the non-natural nucleic acid. Such polymerases can be expressed or manipulated intracellularly to safely incorporate UBP into cells. Accordingly, the present invention includes compositions comprising heterogeneous or recombinant polymerases and methods of using the same.
[0127] Polymerases can be modified using methods related to protein engineering. For example, molecular modeling can be performed based on the crystal structure to identify locations in the polymerase where mutations can be made to alter the target activity. Groups identified as targets for substitution can be analyzed using energy minimization modeling, homology modeling, etc. The selected group may be replaced using conserved amino acid substitutions, e.g., those described in Bordo et al., J Mol Biol 217:721~729 (1991) and Hayes et al., Proc Natl Acad Sci, USA 99:15926~15931 (2002).
[0128] Any of the various polymerases can be used in the methods or compositions described herein, such as those comprising enzymes based on proteins isolated from biological systems and their functional variants. References to specific polymerases, such as those embodied below, will be understood to include their functional variants unless otherwise indicated. In some embodiments, the polymerase is a wild-type polymerase. In some embodiments, the polymerase is a modified or mutated polymerase.
[0129] Polymerases can also be used that have features to improve the entry of non-natural nucleic acids into the active site region and to coordinate with non-natural nucleotides in the active site region. In some embodiments, the modified polymerase has a modified nucleotide binding site.
[0130] In some embodiments, the modified polymerase exhibits specificity for non-natural nucleic acids, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of the wild-type polymerase to non-natural nucleic acids. In some embodiments, the modified or wild-type polymerase exhibits specificity for non-natural nucleic acids containing modified sugars, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of the wild-type polymerase to natural nucleic acids and / or non-natural nucleic acids without modified sugars. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids containing the modified base, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of the wild-type polymerase for natural nucleic acids and / or non-natural nucleic acids that do not contain the modified base. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids containing the triphotate, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of the wild-type polymerase for nucleic acids containing the triphotate and / or non-natural nucleic acids that do not contain the triphotate. For example, modified or wild-type polymerases may exhibit specificity for non-natural nucleic acids containing triphotates, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type polymerase for non-natural nucleic acids containing diphosphates or monophosphates, or phosphate-free nucleic acids, or combinations thereof.
[0131] In some embodiments, the modified or wild-type polymerase has relaxed specificity to non-natural nucleic acids. In some embodiments, the modified or wild-type polymerase has specificity to non-natural nucleic acids and specificity to natural nucleic acids, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of the wild-type polymerase to natural nucleic acids. In some embodiments, the modified or wild-type polymerase has specificity to non-natural nucleic acids containing modified sugars and specificity to natural nucleic acids, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of the wild-type polymerase to natural nucleic acids. In some embodiments, the modified or wild-type polymerase exhibits specificity for non-natural nucleic acids containing the modified base and for natural nucleic acids. It possesses specificity that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type polymerase to natural nucleic acids.
[0132] The absence of exonuclease activity can be a wild-type characteristic or a characteristic given by a variant or engineered polymerase. For example, the exo-minus Klenow fragment is a mutant form of the Klenow fragment lacking 3'-5' proofreading exonuclease activity.
[0133] The methods of the present invention are used to broaden the substrate range of any DNA polymerase that lacks or has rendered endogenous 3' to 5' exonuclease proofreading activity, for example, through mutation. Examples of DNA polymerases include polA, polB (see, e.g., Parrel & Loeb, Nature Struc Biol 2001), polC, polD, polY, polX, and reverse transcriptase (RT), but preferably progressive high-fidelity polymerases (PCT / GB2004 / 004643). In some embodiments, the modified or wild-type polymerase is substantially lacking in 3' to 5' proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase is substantially lacking in 3' to 5' proofreading exonuclease activity with respect to non-native nucleic acids. In some embodiments, the modified or wild-type polymerase has 3'-5' proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase has 3'-5' proofreading exonuclease activity with respect to native nucleic acids and substantially lacks 3'-5' proofreading exonuclease activity with respect to non-native nucleic acids.
[0134] In some embodiments, the modified polymerase has 3'-5' proofreading exonuclease activity that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the proofreading exonuclease activity of the wild-type polymerase. In some embodiments, the modified polymerase has 3'-5' proofreading exonuclease activity with respect to non-natural nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the proofreading exonuclease activity of the wild-type polymerase with respect to natural nucleic acids. In some embodiments, the modified polymerase has 3'-5' proofreading exonuclease activity for non-natural nucleic acids and 3'-5' proofreading exonuclease activity for natural nucleic acids, which is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the proofreading exonuclease activity of wild-type polymerase for natural nucleic acids.
[0135] In some embodiments, the polymerase is characterized by its dissociation rate from nucleic acids. In some embodiments, the polymerase has a relatively low dissociation rate with respect to one or more natural and non-natural nucleic acids. In some embodiments, the polymerase has a relatively high dissociation rate with respect to one or more natural and non-natural nucleic acids. The dissociation rate is the activity of the polymerase, which can be adjusted to control the reaction rate in the manner described herein.
[0136] In some embodiments, polymerase is used to extract specific natural and / or non-natural nucleic acids or When used with natural and / or non-natural nucleic acid collections, it is characterized by its fidelity. Fidelity generally refers to the accuracy with which polymerase incorporates the correct nucleic acid into the growing nucleic acid chain when making a copy of the nucleic acid template. DNA polymerase fidelity can be measured as the ratio of correct to incorrect incorporations of natural and non-natural nucleic acids, for example, when natural and non-natural nucleic acids are present at equal concentrations, because they compete for chain synthesis at the same site in the polymerase-chain-template nucleic acid binary complex. DNA polymerase fidelity is expressed in terms of (k) natural and non-natural nucleic acids. cat / K m ) and (kc) regarding inaccurate natural and non-natural nucleic acids at / K m It can be calculated as the ratio of ) to; in the formula, k cat and K m This is the Michaelis-Menten parameter in steady-state enzyme dynamics (incorporated by reference to Fersht, AR (1985) Enzyme Structure and Mechanism, 2nd edition, p.350, WH Freeman & Co., New (York.) In some embodiments, the polymerase is present or absent in proofreading activity at least about 100, 1,000, 10,000, 100,000, or 1 × 10 6 It has a fidelity value of [value].
[0137] In some embodiments, polymerases or variants from natural sources are screened using assays that detect the incorporation of non-natural nucleic acids having a specific structure. In one embodiment, polymerases can be screened for their ability to incorporation non-natural nucleic acids or UBPs; for example, d5SICSTP, dNaMTP, or d5SICSTP-dNaMTP UBPs. Polymerases, such as heterologous polymerases, that exhibit modified properties with respect to non-natural nucleic acids compared to wild-type polymerases can be used. For example, the modified properties may include, for example, K m , k cat , V maxThis can include polymerase processability in the presence of non-natural nucleic acids (or natural nucleotides), average template read length by polymerase in the presence of non-natural nucleic acids, polymerase specificity of non-natural nucleic acids, binding rate of non-natural nucleic acids, production and release rate of products (pyrophosphate, triphosphate, etc.), branching rate, or any combination thereof. In one embodiment, the modified property is a reduced K with respect to non-natural nucleic acids. m and / or increased k with respect to non-natural nucleic acids cat / K m or V max / K m Similarly, polymerases may, in some cases, exhibit increased binding rates of non-natural nucleic acids, increased product release rates, and / or decreased branching rates compared to wild-type polymerases.
[0138] Simultaneously, polymerases can incorporate native nucleic acids, such as A, C, G, and T, into the growing nucleic acid copy. For example, polymerases may exhibit specific activity with respect to native nucleic acids that is at least about 5% higher than that of the corresponding wild-type polymerase (e.g., 5%, 10%, 25%, 50%, 75%, 100%, or higher), and processability with native nucleic acids in the presence of a template that is at least about 5% higher than that of the wild-type polymerase (e.g., 5%, 10%, 25%, 50%, 75%, 100%, or higher). In some cases, polymerases may exhibit activity with respect to native nucleotides that is at least about 5% higher than that of the wild-type polymerase (e.g., about 5%, 10%, 25%, 50%, 75%, or 100%, or higher), cat / K m or V max / K m This indicates.
[0139] The polymerases used herein, which may have the ability to incorporate non-natural nucleic acids of a specific structure, can also be generated using directed evolution. Nucleic acid synthesis assays can be used to screen polymerase variants that have specificity with respect to any of a variety of non-natural nucleic acids. For example, polymerase variants can be screened for their ability to incorporate non-natural nucleic acids, or UBPs; e.g., d5SICSTP, dNaMTP, or d5SICSTP-dNaMTP UBPs, into nucleic acids. In some embodiments, such assays are, for example, in vitro assays using recombinant polymerase variants. In some embodiments, such assays are, for example For example, an in vivo assay expressing polymerase variants within cells. Such directed evolution methods can be used to screen for any suitable polymerase variant with respect to activity against any of the non-native nucleic acids described herein.
[0140] The modified polymerase of the described composition may optionally be a modified and / or recombinant Φ29-type DNA polymerase. The polymerase may optionally be a modified and / or recombinant Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, or L17 polymerase.
[0141] The modified polymerases of the compositions described may optionally be modified and / or recombinant prokaryotic DNA polymerases, such as DNA polymerase II (PolII), DNA polymerase III (PolIII), DNA polymerase IV (PolIV), and DNA polymerase V (PolV). In some embodiments, the modified polymerases include polymerases that mediate DNA synthesis on non-instructively damaged nucleotides. In some embodiments, genes encoding PolI, PolII (polB), PolIV (dinB), and / or PolV (umuCD) are constitutively expressed or overexpressed in the engineered cells or SSOs. In some embodiments, increased expression or overexpression of PolII contributes to an increased retention rate of non-native base pairs (UBPs) in the engineered cells or SSOs.
[0142] Nucleic acid polymerases generally useful in this invention include DNA polymerase, RNA polymerase, reverse transcriptase, and their variants or altered forms. DNA polymerase and its properties are described in detail, in particular, in *DNA Replication*, 2nd edition, Kornberg and Baker, WH. Freeman, New York, NY (1991). Known conventional DNA polymerases useful in this invention include, but are not limited to, Pyrococcus furiosus (Pfu) DNA polymerase (Lundberg et al., 1991, Gene, 108:1, Stratagene), Pyrococcus woesey (Pwo) DNA polymerase (Hinnisdaels et al., 1996, Biotechniques, 20:186~8, Boehringer Mannheim), Thermus thermophilus (Tth) DNA polymerase (Myers and Gelfand 1991, Biochemistry 30:7661), and Bacillus stearothermophilus DNA polymerase (Stenesh and McGowan, 1977, Biochim). Biophys Acta 475:32), Thermococcus litoralis (TIi) DNA polymerase (also called Vent® DNA polymerase; Cariello et al., 1991, Polynucleotides Res, 19:4193, New England Biolabs), 9°Nm® DNA polymerase (New England Biolabs), Stoffel fragment, Thermo Sequenase® (Amersham Pharmacia Biotech UK), Therminator® (New England Biolabs), Thermotoga maritima (Tma) DNA polymerase (Diaz and Sabino, 1998 Braz J Med.Res, 31:1239), Thermus aquaticus (Taq) DNA polymerase (Chien et al., 1976, J.Bacteoriol, 127:1550) DNA polymerase, Pyrococcus kodakaraensis KOD DNA polymerase (Takagi et al., 1997, Appl.Environ.Microbiol. 63:4504), JDF-3 DNA polymerase (from Thermococcus species, JDF-3, parent application WO 0132887), Pyrococcus GB-D (PGB-D) DNA polymerase (Deep Ve Also known as nt(trademark) DNA polymerase (Juncosa-Ginesta et al., 1994, Biotechniques, 16:820, New England Biolabs), UlTma DNA polymerase (from the thermophilic bacterium Thermotoga maritima; Diaz and Sabino, 1998 Braz J. Med. Res, 31:1239; PE Applied Biosystems), Tgo DNA polymerase (from Thermococcus gorgonarius, Roche Molecular Biochemicals), Escherichia coli DNA polymerase I (Lecomte and Doubleday, 1983, Polynucleotides Res. 11:7505), T7 DNA polymerase (Nordstrom et al., 1981, J Biol. Chem. 256:3112), and archaeon DP1I / DP2 This includes DNA polymerase II (Cann et al., 1998, Proc. Natl. Acad. Sci. USA 95:14250). Both mesothermal and thermophilic polymerases are intended. Thermophilic DNA polymerases include, but are not limited to, ThermoSequenase®, 9°Nm®, Therminator®, Taq, Tne, Tma, Pfu, TfI, Tth, TIi, Stoffel fragment, Vent® and Deep Vent® DNA polymerases, KOD DNA polymerase, Tgo, JDF-3, and their variants, variants, and derivatives. Polymerases that are 3' exonuclease-deficient variants are also intended. Reverse transcriptases useful in this invention include, but are not limited to, reverse transcriptases from HIV, HTLV-I, HTLV-II, FeLV, FIV, SIV, AMV, MMTV, MoMuLV, and other retroviruses (see Levin, Cell 88:5~8 (1997); Verma, Biochim Biophys Acta. 473:1~38 (1977); Wu et al., CRC Crit Rev Biochem. 3:289~347 (1975)).Other examples of polymerases include, but are not limited to, 9°N DNA polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, RB69 DNA polymerase, KOD DNA polymerase, and Vent® DNA polymerase. Gardner et al. (2004) "Comparative Kinetics." This includes Nucleotide Analog Incorporation by Vent DNA Polymerase (J. Biol. Chem., 279(12), 11834-11842; Gardner and Jack, "Determinants of nucleotide sugar recognition in an archaeon DNA polymerase," Nucleic Acids Research, 27(12), 2545-2553). Polymerases isolated from non-thermophilic organisms can be made thermoinactivatable. An example of this is DNA polymerase from phages. It will be understood that polymerases from any of the various sources can be modified to increase or decrease their tolerance to high-temperature conditions. In some embodiments, polymerases can be made thermophilic. In some embodiments, thermophilic polymerases can be made thermoinactivatable. Thermophilic polymerases are typically useful when used under high-temperature conditions or thermal cycling conditions, such as polymerase chain reaction (PCR) techniques.
[0143] In some embodiments, the polymerase is Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, ThermoSequenase®, 9°Nm®, Therminator® DNA polymerase, Tne, Tma, TfI, Tth, TIi, Stoffel fragment, Vent® and Deep Vent® DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, Escherichia coli DNA polymerase I, and E. coli. This includes bacterial DNA polymerase III, archaeal DP1I / DP2 DNA polymerase II, 9°N DNA polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney's mouse leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, and SuperScript® III reverse transcriptase.
[0144] In some embodiments, the polymerase is DNA polymerase 1-Klenow fragment, Vent polymerase, Phusion® DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator® DNA polymerase, POLB polymerase, SP6 RNA polymerase, Escherichia coli DNA polymerase I, Escherichia coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney's mouse leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, or SuperScript® III reverse transcriptase.
[0145] Furthermore, such polymerases can be used for DNA amplification and / or sequencing applications, including real-time uses, in the context of amplification or sequencing, for example, involving the incorporation of non-natural nucleic acid residues into DNA by polymerase. In other embodiments, the non-natural nucleic acid to be incorporated can be identical to a natural residue, for example, in which case the label or other portion of the non-natural nucleic acid is removed by the action of the polymerase during incorporation, or the non-natural nucleic acid may have one or more features that distinguish it from the natural nucleic acid.
[0146] Nucleotide transporter Nucleotide transporters (NTs) are a group of membrane transport proteins that facilitate the transport of nucleoside substrates on cell membranes and vesicles. In some embodiments, there are two types of nucleoside transporters: concentrated nucleoside transporters and equilibrium nucleoside transporters. In some cases, NTs also encompass organic anion transporters (OATs) and organic cation transporters (OCTs). In some cases, the nucleotide transporter is a nucleoside triphosphate transporter.
[0147] In some embodiments, the nucleotide triphosphate transporter (NTT) is derived from bacteria, plants, or algae. In some embodiments, the nucleotide nucleoside triphosphate transporter is TpNTT1, TpNTT2, TpNTT3, TpNTT4, TpNTT5, TpNTT6, TpNTT7, TpNTT8 (T. pseudonana), PtNTT1, PtNTT2, PtNTT3, PtNTT4, PtNTT5, PtNTT6 (P. tricornutum), GsNTT (Galdieria sulphuraria), AtNTT1, AtNTT2 (Arabidopsis saliana) These are thaliana), CtNTT1, CtNTT2 (Chlamydia trachomatis), PamNTT1, PamNTT2 (Protochlamydia amoebophila), CcNTT (Caedibacter caryophilus), and RpNTT1 (Rickettsia prowazekii).
[0148] In some embodiments, NTT is CNT1, CNT2, CNT3, ENT1, ENT2, OAT1, OAT3, or OCT1.
[0149] In some embodiments, NTTs are used to transfer non-natural nucleic acids into organisms, such as cells. In some embodiments, NTTs can be modified so that their nucleotide binding sites are altered to reduce inhibition of steric entry of non-natural nucleic acids into nucleotide binding sites. In some embodiments, NTTs can be modified to increase their interaction with one or more features of non-natural nucleic acids. Such NTTs can be expressed or manipulated in cells to stably transfer UBPs into cells. Accordingly, the present invention includes compositions comprising heterologous or recombinant NTTs and methods for using them.
[0150] NTTs can be modified using methods related to protein engineering. For example, molecular modeling can be performed based on crystal structure to identify locations in NTTs where mutations can be made to modify the target activity or binding site. Residues identified as targets for substitution can be replaced with selected residues using energy minimization modeling, homology modeling, and / or conserved amino acid substitutions, as described in Bordo et al., J Mol Biol 217:721~729 (1991) and Hayes et al., Proc Natl Acad Sci, USA 99:15926~15931 (2002).
[0151] Any of the various NTTs can be used in the methods or compositions described herein, for example, those comprising enzymes based on proteins isolated from biological systems and their functional variants. References to specific NTTs, such as those embodied below, will be understood to include their functional variants unless otherwise indicated. In some embodiments, NTT is wild-type NTT. In some embodiments, NTT is modified or mutant NTT.
[0152] NTTs can also be used that have features to improve the entry of non-natural nucleic acids into cells and to coordinate with non-natural nucleotides at the nucleotide-binding region. In some embodiments, modified NTTs have a modified nucleotide-binding site. In some embodiments, modified or wild-type NTTs have relaxation specificity for non-natural nucleic acids.
[0153] In some embodiments, modified NTT has specificity to non-natural nucleic acids that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT to non-natural nucleic acids. In some embodiments, modified or wild-type NTT has specificity to non-natural nucleic acids containing modified sugars that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT to natural nucleic acids and / or non-natural nucleic acids without modified sugars. In some embodiments, the modified or wild-type NTT has specificity for non-natural nucleic acids containing modified bases, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT for natural nucleic acids and / or non-natural nucleic acids that do not contain modified bases. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids containing triphotates, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT for triphotate-containing nucleic acids and / or non-natural nucleic acids that do not contain triphotates. For example, modified or wild-type NTT may have specificity for non-natural nucleic acids containing triphot phosphates, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT for non-natural nucleic acids containing diphosphates or monophosphates, or not containing phosphates, or a combination thereof.
[0154] In some embodiments, the modified or wild-type NTT has specificity to non-natural nucleic acids and specificity to natural nucleic acids, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT to natural nucleic acids. In some embodiments, the modified or wild-type NTT has specificity to non-natural nucleic acids and specificity to natural nucleic acids, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT to natural nucleic acids. In some embodiments, the modified or wild-type NTT has specificity to non-natural nucleic acids and specificity to natural nucleic acids, including the modified base, which is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, and 99.99% of the specificity of wild-type NTT to natural nucleic acids.
[0155] NTTs can be characterized by their dissociation rates from nucleic acids. In some embodiments, NTTs have relatively low dissociation rates with respect to one or more natural and non-natural nucleic acids. In some embodiments, NTTs have relatively high dissociation rates with respect to one or more natural and non-natural nucleic acids. The dissociation rate is the activity of the NTT, which can be adjusted to control the reaction rate in the methods described herein.
[0156] NTT or its variants from natural sources can be screened using assays that detect the importation of non-natural nucleic acids having a specific structure. In one embodiment, NTT can be screened for its ability to import non-natural nucleic acids or UBPs, such as d5SICSTP, dNaMTP, or d5SICSTP-dNaMTP UBPs. NTT, for example heterologous NTT, can be used that exhibits modified properties with respect to non-natural nucleic acids compared to wild-type NTT. For example, the modified properties may include, for example, K m , k cat , Vmax This can be NTT transfer in the presence of non-natural nucleic acids (or natural nucleotides), average template read length by cells with NTT in the presence of non-natural nucleic acids, specificity of NTT to non-natural nucleic acids, binding rate of non-natural nucleic acids, or product release rate, or any combination thereof. In one embodiment, the modified property is a reduced K for non-natural nucleic acids. m and / or increased k for non-natural nucleic acids cat / K m Or V max / K m Similarly, NTT, in some cases, exhibits increased binding rates, increased product release rates, and / or increased cell transfer rates of non-natural nucleic acids compared to wild-type NTT.
[0157] Simultaneously, NTT can transfer native nucleic acids, such as A, C, G, and T, into cells. For example, NTT may exhibit specific transfer activity for native nucleic acids that is at least about 5% higher (e.g., 5%, 10%, 25%, 50%, 75%, 100%, or higher) than that of the corresponding wild-type NTT. In some cases, NTT may exhibit specific transfer activity for native nucleotides that is at least about 5% higher (e.g., 5%, 10%, 25%, 50%, 75%, 100%, or higher) than that of wild-type NTT. cat / K m or V max / K m This indicates.
[0158] NTTs used herein, which can have the ability to transfer non-natural nucleic acids of a specific structure, can also be generated using directed evolution methods. Nucleic acid synthesis assays can be used to screen NTT variants that are specific to any of a variety of non-natural nucleic acids. For example, NTT variants can be screened for their ability to transfer non-natural nucleic acids or UBPs, such as d5SICSTP, dNaMTP, or d5SICSTP-dNaMTP UBPs, into nucleic acids. In some embodiments, such assays can be used, for example, in vitro assays using recombinant NTT variants. (i) In some embodiments, such assays are, for example, in vivo assays in which an NTT variant is expressed in cells. Such directed evolution methods can be used to screen for any suitable NTT variant with respect to activity against any of the non-native nucleic acids described herein.
[0159] Nucleic acid reagents and tools The nucleic acid reagents used in the methods, cells, or engineered microorganisms described herein include one or more ORFs. The ORFs are from any suitable source and are sometimes from nucleic acid libraries containing genomic DNA, mRNA, reverse transcriptase RNA, or complementary DNA (cDNA), or one or more of the aforementioned, from any species containing the nucleic acid sequence of interest, the protein of interest, or the activity of interest. Non-limiting examples of organisms from which ORFs can be obtained include, for example, bacteria, yeast, fungi, humans, insects, nematodes, cattle, horses, dogs, cats, rats, or mice. In some embodiments, the nucleic acid reagents or other reagents described herein are isolated or purified.
[0160] Nucleic acid reagents sometimes contain nucleotide sequences adjacent to the ORF that are translated in conjunction with the ORF and encode an amino acid tag. The tag-coding nucleotide sequence is positioned at the 3' and / or 5' of the ORF in the nucleic acid reagent, thereby encoding the tag at the C-terminus or N-terminus of the protein or peptide encoded by the ORF. Any tag that does not inhibit in vitro transcription and / or translation is utilized and appropriately selected by those skilled in the art. The tag can facilitate the isolation and / or purification of the desired ORF product from the culture or fermentation medium.
[0161] A nucleic acid or nucleic acid reagent may contain certain elements, such as regulatory elements, which are often selected depending on the intended use of the nucleic acid. Any of the following elements may be included in or excluded from a nucleic acid reagent. A nucleic acid reagent may, for example, contain one or more of the following nucleotide elements: one or more promoter elements, one or more 5' untranslated regions (5'UTR), one or more regions into which the target nucleotide sequence is inserted ("insertion element"), one or more target nucleotide sequences, one or more 3' untranslated regions (3'UTR), and one or more selection elements. A nucleic acid reagent may provide one or more such elements, with other elements being inserted into the nucleic acid before it is introduced into the desired organism. In some embodiments, the provided nucleic acid reagent includes a promoter, a 5'UTR, optionally a 3'UTR, and an insertion element (possibly multiple) into which the target nucleotide sequence is inserted (i.e., cloned) into the nucleic acid reagent. In certain embodiments, the nucleic acid reagent provided comprises a promoter, an insertion element(s), and optionally a 3'UTR, with the 5'UTR / target nucleotide sequence inserted along with the optionally 3'UTR. The elements can be arranged in any order suitable for expression in a selected expression system (e.g., expression in a selected organism, or expression in a cell-free system), and in some embodiments, the nucleic acid reagent comprises the following elements in the 5'-to-3' direction: (1) promoter element, 5'UTR, and insertion element(s); (2) promoter element, 5'UTR, and target nucleotide sequence; (3) promoter element, 5'UTR, insertion element(s), and 3'UTR; and (4) promoter element, 5'UTR, target nucleotide sequence, and 3'UTR.
[0162] Nucleic acid reagents, such as expression cassettes and / or expression vectors, can contain a variety of regulatory elements, including promoters, enhancers, translation initiation sequences, transcription termination sequences, and other elements. The "promoter" is generally relatively fixed with respect to the transcription initiation site. A promoter is one or more sequences of DNA that function when located in a specific place. For example, a promoter can be upstream of a nucleotide acid phosphate transporter nucleic acid segment. A promoter contains core elements required for the basic interaction between RNA polymerase and transcription factors, and may contain upstream and response elements. An enhancer generally refers to a sequence of DNA that functions at an unfixed distance from the transcription start site and can be either 5' or 3" from the transcription unit. Furthermore, enhancers can be within introns and even within the coding sequence itself. They are typically between 10 and 300 in length and function in cis. Enhancers function to increase transcription from the vicinity of the promoter. Like promoters, enhancers often also contain response elements that mediate the regulation of transcription. Enhancers often determine the regulation of expression.
[0163] As described above, nucleic acid reagents may contain one or more 5'UTRs and one or more 3'UTRs. For example, expression vectors used in eukaryotic host cells (e.g., yeast, fungi, insects, plants, animals, humans, or nucleated cells) and prokaryotic host cells (e.g., viruses, bacteria) may contain sequences that signal transcription termination, which may affect mRNA expression. These reagents can be transcribed as polyadenylated segments in the untranslated portion of mRNA-coding tissue factor proteins. The 3" untranslated region also includes the transcription termination site. In some preferred embodiments, the transcription unit includes a polyadenylated region. One benefit of this region is that the transcribed unit is more likely to be processed and transported like mRNA. The identification and use of polyadenylated signals in expression constructs are well established. In some preferred embodiments, homogeneous polyadenylated signals can be used in transgene constructs.
[0164] A 5'UTR may contain one or more elements intrinsically related to the nucleotide sequence from which it originates, and possibly one or more extrinsic elements. A 5'UTR may originate from any suitable nucleic acid, e.g., genomic DNA, plasmid DNA, RNA, or mRNA, or from any suitable organism, e.g., a virus, bacterium, yeast, fungus, plant, insect, or mammal. Those skilled in the art can select elements appropriate for a 5'UTR based on a chosen expression system (e.g., expression in a chosen organism, or expression in a cell-free system). A 5'UTR may sometimes contain one or more of the following elements known to those skilled in the art: enhancer sequences (e.g., transcription or translation), transcription start sites, transcription factor binding sites, translation regulatory sites, translation start sites, translation factor binding sites, accessory protein binding sites, feedback regulator binding sites, Pribnow boxes, TATA boxes, -35 elements, E-boxes (helix-loop-helix binding elements), ribosome binding sites, replicons, internal ribosome entry sites (IRESs), silencer elements, etc. In some embodiments, the promoter element is isolated such that all 5'UTR elements necessary for proper conditioning are contained within the promoter element fragment or within a functional subarrangement of the promoter element fragment.
[0165] The 5'UTR in nucleic acid sequence reagents may contain translational enhancer nucleotide sequences. Translational enhancer nucleotide sequences are often positioned between the promoter and the target nucleotide sequence in the nucleic acid reagent. Translational enhancer sequences often bind to ribosomes and may be 18S rRNA-binding ribonucleotide sequences (i.e., 40S ribosome-binding sequences) or internal ribosome entry sequences (IRESs). IRESs generally form an RNA scaffold with a precisely arranged RNA tertiary structure that contacts the 40S ribosome subunit via several specific intermolecular interactions. Examples of ribosome enhancer sequences are publicly known and can be identified by those skilled in the art (e.g., Mignone et al., Nucleic Acids Research 33:D141-D146). (2005); Paulous et al., Nucleic Acids Research 31:722 - 733 (2003); Akbergenov et al., Nucleic Acids Research 32:239 - 247 (2004); Mignone et al., Genome Biology 3(3): reviews0004.1 - 0001.10 (2002); Gallie, Nucleic Acids Research 30:3401 - 3411 (2002); Shaloiko et al., DOI:10.1002 / bit.20267; and Gallie et al., Nucleic Acids Research 15:3257 - 3273 (1987)).
[0166] Translation enhancer sequences are sometimes eukaryotic sequences such as the Kozak consensus sequence or other sequences (e.g., the hydropholy polyp sequence, GenBank accession number U07128). Translation enhancer sequences are sometimes prokaryotic sequences such as the Shine - Dalgano consensus sequence. In certain embodiments, the translation enhancer sequence is a viral nucleotide sequence. Translation enhancer sequences are sometimes from the 5’UTR of plant viruses such as, for example, Tobacco Mosaic Virus (TMV), Alfalfa Mosaic Virus (AMV); Tobacco Etch Virus (ETV); Potato Virus Y (PVY); Turnip Mosaic (poty) virus, and Pea Seed Borne Mosaic Virus). In certain embodiments, the approximately 67 - base omega sequence from TMV is included in the nucleic acid reagent as a translation enhancer sequence (e.g., lacking guanosine nucleotides and containing a 25 - nucleotide - long poly(CAA) central region).
[0167] The 3’UTR can contain one or more elements endogenous to the nucleotide sequence arising therefrom and, optionally, one or more exogenous elements. The 3’UTR is derived from any suitable nucleic acid such as genomic DNA, plasmid DNA, RNA, or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect, or mammal). One of ordinary skill in the art can select suitable elements for the 3’UTR based on the selected expression system (e.g., expression in the selected organism). The 3’UTR can optionally contain one or more of the following elements known to those of ordinary skill in the art: transcriptional regulatory sites, transcriptional start sites, transcriptional termination sites, transcription factor binding sites, translational regulatory sites, translational termination sites, translational start sites, translation factor binding sites, ribosome binding sites, replicons, enhancer elements, silencer elements, and polyadenosine tails. The 3’UTR often contains a polyadenosine tail, optionally does not contain a polyadenosine tail, and if a polyadenosine tail is present, one or more adenosine moieties are added to or lost from it (e.g., about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 adenosine moieties are added or removed).
[0168] In some embodiments, modifications to the 5'UTR and / or 3'UTR are used to alter promoter activity (e.g., increase, add, decrease, or substantially eliminate it). The alteration of promoter activity can then alter the activity (e.g., enzymatic activity) of a peptide, polypeptide, or protein by altering the transcription of the desired nucleotide sequence(s) from an operablely linked promoter element containing the modified 5' or 3'UTR. For example, a microorganism can be manipulated by genetic modification to express a nucleic acid reagent containing the modified 5' or 3'UTR, which in certain embodiments can add novel activity (e.g., activity not normally seen in the host organism) or increase the expression of existing activity by increasing transcription from an operable homogeneous or heterogeneous promoter linked to the desired nucleotide sequence(s) (e.g., the desired homogeneous or heterogeneous nucleotide sequence). Some embodiments Morphologically, microorganisms can be manipulated by genetic modification to express nucleic acid reagents containing modified 5' or 3' UTRs, which in certain embodiments can reduce the expression of activity by decreasing or substantially eliminating transcription from homogeneous or heterogeneous promoters operably linked to the nucleotide sequence of interest.
[0169] The expression of nucleotide triphosphate transporters from an expression cassette or expression vector can be controlled by any promoter capable of expression in prokaryotic or eukaryotic cells. Promoter elements are typically required for DNA and / or RNA synthesis. Promoter elements often contain a region of DNA that can facilitate the transcription of a particular gene by providing an initiation site for the synthesis of RNA corresponding to the gene. Promoters are generally located near the gene they regulate, upstream of the gene (e.g., 5' of the gene), and in some embodiments, on the same DNA strand as the gene's sense strand. In some embodiments, promoter elements can be isolated from a gene or organism and inserted into a functional conjugate with a polynucleotide sequence, resulting in altered and / or regulated expression. Non-native promoters used for nucleic acid expression (e.g., promoters not typically associated with a given nucleic acid sequence) are often called heterologous promoters. In certain embodiments, a heterologous promoter and / or 5'UTR can be inserted into a functional conjugate with a polynucleotide encoding a polypeptide having the desired activity described herein. As used herein with respect to promoters, the terms “operably linked” and “functionally connected to” refer to the relationship between a coding sequence and a promoter element. A promoter is operably linked or functionally connected to a coding sequence if the expression from the coding sequence via transcription is regulated or controlled by the promoter element. The terms “operably linked” and “functionally connected to” are used synonymously herein with respect to promoter elements.
[0170] Promoters often interact with RNA polymerases. Polymerases are enzymes that catalyze the synthesis of nucleic acids using existing nucleic acid reagents. When the template is a DNA template, the RNA molecule is transcribed before protein synthesis. Enzymes having polymerase activity suitable for use in the methods of the present invention include any polymerase that is active in a selected system having a template selected for protein synthesis. In some embodiments, a promoter (e.g., a heterologous promoter), also referred herein as a promoter element, can be operably ligated to a nucleotide sequence or an open reading frame (ORF). Transcription from the promoter element can catalyze the synthesis of RNA corresponding to the nucleotide sequence or ORF sequence operably ligated to the promoter, thus resulting in the synthesis of a desired peptide, polypeptide, or protein.
[0171] Promoter elements may, in some cases, be responsive to regulatory control. Promoter elements may also, in some cases, be modulated by selective agents. That is, transcription from a promoter element may, in some cases, be initiated, terminated, upregulated, or downregulated in response to changes in the environment, nutrients, or internal conditions, or signals (e.g., thermoinducible promoters, photoregulated promoters, feedback-regulated promoters, hormone-stimulated promoters, tissue-specific promoters, oxygen and pH-stimulated promoters, promoters responsive to selective agents (e.g., kanamycin)). Promoter elements that are frequently affected by the environment, nutrients, or internal signals may be affected (directly or indirectly) by signals that bind at or near the promoter and increase or decrease the expression of target sequences under certain conditions.
[0172] Non-limiting examples of select or modulodentifiers that affect transcription from promoter elements used in embodiments described herein include, but are not limited to: (1) nucleic acid segments encoding products that provide resistance to otherwise toxic compounds (e.g., antibiotics); (2) nucleic acid segments encoding products that are otherwise lacking in receptor cells (e.g., essential products, tRNA genes, acute nutritional markers); (3) nucleic acid segments encoding products that suppress the activity of gene products; (4) nucleic acid segments encoding easily identifiable products (e.g., phenotypic markers, e.g., antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind products otherwise harmful to cell survival and / or function; (6) nucleic acid segments that otherwise inhibit the activity of any of the nucleic acid segments described in No. 1 to 5 above. (1) Nucleic acid segments that bind substrate-modifying products (e.g., antisense oligonucleotides); (2) Nucleic acid segments that can be used to isolate or identify a desired molecule (e.g., specific protein-binding sites); (3) Nucleic acid segments that encode specific nucleotide sequences that can otherwise be non-functional (e.g., for PCR amplification of subpopulations of molecules); (4) Nucleic acid segments that, in the absence of such sequences, directly or indirectly confer tolerance or sensitivity to a particular compound; (5) Nucleic acid segments that encode products that convert toxic or relatively non-toxic compounds in receptor cells into toxic compounds (e.g., herpesthymidine kinase, cytosine deaminase); (6) Nucleic acid segments that inhibit the replication, distribution, or heritability of nucleic acid molecules containing them; and / or (7) Nucleic acid segments that encode conditioned replication function, e.g., replication in a particular host or host cell line or under certain environmental conditions (e.g., temperature and nutritional status).In some embodiments, modifiers or selective agents may be added to alter the existing growth conditions to which the organism is supplied (e.g., growth in a liquid culture, growth in a fermenter, growth on a solid nutrient plate, etc.).
[0173] In some embodiments, the modification of promoter elements can be used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity (e.g., enzyme activity) of a peptide, polypeptide, or protein. For example, a microorganism can be genetically modified to express a nucleic acid reagent that can add novel activity (e.g., activity not normally seen in the host organism) or increase the expression of existing activity by increasing transcription from a homogeneous or heterogeneous promoter (e.g., the homogeneous or heterogeneous nucleotide sequence of interest) operably linked to the nucleotide sequence of interest in some particular embodiments. In some embodiments, a microorganism can be genetically modified to express a nucleic acid reagent that can reduce the expression of activity by decreasing or substantially eliminating transcription from a homogeneous or heterogeneous promoter operably linked to the nucleotide sequence of interest in some particular embodiments.
[0174] Nucleic acids encoding heterologous proteins, such as nucleotide triphosphate transporters, can be inserted into or used in conjunction with any suitable expression system. In some embodiments, the nucleic acid reagent may be stably integrated into the chromosome of a host organism, or in certain embodiments, the nucleic acid reagent may be a deletion of a portion of the host chromosome (e.g., a genetically modified organism in which the alteration of the host genome gives the organism the ability to selectively or preferentially maintain the desired organism that retains the genetic modification). Such nucleic acid reagents (e.g., nucleic acids or genetically modified organisms in which the altered genome gives the organism a selectable characteristic) can be selected to match their ability to lead to the production of the desired protein or nucleic acid molecule. If desired, the nucleic acid reagent may use tRNA with codons that are (i) different from those specified in the natural sequence, to produce the same amino acids, or (ii) unconventional or This can be modified to encode amino acids different from the usual ones, including non-natural amino acids (including detectably labeled amino acids).
[0175] Recombinant expression is effectively realized using an expression cassette, which can be a portion of a vector, such as a plasmid. The vector may include a promoter operably ligated to a nucleic acid encoding a nucleotide triphosphate transporter. The vector may also include other elements required for transcription and translation as described herein. The expression cassette, expression vector, and sequences within the cassette or vector may be heterogeneous to the cell into which the non-native nucleotides come into contact. For example, a nucleotide triphosphate transporter sequence may be heterogeneous to a cell.
[0176] A variety of prokaryotic and eukaryotic expression vectors suitable for holding, encoding, and / or expressing nucleotide triphosphate transporters can be generated. Such expression vectors include, for example, pET, pET3d, pCR2.1, pBAD, pUC, and yeast vectors. The vectors can be used, for example, in a variety of in vivo and in vitro situations. Non-limiting examples of prokaryotic promoters that can be used include SP6, T7, T5, tac, bla, trp, gal, lac, or maltose promoters. Non-limiting examples of eukaryotic promoters that can be used include constitutive promoters, e.g., viral promoters, e.g., CMV, SV40, and RSV promoters, as well as moduloable promoters, e.g., inducible or repressive promoters, e.g., tet promoter, hsp70 promoter, and synthetic promoters regulated by CRE. A vector for bacterial expression is pGEX-5X-3, and for eukaryotic expression is pCIneo-CMV. Viral vectors that can be used include those relating to lentiviruses, adenoviruses, adeno-associated viruses, herpesviruses, vaccinia viruses, polioviruses, AIDS viruses, neurotrophic viruses, Sindbis and other viruses. Any viral family that shares the properties of these viruses is also useful to make suitable for use as a vector. Retroviral vectors that can be used include those described in Verma, American Society for Microbiology, pp. 229-232, Washington (1985). For example, such retroviral vectors may include mouse Maloney's leukemia virus, MMLV, and other retroviruses that express the desired properties. Typically, viral vectors contain unstructured early genes, structural late genes, RNA polymerase III transcripts, inverted end repeats necessary for replication and capsid formation, and promoters that control the transcription and replication of the viral genome.When manipulated as a vector, the virus typically has one or more removed initial genes, and the gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral nucleic acid.
[0177] cloning Any convenient cloning strategy known in the art is used to incorporate elements such as ORFs into nucleic acid reagents. Known methods can be used to insert elements into a template independently of the insertion element, such as (1) cleaving the template at one or more existing restriction enzyme sites to ligate the element of interest, and (2) adding the restriction enzyme site to the template by hybridizing with one or more oligonucleotide primers containing one or more suitable restriction enzyme sites, and amplifying by polymerase chain reaction (as described in more detail herein). Other cloning strategies utilize one or more insertion sites present in or inserted into nucleic acid reagents, such as oligonucleotide primer hybridization sites for PCR and others described herein. In some embodiments, cloning strategies can be combined with genetic manipulation such as recombination (e.g., modified genetic material as further described herein). (Recombination of nucleic acid reagents with the target nucleic acid sequence into the genome of an object). In some embodiments, cloned ORFs (may be multiple) can be used to generate modified or wild-type nucleotide triphosphate transporters and / or polymerases (directly or indirectly) by manipulating microorganisms containing altered nucleotide triphosphate transporter activity or polymerase activity in one or more ORFs of interest.
[0178] Nucleic acids are specifically cleaved by contacting them with one or more specific cleavage agents. These specific cleavage agents often cleave the nucleic acid specifically at a particular site and along a specific nucleotide sequence. Examples of enzyme-specific cleavage agents include, but are not limited to, endonucleases (e.g., DNase (e.g., DNase I, II); RNase (e.g., RNase E, F, H, P); Cleavase® enzyme; Taq DNA polymerase; Escherichia coli DNA polymerase I; and eukaryotic structure-specific endonucleases; mouse FEN-1 endonuclease; type I, II, or III restriction endonucleases, e.g., Acc I, Afl III, Alu I, Alw44 I, Apa I, Asn I, Ava I, Ava II, BamH I, Ban II, Bcl I, Bgl I, Bgl II, Bln I, BsaI, Bsm I, BsmBI, BssH II, BstE II, Cfo I, CIa I, Dde I, Dpn I, Dra I, EcIX I, EcoR I, EcoR I, EcoR II, EcoR V, Hae II, Hae II, Hind II, Hind III, Hpa I, Hpa II, Kpn I, Ksp I, Mlu I, MIuN I, Msp I, Nci I, Nco I, Nde I, Nde II, Nhe I, Not I, Nru I, Nsi I, Pst I, Pvu I, Pvu II, Rsa I, Sac I, Sal I, Sau3A I, Sca I, ScrF I, Sfi I, Sma I, Spe I, Sph I, Ssp I, Stu I, Sty I, Swa I, Taq I, Xba I, Xho This includes (I) glycosylases (e.g., uracil-DNA glycosylase (UDG), 3-methyladenine DNA glycosylase, 3-methyladenine DNA glycosylase II, pyrimidine hydrate-DNA glycosylase, FaPy-DNA glycosylase, thymine mismatch-DNA glycosylase, hypoxanthine-DNA glycosylase, 5-hydroxymethyluracil-DNA glycosylase (HmUDG), 5-hydroxymethylcytosine-DNA glycosylase, or 1,N6-etheno-adenine-DNA glycosylase); exonucleases (e.g., exonuclease III); ribozymes, and DNAzymes. Sample nucleic acids are treated with chemical agents or synthesized using modified nucleotides, and the modified nucleic acids are cleaved. In non-limiting examples, sample nucleic acids are treated with (i) alkylating agents such as methylnitrosourea that generate several alkylating bases, including N3-methyladenine and N3-methylguanine, which are recognized and cleaved by alkylpurine DNA glycosylase; (ii) sodium bisulfate, which causes deamination of cytosine residues of DNA to form uracil residues that can be cleaved by uracil N-glycosylase; and (iii) chemical agents that convert guanine into its oxidized form, 8-hydroxyguanine, which can be cleaved by formamidepyrimidine DNA N-glycosylase. Examples of chemical cleavage processes include, but are not limited to, alkylation (e.g., alkylation of phosphorothioate-modified nucleic acids); cleavage of acid-instability nucleic acids containing P3'-N5'-phosphoramidates; and treatment of nucleic acids with osmium tetroxide and piperidine.
[0179] In some embodiments, nucleic acid reagents include one or more recombinase insertion sites. Recombinase insertion sites are recognition sequences on nucleic acid molecules that are involved in the integration / recombination reaction by recombinant proteins. For example, the recombination site for Cre recombinase is loxP, which is a 34-base pair sequence consisting of two 13-base pair inverted repeats (acting as recombinase binding sites) flanking an 8-base pair core sequence (e.g., Sauer, Curr. Opin. Biotech. 5:521~527 (1994)). Other examples include the attB, attP, attL, and attR sequences, as well as their variants, fragments, variants, and derivatives, which are recognized by recombinant protein λ Int and by co-protein integration host factors (IHF), FIS, and excisionase (Xis) (e.g., U.S. Patent Nos. 5,888,732; 6,143,557; 6,171,861; 6,270,969; 6,277,608; and 6,720,140; U.S. Patent Application Nos. 09 / 517,466 and 09 / 732,914; U.S. Patent Publication No. 2002 / 0007051; and Landy, Curr. Opin. Biotech. 3:699~707 (1993)).
[0180] An example of recombinase cloning nucleic acids is found in the Gateway® system (Invitrogen, California), which includes at least one recombination site for cloning a desired nucleic acid molecule in vivo or in vitro. In some embodiments, the system utilizes a vector containing at least two different site-specific recombination sites, often based on bacteriophage lambda systems (e.g., att1 and att2), and mutated from the wild-type (att0) site. Each mutant site has unique specificity with respect to its congenerate partner att site of the same type (e.g., attB1 for attP1, attL1 for attR1) (i.e., its binding partner recombination site), and does not cross-react with other mutant recombination sites or the wild-type att0 site. Different site specificities enable targeted cloning or ligation of the desired molecule, thus providing the desired orientation of the cloned molecule. Nucleic acid fragments flanked by recombination sites are cloned and subcloned using the Gateway® system by replacing a selectable marker (e.g., ccdB) flanked by att sites on a receptor plasmid molecule, sometimes called a destination vector. The desired clone is then selected by phenotypic transformation of the ccdB-sensitive host strain and positive selection of the marker on the receptor molecule. Similar strategies for negative selection (e.g., the use of toxic genes) can be used in other organisms, such as thymidine kinases (TKs) in mammals and insects.
[0181] Nucleic acid reagents sometimes contain one or more origin of replication (ORI) elements. In some embodiments, the template contains two or more ORIs, one that functions efficiently in one organism (e.g., bacteria) and another that functions efficiently in another organism (e.g., eukaryotes, e.g., yeast). In some embodiments, an ORI can function efficiently in one species (e.g., S. cerevisiae) and another ORI can function efficiently in a different species (e.g., S. pombe). Nucleic acid reagents sometimes also contain one or more transcriptional regulatory sites.
[0182] Nucleic acid reagents, such as expression cassettes or vectors, may contain nucleic acid sequences encoding marker products. The marker products are used to determine whether a gene is delivered to a cell and, if delivered, whether it is expressed. The marker genes in the examples include the E. coli lacZ gene encoding β-galactosidase and green fluorescent protein. In some embodiments, the markers can be selectable markers. If such selectable markers successfully migrate to host cells, the transmuted host cells can survive under selective pressure. There are two widely used, entirely distinct categories of selective regimes. The first category is based on the use of mutant cell lines lacking cellular metabolism and the ability to grow independently of supplemented medium. The second category is dominant selection, which refers to selection schemes used on any cell type and does not require the use of mutant cell lines. These schemes typically use drugs that halt the growth of host cells. Those cells, which would have novel genes, may express proteins that transmit drug resistance and thus survive the selection. An example of such dominant selection is drug-mediated selection. Omycin (Southern et al., J. Molec. Appl. Genet.) Use 1:327 (1982), mycophenicolic acid (Mulligan et al., Science 209:1422 (1980)), or hygromycin (Sugden et al., Mol. Cell. Biol. 5:410~413 (1985)).
[0183] A nucleic acid reagent may contain one or more selection elements (for example, elements for selecting the presence of the nucleic acid reagent, and not for activating a promoter element that can be selectively regulated). The selection elements often utilize known processes for determining whether or not the nucleic acid reagent is present in a cell. In some embodiments, the nucleic acid reagent may contain two or more selection elements, in which case one selection element functions efficiently in one organism and another element functions efficiently in another organism.Examples of selection elements include, but are not limited to, (1) nucleic acid segments encoding products that provide resistance to otherwise toxic compounds (e.g., antibiotics); (2) nucleic acid segments encoding products that are otherwise lacking in receptor cells (e.g., essential products, tRNA genes, acute nutritional markers); (3) nucleic acid segments encoding products that suppress the activity of gene products; (4) nucleic acid segments encoding easily identifiable products (e.g., phenotypic markers, e.g., antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind products otherwise harmful to cell survival and / or function; and (6) nucleic acid segments that otherwise inhibit the activity of any of the nucleic acid segments described in No. 1-5 above (e.g., antisense oligonucleotides). (7) Nucleic acid segments that bind products that modify substrates (e.g., restriction endonucleases); (8) Nucleic acid segments that can be used to isolate or identify a desired molecule (e.g., specific protein binding sites); (9) Nucleic acid segments that encode specific nucleotide sequences that can otherwise be non-functional (e.g., for PCR amplification of subpopulations of molecules); (10) Nucleic acid segments that, in the absence of such segments, directly or indirectly confer tolerance or sensitivity to a particular compound; (11) Nucleic acid segments that encode products that convert a toxic or relatively non-toxic compound in a recipient cell into a toxic compound (e.g., herpesthymidine kinase, cytosine deaminase); (12) Nucleic acid segments that inhibit the replication, distribution, or heritability of nucleic acid molecules containing them; and / or (13) Nucleic acid segments that encode conditioned replication function, e.g., replication in a particular host or host cell line or under certain environmental conditions (e.g., temperature and nutritional status).
[0184] The nucleic acid reagent can be in any form useful for in vivo transcription and / or translation. The nucleic acid may be a plasmid such as a supercoiled plasmid, may be a yeast artificial chromosome (e.g., YAC), may be a linear nucleic acid (e.g., a linear nucleic acid generated by PCR or by restriction digestion), may be single-stranded, or may be double-stranded. The nucleic acid reagent may be produced by an amplification process such as a polymerase chain reaction (PCR) process or a transcription-mediated amplification process (TMA). In TMA, two enzymes are used in an isothermal reaction to generate amplification products detected by luminescence (e.g., Biochemistry 1996 Jun 25;35(25):8429~38). The standard PCR process is known (e.g., U.S. Patent Nos. 4,683,202; 4,683,195; 4,965,188; and 5,656,493) and is generally performed within a plurality of cycles. Each cycle includes heat denaturation where hybridized nucleic acids dissociate; cooling where primer oligonucleotides hybridize; and elongation of the oligonucleotides by a polymerase (i.e., Taq polymerase). An example of the PCR cycle process is to treat the sample at 95°C for 5 minutes; repeat 45 cycles of 1 minute at 95°C, 1 minute at 59°C, 10 seconds, and 1 minute 30 seconds at 72°C; then treat the sample at 7 2°C for 5 minutes. Multiple cycles are frequently performed using a commercially available thermal cycler. The PCR amplification products may be stored at a lower temperature (e.g., 4°C) for some time and may, in some cases, be frozen (e.g., at -20°C) before analysis.
[0185] Manufacturing kits / articles This specification discloses manufacturing kits and articles used in certain embodiments in conjunction with one or more methods described herein. Such kits include a carrier, package, or container partitioned to receive one or more containers, such as vials and tubes, each of which contains one of the individual elements used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. In one embodiment, the containers are formed from a variety of materials, such as glass or plastic.
[0186] In some embodiments, the kit includes packaging material suitable for containing the contents of the kit. In some cases, the packaging material is constructed in a well-known manner to provide a preferably sterile, contaminant-free environment. Packaging materials used herein may include, for example, those conventionally used in commercial kits sold for use with nucleic acid sequencing systems. Exemplary packaging materials include, but are not limited to, glass, plastic, paper, and foil, which are capable of holding the components described herein within fixed limits.
[0187] The packaging material may include labels indicating specific uses for the components. The uses indicated by the labels for the kit may be one or more of the methods described herein, appropriately tailored to specific combinations of components present in the kit. For example, the labels may indicate that the kit is useful for methods of synthesizing polynucleotides or for determining nucleic acid sequences.
[0188] Instructions for using the packaged reagents or components may also be included in the kit. These instructions typically include tangible expressions describing reaction parameters such as the relative amounts of the kit components and the samples to be mixed, the retention period of the reagent / sample mixture, temperature, and buffering conditions.
[0189] It should be understood that not all components necessary for a particular reaction must be present in a given kit. Rather, one or more additional components may be provided from other sources. The instructions provided with the kit may specify any additional components that may be provided and where they can be obtained.
[0190] In some embodiments, a kit is provided that is useful for stably incorporating non-natural nucleic acids into cellular nucleic acids, for example, using a method provided by the present invention for producing genetically modified cells. In some embodiments, the kit described herein comprises genetically modified cells and one or more non-natural nucleic acids. In other embodiments, the kit described herein comprises an isolated and purified plasmid containing a sequence selected from SEQ ID NOs: 1 to 32.
[0191] In additional embodiments, the kit described herein provides cells and a nucleic acid molecule containing a heterogeneous gene for introduction into the cells and thereby providing genetically engineered cells, such as an expression vector containing the nucleic acid of any of the embodiments described above in this paragraph. ru.
[0192] A certain term Unless otherwise defined, all technical and chemical terms used herein have the same meaning as those generally understood by those skilled in the art in the field to which the subject matter described in the claims pertains. It will be understood that the above summary description and the following detailed description are illustrative and descriptive only and do not limit any subject matter described in the claims. In this application, the use of the singular form includes the plural unless otherwise indicated. It should be noted that the singular forms "a," "an," and "the" used herein and in the appended claims include multiple subjects unless otherwise indicated by the context. In this application, the use of "or" means "and / or" unless otherwise indicated. Furthermore, the use of the terms "including" and other forms, such as "include," "includes," and "included," is not limiting.
[0193] As used herein, ranges and quantities can be expressed as "approximately" specific values or ranges. "Approximately" also includes exact quantities. Therefore, "approximately 5 μL" also means "approximately 5 μL" and "5 μL". Generally, the term "approximately" includes quantities that can be predicted to be within experimental error.
[0194] The section titles used herein are for organizational purposes only and should not be construed as limiting the scope of what is described.
[0195] Examples These examples are provided for illustrative purposes only and do not limit the scope of the claims provided herein. [Examples]
[0196] Determining how cells retain or lose UBP within E. coli Under steady-state conditions, DNA containing dNAM-dTPT3 UBP is replicated in vitro with an efficiency approaching that of the complete native counterpart. However, these rates are likely to be limited by product dissociation. In vivo replication is more processive and correspondingly appears to be less limited by product dissociation. Thus, replication of DNA containing UBP within SSO appears to be less efficient than that of completely native DNA and thus may stall the replication fork. Further structural studies have shown that UBP adopts a Watson-Crick-like structure during triphosphate insertion, but once inserted, UBP adopts a cross-strand intercalate structure that induces local helical distortion. 8,9 Cells interpret both stalled replication forks and helical distortions as signs of DNA damage and initiate programs to repair or tolerate nucleotides that are likely to be the source of the impairment and that the inventors suspect may be involved in UBP loss.
[0197] To determine how cells retain or lose UBP, the impact of these pathways being dysfunctional was studied. The results show that neither nucleotide excision repair (NER) nor the SOS response is significantly involved in UBP retention or loss. Conversely, normal replisome polymerases, DNA polymerase III (PolIII), PolII, and methyl-directed mismatch repair (MMR) are all involved in UBP retention; while recombinational repair of stalled replication forks (RER) provides the major pathway for UBP loss. Next, the replisome of SOO was reprogrammed to not only better retain UBP on plasmids but also to confer the ability to stably retain UBP within its chromosome.
[0198] Nucleotide excision repair is not involved in YBP retention or loss In general, *E. coli* responds to DNA damage via direct damage inversion, base excision duplication, NER, MMR, RER, and SOS responses. Since these pathways rely on enzymes that recognize specific forms of DNA damage not mimicked by UBP, neither direct damage inversion nor base excision repair appears to be involved in UBP retention or loss. In contrast, NER, MMR, RER, and SOS responses are induced by fewer structure-specific signals. To begin investigating how cells manage to retain UBP in their DNA, we studied NER, mediated by a replication-independent method using a protein complex that scans DNA for strains resulting from bulk lesions mimicked by UBP. NER's involvement in UBP retention or loss was investigated by deleting uvrC, which encodes an essential component of NER, from parental SSO (*E. coli BL21(DE3)+pACS2 (Figure 4)). Replication of DNA containing dNaM-dTPT3 UBP positioned in two different sequence configurations in plasmids pINF1 and pINF2 was unaffected by uvrC deletion, indicating that NER is not involved in UBP retention or loss (Figure 1B).
[0199] Methyl-directed mismatch repair increases UBP retention. Next, we investigated MMR, which performs a decisive first check on newly synthesized DNA as it emerges from DNA polymerase during replication and is mediated by a protein complex that recognizes helical strain caused by mismatched native nucleotides. Upon detection of mismatches, the MMR complex cleaves the newly synthesized unmethylated chain, leading to gap formation and subsequent DNA resynthesis. In contrast to NER, inactivation of MMR via mutH deletion resulted in reduced UBP retention by both pINF1 and pINF2 (Figure 1B). These results indicate that helical strain associated with UBP is not severe enough to activate MMR or to excise non-native nucleotides, however, strain caused by pairing of non-native and native nucleotides is recognized and processed by MMR. Thus, MMR appears to effectively recognize UBP as native and selectively remove mispaired native nucleotides, thereby supporting stable expansion of the gene alphabet.
[0200] Recombination repair provides a primary pathway for UBP loss. RER is mediated by RecA, which forms a filament on single-stranded DNA before the stunted replication fork, facilitating the formation of a recombinant intermediate and switching to an allotype template for continued DNA replication. The SOS response is triggered when the same RecA filament promotes the cleavage of the SOS repressor LexA, resulting in the desuppression of various genes involved in the tolerance and / or repair of damaged DNA that stunted the fork. We investigated the combined involvement of RER and SOS responses through recA deletion and observed a significant increase in UBP retention by pINF1 (Figure 1B). To further investigate the involvement of RecA, we measured UBP retention in more challenging sequences provided by pINF3, pINF4, and pINF5 using ΔrecA SSO (Figure 1C). In these sequence configurations, the absence of recA resulted in a more dramatic increase in UBP retention.
[0201] To determine whether recA deletion facilitates UBP retention by removing the RER or preventing the induction of the SOS response, we examined an SSO (SSO lexA(S119A)) that is RER-friendly but does not induce the SOS response (Figure 1C). Selective suppression of the SOS response resulted in a moderate increase in UBP retention by pINF3, but this increase was lower than that observed with ΔrecA SSO. In pINF4 and pINF5, selective SOS suppression resulted in only a moderate increase in UBP retention, which was considerably lower than that observed with recA SSO. These results suggest that the majority of UBP loss mediated by RecA is mediated through the RER. This demonstrates that it occurs without the induction of an SOS response.
[0202] PolII is involved in the replication of UBP-containing DNA. The data suggest that much of the UBP loss is mediated via RER, but the slight sequence-specific increase in UBP retention by lexA(S119A)SSO suggests that one or more SOS regulatory proteins may also contribute. We investigated the involvement of three SOS regulatory DNA polymerases, PolII, PolIV, and PolV. Indeed, PolIV and PolV are "damage-overcoming" polymerases, known for their ability to mediate DNA synthesis at "non-instructive" damaged nucleotides. However, deletions of dinB and umuCD (encoding precursors of PolIV and PolV, respectively) did not affect UBP retention in either pINF1 or pINF2 (Figure 1D). In contrast to ΔdinBΔumuDC SSO, deletion of polB (encoding PolII) resulted in a dramatic increase in UBP loss in both pINF1 and pINF2 (Figure 1D). Overall, these data demonstrate that RER constitutes the primary pathway to UBP loss, and PolII provides a significant pathway to UBP retention. PolII generation increases with SOS induction, but the data suggest that its beneficial role is diminished by the detrimental effects of the accompanyingly induced RER.
[0203] DNA polymerase III is also involved in the replication of UBP-containing DNA. The reduced but still detectable retention rate of UBP in ΔpolB SSO strongly suggests that one or both of the remaining DNA polymerases, PolI and PolIII, must also be involved in UBP retention, along with the minimal impact of deletions of the genes encoding PolIV and PolV. When specifically examining whether PolI or PolIII is involved in the replication of UBP-containing DNA, their 3'-5' exonuclease ("proofing") activity may be affected by mutations (PolI, respectively). exo- , polA (D424A, K890R), and PolIII exo- We characterized strains in which dnaQ(D12N) was eliminated or impaired (Figures 4 and 6). While deletion of PolI exonuclease activity had no effect on UBP retention, PolIII exonuclease deletion mutants showed a dramatic reduction in UBP retention. These data clearly demonstrate that, in wild-type cells, PolIII, rather than PolI, is involved in the replication of UBP-containing DNA.
[0204] To determine whether any effect of PolI or PolIII mutants was masked by PolII and / or RER activity, UBP retention was examined in ΔpolB or ΔpolBΔrecA SSO. The results showed that UBP was well retained in ΔpolBΔrecA SSO, demonstrating that polymerases other than Pol II can mediate high levels of UBP retention without competition with RER-mediated loss (Figure 1D). PolIII exonuclease mutants again showed reduced UBP retention in both ΔpolB and ΔpolBΔrecA SSO. However, in contrast to wild-type cells, deletion of PolI exonuclease activity had a significant and opposite effect in ΔpolB and ΔpolBΔrecA SSO, increasing and decreasing their retention, respectively. These data demonstrate that PolIII, in addition to PolII, is involved in UBP retention, as is PolI in the absence of RER.
[0205] Model for DNA replication containing UBP While not adhering to any particular theory, the results described herein suggest the following model for replicating DNA containing dNaM-dTPT3 UBP in E. coli SSO (Figure 2). When a replisome with PolIII encounters a non-native nucleotide during forward reading or lagging strand replication, PolIII incorporates either a native or non-native nucleotide. If a native nucleotide is incorporated, proofreading... The rate of extension competes with and is likely more efficient than this extension, and therefore native nucleotides are generally removed via the proofreading activity of PolIII. However, if the correct UBP is synthesized, the more efficient extension prevents removal, and the replisome continues to synthesize DNA. As it exits polymerase, the nascent double helix is scanned by the MMR complex, further increasing UBP retention by preferentially removing any mispaired native nucleotides that have escaped proofreading.
[0206] Furthermore, since correct UBP elongation appears to be less efficient than natural synthesis, PolIII may dissociate. The stalled fork, as in the case of an elongated strand terminating immediately before a non-natural nucleotide in the template, is currently a substrate for the RER, restarting synthesis using the same natural sequence and thus providing the dominant mechanism for UBP loss. However, in competition with the RecA-mediated RER, PolII can rescue the stalled fork, restarting synthesis with high UBP retention, and then possibly leading to the re-establishment of a normal replication fork, possibly resulting in PolIII. PolI's involvement is more complex. In wild-type cells, PolI does not appear to be involved in the replication of UBP-containing DNA. In contrast, in the absence of PolII and RecA, PolI is not involved, and accordingly, the loss of PolI exonuclease activity results in reduced UBP retention. However, when exonuclease activity is lost, PolI can take over if PolII is lost, and in this case, retention increases due to competition with RER.
[0207] PolII is recognized to have two presumed roles: (1) restarting replication when PolII rescues a stuck fork after PolIII has synthesized mispairs that cannot be efficiently extended; and (2) PolII competing with RER to fill gaps created by NER as part of the cellular response to interstrand crosslinked DNA. Interestingly, the evoked role of PolII in rescuing UBP-stuck replication forks competing with RER is remarkably similar to both aspects of its presumed innate role. However, this effect on replication of UBP-containing DNA is the most significant phenotype observed to date due to its elimination.
[0208] SSO optimization UBP retention can be optimized through manipulation of RecA and PolII. To investigate this possibility, SSO was optimized with or without constitutively expressed PolII at the SOS suppression release level, while depriving recA (ΔrecA and PolII, respectively).+ ΔrecA) (Figure 6). These strains (YZ3) also expressed the optimized PtNTT2 transporter from the chromosomal locus (ΔlacZYA::P lacUV5 -PtNTT2(66-575)) (Figure 4). For comparison, a wild-type strain (WT-Opt) with the same chromosomally integrated transporter was used. SSOs were transformed with pINF1, pINF5, or pINF6 (Figure 3A), and with pINF6, UBP was embedded in sequences whose retention was particularly difficult, and plasmids were recovered from individual colonies to characterize UBP retention. In this case, the introduction of the selection of the solid growth medium enabled the analysis of UBP retention in individual clones, in contrast to the average UBP retention determined in the previous experiment. The distribution of UBP retention was observed for each plasmid in all SSOs, but the distribution shifted to a higher retention rate in ΔrecA-Opt compared to WT-Opt SSOs, especially in PolII + ΔrecA SSOs. Furthermore, PolII + Only ΔrecA SSOs generated clones with undetectable UBP loss in each sequence configuration tested. In particular, this is also true for pINF6, for which the retention rate in wild-type SSOs was undetectable and was only moderate (<60%) when enhanced with Cas9 selection.
[0209] Genetically optimized ΔrecA-Opt and PolII + ΔrecA SSOs were evaluated for their ability to facilitate the integration of UBP into the chromosome. A cassette targeting the sequence GTAXTGA (X = NaM) to the arsB locus was constructed, and using lambda red recombination, the cassette was integrated into the chromosomes of WT-Opt, ΔrecA-Opt, and PolII + ΔrecA SSOs. Screening of the integrants for UBP retention was performed for ΔrecA-Opt and PolII +While clones with 100% retention were identified from ΔrecA SSO, despite considerable effort, the inventors were unable to isolate WT-Opt clones with UBP retention rates higher than 91% (Figure 7), suggesting that significant UBP loss occurred during the required growth process. To characterize the effect of chromosomal integration UBP, a certain amount of metaphase logarithmic phase cells were inoculated into growth media with or without dNaMTP and dTPT3TP (Figures 3B and 8). ΔrecA-Opt and PolII + The ΔrecA integration was poorly developed when non-natural triphosphates were not provided, which was consistent with a model in which RER is required to efficiently bypass non-natural nucleotides in the template. However, this growth defect was almost completely eliminated in both SSOs when dNaMTP and dTPT3TP were provided. Therefore, deletion of recA and overexpression of PolII facilitate high levels of UBP retention in chromosomes that have minimal consequences for fitness.
[0210] Finally, we evaluated whether genetically optimized strains facilitate the long-term stability of chromosomal integrated UBPs. Previous studies have demonstrated that without Cas9-mediated selection for retention, plasmid-derived UBPs are lost during long-term growth. WT-Opt, ΔrecA-Opt, and PolII + ΔrecA constructs were successively passaged over many generations of growth, and UBP retention was characterized (Figure 3C). In WT-Opt, UBP was slowly lost up to approximately 40 generations, then lost more rapidly, with complete loss observed up to 90 generations. The clearly biphasic dynamics of loss suggest that at least one additional process is further involved in RER. Indeed, sequencing revealed a large-scale chromosomal rearrangement that eliminated the PtNTT2 gene at the point of the sharp drop in UBP retention (Figure 10). WT-Opt, ΔrecA-Opt, and PolII + In contrast to both ΔrecA SSOs, PtNTT2 remains intact, and the retention rate of genomic UBP is particularly high for PolII. +ΔrecA remained high during SSO and remained above 55% after 137 generations.
[0211] These results demonstrate that recA deletion not only facilitates UBP retention during replication but also significantly increases transporter stability during long-term growth. The observed retention rate corresponds to a fidelity per doubling of over 99.6%, which corresponds to the loss of chromosomal UBP in only a small percentage (<0.4%) of cells per doubling. Therefore, this error prevention system, along with the Cas9 error elimination system not used in the current study, should enable UBP retention across a wide range of sequence configurations, and thus allow for the preservation of the new information made possible by UBP.
[0212] Since the last common ancestor of all life on Earth, biological information has been preserved using four alphabetic letters: PolII. + The reprogrammed replisomes of ΔrecA SSO represent a significant step toward the unlimited expansion of this alphabet, with the first step mediated through the optimization of the cell itself. While the primary goal of the study was to understand how UBP replicates and use that information to optimize SSO, the results also provide a novel avenue for studying how difficult replication is successfully managed. For example, the data suggest that a considerable proportion of UBP-containing DNA is replicated by PolIII, but also clearly show that a significant amount is not, and in these cases, the data reveal an intriguing competition between PolII-mediated replication restart and RecA-mediated RER. Such competition is considered common in difficult replication and may have contributed to the challenge in identifying the normal role of PolII. Furthermore, the inability of MMR to recognize UBP suggests that helical strain alone is insufficient, and that the process requires specific interactions with nucleic acid bases not available in non-native nucleotides. Finally, the increased genetic stability obtained by the deletion of recA may have significant implications for methods directed by extensions of the gene code via Amber repression, because these methods also suffer from genetic instability for extended growth. 23 Regardless of these intriguing challenges, reprogrammed SSOs should, as demonstrated previously, enable a more stable retention of the increased biological information contained within their chromosomes, and that this information can be retrieved in the form of proteins with non-canonical amino acids, thus providing a platform for achieving a central goal of synthetic biology—the creation of life with new forms and functions. [Examples]
[0213] Methods and materials pINF / UBP-containing DNA composition pINF (Figure S8) has the following modifications, which have already been reported. 3The assembly consisted of a Golden Gate configuration with an insert dsDNA containing pUCX2 and a dNaM-dTPT3 pair. The UBP-containing dsDNA was generated by PCR in 50 μL of a solution containing chemically synthesized UBP-containing oligonucleotide (0.025 ng / μL), primers for introducing the BsaI site and vector homology (1 μM, Table S1), dTPT3TP (100 μM), dNaMTP (100 μM), dNTP (200 μM), MgSO4 (1.2 mM), OneTaq DNA polymerase (0.025 U / μL), and OneTaq standard reaction buffer (1×, New England Biolabs). The reaction was repeated on the MJ Research PTC-200 system following the following temperature regime (times in mm:ss): [94℃ 00:30|25×(94℃ 00:30|47℃ 00:30|68℃ 04:00)]. The resulting UBP-containing dsDNA was purified using DNA Clean & Concentrator-5 (Zymo Research) according to the manufacturer's recommendations. For pINF assembly, pUCX2 (1 μg) and insert DNA were combined in a 1:4 molar ratio in 80 μL of reaction mixture containing ATP (1 mM), T4 DNA ligase (6.65 U / μL, New England Biolabs), BsaI-HF (0.66 U / μL, New England Biolabs), and CutSmart Buffer (1×, New England Biolabs), and subjected to the following temperature regime: [37°C 20 min | 40× (37°C 5 min | 16°C 10 min | 22°C 5 min) | 37°C 20 min | 50°C 15 min | 70°C 30 min]. Then, BsaI-HF (0.33 U / μL) and T5 exonuclease (0.16 U / μL, New England Biolabs) were added, and the reaction mixture was incubated at 37°C for 1 hour to remove any pUCX2 without the insert. The reaction product was purified using DNA Clean & Concentrator-5 according to the manufacturer's recommendations, except that it was mixed with 3 volumes of 1:1 DNA wash:DNA binding buffer and then bound to a silica column.
[0214] A UBP knock-in cassette for the arsB seat (Figure S4) containing UBP, 150bp The dsDNA was generated by PCR overlap of the kanamycin resistance gene in pKD13. 150 bp DNA was generated by PCR in 50 μL using the same reaction solution conditions as above and the following temperature regime (units of time: mm:ss) [98°C 02:00 | 5× (98°C 00:10 | 50°C 00:10 | 68°C 04:00) | 15× (98°C 00:10 | 58°C 00:10 | 68°C 04:00)]. The kanamycin resistance gene amplicon was generated by PCR amplification off pKD13 using Q5 DNA polymerase, as recommended by the manufacturer. Amplification of long DNA (approximately 200 bp or longer) is inhibited by the presence of dTPT3TP. Therefore, overlap assembly PCR of UBP-containing amplicon and kanamycin resistance gene amplicon was performed using the following solution conditions: UBP-containing amplicon (0.02 ng / μL), kanamycin resistance gene amplicon (0.02 ng / μL), primer (1 μM, Table S1), dTPT3TP (5 μM), dNaMT P (100μM), dNTP (200μM), MgSO4 (1.2mM), OneTaq The reaction was carried out on a large scale using DNA polymerase (0.025 U / μL) and OneTaq standard reaction buffer (1×) (2 mL of the reaction mixture was divided into 40 individual 50 μL reactants). The reactants were subjected to the following temperature regime (times in mm:ss): [98°C 02:00 | 5× (98°C 00:10 | 50°C 00:10 | 68°C 04:00) | 15× (98°C 00:10 | 58°C 00:10 | 68°C 04:00)]. These reactants were collected and concentrated using DNA Clean & Concentrator-5 according to the manufacturer's recommendations.
[0215] In vivo UBP replication in gene knockout All gene knockouts (Figure 1 and Figure S2) were assayed for their ability to replicate pINF-derived UBP according to the protocol described below. Electrocompetent cells were pelleted and washed twice with 50 mL of sterile deionized water (diH2O) at 4°C to obtain metaphase logarithmic phase cells (OD). 600 Prepared from 45 mL of culture (0.35-0.7). Washed cells were placed in sterile deionized water at 4°C, with a final OD of 40-60. 600The cells were resuspended. 50 μL of cells were mixed with pINF2ng assembled in Golden Gate and transferred to an electroporation cuvette (2 mm gap, Cat.#FB102, Fisher Scientific). Electroporation was performed using a Gene Pulser II (BioRad) according to the manufacturer's recommendations (voltage 25 kV, capacitor 2.5 μF, resistor 200 Ω). Transformed cells were diluted in 950 μL of 2×YT containing chloramphenicol (33 μg / mL) and potassium phosphate (50 mM, pH 7). 40 μL of diluted cells were further diluted in 200 μL of 2×YT containing chloramphenicol (33 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7), transferred to a 1.5 mL test tube, and collected at 37°C and 230 RPM for 1 hour. 10 μL of the collected cells were diluted in a well of a 96-well plate (see #655161, Greiner Bio-One) in 100 μL of 2×YT containing chloramphenicol (33 μg / mL), ampicillin (100 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and potassium phosphate (50 mM, pH 7). Furthermore, the harvested cells were seeded on 2×YT agar (2%) containing ampicillin (100 μg / mL) and potassium phosphate (50 mM, pH 7) to evaluate transformation efficiency. The 96-well plates and transformation efficiency plates were kept overnight (approximately 12 hours) at 4°C and 37°C, respectively. The transformation efficiency plates were examined to ensure that all samples in the 96-well plates had received at least 50 colony-forming units before refrigeration. The 96-well plates were then transferred to 37°C and 230 RPM. The cells were pelleted, decanted, and refrigerated to 0.6–0.92 OD. 600The cells were frozen after reaching the specified time. In vivo replicated pINF cells were isolated using the ZR plasmid Miniprep-Classic kit (Zymo Research) and a 5 μg silica column (Cat.#D4003, Zymo Research) as recommended by the manufacturer, and proceeded to biotin-shift PCR analysis (see Supplementary Information). This procedure was performed at least triple for each knockout strain, starting from the production of electrocompetent cells.
[0216] Under these conditions, it should be noted that the clones and strains undergo cell doubling of a similar, rather than identical, number during pINF replication experiments. However, due to the pINF-unregulated origin of replication, cell matching doubles between clones, and strains correspond to matching of the number of pINF replication events. Therefore, the data in Figures 1 and 3A are reported as retention percentage values, as opposed to estimated fidelity (see supplementary information for further consideration), and should be interpreted in this manner.
[0217] Testing of clonal pINF The ability of the optimized strain to clone pINF was evaluated as described above with the following modifications (Figure 3A). After harvesting, diluted versions of the harvested cultures were seeded onto 2×YT plates containing agar (2%), carbenicillin (100 μg / mL), chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7). The plates were incubated at 37°C for approximately 12 hours. Individual colonies were collected and transferred to 100 μL of 2×YT plates containing carbenicillin (100 μg / mL), chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7) in the wells of a 96-well plate. The 96-well plates were kept at 4°C for approximately 12 hours, then transferred to 37°C and 230 RPM. Cells were pelleted, decanted, and given an OD of 0.6–0.9. 600The cells were frozen after reaching the specified time. In vivo replicated pINF cells were isolated using the ZR plasmid Miniprep-classic kit, as recommended by the manufacturer, and then subjected to biotin-shift PCR analysis (see Supplementary Information).
[0218] PolII used in these experiments + It should be noted that the ΔrecA strain (Figure 3A) possessed a neocassette at the former recA locus (P_polB(-)lexA-polB+FRT+ΔrecA+KanR+lacZYA::P_lacUV5-ΔΔ(CoOp)col 2.1, Table S1).
[0219] UBP integration in arsB The UBP embedded cassette for the arsB seat was configured as described above and as shown in Figure S4. The embedded cassette is a standard Lambda Red recombinant with the following modifications. 24 The procedure was performed using overnight cultures of strains containing pKD46 (WT-Opt, ΔrecA-Opt, and PolII). + Diluting ΔrecA-Opt (added to 2×YT containing chloramphenicol (5 μg / mL) and KPi (50 mM, pH 7)) into a 2×YT containing ampicillin (100 μg / mL), chloramphenicol (5 μg / mL), and KPi (50 mM, pH 7) results in a concentration of 0.03 OD. 600 The culture was then prepared to approximately 0.1 OD. 600 It was allowed to grow, then induced with 0.4% L-(+)-arabinose, resulting in approximately 0.4 OD. 600Growth was continued until [a certain point]. Electrocompetent cells were prepared from these cultures as described above. 50 μL of electrocompetent cells were mixed with 960 ng of the embedded cassette described above (5 μL at 192 ng / μL) and electroporated as described above. Transformed cells were diluted to a final volume of 1 mL of 2×YT containing chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7), transferred to a 1.5 mL test tube, and harvested at 37°C and 230 RPM for 2 hours. Cells were pelleted and resuspended in 115 μL of 2×YT containing chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7). A 15 μL sample of this cell suspension was seeded onto 2×YT plates containing agar (2%), kanamycin (50 μg / mL), chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7). The plates were incubated at 37°C for 14 to 24 hours. Colonies were harvested and transferred to 500 μL of 2×YT plates containing kanamycin (50 μg / mL), chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7) in a 48-well plate (#677180, see Greiner Bio-One). The plates were refrigerated at 4°C for approximately 12 hours and then incubated at 37°C and 230 RPM, or incubated directly. 0.6~1OD 600 After reaching the target temperature, the culture was sampled as follows: 100 μL was combined with 100 μL of glycerin (50%) and frozen at -80°C; 350 μL was pelletized and frozen for later isolation of genomic DNA; 50 μL was pelletized, washed once with 200 μL of deionized water, pelletized, and resuspended in 200 μL.
[0220] Cell suspensions were analyzed by colony biotin-shift PCR (see Supplementary Information). Genomic DNA was isolated from frozen cell pellets stored for samples showing high colony biotin-shift PCR percent shift values (≥80%) using the PureLink Genomic DNA Mini-Kit (Thermo Fisher Scientific), as recommended by the manufacturer. Genomic DNA was analyzed by biotin-shift PCR (see Supplementary Information). This analysis showed high retention rates for all gene backgrounds (retention rate). B This revealed a retention rate of ≥90%. While these results confirmed successful chromosomal integration of UBP and remarkably high retention rates of UBP in chromosomal DNA, it was suspected that the cells lost their dTPT3TP and dNaMTP culture medium during the integration protocol, given the protocol requirements of incubating cells at high cell density. It is known that actively growing cultures of E. coli degrade extracellular dTPT3TP and dNaMTP into their corresponding diphosphate and monophosphate and nucleoside species. 5 To address this possibility, using a glycerol stock of the highest retention sample, 2×YT100μL containing kanamycin (50μg / mL), chloramphenicol (5μg / mL), dTPT3TP (37.5μM), dNaMTP (150μM), and KPi (50mM, pH7) was inoculated into a 96-well plate. The cultures were incubated at 37°C and 230 RPM for approximately 0.6 OD. 600 The cells were grown to this stage. Cells from this culture were seeded, harvested, grown, and obtained as samples as described above. This "re-seeding" procedure has undetectable chromosome UBP loss (retention rate). B =100%) ΔrecA-Opt and PolII + Clones related to ΔrecA-Opt SSO were rapidly identified. However, despite screening 12 clones for WT-Opt SSO, the retention rate was low. B >91% of clones were not found. Therefore, WT-Opt constructs that did not undergo the reseeding procedure for doubling time and subculturing experiments (retention rate) BWe choose to use ΔrecA-Opt and PolII (=91%). + In the case of ΔrecA-Opt, retention rate with respect to doubling time and passaging experiments. B We selected one clone each that had =100%.
[0221] PolII used in these experiments + It should be noted that the ΔrecA strain (Figures 3B and 3C) did not have a neocassette at the former recA locus (P_polB(-)lexA-polB+ΔrecA+FRT+lacZYA::P_lacUV5-ΔΔ(CoOp)col 1.1, Table S1).
[0222] Determining the stock doubling time Metaphase log-phase cells WT-Opt, ΔrecA-Opt, and PolII + ΔrecA-Opt SSO and their corresponding chromosomal UBP integrations (described above) were prepared using the following procedure: Saturated overnight cultures were inoculated with 2×YT containing chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7) from a glycerol stock stab, and grown overnight at 37°C and 230 RPM (approximately 14 hours). These cells were then fertilized in 500 μL of 2×YT containing chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7) at 0.03 OD. 600 Diluted to OD. Growth, 600 The cells were monitored by the metaphase log phase (0.3-0.5 OD). 600 Once the target is reached, in a 48-well plate, use 500 μL of 2×YT containing chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7), or use 2×YT containing chloramphenicol (5 μg / mL) and KPi (50 mM, pH 7), to reach 0.013 OD. 600 Diluted and grown at 37°C and 230 RPM. OD600 The values were measured every 30 minutes. This procedure was repeated for each strain, starting with inoculation of the culture overnight.
[0223] OD from each experiment 600 By analyzing the data, we obtained the theoretical cell doubling time (Figure 3B and Figure S5). The OD corresponds to the exponential growth phase (0.01-0.9). 600 The measured values were fitted to the following exponential growth model using R version 3.2.4: 25
number
[0224] OD i However, OD at time (t) 600 If so, OD0 is the minimum OD with respect to a given data set. 600 It is a value, C growth C is the growth constant. growth This was applied using the "nls()" command. The doubling time (DT) was calculated using the following equation:
number
[0225] Passaging of strains that carry genomic UBP WT-Opt, ΔrecA-Opt, and PolII + Using glycerol stock stubs of chromosomal UBP integration from ΔrecA-Opt SSO (described above), 2×YT500μL containing kanamycin (50μg / mL), chloramphenicol (5μg / mL), dTPT3TP (37.5μM), dNaMTP (150μM), and KPi (50mM, pH7) was inoculated. Cells were incubated at 37°C and 230 RPM until metaphase logarithmic phase (0.5-0.8 OD) 600The cells were grown until they reached 0.03 OD, and then in a 48-well plate, 0.03 OD was added to 500 μL of 2×YT containing kanamycin (50 μg / mL), chloramphenicol (5 μg / mL), dTPT3TP (37.5 μM), dNaMTP (150 μM), and KPi (50 mM, pH 7). 600 Diluted to 0.03OD and grown at 37°C and 230 RPM. 600 The culture inoculated with [method] was considered the starting point for subculturing (doubling = 0). The culture was then subculturised to 1-1.5 OD, which corresponds to approximately 5 cell doubling. 600 It was grown to this size. 0.03 to 1-1.5 OD 600 This growth up to this point was considered as one "passage," where one passage corresponds to approximately 5 cell doubling. These samples were 1-1.5 OD 600 After reaching 0.03 OD, cells are placed in fresh culture medium of the same composition. 600 Dilution was initiated to start another subculturing. After dilution, the concentration was 1-1.5 OD. 600 The cultures were sampled as follows: 100 μL was combined with 100 μL of glycerin (50%) and frozen at -80°C; 350 μL was pelletized and frozen for subsequent isolation of genomic DNA; and 50 μL was pelletized, washed once with 200 μL of deionized water, pelletized, and resuspended in 200 μL. The subculturing process was repeated for all three strains up to 15 subcultures, which corresponds to a total of approximately 80 cell doublings.
[0226] Colony biotin-shift PCR analysis (see supplementary information) was performed on cell suspension samples throughout the entire passage. This revealed that retention rates decreased to less than 10% in WT-Opt after 15 passages. Therefore, this strain was no longer passaged. In contrast, retention rates were lower in Δrec-Opt and PolII + In ΔrecA-Opt, the retention rate was maintained at 60–80%. Therefore, additional passaging was performed on these strains as described above. The retention rate remained unchanged over a total of 16 passages. These strains were then subjected to four additional passages at a higher dilution factor, corresponding to approximately 13 cell doubling per passage (approximately 0.0001 to 1.5 OD). 600(Growing from). At this point, ΔrecA-Opt and PolII + The ΔrecA-Opt integrated organism underwent approximately 130 cell duplication, and UBP Retention remained above 40% according to colony biotin-shift PCR analysis. Further passages were deemed unnecessary, and the experiment was stopped for more rigorous analysis of genomic DNA samples collected during passage. This experiment was performed in triple replication, starting with inoculation of culture medium in genomic integration glycerol stock stubs.
[0227] After the completion of the passage experiment, genomic DNA was isolated and analyzed by biotin-shift PCR (Figure 3C) (see Supplementary Information). The slow, then rapid, loss of UBP in WT-Opt suggested that multiple processes were involved in UBP loss. lacUV5 -PtNTT2(66-575) was suspected to be mutated during the experiment because PtNTT2 expression causes slight growth defects. 3 Therefore, cells that inactivate transporters through mutation can gain a fitness advantage and quickly dominate experimental populations. This hypothesis was investigated through the isolation of individual clones from the end of WT-Opt passages and PCR analysis of purified genomic DNA (see Supplementary Information and Figure S7). Primers moving between several clones revealed that all genes between cat and insB-4, including PtNTT2(66-575), were deleted in these cells. The insB-4 gene encodes one or two proteins required for the transposition of the IS1 transposon. 26 Sequencing of one clone confirmed that the IS1 insertion at PtNTT2(66-575)(T1495) corresponds to a 15890 base pair deletion.
[0228] After confirming the PtNTT2(66-575) mutation event, the appearance of deletion mutants was evaluated by PCR analysis of genomic DNA samples from WT-Opt incorporated organisms (see Supplementary Information and Figure S7B). This analysis revealed that several amplicons of a size corresponding to the IS-1-mediated PtNTT2(66-575) deletion event appeared in the passaged samples during the rapid phase of UBP loss.
[0229] PolII + One replica of the ΔrecA-Opt embedding was also observed to rapidly lose UBP simultaneously with the WT-Opt embedding, strongly suggesting that this replica was contaminated with WT-Opt cells during passage. This possibility was confirmed using colony PCR analysis, revealing that this replica became contaminated with WT-Opt cells during passage corresponding to the rapid loss of UBP (see Supplementary Information and Figure S6). Therefore, data from this replica were only used from samples without WT-Opt cell contamination.
[0230] Bacterial strains and plasmids All strains used in this study (Table S1; provided as separate supplemental files) were constructed from E. coli-BL21(DE3) via lambda red recombination unless otherwise indicated. Gene knockout cassettes were obtained by PCR amplification of either genomic DNA or pKD13 of Keio collection strains using relevant primers (Table S1) (using OneTaq or Q5 as recommended by the manufacturer (New England Biolabs)). Functional gene knock-in cassettes, polA(D424A, K890R) and PolII+ (Figure S4), were constructed via overlap PCR. Strains were converted to pACS2 or pACS2-dnaQ(D12N) at the lacZYA locus or to PlacUV5 - PtNTT2(66-575) + It was deemed capable of importing XTP via the integration of one of the cat cassettes (Figure S1). pACS2 and P lacUV5 The configuration of -PtNTT2(66-575)+cat has already been described.3 pACS2-dnaQ(D12N) was constructed using a Gibson assembly of PCR amplicons. PtNTT2 function was confirmed in all relevant strains using a radioactive dATP uptake assay.
[0231] Exonuclease-deficient Pol I and III DNA PolI and III are conditionally essential and essential genes, respectively. Therefore, unlike SOS regulatory polymerases, they could not be tested by gene knockout. Instead, 3'-5' exonuclease-deficient mutants were constructed for these enzymes. PolI (polA) was deficient in 3'-5' exonuclease by mutating the active site (D424A) of its exonuclease domain. This was achieved through two phases of lambda red recombination (Figure S4). The first polA was cleaved to form its 5'-3' exonuclease domain (removing both the polymerase and 3'-5' exonuclease domains). The second polymerase and 3'-5' exonuclease domain were reintroduced with the D424A mutation. Due to the length of the genes, PCR mutations were generated in the amplicon used for integration. This resulted in the K890R mutation. However, since K890 is a surface-exposed residue on the disordered loop of the protein, its mutation to arginine was predicted to have minimal impact on protein function. Furthermore, arginine from lysine maintains the approximate charge and size of the residue.
[0232] The DNA PolIII holoenzyme is a multi-enzyme complex containing separate polymerases and 3'-5' exonuclease enzymes. The exonuclease enzyme (dnaQ) is thought to play a structural role in the PolIII homoenzyme in addition to its editing activity. Therefore, dnaQ deletion eliminates PolIII editing activity but also prevents cell growth unless compensatory mutations are added to other parts of the holoenzyme. Accordingly, we choose to test the role of PolIII in UBP replication via the expression of a mutagenic dnaQ mutant (D12N) from the plasmid pACS2+dnaQ(D12N) (Figure S1). Expression of dnaQ(D12N) from a multicopy plasmid has already been demonstrated to produce a dominant mutagenic phenotype in E. coli, regardless of the expression of wild-type DnaQ from chromosomal copies of the gene. pACS2+dnaQ(D12N) expresses dnaQ(D12N) via both the innate gene promoter and the pACS2+dnaQ(D12N).
[0233] Adaptation costs from gene optimization of SSO Deletion of recA clearly results in significantly improved retention of UBP in many sequences. While this is highly desirable, recA deletion retains some adaptive cost. It is known that deletions in recA in strains have lower resistance to DNA damage. However, since all recent applications of SSO occur in highly controlled environments, the inventors do not anticipate this to be a problem. Furthermore, recA deletion increases doubling time when measured in Figure S5. However, these experiments were mainly conducted to show differences in growth rates for strains retaining chromosomal UBP growing in the presence or absence of dNaMTP and dTPT3TP. Several factors complicate the strain fitness directly related to the measured doubling time. The main challenge is that cells in solution may increase OD600 by changing their morphology rather than actually increasing their cell number. Regardless, the doubling time measured for L recA-Opt (approximately 18 minutes longer than WT-Opt) suggests that recA deletion results in a significantly reduced growth rate. However, given the benefits of this correction, this reduced growth rate becomes an acceptable trade-off. It should also be noted that some data points in Figure 8 are difficult to rationalize. For example, the presence of chromosomal UBP affects L recA-Opt and PolII + L recA-Opt appears to reduce the doubling time.
[0234] Biotin-shift analysis The retention of UBP in pINF and chromosomal DNA was measured as described above with the following modifications: All biotin-shift PCRs were performed using primers (1 iiM, Table S1), d5SICSTP (65 iiM), dMMO2bioTP (65 μM), and dNTP (400 μM). iiM), MgSO4 (2.2mM), OneTaq DNA polymerase (0.018U / iiL), DeepVent DNA polymerase (0.007U / iiL, New England Biolabs), SYBR Green I (1×, Thermo Experiments were performed in 15-2L volumes containing Fisher Scientific and OneTaq standard reaction buffer (1×). The amount of sample DNA added to biotin-shift PCR and th...
Claims
1. A manipulated host cell, a. The first nucleic acid molecule containing non-natural nucleotides; b. A second nucleic acid molecule: The nucleotide sequence comprises at least one modified DNA repair response-related protein, wherein the modification is: (A) Deletion of RecA; and / or (B) A mutation in Escherichia coli LexA, the mutation comprising the S119A mutation; The second nucleic acid molecule comprising: Furthermore c. A third nucleic acid molecule encoding a nucleoside triphosphate transporter that imports at least one non-native nucleotide from Pheodactylum tricornutum into the engineered host cell, wherein the third nucleic acid molecule is either integrated into the genome sequence of the engineered host cell or comprises a plasmid encoding the nucleoside triphosphate transporter; The manipulated host cells, including the aforementioned manipulated host cells.
2. The manipulated host cell according to claim 1, wherein the nucleoside triphosphate transporter exhibits higher stability of expression in the manipulated host cell compared to expression in an equivalent manipulated host cell that does not contain the second nucleic acid molecule.
3. The manipulated host cell according to claim 1 or 2, wherein the nucleoside triphosphate transporter comprises N-terminal cleavage, C-terminal cleavage, or cleavage of both ends.
4. The manipulated host cell according to any one of claims 1 to 3, wherein the nucleoside triphosphate transporter comprises a codon-optimized nucleoside triphosphate transporter derived from Pheodactylum tricornutum.
5. The manipulated host cell according to any one of claims 1 to 4, wherein the nucleoside triphosphate transporter comprises PtNTT2.
6. The engineered host cell according to any one of claims 1 to 5, wherein the nucleoside triphosphate transporter or PtNTT2 derived from Pheodactylum tricornutum is under the control of a promoter containing a pSC plasmid or a promoter derived from a lac operon.
7. a. Cas9 polypeptide or its variants; and b. A single guide RNA (sgRNA) comprising a crRNA-tracrRNA scaffold, wherein the combination of the Cas9 polypeptide or its variant and the sgRNA regulates the replication of a first nucleic acid molecule containing non-native nucleotides. The manipulated host cell according to any one of claims 1 to 6, further comprising:
8. The modification comprises the deletion of RecA, as described in any one of claims 1 to 7, the modified host cell.
9. The manipulated host cell according to any one of claims 1 to 8, wherein the deletion of RecA includes an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, an internal deletion, and / or a deletion of the entire gene.
10. The manipulated host cell according to any one of claims 1 to 9, wherein the deletion of RecA includes the deletion of the entire gene.
11. The modified host cell according to any one of claims 1 to 7, wherein the modification includes a mutation in Escherichia coli LexA, and the mutation includes the S119A mutation.
12. The manipulated host cell according to any one of claims 1 to 11, wherein the manipulated host cell is a prokaryotic cell.
13. The manipulated host cell according to any one of claims 1 to 11, wherein the manipulated host cell is an Escherichia coli cell.
14. The manipulated host cell according to any one of claims 1 to 11, wherein the manipulated host cell is an Escherichia coli (DE3) cell.
15. Non-natural nucleotides or non-natural bases of non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-aminopropyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thiothymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo, and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocyto Syn, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azouracil, 6-azocytosine, azacytosine, 6-azothimine, 6-thioguanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenines and guanines;N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazinecytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimide[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazinecytidine (9-(2-aminoethoxy)-H-pyrimide [5,4-b][1,4]benzoxazine-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindolecytidine (H-pyrimido[3',2':4,5]pyrrolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carbazole Xymethylaminomethyluracil, dihydrouracil, β-D-galactosylquosin, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylquosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2 A manipulated host cell according to any one of claims 1 to 14, selected from -methylthio-N6-isopentenyladenine, uracil-5oxyacetic acid, weibtoxosin, pseudouracil, cuosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, and 2,6-diaminopurine.
16. Non-natural nucleotides are 【Chemistry 1】 An engineered host cell according to any one of claims 1 to 14, comprising a non-natural base selected from.
17. The manipulated host cell according to any one of claims 1 to 16, further comprising a non-natural nucleotide and a non-natural sugar moiety.
18. The non-natural sugar portion is (a) Modification at position 2': (i) OH; (ii) substituted lower alkyl, aryl, aralkyl, O-aryl, O-aralkyl, SH, SCH 3 , OCN, Cl, Br, CN, CF 3 , OCF 3 , SOCH 3 , SO 2 CH 3 , ONO 2 , NO 2 , N 3 , or NH 2 F; (iii) O-alkyl, S-alkyl, or N-alkyl; (iv) O-alkenyl, S-alkenyl, or N-alkenyl; (v) O-alkynyl, S-alkynyl, or N-alkynyl; or (vi) O-alkyl-O-alkyl, 2'-F, 2'-OCH 3 , or 2'- O(CH 2 ) 2 OCH 3 And, Here, alkyl, alkenyl, and alkynyl are substituted or unsubstituted C 1 ~C 10 Alkyl, C 2 ~C 10 Alkenil, C 2 ~C 10 Alkinyl, -O[(CH 2 )nO]mCH 3 , -O(CH 2 )nOCH 3 , -O(CH 2 )nNH 2 , -O(CH 2 ) nCH 3 , -O(C H 2 )n-ONH 2 , and -O(CH 2 )nON[(CH 2 ) nCH 3 )] 2 And n and m are between 1 and approximately 10. The modifications including; (b) Modification at the 5' position, including 5'-vinyl or 5'-methyl (R or S); (c) Modifications at the 4' position comprising 4'-S, heterocycloalkyl, heterocycloalkalyl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides; and, (d) The manipulated host cell according to claim 17, selected from any combination of (a) to (c).
19. The manipulated host cell according to any one of claims 1 to 18, wherein the manipulated host cell constitutively expresses or overexpresses DNA polymerase II.
20. The manipulated host cell according to claim 19, wherein DNA polymerase II is encoded by the polB gene.
21. The manipulated host cell according to claim 20, wherein the polB gene is desuppressed.
22. The manipulated host cell according to any one of claims 1 to 21, wherein the manipulated host cell exhibits improved retention of nucleic acid molecules containing non-natural nucleotides during DNA replication, compared to the retention of nucleic acid molecules containing non-natural nucleotides during DNA replication in equivalent manipulated host cells that do not contain the second nucleic acid molecule.
23. A method for increasing the production of a first nucleic acid molecule containing non-natural nucleotides, A step of incubating a manipulated host cell with a number of non-native nucleotides, wherein the manipulated host cell is: The first nucleic acid molecule containing non-natural nucleotides; A nucleoside triphosphate transporter derived from Pheodactylum tricornutum that delivers (imports) at least one non-native nucleotide into the manipulated host cell; and a second nucleic acid molecule: A nucleotide sequence encoding at least one modified DNA repair response-related protein. Including columns, the modification is: (A) Deletion of RecA; and / or (B) A mutation in Escherichia coli LexA, the mutation comprising the S119A mutation; The second nucleic acid molecule, which includes, Here, the manipulated host cell newly synthesizes the multiple non-native nucleotides. This includes the step of incorporating it into a first nucleic acid molecule, The method wherein the manipulated host cell exhibits enhanced retention of non-natural base pairs containing the non-natural nucleotide in the newly synthesized first nucleic acid molecule, compared to the retention in an equivalent manipulated host cell that does not contain the second nucleic acid molecule.
24. A method for preparing polypeptides containing unnatural amino acids, A step of incubating a manipulated host cell with multiple non-natural nucleotides and multiple non-natural amino acids, wherein the manipulated host cell is: The first nucleic acid molecule containing non-natural nucleotides; A nucleoside triphosphate transporter derived from Pheodactylum tricornutum that delivers (imports) at least one non-natural nucleotide into the manipulated host cell; and, The second nucleic acid molecule is: A nucleotide sequence encoding at least one modified DNA repair response-related protein. Including columns, the modification is: (A) a deletion of RecA; and / or (B) a mutation of Escherichia coli LexA, the mutation comprising the mutation comprising the S119A mutation, and the second nucleic acid molecule comprising the second nucleic acid molecule, Here, the manipulated host cell incorporates the plurality of non-native nucleotides into one or more newly synthesized first nucleic acid molecules; and, Here, the manipulated host cell incorporates the multiple non-natural amino acids into one or more newly synthesized polypeptides, thereby producing a polypeptide containing non-natural amino acids; Includes, Here, the at least one modified DNA repair response-related protein and the nucleoside triphosphate transporter enhance the retention of non-natural base pairs, thereby promoting the incorporation of the non-natural amino acid into a newly synthesized polypeptide to produce a polypeptide containing the non-natural amino acid. The aforementioned method.
25. The modification is the method according to claim 23 or 24, wherein the modification includes the deletion of RecA.
26. The method according to any one of claims 23 to 25, wherein the deletion of RecA includes an N-terminal deletion, a C-terminal deletion, a cleavage at both ends, an internal deletion, and / or a deletion of the entire gene.
27. The method according to any one of claims 23 to 26, wherein the deletion of RecA includes the deletion of the entire gene.
28. The method according to claim 23 or 24, wherein the modification includes a mutation in Escherichia coli LexA, and the mutation includes the S119A mutation.
29. The method according to any one of claims 23 to 28, wherein the nucleoside triphosphate transporter includes N-terminal cleavage, C-terminal cleavage, or cleavage at both ends.
30. The method according to any one of claims 23 to 29, wherein the manipulated host cell is a prokaryotic cell.
31. The method according to any one of claims 23 to 29, wherein the manipulated host cell is an Escherichia coli cell.
32. The manipulated host cell is an Escherichia coli (DE3) cell, any one of claims 21 to 29 The method described in section [section number].
33. The method according to any one of claims 23 to 32, wherein the nucleoside triphosphate transporter includes a codon-optimized nucleoside triphosphate transporter derived from Pheodactylum tricornutum.
34. The method according to any one of claims 23 to 33, wherein the nucleoside triphosphate transporter comprises PtNT2.
35. The method according to any one of claims 23 to 34, wherein the nucleoside triphosphate transporter or PtNT2 derived from Pheodactylum tricornutum is under the control of a promoter containing a pSC plasmid or a promoter derived from a lac operon.
36. Non-natural nucleotides or non-natural bases of non-natural nucleotides include 2-aminoadenine-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-aminoadenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, and 3-deazaguani. 3-deazaadenine, 4-thiouracil, 4-thiothymine, uracil-5-yl, hypoxanthin-9-yl(I), 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthin, 5-bromo, and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hy Droxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, and 5-iodouracil, adenine and guanine 6-alkyl derivatives, 6-azapyrimidine, 6-azouracil, 6-azocytosine, azacytosine, 6-azothymine, 6-thioguanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deazaadenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, and 8-hydroxyl-substituted adenines and guanines;N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, substances that enhance the stability of double helix formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, enlarged nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazinecytidine ([5,4-b][1,4]benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimido[5,4-b][1,4]benzothiadin-2(3H)-one), G-clamps, phenoxazinecytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazine-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindolecytidine (H-pyrimido[3',2':4,5]pyrim Lolo[2,3-d]pyrimidine-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueucine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-; The method according to any one of claims 23 to 35, selected from mannosylqueucine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5oxyacetic acid, weibtoxosin, pseudouracil, queucine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, and 2,6-diaminopurine.
37. Non-natural nucleotides are 【Chemistry 2】 The method according to any one of claims 23 to 35, comprising a non-natural base selected from.
38. Non-natural nucleotides are (a) Modification at position 2': (i) OH; (ii) Substituted lower alkyl, Alkaryl, Aralkyl, O-Alkaryl, O-Aralkyl, SH, SCH 3 , OCN, Cl, Br, CN, CF 3 OCF 3 EN SOCH 3 SO 2 CH 3 ONO 2 NO 2 , N 3 , or NH 2 F; (iii) O-alkyl, S-alkyl, or N-alkyl; (iv) O-alkenyl, S-alkenyl, or N-alkenyl; (v) O-alkynyl, S-alkynyl, or N-alkynyl; or (vi) O-alkyl-O-alkyl, 2'-F, 2'-OCH 3 , or 2'- O(CH 2 ) 2 OCH 3 And, Here, alkyl, alkenyl, and alkynyl are substituted or unsubstituted C 1 ~C 10 Alkyl, C 2 ~C 10 Alkenil, C 2 ~C 10 Alkinyl, -O[(CH 2 )nO]mCH 3 , -O(CH 2 )nOCH 3 , -O(CH 2 )nNH 2 , -O(CH 2 ) nCH 3 , -O(C H 2 )n-ONH 2 , and -O(CH 2 )nON[(CH 2 ) nCH 3 )] 2 And n and m are between 1 and approximately 10. The modifications including; (b) Modifications at the 5' position, including 5'-vinyl and 5'-methyl (R or S); (c) Modifications at the 4' position comprising 4'-S, heterocycloalkyl, heterocycloalkalyl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of oligonucleotides, or group for improving the pharmacodynamic properties of oligonucleotides; and, The method according to any one of claims 23 to 37, further comprising a non-natural sugar portion selected from any combination of (d)(a) to (c).
39. The method according to any one of claims 23 to 38, wherein the manipulated host cells constitutively express or overexpress DNA polymerase II.
40. The method according to any one of claims 23 to 39, wherein DNA polymerase II is encoded by the pollB gene.
41. The method according to claim 40, wherein the polB gene is released from repression.
42. The method according to any one of claims 23 to 41, wherein the manipulated host cell exhibits improved retention of nucleic acid molecules containing non-natural nucleotides during DNA replication, compared to the retention of nucleic acid molecules containing non-natural nucleotides during DNA replication in equivalent manipulated host cells that do not contain the second nucleic acid molecule.
43. The method according to any one of claims 23 to 42, wherein the nucleoside triphosphate transporter exhibits higher stability of expression in the engineered host cell compared to expression in an equivalent engineered host cell that does not contain the second nucleic acid molecule.
Citation Information
Patent Citations
Import of unnatural or modified nucleoside triphosphates into cells via nucleic acid triphosphate transporters
US20170029829A1
Novel nucleoside triphosphate transporter and uses thereof
WO2017223528A1