Tuning chloroplast transgene expression from a synthetic operon

By employing nucleic acids with modified regulatory sequences and PPR protein-binding sites, the expression of peptides and proteins in chloroplasts is optimized, addressing the challenge of controlled expression and enhancing plant performance and protein production.

WO2025106522A1PCT designated stage expired Publication Date: 2025-05-22MASSACHUSETTS INST OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/055688
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-21
Filing Date
2024-11-13
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current methods for expressing peptides and proteins in chloroplasts lack control over expression levels and are not typically achieved in their normal cellular environment.

Method used

The use of nucleic acids with modified or non-natural regulatory sequences, specifically featuring a Shine-Dalgamo sequence separated by a spacer sequence from a translation initiation codon, and operably linked with pentatricopeptide repeat (PPR) protein-binding sequences, to optimize chloroplast expression systems for controlled production of gene products.

Benefits of technology

This approach allows for controllable and enhanced expression of peptides and proteins in chloroplasts, improving plant performance, resistance to pathogens, growth, and production of proteins suitable for various applications, including medical and agronomic uses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024055688_22052025_PF_FP_ABST
    Figure US2024055688_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure relate, at least in part, to nucleic acids for expression of engineered genes in chloroplasts. In some embodiments, an engineered gene comprises a heterologous regulatory sequence (e.g., a promoter, a translation regulatory sequence, and / or RNA-binding proteins binding sequences). In some embodiments, nucleic acids described herein are useful for expressing peptides and / or proteins at levels that are controllable and not normally achieved when said genes are in their normal cellular environment. In some embodiments, nucleic acids described herein are useful in, for example, improving plant performance, providing resistance to plant pathogens, increasing plant growth, and production of agents which are therapeutic to a mammalian subject. In some embodiments, nucleic acids and methods described herein may be used for production of genetically engineered plants and / or seeds (e.g., of an edible crop or crops harvested for plant-based materials used in the manufacture of clothing, construction materials, etc.).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TUNING CHLOROPLAST TRANSGENE EXPRESSION FROM A SYNTHETIC OPERON

[0002] RELATED APPLICATIONS

[0003] The application claims the benefit under 35 U.S.C. 119(e) of U.S. Provisional Application number 63 / 598,489 filed November 13, 2023, and U.S. Provisional Application number 63 / 650,318 filed May 21, 2024, each of which is herein incorporated by reference in its entirety.

[0004] FEDERALLY SPONSORED RESEARCH

[0005] This invention was made with government support under DE-AR0001569 awarded by the U.S. Department of Energy. The government has certain rights in the invention.

[0006] BACKGROUND OF INVENTION

[0007] Expression of nucleic acid constructs in biological systems that are amenable to molecular biology engineering provides means for producing peptides and proteins of interest.

[0008] SUMMARY OF INVENTION

[0009] Chloroplasts are photosynthetic organelles that exist in the cells of plants and some protists. The chloroplast genome is approximately 120-180 kb in size and is characterized by highly polyploid, circular, double- stranded DNA which, in land plants, encodes approximately 120 genes involved in chloroplast development or photosynthesis. Chloroplast genetic systems also comprise unique gene expression systems, which are useful for expression of exogenous genes.

[0010] Tuning of gene expression dynamics in chloroplasts, such as the expression of multigene systems, can involve engineering nucleic acids to comprise modified or non-natural regulatory sequences. Regulatory sequences that achieve high expression levels of exogenous genes via effecting transcription and / or translation rates can optimize chloroplast expression systems for production of gene products that are useful in various agronomic and medical applications.

[0011] Embodiments described herein are useful for expressing peptides and / or proteins in chloroplasts at levels that are controllable and not normally achieved when genes encoding the peptides and / or proteins are comprised in their normal cellular environment. These features can be used, for example, to improve plant performance, plant resistance to pathogens, plant growth, and production of proteins that are suitable for use in any system, such as for administration to a mammalian subject or as an agent in reactions or for research or any other purpose. The claimed compositions, methods, and systems provide a new useful system for producing proteins. Embodiments described herein may also be used for the production of genetically engineered plants and / or seeds (e.g., seeds of an edible crop or crops harvested for plant-based materials used in the manufacture of clothing, construction materials, etc.).

[0012] Aspects of the present disclosure relate to a nucleic acid comprising at least one gene operably linked to a translation initiation regulatory sequence positioned 5' to the at least one gene and / or a pentatricopeptide repeat (PPR) protein-binding sequence positioned 5' to the at least one gene, wherein the translation initiation regulatory sequence comprises a Shine- Dalgamo (SD) sequence and a translation initiation codon which are separated by a spacer sequence, wherein the spacer sequence is a nucleotide sequence of 1-100 nucleotides and is heterologous to the SD sequence.

[0013] In some embodiments the spacer sequence is a sequence which does not comprise a G nucleotide, and / or an AT or AU-rich nucleotide sequence. In some embodiments the spacer sequence comprises the following: 1) nucleotides that do not form a secondary structure, 2) nucleotides that do not form a secondary structure with an upstream mRNA stem-loop secondary structure region optionally which comprises a native or an engineered RNA- specific binding protein binding site, and / or 3) nucleotides free of a canonical or a non-canonical translation initiation codon. In some embodiments the at least one gene is operably linked to the translation initiation regulatory sequence and the PPR protein-binding sequence. In some embodiments the at least one gene comprises a first gene linked to a first PPR protein-binding sequence and / or a first translation initiation regulatory sequence and a second gene linked to a second PPR protein-binding sequence and / or a second translation initiation regulatory sequence. In some embodiments the nucleic acid comprises from 5’ to 3’ the PPR protein-binding sequence, the translation initiation regulatory sequence comprised of the SD sequence, a spacer, and the translation initiation codon and the gene and optionally the 3’ end of the gene is linked to a nucleic acid comprising from 5’ to 3’ a second PPR protein-binding sequence, a second translation initiation regulatory sequence comprised of a second SD sequence, a second spacer, and a second translation initiation codon and a second gene. In some embodiments, a spacer sequence comprises nucleotides that do not form a secondary structure, nucleotides that do not corrupt an upstream mRNA stem-loop secondary structure which comprises a native or an engineered RNA-specific binding protein binding site, nucleotides free of a canonical or a non-canonical translation initiation codon, a sequence which does not comprise a G nucleotide, an AT-rich nucleotide sequence, or any combination thereof. In some embodiments, the at least one gene is operably linked to the translation initiation regulatory sequence and the PPR protein-binding sequence.

[0014] In some embodiments, the at least one gene is operably linked to a promoter. In some embodiments, the promoter is native to the at least one gene. In some embodiments, the promoter is heterologous to the at least one gene. In some embodiments, the at least one gene comprises at least one of: a gene involved in a plant biosynthetic pathway; a gene which is capable of stimulating plant growth; a gene encoding an insecticide; a gene encoding a pesticide; or a gene encoding a peptide or protein which is capable of modifying a molecule produced by a plant pathogen. In some embodiments, the at least one gene comprises a nif cluster gene. In some embodiments, the at least one gene encodes a therapeutic peptide or a therapeutic protein. In some embodiments, the at least one gene encodes a reporter.

[0015] In some embodiments, the at least one gene comprises a first gene and a second gene. In some embodiments, the PPR protein-binding sequence operably linked to the first gene and the PPR protein-binding sequence operably linked to the second gene are the same. In some embodiments, the PPR-protein binding site is a PPRIO-binding site or a PPR10GG-binding site. In other embodiments, the PPR protein-binding sequence operably linked to the first gene and the PPR protein-binding sequence operably linked to the second gene are different. In some embodiments, the PPR-protein binding site operably linked to the first gene is a PPRIO-binding site and the PPR-protein binding site operably linked to the second gene is a PPR 10GG-b inding site. In some embodiments, the spacer sequence of the translation initiation regulatory sequence operably linked to the first gene and the spacer sequence of the translation initiation regulatory sequence operably linked to the second gene are the same. In some embodiments, the spacer sequence of the translation initiation regulatory sequence operably linked to the first gene and the spacer sequence of the translation initiation regulatory sequence operably linked to the second gene differ by at least one nucleotide. In some embodiments, the spacer sequence of the translation initiation regulatory sequence operably linked to the first gene and / or the spacer sequence of the translation initiation regulatory sequence operably linked to the second gene comprises 2-20 nucleotides in length. In some embodiments, the first gene and the second gene are operably linked to a single promoter. In some embodiments, the first gene and the second gene are operably linked to respective promoters.

[0016] Aspects of the present disclosure also relate to a plant cell or plant cell population thereof comprising a nucleic acid described herein. In some embodiments, the nucleic acid is comprised in at least one chloroplast of the plant cell. In some embodiments, the plant cell further comprises a PPR-binding protein. In some embodiments, the plant cell or the plant cell population are that of an edible plant (e.g., a cereal plant). In other aspects, the present disclosure relates to a plant or a seed thereof (e.g., an edible plant, such as a cereal plant, or a seed thereof) comprising a nucleic acid described herein.

[0017] Further aspects of the present disclosure relate to methods. In some embodiments, a method comprises contacting one or more plant cells with the nucleic acid. In some embodiments, the one or more plant cells comprises one or more cells of an edible plant. In some embodiments, a method is used for providing fixed nitrogen to a cereal plant, wherein the method comprises contacting one or more cells of the cereal plant with the nucleic acid. In some embodiments, a method comprises administering a therapeutic peptide or a therapeutic protein of to a mammalian subject in need thereof.

[0018] In addition, the present disclosure provides for a composition comprising a nucleic acid, a plant cell or a plant cell population thereof, and / or a plant or a seed thereof described herein. In some embodiments, the present disclosure relates to a kit comprising a nucleic acid, a plant cell or a plant cell population thereof, a plant or a seed thereof, and / or a composition described herein.

[0019] BRIEF DESCRIPTION OF DRAWINGS

[0020] FIGs. 1A-1J show non-limiting embodiments of nucleic acids. FIG. 1A shows a nonlimiting embodiment of a nucleic acid comprising at least one gene which is operably linked to a heterologous translation initiation regulatory sequence. FIG. IB shows non-limiting embodiments of a translation initiation regulatory sequence comprising a Shine-Dalgamo (SD) sequence and a translation initiation codon which are separated by a spacer sequence comprising a variable number of nucleotides. FIG. 1C shows non-limiting embodiments of a proteinbinding site comprising a pentatricopeptide repeat (PPR) protein-binding site for PPR10 and PPR10GG. FIG. ID shows non-limiting embodiments of nucleic acids comprising at least one gene which is operably linked to one or more heterologous regulatory sequences. FIG. IE shows non-limiting embodiments of nucleic acids comprising at least one gene which is operably linked to a heterologous regulatory sequence and a promoter, wherein the promoter is either native or heterologous to the at least one gene. FIG. IF shows a non-limiting embodiment of a nucleic acid comprising two genes which are each operably linked to a heterologous regulatory sequence. FIG. 1G shows non-limiting embodiments of nucleic acids comprising two genes which are each operably linked to a plurality of regulatory sequences. FIG. 1H shows non-limiting embodiments of nucleic acids comprising two genes which are each operably linked to a plurality of regulatory sequences, including promoters, translation initiation regulatory sequences, and / or PPR protein-binding sites. FIG. II shows non-limiting embodiments of nucleic acids comprising a plurality of genes, wherein each gene is operably linked to a native PPR protein binding sequence and a translation initiation regulatory sequence comprising a spacer sequence of variable lengths which separates an SD sequence from a translation initiation codon. FIG. 1J shows non-limiting embodiments of nucleic acids comprising a plurality of genes, wherein each gene is operably linked to an engineered PPR protein binding sequence and a translation initiation regulatory sequence comprising a spacer sequence of variable lengths which separates an SD sequence from a translation initiation codon.

[0021] FIG. 2 shows representative results from analyses of GFP expression in tobacco plant chloroplasts which was tuned by adjusting the length of a spacer sequence that was positioned between an SD sequence and a translation initiation codon operably linked to the GFP gene.

[0022] FIG. 3 shows representative results from analyses of GFP expression in tobacco plant chloroplasts, wherein the gene encoding GFP was operably linked to a PPR10 binding site which was positioned 5' relative to the GFP coding sequence.

[0023] FIG. 4 shows representative results from analyses of GFP expression in tobacco plant chloroplasts, wherein the gene encoding GFP was operably linked to a PPR10GGbinding site which was positioned 5' relative to the GFP coding sequence.

[0024] FIGs. 5A-5E show representative results from analyses of chloroplasts engineered to comprise nif genes linked to a PPR10GGprotein binding site. FIG. 5A is a diagram of nitrogenase formation involving the nifS, nifU, nifM, and nifH genes. FIG. 5B is a diagram of a genetic engineering strategy in plant cells, wherein nifS, nifU, nifM, and nifH genes are integrated into the chloroplast genome and each engineered to comprise a 5' UTR that comprises a binding site for a PPR10GGprotein encoded by the nuclear genome of the plant cell. FIG. 5C shows images of wild-type plants (plants in left-most column), plants having chloroplasts with genomes that comprise integrated nifS, nifU, nifM, and nifH genes comprising a binding site for a PPR10GGprotein (atpHGG) in their 5' UTR (plants in middle column), and plants having chloroplasts with genomes that comprise integrated nifS, nifU, nifM, and nifH genes comprising a binding site for a PPR10GGprotein (atpHGG) encoded by the nuclear genome in their 5' UTR PPR10GGprotein (plants in right-most column). FIG. 5D shows results from Southern blot analyses of samples from two plants corresponding to the right-most column of plants in FIG. 5C and indicates engineered nif genes are integrated into the chloroplast genome. FIG. 5E shows results from Northern blot analyses of samples from two plants corresponding to the right-most column of plants in FIG. 5C and performed using probes for detecting RNA encoded by an aadA gene (a selectable marker), a nifU gene, a nifS gene, and a nifM gene.

[0025] DETAILED DESCRIPTION OF INVENTION

[0026] Aspects of the present disclosure relate, at least in part, to nucleic acids for expression of engineered genes in chloroplasts. Non-limiting embodiments of nucleic acids are shown in FIGs. 1A-1J. In some embodiments, a nucleic acid comprises at least one gene which is operably linked to a regulatory sequence (see, e.g., FIG. 1A). In some embodiments, a regulatory sequence is a translational regulatory sequence. In some embodiments, a translational regulatory sequence is engineered to comprise a spacer sequence is positioned between a Shine-Dalgarno sequence and a translation initiation codon. In some embodiments, a spacer sequence is engineered to comprise a length which is useful for upregulating or downregulating expression of a coding sequence to which it is operably linked (see, e.g., FIG. IB). In some embodiments, a regulatory sequence comprises an RNA-binding protein binding sequence. In some embodiments, the RNA-binding protein is a pentatricopeptide repeat (PPR) protein, such as one that regulates the expression of a nucleic acid upon binding to RNA comprising a PPR proteinbinding sequence (see, e.g., FIG. 1C). In some embodiments, a nucleic acid comprises a plurality of regulatory sequences, such as a promoter, a translation regulatory sequence, and / or a PPR protein-binding sequence (see, e.g., FIGs. ID- IE). In some embodiments, nucleic acids described herein comprise a plurality of genes which are each operably linked to regulatory sequences described herein (see, e.g., FIG. 1F-1J). In some embodiments, nucleic acids described herein are useful for transferring single genes or multi-gene systems to chloroplasts for the purposes of expressing peptides and / or proteins. In some embodiments, a nucleic acid is capable expressing peptides and / or proteins at levels in chloroplast expression systems that are controllable and not normally achieved when said genes are comprised in their normal cellular environment. In some embodiments, increased or decreased expression levels of a gene or a plurality thereof is achieved using one or more regulatory sequences that are heterologous to the gene(s) to which they are operably linked (see, e.g., FIGs. 2-4).

[0027] In some embodiments, nucleic acids comprise genes encoding peptides and / or proteins that are useful in, for example, improving plant performance, plant resistance to pathogens, plant growth, and production of agents that are suitable for administration to a mammalian subject. In some embodiments, cells and cell population comprising nucleic acids described herein may be used for production of genetically engineered plants and / or seeds (or seedlings thereof). In some embodiments, a genetically engineered plant or seed described herein is that of an edible crop. In some embodiments, genetically engineered plants comprising a nucleic acid described herein may be planted and harvested for the purposes of food production or production of plant-based materials used in the manufacture of clothing, construction materials, etc.

[0028] Further aspects of the present disclosure relate to methods of introducing nucleic acids into cells (e.g., plant cells). In some embodiments, a method comprises characterizing the expression a gene encoded by a nucleic acid using one or more detection methods described herein. In some embodiments, a method comprises isolating and / or purifying a peptide or protein encoded by a nucleic acid. In some embodiments, the peptide or protein may be administered to a subject, such as in a method comprising administration of a pharmaceutical composition to a subject. However, in other embodiments, plant cells comprising a nucleic acid described herein may be subjected to conditions capable of giving rise to a plant or a seed (or seedling thereof). In some embodiments, a plant and / or a seed may be formulated in an agricultural composition for delivery to soil, such as soil suitable for germinating a seed and / or cultivating a plant. Moreover, embodiments of the present disclosure further relate to kits which may be useful in the production of nucleic acids, cells, and / or genetically engineered plants or seeds described herein.

[0029] Nucleic Acids In some embodiments, a nucleic acid comprises DNA. In some embodiments, a nucleic acid comprises RNA. In some embodiments, a nucleic acid comprises a combination of DNA and RNA. In some embodiments, a nucleic acid comprises double- stranded DNA. In some embodiments, a nucleic acid comprises single-stranded DNA. In some embodiments, a nucleic acid comprises double-stranded RNA. In some embodiments, a nucleic acid comprises singlestranded RNA. In some embodiments, a nucleic acid comprises a combination of DNA and RNA nucleotides. In some embodiments, a nucleic acid comprises a double- stranded nucleic acid (e.g., DNA, RNA, or a combination of DNA / RNA) comprised of one or more gaps (e.g., stretches of sequence about, 1, 5, 10, 25, 50, 100 or more nucleotides in length) of singlestranded sequence. In some embodiments, a nucleic acid is linear. In some embodiments, a nucleic acid is circular. In some embodiments, a nucleic acid comprises one or more singlestranded nicks.

[0030] Non-limiting examples of nucleic acids include regulatory sequences (e.g., transcription regulatory sequences, such as a transcription initiation sequences and transcription termination sequences, translation regulatory sequences, such as translation initiation sequences and translation termination regulatory sequences, protein-binding sites, promoters, etc.), genes (e.g., protein-coding genes and / or genes encoding RNAs that are not translated, such as transgenes, genes comprised in a vector, chromosome, or genome, etc.), operons (e.g., monocistronic operons and polycistronic operons), expression cassettes, genetic clusters (e.g., a nif cluster), vectors, chromosomes (e.g., linear chromosomes or circular chromosomes), genomes (e.g., viral genomes, such as bacteriophage genomes, bacterial genomes, mammalian cell genomes, plant cell genomes, such as a nuclear genome, a mitochondrial genome, or a chloroplast genome, etc.), and nucleic acids comprising any combination of two or more (e.g., two, three, four, five, or more than five) of the foregoing nucleic acids. In some embodiments, a nucleic acid is an endogenous nucleic acid (e.g., a wildtype gene in its natural host genome). In some embodiments, a nucleic acid is an exogenous nucleic acid (e.g., a nucleic acid which has been introduced into a cell, such as nucleic acid comprising a sequence which is heterologous to the cell it is comprised in).

[0031] In some embodiments, a nucleic acid comprises an operon comprising at least one gene. In some embodiments, a nucleic acid comprises an operon comprising a plurality of genes. In some embodiments, a nucleic acid comprises a plurality of operons (e.g., 2, 3, 4, 5, or more operons) which each comprise at least one gene. In some embodiments, an operon is monocistronic (e.g., an operon comprising one or more genes which are operably linked to a single promoter). In some embodiments, an operon is a polycistronic operon (e.g., an operon comprising a plurality of genes, wherein at least of the genes in the plurality are operably linked to respective promoters).

[0032] In some embodiments, a nucleic acid is a vector. In some embodiments, a vector is a plasmid, such as a circular plasmid, a nanoplasmid, or a minicircle plasmid. In some embodiments, a vector is a linear nucleic acid. In some embodiments, vectors are self- complementary nucleic acids. In some embodiments, a vector is a cosmid or an artificial chromosome. In some embodiments, a vector may comprise a sequence which is capable of stably maintaining the vector in a cell, such as an antibiotic resistance gene, a partitioning system, and / or a sequence which promotes integration of the vector into a genome. In some embodiments, a vector is a high copy number vector.

[0033] In some embodiments, a nucleic acid is an engineered nucleic acid which is produced in a non-natural setting (e.g., in a laboratory or a heterologous host cell which modifies the structure of the nucleic acid), substantially purified from its native host cell, and / or structurally altered using one or more engineering methods (e.g., amplified in vitro by polymerase chain reaction (PCR), recombinant cloning, chemically synthesized, etc.). In some embodiments, an engineered nucleic acid comprises naturally-occurring components (e.g., naturally-occurring gene sequences) which are derived from one or more genetic sources and have been structurally altered.

[0034] In some embodiments, a nucleic acid (e.g., an engineered nucleic acid) comprises at least two nucleotides in length. In some embodiments, a nucleic acid comprises more than 100, 500, 1,000, 5,000, 10,000, or 20,000 nucleotides in length. In some embodiments, a nucleic acid comprises approximately 1-60,000 nucleotides. In some embodiments, a nucleic acid comprises approximately 1-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1,000, 1,000-1,100, 1,100-1,200, 1,200-1,300, 1,300-1,400, 1,400-1,500, 1,500-1,600, 1,600-1,700, 1,700-1,800, 1,800-1,900, 1,900-2,000, 2,000-2,100, 2,100-2,200, 2,200-2,300, 2,300-2,400, 2,400-2,500, 2,500-2,600, 2,600-2,700, 2,700-2,800, 2,800-2,900, 2,900-3,000, 3,000-3,100, 3,100-3,200, 3,200-3,300, 3,300-3,400, 3,400-3,500, 3,500-3,600, 3,600-3,700, 3,700-3,800, 3,900-4,000, 4,000-4,100, 4,100-4,200, 4,200-4,300, 4,300-4,400, 4,400-4,500, 4,500-4,600, 4,600-4,700, 4,700-4,800, 4,800-4,900, 4,900-5,000, 5,000-5,100, 5,100-5,200, 5,200-5,300, 5,300-5,400, 5,400-5,500, 5,500-5,600, 5,600-5,700, 5,700-5,800, 5,800-5,900, 5,900-6,000, 6,000-6,100, 6,100-6,300, 6,200-6,300, 6,300-6,400, 6,400-6,500, 6,500-6,600, 6,600-6,700, 6,700-6,800, 6,800-6,900, 6,900-7,000, 7,000-7,100, 7,100-7,200, 7,200-7,300, 7,300-7,400, 7,400-7,500, 7,500-7,600, 7,600-7,700, 7,700-7,800, 7,800-7,900, 7,900-8,000, 8,000-8,100, 8,100-8,200, 8,200-8,300, 8,300-8,400, 8,400-8,500, 8,500-8,600, 8,600-8,700, 8,700-8,800, 8,800-8,900, 8,900-9,000, 9,000-9,100, 9,100-9,200, 9,200-9,300, 9,300-9,400, 9,400-9,500, 9,500-9,600, 9,600-9,700, 9,700-9,800, 9,800-9,900, 9,900-10,000, 10,000-15,000, 15,000-20,000, 20,000- 30,000, 30,000-40,000, 40,000-50,000, or 50,000-60,000 nucleotides in length. However, in some embodiments, a nucleic acid comprises more than 60,000 nucleotides in length (e.g., 60,000-70,000, 70,000-80,000, 80,000-90,000, 90,000-100,000, or more than 100,000 nucleotides).

[0035] Genes and Gene Products Thereof

[0036] In some embodiments, a nucleic acid comprises at least one gene. The term “gene” may be used herein to refer to a nucleic acid sequence which can be transcribed to produce an RNA. Non limiting examples of RNAs include small-hairpin RNAs (shRNAs), short-interfering RNAs (siRNAs), prokaryotic -interfering RNAs (pro-siRNAs), micro-RNAs (miRNAs), long noncoding RNAs (IncRNAs), Piwi-interacting RNAs (piRNAs), exon-skipping RNAs, enzymatic RNAs, guide RNAs ((gRNAs), e.g., single-guide RNAs (sgRNAs)), small nuclear RNAs (snRNAs), small nucleolar RNAs (snoRNAs), ribosomal RNAs (rRNAs), transfer RNAs (tRNAs), and messenger RNAs (mRNAs). In some embodiments, a gene comprises a sequence encoding an mRNA.

[0037] In some embodiments, a gene described herein is a transgene. The term “transgene” may be used herein to refer to a nucleic acid which can be transcribed to produce an RNA (e.g., an mRNA) and comprises a heterologous nucleic acid sequence, is operably linked to a heterologous nucleic acid sequence, and / or is not normally present in a cell, a genome of a cell, or a compartment of a cell (e.g., an organelle, such as a chloroplast) that the transgene is located in. In some embodiments, a transgene comprises a sequence and / or is operably linked to a sequence which is heterologous relative to the native version of the gene (e.g., a coding sequence which has been codon optimized and / or a gene which is operably linked to a heterologous translation regulatory sequence). In some embodiments, a transgene may be a nuclear gene that has been introduced into a different organelle, such as a chloroplast. In some embodiments, a transgene is a gene which is heterologous to the cell that the transgene is located in (e.g., a gene from one plant species that has been introduced into a cell of a different plant species or a gene from a bacterium that has been introduced in to a plant cell).

[0038] In some embodiments, a gene comprises a sequence encoding an RNA (e.g., an mRNA), wherein the gene is operably linked to one or more regulatory sequences. In some embodiments, a gene comprises a sequence encoding an RNA (e.g., an mRNA), wherein the RNA is operably linked to one or more regulatory sequences and / or comprises one or more regulatory sequences. In some embodiments, an mRNA encoded by a gene comprises one or more regulatory sequences which are capable of promoting the translation of the mRNA into a peptide or protein. In some embodiments, a nucleic acid comprises a plurality of genes (e.g., two, three, four, five, or more than five genes, such as wherein one or more of the genes are operably linked to a regulatory sequence and / or one or more of the genes encode an RNA which can include, but is not limited to, an mRNA). In some embodiments, a nucleic acid comprises two genes. In some embodiments, a nucleic acid comprises three genes. In some embodiments, a nucleic acid comprises four genes. In some embodiments, a nucleic acid comprises five genes. In some embodiments, a nucleic acid comprises six genes. In some embodiments, a nucleic acid comprises seven genes. In some embodiments, a nucleic acid comprises eight genes. In some embodiments, a nucleic acid comprises nine genes. In some embodiments, a nucleic acid comprises ten genes. In some embodiments, a nucleic acid comprises more than ten genes (e.g., eleven genes, twelve genes, thirteen genes, fourteen genes, fifteen genes, or more than fifteen genes). In some embodiments, a nucleic acid comprises at least one gene or a plurality of genes which are amenable for chloroplast transformation (e.g., wherein the nucleic acid comprises one or more regulatory sequences that promote gene expression in a chloroplast).

[0039] In some embodiments, a nucleic comprises at least one gene encoding an amino acid sequence of a peptide or protein. In some embodiments, peptides and / or proteins comprise naturally occurring amino acids (e.g., asparagine (N), glutamine (Q), serine (S), threonine (T), cysteine (C), aspartate (D), glutamine (Q), arginine (R), lysine (K), histidine (H), asparagine (N) glutamine (G), glycine (G), alanine (A), valine (V), leucine (L), methionine (M), isoleucine (I), phenylalanine (F), tyrosine (Y), tryptophan (Y), proline (P) tryptophan (W), phenylalanine (F), and tyrosine (Y)). However, in other embodiments, peptides and / or proteins comprise modified amino acids, such as those modified (e.g., post-translationally in a cell, such as in a chloroplast, or by chemical synthesis) to comprise a carbohydrate group, a hydroxyl group, a phosphate group, a famesyl group, an isofamesyl group, a fatty acid group, a linker for conjugation or functionalization, or other modification.

[0040] In some embodiments, a peptide comprises approximately 50 amino acids or less in length (e.g., 1-10 amino acids, 10-20 amino acids, etc.). In some embodiments, a protein comprises approximately 1,500 amino acids or less in length (e.g., 100-250 amino acids, 250- 500 amino acids, 500-750 amino acids, 750 amino acids, etc.). In other embodiments, a protein comprises more than 1,500 amino acids in length (e.g., 2,000 amino acids, 3,000 amino acids, 4,000 amino acids, 5,000 amino acids, 7,500 amino acids, 10,000 amino acids, etc.). In some embodiments, a protein comprises one polypeptide chain (e.g., one subunit) or a plurality of polypeptide chains (e.g., a dimer, a trimer, a tretramer, etc.).

[0041] In some embodiments, a gene comprises a sequence encoding an amino acid sequence comprising a molecular weight of approximately 200 g / mol or more. In some embodiments, a gene comprises a sequence encoding an amino acid sequence comprising a molecular weight of 500-1,000, 1000,-2,000, 2,000-3,000, 3,000-4,000, 4,000-5,000, 5,000-6,000, 6,000-7,000, 7,000-8,000, 8,000-9,000, or 9,000-10,000 g / mol. In some embodiments, a gene comprises a sequence encoding an amino acid sequence comprising a molecular weight of 10,000-12,000, 12,000-14,000, 14,000-16,000, 16,000-17,000, 17,000-18,000, 18,000-20,000, 20,000-25,000, 25,000-30,000, 30,000-40,000, 40,000-50,000, 50,000-60,000, 60,000-70,000, 70,000-80,000, 80,000-90,000, 90,000-100,000, 100,000-110,000, 110,000-120,000, 120,000-130,000, 130,000- 140,000, 140,000-150,000, 150,000-160,000, 160,000-170,000, 170,000-180,000, 180,000- 190,000, 200,000-220,000, 220,000-240,000, 240,000-260,000, 260,000-280,000, 280,000- 300,000, 300,000-350,000, 350,000-400,000, 400,000-450,000, 450,000-500,000, 500,000- 600,000, 600,000-700,000, 700,000-800,000, 800,000-900,000, 900,000-1,000,000, 1,000,000- 2,500,000, or 2,500,000-5,000,000 g / mol. In some embodiments, a gene comprises a sequence encoding an amino acid sequence comprising a molecular weight of approximately 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 125,000, or 150,000, g / mol. In some embodiments, a gene comprises a sequence encoding an amino acid sequence comprising a molecular weight that is greater than 5,000,000 g / mol.

[0042] The nucleic acid may encode one or more proteins. For instance, the nucleic acid may be comprised of a single gene encoding a single protein or may be a multi-gene system, encoding multiple proteins. Any protein may be encoded by the nucleic acids disclosed herein. For instance, a protein may be a protein that has a utility in a cell (e.g., a protein that functions in a plant cell) or chloroplast (e.g., a protein that functions in a chloroplast) or it may be a protein that has other uses (e.g., a protein that can be produced in a first cell, such as a plant cell, and then contacted with a second cell, such as a different plant cell or a mammalian cell including, but not limited to, a human cell), such as a therapeutic agent (e.g., an agent that is therapeutic in plants or mammals, such as humans). In some embodiments the protein is a reporter protein. In other embodiments the protein is a protein involved in plant functions, such as nitrogen regulation. In some embodiments the protein is an insecticide and or pesticide (optionally encoded by a multi-gene system). In some embodiments the protein may be a biofortification product, such as vitamins and other nutritional products that can increase crop nutritional content (optionally part of multi-gene system). In some embodiments the protein may be a protein which is generated in the chloroplast system disclosed herein and isolated for other purposes, such as therapeutic proteins.

[0043] Non-limiting examples of peptides and proteins include a cell-, organelle-, or tissuetargeting peptide or protein, a cell-penetrating peptide or protein, an antibody (e.g., a monoclonal antibody, a polyclonal antibody, a nanobody, a single-chain antibody, such as an scFv, etc.), an antigen-binding fragment, an antigenic peptide or protein, an enzyme or an enzymatic domain (e.g., a protease, signaling protein, polymerase, transcriptional regulator, nucleases, RNA-guided nuclease, metabolic enzymes, nitrogenase, kinases, phosphatases, lipidtransferases, glycosylases, DNA ligases, ubiquitin ligases, methyltransferases, acetyltransferases, SUMO transferases, proteases, foldases, reductases, lyases, dehydrogenases, phosphorylases, decarobxylases, dephosphorylases, kinases, transferases, synthases, and hydrolases, etc.), a proteinaceous enzyme substrate, a glycoprotein, a lipoprotein, a secreted protein, a spore coat protein, an extracellular matrix protein or a fragment thereof, a viral coat protein, a transmembrane receptor or a fragment thereof, a toxin or a fragment thereof, hormones, receptors, a peptibody, a growth factor, a clotting factor, a cytokine, a chemokine, an activating or inhibitory peptide (e.g., a peptide capable of targeting a cell surface receptor or ion channel), a thrombolytic, a bone morphogenetic protein, an Fc-fusion protein, an anticoagulant, a nucleic acid-binding protein (e.g., transcription regulators and / or translation regulators, such as a pentatricopeptide repeat (PPR) protein described herein), and a detectable marker or a reporter (e.g., cell surface proteins, such as an antibody or antigen-binding fragment thereof, receptors, membrane proteins which become glycosylated upon expression in a cell, etc., a protein which produces a plant pigment, such as chlorophylla, carotenoids, etc., fluorescent molecules, such as Mko2, mBeRFP, mNeonGreen, GFP, EGFP, Superfold GFP, Azami Green, mWasabi, TagGFP, TurboGFP, acGFP, zsGreen, T-sapphire, EBFP, EBFP2, Azurite, mTagBFP, ECFP, Mecfp, Cerulean, mTurquoise, CyPet, AmCyanl, TagCFP, Mtfpl, EYFP, mCitrine, TagYFP, phiYFP, zsYellowl, mBanana, Kusabira Orange, mOrange, dTomato, DsRed, mTangerine, mRuby, mApple, mStrawberry, AsRed2, Mrfpl, mCherry, HcRedl, Irfp720, smURFP, and AQ143, etc.).

[0044] In some embodiments, a gene comprises a sequence encoding a therapeutic peptide or protein. The terms “therapeutic peptide” and “therapeutic protein” refers to a molecule that leads to a physiological change in a cell which may improve a biological process in a cell and / or the organism comprising the cell. In some embodiments, a therapeutic peptide or a therapeutic protein is associated with or expected to at least partially, if not fully, prevent and / or remedy at least one symptom associated with a disease, disorder, or condition which effects plants or mammals (e.g., humans).

[0045] In some embodiments, a peptide or protein which is therapeutic for a plant may one which improves plant function in a given environment (e.g., an agricultural environment). In some embodiments, a peptide or protein which is therapeutic for a plant may be one which is capable of preventing a plant disease, such as a disease caused by a plant pathogen. In some embodiments, a peptide or protein which is therapeutic for a plant may prevent, counteract, or reduce the physiological effects of a nutrient deficiency (e.g., a deficiency arising from being planted in soil that is low in a nutrient, a deficiency that could be caused due to a reduced ability to take up a nutrient, and / or a deficiency that could be caused due to a reduced ability to use a nutrient in a biological process).

[0046] In some embodiments, a peptide or protein which is therapeutic for a mammalian subject may be one which is capable of making a subject resistant to a mammalian pathogen, thus preventing the subject from becoming infected or reducing the severity of one or more symptoms associated with the infection if the subject is exposed to the pathogen. In some embodiments, a peptide or protein which is therapeutic for a mammalian subject may be one that is capable of treating a disease, disorder, or condition.

[0047] Plant Biofortification Genes and Plant Growth-Stimulating Genes In some embodiments, a nucleic acid comprises one or more genes encoding components of a metabolic pathway. In some embodiments, a nucleic acid comprises a full set of genes or a subset of genes (e.g., a group of genes that complements a set of genes provided in a cell and / or on a separate nucleic acid) necessary to form a metabolic pathway. In some embodiments, a metabolic pathway is a biosynthetic pathway or a catabolic pathway.

[0048] In some embodiments, a metabolic pathway is involved in the synthesis or the breakdown of a co-enzyme or a co-factor which is involved in a cellular pathway. In some embodiments, a metabolic pathway is involved in the product of a vitamin or other nutrient. Non-limiting examples of vitamins include vitamin A (carotenoids, such as retinol), vitamin C (ascorbic acid), vitamin E (tocopherols and tocotrienols), vitamin K (phylloquinone), B vitamins (e.g., thiamine, riboflavin, niacin, folate, etc.), and vitamin H (biotin). Non-limiting examples of genes involved in the production of vitamin A include those encoding phytoene synthase, phytoene desaturase, zeta-carotene desaturase, carotenoid isomerase, and beta-carotene hydroxylase. Non-limiting examples of genes involved in the production of vitamin C include those encoding GDP-mannose 3,5-epimerase, L-galactose-1 -phosphate phosphatase, L- galactose-1 -dehydrogenase, and L-galactono-l,4-lactone dehydrogenase. Non-limiting examples of genes involved in the production of vitamin E include those encoding homogentisate phytyltransferase, tocopherol cyclase, and tocopherol methyltransferase. Non-limiting examples of genes involved in the production of vitamin K include MenA (l,4-dihydroxy-2-naphthoate octaprenyltransferase), MenB (l,4-dihydroxy-2-naphthoate polyprenyltransferase), MenC (O- succinylbenzoate synthase), MenD (2- succinyl-6-hydroxy-2,4-cyclohexadiene-l -carboxylate synthase), MenE (2-succinyl-6-hydroxy-2,4-cyclohexadiene-l -carboxylate 2,3-dehydrogenase), and other proteins involved in the conversion of chorismate to menaquinone. Non-limiting examples of genes involved in the production of B vitamins include those encoding thiamine synthase, riboflavin synthase, tryptophan synthase, tryptophan 2-monooxygenase, kynurenine formamidase, kynurenine aminotransferase, kynureninase, phosphoribosylanthranilate isomerase, niacin synthase, pantothenate synthetase, and pyridoxine synthase. Non-limiting examples of genes involved in the production of biotin include those encoding biotin synthase, dethiobiotin synthetase, pimeloyl-CoA synthase, and dethiobiotin deaminase.

[0049] In some embodiments, a nucleic acid comprises one or more genes that encode components of a pathway which produce molecules capable of stimulating plant growth (e.g., a phytohormone). Non-limiting examples of molecules which are capable of stimulating plant growth include auxins, cytokinins, gibberellins, brassinosteroids, abscisic acid, ethylene, jasmonic acid, salicylic acid, strigolactones, polyamines, nitric oxide (NO), peptides (e.g., CLE peptides, expansins, root growth factor 1, RALF peptides, ENOD40, phytosulfokines, systemin, and feronia). In some embodiments, a nucleic acid comprises one or genes involved in the production of molecules which are capable of stimulating plant growth, such as tryptophan aminotransferase-related genes, flavin monooxygenase-like genes, isopentenyltransferase genes, a GA20ox gene, a GA3ox gene, a DET2 gene, BR6ox genes, 9-cis-epoxycarotenoid dioxygenase genes, ACC synthase genes, ACC oxidase genes, lipoxygenase genes, phenylamine ammonia-lyase genes, carotgenoid cleavage dioxygenase genes, nitric oxide synthase genes, EXPA, EXPB, RGF1, CLAVAT3, RALF, PSK genes, prosystemin genes, and the feronia genes.

[0050] In some embodiments, a nucleic acid comprises one or more nitrogen fixation genes. The term “nitrogen fixation gene” may be used herein to refer to a gene which, upon expression, influences the conversion of atmospheric nitrogen (N2) into other usable nitrogenous molecules (e.g., nitrogen in the form of ammonia, nitrites, nitrates, etc.) in bacteria. This term applies to both natural forms (e.g., genes with wildtype coding sequences operably linked to native regulatory sequences, such as a wildtype nif cluster gene from a natural nif cluster of a Rhizobi ) and non-natural forms thereof (e.g., those that are engineered, such as a codon-optimized nif gene in a refactored nif cluster). Accordingly, this term encompasses molecules that influence nitrogen fixation either directly (e.g., a nitrogenase) or indirectly (e.g., a transcription factor or signaling protein which regulates a nitrogenase).

[0051] In some embodiments, nitrogen fixation genes comprise one or more nif cluster genes. Non-limiting examples of nif cluster genes include nif A (encodes a nitrogenase-specific transcriptional activator protein (Nif A), responsible for regulating the expression of other nif genes), nifB (encodes a protein involved in the synthesis of the iron-molybdenum cofactor, essential for the activation of nitrogenase), nifD (encodes one of the two subunits of the nitrogenase molybdenum-iron (MoFe) protein, which actively participates in the reduction of nitrogen to ammonia), nifE (encodes a protein involved in the synthesis of FeMo-co (ironmolybdenum cofactor), a critical element in the nitrogenase MoFe protein), nifH (encodes the nitrogenase iron protein (Fe protein or dinitrogenase reductase), playing a pivotal role in electron transfer within the nitrogenase complex), nifK (encodes the other subunit of the nitrogenase MoFe protein, working in conjunction with NifD to facilitate nitrogen fixation), nifL (encodes for a nitrogenase-specific regulatory protein (NifL), which collaborates with NifA to modulate nitrogen fixation in response to varying oxygen levels), nifM (regulates nitrogenase activity), nifN (encodes a protein involved in the synthesis of FeMo-co, contributing to the structural integrity of the nitrogenase complex), nifQ (encodes a protein involved in the synthesis of the iron-molybdenum cofactor, contributing to the overall functionality of nitrogenase), nifS (encodes a protein involved in the synthesis of the iron-molybdenum cofactor, crucial for the proper functioning of nitrogenase), nifU (regulates nitrogenase activity), ni / V (encodes a protein involved in the synthesis of the iron-molybdenum cofactor, contributing to the maturation of functional nitrogenase), nifW (encodes a protein involved in the synthesis of the iron-molybdenum cofactor, playing a role in nitrogenase activation), and nifX (encodes a protein that may be involved in the assembly or stabilization of the nitrogenase complex, contributing to its overall structural organization).

[0052] In some embodiments, a nif cluster gene is derived from a free-living diazotroph nif cluster, a symbiotic diazotroph nif cluster, a photosynthetic gammaproteobacterial nif cluster, a gammaproteobacterial nif cluster, a cyanobacteria nif cluster, or a firmicutes nif cluster. In some embodiments, a nif cluster gene is derived from a nif cluster from Cyanthoece sp., P. polymyxa, K. oxytoca, A. vinelandii, P. stutzeri, A. bras dense, R. paluslris, R. sphaeroides, G. diazotrophicus, or R. palustris.

[0053] In some embodiments, a nucleic acid comprises one or more refactored nif cluster genes or comprises a refactored nif cluster. The term “refactored nif cluster” refers to a nucleic acid comprising a set of nif cluster genes which are sufficient for nitrogen fixation but have been structurally altered to exhibit expression properties that are not exhibited by native nif clusters. In some embodiments, a refactored nif cluster comprises engineered coding sequences not naturally found in the corresponding native nif cluster. For example, in some embodiments, a refactored nif cluster comprises variant nif genes, such as those that have been codon optimized based on which host cell will be engineered to comprise said refactored nif cluster. Further embodiments nucleic acids comprising nif cluster genes and refactored nif clusters are described in the art, such as U.S. Patent No.: US 11,479,516, US 2020 / 0299637, and Ryu et al. (2020). Control of nitrogen fixation in bacteria that associate with cereals. Nat. Microbiol. 5(2): 314- 330, which are incorporated by reference herein for their disclosures regarding nif cluster structure, refactoring, and methods related to the same.

[0054] Genes for Plant Defense Systems In some embodiments, expression of a nucleic acid is capable of making a plant cell resistant to a pathogen. In some embodiments, a nucleic acid comprises a sequence encoding a peptide or protein (e.g., an enzyme or an enzymatic domain) which is capable of modifying and / or degrading a molecule (e.g., a nucleic acid, peptide, or protein) produced by a pathogen, thus providing a form of immunity to a cell (e.g., a mammalian cell or a plant cell) against the pathogen. In some embodiments, a pathogen is a virus, a bacterium, a fungus, or a parasite. In some embodiments, a pathogen is an agricultural pathogen for example, a fungi belonging to the phylas Ascomycota such as Fusarium spp., Thielaviopsis spp., Verticullium spp., Magnaporthe grisea, Sclerotinia sclerotiorum or Basidiomycota such as Ustilago spp., Rhizoctonia spp., Phakospora pachyrhizi, Puccinia spp., or Armillaria spp., a bacteria such as Burkholderia, xanthomonas spp., Pseudomonas spp., Phytoplasma, Spiroplasma, other organisms such as nematodes, Phytomonas, Cephaleuros, broomrape, mistletoe, dodder, Pythium spp., Phytophthora spp., Plasmodiophora, or Spongospora. In some embodiments, a pathogen is associated with crop diseases, such as banana bunchy top, black bunchy top, black sigatoka, Panama disease, Fusarium head blight, powdery, mildew, barley stem rust, African cassava mosaic disease, bacteria blight, cassava brown streak disease, bacterial blight, Fusarium wilt, Verticillium wilt, Aspergillus ear rot, Giberella stalk and ear rot, grey leaf spot, basal stem rot, bud rot, groundnut rosette disease, potato brown rot, late blight, Phoma stem canker, Sclerotinia stem rot, rice clast, rice bacterial blight, sheath blight, Anthracnose, Turcicum leaf blight, soybean cyst nematode disease, Asian soybean rust, Cercospora leaf spot, rhizomania, Ratoon stuntling, red rot, sweet potato virus disease, late blight, tomato yellow leaf curl, Fusarium head blight, wheat stem rust, wheat yellow rust, anthracnose, or yam mosaic disease.

[0055] In some embodiments, a nucleic acid comprises one or more genes encoding an insecticide or pesticide. Non-limiting examples of insecticides and pesticides include Bacillus thurigiensis toxins, neuropeptides (e.g., allatostatins, tachykinins, and myosuppressins), ion channel modulators (e.g., agents derived from venoms, such as snake venoms and spider venoms), lectins, plant-derived protease inhibitors, plant-derived proteinase inhibitors, and antimicrobial peptides.

[0056] In some embodiments, a nucleic acid comprises one or more genes that are capable of defending a plant against a competitor, a plant pathogen, or a predator. In some embodiments, a nucleic acid may encode one or more genes that produce a pesticide against a plant competitor, such as a weed (e.g., grass weeds, broadleaf weeds, perennial weeds, etc.) and / or a sedge (e.g., yellow nutsedge). In some embodiments, the one or more genes produce allelopathic compounds, essential oils, and / or molecules which promote endophytic colonization of a bacteria or fungus that is pathogenic to the plant competitor. In some embodiments, a nucleic acid comprises one or more genes that produce a molecule which is toxic and / or a deterrent against an herbivore (e.g., alkaloids, such as alkaloids produced by an endophytic bacteria).

[0057] Genes for Production of Peptides or Proteins for Administration to Mammalian Subjects

[0058] In some embodiments, a gene comprises a sequence encoding a peptide or protein which is selected for the purposes of vaccine production against a pathogen. In some embodiments, a gene comprises a sequence encoding a peptide or a protein which is an immunogenic protein or an immunogenic fragment thereof (e.g., an immunogenic peptide) comprising an antigen of a pathogen. In some embodiments, the pathogen is pathogenic (e.g., infects and / or causes one or more symptoms of a disease, disorder, or condition) to mammals (e.g., humans). In some embodiments, an immunogenic protein or an immunogenic fragment thereof (e.g., an immunogenic peptide) comprises an antigen which is derived from Adenoviridae, Picomaviridae, Herpesviridae, Hepadnaviridae, Coronaviridae, Flaviviridae, Retroviridae, Orthomyxoviridae, Paramyxoviridae, Papovaviridae, Polyomavirus, Poxviridae, Rhabdoviridae, Togaviridae, Mycobacterium tuberculosis, Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Candida, Aspergillus, Cryptococcus, Histoplasma, Pneumocytis, Stachybotrus , Bacillus anthracis, Clostridium botulinum, Mycobacterium leprae, Yersinia pestis, Rickettsia prowazekii, Bartonella spp., malaria, amoebiasis, babesiosis, giardiasis, toxoplasmosis, cryptosporidiosis, trichomoniasis, Chagas disease, leishmaniasis, African trypanosomiasis (sleeping sickness), Acanthamoeba keratitis, or primary amoebic meningoencephalitis (naegleriasis). In some embodiments, a nucleic acid described herein may be used for production of an immunogenic peptide or immunogenic protein that can be administered to a subject to produce an antibody response in the subject.

[0059] In some embodiments, a gene comprises a sequence encoding a peptide or protein which is selected for the purposes of producing a peptide or protein hormone. Non-limiting examples of peptide or protein hormones which may be encoded by nucleic acids described herein include insulin, glucagon, growth hormone, thyroid-stimulating hormone, adrenocorticotropic hormone, follicle-stimulating hormone, luteinizing hormone, prolactin, oxytocin, vasopressin, parathyroid hormone, cortisol, erythropoietin, luteinizing hormone p, and aldosterone. In some embodiments, a nucleic acid described herein may be used for production of a peptide or protein hormone that can be administered to a subject to improve hormone levels in the subject, such as a subject having or suspecting of having a disease, disorder, or condition arising from a hormone deficiency.

[0060] Regulatory Sequences

[0061] In some embodiments, a nucleic acid comprises one or more regulatory sequences. The term “regulatory sequence” may be used herein to refer to a nucleic acid sequence that is capable of modulating the expression, stability, and / or levels of an RNA (e.g., an mRNA) and / or a peptide or protein product thereof when operably linked to a gene sequence encoding the RNA. A nucleic acid sequence (e.g., a sequence comprising a gene) and a regulatory sequence may be referred to as “operably linked” when they are associated in such a way (e.g., via a covalent bond) as to place the expression of the nucleic acid sequence under the influence or control of the regulatory sequence. As a non-limiting example, a gene sequence may be referred to as operably linked to a promoter if induction of the promoter results in the transcription of a coding sequence comprised in the gene, if the nature of the linkage between the gene and the promoter does not result in the introduction of a frame-shift mutation, and / or interfere with the ability of the promoter to direct the transcription of the coding sequence.

[0062] Non-limiting examples of regulatory sequences include a transcription regulatory sequence, a splicing regulatory sequence, a translation regulatory sequence, and / or a sequence that regulates post-translational modification of a peptide or protein. Further non-limiting examples of regulatory sequences include a promoter, an enhancer, a silencer, a transcription factor binding sequence, a 5' UTR, a 3' UTR, a translation initiation regulatory sequence, a transcriptional start sequence, a transcription terminator sequence, an acceptor / donor splicing site, a mRNA degradation or decay signal, a polyadenylation signal, a translation initiation codon, a protein binding site (e.g., a pentatricopeptide repeat (PPR) protein binding site), a spacer sequence, a ribosome binding site, a ribozyme, an intron, a translation terminator sequence, and / or a stop codon. In some embodiments, a regulatory sequence is a transcriptional regulatory sequence (e.g., a promoter), a post-transcriptional regulatory sequence (e.g., a splicing regulatory signal or a polyadenylation signal), or a translation regulatory sequence (e.g., a translation initiation regulatory sequence). In some embodiments, a regulatory sequence is native to a nucleic acid or a gene comprised therein. In some embodiments, a regulatory sequence is heterologous to a nucleic acid or a gene comprised therein. As used herein, a gene or an mRNA thereof may comprise a “heterologous regulatory sequence” if the heterologous regulatory sequence: comprises one or more substitutions, deletions, and / or insertions relative to the regulatory sequence which is normally linked to (or “native” to) a gene; is completely different from the regulatory sequence which is normally linked to the gene; and / or is operably linked to a gene which does not normally comprise that regulatory sequence (e.g., a gene which normally does not comprise a splice site in an exon but has been engineered to comprise a non-natural splice site).

[0063] In some embodiments, a heterologous regulatory sequence described herein is capable of upregulating the expression of a gene sequence to which it is operably linked. In some embodiments, a heterologous regulatory sequence is capable of upregulating an RNA (e.g., an mRNA) and / or a peptide or protein product thereof relative to the expression of the RNA and / or peptide or protein in its natural environment (e.g., cellular host) and / or when operably linked to its native regulatory sequence. In some embodiments, heterologous regulatory sequences described herein are capable of upregulating expression of an RNA and / or peptide or protein product thereof by about 1-2 fold, 1-5 fold, 1-10 fold, or 1-25 fold (e.g., wherein the expression of the RNA and / or peptide or protein is upregulated relative to when the native regulatory sequence is operably linked to the gene). In some embodiments, heterologous regulatory sequences described herein are capable of upregulating expression of an RNA and / or peptide or protein product thereof by about 2 fold, 3 fold, 4 fold, 5 fold, 6 fold, 7 fold, 8 fold, 9 fold, 10 fold, 11 fold, 12 fold, 13 fold, 14 fold, 15 fold, 16 fold, 17 fold, 18 fold, 19 fold, 20 fold, 21 fold, 22 fold, 23 fold, 24 fold, or 25 fold (e.g., wherein the expression of the RNA and / or peptide or protein is upregulated relative to when the native regulatory sequence is operably linked to the gene). In some embodiments, heterologous regulatory sequences described herein are capable of upregulating expression of an RNA and / or peptide or protein product thereof by more than 25 fold (e.g., 26-50 fold, 51-75 fold, or more) (e.g., wherein the expression of the RNA and / or peptide or protein is upregulated relative to when the native regulatory sequence is operably linked to the gene).

[0064] In some embodiments, a heterologous regulatory sequence described herein is capable of downregulating the expression of a gene sequence to which it is operably linked. In some embodiments, a heterologous regulatory sequence is capable of downregulating an RNA (e.g., an mRNA) and / or a peptide or protein product thereof relative to the expression of the RNA and / or peptide or protein in its natural environment (e.g., cellular host) and / or when operably linked to its native regulatory sequence. In some embodiments, heterologous regulatory sequences described herein are capable of downregulating expression of an RNA and / or peptide or protein product thereof by more than 25 fold (e.g., 26-50 fold, 51-75 fold, or more) (e.g., wherein the expression of the RNA and / or peptide or protein is downregulated relative to when the native regulatory sequence is operably linked to the gene). In some embodiments, heterologous regulatory sequences described herein are capable of downregulating expression of an RNA and / or peptide or protein product thereof by about 2 fold, 3 fold, 4 fold, 5 fold, 6 fold, 7 fold, 8 fold, 9 fold, 10 fold, 11 fold, 12 fold, 13 fold, 14 fold, 15 fold, 16 fold, 17 fold, 18 fold, 19 fold, 20 fold, 21 fold, 22 fold, 23 fold, 24 fold, or 25 fold (e.g., wherein the expression of the RNA and / or peptide or protein is downregulated relative to when the native regulatory sequence is operably linked to the gene). In some embodiments, heterologous regulatory sequences described herein are capable of downregulating expression of an RNA and / or peptide or protein product thereof by more than 25 fold (e.g., 26-50 fold, 51-75 fold, or more) (e.g., wherein the expression of the RNA and / or peptide or protein is downregulated relative to when the native regulatory sequence is operably linked to the gene).

[0065] Spacer Sequences

[0066] In some embodiments, a nucleic acid comprises at least one spacer sequence. The term “spacer sequence” may be used herein to refer to a stretch of nucleotides that separate (or intervene) two nucleic acid sequences, in the case of the instant technology a stretch of nucleotides between the Shine-Dalgamo sequence and a start codon (i.e. AUG) in the 5' UTR. In some embodiments, a nucleic acid comprises a plurality of spacer sequences. In some embodiments, a nucleic acid comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more spacer sequences.

[0067] In some embodiments, a spacer comprises at least one nucleotide in length. In some embodiments, a spacer comprises two or more nucleotides in length. In some embodiments, a spacer comprises 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In some embodiments, a spacer comprises 10-20 nucleotides in length (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length). In some embodiments, a spacer comprises more than 20 nucleotides in length. In some embodiments, a spacer comprises 10-100 nucleotides in length (e.g., 10-25, 10- 30 10-40, 10-50, 10-60, 20-70, 30-80, 20-90, or 20-100 nucleotides in length) or more than 100 nucleotides in length (e.g., 100-120, 120-140, 140-160, 160-170, 170-180, 180-200, 200-250, 250-300, 300-400, or 400-500 nucleotides). In some embodiments, a spacer comprises 12 nucleotides in length. In some embodiments, a spacer comprises 4-20 nucleotides in length (e.g., 4-5, 4-12, 5-12, 6-15, 7-15, 8-15, 9-15, 10-15, 11-15, 11-13, 11-14, 11-16, 11-17, 11-18, 11-19, 11-20, 12-15, 12-13, 12-14, 12-16, 12-17, 12-18, 12-19, or 12-20 nucleotides in length).

[0068] In some embodiments, a spacer does not comprise a translation initiation codon (e.g., AUG, UUG, or GUG). In some embodiments, a spacer comprises a sequence of mainly A / U / C or A / T / C nucleotides. In some embodiments, a spacer sequence may comprise a mix of A / U / C / G nucleotides, wherein “100%-x%” represents the mix of A / U / C / G nucleotides in the spacer sequence such that “x” represents the proportion of A / U / C nucleotides and “100%-x” represents the proportion of G nucleotides. In some embodiments, the proportion of A / U / C nucleotides (“x”) is approximately 60-70%, 70-80%, 80-90%, 90-95%, or more than 95%. In some embodiments, the proportion of G nucleotides is approximately 10-20%, 10-15%, 10-12%, 9-15%, 8-15%, 7-15%, 6-15%, 5-15%, 5-10%, 4-10%, 3-10%, 2-10%, 1-10%, 4-5%, 3-5%, 2- 5%, 1-5%, 0.5-1% or 0%. In some embodiments, a spacer comprises a sequence of mainly A and / or T / U nucleotides. In some embodiments, a spacer comprises a sequence of at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% A and / or T / U nucleotides or 100% A and / or T / U. In some embodiments, a spacer comprises more A / U nucleotides than G / C nucleotides. In some embodiments, a spacer sequence may comprise a mix of A / U and G / C nucleotides, wherein “100%-x%” represents the mix of A / U and G / C nucleotides in the spacer sequence such that “x” represents the proportion of A / U nucleotides and “100%-x” represents the proportion of G / C nucleotides. In some embodiments, the proportion of A / U nucleotides (“x”) is approximately 60-70%, 70-80%, 80-90%, 90-95%, or more than 95%. In some embodiments, a spacer comprises no G nucleotides. In some embodiments, a spacer comprises only A / U / C nucleotides. In some embodiments, a spacer comprises only A / U nucleotides. In each instance of disclosure of U, a T may also apply. For instance, in the above paragraph where U is mentioned, T may alternatively be recited.

[0069] In some embodiments, a spacer is inserted between two regulatory sequences (e.g., a Shine-Dalgarno sequence and a translation initiation codon), a spacer is inserted between a regulatory sequence (e.g., a promoter) and a gene, and / or a spacer is inserted between two or more genes. In some embodiments, a spacer is inserted in a non-coding region of a gene or genome. In some embodiments, a spacer is in the 5' UTR. In some embodiments, a spacer is inserted to modulate the function of one or more sequences comprised in a nucleic acid, such as to modulate expression of a gene or the ability of a regulatory sequence to regulate a sequence to which it is operably linked. In some embodiments, a spacer may not have a direct structural and / or functional role in a peptide or protein encoded by a gene.

[0070] In some embodiments, a spacer comprises a sequence that is designed to not form base pair-base pair interactions (e.g., hybridize) with nucleotides comprised within the spacer. Such a spacer comprises nucleotides that do not form base pair-base pair interactions that form a secondary structure. Base pairs that can form a secondary structure consist of a first set of 2, 3,

[0071] 4, 5, 6, or more consecutive nucleotides that are separated by at least one nucleotide from a second set of 2, 3, 4, 5, 6, or more consecutive nucleotides of matching base pairs. Matching pairs are A - T / U and G-C as well as related naturally occurring and non-naturally occurring nucleosides. In some embodiments the first set of consecutive nucleotides is separated by 2, 3, 4,

[0072] 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides from the second set of consecutive nucleotides. In some embodiments, a spacer comprises a sequence that is designed to not form base pair-base pair interactions (e.g., hybridize) with nucleotides upstream and / or downstream of the spacer. In some embodiments, a spacer comprises a sequence that is designed to not form base pair-base pair interactions (e.g., hybridize) with nucleotides that are positioned approximately 1-3, 3-6, 6-10, 10-20, or 20-30 nucleotides upstream and / or downstream of the spacer. In some embodiments, a spacer comprises a sequence that is designed to not form base pair-base pair interactions (e.g., hybridize) with nucleotides in a region of secondary structure of RNA (e.g., mRNA) that is located upstream and / or downstream of the spacer. In some embodiments, a spacer comprises a sequence that is designed to not form base pair-base pair interactions (e.g., hybridize) with nucleotides in a stem-loop secondary structure region in an RNA (e.g., an mRNA, such as a stem-loop secondary structure located upstream of a translation initiation codon). In some embodiments, a spacer comprises a sequence that is designed to not form base pair-base pair interactions (e.g., hybridize) with nucleotides in a stem-loop secondary structure in an RNA (e.g., an mRNA, such as a stem- loop secondary structure located upstream of a translation initiation codon) which is necessary for translation initiation (e.g., a stem-loop structure comprising a Shine-Dalgarno sequence and / or a stem-loop structure comprising a protein-binding sequence). In some embodiments, a spacer comprises a sequence that does not corrupt an upstream mRNA stem-loop secondary structure which comprises a native or an engineered RNA-specific binding protein binding site. A sequence that does not corrupt an upstream mRNA stem-loop secondary structure is a sequence that does not form a secondary structure with an upstream mRNA stem-loop secondary structure region, e.g., does not form base pair-base pair interactions that form a secondary structure as discussed above. Various tools are available in the art for predicting the secondary structures of RNA sequences (see, e.g., Mfold) which may be used in designing spacer sequences based on embodiments described herein.

[0073] In some embodiments, a spacer comprises a sequence that lacks a canonical or a non- canonical translation initiation codon. The canonical translation initiation codon is AUG. Non- canonical translation initiation codons are any codons other than AUG. For instance, non- canonical translation initiation codons include but are not limited to CUG, GUG, UUG, ACG, AAG, AGG, AUA, AUU, AUC.

[0074] Translation Initiation Regulatory Sequences

[0075] In some embodiments, a nucleic acid comprises at least one translation initiation regulatory sequence. The term “translation initiation regulatory sequence” may be used herein to refer to a nucleic acid sequence which is capable of modulating the expression, stability, and / or levels of a peptide or protein when operably linked to an mRNA comprising a sequence encoding the peptide or protein. In some embodiments, a heterologous translation initiation regulatory sequence may be used in a nucleic acid to increase (e.g., increase relative to a counterpart nucleic acid comprising a native translation initiation regulatory sequence) or decrease (e.g., decrease relative to a counterpart nucleic acid comprising a native translation initiation regulatory sequence) the rate of translation of a peptide or protein. In some embodiments, a heterologous translation initiation regulatory sequence may be used in a nucleic acid to increase or decrease the rate of translation of a peptide or protein.

[0076] Non-limiting examples of translation initiation regulatory sequences include a Shine- Dalgamo sequence (e.g., one comprising a consensus sequence AGGAG), a Kozack sequence (e.g., one comprising a consensus sequence of GCCRCCAUGG (SEQ ID NO: 1), wherein R is A or G), an internal ribosome entry site (IRES), a cap-proximal element (CPE), upstream open reading frames (uORF), and start codons.

[0077] The term “Shine-Dalgamo sequence” or “SD sequence” may be used herein to refer to a mRNA sequence found in prokaryotes (e.g., bacteria) which functions as a ribosome-binding site during the initiation of translation and helps position the ribosome accurately on an mRNA to initiate protein synthesis. The SD sequence can be partially or fully complementary to a region of the 16S ribosomal RNA (rRNA) component of the small ribosomal subunit. In some embodiments, a SD sequence comprises a sequence of AGGAG, AGGAGG, AGGAGGU, AGGAGG, AGGAGGAGG, AGGAGGUAUG (SEQ ID NO: 2), or AGGAGGAGGAUG (SEQ ID NO: 3). In some embodiments, an SD sequence comprises a sequence capable of forming a stem-loop secondary structure in an RNA (e.g., an mRNA). In some embodiments, a secondary structure formed by an SD sequence promotes binding of a ribosome or a subunit thereof during translation initiation.

[0078] In some embodiments, a translation initiation regulatory sequence comprises an SD sequence and a translation initiation codon. In some embodiments, an SD sequence is positioned 5' (or upstream) relative to the translation initiation codon (or start codon) in an mRNA. In some embodiments, a translation initiation codon comprises the sequence AUG, GUG, or UUG.

[0079] In some embodiments, a naturally occurring or native translation initiation regulatory sequence comprises an SD sequence and a translation initiation codon which are separated by 6- 12 nucleotides (e.g., 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, or 12 nucleotides). In other embodiments, a native translation initiation regulatory sequences comprises an SD sequence and a translation initiation sequence which are separated by 3 nucleotides, 4, nucleotides, 5 nucleotides, or 6 nucleotides.

[0080] In some embodiments, a heterologous translation initiation regulatory sequence comprises a heterologous spacer sequence which is positioned between an SD sequence and a translation initiation codon. In some embodiments, a heterologous spacer sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In some embodiments, a heterologous spacer sequence comprises 12 nucleotides in length. In some embodiments, a heterologous spacer sequence comprises 10-15, 15-20, 20-25, or 25-30 nucleotides in length. In some embodiments, a heterologous spacer sequence comprises 2-20 nucleotides in length.

[0081] In some embodiments, spacer lengths may be adjusted in order to incrementally upregulate the expression of a peptide or protein encoded by a gene described herein. In some embodiments, spacer lengths may be adjusted in order to incrementally downregulate the expression of a peptide or protein encoded by a gene described herein. In some embodiments, a spacer sequence and / or length may be designed to disrupt expression of a gene sequence. In some embodiments, a spacer may be designed to disrupt a Shine-Dalgarno and / or a proteinbinding site in an RNA (e.g., wherein translation of the nucleic acid is decreased relative to a counterpart nucleic acid comprising a spacer having a different length and / or different nucleotide sequence). In some embodiments, a spacer may be designed to disrupt a secondary structure (e.g., a stem loop secondary structure) in an RNA (e.g., a secondary structure which is used for initiating translation of an mRNA) (e.g., wherein translation of the nucleic acid is decreased relative to a counterpart nucleic acid comprising a spacer having a different length and / or different nucleotide sequence).

[0082] In some embodiments, a heterologous spacer sequence comprises a sequence which is not found in native gene regulatory regions, wherein the spacer is heterologous relative to the SD sequence. The heterologous sequence may be added to the regulatory region between the SD region and the start codon as a new nucleotide sequence which is included in that region without making additional changes to the regulatory region. The heterologous sequence may be added to the regulatory region between the SD region and the start codon while one or more native or naturally occurring nucleic acid sequences is deleted or removed. Additionally, the heterologous sequence may be added to the regulatory region between the SD region and the start codon by making one or more substitutions, insertions, and / or deletions relative to the native regulatory sequence to create a heterologous spacer. In some embodiments, a heterologous spacer sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotide substitutions relative to the native regulatory sequence. In some embodiments, a heterologous spacer sequence comprises 10-15, 15-20, 20- 25, or 25-30 nucleotide substitutions relative to the native regulatory sequence. In some embodiments, a heterologous spacer sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 additional nucleotides relative to the native regulatory sequence. In some embodiments, a heterologous spacer sequence comprises 10-15, 15-20, 20-25, or 25-30 additional nucleotides relative to the native regulatory sequence. In some embodiments, a heterologous spacer sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 fewer nucleotides relative to the native regulatory sequence. In some embodiments, a heterologous spacer sequence comprises 10-15, 15-20, 20-25, or 25-30 fewer nucleotides relative to the native regulatory sequence.

[0083] Protein-Binding Sequences

[0084] In some embodiments, a nucleic acid comprises at least one protein binding sequence. The terms “protein-binding sequence” and “protein-binding site” may be used herein to refer to a nucleic acid sequence which is capable of binding to a protein comprising a structure that specifically interacts with the nucleic acid sequence. In some embodiments, a nucleic comprises a plurality of protein-binding sequences. In some embodiments, a nucleic acid comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 protein-binding sequences. In some embodiments, a proteinbinding sequence is operably linked to a gene. In some embodiments, a gene is operably linked to a plurality of protein-binding sites.

[0085] In some embodiments, a protein-binding sequence comprises at least two nucleotides in length. In some embodiments, a protein-binding sequence comprises 2-12 nucleotides in length (e.g., 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, or 12 nucleotides). In some embodiments, a protein-binding sequence comprises more than 12 nucleotides in length (e.g., 13-20, 20-30, etc.).

[0086] In some embodiments, a protein-binding sequence binds to a protein which preferentially or specifically binds to the sequence in an RNA molecule (e.g., an mRNA). Non-limiting examples RNA-binding proteins that may preferentially or specifically bind to a protein-binding sequence are proteins which binds to RNA structural motifs (e.g., secondary structures, such as stem-loop structures), stabilize RNAs (e.g., stabilizes a secondary structure in an RNA and / or prevents the degradation or decay of the RNA), degrades RNA, regulates the trafficking and / or transport of RNA, regulates the translation rate of an RNA, and / or enzymatically modifies (e.g., cuts and / or edits) RNAs (e.g., base editing RNAs). In some embodiments, a protein-binding sequence binds to an RNA-binding protein in a cell. In some embodiments, a protein-binding sequence binds to an RNA-binding protein comprised in a plant cell. In some embodiments, a protein-binding sequence binds to an RNA-binding protein comprised in a chloroplast of a plant cell. Non-limiting examples of an RNA-binding proteins in plant cells include a pentatricopeptide repeat (PPR) protein (e.g., a PPR protein described herein), a chloroplast RNA editing Required 2 (CRR2), a chloroplast RNA splicing 1 (CRS1), an RNA helicase 3 (RH3), a high chlorophyll fluorescence 152 (HCF152), a half-a-tetratricopeptide protein high chlorophyll fluorescence 107 (HCF107), a maturase K (MatK), a plastid transcriptionally active 9 (PTAC), a chloroplast RNA-binding protein 1 (CRP1), an RNA helicase of the chloroplast 1 (RHON1), and a sigma factor 2 (SIG2).

[0087] In some embodiments, a protein-binding sequence in an RNA (e.g., an mRNA) forms a secondary structure that is recognized by the protein which binds the binding sequence. In some embodiments, a protein binding sequence in an RNA comprises a stem-loop secondary structure that binds the protein. In some embodiments, the secondary structure (e.g., a stem-loop structure) is positioned upstream of a translation initiation codon. In some embodiments, the secondary structure (e.g., a stem- loop structure) is positioned downstream of a translation initiation codon. In some embodiments, the secondary structure (e.g., stem-loop structure) is comprised in a Shine-Delgarno sequence or comprises a portion of the Shine-Delgamo sequence.

[0088] In some embodiments, a protein-binding sequence is positioned in a 5' UTR or a 3' UTR of a gene (e.g., a gene described herein). In some embodiments, a nucleic acid comprises a gene which is operably linked to a translation initiation regulatory sequence described herein and a protein-binding site, wherein the protein binding site is posited upstream of the translation initiation regulatory sequence. In some embodiments, a protein-binding sequence comprises a sequence present in the 5' UTR of an RNA (e.g, an mRNA) encoded by a plant gene. In some embodiments, a protein-binding sequence comprises a sequence present in the 5' UTR of an RNA (e.g., an mRNA) encoded by a gene in a chloroplast genome. In some embodiments, a protein-binding sequence is positioned in a coding sequence of a gene.

[0089] In some embodiments, a protein-binding site comprises a pentatricopeptide repeat (PPR) protein-binding site. A “pentatricopeptide repeat (PPR) protein” or “PPR protein” refers to a family of RNA-binding proteins that can play roles in post-transcriptional processes, such as RNA editing, RNA splicing, regulation of RNA stability, and translation. In some embodiments, PPR proteins comprise repeated structural motifs comprising approximately 35 amino acids in length. In some embodiments, repeat structural motifs in PPR proteins fold into three alphahelices and form a helical structure that facilitates RNA binding. PPR proteins are commonly found in mitochondria and chloroplasts where they contribute to the regulation of gene expression. Non-limiting examples of PPR proteins include PPR1, PPR2, PPR3, PPR4, PPR5, PPR10, PPR53, PPR40, MTSF1, OTP70, and MEF11. In some embodiments, a PPR protein is PPR10 or PPR10GG. In some embodiments, a PPR10 protein or a or PPR10GGprotein binds to an atpH RNA sequence or a portion thereof. In some embodiments, a PPR protein is MTSF1. In some embodiments, a MTSF1 protein binds to an atp6 RNA sequence or a portion thereof. In some embodiments, a PPR protein is OTP70. In some embodiments, a OTP70 protein binds to a matR RNA sequence or a portion thereof. In some embodiments, a PPR protein is MEF11. In some embodiments, a MEF11 protein binds to a cox2 RNA sequence or a portion thereof.

[0090] Promoters In some embodiments, a nucleic acid comprises one or more promoters. In some embodiments, a nucleic acid comprises one or more genes which are operably linked to a promoter. In some embodiments, a nucleic acid comprises a plurality of promoters (e.g., two, three, four, five, or more than five promoters). In some embodiments, a nucleic acid comprises a plurality of genes (e.g., two, three, four, five, or more than five genes) which are each operably linked to a promoter. In some embodiments, a single promoter may be operably linked to one gene (e.g., wherein the one gene is located in a monocistronic operon) or a plurality of genes (e.g., wherein the plurality of genes are located in a polycistronic operon). In some embodiments, a nucleic acid comprises a plurality of genes which are each operably linked to a respective promoter. In some embodiments, a nucleic acid comprises at least one gene (e.g., one gene, two genes, three genes, four genes, five genes, or more than five genes) operably linked to a first promoter and at least one (e.g., one gene, two genes, three genes, four genes, five genes, or more than five genes) operably linked to a second promoter

[0091] In some embodiments, a promoter is a plant- specific promoter. In some embodiments, a promoter is a promoter that is active in a chloroplast, such as a chloroplast gene-specific promoter. Non-limiting examples of promoters which may be useful for expression of a nucleic acid described herein in a chloroplast include a Prm promoter, a Patpi promoter, a PpsbA promoter, a P16Srrn promoter, a PpsbA promoter, a PrmLatpB promoter, a ZmPclpP promoter, a Prm promoter, a Prrn(np) promoter, a PNG 1014 promoter, a Pcyc6 promoter, a Nac2 promoter, a PrbcL promoter, a PpsbD promoter, a psaA promoter, an atpl promoter, and a CMV163C promoter.

[0092] In some embodiments, a promoter is a constitutive promoter. In some embodiments, a constitutive promoter maintains constant expression of a gene regardless of the conditions or physiological state of a cell. Non-limiting examples of constitutive promoters include the Herpes Simplex virus (HSV) promoter, the thymidine kinase (TK) promoter, the Simian Virus 40 (SV40) promoter, the Mouse Mammary Tumor Virus (MMTV) promoter, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) (see, e.g., Boshart et al., Cell, 41:521-530 (1985)), the dihydro folate reductase promoter, the P-actin promoter, the phosphoglycerol kinase (PGK) promoter, the CAG promoter, the human elongation factor- 1 alpha (EFla) promoter [Invitrogen], the CaMV 35S promoter, the actin promoter (e.g., Actl or Act2), the UBQ10 promoter, the EFla promoter, the tubulin promoter, nopaline synthase promoter, the AtEFla promoter, the GAPDH promoter, the pTac promoter, the BC1 promoter, the OsActl promoter, and the CAB (Chlorophyll a / b-binding protein) promoter.

[0093] In some embodiments, a promoter is an inducible promoter. In some embodiments, an inducible promoter regulates gene expression in response to exogenously supplied compounds, environmental factors (e.g., temperature), or the presence of a specific physiological state. Inducible promoters and inducible systems are available from a variety of commercial sources, including, without limitation, Invitrogen, Clontech, and Ariad. Many other systems have been described and can be readily selected by one of skill in the art. Non-limiting examples of inducible promoters include the cytochrome P450 gene promoters, heat shock protein gene promoters, metallo thionein gene promoters, hormone-inducible gene promoters, the dexamethasone (Dex) -inducible mouse mammary tumor virus (MMTV) promoter, the T7 polymerase promoter system, the ecdysone insect promoter, the tetracycline-inducible system, the RU486-inducible system, the rapamycin-inducible system, the IPTG-inducible promoter, an alcohol-inducible promoter (e.g., an ethanol-inducible promoter), a glucocorticoid-inducible promoter (GR), a P-Estradiol-inducible promoter, a methyl Jasmonate-inducible promoter (JMT), a copper-inducible promoter (CUP1), a nitrate-inducible promoter, a hydrogen peroxideinducible promoter (e.g., CAT3), an AbaA-inducible promoter, a xylose-inducible promoter, and a phytochrome-interacting factor 3-inducible promoter.

[0094] In some embodiments, a promoter is a tissue- specific promoter. In some embodiments, a tissue-specific promoter is capable of binding to a tissue-specific transcription factor or repressor protein that regulates transcription in a tissue-specific manner. Non-limiting examples of tissue-specific promoters include a root-specific promoter (e.g., an RCc3 promoter), a leafspecific promoter (e.g., an At2S3 promoter), a stem-specific promoter (e.g., an GmSHMT promoter), a seed-specific promoter (e.g., an Oleosin promoter), a pollen- specific promoter (e.g., a LAT52 promoter), an endosperm- specific promoter (e.g., an ZmLegl promoter), a root hairspecific promoter (e.g., a pEXPB7 promoter), a trichome- specific promoter (e.g., a GL1 promoter), a guard cell-specific promoter (e.g., a GC1 promoter), a vascular tissue-specific promoter (e.g., a AtMYB46 promoter), a xylem- specific promoter (e.g., a AtCesA8 promoter), a phloem- specific promoter (e.g., a SUC2 promoter), and an glandular trichome- specific promoter (e.g., a CYP71AV1 promoter).

[0095] Untranslated Regions In some embodiments, a nucleic acid described herein comprises a 5' UTR and / or a 3' UTR. The terms “5' untranslated region” or “5' UTR” may be used to refer to a sequence in an mRNA located upstream of a protein coding sequence which does not encode amino acids present in the peptide or product encoded by the mRNA. In some embodiments, a 5' UTR sequence comprises nucleotides beginning at the transcription start site of a gene and extending through to the translation initiation codon. In some embodiments, a 5' UTR comprises a protein binding site described. In some embodiments, a 5' UTR comprises a translation initiation regulatory sequence described herein. The terms “3' untranslated region” or “3' UTR” may be used to refer to a sequence in an mRNA located downstream of a protein coding sequence which does not encode amino acids present in the peptide or product encoded by the mRNA. In some embodiments, a 3' UTR sequence comprises nucleotides beginning at a stop codon in an mRNA and extending through to end of the polyadenylation signal and / or polyA tail of an mRNA.

[0096] In some embodiments, a nucleic acid described herein comprises a 5' UTR which is capable of regulating expression of a gene encoded by the nucleic acid. In some embodiments, a 5' UTR comprises a sequence which is useful for expression of a nucleic acid described herein in a chloroplast. In some embodiments, a 5' UTR comprises a sequence which is useful for expression of a nucleic acid in a chloroplast comprises a 5' UTR of a chloroplast gene, such as psbA, rbcL, atpl, psbA, atpB, T7gl0, and cry2a.

[0097] In some embodiments, a nucleic acid described herein comprises a 3' UTR which is capable of regulating expression of a gene encoded by the nucleic acid. In some embodiments, a 3' UTR comprises a sequence which is useful for expression of a nucleic acid described herein in a chloroplast. In some embodiments, a 3' UTR comprises a sequence which is useful for expression of a nucleic acid in a chloroplast comprises a 3' UTR of a chloroplast gene, such as rpsl6, psbA, rbcL, rrnB, rbcS, and petD.

[0098] Biological Systems

[0099] Cells and Cell Populations

[0100] In some embodiments, a cell or cell population comprises at least one nucleic acid described herein. In some embodiments, a cell or cell population comprises a plurality of nucleic acids described herein (e.g., one, two, three, four, five, or more than five nucleic acids). In some embodiments, a cell population comprises 2-100, 100-500, 500-1,000, 1,000-5,000, 5,000- 10,000, 10,000-50,000, 50,000-100,000, 100,000-500,000, 500,000-1,000,000, 1,000,000- 5,000,000, 5,000-10,000,000, or more than 10,000,000 cells. In some embodiments, a cell or cell population is a cell or cell population in culture (e.g., an in vitro cell or cell population, such as an isolated cell or cell population). In some embodiments, a cell or cell population is located in a tissue and / or an organism (e.g., a plant tissue).

[0101] In some embodiments, a bacterium or a cell population thereof is useful for producing nucleic acids described herein which can be then transferred to a plant cell expression system.

[0102] In some embodiments, a cell is a plant cell or a cell population thereof. In some embodiments, a plant cell or cell population thereof comprises one or more nucleic acids described herein in the chloroplasts, the mitochondrias, and / or the nucleus. Non-limiting examples of plant cells (e.g., plant cell expression systems) include epidermal cells, parenchyma cells, collenchyma cells, sclerenchyma cells, xylem cells, phloem cells, meristematic cells, guard cells, trichomes, root hair cells, ligule cells, aleurone cells, endosperm cells, and embryo cells. In some embodiments, a plant cell or a cell population thereof are cells of an edible plant, such as cultivated forms of grasses (Poaceae) (e.g., cereals or pseudocereals, such as wheat (e.g., spelt, einkom, emmer, kamut, durum and triticale), rye, barley, rice, wild rice, maize (com), millet, sorghum, teff, fonio oats, amaranth, quinoa, buckwheat pulses), legumes or lentils (e.g., chickpeas, lentils, peas, soybeans, fava beans, mung beans, black-eyed peas, kidney beans, pigeon peas, cowpeas, etc.), oilseeds (e.g., canola, sunflower, soybeans, cottonseed, peanut, sesame, safflower, flaxseed, mustard, olive, etc.), tubers or roots (e.g., potato, sweet potato, cassava, yam, taro, beetroot, turnip, carrot, radish, ginger, etc.), fruits (e.g., apple, banana, orange, mango, pineapple, grape, strawberry, watermelon, avocado, lemon, lime, papaya, tomato, etc.), vegetables (e.g., cucumber, bell pepper, spinach, lettuce, broccoli, cauliflower, eggplant, zucchini, onion, etc.), nuts (e.g., almond, walnut, pistachio, cashew, pecan, hazelnut, macadamia, Brazil nut, pine nut, chestnut, etc.), beverage crops (e.g., coffee, tea, cocoa, sugarcane, hops, etc.), or spices or herbs (e.g., black pepper, vanilla, cinnamon, nutmeg, clove, cardamom, basil, thyme, rosemary, oregano, etc.). In some embodiments, a cell or cell population are cells of a fibrous crop (e.g., cotton, hemp, jute, ramie, kenaf, sisal, abaca, coir, etc.). Further non-limiting examples of plant cells include A rabidopsis thaliana cells, Nicotiana tabacm cells (e.g., BY-2 and NT1 cells), Medicago truncatula cells, Oryza sativa (rice) cells, Solatium tuberosum (potato) cells, Lycopersicon esculentum (tomato) cells, Vitis vinifera (grape) cells, Coffea arabica (coffee) cells, Glycine max (soybean) cells, Zea mays (maize) cells, Linum usitatissimum (flax) cells, Jatropha curcas cells, and Ricinus communis (castor bean) cells. Plants and Seeds Thereof

[0103] In some embodiments, a plant or a seed thereof comprises one or more nucleic acids (e.g., one, two, three, four, five, or more than nucleic acids) described herein. In some embodiments, a plant or a seed comprising one or more cells or cell populations described herein. In some embodiments, a plant or a seed comprising one or more nucleic acids described herein is a genetically engineered or transgenic plant or seed. In some embodiments, a plant is an edible plant or a seed thereof, such as cultivated forms of grasses (Poaceae') (e.g., cereals or pseudocereals, such as wheat (e.g., spelt, einkom, emmer, kamut, durum and triticale), rye, barley, rice, wild rice, maize (corn), millet, sorghum, teff, fonio oats, amaranth, quinoa, buckwheat pulses), legumes or lentils (e.g., chickpeas, lentils, peas, soybeans, fava beans, mung beans, black-eyed peas, kidney beans, pigeon peas, cowpeas, etc.), oilseeds (e.g., canola, sunflower, soybeans, cottonseed, peanut, sesame, safflower, flaxseed, mustard, olive, etc.), tubers or roots (e.g., potato, sweet potato, cassava, yam, taro, beetroot, turnip, carrot, radish, ginger, etc.), fruits (e.g., apple, banana, orange, mango, pineapple, grape, strawberry, watermelon, avocado, lemon, lime, papaya, tomato, etc.), vegetables (e.g., cucumber, bell pepper, spinach, lettuce, broccoli, cauliflower, eggplant, zucchini, onion, etc.), nuts (e.g., almond, walnut, pistachio, cashew, pecan, hazelnut, macadamia, Brazil nut, pine nut, chestnut, etc.), beverage crops (e.g., coffee, tea, cocoa, sugarcane, hops, etc.), or spices or herbs (e.g., black pepper, vanilla, cinnamon, nutmeg, clove, cardamom, basil, thyme, rosemary, oregano, etc.). In some embodiments, a plant is a fibrous crop, or a seed thereof (e.g., cotton, hemp, jute, ramie, kenaf, sisal, abaca, coir, etc.),

[0104] Methods

[0105] Methods of Expressing Nucleic Acids

[0106] In some embodiments, a method comprises contacting at least one cell with one or more nucleic acids (e.g., one, two, three, four, five, or more than five nucleic acids) described herein. In some embodiments, contacting at least one cell with one or more nucleic acids results in the expression of the nucleic acids in the cell or a cell population thereof (e.g., expression in a chloroplast).

[0107] In some embodiments, a method comprises introducing a nucleic acid into a chloroplast of a plant cell. Non-limiting examples of methods for introducing nucleic acids into chloroplasts include particle bombardment (e.g., using gold or tungsten particles), electroporation, biolistic methods, polyethylene glycol (PEG) -mediated transformation, glass bead transformation, nanoparticle-mediate nucleic acid transfer, horizontal cell-cell transfer (e.g., via grafting-based methods, such as cell grafting), use of UV-laser microbeam-mediated nucleic acid delivery, delivery of nucleic acids via single- walled carbon nanotubes (e.g., chitosan- wrapped singlewalled carbon nanotubes), covalent attachment of a cell- and / or chloroplast-penetrating peptide or protein to a nucleic acid, delivery of a nucleic acid via Agro / racterzMm-mediated delivery, delivery of a nucleic acid comprised in an episomal vector or a minichromosome, delivery of a nucleic acid comprised in a geminivirus, or any combination thereof.

[0108] In some embodiments, a nucleic acid described herein is integrated into a chloroplast genome. Non-limiting examples of chloroplast genome sequences which may be used as a target site for integration of nucleic acid described herein include rrnl6-tml, tml-tmA, rbcL-aacD, tmfM-tmG, rps!2-tmV, rps7-ndhB, tmV-tml, tmV-rps!2 / 7, rbcL-accD, tmR-tnrN, tmfM-tmG, rps!6 / rpsl2, rml6-rm23, 16S-23S, psbK, psbT-psbN, ahasWT, rrsB-tml , psbY-psbA, chlL, and tmR-CCG.

[0109] In some embodiments, integration of a nucleic acid into a chloroplast genome comprises homologous recombination-Zhomology-directed repair-based methods. In some embodiments, a nucleic acid described herein is engineered to comprise one or more stretches of sequence that are homologous to a chloroplast genomic site. In some embodiments, a sequence which is homologous to a chloroplast genomic site may comprise 20-2000 nucleotides in length. In some embodiments, a sequence which is homologous to a chloroplast genomic site may comprise approximately 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-125, 125-150, ISO- 175, 175-200, 200-225, 225-250, 250-275, 275-300, 300-325, 325-350, 350-375, 375-400, 400- 425, 425-450, 450-475, 475-500, 500-525, 525-550, 550-575, 575-600, 600-625, 625-650, 650- 675, 675-700, 700-725, 725-750, 750-775, 775-800, 800-825, 825-850, 850-875, 875-900, 900- 925, 925-950, 950-975, 975-1000, 1000-1050, 1050-1100, 1100-1150, 1150-1200, 1200-1250, 1250-1300, 1300-1350, 1350-1400, 1400-1450, 1450-1500, 1500-1550, 1550-1600, 1600-1650, 1650-1700, 1700-1750, 1750-1800, 1800-1850, 1850-1900, 1900-1950, or 1950-2000 nucleotides in length. In some embodiments, a sequence which is homologous to a chloroplast genomic site may comprise at least 75% sequence identity to the target site. In some embodiments, a sequence which is homologous to a chloroplast genomic site may comprise 80%-90%, 90-95%, or 95-99% sequence identity to the target site. In some embodiments, a sequence which is homologous to a chloroplast genomic site may comprise 99.5%-99.99% or more sequence identity to the target site. In some embodiments, a nucleic acid described herein is engineered to comprise at least two stretches of sequences that are homologous to a chloroplast genomic site. In some embodiments, the at least two stretches of sequence comprise an equal number of nucleotides. In some embodiments, the at least two stretches of sequence comprise a nonequal number of nucleotides. In some embodiments, the at least two stretches of sequence may differ by approximately 1-5, 5-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70- 80, 80-90, 90-100, 100-125, 125-150, 150-175, 175-200, 200-225, 225-250, 250-275, 275-300, 300-325, 325-350, 350-375, 375-400, 400-425, 425-450, 450-475, 475-500 or more nucleotides in length. In some embodiments, the lengths of the at least two stretches of sequence may differ by approximately, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more nucleotides in length.

[0110] In some embodiments, a nucleic acid comprising sequences which are homologous to a chloroplast genome may be used to introduce one or more heterologous regulatory sequences described herein into a chloroplast genome (e.g., to modify one or more genes which are native to the chloroplast genome). In some embodiments, a nucleic acid comprising sequences which are homologous to a chloroplast genome may be used to introduce one or more heterologous genes described herein into a chloroplast genome (e.g., to modify one or more genes which are native to the chloroplast genome). In some embodiments, a nucleic acid comprising sequences which are homologous to a chloroplast genome may be used to introduce one or more heterologous genes and the regulatory sequences to which they are operably linked into a chloroplast genome. In some embodiments, a nucleic acid introduced into a chloroplast genome may comprise one or more selectable markers.

[0111] Methods of Purification, Isolation, and Detection

[0112] In some embodiments, a peptide or protein described herein is expressed in chloroplasts and then isolated / purified. In some embodiments, isolating and / or purifying a peptide or protein comprises lysing plant cells comprising the nucleic acid, harvesting lysate comprising the peptide or protein, and subjecting the lysate to one or more methods, such as size-exclusion chromatography, affinity chromatography (e.g., metal affinity chromatography), ion exchange chromatography, high-performance liquid chromatography, dialysis, etc. In some embodiments, isolating and / or purifying a peptide or protein from chloroplasts expressing the nucleic acid may further comprise one or more selection methods to enrich for cells comprising the nucleic acid (e.g., cells expressing a marker or reporter encoded by the nucleic acid). In some embodiments, a peptide or protein can be contacted with one or more cells (e.g., cells in a subject, such as a mammalian subject including, but not limited to, a human subject) after being expressed in a chloroplast and subjected to one or more isolation / purification steps (e.g., wherein the peptide or protein is provided in a pharmaceutical composition and / or administered to a subject, such as a mammalian subject including, but not limited to, a human subject).

[0113] In some embodiments, a method comprises detecting presence and / or expression of a nucleic acid described herein in a chloroplast or a biological sample thereof. Non-limiting examples of methods for detecting a peptide or protein encoded by nucleic acid described herein include immunoblot (e.g., dot blot, 2-D gel electrophoresis, Western Blot, etc.), electrochemiluminescence immunoassay (e.g., Meso-Scale Detection (MSD)), immunohistochemistry (IHC), ELISA (e.g., RCA-based ELISA or RT-PCR-based ELISA), label free immunoassays (e.g., surface plasmon resonance bio layer interferometry), immunoquantitative PCR, bead-based immunoassays, immunoprecipitation, immunostaining, immunoelectrophoresis, flow cytometry, microscopy (e.g., confocal microscopy), and mass spectrometry (e.g., GC-MS, LC-MS, MALDI-TOF-MS). Non-limiting examples of methods for detecting a nucleic acid described herein include polymerase chain reaction (PCR), such as quantitative PCR or real-time qPCR (RT-qPCR), fluorescence in situ hybridization (FISH), microarray, RNA-seq, DNA sequencing (e.g., next-generation sequencing), and DNA gel electrophoresis.

[0114] Agricultural Methods

[0115] In some embodiments, a method comprises producing, planting, and / or growing a genetically engineered plant or a seed thereof. In some embodiments, a method comprises introducing one or more nucleic acids into at least one plant cell using a method described herein, thereby transforming the at least one plant cell. In some embodiments, a method comprises selecting transformants (e.g., by selecting for a marker or reporter), and then expanding the transformations in plant tissue culture. In some embodiments, a method comprising isolating cell populations comprising transformants and cultivating the cell populations by transferring them to growth media and / or to soil. In some embodiments, a method comprises subjecting cell populations in soil to conditions capable of promoting plants to flower and produce seeds. In some embodiments, a method comprises harvesting seeds from the engineered plants. In some embodiments, seeds are contacted with soil (e.g., fresh soil which was not used to produce seeds from engineered plants) and subjected to conditions capable of germinating the seeds and cultivating the resulting plants. In some embodiments, soil is chosen, supplemented, and / or formulated to be suitable for growing of an edible plant. In some embodiments, the edible plant is a cereal plant and the soil is suitable for planting a cereal plant or a seed thereof. In some embodiments, a method described herein comprises the use of one or more plant cells, one or more plant seeds, and / or one or more plants that comprise a nucleic acid comprising at least one nif cluster gene. In some embodiments, the nucleic acid comprises a partial or a complete nif cluster. In some embodiments, the nucleic acid comprises a refactored nif cluster. In some embodiments, a method comprises use of the nucleic acid for providing fixed nitrogen to the cereal plant, the cereal plant seed, and / or to the soil thereof.

[0116] Methods of Administration

[0117] In some embodiments, a method comprises administering a peptide or protein expressed from a nucleic acid (e.g., a nucleic acid in a chloroplast) described herein to one or more cells, such as cells in a subject.

[0118] A “subject” to which administration of a peptide or protein is contemplated refers to a human (e.g., a human of any age group, including a pediatric subject, such as an infant, child, or adolescent, or adult subject, such as a young adult, middle-aged adult, or senior adult) or nonhuman animal. In some embodiments, the non-human animal is a mammal (e.g., primate (e.g., cynomolgus monkey or rhesus monkey), commercially relevant mammal (e.g., cattle, pig, horse, sheep, goat, cat, or dog), or bird (e.g., commercially relevant bird, such as chicken, duck, goose, or turkey)). The non-human animal may be at any stage of development. The non-human animal may be a transgenic animal or genetically engineered animal. In some embodiments, a subject is a “subject in need thereof’ which refers to a subject (e.g., a human subject) having, at risk of having, previously had, or is suspected of having a disease, disorder, or condition.

[0119] In some embodiments, a peptide or a protein administered in a method described herein is a therapeutic peptide or a therapeutic protein. In some embodiments, the therapeutic peptide or the therapeutic protein are administered to a subject having or suspected of having a disease, disorder, or condition. In some embodiments, the therapeutic peptide or the therapeutic protein is therapeutic for the disease, disorder, or the condition.

[0120] In some embodiments, administration of a therapeutic peptide or a therapeutic protein may be used to treat a subject in need thereof. The terms “treatment,” “treat,” and “treating” refer to reversing, alleviating, delaying the onset of, or inhibiting the progress of a disease, disorder, or condition. In some embodiments, treatment may be administered after one or more signs or symptoms of the disease have developed or have been observed. In other embodiments, treatment may be administered in the absence of signs or symptoms of the disease. For example, treatment may be administered to a susceptible subject prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of exposure to a pathogen). Treatment may also be continued after symptoms have resolved, for example, to delay or prevent recurrence.

[0121] In some embodiments, administration of a therapeutic peptide or a therapeutic protein achieves one, two, three, four, or more of the following effects, including, for example: (i) reduction or amelioration the severity of disease, disorder, or condition or symptom associated therewith; (ii) reduction in the duration of a symptom associated with a disease, disorder, or condition; (iii) protection against the progression of a disease or disorder or symptom associated therewith; (iv) regression of a disease, disorder, or condition or symptom associated therewith; (v) protection against the development or onset of a symptom associated with a disease, disorder, or condition; (vi) protection against the recurrence of a symptom associated with a disease; (vii) reduction in the hospitalization of a subject; (viii) reduction in the hospitalization length; (ix) an increase in the survival of a subject with a disease; (x) a reduction in the number of symptoms associated with a disease, disorder, or condition; (xi) an enhancement, improvement, supplementation, complementation, or augmentation of the prophylactic or therapeutic effect(s) of another therapy.

[0122] In some embodiments, administration of a therapeutic peptide or a therapeutic protein is performed intravenously, subcutaneously, intraocularly, intravitreally, parenterally, subcutaneously, intravenously, intracerebro-ventricularly, intramuscularly, intracranially, intrathecally, orally, intraperitoneally, or by oral or nasal inhalation, or by direct injection to one or more cells, tissues, or organs. In some embodiments, direct injection is performed concurrently with a surgical procedure or interventional procedure. In general, the most appropriate route of administration will depend upon a variety of factors including the nature of the therapeutic peptide or the therapeutic protein (e.g., its stability in the environment of the gastrointestinal tract), and / or the condition of the subject (e.g., whether the subject is able to tolerate oral administration, injection, etc.). In some embodiments, a therapeutic peptide or a therapeutic protein is administered to a subject through only one administration route. In some embodiments, multiple administration routes may be exploited (e.g., serially, or simultaneously) for administration of a therapeutic peptide or a therapeutic protein to a subject.

[0123] In some embodiments, a therapeutic peptide or a therapeutic protein is administered to a subject in an effective amount. In some embodiments, an effective amount is a “therapeutically effective amount” which refers to an amount sufficient to provide a therapeutic benefit in the treatment of a disease, disorder, or condition or to delay or minimize one or more symptoms associated with the disease, disorder, or condition. In some embodiments, a therapeutically effective amount means an amount of a therapeutic peptide or a therapeutic protein, alone or in combination with other therapies, which provides a therapeutic benefit in the treatment of the disease, disorder, or condition. In some embodiments, a therapeutically effective amount can be an amount that improves overall therapy, reduces or avoids symptoms, signs, or causes of the disease, disorder, or condition, and / or enhances the therapeutic efficacy of another therapeutic cargo. In some embodiments, an effective amount is an amount effective for producing an immunogenic response against an antigen comprised in a therapeutic peptide or a therapeutic protein.

[0124] Compositions

[0125] Agricultural Compositions

[0126] In some aspects, the disclosure relates to agricultural compositions. In some embodiments, agricultural compositions of the present disclosure comprise one or more nucleic acids, one or more plant cells (e.g., a plant cell population), one or more seeds, and / or one or more plants described herein in addition to an agricultural carrier. In some embodiments, an agricultural composition comprises one or more nucleic acids described herein which is comprised in a biological system described herein, such as one that is useful for delivering the one or more nucleic acids to a plant cell, a plant seed, a plant, or soil thereof (e.g., an Agrobacterium, a plant virus, etc.).

[0127] In some embodiments, an agricultural composition comprises at least 1,000 seeds comprising a nucleic acid described herein, for example, at least 5,000 seeds, at least 10,000 seeds, at least 20,000 seeds, at least 30,000 seeds, at least 50,000 seeds, at least 70,000 seeds, at least 80,000 seeds, at least 90,000 seeds, 100,000 seeds, or more than 100,000 seeds. In some embodiments, agricultural compositions comprise a discrete weight of seeds which comprise a nucleic acid described herein, for example, at least 1 lb, at least 2 lbs, at least 5 lbs, at least 10 lbs, at least 30 lbs, at least 50 lbs, at least 70 lbs, or more than 70 lbs in weight. In some embodiments, agricultural compositions comprise a discrete weight of seeds which comprise a nucleic acid described herein, for example, at least 1 kg, at least 2 kgs, at least 5 kgs, at least 10 kgs, at least 30 kgs, at least 50 kgs, at least 70 kgs, or more than 70 kgs in weight. In some embodiments, at least 10%, for example, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, or more than 95% of the seeds in the population are engineered to comprise a nucleic acid described herein.

[0128] In some embodiments, an agricultural composition comprises a substantially uniform population of plants which are engineered to comprise a nucleic acid described herein. In some embodiments, an agricultural composition comprises at least 100 plants, for example, at least 300 plants, at least 1,000 plants, at least 3,000 plants, at least 10,000 plants, at least 30,000 plants, at least 100,000 plants, or more than 100,000 plants.

[0129] In some embodiments, an agricultural composition comprises one or more carriers which are suitable for combining with a plant cell, a plant, or a seed. In some embodiments, agricultural carriers comprise a solid carrier. In some embodiments, agricultural carriers comprise a solid, such as diatomaceous earth, loam, silica, alginate, clay, bentonite, vermiculite, seed cases, other plant and animal products, or combinations, including granules, pellets, wheat, bran, vermiculite, clay, talc, bentonite, diatomaceous earth, fuller's earth, mineral carriers such as kaolin clay, pyrophyllite, bentonite, montmorillonite, diatomaceous earth, acid white soil, vermiculite, and pearlite, and inorganic salts such as ammonium sulfate, ammonium phosphate, ammonium nitrate, urea, ammonium chloride, and calcium carbonate. In some embodiments, organic fine powders, such as wheat flour, wheat bran, and rice bran may be used.

[0130] In some embodiments, agricultural compositions comprises a liquid carrier. In some embodiments, liquid carriers include vegetable oils such as soybean oil and cottonseed oil, glycerol, ethylene glycol, polyethylene glycol, propylene glycol, polypropylene glycol, or suspensions. In some embodiments, agricultural carriers comprise a liquid solution or suspension. In some embodiments, agricultural carriers comprise liquid diluents, aqueous solutions, petroleum distillates, or other liquid carriers. In some embodiments, agricultural carriers comprise a surfactant. Non-limiting examples of surfactants include nitrogen-surfactant blends such as Prefer 28 (Cenex), Surf-N(US), Inhance (Brandt), P-28 (Wilfarm) and Patrol (Helena); esterified seed oils include Sun-It II (AmCy), MSO (UAP), Scoil (Agsco), Hasten (Wilfarm) and Mes-100 (Drexel); and organo- silicone surfactants include Silwet L77 (UAP), Silikin (Terra), Dyne-Amic (Helena), Kinetic (Helena), Sylgard 309 (Wilbur- Ellis), and Century (Precision).

[0131] In some embodiments, agricultural carriers comprise seeds, seed coats, granular carriers, soil, solid carriers, liquid slurry carriers, and / or liquid suspension carriers. In some embodiments, agricultural carriers comprise biologically compatible dispersing agents, such as nonionic, anionic, amphoteric, or cationic dispersing and emulsifying agents. In some embodiments agricultural carriers comprise wetting agents, a synthetic surfactant, a water-in-oil emulsion, a wettable powder, granules, gels, agar strips or pellets, thickeners, microencapsulated particles, and / or liquids, such as aqueous suspensions.

[0132] In some embodiments, agricultural carriers comprise a tackifier or adherent, such as agents useful for combining plant cells, seeds, and / or plants comprising a nucleic acid described herein with other compounds to yield a coating composition. In some embodiments, such compositions help create coatings around the plant or seed to maintain contact with, for example, a bacterium, such as a endosymbiotic bacterium. Non-limiting examples of agents that may be comprised in agricultural compositions include alginate, gums, starches, lecithins, formononetin, polyvinyl alcohol, alkali formononetinate, hesperetin, polyvinyl acetate, cephalins, Gum Arabic, Xanthan Gum, Mineral Oil, Polyethylene Glycol (PEG), Polyvinyl pyrrolidone (PVP), Arabino-galactan, Methyl Cellulose, PEG 400, Chitosan, Polyacrylamide, Polyacrylate, Polyacrylonitrile, Glycerol, Triethylene glycol, Vinyl Acetate, Gellan Gum, Polystyrene, Polyvinyl, Carboxymethyl cellulose, Gum Ghatti, and polyoxyethylenepolyoxybutylene block copolymers.

[0133] In some embodiments, agricultural carriers are prepared from wettable powders, granules, gels, agar strips or pellets, thickeners, microencapsulated particles, grain or legume products, ground grain or beans, broth or flour derived from grain or beans, starch, sugar, or oil.

[0134] In some embodiments, agricultural carriers comprise a microbial stabilizer, such as a desiccant. As used herein, a "desiccant" can include any compound or mixture of compounds that can be classified as a desiccant regardless of whether the compound or compounds are used in such concentrations that they in fact have a desiccating effect on the liquid inoculant. In some embodiments, such desiccants are compatible with the genetically modified bacteria and promote the ability of the bacteria to survive application on the seeds and to survive desiccation. Non-limiting examples of suitable desiccants include one or more of trehalose, sucrose, glycerol, methylene glycol, non-reducing sugars, and sugar alcohols (e.g., mannitol or sorbitol).

[0135] In some embodiments, agricultural carriers comprise one or more agents, such as a fungicide, an antibacterial agent, an herbicide, a nematicide, an insecticide, a plant growth regulator, a rodenticide, and a nutrient. In some embodiments, such agents are compatible with a seed onto which the formulation is applied.

[0136] Pharmaceutical Compositions

[0137] In some aspects, the disclosure relates to pharmaceutical compositions. In some embodiments, a pharmaceutical composition comprises a therapeutic peptide or a therapeutic protein described herein. In some embodiments, a pharmaceutical composition is useful in a method of treating a subject described herein.

[0138] In some embodiments, a pharmaceutical composition comprises a pharmaceutical excipient. Pharmaceutically acceptable excipients (excipients) are substances other than a therapeutic peptide or a therapeutic protein that are intentionally included in a delivery system. In some embodiments, excipients do not exert or are not intended to exert a therapeutic effect. In some embodiments, excipients may act to a) aid in processing of a therapeutic peptide or a therapeutic protein assembly delivery system during manufacture, b) protect, support or enhance stability, bioavailability or patient acceptability of the API, c) assist in product identification, and / or d) enhance any other attribute of the overall safety, effectiveness, or delivery of a therapeutic peptide or a therapeutic protein during storage or use. In some embodiments, a pharmaceutically acceptable excipient may be an inert substance. In some embodiments, a pharmaceutically acceptable excipient may not be an inert substance. Excipients include, but are not limited to, absorption enhancers, anti-adherents, anti-foaming agents, anti-oxidants, binders, buffering agents, carriers, coating agents, colors, delivery enhancers, delivery polymers, dextran, dextrose, diluents, disintegrants, emulsifiers, extenders, fillers, flavors, glidants, humectants, lubricants, oils, polymers, preservatives, saline, salts, solvents, sugars, suspending agents, sustained release matrices, sweeteners, thickening agents, tonicity agents, vehicles, waterrepelling agents, and wetting agents. In some embodiments, a pharmaceutical composition further comprises additional components commonly found in pharmaceutical compositions. Non-limiting examples of additional components of a pharmaceutical composition include can include, but are not limited to: anti-pruritic s, astringents, local anesthetics, an anti-proliferative compound, an anti-cancer compound, an anti-angiogenesis compound, a steroidal or nonsteroidal anti-inflammatory compound, an immunosuppressant compound, an anti-bacterial compound, an anti-viral compound, a cardiovascular compound, a cholesterol-lowering compound, an anti-diabetic compound, an anti-allergic compound, a contraceptive compound, an pain-relieving compound, an anesthetic compound, an anti-coagulant compound, an enzymeinhibiting compound, a steroidal compound, or an analgesic compound.

[0139] Pharmaceutical compositions of the present disclosure may be suitable for treatment regimens and thereby administered to a subject via a variety of methods described herein. Such compositions may be formulated for use in a variety of therapies, such as, in the amelioration, prevention, and / or treatment of conditions for which a peptide or a protein comprised in the composition is therapeutic. Accordingly, pharmaceutical compositions described herein may be administered to a subject such as human or non-human subjects, a cell in situ in a subject, a cell ex vivo, a cell derived from a subject, or a biological sample (e.g., one derived from a subject).

[0140] For administration of an injectable aqueous solution, the pharmaceutical composition may be suitably buffered, if necessary, and the liquid diluent first rendered isotonic with sufficient saline, polyalcohols, or glucose. For example, one dosage of a therapeutic peptide or a therapeutic protein may be dissolved in an isotonic NaCl solution and optionally added to a larger volume of hypodermoclysis fluid prior to being injected at the proposed site of infusion. In some embodiments, a pharmaceutical composition is provided in a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, and liquid polyethylene glycol), and suitable mixtures thereof. A pharmaceutical composition may also comprise adjuvants such as preservatives, wetting agents, emulsifying agents, and dispersing agents.

[0141] Kits

[0142] In some aspects, the disclosure relates to kits. In some embodiments, a kit comprises one or more nucleic acids, one or more plant cells (e.g., a plant cell population), one or more seeds, one or more plants, and / or one or more compositions described herein. Accordingly, kits provided by the present disclosure may comprise any or all of the materials necessary for: producing a nucleic acid described herein; engineering a plant cell, a seed, or a plant described herein; producing a peptide or protein encoded by a nucleic acid described herein; and / or performing a method described herein.

[0143] In some embodiments, the kits described herein may include one or more containers housing components for performing the methods described herein, and optionally instructions for use. In some embodiments, the components may be prepared sterilely, packaged in a syringe, and shipped refrigerated. Alternatively, in some embodiments, they may be housed in a vial or other container for storage. In some embodiments, a second container may have other components prepared sterilely. Alternatively, in some embodiments, the kits may include the active agents premixed and shipped in a vial, tube, or other container. In some embodiments, the kits may also include other components, depending on the specific application, for example, containers, cell media, salts, buffers, reagents, syringes, needles, a fabric, such as gauze, for applying or removing a disinfecting agent, disposable gloves, a support for agents (e.g., nucleic acids, peptides, proteins, and / or cells) prior to administration, etc.

[0144] In some embodiments, any of the kits described herein may further comprise components needed for inducing uptake of a nucleic acid, peptide, or protein into a cell. In some embodiments, each component of the kits, where applicable, may be provided in liquid form (e.g., in solution) or in solid form, (e.g., a dry powder). In some embodiments, some of the components may be reconstitutable or otherwise processible (e.g., to an active form), for example, by the addition of a suitable solvent or other species (for example, water), which may or may not be provided with the kit.

[0145] In some embodiments, a kit further comprises a set of instructions for carrying out the methods described herein. As used herein, “instructions” can define a component of instruction and / or promotion, and typically involve written instructions on or associated with packaging of this disclosure. In some embodiments, instructions also can include any oral or electronic instructions provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), Internet, and / or web-based communications, etc. In some embodiments, the written instructions may be in a form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which can also reflect approval by the agency of manufacture, use or sale for animal administration. As used herein, “promoted” includes all methods of doing business including methods of education, hospital and other clinical instruction, scientific inquiry, drug discovery or development, academic research, pharmaceutical industry activity including pharmaceutical sales, and any advertising or other promotional activity including written, oral, and electronic communication of any form, associated with this disclosure.

[0146] Additionally, in some embodiments, the kits may include other components depending on the specific application, as described herein. In some embodiments, the kits may have a variety of forms, such as a blister pouch, a shrink-wrapped pouch, a vacuum sealable pouch, a sealable thermoformed tray, or a similar pouch or tray form, with the accessories loosely packed within the pouch, one or more tubes, containers, a box, or a bag. In some embodiments, the kits may be sterilized after the accessories are added, thereby allowing the individual accessories in the container to be otherwise unwrapped. In some embodiments, the kits, or any of its components, can be sterilized using any appropriate sterilization techniques, such as radiation sterilization, heat sterilization, or other sterilization methods known in the art.

[0147] General Techniques

[0148] The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, which are within the skill of the art. Such techniques are explained fully in the literature, such as the following references which are incorporated by reference herein for their disclosures related to methods: Molecular Cloning: A Laboratory Manual, second edition (Sambrook, et al., 1989) Cold Spring Harbor Press; Oligonucleotide Synthesis (M. J. Gait, ed. 1984); Methods in Molecular Biology, Humana Press; Cell Biology: A Laboratory Notebook (J. E. Cellis, ed., 1989) Academic Press; Animal Cell Culture (R. I. Freshney, ed. 1987); Introuction to Cell and Tissue Culture (J. P. Mather and P. E. Roberts, 1998) Plenum Press; Cell and Tissue Culture: Laboratory Procedures (A. Doyle, J. B. Griffiths, and D. G. Newell, eds. 1993-8) J. Wiley and Sons; Methods in Enzymology (Academic Press, Inc.); Handbook of Experimental Immunology (D. M. Weir and C. C. Blackwell, eds.): Gene Transfer Vectors for Mammalian Cells (J. M. Miller and M. P. Calos, eds., 1987); Current Protocols in Molecular Biology (F. M. Ausubel, et al. eds. 1987); PCR: The Polymerase Chain Reaction, (Mullis, et al., eds. 1994); Current Protocols in Immunology (J. E. Coligan et al., eds., 1991); Short Protocols in Molecular Biology (Wiley and Sons, 1999); Immunobiology (C. A. Janeway and P. Travers, 1997); Antibodies (P. Finch, 1997); Antibodies: a practice approach (D. Catty., ed., IRL Press, 1988-1989); Monoclonal antibodies: a practical approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000); Using antibodies: a laboratory manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999); The Antibodies (M. Zanetti and J. D. Capra, eds. Harwood Academic Publishers, 1995); DNA Cloning: A practical Approach, Volumes I and II (D.N. Glover ed. 1985); Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985; Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984»; Animal Cell Culture (R.I. Lreshney, ed. (1986; Immobilized Cells and Enzymes (1RL Press, (1986; and B. Perbal, A practical Guide To Molecular Cloning (1984); L.M. Ausubel et al. (eds.).

[0149] EXAMPLES

[0150] Example 1: Tuning chloroplast transgene expression from a synthetic operon

[0151] This Examples relates to nucleic acids that are responsive to the unique attributes of the chloroplast genetic system. These systems are highly amenable for synthetic biology applications to manipulate traits of agronomic relevance. Many of such applications require multigene expression, with optimal performance determined by balanced expression.

[0152] A repertoire of expression elements is described herein which is useful for expressing genes of interest in chloroplasts. For example, FIGs. 1A-1H show non-limiting embodiments of nucleic acids. FIG. 1A shows a non-limiting embodiment of a nucleic acid comprising at least one gene which is operably linked to a heterologous regulatory sequence. FIG. IB shows nonlimiting embodiments of a translation initiation regulatory sequence comprising a Shine- Dalgamo (SD) sequence and a translation initiation codon which are separated by a spacer sequence comprising a variable number of nucleotides. FIG. 1C shows non-limiting embodiments of a protein-binding site comprising a pentatricopeptide repeat (PPR) proteinbinding site for PPR 10 and PPR10GG. FIG. ID shows non-limiting embodiments of nucleic acids comprising at least one gene which is operably linked to one or more heterologous regulatory sequences. FIG. IE shows non-limiting embodiments of nucleic acids comprising at least one gene which is operably linked to a heterologous regulatory sequence and a promoter, wherein the promoter is either native or heterologous to the at least one gene. FIG. IF shows a non-limiting embodiment of a nucleic acid comprising two genes which are each operably linked to a heterologous regulatory sequence. FIG. 1G shows non-limiting embodiments of nucleic acids comprising two genes which are each operably linked to a plurality of regulatory sequences. FIG. 1H shows non-limiting embodiments of nucleic acids comprising two genes which are each operably linked to a plurality of regulatory sequences, including promoters, translation initiation regulatory sequences, and / or PPR protein-binding sites.

[0153] Nucleic acids were generated that incrementally up / down-regulate the expression of a gene by tuning the translation initiation rate through manipulation of a ribosome binding site in the tobacco chloroplast. Spacing lengths between the Shine-Dalgarno (SD) sequence and the translation initiation codon (ATG) were modified to alter the translation initiation rate of a GFP reporter, leading to a gradient expression of the exogenous gene.

[0154] A series of dicistronic vectors comprising aadA (selectable marker gene conferring resistance to spectinomycin for chloroplast transformation selection) and gfp (reporter gene) were constructed. The vectors were incorporated into tobacco chloroplast genome and transcribed into mRNA, the intergenic region between aadA and gfp harbored RNA nucleotides comprising the PPR10 protein binding site. Under this PPRIO-dependent RNA translational activation construct, expression of GFP was incrementally increased from -13% (spacer 4nt) to -45% (spacer 12nt) and incrementally decreased from -45% (spacer 12nt) to -9% (spacer 18nt). “~%” refers to the GFP abundance out of the total soluble protein in plants (FIG. 2).

[0155] In addition, nucleic acids for multigene expression from bacterial-like, synthetic polycistronic operons in chloroplast were engineered. A series of polycistronic vectors carrying the aadA (selectable marker gene conferring resistance to spectinomycin for chloroplast transformation selection) and BFP, Sapphire, mKO and RFP (reporter genes) were constructed. Once the vectors were incorporated into the chloroplast genome and transcribed into mRNA, the intergenic regions between aadA and bfp, bfp and sapphire, sapphire and mKO, mKO and rfp all harbored RNA nucleotides comprising a native PPR 10 protein binding site and spacer variants. Under this PPRIO-dependent RNA translational activation construct for multigene expression, the expression of all four reporter genes were captured by confocal microscopy, which showed that native PPR10 protein can boost efficient translation of a polycistronic mRNA (FIG. 3).

[0156] A polycistronic vector comprising aadA (selectable marker gene conferring resistance to spectinomycin for chloroplast transformation selection) and BFP, Sapphire, mKO and RFP (reporter genes) were constructed. Once the vectors were incorporated into the chloroplast genome and transcribed into mRNA, the intergenic regions between aadA and bfp, bfp and sapphire, sapphire and mKO, mKO and rfp all harbored RNA nucleotides comprising an engineered PPR10GGprotein cognate binding site and spacer variants. Under this PPR 1 (Independent RNA translational activation construct for multigene expression, the expression of all four reporter genes were captured by confocal microscopy which showed that PPR10GGprotein can boost efficient translation of a polycistronic mRNA (FIG. 4).

[0157] The strategic approach and the results described herein indicated expression of a gene of interest can be modulated via adjusting the spacing length between SD sequence and the ATG codon. Multigene activation by native and engineered PPR proteins or other RNA specific binding proteins can also regulate engineered nucleic acid expression within the plastids (chloroplast within the leaf cells, amyloplast within the root cells, etc.) in other plant species that are amenable for chloroplast engineering.

[0158] Example 2: Expression of Nif Cluster Genes in Engineered Plant Cells

[0159] This Example relates to engineering plant cells to express nif cluster genes from sequences encoded by the chloroplast genome. This two-component system relies on 1) nucleus- encoded PPR 1 OGG protein and 2) chloroplast genome-encoded 4 nif genes (nifH, nifM, nifU, nifS plus an aadA selectable marker gene) constructs. A 5' UTR of each zzz / gcnc construct contains the PPR 1 OGG protein binding site (atpHGG site). The translation of each Nif protein from the mRNA in the chloroplast is activated by the PPR10GG protein binding to the 5' UTRs via the atpHGG site.

[0160] FIG. 5A is a diagram of nitrogenase formation involving the nifS, nifU, nifM, and nifH genes. FIG. 5B is a diagram of a genetic engineering strategy in plant cells, wherein nifS, nifU, nifM, and nifH genes are integrated into the chloroplast genome and each engineered to comprise a 5' UTR that comprises a binding site for a PPR10GGprotein encoded by the nuclear genome of the plant cell. Except for the wild-type plants, the two plants comprising chloroplast genome-encoded nif genes can transcribe the mRNA encoding nifH, nifS, nifU, and nifM proteins. Translation of each Nif protein from the mRNA in the chloroplast is activated by the PPR10GGprotein binding to the 5' UTRs via the atpHGGsite. Thus, engineered plants comprising nuclear-encoded PPR10GGprotein activate the translation at the mRNA level and produce nifH, nifS, nifU, and nifM proteins.

[0161] FIG. 5C shows images of wild-type plants (plants in left-most column), plants having chloroplasts with genomes that comprise integrated nifS, nifU, nifM, and nifH genes comprising a binding site for a PPR10GGprotein (atpHGG) in their 5' UTR (plants in middle column), and plants having chloroplasts with genomes that comprise integrated nifS, nifU, nifM, and nifH genes comprising a binding site for a PPR10GGprotein (atpHGG) encoded by the nuclear genome in their 5' UTR PPR10GGprotein (plants in right-most column). FIG. 5D shows results from Southern blot analyses of samples from two plants corresponding to the right-most column of plants in FIG. 5C and indicates engineered nif genes are integrated into the chloroplast genome. FIG. 5E shows results from Northern blot analyses of samples from two plants corresponding to the right-most column of plants in FIG. 5C and performed using probes for detecting RNA encoded by an aadA gene (a selectable marker), a nifU gene, a nifS gene, and a nifM gene. The results indicate chloroplast genome-encoded nif genes are transcribed in engineered plant cells.

[0162] Alternative Embodiments

[0163] Embodiment 1. A nucleic acid comprising at least one gene operably linked to a translation initiation regulatory sequence positioned 5' to the at least one gene and / or a pentatricopeptide repeat (PPR) protein-binding sequence positioned 5' to the at least one gene, wherein the translation initiation regulatory sequence comprises a Shine-Dalgarno (SD) sequence and a translation initiation codon which are separated by a spacer sequence, wherein the spacer sequence is a nucleotide sequence of 1-100 nucleotides in length, optionally wherein the spacer sequence is 5-20, 11-20, 12-20, 13-20, 14-20, or 12 nucleotides in length, and is heterologous to the SD sequence, optionally wherein the spacer sequence comprises the following:

[0164] 1) nucleotides that do not form a secondary structure,

[0165] 2) nucleotides that do not form a secondary structure with an upstream mRNA stem-loop secondary structure region optionally which comprises a native or an engineered RNA-specific binding protein binding site,

[0166] 3) nucleotides free of a canonical or a non-canonical translation initiation codon, and optionally

[0167] 4) a sequence which does not comprise a G nucleotide, and / or

[0168] 5) an AT or AU-rich nucleotide sequence.

[0169] INCORPORATION BY REFERENCE

[0170] The present application refers to various issued patent, published patent applications, scientific journal articles, and other publications, all of which are incorporated herein by reference. The details of one or more embodiments of the invention are set forth herein. Other features, objects, and advantages of the invention will be apparent from the Detailed Description, the Figures, the Examples, and the Claims.

[0171] EQUIVALENTS

[0172] While several inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.

[0173] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0174] All references, patents and patent applications disclosed herein are incorporated by reference with respect to the subject matter for which each is cited, which in some cases may encompass the entirety of the document.

[0175] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0176] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0177] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0178] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0179] It should also be understood that, unless clearly indicated to the contrary, in any methods claimed herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.

[0180] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03. It should be appreciated that embodiments described in this document using an open-ended transitional phrase (e.g., “comprising”) are also contemplated, in alternative embodiments, as “consisting of’ and “consisting essentially of’ the feature described by the open-ended transitional phrase. For example, if the disclosure describes “a composition comprising A and B”, the disclosure also contemplates the alternative embodiments “a composition consisting of A and B” and “a composition consisting essentially of A and B”.

Claims

CLAIMSWhat is claimed is:

1. A nucleic acid comprising at least one gene operably linked to a translation initiation regulatory sequence positioned 5' to the at least one gene and / or a pentatricopeptide repeat (PPR) protein-binding sequence positioned 5' to the at least one gene, wherein the translation initiation regulatory sequence comprises a Shine-Dalgamo (SD) sequence and a translation initiation codon which are separated by a spacer sequence, wherein the spacer sequence is a nucleotide sequence of 1-100 nucleotides in length, optionally wherein the spacer sequence is 5- 20, 11-20, 12-20, 13-20, 14-20, or 12 nucleotides in length, and is heterologous to the SD sequence, optionally wherein the spacer sequence is a sequence which does not comprise a G nucleotide, and / or an AT or AU-rich nucleotide sequence, optionally wherein the spacer sequence comprises the following:1) nucleotides that do not form a secondary structure,2) nucleotides that do not form a secondary structure with an upstream mRNA stem-loop secondary structure region optionally which comprises a native or an engineered RNA- specific binding protein binding site, and / or3) nucleotides free of a canonical or a non-canonical translation initiation codon.

2. The nucleic acid of claim 1, wherein the at least one gene is operably linked to the translation initiation regulatory sequence and the PPR protein-binding sequence.

3. The nucleic acid of claim 1 or 2, wherein the at least one gene comprises a first gene linked to a first PPR protein-binding sequence and / or a first translation initiation regulatory sequence and a second gene linked to a second PPR protein-binding sequence and / or a second translation initiation regulatory sequence.

4. The nucleic acid of claim 3, wherein the first PPR protein-binding sequence and the second PPR protein-binding sequence are the same and wherein the PPR-protein binding sequence is a PPRIO-binding site or a PPR10GG-binding site.

5. The nucleic acid of claim 3, wherein the first PPR protein-binding sequence and the second PPR protein-binding sequence are different from one another.

6. The nucleic acid of claim 5, wherein the first PPR-protein binding sequence is a PPR 10- binding site and the second PPR-protein binding site is a PPR10GG-binding site.

7. The nucleic acid of any one of claims 3-6, wherein a first spacer sequence of the first translation initiation regulatory sequence operably linked to the first gene and a second spacer sequence of the second translation initiation regulatory sequence operably linked to the second gene are the same.

8. The nucleic acid of any one of claims 3-6, wherein a first spacer sequence of the first translation initiation regulatory sequence operably linked to the first gene and a second spacer sequence of the second translation initiation regulatory sequence operably linked to the second gene differ by at least one nucleotide.

9. The nucleic acid of any one of claims 3-8, wherein a first spacer sequence of the first translation initiation regulatory sequence operably linked to the first gene and a second spacer sequence of the second translation initiation regulatory sequence operably linked to the second gene comprise 2-20 nucleotides in length.

10. The nucleic acid of any one of claims 1-9, wherein the at least one gene is operably linked to a promoter.

11. The nucleic acid of claim 10, wherein the promoter is native to the at least one gene.

12. The nucleic acid of claim 10, wherein the promoter is heterologous to the at least one gene.

13. The nucleic acid of any one of claims 3-9, wherein the first gene and the second gene are operably linked to a single promoter.

14. The nucleic acid of any one of claims 3-9, wherein the first gene and the second gene are operably linked to respective promoters.

15. The nucleic acid of any one of claims 1-14, wherein the at least one gene comprises at least one of: a gene involved in a plant biosynthetic pathway; a gene which is capable of stimulating plant growth; a gene encoding an insecticide; a gene encoding a pesticide; or a gene encoding a peptide or protein which is capable of modifying a molecule produced by a plant pathogen.

16. The nucleic acid of any one of claims 1-15, wherein the at least one gene comprises a nif cluster gene.

17. The nucleic acid of any one of claims 1-15, wherein the at least one gene encodes a therapeutic peptide or a therapeutic protein or a reporter.

18. The nucleic acid of claim 1, wherein the nucleic acid comprises from 5’ to 3’ the PPR protein-binding sequence, the translation initiation regulatory sequence comprised of the SD sequence, a spacer, and the translation initiation codon and the gene and optionally the 3’ end of the gene is linked to a nucleic acid comprising from 5’ to 3’ a second PPR protein-binding sequence, a second translation initiation regulatory sequence comprised of a second SD sequence, a second spacer, and a second translation initiation codon and a second gene.

19. A plant cell, or plant cell population thereof, comprising the nucleic acid of any one of claims 1-18.

20. The plant cell, or the plant cell population thereof, of claim 19, wherein the plant cell is a cell of an edible plant.21 The plant cell, or the plant cell population thereof, of claim 19 or 20, wherein the plant cell is a cereal plant cell.

22. The plant cell, or the plant cell population thereof, of any one of claims 19-21, wherein the nucleic acid is comprised in at least one chloroplast of the plant cell.

23. The plant cell, or the plant cell population thereof, of any one of claims 19-22, wherein the plant cell further comprises the PPR-binding protein.

24. A plant, or a seed thereof, comprising the nucleic acid of any one of claims 1-18 or the plant cell, or the plant cell population thereof, of any one of claims 19-23.

25. The plant, or the seed thereof, of claim 24, wherein the plant is an edible plant.

26. The plant, or the seed thereof, of claim 24 or 25, wherein the plant is a cereal plant.

27. A method comprising contacting one or more plant cells with the nucleic acid of any one of claims 1-18.

28. The method of claim 27, wherein the one or more plant cells comprise one or more cells of an edible plant.

29. A method for providing fixed nitrogen to a cereal plant, wherein the method comprises contacting one or more cells of the cereal plant with the nucleic acid of claim 1 .

30. A composition comprising the nucleic acid of any one of claims 1-18, the plant cell, or the plant cell population thereof, of any one of claims 19-23, or the plant, or the seed thereof, of any one of claims claim 24-26.

31. A kit comprising the nucleic acid of any one of claims 1-18, the plant cell, or the plant cell population thereof, of any one of claims 19-23, the plant, or the seed thereof, of any one of claims claim 24-26, or the composition of claim 30.

32. A method for promoting gene expression in an engineered plant cell, comprising contacting one or more engineered plant cells comprising a nuclear-encoded pentatricopeptiderepeat (PPR10GG) protein with a nucleic acid comprising at least one gene operably linked to a translation initiation regulatory sequence comprising a pentatricopeptide repeat (PPR10GG) protein-binding sequence which is positioned within a 5' UTR and 5' to the at least one gene to promote expression of the at least one gene.

33. The method of claim 32, wherein the translation initiation regulatory sequence comprises a Shine-Dalgamo (SD) sequence and a translation initiation codon which are separated by a spacer sequence, wherein the spacer sequence is a nucleotide sequence of 1-100 nucleotides, optionally wherein the spacer sequence is 5-20, 11-20, 12-20, 13-20, 14-20, or 12 nucleotides in length, and is heterologous to the SD sequence, optionally wherein the spacer sequence is a sequence which does not comprise a G nucleotide, and / or an AT or AU-rich nucleotide sequence, and optionally wherein the spacer sequence comprises the following:1) nucleotides that do not form a secondary structure,2) nucleotides that do not corrupt an upstream mRNA stem-loop secondary structure which comprises a native or an engineered RNA- specific binding protein binding site, and / or3) nucleotides free of a canonical or a non-canonical translation initiation codon.

Citation Information

Patent Citations

  • Method for gene expression

    US20080096256A1

  • FUSION PROTEIN FOR IMPROVING PROTEIN EXPRESSION FROM TARGET mRNA

    US20190309029A1

  • Genome editing method

    US20200029538A1

  • Compositions and methods for improving plastid transformation efficiency in higher plants

    US20220220493A1