Compositions and methods for alteration of seed fatty acid composition in plants
Genomic modifications in the FATB gene to alter substrate specificity in soybean seeds increase MCFA production, addressing the low MCFA content in soybean oil for biofuel and sustainable aviation fuel applications.
Patent Information
- Application Number
- PCT/US2025/036156
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Soybean oil produced in the US contains negligible amounts of medium-chain fatty acids (MCFAs), limiting its use in biofuel and sustainable aviation fuel production.
Introduce genomic modifications in the fatty acyl-ACP thioesterase B (FATB) gene to encode a modified FATB polypeptide with altered substrate specificity, increasing the production of medium-chain fatty acids in soybean seeds.
The modified FATB polypeptide enhances the content of medium-chain fatty acids in soybean seeds, producing seeds with increased MCFA content compared to control seeds.
Smart Images

Figure US2025036156_08012026_PF_FP_ABST
Abstract
Description
Docket # 212436-WO-SEC-1 COMPOSITIONS AND METHODS FOR ALTERATION OF SEED FATTY ACID COMPOSITION IN PLANTS REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0001] The official copy of the sequence listing is submitted electronically via Patent Center as an XML formatted sequence listing with a file named 212436A_SequenceListing created on July 1, 2025, and having a size of 354,387 bytes and is filed concurrently with the specification. The sequence listing comprised in this XML formatted document is part of the specification and is herein incorporated by reference in its entirety. BACKGROUND
[0002] Plant oils are a major product of oil seed crops such as soybean, sunflower, and canola. Oils, such as soybean oil, produced in the US are extracted from seeds and have a major use in food products such as cooking oils, shortenings, and margarines. There is a growing interest in medium-chain fatty acids (MCFAs) like C12:0, ideal for biofuel and sustainable aviation fuel (SAF) production. However, soybean oil contains negligible amounts of MCFAs.
[0003] Accordingly, there is a need to develop compositions and methods to increase MCFA content. This disclosure provides such compositions and methods. SUMMARY
[0004] Provided are seeds, plants, and plant cells comprising an introduced genomic modification in a fatty acyl-ACP thioesterase B (FATB) gene to encode a modified FATB polypeptide, the modified FATB polypeptide having altered substrate specificity as compared to a corresponding non-modified FATB polypeptide. In certain embodiments, the seeds, cells, or seed from the plants have an increased MCFA content as compared to a control seed, cell, or plant.
[0005] Also provided are modified FATB polypeptides and polynucleotides encoding modified FATB polypeptides comprising an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to positionDocket # 212436-WO-SEC-1 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine, , a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof. In certain embodiments, expression of the modified FATB polypeptide in a plant cell increases the medium chain fatty acid content as compared to a control plant cell.
[0006] Further provided are methods of producing a soybean plant producing seeds having increased medium chain fatty acid content comprising introducing into a regenerable soybean plant cell a genomic modification in a fatty acyl-ACP thioesterase B (FATB) gene to encode a modified FATB polypeptide having altered substrate specificity as compared to a corresponding non-modified FATB polypeptide and generating the plant, wherein the plant comprises the genomic modification and produces a seed having increased medium chain fatty acid content as compared to a seed of a control plant not comprising the genomic modification.Docket # 212436-WO-SEC-1
[0007] Also provided are seed oil compositions produced from the seeds, plants and cells described herein along with methods for producing the seed oil compositions. BRIEF DESCRIPTION OF THE DRAWINGS AND THE SEQUENCE LISTING
[0008] The disclosure can be more fully understood from the following detailed description and the accompanying drawings and Sequence Listing, which form a part of this application. The sequence descriptions (Table 1) and sequence listing attached hereto comply with the rules governing nucleotide and amino acid sequence disclosures in patent applications as set forth in 37 C.F.R. §§1.831-1.835.
[0009] Fig.1 provides a representation of a representative soybean thioesterase structure (GmFATB1A). The structure, which covers residues 125-407 of SEQ ID NO: 2, is modeled by AlphaFold 2 with a pLDDT of 90. Secondary structure elements such as β-strands are depicted as arrowed ribbons, and α-helices as helical ribbons. The N-terminal domain consists of βN-β1- α1-α2-β3-β4-β5-β2-α3, and the C-terminal domain comprises β6-α4-α5-β8-β9-β10-β7. The bound palmitic acid is shown as large atomic spheres in the middle, labeled as PALM. The catalytic residues are drawn as sticks at the joint of the N- and C-terminal domains. The lipid length determinant residues are depicted by small spheres and labeled. The substitutions between the long-chain GmFATB1A and the medium-chain GmFATB4A (SEQ ID NO: 14) are shown with the residue volume size attached. For example, A167 / C140 (89 / 109 Å3) means that A167 (89Å3 in GmFATB1A) is replaced by C140 (109Å3 in GmFATB4A).
[0010] Figs.2A-2D provide a sequence alignment between the thioesterase amino acid sequences of EgFATB3 (SEQ ID NO: 86), CnFATB2 (SEQ ID NO: 88), CvFATB2 (SEQ ID NO: 90), GmFATB1A (SEQ ID NO: 2), GmFATB1B (SEQ ID NO: 4), GmFATB2A (SEQ ID NO: 6), GmFATB2B (SEQ ID NO: 8), GmFATB3A (SEQ ID NO: 10), GmFATB3B (SEQ ID NO: 12), GmFATB5A (SEQ ID NO: 18), GmFATB5B (SEQ ID NO: 20), GmFATB4B (SEQ ID NO: 16), GmFATB4A (SEQ ID NO: 14), EgFATB1 (SEQ ID NO: 22), UcFATB1 (SEQ ID NO: 24), CvFATB1 (SEQ ID NO: 26), CvFATB2 (SEQ ID NO: 28), and CnFATB3 (SEQ ID NO: 84). The acyl tunnel residues are highlighted by black stars and the identified acyl chain length determining residues are boxed. The horizontal line separates the LCFA above the line and MCFA below the line. GmFatB4B has weak MCFA characteristics.Docket # 212436-WO-SEC-1 Table 1: Sequence Listing Description Amino NucleotideAcid Name SpeciesSEQ ID NO:SEQ IDDocket # 212436-WO-SEC-1 GM-TE2-CR5 Glycine max 52 Domain swap variant 1 Synthetic construct 53 54 D min w vrint 2 S nthti ntr t 55 56Docket # 212436-WO-SEC-1 without chloroplast transit peptide GmFATB1A_A167C_V219F Synthetic construct 121 122 GmFATB1A A118C V170FDocket # 212436-WO-SEC-1 GmFATB1A_V219F_N225I_L259F_I262L_ D278P_D284N_L291FSynthetic construct 191 192GmFATB1A V170F N176I L210F I213LDETAILED DESCRIPTION
[0011] The present disclosure describes polynucleotides encoding plant (e.g., soybean) fatty acyl-ACP (Acyl Carrier Protein) thioesterase polypeptides that are modified to alter the substrate specificity of the encoded polypeptide. For example, long-chain FATB polypeptides are modified to function as medium-chain FATB polypeptides. The present disclosure also provides plants, plant cells and seeds expressing the modified FATB polypeptides. The present disclosure further describes modified plant seeds (e.g., soybean seeds) having an introduced genomic modification of a gene encoding a fatty acyl-ACP thioesterase and an increase in the amount of medium chain fatty acids as compared to a control seed or plant.Docket # 212436-WO-SEC-1
[0012] Fatty acyl-ACP (Acyl Carrier Protein) thioesterases (FATs or TEs), including FATA and FATB are key enzymes in plant fatty acid biosynthesis. Thioesterases catalyze the hydrolysis of the thioester bond in acyl-ACP, terminating the acyl chain elongation process, and determining the chain length of fatty acid product. FATB thioesterases primarily hydrolyze saturated acyl- ACPs, while FATA thioesterases prefer unsaturated acyl-ACPs. This difference in substrate specificity influences the saturation level and chain length of the produced fatty acids. These enzymes exhibit diversity in their enzymatic specificity and activity, acting on different acyl- ACP substrates of varying carbon lengths. Soybeans, which produce predominantly long-chain (C16-C18) fatty acids, comprise 12 FAT genes. Among these, two genes are FATA type which function in unsaturated fatty acid biosynthesis, while ten genes are FATB type that are specific for saturated fatty acid biosynthesis.
[0013] FATB type thioesterases can also be characterized based on the substrate specificity of the enzyme. Long-chain fatty acid (LCFA) FATB type thioesterases predominantly generate C16:0-C18:0 fatty acids while medium-chain fatty acid (MCFA) FATB type thioesterases predominantly generate C6:0-C14:0 fatty acids.
[0014] As used herein “medium chain fatty acids” “MCFAs” “medium chain length fatty acid” or the like refers to a fatty acids having an acyl chain of 6 to 14 carbons. The acyl chain is preferably saturated but may be modified (e.g., comprise a double bond). Examples of MCFAs include, but are not limited to, hexanoic or caproic acid (C6:0), octanoic or caprylic acid (C8:0), decanoic or capric acid (C10:0), dodecanoic or lauric acid (C12:0) and myristic acid (C14:0).
[0015] As used herein “long chain fatty acids” “LCFAs” “long chain length fatty acid” or the like refers to a fatty acids having an acyl chain of 16 or more carbons. The acyl chain is preferably saturated but may be modified (e.g., comprise a double bond). Examples of LCFAs include, but are not limited to, palmitic acid (C16:0), stearic acid (C18:0), oleic acid (C18:1), and linoleic acid (C18:2).
[0016] Provided are FATB polynucleotides encoding modified FATB polypeptides having altered substrate specificity as compared to a corresponding control FATB polypeptide (e.g., corresponding non-modified FATB polypeptide). In certain embodiments, the modified FATB polypeptide has increased specificity for MCFA production as compared to a corresponding control FATB polypeptide. In certain embodiments, expression of the modified FATB polypeptide in a plant cell or plant seed increases the total MCFA content as compared to aDocket # 212436-WO-SEC-1 control plant cell or plant seed. In certain embodiments, the modified FATB polypeptide comprises at least one amino acid deletion, insertion, or substitution as compared to the corresponding wild-type FATB polypeptide.
[0017] As used herein an “amino acid deletion,” “deletion mutation,” or the like, refers to a mutation in which the indicated amino acid residue is removed from the polypeptide sequence, so that, when aligned to the reference sequence (e.g., corresponding wild-type polypeptide) the mutated sequence does not have an amino acid corresponding to the indicated position of the reference sequence. An “amino acid insertion,” “insertion mutation,” or the like, refers to a mutation in which at least one amino acid residue is added to the polypeptide sequence, so that, when aligned to the reference sequence (e.g., corresponding wild-type polypeptide) the mutated sequence contains an additional amino acid corresponding to the indicated position of the reference sequence.
[0018] An “amino acid substitution,” “substitution mutation,” or the like, refers to a modification in which the indicated amino acid residue is replaced with a different amino acid residue, so that, when aligned to the reference sequence (e.g., corresponding wild-type polypeptide) the mutated sequence does not have the same amino acid at the indicated position. When the amino acid residue is substituted for a residue that has similar properties (e.g., size, charge, and / or hydrophobicity) the substitution is referred to as a conservative amino substitution. Conservative amino acid substitutions are well known in the art. For example, the following six groups contain amino acids that are considered to be conservative substitutions for one another: 1) Alanine (A), Serine (S), Threonine (T); 2) Aspartic acid (D), Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); and 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W). Alternatively, when the amino acid residue is substituted for an amino acid that has dissimilar properties the modification is referred to as a radical amino acid substitution.
[0019] The type of amino acid substitution (i.e., conservative or radical) in the FATB polypeptides provided herein is not particularly limited, so long as the substrate specificity of the modified polypeptide is altered, such that the FATB variant polypeptides provided herein may contain all conservative amino acid substitutions, all radical amino acid substitutions, or a combination of radical and conservative amino acid substitutions.Docket # 212436-WO-SEC-1
[0020] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof.
[0021] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises aDocket # 212436-WO-SEC-1 cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2 and a methionine (M), cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 167 of SEQ ID NO: 2 and a phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2.
[0022] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2 and a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 167 of SEQ ID NO: 2 and a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2.
[0023] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32,Docket # 212436-WO-SEC-1 or 34 and comprises a cysteine (C) at the position corresponding to position 167 of SEQ ID NO: 2 a cysteine (C) at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2.
[0024] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, and an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C) at the position corresponding to position 167 of SEQ ID NO: 2 a cysteine (C) at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 171 of SEQ ID NO: 2, and an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2.
[0025] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, and a leucine (L) at the position corresponding to position 270 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an aminoDocket # 212436-WO-SEC-1 acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C) at the position corresponding to position 167 of SEQ ID NO: 2 a phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, and a leucine (L) at the position corresponding to position 270 of SEQ ID NO: 2.
[0026] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, and an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C) at the position corresponding to position 167 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C) at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2, and an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2.
[0027] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ IDDocket # 212436-WO-SEC-1 NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine (C) at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 174 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2.
[0028] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F)Docket # 212436-WO-SEC-1 or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine (C) at the position corresponding to position 171 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an alanine (A) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2.
[0029] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C) at the position corresponding to position 174 of SEQ ID NO: 2, aDocket # 212436-WO-SEC-1 phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a cysteine (C) at the position corresponding to position 167 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C) at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 219 of SEQ ID NO: 2.
[0030] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a lysine (K) or glutamine (Q) at the position corresponding to position 172 of SEQ ID NO: 2, and a glycine (G) or alanine (A) at the position corresponding to position 179 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a lysine (K) at the position corresponding to position 172 of SEQ ID NO: 2, and a glycine (G) at the position corresponding to position 179 of SEQ ID NO: 2.
[0031] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises an isoleucine (I) at the position corresponding to position 166 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 176 of SEQ ID NO: 2, a glycine (G) or alanine (A) at the position corresponding to position 179 of SEQ ID NO: 2, and a serine (S) at the position corresponding toDocket # 212436-WO-SEC-1 position 180 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises an isoleucine (I) at the position corresponding to position 166 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 176 of SEQ ID NO: 2, an alanine (A) at the position corresponding to position 179 of SEQ ID NO: 2, and a serine (S) at the position corresponding to position 180 of SEQ ID NO: 2.
[0032] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises an proline (P) at the position corresponding to position 101 of SEQ ID NO: 2.
[0033] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 70 of SEQ ID NO: 2 and a methionine (M) at the position corresponding to position 107 of SEQ ID NO: 2.
[0034] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 176 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 210 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 213 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 229 of SEQ ID NO: 2, an asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, and a phenylalanine (F) at the position corresponding to position 242 of SEQ ID NO: 2.Docket # 212436-WO-SEC-1
[0035] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2 and a glycine (G) or alanine (A) at the position corresponding to position 179 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2 and a glycine (G) at the position corresponding to position 179 of SEQ ID NO: 2.
[0036] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a lysine (K) or glutamine (Q) at the position corresponding to position 172 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 176 of SEQ ID NO: 2 and a glycine (G) or alanine (A) at the position corresponding to position 179 of SEQ ID NO: 2. In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a glutamine (Q) at the position corresponding to position 172 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 176 of SEQ ID NO: 2 and a glycine (G) at the position corresponding to position 179 of SEQ ID NO: 2.
[0037] In certain embodiments, the modified FATB polypeptide comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 54, 56, 58, 62, 64, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110,Docket # 212436-WO-SEC-1 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216 or 218, and comprises at least one modification described herein.
[0038] In certain embodiments, the modified FATB polypeptides described herein comprise a signal peptide operably linked to the modified FATB polypeptide. In certain embodiments, the signal peptide is operably linked at the N-terminus of the modified FATB polypeptide. In certain embodiments, the signal peptide is operably linked at the C-terminus of the modified FATB polypeptide. In certain embodiments, the modified FATB polypeptide comprises 2 or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more) signal peptides. The 2 or more signal peptide may be operably linked the N-terminus, the C-terminus, or a combination thereof. In certain embodiments, the signal peptide comprises an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 66, 68, 70, 72, 74, 76, 78, 80 and 82. In certain embodiments, the modified FATB polypeptide comprises a linker sequence between the modified FATB polypeptide and the signal peptide. The sequence and length of the linker is not particularly limited so long as the signal peptide can direct the modified FATB polypeptide to the desired location.
[0039] In certain embodiments, the polynucleotide encoding the modified FATB polypeptide is operably linked to at least one (e.g., at least 1, 2, 3, 4, 5, 6, 7 or more) regulatory element. In certain embodiments, the regulatory element is a promoter. In certain embodiments, the regulatory element is the native promoter. In certain embodiments, the regulatory element is a heterologous regulatory element (e.g., heterologous promoter). In certain embodiments, the heterologous regulatory element is heterologous to the polynucleotide sequence encoding the polypeptide. In certain embodiments, in which the polynucleotide operably linked to the heterologous regulatory element is introduced in a cell the regulatory element is heterologous to the cell.
[0040] As used herein “operably linked” is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a polynucleotide of interest and a regulatory sequence (e.g., a promoter) is a functional link that allows for expression of the polynucleotide of interest. Operably linked elements may be contiguous or non-contiguous.Docket # 212436-WO-SEC-1 When used to refer to the joining of two protein coding regions, operably linked is intended that the coding regions are in the same reading frame.
[0041] As used herein “regulatory element” generally refers to a transcriptional regulatory element involved in regulating the transcription of a nucleic acid molecule such as a gene or a target gene. The regulatory element is a nucleic acid and may include a promoter, an enhancer, an intron, expression modulating elements (EMEs), a 5’-untranslated region (5’-UTR, also known as a leader sequence), or a 3’-UTR, or a combination thereof. A regulatory element may act in "cis" or "trans", and generally it acts in "cis", i.e., it activates expression of genes located on the same nucleic acid molecule, e.g., a chromosome, where the regulatory element is located.
[0042] An “enhancer” element is any nucleic acid molecule that increases transcription of a nucleic acid molecule when functionally linked to a promoter regardless of its relative position. Various enhancers are known in the art including for example, introns with gene expression enhancing properties in plants, the ubiquitin intron (i.e., the maize ubiquitin intron 1 (see, for example, NCBI sequence S94464)), the omega enhancer or the omega prime enhancer (Gallie, et al., (1989) Molecular Biology of RNA ed. Cech (Liss, New York) 237-256 and Gallie, et al., (1987) Gene 60:217-25), the CaMV 35S enhancer (see, e.g., Benfey, et al., (1990) EMBO J. 9:1685-96) and the enhancers of US Patent Number 7,803,992 may also be used. The above list of transcriptional enhancers is not meant to be limiting. Any appropriate transcriptional enhancer can be used in the embodiments described herein.
[0043] A “repressor” (also sometimes called herein silencer) is defined as any nucleic acid molecule which inhibits the transcription when functionally linked to a promoter regardless of relative position. The term "cis-element" generally refers to a transcriptional regulatory element that affects or modulates expression of an operably linked transcribable polynucleotide, where the transcribable polynucleotide is present in the same DNA sequence. A cis-element may function to bind transcription factors, which are trans-acting polypeptides that regulate transcription.
[0044] An “intron” is an intervening sequence in a gene that is transcribed into RNA but is then excised in the process of generating the mature mRNA. The term is also used for the excised RNA sequences. An “exon” is a portion of the sequence of a gene that is transcribed and is found in the mature mRNA derived from the gene but is not necessarily a part of the sequence that encodes the final gene product. The 5' untranslated region (5’UTR) (also known as aDocket # 212436-WO-SEC-1 translational leader sequence or leader RNA) is the region of an mRNA that is directly upstream from the initiation codon. This region is involved in the regulation of translation of a transcript by differing mechanisms in viruses, prokaryotes and eukaryotes. The “3' non-coding sequences” refer to DNA sequences located downstream of a coding sequence and include polyadenylation recognition sequences and other sequences encoding regulatory signals capable of affecting mRNA processing or gene expression. The polyadenylation signal is usually characterized by affecting the addition of polyadenylic acid tracts to the 3' end of the mRNA precursor.
[0045] As used herein “promoter” refers to a region of DNA upstream from the start of transcription and involved in recognition and binding of RNA polymerase and other proteins to initiate transcription. A “plant promoter” is a promoter capable of initiating transcription in plant cells. In certain embodiments, the polynucleotides described herein are operably linked to a promoter that drives expression in a plant cell. Any promoter known in the art can be used in the methods of the present disclosure including, but not limited to, constitutive promoters, pathogen- inducible promoters, wound-inducible promoters, tissue-preferred promoters, and chemical- regulated promoters. The choice of promoter may depend on the desired timing and location of expression in the transformed plant as well as other factors, which are known to those of skill in the art. Such constitutive promoters include, for example, the core promoter of the Rsyn7 promoter and other constitutive promoters disclosed in WO 99 / 43838 and U.S. Patent No. 6,072,050; the core CaMV 35S promoter; rice actin; ubiquitin; pEMU; MAS; ALS; and the like. Other constitutive promoters include, for example, those disclosed in U.S. Patent Nos. 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; and 6,177,611, which are known in the art, and can be contemplated for use in the present disclosure.
[0046] Generally, it can be beneficial to express the gene from an inducible promoter, particularly from a pathogen-inducible promoter. Such promoters include those from pathogenesis-related proteins (PR proteins), which are induced following infection by a pathogen, e.g., PR proteins, SAR proteins, beta-l,3-glucanase, chitinase, etc.
[0047] Of interest are promoters that are expressed locally at or near the site of pathogen infection. Additionally, as pathogens find entry into plants through wounds or insect damage, a wound-inducible promoter can be used in the constructions of the disclosure. Such wound- inducible promoters include potato proteinase inhibitor (pin II) gene, wunl and wun2, winl and win2, systemin, WIP1, MPI gene, and the like.Docket # 212436-WO-SEC-1
[0048] Chemical-regulated promoters can be used to modulate the expression of a gene in a plant through the application of an exogenous chemical regulator. Depending upon the objective, the promoter can be a chemical-inducible promoter, where application of the chemical induces gene expression, or a chemical-repressible promoter, where application of the chemical represses gene expression. Chemical-inducible promoters are known in the art and include, but are not limited to, the maize In2-2 promoter, which is activated by benzenesulfonamide herbicide safeners, the maize GST promoter, which is activated by hydrophobic electrophilic compounds that are used as pre-emergent herbicides, and the tobacco PR-la promoter, which is activated by salicylic acid. Other chemical-regulated promoters of interest include steroid-responsive promoters (e.g., the glucocorticoid-inducible promoter, and tetracycline-inducible and tetracycline-repressible promoters).
[0049] Tissue-preferred promoters can be utilized to target enhanced expression of the target genes or proteins within a particular plant tissue. Such tissue-preferred promoters include, but are not limited to, leaf-preferred promoters, root-preferred promoters, seed-preferred promoters, and stem-preferred promoters. Tissue-preferred promoters include Yamamoto et al. (1997) Plant J. 12(2): 255 -265; Kawamata et al. (1997) Plant Cell Physiol.38(7):792-803; Hansen et al. (1997) Mol. Gen Genet.254(3):337-343; Russell et al. (1997) Transgenic Res.6(2): 157-168; Rinehart et al. (1996) Plant Physiol.112(3): 1331-1341; Van Camp et al. (1996) Plant Physiol. 112(2):525-535; Canevascini et al. (1996) Plant Physiol.112(2):513-524; Yamamoto et al. (1994) Plant Cell Physiol.35(5):773-778; Lam (1994) Results Probl. Cell Differ.20: 181-196; Orozco et al. (1993) Plant Mol Biol.23(6): 1129-1138; Matsuoka et al. (1993) Proc Natl. Acad. Sci. USA 90(20):9586-9590; and Guevara-Garcia et al. (1993) Plant J.4(3):495-505. Such promoters can be modified.
[0050] Leaf-specific promoters are known in the art. See, for example, Yamamoto et al. (1997) Plant J.12(2)255-265; Kwon et al. (1994) Plant Physiol.105:357-67; Yamamoto et al. (1994) Plant Cell Physiol.35(5):773-778; Gotor et al. (1993) Plant J.3:509-18; Orozco et al. (1993) Plant Mol. Biol.23(6): 1129-1138; and Matsuoka et al. (1993) Proc. Natl. Acad. Sci. USA 90(20):9586-9590.
[0051] "Seed-preferred" promoters include both "seed-specific" promoters (those promoters active during seed development such as promoters of seed storage proteins) as well as "seed- germinating" promoters (those promoters active during seed germination). Such seed-preferredDocket # 212436-WO-SEC-1 promoters include, but are not limited to, Ciml (cytokinin-induced message), cZ19Bl (maize 19 kDa zein), milps (myo-inositol-1-phosphate synthase), and celA (cellulose synthase) (see WO 00 / 11177, herein incorporated by reference). Gama-zein is a preferred endosperm-specific promoter. Glob-1 is a preferred embryo-specific promoter. For dicots, seed-specific promoters include, but are not limited to, bean β-phaseolin, napin, β-conglycinin, soybean lectin, cruciferin, and the like. For monocots, seed-specific promoters include, but are not limited to, maize 15 kDa zein, 22 kDa zein, 27 kDa zein, g-zein, waxy, shrunken 1, shrunken 2, globulin 1, etc. See also WO 00 / 12733, where seed-preferred promoters from endl and end2 genes are disclosed; herein incorporated by reference.
[0052] In certain embodiments, the polynucleotides of the present disclosure can involve the use of the intact, native FATB genes, wherein the expression is driven by a cognate 5' upstream promoter sequence(s).
[0053] Also contemplated are synthetic promoters which include a combination of one or more heterologous regulatory elements.
[0054] In certain embodiments, the FATB polynucleotides encoding the modified FATB polypeptides described above are inserted into a recombinant DNA construct. In certain embodiments, the recombinant DNA construct further comprises at least one regulatory element. In certain embodiments, the at least one regulatory element of the recombinant DNA construct comprises a promoter. In certain embodiments, the recombinant DNA construct, described herein is expressed in a plant or seed. In certain embodiment, the plant or seed is a soybean plant or soybean seed.
[0055] As used herein, a “recombinant DNA construct” comprises two or more operably linked DNA segments which are not found operably linked in nature. Non-limiting examples of recombinant DNA constructs include a polynucleotide of interest operably linked to heterologous sequences, also referred to as “regulatory elements,” which aid in the expression, autologous replication, and / or genomic insertion of the sequence of interest. Such regulatory elements include, for example, promoters, termination sequences, enhancers, etc., or any component of an expression cassette; a plasmid, cosmid, virus, autonomously replicating sequence, phage, or linear or circular single-stranded or double-stranded DNA or RNA nucleotide sequence; and / or sequences that encode heterologous polypeptides.Docket # 212436-WO-SEC-1
[0056] The modified FATB polypeptides described herein can be provided for expression in a plant of interest or an organism of interest. The cassette can include 5' and 3' regulatory sequences operably linked to the FATB polynucleotide.
[0057] The promoter of the recombinant DNA constructs provided herein can be any type or class of promoter known in the art, such that any one of a number of promoters can be used to express the various modified FATB sequences disclosed herein, including the native promoter of the polynucleotide sequence of interest. The promoters for use in the recombinant DNA constructs of the invention can be selected based on the desired outcome.
[0058] In certain embodiments, the polynucleotides encoding the modified FATB polypeptides described herein are provided in expression cassettes (e.g., a plasmid, cosmid, virus, autonomously replicating sequence, phage, or linear or circular single-stranded or double- stranded DNA or RNA nucleotide sequence) for expression in a plant of interest or any organism of interest. The cassette can include 5' and 3' regulatory sequences operably linked to a polynucleotide encoding the modified FATB polypeptide. The cassette may additionally contain at least one additional gene to be cotransformed into the plant. Alternatively, the additional gene(s) can be provided on multiple expression cassettes. Such an expression cassette is provided with a plurality of restriction sites and / or recombination sites for insertion of the polynucleotide encoding the modified FATB polypeptide to be under the transcriptional regulation of the regulatory regions. The expression cassette may additionally contain selectable marker genes.
[0059] The expression cassette can include in the 5'-3' direction of transcription, a transcriptional and translational initiation region (e.g., a promoter), a polynucleotide encoding the modified FATB polypeptide, and a transcriptional and translational termination region (e.g., termination region) functional in plants. The regulatory regions (e.g., promoters, transcriptional regulatory regions, and translational termination regions) and / or the polynucleotide encoding the modified FATB polypeptide may be native / analogous to the host cell or to each other. Alternatively, the regulatory regions and / or the polynucleotide encoding the modified FATB polypeptide may be heterologous to the host cell or to each other.
[0060] The termination region may be native with the transcriptional initiation region, with the plant host, or may be derived from another source (i.e., foreign or heterologous) than the promoter, the polynucleotide encoding the modified FATB polypeptide, the plant host, or any combination thereof.Docket # 212436-WO-SEC-1
[0061] The expression cassette may additionally contain a 5' leader sequences. Such leader sequences can act to enhance translation. Translation leaders are known in the art and include viral translational leader sequences.
[0062] Generally, the expression cassette can comprise a selectable marker gene for the selection of transformed cells. Selectable marker genes are utilized for the selection of transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), as well as genes conferring resistance to herbicidal compounds, such as glyphosate, glufosinate ammonium, bromoxynil, imidazolinones, and 2,4-dichlorophenoxyacetate (2,4-D). The above list of selectable marker genes is not meant to be limiting. Any selectable marker gene can be used in the present disclosure.
[0063] In preparing the expression cassette, the various DNA fragments may be manipulated, to provide for the DNA sequences in the proper orientation and, as appropriate, in the proper reading frame. Toward this end, adapters or linkers may be employed to join the DNA fragments or other manipulations may be involved to provide for convenient restriction sites, removal of superfluous DNA, removal of restriction sites, or the like. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transitions and transversions, may be involved.
[0064] Also provided herein are plant seeds and plant cells (e.g., soybean seeds and soybean cells) comprising at least one polynucleotide encoding a modified FATB polypeptide described herein such as, for example, a modified FATB polypeptide comprising an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, aDocket # 212436-WO-SEC-1 leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof.
[0065] Also provided herein are plant seeds and plant cells (e.g., soybean seeds and soybean cells) comprising an introduced genomic modification of a gene encoding an endogenous FATB4 (e.g., FATB4A or FATB4B) polypeptide comprising an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 14 and 16.
[0066] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell (e.g., soybean seed or soybean cell) comprises a medium chain fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content in the plant seed. The type of MCFA (e.g., C14:0, C12:0, C10:0, C8:0, or C6:0) of the plant seed or plant cell is not particularly limited such that the MCFA content of the seed or cell may be any combination of MCFAs.
[0067] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell (e.g., soybean seed or soybean cell) comprises C14:0 fatty acid content in an amountDocket # 212436-WO-SEC-1 of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content in the plant seed or plant cell. In certain embodiments, the plant seed or plant cell comprises C14:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total MCFA content in the plant seed or plant cell.
[0068] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell (e.g., soybean seed or soybean cell) comprises C12:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content in the plant seed or plant cell. In certain embodiments, the plant seed or plant cell comprises C12:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total MCFA content in the plant seed or plant cell.
[0069] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises C10:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content in the plant seed or plant cell. In certain embodiments, the plant seed or plant cell comprises C10:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total MCFA content in the plant seed or plant cell.
[0070] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises C8:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content in the plant seed or plant cell. In certain embodiments, the plant seed or plant cell comprises C8:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total MCFA content in the plant seed or plant cell.Docket # 212436-WO-SEC-1
[0071] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises C6:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content in the plant seed or plant cell. In certain embodiments, the plant seed or plant cell comprises C6:0 fatty acid content in an amount of at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total MCFA content in the plant seed or plant cell.
[0072] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises at least, or at least about, a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and less than an 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 or 25 percentage point increase in the medium chain fatty acid content as compared to a control plant seed or plant cell (e.g., plant seed or plant cell not comprising the genetic modification).
[0073] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises at least, or at least about, a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and less than an 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 or 25 percentage point increase in the C14:0 fatty acid content as compared to a control plant seed or plant cell (e.g., plant seed or plant cell not comprising the genetic modification).
[0074] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises at least, or at least about, a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and less than an 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 or 25 percentage point increase in the C12:0 fatty acid content as compared to a control plant seed or plant cell (e.g., plant seed or plant cell not comprising the genetic modification).
[0075] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises at least, or at least about, a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and less than an 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 or 25 percentage point increase in the C10:0 fatty acid content as compared to a control plant seed or plant cell (e.g., plant seed or plant cell not comprising the genetic modification).
[0076] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises at least, or at least about, a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and less than an 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 or 25 percentage pointDocket # 212436-WO-SEC-1 increase in the C8:0 fatty acid content as compared to a control plant seed or plant cell (e.g., plant seed or plant cell not comprising the genetic modification).
[0077] In certain embodiments of the compositions and methods described herein, the plant seed or plant cell comprises at least, or at least about, a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and less than an 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30 or 25 percentage point increase in the C6:0 fatty acid content as compared to a control plant seed or plant cell (e.g., plant seed or plant cell not comprising the genetic modification).
[0078] As used herein, "percentage point" (pp) difference, change, increase or decrease refers to the arithmetic difference of two percentages, e.g. [transgenic or genetically modified value (%) - control value (%)] = percentage points. For example, a modified seed may contain 20% by weight of a component and the corresponding unmodified control seed may contain 15% by weight of that component. The difference in the component between the control and modified seed would be expressed as 5 percentage points.
[0079] In certain embodiments, an endogenous FATB gene (e.g., FATB1A, FATB1B, FATB2A, FATB2B, FATB3A, FATB3B, FATB4B, or FATB5A) is modified to encode a modified FATB polypeptide described herein.
[0080] In certain embodiments, the genomic modification is introduced into an endogenous or native fatty acyl-ACP thioesterase B 4 (FATB4) gene encoding a FATB4 polypeptide. In certain embodiments, the genomic modification of the FATB4 gene increases the expression and / or activity of the encoded FATB4 polypeptide. In certain embodiments, the genomic modification is introduced in a regulatory domain of the FATB4 gene. In certain embodiments, the genomic modification is introduced in a coding sequence of the FATB4 gene. In certain embodiments, the genomic modification comprises the insertion of a regulatory enhancer or promoter sequence operably linked to a polynucleotide encoding the FATB4 polypeptide, such that, for example, in certain embodiments, the polynucleotide encoding the FATB4 polypeptide is operably linked to a heterologous promoter sequence. In certain embodiments, the genomic modification comprises a substitution in a FATB4 gene regulatory enhancer or promoter sequence, the substitution increasing expression of the encoded FATB4 polypeptide. In certain embodiments, the FATB4 gene comprises a nucleotide sequence that is at least, or at least about, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13, 15, 33 or 47-48. In certain embodiments, the FATB4Docket # 212436-WO-SEC-1 gene encodes a FATB4 polypeptide comprising an amino acid sequence that is at least, or at least about, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 14, 16, or 34.
[0081] As used herein “endogenous gene” refers to a gene that is original to a host plant and can be used synonymously with “host genomic DNA,” “pre-existing DNA,” and the like. Moreover, for the purposes herein, an endogenous gene includes coding DNA and genomic DNA within and surrounding the coding DNA, such as for example, the promoter, intron, and terminator sequences. In certain embodiments, the endogenous FATB gene comprises a nucleotide sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% with any of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 29, 31, and 44-48.
[0082] Methods to modify or alter endogenous genomic DNA are known in the art. For example, a pre-existing or endogenous sequence in a host plant can be modified or altered in a site-specific fashion using one or more site-specific engineering systems. In certain embodiments, one or more modifications is a targeted genetic modification. In certain embodiments, the endogenous gene may be modified by a CRISPR associated (Cas) endonuclease, a Zn-finger nuclease- mediated system, a meganuclease-mediated system, an oligonucleobase-mediated system, or any gene modification system known to one of ordinary skill in the art.
[0083] As used herein, a targeted genetic modification refers to the direct manipulation of an organism’s genes. The targeted modification may be introduced using any technique known in the art, such as, for example, plant breeding, genome editing, or single locus conversion.
[0084] In certain embodiments, the polynucleotides encoding the modified FATB polypeptides are introduced into the plant, plant cell, plant part, ore seed using a nucleic acid construct, or an expression cassette described herein. In certain embodiments, the polynucleotides are stably expressed in the plant, plant cell, plant part, or seed using a stable transformation technique.
[0085] In certain embodiments, the plant cells or plant parts are grown into plants. The method for generating the plants from the plant cells or plant parts is not particularly limited. These plants can then be grown, and either pollinated with the same strain or different strains, and the resulting progeny having constitutive expression of the desired phenotypic characteristic identified. Two or more generations can be grown to ensure that expression of the desired phenotypic characteristic is stably maintained and inherited and then seeds harvested to ensureDocket # 212436-WO-SEC-1 expression of the desired phenotypic characteristic has been achieved. In some aspects of the present disclosure, the transformed seed, genome modified seed, or transgenic seed having a modified FATB, nucleotide construct, or an expression cassette is stably incorporated into their genome.
[0086] As used herein, “polynucleotide” includes reference to a deoxyribopolynucleotide, ribopolynucleotide or analogs thereof that have the essential nature of a natural ribonucleotide in that they hybridize, under stringent hybridization conditions, to substantially the same nucleotide sequence as naturally occurring nucleotides and / or allow translation into the same amino acid(s) as the naturally occurring nucleotide(s). A polynucleotide can be full-length or a subsequence of a structural or regulatory gene. Unless otherwise indicated, the term includes reference to the specified sequence as well as the complementary sequence thereof. Thus, DNAs or RNAs with backbones modified for stability or for other reasons are “polynucleotides” as that term is intended herein. Moreover, DNAs or RNAs comprising unusual bases, such as inosine, or modified bases, such as tritylated bases, to name just two examples, are polynucleotides as the term is used herein. It will be appreciated that a great variety of modifications have been made to DNA and RNA that serve many useful purposes known to those of skill in the art. The term polynucleotide as it is employed herein embraces such chemically, enzymatically or metabolically modified forms of polynucleotides, as well as the chemical forms of DNA and RNA characteristic of viruses and cells, including inter alia, simple and complex cells.
[0087] As used herein “encoding,” “encoded,” or the like, with respect to a specified nucleic acid, means comprising the information for translation into the specified protein. A nucleic acid encoding a protein may comprise non-translated sequences (e.g., introns) within translated regions of the nucleic acid, or may lack such intervening non-translated sequences (e.g., as in cDNA). The information by which a protein is encoded is specified by the use of codons. Typically, the amino acid sequence is encoded by the nucleic acid using the “universal” genetic code. However, variants of the universal code, such as is present in some plant, animal and fungal mitochondria, the bacterium Mycoplasma capricolum (Yamao, et al., (1985) Proc. Natl. Acad. Sci. USA 82:2306-9) or the ciliate Macronucleus, may be used when the nucleic acid is expressed using these organisms.
[0088] The terms “polypeptide,” “peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one orDocket # 212436-WO-SEC-1 more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers.
[0089] As used herein "percent (%) sequence identity" with respect to a reference sequence (subject) is determined as the percentage of amino acid residues or nucleotides in a candidate sequence (query) that are identical with the respective amino acid residues or nucleotides in the reference sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any amino acid conservative substitutions as part of the sequence identity. Alignment for purposes of determining percent sequence identity can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software such as BLAST, BLAST-2. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (e.g., percent identity of query sequence = number of identical positions between query and subject sequences / total number of positions of query sequence ×100).
[0090] Unless otherwise stated, sequence identity / similarity values provided herein refer to the value obtained using the BLAST 2.0 suite of programs using default parameters (Altschul, et al., (1997) Nucleic Acids Res.25:3389-402).
[0091] In certain embodiments the plant cells and plant seeds of the methods and compositions described herein comprise a genomic modification increasing expression or activity of a FATB4 polypeptide and at least one polynucleotide encoding a modified FATB polypeptide described herein such as, for example, a modified FATB polypeptide comprising an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C),Docket # 212436-WO-SEC-1 isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof.
[0092] In certain embodiments the plant cells and plant seeds of the methods and compositions described herein comprise a genomic modification increasing expression or activity of a FATB4 polypeptide and further comprise a genomic modification of a non-FATB4 FATB gene (e.g., FATB1, FATB2), the genomic modification decreasing expression or activity of the encoded non-FATB4 polypeptide. In certain embodiments, the non-FATB4 FATB gene comprises a nucleotide sequence that is at least, or at least about, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 1, 3, 5, 7, 9, 11, 17, 19, 29, 31, or 44-46. In certain embodiments, the non-FATB4 FATB gene encodes a FATB polypeptide comprising an amino acid sequence that is at least, or at least about, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34. In certain embodiments, the genomic modification is introduced into an endogenous or native fatty acyl-ACP thioesterase B 4 (FATB4) gene encoding a FATB4 polypeptide.Docket # 212436-WO-SEC-1
[0093] As used herein, “increase in expression” “increased expression” or the like refers to any detectable gain in expression of a gene and / or the corresponding polypeptide. As used herein, “increase in activity” “increased activity” and the like refers to any detectable gain in activity (e.g., enzymatic activity or altered substrate specificity) of the polypeptide. The method by which the activity of a polypeptide described herein is increased is not particularly limited and can be done using methods known in the art such as increasing expression of the gene encoding the polypeptide (e.g., transgenic expression such as by transformation of the plant by introduction of a heterologous sequence, such as a coding sequence operably linked to one or more heterologous regulatory elements, promoter swap, gene modification) or a targeted modification of the gene encoding the polypeptide to, for example, remove a self-regulatory domain.
[0094] As used herein, “decrease in expression” “decreased expression” or the like refers to any detectable reduction in expression of a gene and / or the corresponding polypeptide. Similarly, “decrease in activity” “decreased activity” or the like refers to any detectable reduction in the activity (e.g., enzymatic activity) of the encoded polypeptide. The method by which the expression or activity of a gene or polypeptide described herein is decreased is not particularly limited and can be done using methods known in the art such at RNAi, gene knockdown, gene knockout, or targeted amino acid modification.
[0095] As used herein a “gene knockout” is used to refer to gene in which there is no detectable expression of the mRNA or protein encoded by the gene, whereas “gene knockdown” is used to refer to a gene in which there is reduced expression of the mRNA or protein encoded by the gene. As used herein, “decreased expression” encompasses both gene knockout and gene knockdown.
[0096] In certain embodiments the plant cells and plant seeds of the methods and compositions described herein further comprise a genomic modification increasing oil content. The modification to enhance oil content is not particularly limited and may be any modification known in the art, or described herein, that can increase total oil content as compared to a control seed (e.g., a seed not comprising the modification).
[0097] In certain embodiments, the modification to enhance oil content is a modification in a gene, including modifications in endogenous coding and regulatory sequences of the gene, encoding at least one of (i) a modification increasing expression and / or activity of a Sugars WillDocket # 212436-WO-SEC-1 Eventually be Exported Transporter (SWT) polypeptide, (ii) a modification increasing expression and / or activity of a sucrose transporter (SUT) polypeptide, (iii) a modification decreasing expression, activity, and / or stability of an endogenous Mother of Flowering Time (MFT) polypeptide, (iv) a modification increasing expression and / or activity of an MFT network gene, (v) a modification increasing expression and / or activity of an ABI3 polypeptide, (vi) a modification increasing expression and / or activity of an ODP1 polypeptide, (vii) a modification introducing a high oil DGAT variant, (viii) a modification decreasing expression, activity, and / or stability of an endogenous raffinose synthase (RS) polypeptide, or any combination thereof.
[0098] In certain embodiments, the at least one additional modification is introduced by genome editing. In certain embodiments of the compositions and methods described herein, the at least one additional modification is introduced by breeding such as, for example, by crossing a first plant comprising a seed expressing a modified FATB polypeptide described herein and having an increase in MCFA content with a second plant comprising an introduced high oil gene, a modification of a FATB4 gene or both. In certain embodiments, the progeny is backcrossed to the first plant to produce seed expressing a modified FATB polypeptide and comprising an increase in MCFA content and a high oil gene, modified FATB4 gene or both.
[0099] In certain embodiments, the soybean plant cell, soybean seed or seed of the soybean plant of the compositions and methods described herein comprises at least or at least about a 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, or 8.0 percentage point increase and less than about 10.0, 9.5, 9.0, 8.5, 8.0, 7.5, 7.0, 6.5, 6.0, 5.5, 5.0 or 4.5 percentage point increase in fatty acid or oil content measured on a dry weight basis relative to a control seed.
[0100] In certain embodiments, the soybean plant cell, soybean seed or seed of the soybean plant of the compositions and methods described herein comprises a total seed weight greater than or within 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, or 20% as compared to a control seed (e.g., an unmodified seed).
[0101] Also provided herein are seeds (e.g., soybean) and cells (e.g., soybean cells) comprising a polynucleotide encoding a FATB4 polypeptide operably linked to a heterologous regulatory element. In certain embodiments, the soybean seed or soybean cell has an increase in the amount of medium chain fatty acids as compared to a control seed or plant. In certain embodiments, theDocket # 212436-WO-SEC-1 heterologous regulatory element comprises a heterologous promoter such as, for example, a promoter described herein. In certain embodiments, the polynucleotide is introduced into the cell in a recombinant construct. Further provided are legume seeds (e.g., soybean or peanut seeds) or legume plants (e.g., soybean o peanut plants) comprising a polynucleotide encoding a FATB4 polypeptide comprising an amino acid sequence that is at least, or at least about, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 14, 16, or 34 or a homolog thereof operably linked to a regulatory element. In certain embodiments, the regulatory element comprises a heterologous promoter such as, for example, a promoter described herein. In certain embodiments, the polynucleotide is introduced into the cell in a recombinant construct.
[0102] Also provided herein are plants (e.g., soybean plants) comprising a plant cell described herein and plants producing the seeds described herein. Further provided are methods of plant breeding comprising crossing a plant described herein with a second plant.
[0103] As used herein, the term “plant” includes plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant calli, plant clumps, and plant cells that are intact in plants or parts of plants such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, kernels, ears, cobs, husks, stalks, roots, root tips, anthers, and the like. Grain is intended to mean the mature seed produced by commercial growers for purposes other than growing or reproducing the species. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the disclosure, provided that these parts comprise the modified FATB gene.
[0104] In certain embodiments, the plants of the compositions and methods described herein are elite lines. As used herein, “elite line” refers to any line that has resulted from breeding and selection for superior agronomic performance that allows a producer to harvest a product of commercial significance. Numerous elite lines are available and known to those of skill in the art of plant breeding.
[0105] In certain embodiments, the plant of the compositions and methods described herein has a yield that is greater than or within 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, or 5%, as compared to the corresponding control plant (e.g., an unmodified plant or a plant comprising one of the two or more modifications). As used herein, “yield” refers to the amount of agricultural production harvested per unit of land and may include reference to bushels per acre or kilograms per hectare of a crop at harvest, as adjusted for grain moisture. Grain moisture is measured in theDocket # 212436-WO-SEC-1 grain at harvest. The adjusted test weight of grain is determined to be the weight in pounds per bushel or kilogram, adjusted for grain moisture level at harvest.
[0106] Any of the modified plants described herein may be grown in a field in a double cropping system with other plants of the same or different species, such that, for example, two or more crops are harvested per 12-month period.
[0107] Also provided herein are methods of producing a soybean plant producing seeds having increased medium chain fatty acid content comprising introducing into a regenerable plant cell a polynucleotide encoding any of the modified FATB polypeptides described herein; and generating the plant, wherein the plant comprises the polynucleotide encoding the modified FATB polypeptide and produces a seed having an increased amount of MCFAs as compared to seed of a plant not comprising the modified FATB polypeptide. In certain embodiments, the method further comprises introducing at least one additional modification increasing expression or activity of a FATB4 polypeptide or a modification associated with increased seed oil content, or any combination thereof, such as for example, the modifications described herein. The method for introducing the at least one additional modification may be any method known in the art or described herein. In certain embodiments, the at least one additional modification is introduced by genome editing or transformation. In certain embodiments, the at least one additional modification is introduced by crossing the plant produced by the method with a second plant comprising the at least one additional modification, harvesting the seed produced thereby, and generating a progeny plant, the progeny plant comprising the modified FATB polypeptide and the at least one additional modification.
[0108] Also provided herein are methods of producing a plant (e.g., soybean) producing seeds having increased medium chain fatty acid content comprising introducing into a regenerable soybean plant cell a genomic modification of a FATB gene, the genomic modification increasing the expression or activity of a FATB polypeptide and generating the plant, wherein the plant comprises the genomic modification and produces a seed having increased medium chain fatty acid content as compared to a seed of a control plant not comprising the genomic modification. The modification increasing the expression or activity of a FATB polypeptide may be any modification described herein. In certain embodiments, the genomic modification increases the expression or activity of a FATB4 (e.g., FATB4A or FATB4B) polypeptide.Docket # 212436-WO-SEC-1
[0109] Various methods can be used to introduce the polynucleotide sequences into a plant, plant part, plant cell, seed, and / or grain. "Introducing" is intended to mean presenting to the plant, plant cell, seed, and / or grain the inventive polynucleotide or resulting polypeptide in such a manner that the sequence gains access to the interior of a cell of the plant. The methods of the disclosure do not depend on a particular method for introducing a sequence into a plant, plant cell, seed, and / or grain, only that the polynucleotide or polypeptide gains access to the interior of at least one cell of the plant.
[0110] In certain embodiments, the method for introducing the modified FATB polypeptide and / or the modification associated with increased FATB (e.g., FATB4) expression or activity, and / or increased oil, or any other polypeptide disclosed herein, comprises transforming a regenerable plant cell with a nucleic acid construct or expression cassette comprising a polynucleotide described herein. The transformation technique of the methods is not particularly limited and includes both stable transformation methods and transient transformation methods.
[0111] "Stable transformation" is intended to mean that the polynucleotide introduced into a plant integrates into the genome of the plant of interest and is capable of being inherited by the progeny thereof. "Transient transformation" is intended to mean that a polynucleotide is introduced into the plant of interest and does not integrate into the genome of the plant or organism, or a polypeptide is introduced into a plant or organism.
[0112] Transformation protocols as well as protocols for introducing polypeptides or polynucleotide sequences into plants may vary depending on the type of plant or plant cell, i.e., monocot or dicot, targeted for transformation. Suitable methods of introducing polypeptides and polynucleotides into plant cells include microinjection (Crossway et al. (1986) Biotechniques 4:320-334), electroporation (Riggs et al. (1986) Proc. Natl. Acad. Sci. USA 83:5602-5606), Agrobacterium-mediated transformation (U.S. Patent No.5,563,055 and U.S. Patent No. 5,981,840), Ochrobacterium-mediated transformation (U.S. Patent Application Publication 2018 / 0216123 and WO20 / 092494) direct gene transfer (Paszkowski et al. (1984) EMBO J. 3:2717-2722), and ballistic particle acceleration (see, for example, U.S. Patent Nos.4,945,050; U.S. Patent No.5,879,918; U.S. Patent No.5,886,244; and, 5,932,782; Tomes et al. (1995) in Plant Cell, Tissue, and Organ Culture: Fundamental Methods, ed. Gamborg and Phillips (Springer-Verlag, Berlin); McCabe et al. (1988) Biotechnology 6:923-926); and Lec1 transformation (WO 00 / 28058). D'Halluin et al. (1992) Plant Cell 4:1495-1505Docket # 212436-WO-SEC-1 (electroporation); Li et al. (1993) Plant Cell Reports 12:250-255 and Christou and Ford (1995) Annals of Botany 75:407-413 (rice); Osjoda et al. (1996) Nature Biotechnology 14:745-750 (maize via Agrobacterium tumefaciens); all of which are herein incorporated by reference.
[0113] Methods are known in the art for the targeted insertion of a polynucleotide at a specific location in the plant genome and can be used to introduce any of the coding or regulatory sequences disclosed herein. In one embodiment, the insertion of the polynucleotide at a desired genomic location is achieved using a site-specific recombination system. See, for example, WO99 / 25821, WO99 / 25854, WO99 / 25840, WO99 / 25855, and WO99 / 25853, all of which are herein incorporated by reference. Briefly, the polynucleotide disclosed herein can be contained in a transfer cassette flanked by two non-recombinogenic recombination sites. The transfer cassette is introduced into a plant having stably incorporated into its genome a target site which is flanked by two non-recombinogenic recombination sites that correspond to the sites of the transfer cassette. An appropriate recombinase is provided, and the transfer cassette is integrated at the target site. The polynucleotide of interest is thereby integrated at a specific chromosomal position in the plant genome.
[0114] One of skill will recognize that after the expression cassette containing the inventive polynucleotide is stably incorporated in transgenic plants and confirmed to be operable, it can be introduced into other plants by sexual crossing. Any of a number of standard breeding techniques can be used, depending upon the species to be crossed.
[0115] Parts obtained from the regenerated plants described herein, such as flowers, seeds, leaves, branches, fruit, and the like are included, provided that these parts comprise cells comprising the inventive polynucleotide. Progeny and variants, and mutants of the regenerated plants are also included, provided that these parts comprise the introduced nucleic acid sequences.
[0116] In one embodiment, a homozygous transgenic plant can be obtained by sexually mating (selfing) a heterozygous transgenic plant that contains a single added heterologous nucleic acid, germinating some of the seed produced and analyzing the resulting plants produced. Back- crossing to a parental plant and out-crossing with a non-transgenic plant are also contemplated.
[0117] In certain embodiments, the method for introducing the modified FATB polypeptide and / or the modification associated with increased FATB (e.g., FATB4) expression or activity, and / or increased oil, into the regenerable plant cell comprises using genome editing technologies.Docket # 212436-WO-SEC-1 In certain embodiments, the method comprises editing the endogenous gene or a previously introduced gene.
[0118] In certain embodiments, an endogenous FATB sequence or other sequence disclosed herein is modified using a genome editing technology to encode the modified FATB protein or modified version of the corresponding polypeptide disclosed herein. In certain embodiments, two or more endogenous FATB sequences are modified using genome editing technology to encode a modified FATB protein. In certain embodiments, the two or more endogenous FATB sequences are modified using genome editing technology to encode a modified FATB polypeptide comprising the same amino acid sequences. In certain embodiments, the two or more endogenous FATB sequences are modified using genome editing technology to encode a modified FATB polypeptide comprising the different amino acid sequences. In embodiments, when three or more endogenous FATB sequences are modified using genome editing technology, the modified FATB polypeptides may comprise the same sequence, different sequences, or a combination thereof.
[0119] The genome editing technology for use in the methods and compositions described herein is not particularly limited and may be any genome editing technique that allows for the modification or targeted introduction of the desired polynucleotide.
[0120] In certain embodiments the genome editing technique uses an enzyme selected from the group consisting of a polynucleotide-guided endonuclease, CRISPR-Cas endonucleases, base editing deaminases, zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), or an engineered site-specific meganuclease.
[0121] In certain embodiments, the genome modification may be facilitated through the induction of a double-stranded break (DSB) or single-strand break, in a defined position in the genome near the desired alteration. DSBs can be induced using any DSB-inducing agent available, including, but not limited to, TALENs, meganucleases, zinc finger nucleases, Cas9- gRNA systems (based on bacterial CRISPR-Cas systems), guided cpf1 endonuclease systems, and the like. In some embodiments, the introduction of a DSB can be combined with the introduction of a polynucleotide modification template.
[0122] In certain embodiments, the method comprises: (a) providing a guide RNA, at least one polynucleotide modification template, and at least one Cas endonuclease to the regenerable plant cell, wherein the at least one Cas endonuclease introduces a double stranded break at anDocket # 212436-WO-SEC-1 endogenous gene to be modified (e.g., FATB1A, FATB1B, FATB2A, FATB2B, FATB3A, FATB3B, FATB4A, FATB4B, FATB5A or FATB5B) in the plant cell, and wherein the polynucleotide modification template generates a modified gene that encodes any of the modifications described herein; (b) obtaining a plant from the plant cell; and (c) generating a progeny plant.
[0123] Double-strand breaks induced by double-strand-break-inducing agents, such as endonucleases that cleave the phosphodiester bond within a polynucleotide chain, can result in the induction of DNA repair mechanisms, including the non-homologous end-joining pathway, and homologous recombination. Endonucleases include a range of different enzymes, including restriction endonucleases (see e.g. Roberts et al., (2003) Nucleic Acids Res 1:418-20), Roberts et al., (2003) Nucleic Acids Res 31:1805-12, and Belfort et al., (2002) in Mobile DNA II, pp.761- 783, Eds. Craigie et al., (ASM Press, Washington, DC)), meganucleases (see e.g., WO 2009 / 114321; Gao et al. (2010) Plant Journal 1:176-187), TAL effector nucleases or TALENs (see e.g., US20110145940, Christian, M., T. Cermak, et al.2010. Targeting DNA double-strand breaks with TAL effector nucleases. Genetics 186(2): 757-61 and Boch et al., (2009), Science 326(5959): 1509-12), zinc finger nucleases (see e.g. Kim, Y. G., J. Cha, et al. (1996). "Hybrid restriction enzymes: zinc finger fusions to FokI cleavage”), and CRISPR-Cas endonucleases (see e.g. WO2007 / 025097 application published March 1, 2007).
[0124] Once a double-strand break is induced in the genome, cellular DNA repair mechanisms are activated to repair the break. There are two DNA repair pathways. One is termed nonhomologous end-joining (NHEJ) pathway (Bleuyard et al., (2006) DNA Repair 5:1-12) and the other is homology-directed repair (HDR). The structural integrity of chromosomes is typically preserved by NHEJ, but deletions, insertions, or other rearrangements (such as chromosomal translocations) are possible (Siebert and Puchta, 2002, Plant Cell 14:1121-31; Pacher et al., 2007, Genetics 175:21-9. The HDR pathway is another cellular mechanism to repair double-stranded DNA breaks and includes homologous recombination (HR) and single- strand annealing (SSA) (Lieber.2010 Annu. Rev. Biochem.79:181-211).
[0125] In addition to the double-strand break inducing agents, site-specific base conversions can also be achieved to engineer one or more nucleotide changes to create one or more modifications described herein into the genome. These include for example, a site-specific base edit mediated by an C•G to T•A or an A•T to G•C base editing deaminase enzymes (Gaudelli et al.,Docket # 212436-WO-SEC-1 Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al. “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems.” Science 353 (6305) (2016); Komor et al. “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage.” Nature 533 (7603) (2016):420-4.
[0126] In the methods described herein, the endogenous gene may be modified by a CRISPR associated (Cas) endonuclease, a Zn-finger nuclease-mediated system, a meganuclease-mediated system, an oligonucleobase-mediated system, or any gene modification system known to one of ordinary skill in the art.
[0127] In certain embodiments the endogenous gene is modified by a CRISPR associated (Cas) endonuclease.
[0128] Class I Cas endonucleases comprise multisubunit effector complexes (Types I, III, and IV), while Class 2 systems comprise single protein effectors (Types II, V, and VI) (Makarova et al.2015, Nature Reviews Microbiology Vol.13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13; Haft et al., 2005, Computational Biology, PLoS Comput Biol 1(6): e60; and Koonin et al.2017, Curr Opinion Microbiology 37:67-78). In Class 2 Type II systems, the Cas endonuclease acts in complex with a guide polynucleotide.
[0129] Accordingly, in certain embodiments of the methods described herein the Cas endonuclease forms a complex with a guide polynucleotide (e.g., guide polynucleotide / Cas endonuclease complex).
[0130] As used herein, the term “guide polynucleotide”, relates to a polynucleotide sequence that can form a complex with a Cas endonuclease, including the Cas endonucleases described herein, and enables the Cas endonuclease to recognize, optionally bind to, and optionally cleave a DNA target site. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (a RNA-DNA combination sequence). The guide polynucleotide may further comprise a chemically modified base, such as, but not limited, to Locked Nucleic Acid (LNA), 5-methyl dC, 2,6-Diaminopurine, 2’-Fluoro A, 2’-Fluoro U, 2'-O-Methyl RNA, Phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 (hexaethylene glycol chain) molecule, or 5’ to 3’ covalent linkage resulting in circularization.Docket # 212436-WO-SEC-1
[0131] In certain embodiments, the Cas endonuclease forms a complex with a guide polynucleotide (e.g., gRNA) that directs the Cas endonuclease to cleave the DNA target to enable target recognition, binding, and cleavage by the Cas endonuclease. The guide polynucleotide (e.g., gRNA) may comprise a Cas endonuclease recognition (CER) domain that interacts with the Cas endonuclease, and a Variable Targeting (VT) domain that hybridizes to a nucleotide sequence in a target DNA. In certain embodiments, the guide polynucleotide (e.g., gRNA) comprises a CRISPR nucleotide (crNucleotide; e.g., crRNA) and a trans-activating CRISPR nucleotide (tracrNucleotide; e.g., tracrRNA) to guide the Cas endonuclease to its DNA target. The guide polynucleotide (e.g., gRNA) comprises a spacer region complementary to one strand of the double strand DNA target and a region that base pairs with the tracrNucleotide (e.g., tracrRNA), forming a nucleotide duplex (e.g. RNA duplex).
[0132] In certain embodiments, the gRNA is a “single guide RNA” (sgRNA) that comprises a synthetic fusion of crRNA and tracrRNA. In many systems, the Cas endonuclease-guide polynucleotide complex recognizes a short nucleotide sequence adjacent to the target sequence (protospacer), called a “protospacer adjacent motif” (PAM).
[0133] The terms “single guide RNA" and “sgRNA” are used interchangeably herein and relate to a synthetic fusion of two RNA molecules, a crRNA (CRISPR RNA) comprising a variable targeting domain (linked to a tracr mate sequence that hybridizes to a tracrRNA), fused to a tracrRNA (trans-activating CRISPR RNA). The single guide RNA can comprise a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment of the type II CRISPR / Cas system that can form a complex with a type II Cas endonuclease, wherein said guide RNA / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, optionally bind to, and optionally nick or cleave (introduce a single or double-strand break) the DNA target site.
[0134] The nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide can comprise a RNA sequence, a DNA sequence, or a RNA-DNA combination sequence. In one embodiment, the nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide can be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89,Docket # 212436-WO-SEC-1 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleotides in length. In one embodiment, the nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide can comprise a tetraloop sequence, such as, but not limiting to a GAAA tetraloop sequence.
[0135] The term “variable targeting domain” or “VT domain” is used interchangeably herein and includes a nucleotide sequence that can hybridize (is complementary) to one strand (nucleotide sequence) of a double strand DNA target site. The percent complementation between the first nucleotide sequence domain (VT domain) and the target sequence can be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. The variable targeting domain can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length. In some embodiments, the variable targeting domain comprises a contiguous stretch of 12 to 30 nucleotides. The variable targeting domain can be composed of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence, or any combination thereof.
[0136] The term “Cas endonuclease recognition domain” or “CER domain” (of a guide polynucleotide) is used interchangeably herein and includes a nucleotide sequence that interacts with a Cas endonuclease polypeptide. A CER domain comprises a (trans-acting) tracrNucleotide mate sequence followed by a tracrNucleotide sequence. The CER domain can be composed of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence (see for example US20150059010A1, published 26 February 2015), or any combination thereof.
[0137] A “protospacer adjacent motif” (PAM) as used herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide / Cas endonuclease system described herein. In certain embodiments, the Cas endonuclease may not successfully recognize a target DNA sequence if the target DNA sequence is not adjacent to, or near, a PAM sequence. In certain embodiments, the PAM precedes the target sequence (e.g., Cas12a). In certain embodiments, the PAM follows the target sequence (e.g., S. pyogenes Cas9). The sequence and length of a PAM herein can differ depending on the Cas protein or Cas protein complex used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.Docket # 212436-WO-SEC-1
[0138] As used herein, the terms “guide polynucleotide / Cas endonuclease complex”, “guide polynucleotide / Cas endonuclease system”, “ guide polynucleotide / Cas complex”, “guide polynucleotide / Cas system” and “guided Cas system” “polynucleotide-guided endonuclease”, and “PGEN” are used interchangeably herein and refer to at least one guide polynucleotide and at least one Cas endonuclease, that are capable of forming a complex, wherein said guide polynucleotide / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) the DNA target site. A guide polynucleotide / Cas endonuclease complex herein can comprise Cas protein(s) and suitable polynucleotide component(s) of any of the known CRISPR systems (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al.2015, Nature Reviews Microbiology Vol.13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13). In certain embodiments, the guide polynucleotide / Cas endonuclease complex is provided as a ribonucleoprotein (RNP), wherein the Cas endonuclease component is provided as a protein and the guide polynucleotide component is provided as a ribonucleotide.
[0139] Examples of Cas endonucleases for use in the methods described herein include, but are not limited to, Cas9 and Cpf1. Cas9 (formerly referred to as Cas5, Csn1, or Csx12) is a Class 2 Type II Cas endonuclease (Makarova et al.2015, Nature Reviews Microbiology Vol.13:1-15). A Cas9-gRNA complex recognizes a 3’ PAM sequence (NGG for the S. pyogenes Cas9) at the target site, permitting the spacer of the guide RNA to invade the double-stranded DNA target, and, if sufficient homology between the spacer and protospacer exists, generate a double-strand break cleavage. Cas9 endonucleases comprise RuvC and HNH domains that together produce double strand breaks, and separately can produce single strand breaks. For the S. pyogenes Cas9 endonuclease, the double-strand break leaves a blunt end. Cpf1 is a Clas 2 Type V Cas endonuclease, and comprises nuclease RuvC domain but lacks an HNH domain (Yamane et al., 2016, Cell 165:949-962). Cpf1 endonucleases create “sticky” overhang ends.
[0140] Some uses for Cas9-gRNA systems at a genomic target site include, but are not limited to, insertions, deletions, substitutions, or modifications of one or more nucleotides at the target site; modifying or replacing nucleotide sequences of interest (such as a regulatory elements); insertion of polynucleotides of interest; gene knock-out; gene-knock in; modification of splicing sites and / or introducing alternate splicing sites; modifications of nucleotide sequences encodingDocket # 212436-WO-SEC-1 a protein of interest; amino acid and / or protein fusions; and gene silencing by expressing an inverted repeat into a gene of interest.
[0141] The terms “target site”, “target sequence”, “target site sequence, ”target DNA”, “target locus”, “genomic target site”, “genomic target sequence”, “genomic target locus” and “protospacer”, are used interchangeably herein and refer to a polynucleotide sequence such as, but not limited to, a nucleotide sequence on a chromosome, episome, a locus, or any other DNA molecule in the genome (including chromosomal, chloroplastic, mitochondrial DNA, plasmid DNA) of a cell, at which a guide polynucleotide / Cas endonuclease complex can recognize, bind to, and optionally nick or cleave . The target site can be an endogenous site in the genome of a cell, or alternatively, the target site can be heterologous to the cell and thereby not be naturally occurring in the genome of the cell, or the target site can be found in a heterologous genomic location compared to where it occurs in nature. As used herein, terms “endogenous target sequence” and “native target sequence” are used interchangeable herein to refer to a target sequence that is endogenous or native to the genome of a cell and is at the endogenous or native position of that target sequence in the genome of the cell. An “artificial target site” or “artificial target sequence” are used interchangeably herein and refer to a target sequence that has been introduced into the genome of a cell. Such an artificial target sequence can be identical in sequence to an endogenous or native target sequence in the genome of a cell but be located in a different position (i.e., a non-endogenous or non-native position) in the genome of a cell. An “altered target site”, “altered target sequence”, “modified target site”, “modified target sequence” are used interchangeably herein and refer to a target sequence as disclosed herein that comprises at least one alteration when compared to non-altered target sequence. Such “alterations” include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) – (iii).
[0142] A “polynucleotide modification template” is also provided that comprises at least one nucleotide modification when compared to the nucleotide sequence to be edited. A nucleotide modification can be at least one nucleotide substitution, addition, deletion, or chemical alteration. Optionally, the polynucleotide modification template can further comprise homologous nucleotide sequences flanking the at least one nucleotide modification, wherein the flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.Docket # 212436-WO-SEC-1
[0143] In certain embodiments of the methods disclosed herein, a polynucleotide of interest is inserted at a target site and provided as part of a “donor DNA” molecule. As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into the target site of a Cas endonuclease. The donor DNA construct further comprises a first and a second region of homology that flank the polynucleotide of interest. The first and second regions of homology of the donor DNA share homology to a first and a second genomic region, respectively, present in or flanking the target site of the cell or organism genome. The donor DNA can be tethered to the guide polynucleotide. Tethered donor DNAs can allow for co- localizing target and donor DNA, useful in genome editing, gene insertion, and targeted genome regulation, and can also be useful in targeting post-mitotic cells where function of endogenous HR machinery is expected to be highly diminished (Mali et al., 2013, Nature Methods Vol.10: 957-963). The amount of homology or sequence identity shared by a target and a donor polynucleotide can vary and includes total lengths and / or regions.
[0144] The process for editing a genomic sequence at a Cas9-gRNA double-strand-break site with a modification template generally comprises: providing a host cell with a Cas9-gRNA complex that recognizes a target sequence in the genome of the host cell and is able to induce a double-strand-break in the genomic sequence, and at least one polynucleotide modification template comprising at least one nucleotide alteration when compared to the nucleotide sequence to be edited. The polynucleotide modification template can further comprise nucleotide sequences flanking the at least one nucleotide alteration, in which the flanking sequences are substantially homologous to the chromosomal region flanking the double-strand break. Genome editing using double-strand-break-inducing agents, such as Cas9-gRNA complexes, has been described, for example in US20150082478, WO2015026886, WO2016007347, and WO2016025131.
[0145] To facilitate optimal expression and nuclear localization for eukaryotic cells, the gene comprising the Cas endonuclease may be optimized as described in WO2016186953 published 24 November 2016, and then delivered into cells as DNA expression cassettes by methods known in the art. In certain embodiments, the Cas endonuclease is provided as a polypeptide. In certain embodiments, the Cas endonuclease is provided as a polynucleotide encoding a polypeptide. In certain embodiments, the guide RNA is provided as a DNA molecule encoding one or more RNA molecules. In certain embodiments, the guide RNA is provided as RNA orDocket # 212436-WO-SEC-1 chemically modified RNA. In certain embodiments, the Cas endonuclease protein and guide RNA are provided as a ribonucleoprotein complex (RNP).
[0146] In certain embodiments of the inventive methods described herein the endogenous gene is modified by a zinc-finger-mediated genome editing process. The zinc-finger-mediated genome editing process for editing a chromosomal sequence includes for example: (a) introducing into a cell at least one nucleic acid encoding a zinc finger nuclease that recognizes a target sequence in the chromosomal sequence and is able to cleave a site in the chromosomal sequence, and, optionally, (i) at least one donor polynucleotide that includes a sequence for integration flanked by an upstream sequence and a downstream sequence that exhibit substantial sequence identity with either side of the cleavage site, or (ii) at least one exchange polynucleotide comprising a sequence that is substantially identical to a portion of the chromosomal sequence at the cleavage site and which further comprises at least one nucleotide change; and (b) culturing the cell to allow expression of the zinc finger nuclease such that the zinc finger nuclease introduces a double-stranded break into the chromosomal sequence, and wherein the double-stranded break is repaired by (i) a non-homologous end-joining repair process such that an inactivating mutation is introduced into the chromosomal sequence, or (ii) a homology-directed repair process such that the sequence in the donor polynucleotide is integrated into the chromosomal sequence or the sequence in the exchange polynucleotide is exchanged with the portion of the chromosomal sequence.
[0147] A zinc finger nuclease includes a DNA binding domain (i.e., zinc finger) and a cleavage domain (i.e., nuclease). The nucleic acid encoding a zinc finger nuclease may include DNA or RNA. Zinc finger binding domains may be engineered to recognize and bind to any nucleic acid sequence of choice. See, for example, Beerli et al. (2002) Nat. Biotechnol.20:135-141; Pabo et al. (2001) Ann. Rev. Biochem.70:313-340; Choo et al. (2000) Curr. Opin. Struct. Biol.10:411- 416; and Doyon et al. (2008) Nat. Biotechnol.26:702-708; Santiago et al. (2008) Proc. Natl. Acad. Sci. USA 105:5809-5814; Urnov, et al., (2010) Nat Rev Genet.11(9):636-46; and Shukla, et al., (2009) Nature 459 (7245):437-41. An engineered zinc finger binding domain may have a novel binding specificity compared to a naturally occurring zinc finger protein. As an example, the algorithm of described in U.S. Pat. No.6,453,242 may be used to design a zinc finger binding domain to target a preselected sequence. Nondegenerate recognition code tables may also be used to design a zinc finger binding domain to target a specific sequence (Sera et al.Docket # 212436-WO-SEC-1 (2002) Biochemistry 41:7074-7081). Tools for identifying potential target sites in DNA sequences and designing zinc finger binding domains may be used (Mandell et al. (2006) Nuc. Acid Res.34:W516-W523; Sander et al. (2007) Nuc. Acid Res.35:W599-W605).
[0148] An exemplary zinc finger DNA binding domain recognizes and binds a sequence having at least about 80% sequence identity with the desired target sequence. In other embodiments, the sequence identity may be about 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0149] A zinc finger nuclease also includes a cleavage domain. The cleavage domain portion of the zinc finger nucleases may be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which a cleavage domain may be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, 2010-2011 Catalog, New England Biolabs, Beverly, Mass.; and Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Additional enzymes that cleave DNA are known (e.g., S1 Nuclease; mung bean nuclease; pancreatic DNase I; micrococcal nuclease; yeast HO endonuclease). One or more of these enzymes (or functional fragments thereof) may be used as a source of cleavage domains.
[0150] In certain embodiments of the methods described herein the endogenous gene is modified by using “custom" meganucleases produced to modify plant genomes (see e.g., WO 2009 / 114321; Gao et al. (2010) Plant Journal 1:176-187). The term "meganuclease" generally refers to a naturally occurring homing endonuclease that binds double-stranded DNA at a recognition sequence that is greater than 12 base pairs and encompasses the corresponding intron insertion site. Naturally occurring meganucleases can be monomeric (e.g., I-SceI) or dimeric (e.g., I-CreI). The term meganuclease, as used herein, can be used to refer to monomeric meganucleases, dimeric meganucleases, or to the monomers which associate to form a dimeric meganuclease.
[0151] Naturally occurring meganucleases, for example, from the LAGLIDADG family, have been used to effectively promote site-specific genome modification in plants, yeast, Drosophila, mammalian cells and mice. Engineered meganucleases such as, for example, LIG-34 meganucleases, which recognize and cut a 22 basepair DNA sequence found in the genome of Zea mays (maize) are known (see e.g., US 20110113509).
[0152] In certain embodiments of the methods described herein the endogenous gene is modified by using TAL endonucleases (TALEN). TAL (transcription activator-like) effectors from plantDocket # 212436-WO-SEC-1 pathogenic Xanthomonas are important virulence factors that act as transcriptional activators in the plant cell nucleus, where they directly bind to DNA via a central domain of tandem repeats. A transcription activator-like (TAL) effector-DNA modifying enzymes (TALE or TALEN) are also used to engineer genetic changes. See e.g., US20110145940, Boch et al., (2009), Science 326(5959): 1509-12. Fusions of TAL effectors to the FokI nuclease provide TALENs that bind and cleave DNA at specific locations. Target specificity is determined by developing customized amino acid repeats in the TAL effectors.
[0153] In certain embodiments of the methods described herein the endogenous gene is modified by using base editing, such as an oligonucleobase-mediated system. In addition to the double- strand break inducing agents, site-specific base conversions can also be achieved to engineer one or more nucleotide changes to create one or more EMEs described herein into the genome. These include for example, a site-specific base edit mediated by a C•G to T•A or an A•T to G•C base editing deaminase enzymes (Gaudelli et al., Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al. “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems.” Science 353 (6305) (2016); Komor et al. “Programmable editing of a target base in genomic DNA without double- stranded DNA cleavage.” Nature 533 (7603) (2016):420-4. Catalytically dead dCas9 fused to a cytidine deaminase or an adenine deaminase protein becomes a specific base editor that can alter DNA bases without inducing a DNA break. Base editors convert C->T (or G->A on the opposite strand) or an adenine base editor that would convert adenine to inosine, resulting in an A->G change within an editing window specified by the gRNA.
[0154] Further provided are methods of plant breeding comprising crossing any of the plants (e.g., soybean plant) described herein with a second plant to produce a progeny seed comprising a polynucleotide encoding a modified FATB polypeptide described herein and optionally further comprising at least one modification associated with increased FATB4 expression or activity and / or increased oil, described herein. In certain embodiments, a plant is produced from the progeny seed.
[0155] The seeds of the compositions and methods described herein comprising the modification of a FATB gene and an increase in MCFA content can be processed to produce oil. Methods of processing the seeds to produce oil are provided which include one or more steps of dehulling the seeds, crushing the seeds, heating the seeds, such as with steam, extracting the oil, roasting,Docket # 212436-WO-SEC-1 and extrusion. Processing and oil extraction can be done using solvents or mechanical extraction. In certain embodiments, the MCFA content of the oil produced from the seeds of the compositions and methods described herein comprises at least, or at least about 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, or 60% and less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the total fatty acid content of the oil.
[0156] Products formed following processing include, without limitation, soybean oil, aviation fuel, or biofuel. Crude or partially degummed oil can be further processed by one or more of degumming, alkali treatment, silica absorption, vacuum bleaching, hydrogenation, interesterification, filtration, deodorization, physical refining, refractionation, and optional blending to produce refined bleached deodorized (RBD) oil.
[0157] The oil can be used in animal feed and in food products for human consumption. Provided are food products and animal feed comprising oils which contain or are derived from the modified polynucleotides and modified polypeptides. The food products and animal feed may comprise nucleotides comprising one or more of the modified alleles disclosed herein and the modified polynucleotides, polypeptides and plant cell disclosed herein.
[0158] The following are examples of specific embodiments of some aspects of the invention. The examples are offered for illustrative purposes only and are not intended to limit the scope of the invention in any way. EXAMPLE 1
[0159] This example demonstrates the in vivo activity of thioesterases in E. coli.
[0160] Soybean contains 10 genes encoding fatty acyl-ACP thioesterases B (FATB). Five of the soybean thioesterase genes, FATB1A, FATB1B, FATB2A, FATB2B, and FATB5B, are highly expressed in seeds, whereas five of the soybean thioesterase genes FATB3A, FATB3B, FATB4A, FATB4B, and FATB5A have low seed expression (Table 2). Table 2: Expression Profile of Soybean FATB Genes V1R2 R5 Pod R5 Whole R6 Whole V1 V5 MatureDocket # 212436-WO-SEC-1 FATB1B 958 804.5 1010.3 560 262.25 687.3 497.25 SEQ ID NO: 3 ur
[0161] To determine the in vivo fatty acid activity of these thioesterases (TE), mature protein sequences were predicted and the corresponding coding DNA sequences were cloned into pET28A vector with an N-terminal 6xHIS, followed by transformation into E. coli strain K27 containing a mutant fadD gene resulting in impaired beta oxidation and accumulation of free fatty acids in the growth medium. As a positive control, FATB genes from Palm oil (EgFatB1) and California bay laurel (UcFatB1) were cloned and expressed as above. E. coli expressing empty pET28a vector was used as a negative control.
[0162] Three independent colonies were cultured at 37°C for 18 hours in 2 ml of LB medium supplemented with 50 µg / ml Kanamycin. These cultures were diluted in M9 minimal medium supplemented with 1% glucose, 0.01% Biotin, 0.01% vitamin B, and 50 µg / ml Kanamycin. When the culture reached OD600 of 0.5~0.6, FATB expression was induced by adding 0.1mM isopropyl-D-thiogalactopyranoside (IPTG) and incubated at 30°C for 24 hours. Cultures were centrifuged and the free fatty acids in the supernatant were extracted and analyzed by Gas Chromatography (GC) following methods described in Voelker et al.1994 with minor modifications. The quantity of individual fatty acids was calculated relative to the internalDocket # 212436-WO-SEC-1 standard and molar concentration was determined by multiplying with the molecular weight correction factor (MWF) of each fatty acid.
[0163] As shown in Table 3, E. coli cells expressing the empty vector (EV) accumulated mostly LCFA like C16:0 and C16:1 in the medium, accounting for more than 75%, and negligible amounts of short and MCFA acids (~12%). A similar trend was observed in cells expressing FATB1A (SEQ ID NO: 30) and FATB2A (SEQ ID NO: 32), albeit with a slightly higher total LCFA content (more than 80%) as compared to EV. Cells expressing FATB4A (SEQ ID NO: 34) accumulated more than 60% MCFA, most of which was C8:0 (~38%) and C12:0 (~28%). This fatty acid profile was similar but not identical to that of E.coli cells expressing EgFatB1 (SEQ ID NO: 36) or UcFATB1(SEQ ID NO: 38), which accumulated mostly C12:0. Although lower than both EgFATB1 and UcFATB1, GmFatB4A showed in vivo activity as measured by the total medium chain fatty acid yield that was significantly higher than that of either GmFATB1A or GmFATB2A. This result suggests that FATB4A prefers medium-chain acyl- ACP, such as C8:0-ACP and C12:0-ACP. Table 3: Fatty Acid Composition in the Culture Medium of E.coli Cells Expressing FatB Genes Percentage of individual FATOTAL FAGene ID C6:0 C8 C10 C12 C14 C16 C16:1 C18 C18:1 C18:3 Others(µg / mL)Empty Vector 0.98 1.99 0.91 4.44 8.34 48.76 12.99 11.79 5.47 1.08 3.25 5.74 GmFATB1A 1.41 3.50 1.50 4.66 12.01 40.87 11.56 12.53 5.71 2.56 3.43 3.79 (SEQ ID NO:30) GmFATB2A 2.79 3.40 1.79 4.81 11.66 43.07 5.63 14.11 9.15 0.61 2.97 1.41 (SEQ ID NO:32) GmFATB4A 5.62* 37.52* 4.86* 27.90* 4.28 7.05* 1.95 4.01 1.33 3.03 2.40 9.02 (SEQ ID NO:34) EgFATB1 0.16 3.15 1.98 45.34* 15.64* 5.11* 24.91 0.67* 1.19 1.36 0.49* 60.43* (SEQ ID NO:36) UcFATB1 0.38 1.53 10.72* 57.87* 3.03* 5.82* 2.26* 2.00* 1.06 12.58* 2.69 8.55 (SEQ ID NO:38) Fatty acid composition and total FA content in the culture medium of E. coli K27 expressing the indicated FATB (without the chloroplast transit peptide) polypeptide. Values are the mean and values followed by an asterisk are significantly higher or lower than the empty vector for a given fatty acid. The column for others includes values for 14:1, 17:0, 17:1, and 18:2.Docket # 212436-WO-SEC-1 EXAMPLE 2
[0164] This example demonstrates structural modeling of a soybean thioesterase.
[0165] Soybeans, which produce dominantly long-chain (C16-C18) fatty acids, harbor 12 FAT genes 10 of which are specific for saturated fatty acid biosynthesis (FATB type). These FATs are highly similar to well-characterized plant FATs (with pair-wise sequence identity >40%) and contain conserved “hotdog” domains. Hotdog superfamily proteins, typically consisting of a long helix partially wrapped by a 5-stranded β-sheet, are ubiquitous from bacteria to eukaryotes. Members of this superfamily are heavily involved in fatty acid biosynthesis, regulation, and metabolism (Dillon et al., 2004). In prokaryotes, hotdog proteins often form a homodimer as a functional unit, while in eukaryotes, two tandem hotdog domains are fused into one reading frame. AlphaFold2 models of all 12 soybean FATs were built (Jumper et al.2021). The AlphaFold2 structure of a soybean FAT, GmFATB1A (SEQ ID NO: 2; residue F125-V402), as an example (Figure 1), exhibits an α / β structure with two similar hotdog domains. Its N-terminal domain (125-282) consists of 3 helices and 6 β-strands with the long α1 partially wrapped by a curved β-sheet (β2-β5-β4-β3-β1-βN), while its C-terminal domain has a long α4 surrounded by a β-sheet (β6-β8-β9-β10-β7). The two central β-sheets of the domains in GmFATB1A are aligned in an antiparallel manner, forming a large 11-strand sheet. The long helices α1 and α4 interact with each other with N-terminal ends in a lap joint manner. It has been demonstrated in California bay FAT (UcFATB1, SEQ ID NO: 24) that the catalytic active site is located at the inter-domain joint and is defined by conserved catalytic residues D311-N313-H315 in a loop of β6-α4 and E349 (Feng et al.2017). Visual inspection further reveals the substrate binding orientation around this catalytic center. The substrate acyl-ACP’s ACP likely binds to the back of the β-sheet. The phosphopantetheine (a prosthetic group for acyl-ACP) crosses the central sheet through the inter-domain wedge between β2 and β7, its thioester bond is positioned under the catalytic loop, and finally, the acyl chain is placed in a tunnel along α1 delineated by palmitic acid. The N-domain curved sheet and long α1 serve as two walls for the acyl-chain binding tunnel, which is further covered by α2 and α3 along with their linking loops.
[0166] The acyl-length determining factor in the FAT product of GmFATB1A (SEQ ID NO: 2), is the lipid binding site. This site is a tunnel formed between the N-domain’s α1 and a curved central sheet lined by residues M160-Q164-A167-L168-H170-V171-A174-L176-F181-G182- T184-M187 from α1-loop-α2, W194-V196-M199-T217-V219-M227-R229-W231-S247-W249Docket # 212436-WO-SEC-1 from the β-sheet, and I270-Y273-F274 from α3 (marked as black stars in Fig.2B, 2C, and 2D). Most of these residues are hydrophobically conserved, and a few are polar Q164-H170-T184- T217-S247, consistent with acyl hydrophobic chain binding. Surprisingly, a charged residue, R229, is at the center. The R229 guanidium group forms hydrogen bonds with T217 OG1 and H270 ND1, partially neutralizing its positive charge to favor hydrophobic interaction with acyl chain. Of the palmitic acid binding residues, six positions A167-V171-A174-V219-M227-R229 directly interact with C12-C13 to -C18 at the methyl end, likely playing a role in acyl-length specificity (Fig.1). An alignment between characterized plant FATBs (Figs.2A-2D) indicates that three residues, A167, A174 and V219 (boxed residues in Fig.2B), in GmFATB1A (SEQ ID NO: 2), play a crucial role in determining lipid specificity. The long-chain FATs have small residues Ala, Ala, and Val, while in the medium-chain enzymes, these positions are occupied by bulkier amino acids Thr / Ser, Ile / Leu, and Phe / Ile / Leu, respectively. The bulkier residues shrink the tunnel size and are predicted to restrict or block long-chain substrates’ access to tunnel active sites. GmFATB4A has larger residues, C140, C147, and F192, corresponding to GmFatB1A’s A167, A174, and V219, respectively. In detail, GmFATB1A’s A167, with a residue volume of 89 Å3, is replaced by a bigger C140 with a volume of 109 Å3in GmFATB4A, V171 (140Å3) by M144 (163Å3), A174(89Å3) by C147(109Å3), V219 (140Å3) by F192 (190Å3), and M227 (163Å3) by I200 (168Å3) (Fig.1).
[0167] To alter the substrate specificity of a long chain FATB (e.g., FATB1A, FATB1B, FATB2A, FATB2B, FATB3A, FATB3B, FATB4B, and FATB5A) to produce medium chain fatty acids, the lipid methyl end-recognition residues of the long-chain FATB are modified to contain a larger amino acid at one or more positions corresponding to position A167, V171, A174, V219, M227 and R229 in SEQ ID NO: SEQ ID NO: 2. EXAMPLE 3
[0168] This example demonstrates the alteration of fatty acid profile and activity of thioesterases by domain swapping.
[0169] The functional domains of mature tropical thioesterases (Palm kernel, EgFATB1, SEQ ID NO: 36 or California bay laurel, UcFATB1, SEQ ID NO: 38) were swapped with those of soybean thioesterase (SEQ ID NO: 32). The resulting variants were cloned and expressedDocket # 212436-WO-SEC-1 following the method in Example 1, and their fatty acid profiles were determined as outlined in Example 1.
[0170] As shown in Table 4, expressing SEQ ID NO: 54 (created by replacing 161 N-terminal residues of GmFATB2A with the corresponding residues from EgFATB1) led to a MCFA profile similar to that of EgFATB1 and enhanced the activity. Expressing SEQ ID NO: 56 (created by replacing 131 N-terminal residues of UcFATB1 with corresponding residues from GmFatB2A) resulted in a long chain fatty acid profile and improved activity. The replacement of 92 C-terminal residues of EgFATB1 with the corresponding residues from GmFATB2A resulted in SEQ ID NO 58, which exhibited a fatty acid profile similar to that of EgFATB1 when expressed in E. coli. The replacement of 85 or 195 C-terminal residues of GmFatB2A with the corresponding residues from EgFATB1 resulted in SEQ ID NOs: 59 and 61, respectively. These sequences resulted in a fatty acid profile, primarily consisting of C14:0 ~ C16:1, when expressed in E. coli. Swapping 100 middle domain residues of GmFATB2A with corresponding residues from GmFATB1 to create SEQ ID NO 63 significantly improved the activity and shifted the profile towards C14:0 and C16:1. Replacing the N-terminal residues of GmFATB2A with shorter domain swaps from medium chain-preferring TEs resulted in a long-chain fatty acid profile similar to that of GmFATB2A. This is exemplified by the fatty acid profiles of SEQ ID NO: 94 (created by replacing 76 N-terminal residues of GmFATB2A with the corresponding residues from EgFATB1), SEQ ID NO: 98 (created by replacing 57 mid N-terminal residues of GmFATB2A with the corresponding residues from GmFATB4A), and SEQ ID NO: 100 (created by replacing mid 57 N-terminal residues of GmFATB2A with the corresponding residues from UcFATB1). In SEQ ID NO: 96, in which 82 residues of GmFATB2A were replaced with the corresponding residues from the mid N-terminal region of EgFATB, a medium-chain fatty acid (MCFA) profile was observed, although it was less pronounced compared to that of SEQ ID NO: 54 (which involved a longer swap). Table 4A: Fatty Acid Composition and total FA content in the Culture Medium of E. coli Cells Expressing Domain Swapped FatB Genes Percentage of individual FADocket # 212436-WO-SEC-1 36 7.02 ± 1.82* 6.40 ± 0.71* 54.02 ± 2.88* 7.57 ± 0.68 32 4.05 ± 0.86 4.44 ± 1.08* 5.34 ± 0.39 9.67 ± 0.63 * *y p . lls Expressing Domain Swapped FatB Genes Percentage of individual FA SEQ ID NO:Docket # 212436-WO-SEC-1 96 30.56 ± 5.55* 13.91 ± 2.97 10.58 ± 4.45 1.87 ± 0.68 1.85 ± 0.71* 98 48.70 ± 2.02 12.69 ± 4.04 10.66 ± 3.76 2.36 ± 1.01 5.10 ± 1.17lls Expressing Domain Swapped FatB Genes SEQ ID NO: Total FA (µg / Ml) atty acids produced by E. coli K27expressing the indicated domain swap variants of mature acyl-ACP thioesterases (without chloroplast transit peptides). Values are mean ± standard error. Asterisks (*) indicate values significantly higher or lower than empty vector control for each fatty acid. The "Others" column includes 6:0, 14:1, 17:0, 17:1, and 18:2 fatty acids. EXAMPLE 4
[0171] This example demonstrates the physiological functionality of FATB4 in plants.Docket # 212436-WO-SEC-1
[0172] To confirm whether GmFATB4A is functional in plants, transgenic vectors were made to express either the GmFATB1A (SEQ ID NO:2) or GmFATB4A (SEQ ID: 14) in plants. N. bentnhamiana leaves were agro-infiltrated with either an empty vector (control), the GmFATB1A construct, or the GmFATB4A construct. Leaf discs were collected four days post- infiltration and lyophilized for three days. The fatty acid profile of hexane oil extracted from leaf discs was determined by gas chromatography. N. bentnhamiana leaves expressing GmFATB4A accumulated C12:0 (~0.6%) and C14:0 (1.9%) fatty acids, with no significant change in C16:0 levels compared to leaves expressing the empty vector control, which produced no detectable C12:0 or C14:0. By contrast, leaves expressing GmFATB1A accumulated up to 4.7% C14:0 and a significantly higher amount of C16:0 relative to both the empty vector control and GmFATB4A (Table 5A). These results suggest that GmFATB4A, although minimally expressed in soybean (Table 2), is functional in plants and exhibits a preference for shorter-chain acyl-ACP substrates, in contrast to GmFATB1A, which favors longer-chain substrates. In addition, transgenic vectors to express GmFATB2A / EgFATB1v161N (SEQ ID NO: 92), GM- FATB2A A170S / V174F / A177C / V222F (SEQ ID NO: 134), GM-FATB4A_F192V (SEQ ID NO:142), or CH-FATB2 (SEQ ID NO: 174) were transiently expressed in N.benthamiana as described above. As shown in table 5B, expression of GmFATB2A / EgFATB1v161N and GM- FATB4A_F192V significantly increased MCFA accumulation in N. benthamiana leaves relative to empty vector control. Overexpression of GM-FATB2A A170S / V174F / A177C / V222F significantly increased C14:0 accumulation but not overall MCFA. The fatty acid profile of GM- FATB2A A170S / V174F / A177C / V222F (shown to have specific activity toward medium chain substrates, Tables 8, 10, and 11) was similar to that of ChFATB2 which is reported to have a C8:0 and C10:0 specificity, but none of these fatty acids accumulated in leaves in our experimental conditions. Table 5A: Fatty acid composition of Nicotiana benthamiana leaves expressing GM-FATB1A and GM-FATB4A C8 C12 C14 C16 C18 C18:1 C18:2 C18:3 Others %MCFA Empty vector 0 0 0 21.11 2.65 2.08 14.38 52.2 7.58* 0 GmFatB1A 0 0 4.65* 52.00* 3.32* 0.88* 9.23* 27.08* 2.83* 4.65*Docket # 212436-WO-SEC-1 (SEQ ID NO:2) GmFatB4A 0.04 0.61* 1.94* 21.78 3.08* 1.64* 11.86* 52.52 6.52* 2.59* (SEQ ID NO:14) EgFatB1 0 4.43* 15.80* 28.44* 2.59 0.92* 10.56* 33.62* 3.62* 20.24* (SEQ ID NO:22) Fatty acid profile of hexane oil extracts from leaf discs. Values are mean, and values followed by an asterisk are significantly higher or lower than the empty vector control for a given fatty acid. The column for others includes values for 6:0, 10:0, 14:1, 17:0, and 17:1 fatty acid. Table 5B: Fatty acid composition of Nicotiana benthamiana leaves expressing GM-FATB2A and GM-FATB4A variants C12:0 C14:0 C16:0 C18:0 C18:1 C18:2 C18:3 others MCFA Em t v t r 015 004 2155 317 146 1313 5385 665 019 * * * *Docket # 212436-WO-SEC-1 Fatty acid profile of hexane oil extracts from leaf discs. Values are mean, and values followed by an asterisk are significantly higher or lower than the empty vector control for a given fatty acid. The column for others includes values for 14:1, 15:0, 16:1, 17:0, and 17:1 fatty acid. EXAMPLE 5
[0173] This example demonstrates increasing soybean seed medium chain fatty acid content by genome editing.
[0174] There are two gene knockout designs as shown in Table 6. Table 6: Guide RNAs and their corresponding target genes gRNA Target Gene 1 Target Gene 2
[0005] n xper ment , t e G - -C was used to create rames t noc outs n both Glyma.04g151600 (FATB2A) and Glyma.06g211300 (FATB2B) genes. In the experiment 2, the GM-TE2-CR3 and GM-TE2-CR4 were used to create frameshift knockouts in both Glyma.04g151600 (FATB2A) and Glyma.05g012300 (FATB1A) genes. The transformation experiments were carried out by Agrobacterium-mediated embryo axis transformation process. T0 plants were molecularly characterized for the gene edits. The segregating T1 seeds were planted and molecularly characterized to identify homozygous T1 plants. The T2 seeds were evaluated for the fatty acid composition using NIR spectroscopy. As shown in Table 7, knocking out Glyma.04g151600 (FATB2A) and Glyma.06g211300 (FATB2B) resulted in no significant changes in seed oil composition relative to WT. Seeds in which both Glyma.04g151600 (FATB2A) and Glyma.05g012300 (FATB1A) were knocked out showed significant reduction in C16:0 and total saturates compared to WT or segregating null seeds (Table 7). As described in Example 3, the seeds from experiment 1 and 2 can be used as transformation initiation materials for transgenic or genome editing expression of MCFA-preferred FATB variants.
[0176] Genome editing can also be designed and carried out to use a template-based double strand break repair process to either insert or replace the native FATB4 promoter at its nativeDocket # 212436-WO-SEC-1 locus with a strong seed or cotyledon specific promoter, such as soybean conglycinin promoter or glycinin 1 promoter. In this case, gRNAs can be designed at the 5’end of FATB4 gene promoter, and near or close to the 5’UTR of the FATB4A gene (SEQ ID NO: 48) or FATB4B to enable the stronger seed-preferred promoter insertion or promoter swap.
[0177] Alternatively, the FATB4A genomic sequences, including its exons and introns, can be used to replace the exons / introns of seed storage protein genes, such as one or more of the seven conglycinin subunit genes and five glycinin genes at their native loci. Specific gRNAs need to be designed at or near the 5UTRs and 3’UTR of these native seed storage protein genes to enable gene swap. Transformation will be carried out by Agrobacterium-mediated embryo axis transformation process. T0 plants will be molecularly characterized for the gene edits. The segregating T1 seeds will be planted and molecularly characterized to identify homozygous T1 plants. The T2 seeds will be evaluated for the fatty acid compositions, especially for the accumulation of MCFA in soybean seeds. Table 7: Fatty acid composition of GM-FATB1A / GM-FATB2A- or GM-FATB2A / GM- FATB2B-knockout soybean T2 seed. Target Gene Event Zygosity 16:0 16:1 18:0 18:1 18:2 18:3 Total Sats NA N / A WT 11.33 0.15 3.43 24.22 52.1 8.73 14.77 GmFATB2A (SEQ ID NO:6) 32 Hom.hom 11.25 0.15 3.55 21.9 54.15 8.95 14.8 and GmFATB2B (SEQ ID NO:8) GmFATB1A 1 Null 11.7 0.15 3.6 21.15 54.15 9.3 15.3 (SEQ ID NO:2) and GmFATB2A 1 Hom.hom 4.20* 0.15 2.70* 25.8 57.25 9.9 6.90* (SEQ ID NO:6) 2 Hom.hom 4.23* 0.27 2.77* 24.43 58.53* 9.77 7.00* 3 Hom.hom 5.17* 0.23 2.87* 23.55 58.20* 9.98 8.03* The oil composition of greenhouse-grown T2 seeds was determined by NIR spectroscopy. Values represent means, and values followed by an asterisk (*) are significantly higher or lower than wild-type or null seeds for the given fatty acid. EXAMPLE 6Docket # 212436-WO-SEC-1
[0178] This example demonstrates increasing soybean seed medium chain fatty acid content by altering FATB substrate specificity.
[0179] In Example 2, four residues (A167, V171, A174, and V219) in GmFATB1A or (A170, V174, A177, and V222) in GmFATB2A are identified as chain-length determinants due to their lipid binding site positions. Other residues (L176, W194, M227, R229, W249, I270, and Y273) within the same site are predicted to also directly interact with C12-C13 to C18 at the methyl end of the substrate. Mutagenesis can generate variants with residue substitutions, individually or in combination. To shift GmFatB1A or GmFATB2A specificity from long-chain to medium- chain fatty acids, smaller or hydrophobic residues are replaced with larger or hydrophilic residues such as cysteine, methionine, phenylalanine, serine, tyrosine, and isoleucine, specifically corresponding to residues from GmFATB4A (C140, M144, C147, and F192), or corresponding residues from other medium-chain thioesterases. Resulting variants such as A167C, A167S, A174C, V219F, V171M, M227I, and W249Y are introduced into E. coli, with fatty acid profiles determined as in Example 1 for initial testing.
[0180] The coding DNA sequences corresponding to mature protein sequences, but containing specific amino acid substitutions, were cloned as described in Example 1. Data were obtained and analyzed following the procedures outlined in Example 1. As shown in Table 8, expression of variant A118S or A118S / V170F of GmFATB1A in E.coli resulted in a significant increase in C14:0, shifting the fatty acid profile toward shorter chain fatty acids. E.coli cells expressing variants of GmFATB4A F136V or C84A / F136V accumulated more C12:0 as compared to GmFATB4A WT, which accumulated more C8:0, whereas variant M884F accumulated more C6:0. E.coli cells expressing variants of GmFATB2A (A114S / V118F / A121C / V166F) accumulated more than 50% MCFA in the growth medium. Replacing the corresponding four residues in GmFATB4A (C84A / M88V / C91A / / F136V) (or in medium-chain specific EG-FATB1 or UC-FATB1) with smaller or hydrophobic residues resulted in accumulation of longer chain fatty acids (Table 8).
[0181] This result demonstrates that these four amino acid residues are involved in determining the chain length specificity of thioesterase. Table 8: Fatty Acid Composition in the Culture Medium of E. coli Cells Expressing Amino Acid Substitution VariantsDocket # 212436-WO-SEC-1 SEQ IDPercentage of individual FATotal FANO: C6 C8 C10 C12 C14 C16 C16:1 C18 C18:3 Others (µg / mL) ngt e ndcated w d-type or substtuton varants o mature (wt out c oropast trans t peptde) acy-ACP thioesterases. Values are means, and values followed by an asterisk (*) are significantly higher or lower than empty vector control for each fatty acid. The "Others" column includes C14:1, C17:0, C17:1, C18:1, and C18:2 fatty acids.
[0182] To increase MCFA accumulation in soybean, transgenic vectors were made with the GmFATB2A / EgFATB1v161N (SEQ ID NO: 91), GM-FATB2A A170S / V174F / A177C / V222F (SEQ ID NO: 133, Table 8), GM-FATB4A_F192V (SEQ ID NO: 141), and CH-FATB2 (SEQ ID NO: 173) under the control of a soybean seed storage protein beta-conglycinin promoterDocket # 212436-WO-SEC-1 (SEQ ID: 39) and a Phaseolin terminator (SEQ ID NO: 40). T0 plants will be molecularly characterized to get single copy events. The segregating T1 seeds and homozygous T2 seeds will be evaluated for the accumulation of MCFA in soybean seeds. Increased MCFA accumulation to over 10% is expected in seed expressing each of the above gene variants. The transgenics described above can also be carried out in FATB knockout backgrounds as a gene replacement to further increase medium chain fatty acids. These FATB knockout experiments are described in Example 5.
[0183] Genome editing is designed and carried out to use a template-based double strand break repair process to introduce the amino acid replacements in either FATB2A (such as the A170S / V174F / A177C / V222F described above) or FATB1A gene (such as A167S or A167S / V170F described above). Alternatively, FATB4A_F192V, gene replacement can be made to increase C12:0 (as shown in E. coli, Table 8) along with a promoter swap (such as beta- conglycinin) to increase seed specific gene expression. In these cases, the donor DNA will contain nucleotide changes to encode new amino acids, and gRNAs are designed near the amino acid replacements sites to enable efficient donor DNA replacement by double strand break repair mechanism. EXAMPLE 7
[0184] An AI-based strategy was employed to shift the substrate acyl chain length selectivity of acyl-ACP thioesterase, primarily using three approaches, (1) reference-guided mutation selection and combination, (2) prediction of mutational effect on substrate selectivity with contrastive fitness-learned protein language model (Zhao et al.2024, International Conference on Research in Computational Molecular Biology, 470-474), and (3) zero-shot prediction of mutation tolerance with SaProt – a structure-aware model with a specialized vocabulary (Su et al.2023, bioRxiv, 2023-10).
[0185] Drawing on prior studies on substrate selectivity engineering of acyl-ACP thioesterases (Jing et al 2018, Nature Communications, 9(1), 860; Jing et al.2024, Frontiers in Bioengineering and Biotechnology, 12, 1379121; Hernandez Lozada et al.2018, ACS Synthetic Biology, 7(9), 2205-2215), key mutations known to enhance medium-chain fatty acid (MCFA) production were identified. These mutations, along with all possible substitutions at the corresponding residues, were evaluated for their predicted impact on substrate selectivity and their likelihood of beingDocket # 212436-WO-SEC-1 tolerated. The substrate selectivity prediction model was trained using contrastive fitness learning on ESM-1v (650M parameters) with internal fatty acid production data from GmFatB1A variants (sequences 104, 108, 112, 116, 120, 124, 128, and 132). SaProt (650M, AF2-based) was used to assess mutation tolerance.
[0186] Based on these predictions, identified mutations were selected unless they indicated significantly negative effects on both properties – in such cases, the most favorable mutation at that residue was chosen instead. Combinations of these selected mutations were then determined by integrating previously tested combinations from the literature and model-based predictions of their effects on substrate selectivity. The coding DNA sequences corresponding to the mature protein sequence of GmFATB1A, but containing specific amino acid substitutions, were cloned as described in Example 1. Data were obtained and analyzed following the procedures outlined in Example 1. When expressed in E. coli, several variants showed increased medium-chain fatty acid (MCFA) composition in the culture medium relative to the wild-type (WT) GmFATB1A. Notably, GmFATB1A_V170F / G172K / R179G (SEQ ID NO: 206) and GmFATB1A_V166I / V170F / N176I / R179A / R180S (SEQ ID NO: 214) significantly increased C8:0 and C10:0, and SEQ ID NO: 214 also significantly increased C12:0, resulting in over 50% MCFA. Variants GmFATB1A_A101P (SEQ ID NO: 182), GmFATB1A_StartAt43_D70S / I107M (SEQ ID NO: 186), and GmFATB1A_V170F / N176I / L210F / I213L / D229P / D235N / L242F (SEQ ID NO: 194) increased C8:0 and C14:0, while variants GmFATB1A_StartAt61 (SEQ ID NO: 190) and GmFATB1A_V170F / R179G (SEQ ID NO: 202) increased C14:0. In contrast, the variant GmFATB1A_V170F / G172Q / N176I / R179G (SEQ ID NO: 218) did not result in significant changes in the fatty acid profile compared to the WT (Tables 9A-9C). Table 9A: Fatty Acid Content in the Culture Medium of E. coli Cells Expressing AI-guided Amino Acid Substitution Variants Percentage of individual FA SEQ ID NO: C6:0 C8:0 C10:0 C12:0 C14:0 Empty vector 0.00 ± 0.00* 4.88 ± 1.03 3.02 ± 0.91 5.85 ± 0.99* 11.24 ± 1.03* 30 1.91 ± 0.41 2.29 ± 0.47 4.50 ± 0.24 9.54 ± 0.09 15.38 ± 0.45Docket # 212436-WO-SEC-1 182 2.80 ± 0.68 16.48 ± 0.21* 6.92 ± 0.03* 3.05 ± 0.50* 31.01 ± 1.51* 186 3.51 ± 0.73 22.03 ± 4.08* 8.30 ± 0.75* 3.94 ± 0.51* 20.05 ± 3.75 190 1.23 ± 0.61 8.69 ± 3.16 4.67 ± 0.60 2.31 ± 0.23* 24.83 ± 1.87* 194 2.08 ± 0.25 19.23 ± 2.98* 8.65 ± 0.64* 9.25 ± 0.96 19.62 ± 1.87 198 3.27 ± 0.71 22.71 ± 1.69* 7.47 ± 0.93* 8.05 ± 1.04 14.22 ± 0.35 202 1.30 ± 0.17 11.28 ± 1.92* 6.68 ± 1.09 7.26 ± 1.01 24.10 ± 0.38* 206 2.25 ± 0.14 20.25 ± 1.85* 17.97 ± 3.11* 13.36 ± 3.40 12.76 ± 0.08* 210 2.77 ± 0.40 14.43 ± 1.95* 5.08 ± 0.65 5.84 ± 0.22* 9.63 ± 1.07* 214 1.10 ± 0.23 5.28 ± 0.03* 21.37 ± 5.74* 34.19 ± 0.65* 4.32 ± 0.59* 218 2.55 ± 0.26 8.45 ± 3.25 6.00 ± 1.41 11.97 ± 1.03 13.19 ± 0.97 Table 9B: Fatty Acid Content in the Culture Medium of E. coli Cells Expressing AI-guided Amino Acid Substitution Variants Percentage of individual FA SEQ ID NO: C16:0 C16:1 C18:0 C18:3 Others Empty vector 55.85 ± 1.87* 4.54 ± 0.52* 11.13 ± 1.36* 0.00 ± 0.00* 3.48 ± 0.09 30 49.10 ± 0.48 3.05 ± 0.07 4.47 ± 0.15 2.33 ± 0.06 5.50 ± 1.81 182 10.94 ± 3.03* 24.53 ± 4.20* 3.99 ± 2.01 0.22 ± 0.22* 0.17 ± 0.17 186 22.82 ± 2.92* 10.57 ± 3.88 8.17 ± 2.23 0.00 ± 0.00* 0.28 ± 0.28* 190 38.98 ± 1.73* 13.38 ± 0.80* 4.24 ± 0.28 0.18 ± 0.18* 3.56 ± 2.07 194 26.12 ± 0.45* 8.66 ± 1.45* 4.02 ± 0.88 0.30 ± 0.15* 1.88 ± 0.20 198 28.96 ± 0.92* 7.44 ± 0.57* 6.45 ± 0.40* 0.00 ± 0.00* 9.64 ± 7.61 202 25.16 ± 0.81* 18.80 ± 1.40* 2.99 ± 0.08* 0.62 ± 0.32* 1.94 ± 0.31 206 19.86 ± 3.70* 5.75 ± 1.21 7.10 ± 0.53* 0.80 ± 0.80 0.97 ± 0.97 210 45.20 ± 2.92 3.45 ± 0.07* 9.80 ± 3.02 0.53 ± 0.53* 1.64 ± 0.27Docket # 212436-WO-SEC-1 214 18.27 ± 6.80* 0.83 ± 0.40* 2.47 ± 0.20* 17.63 ± 2.80* 3.39 ± 1.85 218 40.18 ± 4.02 3.66 ± 0.55 7.65 ± 2.47 2.10 ± 0.14 4.57 ± 1.33 Table 9C: Fatty Acid Content in the Culture Medium of E. coli Cells Expressing AI-guided Amino Acid Substitution Variants SEQ ID NO Total FA (µg / mL) Empty vector 1.70 ± 0.19* 30 5.27 ± 0.07 182 4.33 ± 2.03 186 1.86 ± 0.49* 190 3.37 ± 0.18* 194 3.86 ± 0.73 198 2.63 ± 0.05* 202 5.05 ± 0.17 206 2.12 ± 0.03* 210 2.68 ± 0.62* 214 10.66 ± 1.61* 218 5.22 ± 0.36 Tables 9A-9C: Data represents the molar percentages and total concentrations of fatty acids accumulated in the culture medium of E. coli expressing the indicated wild-type (GmFATB1A) or AI-guided substitution variants of mature (without chloroplast transit peptide) acyl-ACP thioesterases. Values are means ± standard error. Asterisks (*) indicate values significantly higher or lower than empty vector control for each fatty acid. The "Others" column includes C14:1, C17:0, C17:1, C18:1, and C18:2 fatty acids.
[0187] To increase MCFA accumulation in soybean, transgenic vectors are made with a high medium-chain fatty acid AI-guided variant FATB variant sequence under the control of a promoter, for example the soybean seed storage protein beta-conglycinin promoter (SEQ ID: 39).Docket # 212436-WO-SEC-1 T0 plants are molecularly characterized to get single copy events. The segregating T1 seeds and homozygous T2 seeds are evaluated for the accumulation of MCFA in soybean seeds. Increased MCFA accumulation to over 10% is expected in seed expressing each of the above high medium-chain fatty acid AI-guided FATB variants. The transgenics described above can also be carried out in FATB knockout backgrounds as a gene replacement to further increase medium chain fatty acids. These FATB knockout experiments are described in Example 5.
[0188] Alternatively, genome editing is designed and carried out to use a template-based double strand break repair process to introduce the high medium-chain fatty acid AI-guided variant amino acid replacements in an endogenous FATB genes such as the FATB2A or FATB1A gene or to introduce the high medium-chain fatty acid AI-guided variant amino acid by SDN3. T0 plants are molecularly characterized to get single copy events. The segregating T1 seeds and homozygous T2 seeds are evaluated for the accumulation of MCFA in soybean seeds. Increased MCFA accumulation to over 10% is expected in seed expressing each of the above high medium-chain fatty acid AI-guided variant FATB variants. EXAMPLE 8
[0189] This example demonstrates increasing soybean seed medium chain fatty acid content.
[0190] To increase MCFA content, a combination of one or more high oil genes, such as MFT, SWT, SUT, ODP1, and / or DGAT are introduced into a soybean plant along with a genetically edited GmFATB gene (such as GmFATB1A or GmFATB4) that results in increased MCFA. The high medium chain fatty acid FATB variants, such as GmFATB2A (A170S / V174F / A177C / V222F), described in Examples 6 and 7 will be transformed into stable soybean using a seed specific promoter to produce MCFAs in soybean seeds. These variants can be expressed in KASII knockout background for potential increase in seed MCFA content. Additionally, the corresponding residues for altering the specificity of GmFATB variants with MCFA specificity can be edited into the native GmFATB gene, such as GmFATB1A in a high oil soybean background using CRISPR-Cas9 or other suitable gene editing technologies. EXAMPLE 9
[0191] This example demonstrates the specific activity and kinetic properties of thioesterases and variants on different Acyl-ACP substratesDocket # 212436-WO-SEC-1
[0192] The specific activity of GmFATB4A and other TEs were tested on different substrates C8:0-ACP, C12:0-ACP, C14:0-ACP, and C16:0-ACP, while the kinetic properties (Km, Vmax and Kcat) were estimated for C8:0-ACP and C16:0-ACP.
[0193] The proteins used for enzymes assays were Escherichia coli Acyl Carrier Protein (ACP), Vibrio harveyi Acyl-Acyl Carrier Protein Synthetase (AasS), Bacillus subtilis 4'- Phosphopantetheinyl Transferase (SfP), and soybean thioesterases or .
[0194] DNA sequences encoding mature proteins were cloned into vectors. Constructs were transformed into E. coli BAP1 (ACP) or BL21(DE3) (all other proteins) competent cells. Single colonies were grown in 1 L LB medium containing 50 µg / mL kanamycin at 37°C until OD₆₀₀ = 0.6-0.8. Protein expression was induced with 0.1 mM IPTG (for thioesterases) or 1 mM IPTG (all other proteins) for 18-24 h at 20°C (thioesterases) or 30°C (all other proteins).
[0195] Protein Purification: Cells were harvested by centrifugation (6,000 × g, 20 min, 4°C) and stored at -80°C. Cell pellets were resuspended in bacterial protein extraction reagent at 5 mL / g pellet, supplemented with 1mM EDTA, and incubated with gentle agitation for 2 hours at room temperature. Lysates were clarified by centrifugation (16,000 × g, 45 min, 4°C). Soluble proteins were purified using Ni-NTA or Co-NTA affinity chromatography according to manufacturer's protocols and protein concentrations were determined. Purified proteins were buffer-exchanged into 100 mM sodium phosphate (pH 8.0) containing 10% glycerol using centrifugal filters and stored at -80°C.
[0196] ACP processing: His-tagged ACP was treated with TEV protease (1:100 molar ratio) overnight at 30°C in 15 mL falcon tubes with gentle agitation. Cleaved ACP was separated from uncleaved protein and protease by Ni-NTA chromatography (cleaved ACP in flow-through) and buffer-exchanged as described above.
[0197] Holo-ACP synthesis: Apo-ACP (200 µM) was converted to holo-ACP using SfP (5 µM), coenzyme A (5 mM), TCEP (5 mM), and MgSO₄ (10 mM) in 100 mM sodium phosphate (pH 8.0) at 37°C overnight. Complete conversion was verified by urea-PAGE (18% polyacrylamide, 4 M urea), after which holo-ACP was separated from other proteins by Ni-NTA chromatography (holo-ACP in flow-through) and buffer-exchanged as described above. Acyl-ACP synthesis: To 50 µM of holo-ACP was added; AasS (5 µM), ATP (10 mM), TCEP (2.5 mM), glycerol (4%), and fatty acid substrate (200 mM octanoate, 100 mM laurate, 50 mM myristate, or 40 mM palmitate) were added in a final volume of 2 mL. Reactions were incubatedDocket # 212436-WO-SEC-1 at 37°C for 24 h with shaking (250 rpm). Unreacted CoA and traces of holo-ACP were quenched with excess DTNB for 30 min with gentle shaking at room temperature. Acyl-ACP products were purified by Ni-NTA chromatography (flow-through collection) and buffer-exchanged into 100 mM sodium phosphate (pH 7.4) containing 10% glycerol. Complete acyl-ACP formation was confirmed by urea-PAGE.
[0198] The specific thioesterase activity of the enzymes was determined by tracking the formation of holo-ACP using 5,5'-dithiobis(2-nitrobenzoic acid) (DTNB). TNB formation was monitored for 40 minutes at 412 nm using a 96-well plate reader at 25°C, with all assay reactions carried out in 100 mM sodium phosphate buffer (pH 7.4). Prior to each assay, acyl-ACP substrate, BSA, and DTNB (all prepared in 100 mM Na₂HPO₄, pH 7.4) were combined in the same buffer and incubated with shaking at room temperature for 30 minutes to ensure complete mixing.
[0199] Specific activity assays: Reactions were performed in 200 µL total volume by transferring 180 µL of the pre-incubated substrate mixture to 96-well plate wells, each containing 20 µL of thioesterase (dissolved in 100 mM Na₂HPO₄, pH 8.0). The final concentrations in each reaction were: 30 µM acyl-ACP substrate, 10 µg / mL BSA, 250 µM DTNB, and 2 µM thioesterase. OD readings were subsequently converted into holo-ACP concentration using the molar extinction coefficient of TNB (14,150 M⁻¹cm⁻¹), and initial reaction rates were determined from linear portions of progress curves to express specific activity in µmol / min / mg.
[0200] Kinetic parameter determination: To determine enzyme kinetic properties, the same general procedure was followed with the following modifications: total reaction volume was reduced to 100 µL (90 µL substrate mixture + 10 µL enzyme), and varying substrate concentrations (0-160 µM acyl-ACP) were tested while maintaining constant concentrations of other components. The concentration of free holo-ACP released during the reaction was estimated from a holo-ACP standard curve run in parallel with the experiment. Initial reaction rates were determined from linear portions of progress curves and subsequently fitted to the Michaelis-Menten equation to determine Km and Vmax values.
[0201] GmFATB4A showed a strong specific activity toward C8:0-ACP and C12:0-ACP while GmFATB1A and GmFATB2A showed a preference for 14:0-ACP and C16:0-ACP substrates. California Bay (UC-FATB1) showed preference for C12:0-ACP and C14:0-ACP asDocket # 212436-WO-SEC-1 expected.
[0202] Substituting specific residues (C84A_M88V_C91A_F136V) shifted the preference of GmFATB4A to C16:0-ACP substrates. By contrast, substituting the corresponding residues in GmFATB2A resulted in a C8:0-ACP preference (Table 10). The substrate preference of GmFATB1A was not significantly altered by these changes.
[0203] As shown in Table 11, the specific activity of the wild-type (WT) enzymes - GmFATB1A, GmFATB2A (on C8:0-ACP), and GmFATB4A (on C16:0-ACP) - was too low to obtain meaningful kinetic parameters under the assay conditions described above. GmFATB1A exhibited moderate affinity for C16:0-ACP, with a Km of approximately 28 μM and a turnover number of about 0.04. In contrast, the Km of GmFATB2A was ten times higher than that of GmFATB1A, indicating much lower affinity for C16:0-ACP substrates, although it displayed a higher turnover number when sufficient substrate was available. GmFATB4A showed moderate affinity for C8:0-ACP, with a Km of around 37 μM and a turnover number greater than 0.2 μM / s, but had low affinity for C16:0-ACP, as indicated by a high Km of 89 μM.
[0204] Amino acid substitution variants demonstrated altered substrate preferences. Both GmFATB1A_A118S / V122F / A125C / V170F and GmFATB2A_A114S / V118F / A121C / V166F showed increased affinity for C8:0-ACP substrates relative to their respective WT enzymes. For GmFATB2A_A114S / V118F / A121C / V166F, the specific activity on C16:0-ACP was too low to determine reliable kinetic properties, whereas GmFATB1A_A118S / V122F / A125C / V170F appeared to have higher affinity (low Km, 2.5 μM) for C16:0-ACP compared to WT (Km = 29 μM). This is consistent with the E. coli data (Table 8), which showed moderate accumulation of MCFA and high amount of C16:0, suggesting GmFATB1A has different characteristics relative to GmFATB2A. This may also explain why knocking out GM-FATB2A / GMFATB2B does not result in a reduction in C16:0 while seeds without GMFATB1A / GMFATB2A have significantly reduced C16:0 (Table 7). The variant GmFATB4A_C84A / M88V / C91A / F136V exhibited increased affinity for C16:0-ACP, but its specific activity on C8:0-ACP was too low to allow reliable kinetic measurements.
[0205] These kinetic data indicate that GmFATB1A and GmFATB2A preferentially utilize C16:0-ACP substrates, whereas GmFATB4A favors C8:0-ACP. Importantly, the four amino acid positions targeted in these variants play a critical role in determining substrate specificity. The engineered variants, such as GmFATB1A_A118S / V122F / A125C / V170F,Docket # 212436-WO-SEC-1 GmFATB2A_A114S / V118F / A121C / V166F, and GmFATB4A_C84A / M88V / C91A / F136V, exhibited a switch in substrate preference compared to their respective wild-type enzymes. This demonstrates that specific amino acid substitutions can effectively alter the chain-length specificity of these thioesterases, providing a rational strategy for engineering enzymes with desired fatty acid profiles. Table 10: Specific activity of GmFATB4A and other thioesterases on different Acyl-ACP substrates C8:0-ACP C12:0-ACP C14:0-ACP C16:0-ACP GM-FATB1A 79.44 ± 1.92 156.72 ± 13.02 362.75 ± 41.89 407.16 ± 11.26 GmFatB1A A118S_V122F_A125C_V170F 104.09 ± 6.31 142.24 ± 16.85 102.91 ± 28.89 116.03 ± 6.77 GmFatB2A 7.40 ± 4.07 9.04 ± 1.02 5.75 ± 2.91 43.55 ± 22.58 GmFATB2A A114S / V118F / A121C / V166F 195.58 ± 10.46 55.61 ± 19.23 46.84 ± 13.36 83.00 ± 16.08 GmFatB4A 316.57 ± 26.30 432.21 ± 17.42 298.77 ± 22.04 88.44 ± 4.43 GmFatB4A C84A_M88V_C91A_F136V 135.59 ± 7.79 198.79 ± 13.19 189.39 ± 22.49 367.64 ± 5.81 UcFatB1 56.84 ± 0.73 463.55 ± 11.89 439.95 ± 30.01 268.61 ± 11.49 The specific activity estimates are in µmole / min / mg enzyme. Values are mean ± standard error from 4 technical replicates Table 11: Kinetic parameters of GM-FATB1A, GM-FATB2A, GM-FATB4A and their 4-amino acid substitution variants Enzyme Substrate Vmax (µM / s) Km (µM) kcat (s-1)Docket # 212436-WO-SEC-1 C8:0 nd nd nd GmFatB2A eties
[0206] All publications and patent applications in this specification are indicative of the level of ordinary skill in the art to which this invention pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated by reference.
[0207] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Unless mentioned otherwise, the techniques employed or contemplated herein are standard methodologies well known to one of ordinary skill in the art. The materials, methods and examples are illustrative only and not limiting.
[0208] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.Docket # 212436-WO-SEC-1
[0209] Units, prefixes and symbols may be denoted in their SI accepted form. Unless otherwise indicated, nucleic acids are written left to right in 5’ to 3’ orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively. Numeric ranges are inclusive of the numbers defining the range. Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
Claims
Docket # 212436-WO-SEC-1 We claim:
1. A soybean seed comprising an introduced genomic modification in an endogenous fatty acyl- ACP thioesterase B (FATB) gene to encode a modified FATB polypeptide, the modified FATB polypeptide having altered substrate specificity as compared to a corresponding non- modified FATB polypeptide.
2. The soybean seed of claim 1, wherein the substrate specificity of the modified FATB polypeptide is altered to produce medium chain fatty acids (MCFAs).
3. The soybean seed of claim 1 or 2, wherein at least 2% of total fatty acid content in the soybean seed comprises a medium chain fatty acid.
4. The soybean seed of any one of claims 1-3, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the positionDocket # 212436-WO-SEC-1 corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof.
5. The soybean seed of claim 4, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 32, or 34 and comprises a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a phenylalanine at the position corresponding to position 171 of SEQ ID NO:
2.
6. The soybean seed of claim 4, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprises a. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2; b. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a tyrosine at the position corresponding to position 249 of SEQ ID NO: 2; c. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; d. comprises a cysteine at the position corresponding to position 167 of SEQ ID NO: 2 a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; e. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2, and a leucine at the position corresponding to position 270 of SEQ ID NO: 2;Docket # 212436-WO-SEC-1 f. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; g. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 174 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 219 of SEQ ID NO: 2, a glycine at the position corresponding to position 228 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; h. a threonine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, a leucine at the position corresponding to position 174 of SEQ ID NO: 2, a leucine at the position corresponding to position 196 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 225 of SEQ ID NO: 2, an alanine at the position corresponding to position 228 of SEQ ID NO: 2, a serine at the position corresponding to position 229 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 259 of SEQ ID NO: 2, a leucine at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; i. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; or j. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2, aDocket # 212436-WO-SEC-1 cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO:
2.
7. The soybean seed of any one of claims 1-6, wherein the soybean seed further comprises at least one additional genome modification to increase the expression or activity of the FATB polypeptide or decrease expression of a long-chain fatty acid FATB polypeptide.
8. The soybean seed of claim 7, wherein an endogenous long-chain fatty acid FATB type thioesterase is replaced with a medium-chain fatty acid FATB type thioesterase or variant thereof.
9. The soybean seed of any one of claims 1-8, wherein the seed further comprises a modification to enhance seed oil content.
10. The soybean seed of claim 9, wherein the modification to enhance seed oil content being in a gene encoding at least one of (i) a modification increasing expression and / or activity of a Sugars Will Eventually be Exported Transporter (SWT) polypeptide, (ii) a modification increasing expression and / or activity of a sucrose transporter (SUT) polypeptide, (iii) a modification decreasing expression, activity, and / or stability of an endogenous Mother of Flowering Time (MFT) polypeptide, (iv) a modification increasing expression and / or activity of an MFT network gene, (v) a modification increasing expression and / or activity of an ABI3 polypeptide, (vi) a modification increasing expression and / or activity of an ODP1 polypeptide, (vii) a modification introducing a high oil DGAT variant, (viii) a modification decreasing expression, activity, and / or stability of an endogenous raffinose synthase (RS) polypeptide, or any combination thereof.
11. A soybean seed comprising a medium chain fatty acid content in an amount of at least 2% of total fatty acid content in the seed and a genomic modification, the genomic modification introduced into an endogenous gene encoding a fatty acyl-ACP thioesterase B (FATB) polypeptide.
12. The soybean seed of claim 11, wherein the genomic modification alters the substrate specificity of the encoded FATB polypeptide.
13. The soybean seed of claim 11 or 12, wherein the genomic modification introduces at least one deletion, insertion, or substitution in a coding sequence of the endogenous FATB gene toDocket # 212436-WO-SEC-1 produce a modified FATB polypeptide comprising at least one amino acid modification as compared to the wild-type FATB polypeptide sequence.
14. The soybean seed of claim 13, wherein the modified FATB polypeptide comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof.
15. The soybean seed of claim 13 or 14, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprisesDocket # 212436-WO-SEC-1 a. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2; b. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a tyrosine at the position corresponding to position 249 of SEQ ID NO: 2; c. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; d. comprises a cysteine at the position corresponding to position 167 of SEQ ID NO: 2 a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; e. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2, and a leucine at the position corresponding to position 270 of SEQ ID NO: 2; f. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; g. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 174 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 219 of SEQ ID NO: 2, a glycine at the position corresponding to position 228 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2;Docket # 212436-WO-SEC-1 h. a threonine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, a leucine at the position corresponding to position 174 of SEQ ID NO: 2, a leucine at the position corresponding to position 196 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 225 of SEQ ID NO: 2, an alanine at the position corresponding to position 228 of SEQ ID NO: 2, a serine at the position corresponding to position 229 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 259 of SEQ ID NO: 2, a leucine at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; i. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; or j. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO:
2.
16. The soybean seed of any one of claims 11-15, wherein the seed further comprises a modification to enhance seed oil content.
17. The soybean seed of claim 16, wherein the modification to enhance seed oil content comprises at least one of (i) a modification increasing expression and / or activity of a Sugars Will Eventually be Exported Transporter (SWT) polypeptide, (ii) a modification increasing expression and / or activity of a sucrose transporter (SUT) polypeptide, (iii) a modification decreasing expression, activity, and / or stability of an endogenous Mother of Flowering Time (MFT) polypeptide, (iv) a modification increasing expression and / or activity of an MFT network gene, (v) a modification increasing expression and / or activity of an ABI3 polypeptide, (vi) a modification increasing expression and / or activity of an ODP1 polypeptide, (vii) a modification introducing a high oil DGAT variant, (viii) a modificationDocket # 212436-WO-SEC-1 decreasing expression, activity, and / or stability of an endogenous raffinose synthase (RS) polypeptide, or any combination thereof.
18. A soybean plant which produces the soybean seed of any one of claims 1-17.
19. A polynucleotide encoding a modified FATB polypeptide comprising an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof, wherein expression of theDocket # 212436-WO-SEC-1 modified FATB polypeptide in a plant cell increase the medium chain fatty acid content as compared to a control plant cell.
20. The polynucleotide of claim 19, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprises a. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2; b. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a tyrosine at the position corresponding to position 249 of SEQ ID NO: 2; c. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; d. comprises a cysteine at the position corresponding to position 167 of SEQ ID NO: 2 a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; e. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2, and a leucine at the position corresponding to position 270 of SEQ ID NO: 2; f. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; g. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at theDocket # 212436-WO-SEC-1 position corresponding to position 174 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 219 of SEQ ID NO: 2, a glycine at the position corresponding to position 228 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; h. a threonine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, a leucine at the position corresponding to position 174 of SEQ ID NO: 2, a leucine at the position corresponding to position 196 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 225 of SEQ ID NO: 2, an alanine at the position corresponding to position 228 of SEQ ID NO: 2, a serine at the position corresponding to position 229 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 259 of SEQ ID NO: 2, a leucine at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; i. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; or j. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO:
2.
21. A plant cell comprising the polynucleotide of any one of claims 20-21.
22. The plant cell of claim 21, wherein the plant cell comprises a C12:0 medium chain fatty acid in an amount of at least 1% of total fatty acid content in the cell.
23. The plant cell of claim 21 or 22, wherein the plant further comprises a genome modification of a FATB4 gene, the genome modification increasing expression or activity of the encoded FATB4 polypeptide.Docket # 212436-WO-SEC-1 24. The plant cell of any one of claims 21-23, wherein the plant cell further comprises a modification to enhance oil content.
25. The plant cell of claim 24, wherein the modification to enhance oil content being in a gene encoding at least one of (i) a modification increasing expression and / or activity of a Sugars Will Eventually be Exported Transporter (SWT) polypeptide, (ii) a modification increasing expression and / or activity of a sucrose transporter (SUT) polypeptide, (iii) a modification decreasing expression, activity, and / or stability of an endogenous Mother of Flowering Time (MFT) polypeptide, (iv) a modification increasing expression and / or activity of an MFT network gene, (v) a modification increasing expression and / or activity of an ABI3 polypeptide, (vi) a modification increasing expression and / or activity of an ODP1 polypeptide, (vii) a modification introducing a high oil DGAT variant, (viii) a modification decreasing expression, activity, and / or stability of an endogenous raffinose synthase (RS) polypeptide, or any combination thereof.
26. A plant comprising the plant cell of any one of claims 21-25.
27. The plant of claim 26, wherein at least 2% of the total fatty acid content in a seed of the plant comprises a medium chain fatty acid.
28. A seed comprising the plant cell of any one of claims 21-25, wherein at least 2% of the total fatty acid content of the seed comprises a medium chain fatty acid.
29. A method of plant breeding, the method comprising crossing the soybean plant of any one of claims 18 and 26-27 with a second soybean plant to produce progeny seed.
30. A method of producing a soybean plant producing seeds having increased medium chain fatty acid content, the method comprising: a. introducing into a regenerable soybean plant cell a genomic modification in a fatty acyl-ACP thioesterase B (FATB) gene to encode a modified FATB polypeptide having altered substrate specificity as compared to a corresponding non-modified FATB polypeptide; and b. generating the plant, wherein the plant comprises the genomic modification and produces a seed having increased medium chain fatty acid content as compared to a seed of a control plant not comprising the genomic modification.Docket # 212436-WO-SEC-1 31. The method of claim 30, wherein the substrate specificity of the modified FATB polypeptide is altered to produce medium chain fatty acids (MCFAs).
32. The method of claim 30 or 31, wherein the seed of the generated plant comprises a C12:0 medium chain fatty acid in an amount of at least 1% of total fatty acid content in the cell.
33. The method of any one of claims 30-32, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprises a serine (S) at the position corresponding to position 119 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 150 of SEQ ID NO: 2, a methionine (M) at the position corresponding to position 156 of SEQ ID NO: 2, a cysteine (C), serine (S), or threonine (T) at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 170 of SEQ ID NO: 2, a methionine (M), a cysteine (C) or phenylalanine (F) at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine (C), isoleucine (I), or leucine (L) at the position corresponding to position 174 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 196 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 215 of SEQ ID NO: 2, a phenylalanine (F) or an isoleucine (I) at the position corresponding to position 219 of SEQ ID NO: 2, a lysine (K) or a glutamine (Q) at the position corresponding to position 221 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 225 of SEQ ID NO: 2, an isoleucine (I) at the position corresponding to position 227 of SEQ ID NO: 2, an alanine (A) or a glycine (G) at the position corresponding to position 228 of SEQ ID NO: 2, a serine (S) at the position corresponding to position 229 of SEQ ID NO: 2, a asparagine (N) at the position corresponding to position 235 of SEQ ID NO: 2, a tyrosine (Y) at the position corresponding to position 249 of SEQ ID NO: 2, a phenylalanine (F) or a tyrosine (Y) at the position corresponding to position 259 of SEQ ID NO: 2, a leucine (L) at the position corresponding to position 262 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine (F) or histidine (H) at the position corresponding to position 273 of SEQ ID NO: 2, a proline (P) at the position corresponding to position 278 of SEQ ID NO: 2, a asparagine (N) or serine (S) at the position corresponding to position 284 of SEQ ID NO: 2, a phenylalanine (F) at the position corresponding to position 291 of SEQ ID NO: 2, or any combination thereof.Docket # 212436-WO-SEC-1 34. The method of claim 33, wherein the modified FATB polypeptide comprises an amino acid sequence having at least 90% sequence identity to any of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, or 32 and comprises a. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2; b. a serine at the position corresponding to position 167 of SEQ ID NO: 2 and a tyrosine at the position corresponding to position 249 of SEQ ID NO: 2; c. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; d. comprises a cysteine at the position corresponding to position 167 of SEQ ID NO: 2 a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; e. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2, and a leucine at the position corresponding to position 270 of SEQ ID NO: 2; f. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, and an isoleucine at the position corresponding to position 227 of SEQ ID NO: 2; g. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 174 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 219 of SEQ ID NO: 2, a glycine at the positionDocket # 212436-WO-SEC-1 corresponding to position 228 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; or h. a threonine at the position corresponding to position 167 of SEQ ID NO: 2, a cysteine at the position corresponding to position 171 of SEQ ID NO: 2, a leucine at the position corresponding to position 174 of SEQ ID NO: 2, a leucine at the position corresponding to position 196 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2, an isoleucine at the position corresponding to position 225 of SEQ ID NO: 2, an alanine at the position corresponding to position 228 of SEQ ID NO: 2, a serine at the position corresponding to position 229 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 259 of SEQ ID NO: 2, a leucine at the position corresponding to position 270 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 273 of SEQ ID NO: 2, and a phenylalanine at the position corresponding to position 291 of SEQ ID NO: 2; i. a cysteine at the position corresponding to position 167 of SEQ ID NO: 2, a methionine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO: 2; or j. a serine at the position corresponding to position 167 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 171 of SEQ ID NO: 2, a cysteine at the position corresponding to position 174 of SEQ ID NO: 2, a phenylalanine at the position corresponding to position 219 of SEQ ID NO:
2.
35. The method of any one of claims 30-34, wherein the method further comprises introducing at least one additional genome modification to increase the expression or activity of the FATB polypeptide or decrease expression of a long-chain fatty acid FATB polypeptide.
36. The method of claim 35, wherein an endogenous long-chain fatty acid FATB type thioesterase is replaced with a medium-chain fatty acid FATB type thioesterase or variant thereofDocket # 212436-WO-SEC-1 37. The method of any one of claims 30-36, wherein the method further comprises introducing a modification to enhance seed oil content.
38. The method of claim 37, wherein the modification to enhance seed oil content being in a gene encoding at least one of (i) a modification increasing expression and / or activity of a Sugars Will Eventually be Exported Transporter (SWT) polypeptide, (ii) a modification increasing expression and / or activity of a sucrose transporter (SUT) polypeptide, (iii) a modification decreasing expression, activity, and / or stability of an endogenous Mother of Flowering Time (MFT) polypeptide, (iv) a modification increasing expression and / or activity of an MFT network gene, (v) a modification increasing expression and / or activity of an ABI3 polypeptide, (vi) a modification increasing expression and / or activity of an ODP1 polypeptide, (vii) a modification introducing a high oil DGAT variant, (viii) a modification decreasing expression, activity, and / or stability of an endogenous raffinose synthase (RS) polypeptide, or any combination thereof.
39. The method of any one of claims 30-38, wherein at least 2% of the total fatty acid in the seed of the generated plant comprises a medium chain fatty acid.
40. A seed oil composition produced from the seed of any one of claims 1-17 and 28.
41. A seed oil composition produced from the seeds of the plants of any one of claims 18 and 26- 27.
42. A method of producing an oil compositions, the method comprising crushing the seed of any one of claims 1-17 and 28 or seeds of the plants of any one of claims 18 and 26-27 and extracting oil from the crushed seed to form the oil composition.
Citation Information
Patent Citations
Variant thioesterases and methods of use
US20160032332A1
Soybean Lines with Low Saturated Fatty Acid and High Oleic Acid Contents
US20220380789A1
Cited By
Application and Creation Methods of Gene Expression Inhibitors in Improving Soybean Quality
CN122405730A