Engineered sesquiterpene synthases
Patent Information
- Application Number
- JP2024531183
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-02
- Filing Date
- 2022-11-18
- Publication Date
- 2025-10-31
AI Technical Summary
The purification of sesquiterpenes from natural sources and de novo chemical synthesis of guayene results in high production costs and low yields, and the structural complexity of guayene limits its extraction, making it difficult to obtain in high yields.
Engineered sesquiterpene synthases with specific amino acid substitutions, such as those at positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and/or 533, are introduced to increase the production of alpha-guayene relative to other sesquiterpene products.
The engineered sesquiterpene synthases significantly enhance alpha-guayene production, achieving higher yields compared to wild-type enzymes, addressing the limitations of natural extraction methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Application No. 63 / 283,194, entitled "ENGINEERED SESQUITERPENE SYNTHASES," filed on November 24, 2021, and U.S. Provisional Application No. 63 / 285,468, entitled "ENGINEERED SESQUITERPENE SYNTHASES," filed on December 2, 2021, the entire disclosures of each of which are incorporated herein by reference in their entireties.
[0002] Reference to Electronic Sequence Listing Contents of the Electronic Sequence Listing (G091970079WO00-SEQ-KVC.xml; size: 300,953 bytes; creation date: 202 15 Nov. 2002, which is incorporated herein by reference in its entirety.
[0003] FIELD OF THEINVENTION The present disclosure relates to sesquiterpene synthases, host cells containing the sesquiterpene synthases, and methods for making sesquiterpenes. [Background technology]
[0004] Terpenes are a diverse class of organic compounds constructed from five-carbon building blocks and encompass at least 400 distinct structural families. Being structurally diverse, terpenes have numerous roles, including acting as pheromones, antioxidants, and antimicrobial agents. Monoterpenes (C10), sesquiterpenes (C15), diterpenes (C20), and triterpenes (C30) make up the majority of terpenes. Guaiene is a type of sesquiterpene formed from three isoprene units and has the molecular formula C 15 H 24Examples include compounds such as alpha-guayene, delta-guayene (alpha-bulnesene), aciphyllene, beta-guayene, gamma-guayene and gamma-gurjunene, all of which have a guaiane sesquiterpene skeleton. Guaiene has been used in numerous contexts, such as in the production of fragrances and flavorings. Guaiene can be extracted from plants, but these resources are limited as many of these plants are endangered species. Furthermore, the large variety of sesquiterpene isomers often prevents extraction in high yields from naturally occurring sources, while the structural complexity of guayene often limits de novo chemical synthesis. Summary of the Invention [Means for solving the problem]
[0005] An embodiment of the disclosure relates to a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, at least one of the amino acid substitutions being at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0006] Further aspects of the disclosure relate to a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, at least one of the amino acid substitutions being at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1; and at least one of the amino acid substitutions being at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0007] In some embodiments, the host cell is capable of producing sesquiterpene products, and at least 50% of the total sesquiterpene products produced by the host cell is alpha-guayene. In some embodiments, at least 15% of the total sesquiterpene products produced by the host cell is aciphyllene.
[0008] A further aspect of the disclosure is a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:1, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, wherein the two or more amino acid substitutions are at least one of 72, 122, 124, 153, 191, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 2 The host cell is at a position corresponding to positions 5, 274, 275, 289, 290, 291, 292, 293, 295, 301, 346, 368, 398, 404, 406, 407, 442, 480, 494, 507, 509, 512, 526 and / or 533, and the host cell is capable of producing a sesquiterpene product, and at least 50% of the total sesquiterpene product produced by the host cell is alpha-guayene. In some embodiments, the sesquiterpene synthase comprises 4, 5, 6, 7, 8, 9, or more than 9 amino acid substitutions compared to SEQ ID NO:1.
[0009] In some embodiments, the sesquiterpene synthase comprises an isoleucine (I) residue at a position corresponding to position 72 of SEQ ID NO:1; an asparagine (N) residue at a position corresponding to position 122 of SEQ ID NO:1; a serine (S) residue at a position corresponding to position 124 of SEQ ID NO:1; an S residue at a position corresponding to position 153 of SEQ ID NO:1; an N residue at a position corresponding to position 191 of SEQ ID NO:1; an I residue at a position corresponding to position 201 of SEQ ID NO:1; a glutamic acid (E) residue at a position corresponding to position 205 of SEQ ID NO:1; an S residue at a position corresponding to position 274 of SEQ ID NO:1; a glycine (G) residue at a position corresponding to position 275 of SEQ ID NO:1; a histidine (H) residue at a position corresponding to position 289 of SEQ ID NO:1; a lysine (K) residue at a position corresponding to position 290 of SEQ ID NO:1; a valine (V) residue at a position corresponding to position 291 of SEQ ID NO:1; a phenylalanine (F) or methionine (M) residue at a position corresponding to position 292 of SEQ ID NO:1; a glutamine (Q), I, or N residue at a position corresponding to position 293 of SEQ ID NO:1. a V residue at a position corresponding to position 295 of SEQ ID NO:1; an S residue at a position corresponding to position 301 of SEQ ID NO:1; an E residue at a position corresponding to position 346 of SEQ ID NO:1; a cysteine (C) residue at a position corresponding to position 368 of SEQ ID NO:1; a C residue at a position corresponding to position 398 of SEQ ID NO:1; a tryptophan (W) residue at a position corresponding to position 404 of SEQ ID NO:1; a leucine (L) residue at a position corresponding to position 406 of SEQ ID NO:1; a G residue at a position corresponding to position 407 of SEQ ID NO:1; an L residue at a position corresponding to position 442 of SEQ ID NO:1; an I residue at a position corresponding to position 480 of SEQ ID NO:1; an E residue at a position corresponding to position 494 of SEQ ID NO:1; a tryptophan (W) residue at a position corresponding to position 507 of SEQ ID NO:1; an alanine (A) residue at a position corresponding to position 509 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; an F residue at a position corresponding to position 526 of SEQ ID NO:1; an N residue at a position corresponding to position 533 of SEQ ID NO:1; or any combination thereof.
[0010] In some embodiments, the sesquiterpene synthase contains an I residue at a position corresponding to position 72 of SEQ ID NO:1; an N residue at a position corresponding to position 122 of SEQ ID NO:1; an S residue at a position corresponding to position 124 of SEQ ID NO:1; an S residue at a position corresponding to position 153 of SEQ ID NO:1; an N residue at a position corresponding to position 191 of SEQ ID NO:1; an I residue at a position corresponding to position 201 of SEQ ID NO:1; an E residue at a position corresponding to position 205 of SEQ ID NO:1; an S residue; a G residue at a position corresponding to position 275 of SEQ ID NO:1; an L, T, S, H, M or D residue at a position corresponding to position 289 of SEQ ID NO:1; a K residue at a position corresponding to position 290 of SEQ ID NO:1; an F, L, T, V or C residue at a position corresponding to position 291 of SEQ ID NO:1; an A, Q, C, Y, H, E, F, M, W, T or F residue at a position corresponding to position 292 of SEQ ID NO:1; an L, V, T, Y, C, F, W, Q, I, N or M residue at a position corresponding to position 293 of SEQ ID NO:1; an E, D, N, W, G, V or I residue at a position corresponding to position 295 of SEQ ID NO:1; an S residue at a position corresponding to position 301 of SEQ ID NO:1; an E residue at a position corresponding to position 346 of SEQ ID NO:1; a C residue at a position corresponding to position 368 of SEQ ID NO:1; a C residue at a position corresponding to position 398 of SEQ ID NO:1; a W residue at a position corresponding to position 404 of SEQ ID NO:1; an L, N, W or T residue at a position corresponding to position 406 of SEQ ID NO:1; a G residue at a position corresponding to position 407 of SEQ ID NO:1; an L residue at a position corresponding to position 442 of SEQ ID NO:1; an I residue at a position corresponding to position 480 of SEQ ID NO:1; an E residue at a position corresponding to position 494 of SEQ ID NO:1; a W residue at a position corresponding to position 507 of SEQ ID NO:1; an A residue at a position corresponding to position 509 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; an F residue at a position corresponding to position 526 of SEQ ID NO:1; an N residue at a position corresponding to position 533 of SEQ ID NO:1; or any combination thereof.
[0011] In some embodiments, the sesquiterpene synthase has the following amino acid substitutions relative to SEQ ID NO:1: N289H, V292F, G293N, T295V, K404W, F406L, and F512L; N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L; N289H, F406L, F512L, D191N, and A205E; I291V, V292F, G293Q, T295V, K404W, F406L, I507W, and F512L; N289H, V292F, T295V ,K404W,F406L,and F512L;N289H,I291V,V292F,G293N,T295V,K404W,F406L,and F512L;V292F,G293Q,K404W,and F406L;N289H,V292F,G293I,T295V,F406L,I407G,and F512L;V292F,G293Q,K404W,F406L,I407G,and F512L;N289H,I291V,V292F,G293I,T295V,F406L,I407G,and F512L;V292F,T 295V, K404W, F406L, I507W, and F512L;N289H, V292M, G293I, F406L, F512L, and M480I;N289H, I291V, V292F, G293Q, T295V, K404W, and F406L;V292F, G293Q, T295V, F406L, and F512L;V292F, T295V, K404W, F406L, and F512L;I291V, V292F, G293I, K404W, F406L, and I507W;V292F, G293I, T295V, F40 6L, I407G, and F512L;N289H, V292F, G293I, T295V, F406L, I507W, and F512L;N289H, I291V, V292F, G293I, T295V, K404W, F406L, and F512L;N289H, V292F, G293Q, T295V, F406L, I407G, and F512L;N289H, I291V, V292F, G293N, T295V, K404W, and F406L;N289H, V292F, G293Q, F406L, I507W, and F512L;I291V, V292F, G293N, T295V, K404W, F406L, I507W, and F512L;N289H, V292F, G293N, K404W, F406L, I507W, and F512L;N289H, V292F, G293I, T295V, F406L, I407G, I507W, and F512L;N289H, I291V, V292F, G293N, T295V, F406L, and F512L;I291V, V292F, G293Q, K404W, F406L, I407G, and F512L;I291V, V292F ,G293Q,T295V,K404W,F406L,I407G,and F512L;N289H,V292F,T295V,F406L,and F512L;N289H,V292F,K404W,F406L,and F512L;V292F,G293N,T295V,F406L,and F512L;N289H,I291V,V292F,G293N,T295V,F406L,I407G,I507W,and F512L;N289H,I291V,V292F,G293I,T295V,F406L,and F512L;N289H,I 291V, V292F, G293Q, T295V, F406L, and I507W;N289H, V292F, G293I, F406L, I407G, I507W, and F512L;N289H, V292F, G293Q, T295V, K404W, F406L, and I507W;I291V, V292F, G293Q, T295V, F406L, I507W, and F512L;S122N, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L;P124S, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L;T153S, N289H, I291V, V292F, G293Q, T295V, D346E, K404W, F406L, and F512L;P124S, N289H, I291V, V292F, G293Q, T295V, K404W, and F406L;R290K, V292F, T295V, K404W, F406L, and F512L;N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L;N289H, V292F, G293N, T295V, K404W, F406L, Y442L, and F512L;T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, T494E, and F512L;T72I, N289H, V292F, T295V, K404W, F406L, T494E, and F512L;T72I, N289H, V292F, G293I, T295V, F406L, I407G, Y442L, I507W, and F512 L; P124S, V292F, T295V, K404W, F406L, I507W, and F512L; N289H, V292F, T295V, K404W, F406L, Y442L, and F512L; N289H, F406L, F512L, and T295V; N289H, F406L, F512L, T295V, and G274S; N289H, F406L, F512L, T295V, and V201I; or N289H, F406L, F512L, T295V, and I398C.;
[0012] In some embodiments, the sesquiterpene synthase comprises a sequence at least 90% identical to any one of SEQ ID NOs: 3-39 or 77-92.
[0013] In some embodiments, the sesquiterpene synthase comprising the amino acid substitutions N289H, V292F, G293N, T295V, K404W, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:3; the sesquiterpene synthase comprising the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:4; the sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, D191N, and A20 A sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293Q, T295V, K404W, F406L, I507W, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:6; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, K404W, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:7; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, The sesquiterpene synthase comprising G293N, T295V, K404W, F406L, and F512L comprises the sequence of SEQ ID NO:8; the sesquiterpene synthase comprising, relative to SEQ ID NO:1, the amino acid substitutions V292F, G293Q, K404W, and F406L comprises the sequence of SEQ ID NO:9; the sesquiterpene synthase comprising, relative to SEQ ID NO:1, the amino acid substitutions N289H, V292F, G293I, T295V, F406L, I407G, and F512L comprises the sequence of SEQ ID NO:10; the sesquiterpene synthase comprising, relative to SEQ ID NO:1, the amino acid substitutions V292F , G293Q, K404W, F406L, I407G, and F512L, comprises the sequence of SEQ ID NO:11; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293I, T295V, F406L, I407G, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:12; a sesquiterpene synthase that includes the amino acid substitutions V292F, T295V, K404W, F406L, I507W, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:13;A sesquiterpene synthase comprising the amino acid substitutions N289H, V292M, G293I, F406L, F512L, and M480I, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:14; a sesquiterpene synthase comprising the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, K404W, and F406L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:15; a sesquiterpene synthase comprising the amino acid substitutions V292F, G293Q, T295V, F406L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:16. comprises the sequence of SEQ ID NO:16; a sesquiterpene synthase that comprises the amino acid substitutions V292F, T295V, K404W, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:17; a sesquiterpene synthase that comprises the amino acid substitutions I291V, V292F, G293I, K404W, F406L, and I507W compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:18; a sesquiterpene synthase that comprises the amino acid substitutions V292F, G293I, T295V, F406L, I407G, and F512L compared to SEQ ID NO:1 The terpene synthase comprises the sequence of SEQ ID NO:19; the sesquiterpene synthase comprises the amino acid substitutions N289H, V292F, G293I, T295V, F406L, I507W, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:20; the sesquiterpene synthase comprises the amino acid substitutions N289H, I291V, V292F, G293I, T295V, K404W, F406L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:21; the sesquiterpene synthase comprises the amino acid substitutions N289H, V292F, G293I, T295V, K404W, F406L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:22; , G293Q, T295V, F406L, I407G, and F512L, comprises the sequence of SEQ ID NO:22; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293N, T295V, K404W, and F406L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:23; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293Q, F406L, I507W, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:24;A sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293N, T295V, K404W, F406L, I507W, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:25; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293N, K404W, F406L, I507W, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:26; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293I, T295V, F406L, I407G, I507W, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:27. and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:27; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293N, T295V, F406L, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:28; a sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293Q, K404W, F406L, I407G, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:29 .... A sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, F406L, I407G, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:30; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:31; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, K404W, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:32; a sesquiterpene synthase that includes the amino acid substitutions V292F, G295V, I407G, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:33. a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293N, T295V, F406L, I407G, I507W, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:34; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293I, T295V, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:35;A sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, F406L, and I507W, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:36; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293I, F406L, I407G, I507W, and F512L, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:37; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293Q, T295V, K404W, F406L, and I507W, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:38; A sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293Q, T295V, F406L, I507W, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:39; a sesquiterpene synthase that includes the amino acid substitutions S122N, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:84; A sesquiterpene synthase comprising the amino acid substitutions P124S, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L comprises the sequence of SEQ ID NO: 80; a sesquiterpene synthase comprising the amino acid substitutions T153S, N289H, I291V, V292F, G293Q, T295V, D346E, K404W, F406L, and F512L compared to SEQ ID NO: 1 comprises the sequence of SEQ ID NO: 82; a sesquiterpene synthase that includes the amino acid substitutions R290K, V292F, T295V, K404W, F406L, and F512L, relative to SEQ ID NO:1, includes the sequence of SEQ ID NO:83; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L, relative to SEQ ID NO:1, includes the sequence of SEQ ID NO:78;A sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293N, T295V, K404W, F406L, Y442L, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:77; a sesquiterpene synthase that includes the amino acid substitutions T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, T494E, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:79; , N289H, V292F, T295V, K404W, F406L, T494E, and F512L, comprises the sequence of SEQ ID NO: 88; sesquiterpene synthases that include the amino acid substitutions T72I, N289H, V292F, G293I, T295V, F406L, I407G, Y442L, I507W, and F512L, compared to SEQ ID NO: 1, comprise the sequence of SEQ ID NO: 86; and sesquiterpene synthases that include the amino acid substitutions P124S, V292F, T295V, K404W, F406L, T494E, and F512L, compared to SEQ ID NO: 1, comprise the sequence of SEQ ID NO: 87. a sesquiterpene synthase that includes, relative to SEQ ID NO:1, the amino acid substitutions N289H, V292F, T295V, K404W, F406L, Y442L, and F512L comprises the sequence of SEQ ID NO:85; a sesquiterpene synthase that includes, relative to SEQ ID NO:1, the amino acid substitutions N289H, F406L, F512L, and T295V comprises the sequence of SEQ ID NO:87; a sesquiterpene synthase that includes, relative to SEQ ID NO:1, the amino acid substitutions N289H, F406L, F512L, and T295V comprises the sequence of SEQ ID NO:89; A sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, T295V, and G274S comprises the sequence of SEQ ID NO:90; a sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, T295V, and V201I, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:91; or a sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, T295V, and I398C, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:92.
[0014] In some embodiments, the heterologous nucleic acid comprises a sequence that is at least 90% identical to any one of SEQ ID NOs: 40-76 or 93-109.
[0015] In some embodiments, the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to 294, 296, 297, 403, 444, 515, and / or 525 in SEQ ID NO:1.
[0016] In some embodiments, the sesquiterpene synthase comprises an L, C, Y, V, A, or S residue at a position corresponding to position 294 of SEQ ID NO:1; an L, A, proline (P), Y, N, F, or R residue at a position corresponding to position 296 of SEQ ID NO:1; an E, Y, I, lysine (K), M, or H residue at a position corresponding to position 297 of SEQ ID NO:1; an M, Q, N, S, T, A, E, H, C, or V residue at a position corresponding to position 403 of SEQ ID NO:1; an A or N residue at a position corresponding to position 444 of SEQ ID NO:1; an H, A, E, or Q residue at a position corresponding to position 515 of SEQ ID NO:1; an H, C, L, or N residue at a position corresponding to position 525 of SEQ ID NO:1; or any combination thereof.
[0017] In some embodiments, the host cell is capable of producing more alpha-guayene than a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO: 1. In some embodiments, the host cell produces more alpha-guayene than delta-guayene.
[0018] In some embodiments, the sesquiterpene synthase comprises: (a) one or more amino acid substitutions in a first region of the active site compared to SEQ ID NO:1, the one or more amino acid substitutions in the first region are at positions corresponding to positions 295, 291, 406, 512, and / or 519 of SEQ ID NO:1, and the one or more amino acid substitutions in the first region comprise a residue having a smaller side chain than the side chain of the amino acid at positions 295, 291, 406, 512, and / or 519 of SEQ ID NO:1, respectively; (b) an amino acid substitution in a second region of the active site compared to SEQ ID NO:1, the amino acid substitution in the second region corresponds to position 292 of SEQ ID NO:1, and the amino acid substitution in the second region comprises a larger side chain than the side chain of the amino acid at position 292 of SEQ ID NO:1; or (c) any combination thereof. In some embodiments, the side chain in (a) is hydrophobic.
[0019] In some embodiments, the sesquiterpene synthase comprises a V residue at a position corresponding to position 295 of SEQ ID NO:1; a C residue at a position corresponding to position 291 of SEQ ID NO:1; an L residue at a position corresponding to position 406 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; an F residue at a position corresponding to position 292 of SEQ ID NO:1; or any combination thereof.
[0020] In some embodiments, the sesquiterpene synthase further comprises one or more amino acid substitutions relative to SEQ ID NO:1 at one or more positions corresponding to 294, 296, 297, 403, 444, 515, and / or 525 in SEQ ID NO:1.
[0021] In some embodiments, the sesquiterpene synthase has an L, T, S, D, or M residue at a position corresponding to position 289 of SEQ ID NO:1; an F, L, T, or C residue at a position corresponding to position 291 of SEQ ID NO:1; an A, Q, C, Y, H, E, F, T, or W residue at a position corresponding to position 292 of SEQ ID NO:1; an L, V, T, Y, C, F, W, or M residue at a position corresponding to position 293 of SEQ ID NO:1; an L, C, Y, V, A, or S residue at a position corresponding to position 294 of SEQ ID NO:1; an E, D, N, W, G, or I residue at a position corresponding to position 295 of SEQ ID NO:1; an L, A, P, Y, N, arginine (R), or F residue at the corresponding position; an E, Y, I, K, M, or H residue at a position corresponding to position 297 of SEQ ID NO:1; an M, Q, N, V, C, H, E, A, T, or S residue at a position corresponding to position 403 of SEQ ID NO:1; an L, T, W, or N residue at a position corresponding to position 406 of SEQ ID NO:1; an A or N residue at a position corresponding to position 444 of SEQ ID NO:1; an H, A, E, or Q residue at a position corresponding to position 515 of SEQ ID NO:1; an H, C, L, or N residue at a position corresponding to position 525 of SEQ ID NO:1; or any combination thereof.
[0022] In some embodiments, the sesquiterpene synthase contains a Q, C, V, F, A, I, H, G, W, or Y residue at a position corresponding to position 289 of SEQ ID NO:1; a V, M, or A residue at a position corresponding to position 291 of SEQ ID NO:1; an I, M, S, L, G, or N residue at a position corresponding to position 292 of SEQ ID NO:1; an I, Q, S, N, or E residue at a position corresponding to position 293 of SEQ ID NO:1; an L, V, A, or S residue at a position corresponding to position 295 of SEQ ID NO:1; a K or Q residue at a position corresponding to position 6; a T or A residue at a position corresponding to position 297 of SEQ ID NO:1; a P, F, I, L, G, or D residue at a position corresponding to position 403 of SEQ ID NO:1; a G, S, A, I, Y, M, H, V, Q, or C residue at a position corresponding to position 406 of SEQ ID NO:1; an H residue at a position corresponding to position 444 of SEQ ID NO:1; an M, C, F, G, N, S, or I residue at a position corresponding to position 515 of SEQ ID NO:1; or any combination thereof.
[0023] Further aspects of the disclosure relate to non-naturally occurring sesquiterpene synthases, wherein the sequence of the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, at least one of the amino acid substitutions being at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0024] Further embodiments of the disclosure relate to non-naturally occurring sesquiterpene synthases, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, with at least one amino acid substitution at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1; and at least one amino acid substitution at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0025] A further embodiment of the disclosure is a non-naturally occurring sesquiterpene synthase, wherein the sesquiterpene synthase comprises an amino acid sequence at least 90% identical to the sequence of SEQ ID NO:1, wherein the sesquiterpene synthase comprises an amino acid sequence having two or more amino acid substitutions compared to SEQ ID NO:1, wherein the two or more amino acid substitutions are at least one of 72, 122, 124, 153, 191, 201, 205, 274, 275, 289, 290, 291, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 37 and wherein the sesquiterpene synthase is capable of producing a sesquiterpene product, and at least 50% of the sesquiterpene product produced by the sesquiterpene synthase is alpha-guayene.
[0026] In some embodiments, the amino acid sequence of the sesquiterpene synthase is at least 90% identical to any one of SEQ ID NOs: 3-39 or 77-92.
[0027] A further aspect of the present disclosure relates to a non-naturally occurring nucleic acid encoding a sesquiterpene synthase, wherein the non-naturally occurring nucleic acid comprises a sequence at least 90% identical to SEQ ID NOs: 40-76 or 93-109.
[0028] A further aspect of the present disclosure relates to a method for producing a sesquiterpene, comprising culturing a host cell associated with the present disclosure in a culture medium in the presence of an FPP substrate, and optionally isolating or recovering the sesquiterpene from the host cell and / or the culture medium. In some embodiments, the sesquiterpene is alpha-guayene. In some embodiments, the sesquiterpene is aciphyllene.
[0029] In some embodiments, at least 50% of the sesquiterpenes isolated or recovered from the host cell or culture medium are alpha-guayene. In some embodiments, at least 15% of the sesquiterpenes isolated or recovered from the host cell or culture medium are acifylene. In some embodiments, alpha-guayene is recovered from the culture medium. In some embodiments, acifylene is recovered from the culture medium. In some embodiments, the method further comprises obtaining a composition comprising alpha-guayene. In some embodiments, the method further comprises obtaining a composition comprising acifylene.
[0030] Further aspects of the present disclosure relate to sesquiterpenes obtainable from the host cells and / or methods related to the present disclosure. In some embodiments, the sesquiterpene is alpha-guayene. In some embodiments, the sesquiterpene is aciphyllene.
[0031] A further aspect of the present disclosure relates to a culture medium comprising the sesquiterpenes associated with the present disclosure.
[0032] Further aspects of the present disclosure relate to compositions comprising sesquiterpenes related to the present disclosure. In some embodiments, the sesquiterpene is alpha-guayene. In some embodiments, the sesquiterpene is aciphyllene.
[0033] Further aspects of the present disclosure relate to compositions comprising (a) sesquiterpenes, wherein at least 50% of the sesquiterpenes are alpha-guayene, and (b) one or more additional components comprising a fermentation medium, a cell culture supernatant, and / or a hydrophobic overlay. In some embodiments, at least 15% of the sesquiterpenes are acifylene. In some embodiments, the alpha-guayene is produced using a microbial host cell. In some embodiments, the alpha-guayene is produced using an in-vitro or in-vivo system. In some embodiments, about 50% to about 90% of the sesquiterpenes are alpha-guayene. In some embodiments, the composition further comprises delta-guayene. In some embodiments, the composition further comprises acifylene. In some embodiments, about 50% to about 10% of the sesquiterpenes are alpha-guayene.
[0034] In some embodiments, the composition further comprises one or more non-terpene components or one or more additional terpene components.In some embodiments, the one or more non-terpene components comprise FPP.In some embodiments, the one or more additional terpene components comprise delta-guaiene, beta-guaiene, gamma-guaiene, germacrene A, aciphyllene, and / or alpha-humulene, as determined by GC.
[0035] A further aspect of the present disclosure relates to a method of making a composition related to the present disclosure, comprising culturing a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, and at least one of the amino acid substitutions is at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0036] Further aspects of the present disclosure relate to methods of making a composition associated with the present disclosure, comprising culturing a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, wherein at least one amino acid substitution is at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1; and at least one amino acid substitution is at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0037] Further aspects of the present disclosure relate to methods of making a composition associated with the present disclosure comprising culturing a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises an amino acid sequence at least 90% identical to the sequence of SEQ ID NO:1, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions relative to SEQ ID NO:1, wherein the two or more amino acid substitutions are at positions corresponding to positions 72, 122, 124, 153, 191, 201, 205, 274, 275, 289, 290, 291, 292, 293, 295, 301, 346, 368, 398, 404, 406, 407, 442, 480, 494, 507, 509, 512, 526 and / or 533 in SEQ ID NO:1.
[0038] A further aspect of the present disclosure relates to a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, and at least one of the amino acid substitutions is at a position corresponding to positions 44, 212, 217, 293, 295, 404, 515, and / or 542 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:1.
[0039] Further aspects of the disclosure relate to a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, at least one of the amino acid substitutions being at a position corresponding to positions 44, 212, 217, 293, 295, 404, 515, and / or 542 in SEQ ID NO:1; and at least one of the amino acid substitutions being at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:1. In some embodiments, the host cell is capable of producing sesquiterpene products, and at least 50% of the total sesquiterpene products produced by the host cell is alpha-guayene. In some embodiments, at least 15% of the total sesquiterpene products produced by the host cell is aciphyllene.
[0040] A further aspect of the disclosure is a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises an amino acid sequence at least 90% identical to the sequence of SEQ ID NO:1, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, wherein the two or more amino acid substitutions are at least one of 23, 44, 72, 86, 111, 118, 134, 147, 188, 201, 212, 217, 220, 222, 224, 226, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 260, 261, 262, 263, 264, 265, 266, 267, 268, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 290, 300, 301, 302, 303, 304, 305, 3 , 224, 252, 255, 289, 290, 291, 292, 293, 295, 346, 381, 390, 404, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 515, 516, 519, 542, wherein the host cell is capable of producing a sesquiterpene product, and at least 50% of the total sesquiterpene products produced by the host cell is alpha-guayene. In some embodiments, the sesquiterpene synthase comprises 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or more than 27 amino acid substitutions compared to SEQ ID NO:1.
[0041] In some embodiments, the sesquiterpene synthase comprises an aspartic acid (D) residue at a position corresponding to position 23 of SEQ ID NO:1; a valine (V) residue at a position corresponding to position 44 of SEQ ID NO:1; an isoleucine (I) residue at a position corresponding to position 72 of SEQ ID NO:1; a glutamic acid (E) residue at a position corresponding to position 86 of SEQ ID NO:1; a leucine (L) residue at a position corresponding to position 111 of SEQ ID NO:1; a glutamine (Q) residue at a position corresponding to position 118 of SEQ ID NO:1; a D residue at a position corresponding to position 134 of SEQ ID NO:1; a V residue at a position corresponding to position 147 of SEQ ID NO:1; a Q residue at a position corresponding to position 188 of SEQ ID NO:1; a serine (S) residue at a position corresponding to position 201 of SEQ ID NO:1; a phenylalanine (F) residue at a position corresponding to position 212 of SEQ ID NO:1; an E residue at a position corresponding to position 217 of SEQ ID NO:1; an L residue at a position corresponding to position 224 of SEQ ID NO:1; an L residue at a position corresponding to position 252 of SEQ ID NO:1; an alanine (A) residue at a position corresponding to position 255 of SEQ ID NO:1; a lysine (K) residue at a position corresponding to position 290 of SEQ ID NO:1; a histidine (H) residue at a position corresponding to position 290 of SEQ ID NO:1; a V residue at a position corresponding to position 291 of SEQ ID NO:1; a cysteine (C) residue at a position corresponding to position 292 of SEQ ID NO:1; an F residue at a position corresponding to position 292 of SEQ ID NO:1; an asparagine (N) residue at a position corresponding to position 293 of SEQ ID NO:1; a Q residue at a position corresponding to position 293 of SEQ ID NO:1; an I residue at a position corresponding to position 295 of SEQ ID NO:1; a V residue at a position corresponding to position 295 of SEQ ID NO:1; an E residue at a position corresponding to position 381 of SEQ ID NO:1; a tryptophan (W) residue at a position corresponding to position 390 of SEQ ID NO:1; a W residue at a position corresponding to position 404 of SEQ ID NO:1; an L residue at a position corresponding to position 406 of SEQ ID NO:1; an S residue at a position corresponding to position 419 of SEQ ID NO:1; an I residue at a position corresponding to position 433 of SEQ ID NO:1; an L ...42 of SEQ ID NO:1; a W residue at a position corresponding to position 443 of SEQ ID NO:1;a threonine (T) residue at a position corresponding to position 443 of SEQ ID NO:1; an N residue at a position corresponding to position 444 of SEQ ID NO:1; a G residue at a position corresponding to position 448 of SEQ ID NO:1; an L residue at a position corresponding to position 458 of SEQ ID NO:1; a V residue at a position corresponding to position 458 of SEQ ID NO:1; a K residue at a position corresponding to position 467 of SEQ ID NO:1; an E residue at a position corresponding to position 494 of SEQ ID NO:1; an N residue at a position corresponding to position 499 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; a methionine (M) residue at a position corresponding to position 515 of SEQ ID NO:1; an A residue at a position corresponding to position 516 of SEQ ID NO:1; a V residue at a position corresponding to position 516 of SEQ ID NO:1; an L residue at a position corresponding to position 519 of SEQ ID NO:1; a V residue at a position corresponding to position 542 of SEQ ID NO:1; or any combination thereof.
[0042] In some embodiments, the sesquiterpene synthase has the following amino acid substitutions relative to SEQ ID NO:1: T72I, N289H, R290K, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, T494E, and F512L; L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L; L44V, T72I, T86E, G118Q, T134D, W147L, Q212F, Q217E, S224L, F 252L, P255A, N289H, R290K, I291V, V292F, G293Q, T295V, D346E, Y390F, K404W, F406L, L419S, Y442L, T458V, R467K, F512L, and R542V;T23D, L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L;V201S, N289H, I291V, V292F, G293N, T295V, Y381W, K404W, F406L, T49 4E, and F512L;V201S, V292C, T295I, and D444N;V201S, V292C, T295I, D444N, and L516V;L44V, T72I, V201S, V292C, T295I, D444N, and L516V;L44V, T72I, V201S, V292C, T295I, D444N, and L516A;L44V, T72I, V201S, V292C, T295I, I443W, and D444N;L44V, T72I, V201S, V292C, T295I, D444N, and M 519L;W147V, C188Q, V201S, S224L, V292C, T295I, and D444N;L44V, T72I, V111L, V201S, S224L, V292C, T295I, I443T, D444N, T448G, and M519L;L44V, T72I, V111L, V201S, V292C, T295I, V433I, I443T, D444N, T448G, D499N, and L516V;L44V, T72I, V201S, V292C, T295I, D444N, L516V, and M519L;L44V, T72I, V201S, V292C, T295I, D444N, V515M, and L516V;V111L, V201S, S224L, V292C, T295I, I443T, D444N, and L516V;V201S, V292C, T295I, D444N, and L516V;V201S, V292C, T2 95I, D444N, T458L, and L516A; V201S, V292C, T295I, I443T, D444N, and L516V; V201S, V292C, T295I, V433I, D444N, and L516V; or V201S, V292C, T295I, V433L, D444N, and L516V.
[0043] In some embodiments, the sesquiterpene synthase comprises a sequence at least 90% identical to any one of SEQ ID NOs:110-131.
[0044] In some embodiments, the sesquiterpene synthase comprising the amino acid substitutions T72I, N289H, R290K, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, T494E, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:110; the sesquiterpene synthase comprising the amino acid substitutions L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:110. A sesquiterpene synthase comprising the sequence of SEQ ID NO:111; and comprising the amino acid substitutions L44V, T72I, T86E, G118Q, T134D, W147L, Q212F, Q217E, S224L, F252L, P255A, N289H, R290K, I291V, V292F, G293Q, T295V, D346E, Y390F, K404W, F406L, L419S, Y442L, T458V, R467K, F512L, and R542V compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:112; and A sesquiterpene synthase comprising the amino acid substitutions T23D, L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L comprises the sequence of SEQ ID NO:113; compared to SEQ ID NO:1. A sesquiterpene synthase comprising the amino acid substitutions V201S, N289H, I291V, V292F, G293N, T295V, Y381W, K404W, F406L, T494E, and F512L comprises the sequence of SEQ ID NO:114; compared to SEQ ID NO:1. , a sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, and D444N, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:115; a sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, D444N, and L516V, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:116; a sesquiterpene synthase comprising the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and L516V, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:117;A sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and L516A, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:118; a sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, I443W, and D444N, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:119; a sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and M519L, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:119. The sesquiterpene synthase comprises the sequence of SEQ ID NO:120; the sesquiterpene synthase comprises the amino acid substitutions W147V, C188Q, V201S, S224L, V292C, T295I, and D444N, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:121; the sesquiterpene synthase comprises the amino acid substitutions L44V, T72I, V111L, V201S, S224L, V292C, T295I, I443T, D444N, T448G, and M519L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:122; In all, the sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V111L, V201S, V292C, T295I, V433I, I443T, D444N, T448G, D499N, and L516V comprises the sequence of SEQ ID NO:123; the sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, L516V, and M519L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:124; A sesquiterpene synthase that includes the amino acid substitutions V111L, V201S, S224L, V292C, T295I, I443T, D444N, and L516V, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:126; a sesquiterpene synthase that includes the amino acid substitutions V201S, V292C, T295I, D444N, and L516V, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:127;A sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, D444N, T458L, and L516A, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:128; a sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, I443T, D444N, and L516V, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:129; a sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, V433I, D444N, and L516V, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:130; or a sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, V433L, D444N, and L516V, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:131.
[0045] In some embodiments, the heterologous nucleic acid comprises a sequence that is at least 90% identical to any one of SEQ ID NOs:133-154.
[0046] In some embodiments, the host cell is capable of producing more alpha-guayene than a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO: 1. In some embodiments, the host cell produces more alpha-guayene than delta-guayene. In some embodiments, the host cell produces more aciphyllene than delta-guayene.
[0047] A further embodiment of the present disclosure relates to a non-naturally occurring sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, and at least one of the amino acid substitutions is at a position corresponding to positions 44, 212, 217, 293, 295, 404, 515, and / or 542 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:1.
[0048] A further embodiment of the disclosure is a non-naturally occurring sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, with at least one amino acid substitution at a position corresponding to positions 44, 212, 217, 293, 295, 404, 515, and / or 542 in SEQ ID NO:1; No. 1, wherein the non-naturally occurring sesquiterpene synthase is located at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in the present invention.
[0049] A further embodiment of the disclosure is a non-naturally occurring sesquiterpene synthase, wherein the sesquiterpene synthase comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:1, wherein the sesquiterpene synthase comprises an amino acid sequence having two or more amino acid substitutions compared to SEQ ID NO:1, wherein the two or more amino acid substitutions are at least one of 23, 44, 72, 86, 111, 118, 134, 147, 188, 201, 212, 217, 224, 252, 255, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355 The present invention relates to a non-naturally occurring sesquiterpene synthase wherein the sesquiterpene synthase is at a position corresponding to positions 291, 292, 293, 295, 346, 381, 390, 404, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 515, 516, 519, or 542, wherein the sesquiterpene synthase is capable of producing a sesquiterpene product, and at least 50% of the sesquiterpene product produced by the sesquiterpene synthase is alpha-guayene.
[0050] In some embodiments, the amino acid sequence of the sesquiterpene synthase is at least 90% identical to any one of SEQ ID NOs:110-131.
[0051] A further aspect of the present disclosure relates to a non-naturally occurring nucleic acid encoding a sesquiterpene synthase, the non-naturally occurring nucleic acid comprising a sequence at least 90% identical to any one of SEQ ID NOs: 133-154.
[0052] A further aspect of the present disclosure relates to a method for producing a sesquiterpene, comprising culturing a host cell associated with the present disclosure in a culture medium in the presence of an FPP substrate, and optionally isolating or recovering the sesquiterpene from the host cell and / or the culture medium. In some embodiments, the sesquiterpene is alpha-guayene. In some embodiments, the sesquiterpene is acifylene. In some embodiments, 50% of the sesquiterpenes isolated or recovered from the host cell or culture medium are alpha-guayene. In some embodiments, at least 15% of the sesquiterpenes isolated or recovered from the host cell or culture medium are acifylene. In some embodiments, alpha-guayene is recovered from the culture medium. In some embodiments, acifylene is recovered from the culture medium. In some embodiments, the method further comprises obtaining a composition comprising alpha-guayene. In some embodiments, the method further comprises obtaining a composition comprising acifylene.
[0053] Further aspects of the present disclosure relate to sesquiterpenes obtainable from the host cells or methods related to the present disclosure. In some embodiments, the sesquiterpene is alpha-guayene. In some embodiments, the sesquiterpene is aciphyllene.
[0054] A further aspect of the present disclosure relates to a culture medium comprising the sesquiterpenes associated with the present disclosure.
[0055] Further aspects of the present disclosure relate to compositions comprising sesquiterpenes related to the present disclosure. In some embodiments, the sesquiterpene is alpha-guayene. In some embodiments, the sesquiterpene is aciphyllene.
[0056] Further aspects of the present disclosure relate to compositions comprising (a) sesquiterpenes, wherein at least 50% of the sesquiterpenes are alpha-guayene, and (b) one or more additional components comprising a fermentation medium, a cell culture supernatant, and / or a hydrophobic overlay. In some embodiments, at least 15% of the sesquiterpenes are acifylene. In some embodiments, the alpha-guayene is produced using a microbial host cell. In some embodiments, the alpha-guayene is produced using an in vitro or in vivo system. In some embodiments, about 50% to about 90% of the sesquiterpenes are alpha-guayene. In some embodiments, the composition further comprises delta-guayene. In some embodiments, the composition further comprises acifylene. In some embodiments, about 50% to about 10% of the sesquiterpenes are alpha-guayene. In some embodiments, at least 15% of the sesquiterpenes are acifylene.
[0057] In some embodiments, the composition further comprises one or more non-terpene components or one or more additional terpene components.In some embodiments, the one or more non-terpene components comprise FPP.In some embodiments, the one or more additional terpene components comprise delta-guaiene, beta-guaiene, gamma-guaiene, germacrene A, aciphyllene, and / or alpha-humulene, as determined by GC.
[0058] Further aspects of the present disclosure relate to methods of making a composition associated with the present disclosure, comprising culturing a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, wherein at least one of the amino acid substitutions is at a position corresponding to positions 44, 212, 217, 293, 295, 404, 448, 515, and / or 542 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:1.
[0059] A further aspect of the disclosure is a method of making a composition related to the disclosure, comprising culturing a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, wherein at least one amino acid substitution is at 44, 212, 217, 293, 295, 404, 515, and / or 542 in SEQ ID NO:1. at least one amino acid substitution is at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:1.
[0060] A further aspect of the present disclosure is a method of making a composition related to the present disclosure, comprising culturing a host cell comprising a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises an amino acid sequence at least 90% identical to the sequence of SEQ ID NO:1, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO:1, and wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions. is at a position corresponding to positions 23, 44, 72, 86, 111, 118, 134, 147, 188, 201, 212, 217, 224, 252, 255, 289, 290, 291, 292, 293, 295, 346, 381, 390, 404, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 515, 516, 519, or 542 in SEQ ID NO:1.
[0061] Each of the composition and method limitations described in this disclosure may encompass the various described embodiments. It is therefore anticipated that each of the limitations of the invention including any one element or combination of elements may be included in each aspect of the invention. The disclosure of the invention is not limited in its application to the details of construction or the arrangement of components set forth in the following description or illustrated in the drawings.
[0062] The accompanying drawings are not intended to be drawn to scale. The drawings are merely illustrative and are not necessary for the operability of the present disclosure. For clarity, not every component is labeled in every drawing. In the drawings: [Brief description of the drawings]
[0063] [Figure 1] 1 depicts the profile of sesquiterpenes produced by delta-guayene synthase 2 from A. crassna. The protein sequence of delta-guayene synthase 2 from A. crassna is available at UniProtKB Accession No. D0VMR7 and is provided in this disclosure as SEQ ID NO:1. [Diagram 2] 2 depicts the results from screening of S. cerevisiae strains containing sesquiterpene synthases with amino acid substitutions compared to delta-guaiene synthase 2 (SEQ ID NO: 1) described in Example 1. A strain expressing GFP was used as a negative control. A strain expressing delta-guaiene synthase 2 (SEQ ID NO: 1) labeled as "WT D0VMR7" was used as a positive control. On the graph, strains with a higher percentage of alpha-guaiene to total sesquiterpene products than that of the positive control are shown. [Diagram 3] 3 depicts results from screening S. cerevisiae strains containing sesquiterpene synthases containing amino acid substitutions compared to delta-guayene synthase 2 (SEQ ID NO: 1) described in Example 2. Results from a secondary screen are shown. A strain expressing delta-guayene synthase 2 (SEQ ID NO: 1) labeled as "WT D0VMR7" was used as a positive control. [Figure 4] 4 depicts results from screening S. cerevisiae strains containing sesquiterpene synthases containing amino acid substitutions compared to delta-guayene synthase 2 (SEQ ID NO: 1) described in Example 3. Results from a secondary screen are shown. A strain expressing delta-guayene synthase 2 (SEQ ID NO: 1) labeled as "WT D0VMR7" was used as a positive control. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0064] Sesquiterpenes such as guaiene are widely used in the fragrance and flavor industry, but purification of sesquiterpenes from natural sources and de novo chemical synthesis often result in high production costs and / or low yields. This disclosure is premised, at least in part, on the unexpected discovery that amino acid substitutions in delta-guaiene synthase 2 protein (SEQ ID NO: 1) from A. crassuna can result in increased alpha-guaiene production relative to total sesquiterpene products, or increased alpha-guaiene production relative to the production of wild-type delta-guaiene synthase of SEQ ID NO: 1. Thus, this disclosure provides, in some embodiments, engineered enzymes for increased alpha-guaiene production, host cells expressing engineered enzymes, and methods for producing alpha-guaiene using such enzymes and host cells.
[0065] Terpenes As used in this disclosure, unless otherwise specified, terpene is isoprene (i.e., C 5 H 8 ) units and their derivatives. The terms "isoprenoid", "terpene" and "terpenoid" are used interchangeably in this application. For example, terpenes are compounds having the molecular formula (C 5 H 8 ) nTerpenes include pure hydrocarbons having the formula: (where n represents the number of isoprene subunits). Terpenes also include oxygen-containing compounds (often referred to as terpenoids). Terpenes are structurally diverse compounds, for example, cyclic (e.g., monocyclic, polycyclic, homocyclic and heterocyclic compounds) or acyclic (e.g., linear and branched compounds). In some embodiments, terpenes can be aroma compounds. As used in this disclosure, aroma compounds refer to compounds that have a characteristic odor. Any method known in the art, such as high performance liquid chromatography (HPLC), gas chromatography (GC) (e.g., gas chromatography combined with mass spectrometry (GC / MS) or gas chromatography combined with flame ionization detector (GC / FID)), can be used to identify the terpene of interest.
[0066] Non-limiting examples of terpenes include monoterpenes, sesquiterpenes, diterpenes, sesterterpenes, triterpenes, and tetraterpenes. Monoterpenes contain 10 carbons. Non-limiting examples of monoterpenes include, but are not limited to, myrcene, methanol, carvone, hinokitiol, linalool, limonene, sabinene, thujene, carene, borneol, eucalyptol, and camphene. Sesquiterpenes contain 15 carbons. As used in this disclosure, sesquiterpenes include sesquiterpene hydrocarbons and sesquiterpene alcohols (sesquiterpenols). Non-limiting examples of sesquiterpenes include, but are not limited to, delta-cadinene, epicuvenol, tau-cadinol, alpha-cadinol, gamma-selinene, 10-epi-gamma-eudesmol, gamma-eudesmol, alpha / beta-eudesmol, juniper camphor, 7-epi-alpha-eudesmol, cryptomeridiol isomer 1, cryptomeridiol isomer 2, cryptomeridiol isomer 3, humulene, alpha-guayene, delta-guayene (alpha-bulnesene), aciphyllene, beta-guayene, gamma-guayene, Guaiene is formed from three isoprene units and is often represented by the molecular formula C. 15 H 24A type of sesquiterpene having a guaiane sesquiterpene skeleton, examples of which include compounds such as alpha-guaiane, delta-guaiane (alpha-bulnesene), aciphyllene, beta-guaiane, gamma-guaiane, and gamma-gurjunene, all of which have a guaiane sesquiterpene skeleton. A diterpene contains 20 carbons. Non-limiting examples of diterpenes include, but are not limited to, cembrene and sclareol. A sesterterpene contains 25 carbons. A non-limiting example of a sesterterpene is geranylfarnesol. A triterpene contains 30 carbons. Non-limiting examples of triterpenes include squalene, polypodatetraene, malabaricane, lanostane, cucurbitacin, hopane, oleanane, and ursolic acid. A tetraterpene contains 40 carbons. Non-limiting examples of tetraterpenes include carotenoids, such as xanthophylls and carotenes.See, for example, WO2019 / 161141.In some embodiments, the isoprenoid is a cannabinoid.See, for example, WO2020 / 176547.
[0067] Terpene synthase As used in this disclosure, "terpene synthase" refers to a protein that can use prenyl diphosphate as substrate to produce terpene.At least two types of terpene synthase have been characterized: typical terpene synthase and isoprenyl diphosphate synthase type terpene synthase.Typical terpene synthase is found in prokaryotes (e.g., bacteria) and eukaryotes (e.g., plants, fungi and amoeba), while isoprenyl diphosphate synthase type terpene synthase is found in insects (see, for example, Chen et al., Proc Natl Acad Sci US A. 2016;113(43):12132-12137, which is incorporated herein by reference in its entirety). In typical terpene synthases, a number of highly conserved structural motifs have been reported, such as the aspartic acid-rich "DDxx(x)D / E" motif and the "NDxxSxxxD / E" motif, both of which are involved in regulating substrate binding (see, for example, Starks et al., Science. 1997 Sep 19;277(5333):1815-20; and Christianson et al., Curr Opin Chem Biol. 2008 Apr;12(2):141-50, the contents of each of which are incorporated herein by reference in their entirety). Also see, for example, WO2019 / 161141 and WO2020 / 176547.
[0068] Terpene synthase can be classified according to the type of terpene it produces.Terpene synthase can include, for example, monoterpene synthase, diterpene synthase, and sesquiterpene synthase.Specific non-limiting examples of monoterpene synthase and sesquiterpene synthase can be found, for example, in Degenhardt et al., Phytochemistry. 2009 Oct-Nov;70(15-16):1621-37, the entirety of which is incorporated herein by reference.
[0069] Monoterpene synthases catalyze the formation of 10-carbon monoterpenes. Generally, monoterpene synthases use geranyl diphosphate (GPP) as a substrate. Non-limiting examples of monoterpene synthases include myrcene synthase (UniProtKB Accession No. O24474), (R)-limonene synthase (UniprotKB Accession No. Q2XSC6), (E)-beta-ocimene synthase (UniProtKB Accession No. Q5CD81) and limonene synthase (UniProtKB Accession No. Q9FV72).
[0070] Diterpene synthase catalyzes the formation of 20-carbon diterpenes. Generally, diterpene synthase uses geranylgeranyl diphosphate as substrate. Non-limiting examples of diterpene synthase include cis-abienol synthase (UniProtKB Accession No. H8ZM73), sclareol synthase (UniProtKB Accession No. K4HYB0) and abietadiene synthase (UniProtKB Accession No. Q38710). For example, see Gong et al., Nat Prod Bioprospect. 2014;4(2):59-72, the entirety of which is incorporated herein by reference.
[0071] Sesquiterpene synthase catalyzes the formation of sesquiterpenes containing 15 carbon atoms. Generally, sesquiterpene synthase converts farnesyl diphosphate (FDP) to sesquiterpenes. Non-limiting examples of sesquiterpene synthases include (+)-delta-cadinene synthase (UniProtKB Accession No. Q9SAN0), UniProtKB Accession No. A0A067FTE8, beta-eudesmol synthase (UniProtKB Accession No. B1B1U4), (+)-delta-cadinene synthase isozyme XC14 (UniProtKB Accession No. Q39760), (+)-delta-cadinene synthase isozyme XC1 (UniProtKB Accession No. Q39761), (+)-delta-cadinene synthase isozyme A (UniProtKB Accession No. Q43714), sesquiterpene synthase 2 (UniProtKB Accession No. Q9F Q26), putative delta-guaiene synthase (UniProtKB accession number A0A0A0QUT9), delta-guaiene synthase 1 (UniProtKB accession number D0VMR6), alpha-zingiberene synthase (UniProtKB accession number Q5SBP4), (Z)-gamma-bisabolene synthase 1 (UniProtKB accession number Q9T0J9), A0A067D5M4, delta-elemene synthase (UniProtKB accession number A0A097ZIE0), ShoBecSQTS1, A0A068UHT0, terpene synthase (UniProtKB accession number G5CV47), A0A068VE40, and A0A068VI46.
[0072] The present disclosure also encompasses polyfunctional (e.g., capable of producing more than one sesquiterpene) sesquiterpene synthases. In some embodiments, the sesquiterpene synthase is capable of producing delta-cadinene and alpha-cadinol. In some embodiments, the sesquiterpene synthase is capable of producing delta-cadinene, tau-cadinol, and alpha-cadinol. In some embodiments, the sesquiterpene synthase is capable of producing alpha-guayene and delta-guayene. In some embodiments, the sesquiterpene synthase is capable of producing beta-caryophyllene and humulene. In some embodiments, the sesquiterpene synthase is capable of producing alpha-guayene, delta-guayene, beta-elemene, neointermediol, and / or humulene. Beta-elemene is further described in WO2005 / 052163, which is incorporated by reference therein. In some embodiments, the sesquiterpene synthase is capable of producing acifylene.In some embodiments, the sesquiterpene synthase is capable of producing alpha-guayene and acifylene.In some embodiments, the most abundant sesquiterpene produced by the sesquiterpene synthase is alpha-guayene, and the second most abundant sesquiterpene produced by the sesquiterpene synthase is acifylene.
[0073] The present disclosure also encompasses fragments of sesquiterpene synthase.As used in this disclosure, fragments of sesquiterpene synthase refer to parts of sesquiterpene synthase that are smaller than the full-length molecule.Fragments of sesquiterpene synthase of the present disclosure can include biologically active parts of enzymes, such as catalytic domains.Functional fragments of sesquiterpene synthase refer to fragments of sesquiterpene synthase that have the same type of activity as full-length sesquiterpene synthase, but the level of activity of the fragments may change compared to the level of activity of full-length sesquiterpene synthase.
[0074] Alpha-guayene production Guaiene has the molecular formula C 15 H 24 Non-limiting examples of guaienes include alpha-guaienes (α-guaienes; CAS Registry No. 3691-12-1), beta-guaienes (β-guaienes; CAS Registry No. 88-84-6), alpha-bulnesenes (also known as delta-guaienes (δ-guaienes; CAS Registry No. 3691-11-0), gamma-guaienes (γ-guaienes; CAS Registry No. 145267-53-4), aciphyllene (CAS Registry No. 87745-31-1), and gamma-gurjunenes (γ-gurjunenes; CAS Registry No. 22567-17-5).
[0075] In some embodiments, the sesquiterpene synthase associated with the present disclosure is the delta-guayene synthase 2 protein from A. krassuna. The amino acid sequence of the delta-guayene synthase 2 protein from A. krassuna is available under UniProtKB Accession No. D0VMR7, which is provided herein as SEQ ID NO:1: MSSAKLGSASEDVSRRDANYHPTVWGDFFLTHSSNFLENNDSILEKHEELKQEVRNLLVV ETSDLPSKIQLTDEIIRLGVGYHFETEIKAQLEKLHDHQLHLNFDLLTTSVWFRLLRGHG FSIPSDVFKRFKNTKGEFETEDARTLWCLYEATHLRVDGEDILEEAIQFSRKRLEALLPK LSFPLSECVRDALHIPYHRNVQRLAARQYIPQYDAEQTKIESLSLFAKIDFNMLQALHQS ELREASRWWKEFDFPSKLPYARDRIAEGYYWMMGAHFEPKFSLSRKFLNRIVGITSLIDD TYDVYGTLEEVTLFTEAVERWDIEAVKDIPKYMQVIYIGMLGIFEDFKDNLINARGKDYC IDYAIEVFKEIVRSYQREAEYFHTGYVPSYDEYMENSIISGGYKMFIILMLIGRGEFELK ETLDWASTIPEMVKASSLIARYIDDLQTYKAEEERGETVSAVRCYMREFGVSEEQACKKM REMIEIEWKRLNKTTLEADEISSSVVIPSLNFTRVLEVMYDKGDGYSDSQGVTKDRIAAL LRHAIEI (SEQ ID NO: 1)
[0076] A non-limiting example of a nucleic acid sequence that encodes SEQ ID NO:1 is provided herein as SEQ ID NO:2. atgtcttcagctaaactgggaagtgcctccgaggacgtttctcgtcgggatgcaaattacca ccctactgtatggggtgatttctttttgacgcatagctccaacttcctcgaaaaacaatgact cgatccttgaaaagcacgaagaattgaagcaggaagtgaggaacctattagtcgttgaaaca tctgacttgccatctaaaattcaattgaccgatgagataattagattgggtgttggctatca ttttgaaactgaaatcaaggctcaattagaaaagttgcacgaccaccaattgcatttaaact tcgatcttttgaccacttcggtctggttcagacttttgcgaggtcacggtttttccattcca tcagatgtttttaagagattcaaaaacaccaagggtgaatttgagactgaagacgcgagaac attatggtgtttgtacgaagccacccatcttagagttgacggggaggatatcctagaagagg ctatacaattctctcgtaagagattggaagctctgttacccaaattgagcttcccattgtcc goesgcgtcagagatgctctgcatattccttaccacagaaatgtacaaagattggccgctag acaaatatatcccacaatacgatgctgaacagactaagattgaatctctatctttattcgcca agatcgactttaacatgttgcaagctttgcaccaaagtgaactgagagaagcctcacgctggt ggaaaaatttgacttcccatccaagttaccttacgcaagggatagaattgctgaaggttatt actggatgatgggtgctcattttgaaccgaagttctctttgtctcgtaagttcttgaatagaa ttgtcggtatcacttcgttaattgatgacacgtatgatgtttacgggacccttgaagaggtga cactattcactgaagcggtcgaacgttgggacattgaggcagttaaagatatcccaaagtaca tgcaagtcatctacatcggtatgttgggtatctttgaagactttaaagataacttaattaatg ctcgaggtaaagattactgtatcgactatgctattgaagtctttaaggaaattgttcgttcct atcaaagagaagctgaatatttccacaccggctacgttccaagttacgatgaatacatggaga atagtaataatctggaggttacaagatgttcattattcttatgttgattggtagaggggaat ttgaattgaaggaaacactagactgggcctcaactattcctgaaatggttaaggctagttctc tcattgctagatacatcgacgatctccagacctataaggcagaggaggaaagaggtgaaactg tatcagctgttagatgttacatgagagaatttggtgtgtctgaggaacaagcttgcaagaaaa tgcgtgaaatgatcgaaattgaatggaagcgattgaacaagacgactttggaggcagacgaaa tatcctccagtgttgtaataccatctctcaacttcaccagagttttggaagtcatgtatgaca aaggtgatggttactctgattctcaaggtgtcacaaaggatagaatcgccgctttacttagac acgcaattgaaatc (SEQ ID NO: 2)
[0077] Other non-limiting examples of sesquiterpene synthases include those associated with UniProt Accession No. D0VMR6, UniProt Accession No. D0VMR8, and UniProt Accession No. Q49SP3.
[0078] The sesquiterpene synthase may comprise a sequence that is at least 50% identical (e.g., at least 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more than 99%, including all values in between) to SEQ ID NO:1 or a conservatively substituted version thereof. In certain embodiments, the sesquiterpene synthase comprises SEQ ID NO:1 or a conservatively substituted version thereof. In certain embodiments, the sesquiterpene synthase consists of or essentially consists of SEQ ID NO:1 or a conservatively substituted version thereof.
[0079] Sesquiterpene synthase mutants for increased production of alpha-guayene Naturally occurring sesquiterpene synthases that produce guaiene generally produce a mixture of different sesquiterpene products. For example, as shown in Example 1, expression of the naturally occurring delta-guaiene synthase 2 protein (SEQ ID NO: 1) from A. crassuna in a host cell resulted in the production of a mixture of sesquiterpene products that contained about 14.6% alpha-guaiene, as determined using gas chromatography (GC) (Table 4).
[0080] It would be advantageous to be able to affect the amount or ratio of various guaiene products produced by a sesquiterpene synthase. In particular, given the importance of alpha-guaiene to fragrances and flavors, it would be advantageous to be able to produce an increased amount of alpha-guaiene or an increased amount of alpha-guaiene relative to other sesquiterpene products. Aspects of the present disclosure relate to the surprising identification of synthetic sesquiterpene synthases that produce an increased amount of alpha-guaiene, or an increased ratio of alpha-guaiene to total sesquiterpene products, compared to those produced by a control sesquiterpene synthase. In some embodiments, the control sesquiterpene synthase comprises the sequence of SEQ ID NO: 1. In some embodiments, the sesquiterpene synthase associated with the present disclosure produces more alpha-guaiene than one or more other sesquiterpene products, such as delta-guaiene. As shown in Examples 1-3, it was surprisingly found that mutant sesquiterpene synthases containing amino acid substitutions relative to SEQ ID NO:1 produced an increased percentage of alpha-guayene relative to total sesquiterpene products compared to a sesquiterpene synthase containing the sequence of SEQ ID NO:1 (Tables 4, 6 and 7).
[0081] In some embodiments, the sequences of the sesquiterpene synthases associated with the present disclosure include one or more amino acid substitutions compared to SEQ ID NO:1, the one or more amino acid substitutions being 23, 44, 72, 86, 111, 118, 122, 124, 134, 147, 153, 188, 191, 201, 205, 212, 217, 224, 252, 255, 260, 262, 264, 266, 268, 270, 272, 276, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, and / or 542.
[0082] In some embodiments, the sequences of sesquiterpene synthases associated with the present disclosure contain one or more amino acid substitutions compared to SEQ ID NO:1, and at least one of the amino acid substitutions is at a position corresponding to positions 44, 122, 212, 217, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, 533, and / or 542 in SEQ ID NO:1. In some embodiments, the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to 23, 72, 86, 111, 118, 124, 134, 147, 153, 188, 191, 201, 205, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 448, 458, 467, 480, 494, 499, 509, 512, 516, 519, and / or 526 in SEQ ID NO:1.
[0083] In some embodiments, the sequence of the sesquiterpene synthase associated with the present disclosure comprises two or more amino acid substitutions compared to SEQ ID NO:1, (i) at least one of the amino acid substitutions is at a position corresponding to positions 122, 274, 275, 293, 295, 301, 368, 398, 404, 407, 507, 515, 525, and / or 533 in SEQ ID NO:1; (ii) at least one of the amino acid substitutions is at a position corresponding to positions 72, 124, 153, 191, 201, 205, 289, 290, 291, 292, 346, 406, 442, 480, 494, 509, 512, and / or 526 in SEQ ID NO:1.
[0084] In some embodiments, the sequences of the sesquiterpene synthases associated with the present disclosure comprise two or more amino acid substitutions compared to SEQ ID NO:1, and the two or more amino acid substitutions are at positions corresponding to 72, 122, 124, 153, 191, 201, 205, 274, 275, 289, 290, 291, 292, 293, 295, 301, 346, 368, 398, 404, 406, 407, 442, 480, 494, 507, 509, 512, 526 and / or 533 in SEQ ID NO:1.
[0085] Although one or more amino acid substitutions are described in this disclosure for the sequence of SEQ ID NO:1, it is understood that the disclosure encompasses the same one or more substitutions in the corresponding residues in other sesquiterpene synthase sequences. A person skilled in the art will be able to determine which one or more residues in different sesquiterpene synthases correspond to any one or more residues in the sequence of SEQ ID NO:1. As used in this application, a residue (e.g., a nucleic acid residue or an amino acid residue) in a first sequence is said to correspond to a position or residue (e.g., a nucleic acid residue or an amino acid residue) in a second sequence if the residue in the first sequence is in a position that is the counterpart of the residue in the second sequence when the first sequence and the second sequence are aligned using amino acid sequence alignment tools known in the art.
[0086] For example, SEQ ID NO:1 has a valine (V) residue at amino acid position 292. Some sesquiterpene synthases have an isoleucine (I) residue instead of a V residue at the corresponding position. It is understood that in sesquiterpene synthases having an I residue at the position corresponding to amino acid 292 in SEQ ID NO:1, the amino acid substitutions disclosed herein for that position are still applicable. Similarly, SEQ ID NO:1 has a serine (S) residue at amino acid position 296. Some sesquiterpene synthases have a proline (P) residue instead of an S residue at the corresponding position. It is understood that in sesquiterpene synthases having a P residue at the position corresponding to amino acid 296 in SEQ ID NO:1, the amino acid substitutions disclosed herein for that position (except for the S to P substitution) are still applicable.
[0087] In some embodiments, the sesquiterpene synthase comprises an aspartic acid (D) residue at a position corresponding to position 23 of SEQ ID NO:1; a valine (V) residue at a position corresponding to position 44 of SEQ ID NO:1; an isoleucine (I) residue at a position corresponding to position 72 of SEQ ID NO:1; a glutamic acid (E) residue at a position corresponding to position 86 of SEQ ID NO:1; a leucine (L) residue at a position corresponding to position 111 of SEQ ID NO:1; a glutamine (Q) residue at a position corresponding to position 118 of SEQ ID NO:1; an asparagine (N) residue at a position corresponding to position 122 of SEQ ID NO:1. a serine (S) residue at a position corresponding to position 124 of SEQ ID NO:1; a D residue at a position corresponding to position 134 of SEQ ID NO:1; a V residue at a position corresponding to position 147 of SEQ ID NO:1; an L residue at a position corresponding to position 147 of SEQ ID NO:1; an S residue at a position corresponding to position 153 of SEQ ID NO:1; a Q residue at a position corresponding to position 188 of SEQ ID NO:1; an N residue at a position corresponding to position 191 of SEQ ID NO:1; an S residue at a position corresponding to position 201 of SEQ ID NO:1; an I residue at a position corresponding to position 205 of SEQ ID NO:1 an E residue at position 217 of SEQ ID NO:1; an L residue at position 224 of SEQ ID NO:1; an L residue at position 252 of SEQ ID NO:1; an alanine (A) residue at position 255 of SEQ ID NO:1; an S residue at position 274 of SEQ ID NO:1; a glycine (G) residue at position 275 of SEQ ID NO:1; a lysine (K) residue at position 289 of SEQ ID NO:1; a histidine (H) residue at a position corresponding to position 290 of SEQ ID NO:1; an H residue at a position corresponding to position 290 of SEQ ID NO:1; a K residue at a position corresponding to position 290 of SEQ ID NO:1; a V residue at a position corresponding to position 291 of SEQ ID NO:1; an F or methionine (M) residue at a position corresponding to position 292 of SEQ ID NO:1; a cysteine (C) residue at a position corresponding to position 292 of SEQ ID NO:1; a Q, I, or N residue at a position corresponding to position 293 of SEQ ID NO:1; an I residue at a position corresponding to position 295 of SEQ ID NO:1; a V residue at a position corresponding to position 295 of SEQ ID NO:1;an S residue at a position corresponding to position 301 of SEQ ID NO:1; an E residue at a position corresponding to position 346 of SEQ ID NO:1; a C residue at a position corresponding to position 368 of SEQ ID NO:1; a tryptophan (W) residue at a position corresponding to position 381 of SEQ ID NO:1; an F residue at a position corresponding to position 390 of SEQ ID NO:1; a C residue at a position corresponding to position 398 of SEQ ID NO:1; a W residue at a position corresponding to position 404 of SEQ ID NO:1; an L residue at a position corresponding to position 406 of SEQ ID NO:1; a G residue at a position corresponding to position 407 of SEQ ID NO:1; an S residue at a position corresponding to position 419 of SEQ ID NO:1; an I residue at a position corresponding to position 433 of SEQ ID NO:1; an L residue at a position corresponding to position 433 of SEQ ID NO:1; an L residue at a position corresponding to position 442 of SEQ ID NO:1; a W residue at a position corresponding to position 443 of SEQ ID NO:1; a threonine (T) residue at a position corresponding to position 443 of SEQ ID NO:1; an N residue at a position corresponding to position 444 of SEQ ID NO:1; a G residue at a position corresponding to position 458 of SEQ ID NO:1; an L residue at a position corresponding to position 458 of SEQ ID NO:1; a V residue at a position corresponding to position 458 of SEQ ID NO:1; a K residue at a position corresponding to position 467 of SEQ ID NO:1; an I residue at a position corresponding to position 480 of SEQ ID NO:1; an E residue at a position corresponding to position 494 of SEQ ID NO:1; an N residue at a position corresponding to position 499 of SEQ ID NO:1; a W residue at a position corresponding to position 507 of SEQ ID NO:1; an A residue at a position corresponding to position 509 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; an M residue at a position corresponding to position 515 of SEQ ID NO:1; an A residue at a position corresponding to position 516 of SEQ ID NO:1; a V residue at a position corresponding to position 516 of SEQ ID NO:1; an L residue at a position corresponding to position 519 of SEQ ID NO:1; an F residue at a position corresponding to position 526 of SEQ ID NO:1; an N residue at a position corresponding to position 533 of SEQ ID NO:1; a V residue at a position corresponding to position 542 of SEQ ID NO:1; or any combination thereof.
[0088] In some embodiments, the sesquiterpene synthase contains a D residue at a position corresponding to position 23 of SEQ ID NO:1; a V residue at a position corresponding to position 44 of SEQ ID NO:1; an I residue at a position corresponding to position 72 of SEQ ID NO:1; an E residue at a position corresponding to position 86 of SEQ ID NO:1; an L residue at a position corresponding to position 111 of SEQ ID NO:1; a Q residue at a position corresponding to position 118 of SEQ ID NO:1; an N residue at a position corresponding to position 122 of SEQ ID NO:1; an S residue at a position corresponding to position 124 of SEQ ID NO:1; a D residue at a position corresponding to position 134 of SEQ ID NO:1; a V or L residue at a position corresponding to position 147 of sequence number 1; an S residue at a position corresponding to position 153 of sequence number 1; a Q residue at a position corresponding to position 188 of sequence number 1; an N residue at a position corresponding to position 191 of sequence number 1; an I or S residue at a position corresponding to position 201 of sequence number 1; an E residue at a position corresponding to position 205 of sequence number 1; an F residue at a position corresponding to position 212 of sequence number 1; an E residue at a position corresponding to position 217 of sequence number 1; an L residue at a position corresponding to position 224 of sequence number 1; a position corresponding to position 252 of sequence number 1 an L residue at a position corresponding to position 255 of SEQ ID NO:1; an A residue at a position corresponding to position 274 of SEQ ID NO:1; a G residue at a position corresponding to position 275 of SEQ ID NO:1; an L, T, S, H, M, K or D residue at a position corresponding to position 289 of SEQ ID NO:1; a K or H residue at a position corresponding to position 290 of SEQ ID NO:1; an F, L, T, V or C residue at a position corresponding to position 291 of SEQ ID NO:1; an A, Q, C, Y, H, E, F, M, W, T or F residue at a position corresponding to position 292 of SEQ ID NO:1; an L, V, T, Y, C, F, W, Q, I, N, or M residue at a position corresponding to position 295 of SEQ ID NO:1; an E, D, N, W, G, V, or I residue at a position corresponding to position 301 of SEQ ID NO:1; an E residue at a position corresponding to position 346 of SEQ ID NO:1; a C residue at a position corresponding to position 368 of SEQ ID NO:1; a W residue at a position corresponding to position 381 of SEQ ID NO:1; an F residue at a position corresponding to position 390 of SEQ ID NO:1; a C residue at a position corresponding to position 398 of SEQ ID NO:1; a W residue at a position corresponding to position 404 of SEQ ID NO:1;an L, N, W, or T residue at a position corresponding to position 406 of SEQ ID NO:1; a G residue at a position corresponding to position 407 of SEQ ID NO:1; an S residue at a position corresponding to position 419 of SEQ ID NO:1; an I or L residue at a position corresponding to position 433 of SEQ ID NO:1; an L residue at a position corresponding to position 442 of SEQ ID NO:1; a W or T residue at a position corresponding to position 443 of SEQ ID NO:1; an N residue at a position corresponding to position 444 of SEQ ID NO:1; a G residue at a position corresponding to position 448 of SEQ ID NO:1; an L or V residue at a position corresponding to position 458 of SEQ ID NO:1; a K residue at a position corresponding to position 467 of SEQ ID NO:1; an I residue at a position corresponding to position 480 of SEQ ID NO:1; an E residue at a position corresponding to position 494 of SEQ ID NO:1; an N residue at a position corresponding to position 499 of SEQ ID NO:1; a W residue at a position corresponding to position 507 of SEQ ID NO:1; an A residue at a position corresponding to position 509 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; an M residue at a position corresponding to position 515 of SEQ ID NO:1; an A or V residue at a position corresponding to position 516 of SEQ ID NO:1; an L residue at a position corresponding to position 519 of SEQ ID NO:1; an F residue at a position corresponding to position 526 of SEQ ID NO:1; an N residue at a position corresponding to position 533 of SEQ ID NO:1; a V residue at a position corresponding to position 542 of SEQ ID NO:1; or any combination thereof.
[0089] In some embodiments, the sesquiterpene synthase has the following amino acid substitutions relative to SEQ ID NO:1: N289H, V292F, G293N, T295V, K404W, F406L, and F512L; N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L; N289H, F406L, F512L, D191N, and A205E; I291V, V292F, G293Q, T295V, K404W, F406L, I507W, and F512L; N289H, V292F, T295V ,K404W,F406L,and F512L;N289H,I291V,V292F,G293N,T295V,K404W,F406L,and F512L;V292F,G293Q,K404W,and F406L;N289H,V292F,G293I,T295V,F406L,I407G,and F512L;V292F,G293Q,K404W,F406L,I407G,and F512L;N289H,I291V,V292F,G293I,T295V,F406L,I407G,and F512L;V292F,T 295V, K404W, F406L, I507W, and F512L;N289H, V292M, G293I, F406L, F512L, and M480I;N289H, I291V, V292F, G293Q, T295V, K404W, and F406L;V292F, G293Q, T295V, F406L, and F512L;V292F, T295V, K404W, F406L, and F512L;I291V, V292F, G293I, K404W, F406L, and I507W;V292F, G293I, T295V, F40 6L, I407G, and F512L;N289H, V292F, G293I, T295V, F406L, I507W, and F512L;N289H, I291V, V292F, G293I, T295V, K404W, F406L, and F512L;N289H, V292F, G293Q, T295V, F406L, I407G, and F512L;N289H, I291V, V292F, G293N, T295V, K404W, and F406L;N289H, V292F, G293Q, F406L, I507W, and F512L;I291V, V292F, G293N, T295V, K404W, F406L, I507W, and F512L;N289H, V292F, G293N, K404W, F406L, I507W, and F512L;N289H, V292F, G293I, T295V, F406L, I407G, I507W, and F512L;N289H, I291V, V292F, G293N, T295V, F406L, and F512L;I291V, V292F, G293Q, K404W, F406L, I407G, and F512L;I291V, V292F ,G293Q,T295V,K404W,F406L,I407G,and F512L;N289H,V292F,T295V,F406L,and F512L;N289H,V292F,K404W,F406L,and F512L;V292F,G293N,T295V,F406L,and F512L;N289H,I291V,V292F,G293N,T295V,F406L,I407G,I507W,and F512L;N289H,I291V,V292F,G293I,T295V,F406L,and F512L;N289H,I 291V, V292F, G293Q, T295V, F406L, and I507W;N289H, V292F, G293I, F406L, I407G, I507W, and F512L;N289H, V292F, G293Q, T295V, K404W, F406L, and I507W;I291V, V292F, G293Q, T295V, F406L, I507W, and F512L;S122N, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L;P124S, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L;T153S, N289H, I291V, V292F, G293Q, T295V, D346E, K404W, F406L, and F512L;P124S, N289H, I291V, V292F, G293Q, T295V, K404W, and F406L;R290K, V292F, T295V, K404W, F406L, and F512L;N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L;N289H, V292F, G293N, T295V, K404W, F406L, Y442L, and F512L;V292F, G293Q, and K404W;T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, T494E, and F512L;T72I, N289H, V292F, T295V, K404W, F406L, T494E, and F512L;T72I, N289H, V292F, G293I, T295V, F406L, I407G, Y442L, I507W, and F512L;A27 5G, G293W, I294V, T295W, T301S, F368C, S509A, F512L, Y526F, and T533N;P124S, V292F, T295V, K404W, F406L, I507W, and F512L;N289H, V292F, T295V, K404W, F406L, Y442L, and F512L;N289H, F406L, F512L, and T295V;N289H, F406L, F512L, T295V, and G274S;N289H, F406L, F512L, T295V, and V201I;N289H, F406L, F512L, T295V, and I398C;T72I, N289H, R290K, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, T494E, and F512L;L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L;L44V, T72I, T86E, G118Q, T134D, W147L, Q212F, Q217E, S224L, F252L, P255A, N289H, R290K, I291V, V292F, G293Q, T295V, D346E, Y390F, K404W, F406L, L419S, Y442L, T458V, R467K, F512L, and R542V;T23D, L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L;V201S, N289H, I291V, V292F, G293N, T295V, Y381W, K404W, F406L, T494E, and F512L;V201S, V292C, T295I, and D444N;V201S, V292C, T295I, D444N, and L516V;L44V, T72I, V201S, V292C, T295I, D444N, and L516V;L44V, T72I, V201S, V292C, T295I, D444N, and L516A;L44V, T72I, V201S, V292C, T295I, I443W, and D444N;L44V, T72I, V201S, V292C , T295I, D444N, and M519L;W147V, C188Q, V201S, S224L, V292C, T295I, and D444N;L44V, T72I, V111L, V201S, S224L, V292C, T295I, I443T, D444N, T448G, and M519L;L44V, T72I, V111L, V201S, V292C, T295I, V433I, I443T, D444 N, T448G, D499N, and L516V;L44V, T72I, V201S, V292C, T295I, D444N, L516V, and M519L;L44V, T72I, V201S, V292C, T295I, D444N, V515M, and L516V;V111L, V201S, S224L, V292C, T295I, I443T, D444N, and L516V;V201S, V292C , T295I, D444N, and L516V; V201S, V292C, T295I, D444N, T458L, and L516A; V201S, V292C, T295I, I443T, D444N, and L516V; V201S, V292C, T295I, V433I, D444N, and L516V; or V201S, V292C, T295I, V433L, D444N, and L516V.
[0090] In some embodiments, the sesquiterpene synthase comprising the amino acid substitutions N289H, V292F, G293N, T295V, K404W, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:3; the sesquiterpene synthase comprising the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:4; the sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, D191N, and A20 A sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293Q, T295V, K404W, F406L, I507W, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:6; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, K404W, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:7; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, The sesquiterpene synthase comprising G293N, T295V, K404W, F406L, and F512L comprises the sequence of SEQ ID NO:8; the sesquiterpene synthase comprising, relative to SEQ ID NO:1, the amino acid substitutions V292F, G293Q, K404W, and F406L comprises the sequence of SEQ ID NO:9; the sesquiterpene synthase comprising, relative to SEQ ID NO:1, the amino acid substitutions N289H, V292F, G293I, T295V, F406L, I407G, and F512L comprises the sequence of SEQ ID NO:10; the sesquiterpene synthase comprising, relative to SEQ ID NO:1, the amino acid substitutions V292F , G293Q, K404W, F406L, I407G, and F512L, comprises the sequence of SEQ ID NO:11; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293I, T295V, F406L, I407G, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:12; a sesquiterpene synthase that includes the amino acid substitutions V292F, T295V, K404W, F406L, I507W, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:13;A sesquiterpene synthase comprising the amino acid substitutions N289H, V292M, G293I, F406L, F512L, and M480I, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:14; a sesquiterpene synthase comprising the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, K404W, and F406L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:15; a sesquiterpene synthase comprising the amino acid substitutions V292F, G293Q, T295V, F406L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:16. comprises the sequence of SEQ ID NO:16; a sesquiterpene synthase that comprises the amino acid substitutions V292F, T295V, K404W, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:17; a sesquiterpene synthase that comprises the amino acid substitutions I291V, V292F, G293I, K404W, F406L, and I507W compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:18; a sesquiterpene synthase that comprises the amino acid substitutions V292F, G293I, T295V, F406L, I407G, and F512L compared to SEQ ID NO:1 The terpene synthase comprises the sequence of SEQ ID NO:19; the sesquiterpene synthase comprises the amino acid substitutions N289H, V292F, G293I, T295V, F406L, I507W, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:20; the sesquiterpene synthase comprises the amino acid substitutions N289H, I291V, V292F, G293I, T295V, K404W, F406L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:21; the sesquiterpene synthase comprises the amino acid substitutions N289H, V292F, G293I, T295V, K404W, F406L, and F512L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:22; , G293Q, T295V, F406L, I407G, and F512L, comprises the sequence of SEQ ID NO:22; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293N, T295V, K404W, and F406L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:23; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293Q, F406L, I507W, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:24;A sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293N, T295V, K404W, F406L, I507W, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:25; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293N, K404W, F406L, I507W, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:26; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293I, T295V, F406L, I407G, I507W, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:27. and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:27; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293N, T295V, F406L, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:28; a sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293Q, K404W, F406L, I407G, and F512L, compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:29 .... A sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, F406L, I407G, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:30; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:31; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, K404W, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:32; a sesquiterpene synthase that includes the amino acid substitutions V292F, G295V, I407G, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:33. a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293N, T295V, F406L, I407G, I507W, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:34; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293I, T295V, F406L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:35;A sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, F406L, and I507W, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:36; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293I, F406L, I407G, I507W, and F512L, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:37; a sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293Q, T295V, K404W, F406L, and I507W, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:38; A sesquiterpene synthase that includes the amino acid substitutions I291V, V292F, G293Q, T295V, F406L, I507W, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:39; a sesquiterpene synthase that includes the amino acid substitutions S122N, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:84; A sesquiterpene synthase comprising the amino acid substitutions P124S, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, and F512L comprises the sequence of SEQ ID NO: 80; a sesquiterpene synthase comprising the amino acid substitutions T153S, N289H, I291V, V292F, G293Q, T295V, D346E, K404W, F406L, and F512L compared to SEQ ID NO: 1 comprises the sequence of SEQ ID NO: 82; a sesquiterpene synthase that includes the amino acid substitutions R290K, V292F, T295V, K404W, F406L, and F512L, relative to SEQ ID NO:1, includes the sequence of SEQ ID NO:83; a sesquiterpene synthase that includes the amino acid substitutions N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L, relative to SEQ ID NO:1, includes the sequence of SEQ ID NO:78;A sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, G293N, T295V, K404W, F406L, Y442L, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:77; a sesquiterpene synthase that includes the amino acid substitutions T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, T494E, and F512L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:79; A sesquiterpene synthase that includes the amino acid substitutions T72I, N289H, V292F, G293I, T295V, F406L, I407G, Y442L, I507W, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:86; a sesquiterpene synthase that includes the amino acid substitutions P124S, V292F, T295V, K404W, F406L, T494E, and F512L, compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:87; A sesquiterpene synthase that includes the amino acid substitutions N289H, V292F, T295V, K404W, F406L, Y442L, and F512L, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:87; a sesquiterpene synthase that includes the amino acid substitutions N289H, F406L, F512L, and T295V, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:89; A sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, T295V, and G274S comprises the sequence of SEQ ID NO:90; a sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, T295V, and V201I, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:91; a sesquiterpene synthase comprising the amino acid substitutions N289H, F406L, F512L, T295V, and I398C, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:92;A sesquiterpene synthase that includes the amino acid substitutions T72I, N289H, R290K, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, T494E, and F512L, as compared to SEQ ID NO:1, includes the sequence of SEQ ID NO:110; A sesquiterpene synthase comprising L, Y442L, and F512L comprises the sequence of SEQ ID NO:111; and the amino acid substitutions L44V, T72I, T86E, G118Q, T134D, W147L, Q212F, Q217E, S224L, F252L, P255A, N289H, R290K, I291V, V292F, G293Q, T295V, D346E, Y390F, K404W, F406L, compared to SEQ ID NO:1; A sesquiterpene synthase that includes the amino acid substitutions L419S, Y442L, T458V, R467K, F512L, and R542V comprises the sequence of SEQ ID NO:112; a sesquiterpene synthase that includes the amino acid substitutions T23D, L44V, T72I, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:113; a sesquiterpene synthase that includes the amino acid substitutions V201S, N289H, I291V, V292F, G293Q, T295V, K404W, F406L, Y442L, and F512L compared to SEQ ID NO:1 comprises the sequence of SEQ ID NO:114; A sesquiterpene synthase that includes the amino acid substitutions V201S, V292C, T295I, and D444N compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:115; a sesquiterpene synthase that includes the amino acid substitutions V201S, V292C, T295I, D444N, and L516V compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:116. a sesquiterpene synthase that, compared to SEQ ID NO:1, includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and L516V, comprises the sequence of SEQ ID NO:117; a sesquiterpene synthase that, compared to SEQ ID NO:1, includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and L516A, comprises the sequence of SEQ ID NO:118; a sesquiterpene synthase that, compared to SEQ ID NO:1, includes the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and L516A, comprises the sequence of SEQ ID NO:119; , I443W, and D444N comprise the sequence of SEQ ID NO:119; a sesquiterpene synthase that comprises the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, and M519L, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:120; a sesquiterpene synthase that comprises the amino acid substitutions W147V, C188Q, V201S, S224L, V292C, T295I, and D444N, relative to SEQ ID NO:1, comprises the sequence of SEQ ID NO:121;A sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V111L, V201S, S224L, V292C, T295I, I443T, D444N, T448G, and M519L compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:122; a sesquiterpene synthase that includes the amino acid substitutions L44V, T72I, V111L, V201S, V292C, T295I, V433I, I443T, D444N, T448G, D499N, and L516V compared to SEQ ID NO:1 includes the sequence of SEQ ID NO:123; The sesquiterpene synthase comprising the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, L516V, and M519L, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:124; the sesquiterpene synthase comprising the amino acid substitutions L44V, T72I, V201S, V292C, T295I, D444N, V515M, and L516V, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:125; the sesquiterpene synthase comprising the amino acid substitutions V111L, V201S, S224L, V292C, T295I, I4 A sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, D444N, and L516V, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:127; a sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, D444N, T458L, and L516A, as compared to SEQ ID NO:1, comprises the sequence of SEQ ID NO:128; The sesquiterpene synthase comprising V292C, T295I, I443T, D444N, and L516V comprises the sequence of SEQ ID NO: 129; the sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, V433I, D444N, and L516V, relative to SEQ ID NO: 1, comprises the sequence of SEQ ID NO: 130; or the sesquiterpene synthase comprising the amino acid substitutions V201S, V292C, T295I, V433L, D444N, and L516V, relative to SEQ ID NO: 1, comprises the sequence of SEQ ID NO: 131;
[0091] In some embodiments, the sesquiterpene synthase further comprises one or more amino acid substitutions compared to SEQ ID NO:1, wherein the one or more amino acid substitutions are at one or more positions corresponding to positions 294, 296, 297, 403, 444, 515, and / or 525 in SEQ ID NO:1. an L, C, Y, V, A, or S residue at a position corresponding to position 294 of SEQ ID NO:1; an L, A, proline (P), Y, N, F or R residue at a position corresponding to position 296 of SEQ ID NO:1; an E, Y, I, lysine (K), M, or H residue at a position corresponding to position 297 of SEQ ID NO:1; an M, Q, N, S, T, A, E, H, C, or V residue at a position corresponding to 403 of SEQ ID NO:1; an A or N residue at a position corresponding to position 444 of SEQ ID NO:1; an H, A, E, or Q residue at a position corresponding to position 515 of SEQ ID NO:1; an H, C, L, or N residue at a position corresponding to position 525 of SEQ ID NO:1; or any combination thereof Includes.
[0092] In some embodiments, the sesquiterpene synthase is a Q, C, V, F, A, I, H, G, W, or Y residue at a position corresponding to 289 of SEQ ID NO:1; a V, M, or A residue at a position corresponding to position 291 of SEQ ID NO:1; an I, M, S, L, G, or N residue at a position corresponding to position 292 of SEQ ID NO:1; an I, Q, S, N, or E residue at a position corresponding to position 293 of SEQ ID NO:1; an L, V, A, or S residue at a position corresponding to position 295 of SEQ ID NO:1; a K or Q residue at a position corresponding to position 296 of SEQ ID NO:1; a T or A residue at a position corresponding to position 297 of SEQ ID NO:1; a P, F, I, L, G, or D residue at a position corresponding to 403 of SEQ ID NO:1; a G, S, A, I, Y, M, H, V, Q, or C residue at a position corresponding to 406 of SEQ ID NO:1; an H residue at a position corresponding to position 444 of SEQ ID NO:1; an M, C, F, G, N, S, or I residue at a position corresponding to position 515 of SEQ ID NO:1; or any combination thereof Does not include.
[0093] Table 1 provides examples of amino acid substitutions in SEQ ID NO:1 that are relevant to the present disclosure.
[0094] [Table 1-1] [Table 1-2]
[0095] In some embodiments, the sesquiterpene synthase comprises a sequence that is at least 50% identical (e.g., at least 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, including all values therebetween) to any one of SEQ ID NOs: 3-39, 77-92 or 110-131 or a conservatively substituted variant thereof. In certain embodiments, the sesquiterpene synthase comprises a sequence of any one of SEQ ID NOs: 3-39, 77-92 or 110-131 or a conservatively substituted version thereof. In certain embodiments, the sesquiterpene synthase consists or consists essentially of any one of the sequences of SEQ ID NOs: 3-39, 77-92, or 110-131, or conservatively substituted variants thereof.
[0096] In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces one or more sesquiterpenes including alpha-guayene and / or delta-guayene. In some embodiments, a sesquiterpene synthase associated with the disclosure, or a host cell expressing a sesquiterpene synthase associated with the disclosure, produces at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 104%, 105%, 106%, 107%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% alpha-guayene, or 100% alpha-guayene, as a percentage of total sesquiterpene products, as determined by GC. In some embodiments, the sesquiterpene synthase related to the present disclosure, or the host cell expressing the sesquiterpene synthase related to the present disclosure, produces at least the recited percentages more alpha-guayene than the wild-type enzyme of SEQ ID NO: 1 or the host cell expressing the wild-type enzyme, as determined by GC.In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces at least 50% alpha-guayene as a percentage of total sesquiterpene products as determined by GC, or produces at least 50% more alpha-guayene as determined by GC than that produced by the wild-type enzyme of SEQ ID NO:1 or a host cell expressing the wild-type enzyme.
[0097] In some embodiments, a sesquiterpene synthase associated with the present disclosure, or a host cell expressing a sesquiterpene synthase associated with the present disclosure, produces about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, , 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69% or 70% delta-guayene, or produce 100% delta-guayene as a percentage of total sesquiterpene products as determined by GC.
[0098] In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces more alpha-guayene than delta-guayene. For example, in some embodiments, a sesquiterpene synthase associated with the disclosure, or a host cell expressing a sesquiterpene synthase associated with the disclosure, has a sesquiterpene synthase that is at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 104%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81% , 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 150%, 200%, 250%, 300%, 350%, 400%, 450%, or 500% more alpha-guayene. In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces at least 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times more alpha-guayene than delta-guayene as determined by GC.
[0099] In some embodiments, a sesquiterpene synthase associated with the disclosure, or a host cell expressing a sesquiterpene synthase associated with the disclosure, produces about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 108%, 109%, 109%, 109%, 1 Produce 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69% or 70% aciphyllenes, or produce 100% aciphyllenes as a percentage of total sesquiterpene products as determined by GC.
[0100] In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces more alpha-guayene than acifilene. For example, in some embodiments, a sesquiterpene synthase associated with the present disclosure, or a host cell expressing a sesquiterpene synthase associated with the present disclosure, has at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 108%. %,48%,49%,50%,51%,52%,53%,54%,55%,56%,57%,58%,59%,60%,61%,62%,63%,64%,65%,66%,67%,68%,69%,70%,71%,72%,73%,74%,75%,76%,77%,78%,79%,80%,81%, In some embodiments, the sesquiterpene synthases related to the present disclosure, or host cells expressing the sesquiterpene synthases related to the present disclosure, produce at least 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times more alpha-guayene than aciphylene as determined by GC.
[0101] In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces more alpha-guayene than produced by the wild-type enzyme of SEQ ID NO:1 or a host cell expressing the wild-type enzyme. For example, in some embodiments, a sesquiterpene synthase related to the disclosure, or a host cell expressing a sesquiterpene synthase related to the disclosure, has a nucleotide sequence that is at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1090%, 1091%, 1092%, 109 3%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79% , 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 150%, 200%, 250%, 300%, 350%, 400%, 450%, or 500% more alpha-guayene. In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces at least 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times more alpha-guayene than that produced by the wild-type enzyme of SEQ ID NO: 1 or a host cell expressing the wild-type enzyme, as determined by GC.
[0102] In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces a higher ratio of alpha-guayene to delta-guayene than that produced by a sesquiterpene synthase comprising the sequence of SEQ ID NO: 1 or a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO: 1. In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces more alpha-guayene as a percentage of total sesquiterpene products than that produced by a sesquiterpene synthase comprising the sequence of SEQ ID NO: 1 or a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO: 1. In some embodiments, the sesquiterpene synthase related to the present disclosure or the host cell expressing the sesquiterpene synthase related to the present disclosure produces less delta-guayene as a percentage of total sesquiterpene products than that produced by the sesquiterpene synthase comprising the sequence of SEQ ID NO: 1 or the host cell expressing the sesquiterpene synthase comprising the sequence of SEQ ID NO: 1. In some embodiments, the sesquiterpene synthase related to the present disclosure or the host cell expressing the sesquiterpene synthase related to the present disclosure produces more alpha-guayene than that produced by the sesquiterpene synthase comprising the sequence of SEQ ID NO: 1 or the host cell expressing the sesquiterpene synthase comprising the sequence of SEQ ID NO: 1. In some embodiments, the sesquiterpene synthase related to the present disclosure or the host cell expressing the sesquiterpene synthase related to the present disclosure produces more aciphyllene than that produced by the sesquiterpene synthase comprising the sequence of SEQ ID NO: 1 or the host cell expressing the sesquiterpene synthase comprising the sequence of SEQ ID NO: 1.In some embodiments, a sesquiterpene synthase related to the present disclosure, or a host cell expressing a sesquiterpene synthase related to the present disclosure, produces more alpha-guayene and acifylene than produced by a sesquiterpene synthase comprising the sequence of SEQ ID NO:1 or a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO:1.
[0103] For example, in some embodiments, a sesquiterpene synthase associated with the present disclosure, or a host cell expressing a sesquiterpene synthase associated with the present disclosure, produces at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 120%, 122%, 130%, %, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109 ... %, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 150%, 200%, 250%, 300%, 350%, 400%, 450%, or 500% more alpha-guayene.In some embodiments, a sesquiterpene synthase associated with the disclosure, or a host cell expressing a sesquiterpene synthase associated with the disclosure, produces at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, or more sesquiterpene products as a percentage of total sesquiterpene products than a sesquiterpene synthase comprising the sequence of SEQ ID NO:1 or a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO:1. 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77% %, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 150%, 200%, 250%, 300%, 350%, 400%, 450%, or 500% less delta-guayene.In some embodiments, a sesquiterpene synthase associated with the disclosure, or a host cell expressing a sesquiterpene synthase associated with the disclosure, produces at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, or more sesquiterpene products as a percentage of total sesquiterpene products than a sesquiterpene synthase comprising the sequence of SEQ ID NO:1 or a host cell expressing a sesquiterpene synthase comprising the sequence of SEQ ID NO:1. 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77% , 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 150%, 200%, 250%, 300%, 350%, 400%, 450%, or 500% more alpha-guayene.
[0104] It is understood that in some embodiments, the total sesquiterpene products in a sample can be measured, and the amount of alpha-guayene or delta-guayene can be calculated relative to the amount of total sesquiterpene products in a sample. It is understood that in some embodiments, the total sesquiterpene products in a sample can be measured, and the amount of alpha-guayene or acifylene can be calculated relative to the amount of total sesquiterpene products in a sample. In other embodiments, less than the total amount of sesquiterpene products in a sample may be measured. For example, in some embodiments, only the amount of alpha-guayene and delta-guayene in a sample may be measured. In some embodiments, only the amount of alpha-guayene and acifylene in a sample may be measured. In such embodiments, the amount of alpha-guayene relative to the combined amount of alpha-guayene and delta-guayene in a sample, or the amount of alpha-guayene relative to the amount of delta-guayene in a sample may be calculated.
[0105] In some embodiments, the amount of alpha-guaiene produced by the host cell, cell culture, or enzyme related to the present disclosure is at least 1 g / L. In some embodiments, the amount of alpha-guaiene produced by the host cell, cell culture, or enzyme related to the present disclosure is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, or 5.0 g / L.
[0106] In some embodiments, the amount of aciphylene produced by the host cells, cell cultures, or enzymes related to the present disclosure is at least 1 g / L. In some embodiments, the amount of alpha-guayene produced by the host cells, cell cultures, or enzymes related to the present disclosure is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, or 5.0 g / L.
[0107] The mass of the sesquiterpene can be determined by any method known in the art. In some embodiments, the mass of the sesquiterpene is determined by GCMS.
[0108] In some embodiments, the amino acid substitution is introduced into a residue at or near the active site of the sesquiterpene synthase. Mutations at or near the active site can change the structural conformation or catalytic activity of the enzyme, biasing the enzyme to produce one product over another. In some embodiments, the amino acid substitution results in a hydrophilic (polar) amino acid residue at or near the active site. "Hydrophilic residue" as used in this disclosure refers to an amino acid that has a positive charge that attracts water molecules. In some embodiments, the amino acid substitution results in a hydrophobic (non-polar) amino acid residue at or near the active site. "Hydrophobic residue" as used in this disclosure refers to a non-polar amino acid that repels water molecules. In some embodiments, the amino acid substitution results in the introduction of an aromatic amino acid residue at or near the active site, which can help modulate cation-dipole interactions to affect reaction specificity. Examples of aromatic amino acid residues include tyrosine, tryptophan, and phenylalanine.
[0109] Without wishing to be bound by any theory, the shape of the active site pocket of sesquiterpene synthase may affect the ratio of different products produced by sesquiterpene synthase. In the reaction catalyzed by delta-guaiene synthase 2 protein (UniProt accession number D0VMR7; SEQ ID NO: 1) from A. crassuna, the final step involves proton abstraction from either the pro-delta carbon, which produces delta-guaiene as product, or the pro-alpha carbon, which produces alpha-guaiene as product. In some embodiments, the mutant sesquiterpene synthase can produce increased amounts of alpha-guaiene, which is at least partially due to changing the substrate binding mode of sesquiterpene synthase to allow easier access to the pro-alpha carbon from catalytic residue Tyr-520.
[0110] In some embodiments, the sesquiterpene synthase associated with the present disclosure comprises: (a) one or more amino acid substitutions in a first region of the active site compared to SEQ ID NO:1, wherein the one or more amino acid substitutions in the first region are at positions corresponding to positions 295, 291, 406, 512, and / or 519 of SEQ ID NO:1, and the one or more amino acid substitutions in the first region comprise a residue having a smaller side chain than the side chain of the amino acid at positions 295, 291, 406, 512, and / or 519 of SEQ ID NO:1, respectively; (b) an amino acid substitution in a second region of the active site compared to SEQ ID NO:1, wherein the amino acid substitution in the second region corresponds to position 292 of SEQ ID NO:1, and the amino acid substitution in the second region comprises a larger side chain than the side chain of the amino acid at position 292 of SEQ ID NO:1; or any combination thereof.
[0111] In some embodiments, the side chain in (a) is hydrophobic. In some embodiments, the amino acid residue with a smaller side chain in (a) is selected from the group consisting of alanine (A), serine (S), threonine (T), and valine (V). In some embodiments, the amino acid residue with a larger side chain in (b) is selected from the group consisting of arginine (R), phenylalanine (F), tyrosine (Y), and tryptophan (W).
[0112] In some embodiments, the sesquiterpene synthase associated with the present disclosure is a V residue at a position corresponding to position 295 of SEQ ID NO:1; a C residue at a position corresponding to position 291 of SEQ ID NO:1; an L residue at a position corresponding to position 406 of SEQ ID NO:1; an L residue at a position corresponding to position 512 of SEQ ID NO:1; an F residue at a position corresponding to position 292 of SEQ ID NO:1; or any combination thereof; Includes.
[0113] Mutants Aspects of the present disclosure relate to polynucleotides that encode any of the recombinant polypeptides related to the present disclosure, such as sesquiterpene synthases. Mutants of the polynucleotide or polypeptide sequences described in this application are also encompassed by the present disclosure. A "mutant" polynucleotide, as used in this disclosure, refers to a polynucleotide whose nucleic acid sequence differs from that of a reference polynucleotide by one or more changes in the nucleic acid sequence. A "mutant" polypeptide, as used in this disclosure, refers to a polypeptide whose amino acid sequence differs from that of a reference polypeptide by one or more changes in the amino acid sequence.
[0114] Mutant polynucleotides or polypeptides can be constructed synthetically. Typically, the polynucleotides or polypeptides from which the mutants are derived are wild-type polynucleotides, wild-type polypeptides, or wild-type polynucleotides or polypeptide domains. However, the mutants suitable for use in the present disclosure can also be derived from homologs, orthologs, or paralogs of wild-type polynucleotides, wild-type polypeptides, or wild-type polynucleotides or polypeptide domains, or from synthetic polynucleotides or polypeptides. The changes in the nucleic acid and / or amino acid sequence can include substitutions, insertions, deletions, N-terminal truncations, C-terminal truncations, N-terminal additions, C-terminal additions, or any combination of these changes, and can be present at one or more positions.
[0115] A variant may have at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the reference sequence, inclusive of all values in between.
[0116] The term "sequence identity" refers to the relationship between two polypeptide or polynucleotide sequences, as determined by sequence comparison (alignment), unless otherwise specified, as known in the art. In some embodiments, sequence identity is determined over the entire length of the sequence, while in other embodiments, sequence identity is determined over a region of the sequence.
[0117] Identity may also refer to the degree of sequence relatedness between two sequences, as determined by the number of matches between a series of two or more residues (e.g., nucleic acid or amino acid residues). Identity measures the percent of identical matches between two or more sequences, including gap alignments (if present), as assigned by a particular mathematical model, algorithm, or computer program.
[0118] The identity of related polypeptide or nucleic acid sequences can be easily calculated by any method known to those skilled in the art. The "percent identity" of two sequences (e.g., nucleic acid or amino acid sequences) can be determined, for example, using the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, modified as described in Karlin and Altschul Proc. Natl. Acad. Sci. USA 90:5873-77, 1993. Such an algorithm is incorporated in the NBLAST® and XBLAST® programs (version 2.0) of Altschul et al., J. Mol. Biol. 215:403-10, 1990. BLAST® protein searches can be performed, for example, using the XBLAST program, score=50, wordlength=3, to obtain amino acid sequences homologous to the protein molecules of the present invention. When gaps exist between the two sequences, Gapped BLAST® can be utilized, for example, as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When utilizing BLAST® and Gapped BLAST® programs, the default parameters of the respective programs (e.g., XBLAST® and NBLAST®) can be used, or the parameters can be appropriately adjusted, as would be understood by one of skill in the art.
[0119] Another local alignment technique that can be used is, for example, based on the Smith-Waterman algorithm (Smith, TF & Waterman, MS (1981) "Identification of common molecular subsequences." J. Mol. Biol. 147:195-197). A general global alignment technique that can be used is, for example, the Needleman-Wunsch algorithm (Needleman, SB & Wunsch, CD (1970) "A general method applicable to the search for similarities in the amino acid sequences of two proteins." J. Mol. Biol. 48:443-453), which is based on dynamic programming.
[0120] More recently, the Fast Optimal Global Sequence Alignment Algorithm (FOGSAA) has been developed, which is said to produce global alignment of nucleic acid and amino acid sequences faster than other optimal global alignment methods such as the Needleman-Wunsch algorithm.In some embodiments, the identity of two polypeptides is determined by aligning two amino acid sequences, calculating the number of identical amino acids, and dividing by the length of one of the amino acid sequences.In some embodiments, the identity of two nucleic acids is determined by aligning two nucleotide sequences, calculating the number of identical nucleotides, and dividing by the length of one of the nucleic acids.
[0121] For alignment of multiple sequences, computer programs such as Clustal Omega (Sievers et al., Mol Syst Biol. 2011 Oct 11;7:539) may be used.
[0122] In a preferred embodiment, a sequence, such as a nucleic acid or amino acid sequence, is found to have a specified percent identity to a reference sequence, such as a sequence disclosed and / or listed in the claims, when sequence identity is determined using the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990 (e.g., the BLAST®, NBLAST®, XBLAST® or Gapped BLAST® programs), modified as described in Karlin and Altschul Proc. Natl. Acad. Sci. USA 90:5873-77, 1993, using the default parameters of each program.
[0123] In some embodiments, a sequence, such as a nucleic acid or amino acid sequence, is found to have a specified percent identity to a reference sequence, such as a sequence disclosed and / or recited in the claims, when sequence identity is determined using the Smith-Waterman algorithm (Smith, TF & Waterman, MS (1981) "Identification of common molecular subsequences." J. Mol. Biol. 147:195-197) or the Needleman-Wunsch algorithm (Needleman, SB & Wunsch, CD (1970) "A general method applicable to the search for similarities in the amino acid sequences of two proteins." J. Mol. Biol. 48:443-453).
[0124] In some embodiments, a sequence, such as a nucleic acid or amino acid sequence, is found to have a specified percent identity to a reference sequence, such as a sequence disclosed and / or recited in the claims, when sequence identity is determined using the Fast Optimal Global Sequence Alignment Algorithm (FOGSAA).
[0125] In some embodiments, a sequence, such as a nucleic acid or amino acid sequence, is found to have a specified percent identity to a reference sequence, such as a sequence disclosed and / or listed in the claims, when sequence identity is determined using Clustal Omega (Sievers et al., Mol Syst Biol. 2011 Oct 11;7:539).
[0126] The functional variant of sesquiterpene synthase and any other protein disclosed in this application is also included in this disclosure.As used in this disclosure, the functional variant of sesquiterpene synthase refers to a sesquiterpene synthase that has a sequence different from that of reference sesquiterpene synthase but maintains at least one activity of reference sesquiterpene synthase partially or completely.In some embodiments, the functional variant of sesquiterpene synthase enhances one or more activities of reference sesquiterpene synthase.For example, the functional variant can bind one or more of the same substrates (e.g., farnesyl diphosphate, or its precursor) or produce one or more of the same products (e.g., alpha-guayene).
[0127] The variant sequence, including functional variants, may be a homologous sequence. Homologous sequences include, but are not limited to, paralogous sequences, orthologous sequences, or sequences resulting from convergent evolution. Paralogous sequences result from gene duplication in the genome of a species, while orthologous sequences diverge after a speciation event. Two different species may have evolved independently, but each may contain a sequence that shares a certain percent identity with a sequence from another species as a result of convergent evolution. A functional homolog of a reference sesquiterpene synthase, as used in this disclosure, partially or completely maintains at least one activity of the reference sesquiterpene synthase. In some embodiments, a functional homolog of a sesquiterpene synthase enhances one or more activities of the reference sesquiterpene synthase. For example, a functional homolog can bind one or more of the same substrates (e.g., farnesyl diphosphate, or precursors thereof) or produce one or more of the same products (e.g., alpha-guayene).
[0128] Functional variants may be variants of naturally occurring sequences. Functional variants can also be created through site-directed mutagenesis of polypeptide coding sequences or by combining domains from different naturally occurring polypeptide coding sequences ("domain swapping"). Techniques for modifying genes encoding functional variants described in this disclosure are known in the art, including, inter alia, directed evolution, site-directed mutagenesis and random mutagenesis techniques, and such techniques can be useful, for example, to increase the specific activity of polypeptides, change substrate specificity, change expression levels, change subcellular location, or modify polypeptide:polypeptide interactions in a desired manner.
[0129] Variants and homologs can be identified by analysis of polynucleotide and polypeptide sequence alignments. For example, variants and homologs of a polynucleotide sequence that encode derivative polypeptides, etc. can be identified by querying a database of polynucleotide or polypeptide sequences.
[0130] Hybridization can also be used to identify functional variants or functional homologs and / or as a measure of homology between two polynucleotide sequences. Polynucleotide sequences encoding any of the polypeptides disclosed in the present application, or portions thereof, can be used as hybridization probes according to standard hybridization techniques. Hybridization of the probe to DNA or RNA from a test source (e.g., mammalian cells) indicates the presence of the related DNA or RNA in the test source. Hybridization conditions are known to those skilled in the art and can be found, for example, in Current Protocols in Molecular Biology, John Wiley & Sons, NY, 6.3.1-6.3.6, 1991. In some embodiments, moderate hybridization conditions include hybridization in 2x sodium chloride / sodium citrate (SSC) at 30°C, followed by washing at 50°C, 1x SSC, 0.1% SDS. In some embodiments, highly stringent conditions comprise hybridization in 6x sodium chloride / sodium citrate (SSC) at 45°C, followed by a wash at 65°C in 0.2x SSC, 0.1% SDS.
[0131] Sequence analysis to identify functional variants or functional homologs may also include BLAST, reciprocal BLAST, or PSI-BLAST analysis of non-redundant databases using related amino acid sequences as reference sequences. Amino acid sequences are inferred from polynucleotide sequences in some cases. In some embodiments, polypeptides with sequence identity greater than 40% can be identified as candidates for further evaluation for suitability for use according to the present disclosure. Amino acid sequence similarity allows for conservative amino acid substitutions, such as replacing one hydrophobic residue with another, or replacing one polar residue with another. If necessary, manual inspection of such candidates may be performed to limit the number of candidates to be further evaluated. Manual inspection can be performed, for example, by selecting candidates that appear to have conserved functional domains.
[0132] In some embodiments, the polypeptide variant (e.g., a variant of a sesquiterpene synthase variant or any protein variant related to the present disclosure) comprises a domain that shares a secondary structure (e.g., an alpha helix, a beta sheet) with a reference polypeptide (e.g., a reference sesquiterpene synthase or any protein variant related to the present disclosure). In some embodiments, the polypeptide variant (e.g., a variant of a sesquiterpene synthase variant or any protein variant related to the present disclosure) shares a tertiary structure with a reference polypeptide (e.g., a reference sesquiterpene synthase or any protein variant related to the present disclosure). In some embodiments, the reference polypeptide is a sesquiterpene synthase comprising the sequence of SEQ ID NO:1. As non-limiting examples, a variant polypeptide may have low primary sequence identity (e.g., less than 80%, less than 75%, less than 70%, less than 65%, less than 60%, less than 55%, less than 50%, less than 45%, less than 40%, less than 35%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, or less than 5% sequence identity) compared to a reference polypeptide, but share one or more secondary structures (e.g., but are not limited to, a loop, an alpha helix, or a beta sheet, or have the same tertiary structure as the reference polypeptide. For example, a loop may be located between a beta sheet and an alpha helix, between two alpha helices, or between two beta sheets. Homology modeling can be used to compare two or more tertiary structures.
[0133] Mutations can be made in nucleotide sequences by various methods known to those skilled in the art.For example, mutations can be made by PCR-directed mutagenesis, by site-specific mutagenesis by the method of Kunkel (Kunkel, Proc. Nat. Acad. Sci. USA 82: 488-492, 1985), by chemical synthesis of the gene encoding the polypeptide, by gene editing tools, or by insertion, such as insertion of a tag (e.g., HIS tag or GFP tag).Mutations can include, for example, substitutions, deletions, additions, insertions, and translocations, which are generated by any method known in the art. Methods for producing mutations can be found in references such as Molecular Cloning: A Laboratory Manual, J. Sambrook, et al., eds., Fourth Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 2012, or Current Protocols in Molecular Biology, FM Ausubel, et al., eds., John Wiley & Sons, Inc., New York, 2010.
[0134] In some embodiments, methods for producing variants include cyclic permutation (Yu and Lutz, Trends Biotechnol. 2011 Jan;29(1):18-25). In cyclic permutation, the linear primary sequence of a polypeptide can be circularized (e.g., by connecting the N-terminal and C-terminal ends of the sequence) and the polypeptide can be truncated ("broken") at different positions. Thus, the linear primary sequence of the new polypeptide can have low sequence identity (e.g., less than 80%, less than 75%, less than 70%, less than 65%, less than 60%, less than 55%, less than 50%, less than 45%, less than 40%, less than 35%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 5%, including all values therebetween) as determined by linear sequence alignment methods (e.g., Clustal Omega or BLAST). However, the topological analysis of two proteins may reveal whether the tertiary structures of two polypeptides are similar or dissimilar. Without being bound to a particular theory, the mutant polypeptides that have similar tertiary structures to the reference polypeptides created through cyclic substitution of the reference polypeptide may share similar functional characteristics (e.g., enzyme activity, enzyme kinetics, substrate specificity or product specificity). In some cases, cyclic substitution may change the secondary, tertiary or quaternary structure, producing proteins with different functional characteristics (e.g., increased or decreased enzyme activity, different substrate specificity or different product specificity). See, for example, Yu and Lutz, Trends Biotechnol. 2011 Jan;29(1):18-25.
[0135] It is to be understood that in a circularly permuted protein, the linear amino acid sequence of the protein is expected to differ from that of a reference protein that is not circularly permuted. However, a person skilled in the art will be able to determine which residues in a circularly permuted protein correspond to those in a reference protein that is not circularly permuted, for example by aligning the sequences and detecting conserved motifs and / or by comparing the protein structures or predicted structures, for example by homology modeling.
[0136] In some embodiments, the algorithm for determining the percent identity between the sequence of interest and the reference sequence described in this application accounts for the presence of cyclic substitution between sequences.The presence of cyclic substitution can be detected using any method known in the art, such as, for example, RASPODOM (Weiner et al., Bioinformatics. 2005 Apr 1;21(7):932-7).In some embodiments, the presence of cyclic substitution is corrected (e.g., the domain in at least one sequence is rearranged) before calculating the percent identity between the sequence of interest and the sequence described in this application.The claims of this application shall be understood to include sequences whose percent identity to the reference sequence is calculated after taking into account possible cyclic substitutions of the sequence.
[0137] Functional variants or functional homologs can be identified using any method known in the art, for example, the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, supra, can be used to identify homologous proteins with known function.
[0138] Putative functional variants or homologs can also be identified by searching for polypeptides with functionally annotated domains. Databases such as Pfam (Sonnhammer et al., Proteins. 1997 Jul;28(3):405-20) can be used to identify polypeptides with specific domains.
[0139] Homology modeling can also be used to identify amino acid residues that can be mutated without affecting function. Non-limiting examples of such methods include the use of position-specific weight matrices (PSSMs) and energy minimization protocols. See, e.g., Stormo et al., Nucleic Acids Res. 1982 May 11;10(9):2997-3011.
[0140] The PSSM may be coupled with the calculation of a Rosetta energy function, which determines the difference between the wild type and a single point mutant. Without being bound to a particular theory, potentially stabilizing mutations are desirable for protein engineering (e.g., producing functional homologs). In some embodiments, potentially stabilizing mutations have a ΔΔG of less than -0.1 (e.g., less than -0.2, less than -0.3, less than -0.35, less than -0.4, less than -0.45, less than -0.5, less than -0.55, less than -0.6, less than -0.65, less than -0.7, less than -0.75, less than -0.8, less than -0.85, less than -0.9, less than -0.95, or less than -1.0) Rosetta Energy Units (Reu). calc See, e.g., Goldenzweig et al., Mol Cell. 2016 Jul 21;63(2):337-346. doi: 10.1016 / j.molcel.2016.06.012.
[0141] In some embodiments, the coding sequence of a sesquiterpene synthase or any protein related to the disclosure corresponds to a reference coding sequence: , 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more mutations at 100 positions. In some embodiments, the coding sequence of a sesquiterpene synthase or any protein coding sequence related to the disclosure has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 108, 109, 109, 110, 111, 112, 113, Mutations include those in 3, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more codons. As would be understood by one of skill in the art, mutations in a codon may or may not change the amino acid encoded by the codon due to the degeneracy of the genetic code. In some embodiments, the one or more mutations in the coding sequence do not change the amino acid sequence of the coding sequence compared to the amino acid sequence of a reference polypeptide.
[0142] In some embodiments, one or more mutations in a sesquiterpene synthase sequence or other recombinant protein sequence related to the present disclosure change the amino acid sequence of the polypeptide compared to the amino acid sequence of a reference polypeptide. In some embodiments, one or more mutations change the amino acid sequence of the recombinant polypeptide compared to the amino acid sequence of the reference polypeptide and change (enhance or decrease) the activity of the polypeptide compared to the reference polypeptide.
[0143] Assays for determining and quantifying the activity of enzymes and / or enzyme variants are described herein and are known in the art.As an example, the activity of enzymes and / or enzyme variants can be determined by incubating purified enzymes or enzyme variants, or extracts from host cells or whole recombinant host organisms producing enzymes or enzyme variants, with appropriate substrates under appropriate conditions, and analyzing the reaction products (e.g., by gas chromatography (GC) or HPLC analysis).Further details regarding the assay of the activity of enzymes and / or enzyme variants and the analysis of the reaction products are provided in the Examples.These assays include producing the enzyme variants in recombinant host cells.
[0144] The activity of any of the enzymes described in this application, including specific activity, can be measured using methods known in the art. As non-limiting examples, the activity of an enzyme can be determined by measuring its substrate specificity, the product produced, the concentration of the product produced, or any combination thereof.
[0145] The term "activity" as used in this disclosure refers to the ability of an enzyme to react with a substrate to provide a target product. The activity of an enzyme can be determined in an activity test via measuring the increase in one or more target products, the decrease in one or more substrates (as well as starting materials), or via measuring a combination of these parameters as a function of time. The "specific activity" of an enzyme as used in this application refers to the amount (e.g., concentration) of a particular product produced for a given amount (e.g., concentration) of enzyme per unit time.
[0146] "Biological activity", as used in this disclosure, refers to any activity that a polypeptide may exhibit, including, but not limited to, enzymatic activity; binding activity to another compound (e.g., binding to another polypeptide, particularly binding to a receptor, or binding to a nucleic acid); inhibitory activity (e.g., enzyme inhibitory activity); activating activity (e.g., enzyme activating activity); or toxic effect. In some embodiments, a functional variant polypeptide exhibits a relevant activity that is at least 10% of the activity of a parent or reference polypeptide.
[0147] In some embodiments, functional variants of the enzymes associated with the present disclosure produce a yield that is superior to that of a reference or parent enzyme (e.g., a wild-type enzyme or a reference enzyme variant). The term "yield" as used in this disclosure refers to grams of recoverable product per gram of feedstock (which can be calculated as a percentage of molar conversion).
[0148] In some embodiments, functional variants of an enzyme related to the present disclosure exhibit altered (eg, increased) productivity compared to a reference or parent enzyme (eg, a wild-type enzyme or a reference enzyme variant).
[0149] As used in this disclosure, the "productivity" of a mutant sesquiterpene synthase refers to the fold increase in the production of a desired product by the mutant sesquiterpene synthase compared to the production of the desired product by a reference or parent enzyme (e.g., a wild-type enzyme or a reference enzyme variant). For example, if the desired product is alpha-guayene, the productivity of the mutant sesquiterpene synthase refers to the fold increase in the production of alpha-guayene by the mutant sesquiterpene synthase compared to the production of alpha-guayene by a reference or parent enzyme (e.g., a wild-type enzyme or a reference enzyme variant).
[0150] In some embodiments, functional variants of the enzymes related to the present disclosure exhibit a target yield superior to the reference or parent enzyme. The term "target yield" refers to grams of recoverable product per gram of feed (which can be calculated as a percentage of molar conversion).
[0151] In some embodiments, functional variants of the enzymes related to the present disclosure exhibit altered (e.g., increased) target productivity compared to a reference or parent enzyme. The term "target productivity" refers to the amount of recoverable target product in grams per liter of fermentation volume per hour of bioconversion reaction time (i.e., time after substrate addition).
[0152] In some embodiments, functional variants of the enzymes related to the present disclosure exhibit an altered target yield index compared to the reference or parent enzyme. The term "target yield index" refers to the ratio of the resulting product concentration in the culture medium to the concentration of the variant / derivative (e.g., purified enzyme or an extract from a recombinant host cell expressing the desired enzyme).
[0153] In some embodiments, functional variants of enzymes associated with the present disclosure exhibit an altered (e.g., increased) fold in enzymatic activity compared to a reference or parent enzyme (e.g., any one of SEQ ID NOs: 1, 3-39, 77-92, or 110-131). In some embodiments, the increase in activity is at least 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 fold or greater.
[0154] In some embodiments, functional variants of the enzymes related to the present disclosure exhibit altered (e.g., increased) target productivity compared to a reference or parent enzyme. The term "target productivity" refers to the amount of recoverable target product in grams per liter of fermentation volume per hour of bioconversion reaction time (i.e., time after substrate addition).
[0155] Those skilled in the art will understand that mutations in sequences encoding recombinant polypeptides may result in conservative amino acid substitutions that provide functionally equivalent variants of the aforementioned polypeptides, such as variants that retain the activity of the polypeptide. As used in this application, a "conservative amino acid substitution" refers to an amino acid substitution that does not alter the relative charge or size characteristics or functional activity of the protein in which the amino acid substitution is made.
[0156] Thus, the term "conservative substitution", as used in this disclosure, means that an amino acid is replaced by another amino acid listed in the same group of the six standard amino acid groups set forth below.
[0157] (1) Hydrophobic (non-polar): Met, Ala, Val, Leu, Ile, Gly, Pro, Trp, Phe; (2) Neutral hydrophilic: Cys, Ser, Thr; Asn, Gln, Tyr; (3) Acidic: Asp, Glu; (4) Basic: His, Lys, Arg; (5) Residues that affect chain orientation: Gly, Pro; (6) Aromatic: Trp, Tyr, Phe.
[0158] For example, replacement of Asp with Glu retains one negative charge in the modified polypeptide. In addition, glycine and proline may be substituted for each other based on their ability to disrupt alpha-helices. Some preferred conservative substitutions within the above six groups are within the following subgroups: (i) Ala, Val, Leu and Ile; (ii) Ser and Thr; (ii) Asn and Gln; (iv) Lys and Arg; and (v) Tyr and Phe. A skilled scientist can easily construct polynucleotide sequences that code for conservatively substituted amino acid variants, taking into account the known genetic code and recombinant and synthetic DNA techniques.
[0159] A "non-conservative substitution" or "non-conservative amino acid exchange," as used herein, is defined as an amino acid being replaced by another amino acid listed in a different group of the six standard amino acid groups (1)-(6) as set forth above. In some embodiments, mutants of enzymes related to the present disclosure are prepared using non-conservative substitutions that alter the biological function of the mutant.
[0160] For ease of reference, the one-letter amino acid symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission are shown below: For reference purposes, the three-letter codes are also provided:
[0161] [Table 2]
[0162] Amino acid changes, such as amino acid substitutions, may be introduced using known protocols of recombinant genetic techniques, including PCR, gene cloning, site-directed mutagenesis of cDNA, transfection of host cells, and in vitro transcription, which can be used to introduce such changes into sequences to generate mutant / derivative enzymes. Mutants containing amino acid changes can be screened for functional activity. In some cases, amino acids are characterized by their R groups (see, for example, Table 3). For example, amino acids may contain non-polar aliphatic R groups, positively charged R groups, negatively charged R groups, non-polar aromatic R groups, or polar uncharged R groups. Non-limiting examples of amino acids containing non-polar aliphatic R groups include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of amino acids containing positively charged R groups include lysine, arginine, and histidine. Non-limiting examples of amino acids containing negatively charged R groups include aspartic acid and glutamic acid. Non-limiting examples of amino acids containing a non-polar aromatic R group include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of amino acids containing a polar, uncharged R group include serine, threonine, cysteine, proline, asparagine, and glutamine.
[0163] Non-limiting examples of functionally equivalent variants of polypeptides may include conservative amino acid substitutions in the amino acid sequences of the proteins disclosed in the present application. Conservative amino acid substitutions include substitutions made between amino acids within the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D. Additional non-limiting examples of conservative amino acid substitutions are provided in Table 3.
[0164] In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more than 20 residues may be altered when preparing a variant polypeptide. In some embodiments, amino acids are replaced with conservative amino acid substitutions.
[0165] [Table 3]
[0166] In some embodiments of the present disclosure, the amino acid at a particular position in a protein may be replaced with an amino acid having a different molecular weight.For example, in some embodiments, the amino acid at a particular position in a protein may be replaced with a "larger" amino acid, where the larger amino acid refers to the amino acid with a larger molecular weight.In other embodiments, the amino acid at a particular position in a protein may be replaced with a "smaller" amino acid, where the smaller amino acid refers to the amino acid with a smaller molecular weight.The amino acids ranked from smallest to largest based on molecular weight are as follows: G, A, S, P, V, T, C, I, L, N, D, E, K, Q, M, H, F, R, Y, and W.
[0167] Amino acid substitutions in the amino acid sequence of a polypeptide to produce a recombinant polypeptide variant having desired properties and / or activity can be made by altering the coding sequence of the polypeptide. Similarly, conservative amino acid substitutions in the amino acid sequence of a polypeptide to produce a functionally equivalent variant of the polypeptide are typically made by altering the coding sequence of a recombinant polypeptide (e.g., a sesquiterpene synthase, or any other protein related to the present disclosure).
[0168] Expression of Nucleic Acids in Host Cells Aspects of the present disclosure relate to recombinant expression of genes encoding proteins, functional modifications and variants thereof, as well as related uses. For example, the methods described in this application can be used to produce alpha-guayene.
[0169] The term "heterologous" with respect to polynucleotides, such as polynucleotides containing genes, is used interchangeably with the terms "exogenous" and "recombinant" and refers to a polynucleotide that is artificially provided to a biological system; a polynucleotide that is modified in a biological system; or a polynucleotide whose expression or regulation is manipulated in a biological system. The heterologous polynucleotide introduced into or expressed in a host cell may be a polynucleotide derived from a different organism or species than the host cell, or may be a synthetic polynucleotide, or may be a polynucleotide that is also endogenously expressed in the same organism or species as the host cell. For example, a polynucleotide that is endogenously expressed in a host cell may be considered heterologous if it does not naturally occur in the host cell; if it is recombinantly expressed either stably or transiently in the host cell; if it is modified in the host cell; if it is selectively edited in the host cell; if it is expressed in a copy number different from the copy number naturally occurring in the host cell; or if it is expressed in a non-natural manner in the host cell, for example, by manipulating the regulatory region that controls the expression of the polynucleotide. In some embodiments, the heterologous polynucleotide is a polynucleotide that is endogenously expressed in a host cell, but its expression is driven by a promoter that does not naturally regulate the expression of the polynucleotide. In other embodiments, the heterologous polynucleotide is a polynucleotide that is endogenously expressed in a host cell, but its expression is driven by a promoter that naturally regulates the expression of the polynucleotide, but the promoter or another regulatory region is modified. In some embodiments, the promoter is activated or suppressed by recombination. For example, gene editing-based technology can be used to regulate the expression of a polynucleotide, such as an endogenous polynucleotide, from a promoter, such as an endogenous promoter. See, for example, Chavez et al., Nat Methods. 2016 Jul; 13(7): 563-567. The heterologous polynucleotide may comprise a wild-type sequence or a mutant sequence when compared to a reference polynucleotide sequence.
[0170] Any of the recombinant polypeptides, such as a nucleic acid encoding a sesquiterpene synthase, or any protein related to the present disclosure, may be incorporated into any suitable vector via any method known in the art. For example, the vector may be an expression vector, including, but not limited to, a viral vector (e.g., a lentiviral, retroviral, adenoviral, or adeno-associated viral vector), any vector suitable for transient expression, any vector suitable for constitutive expression, or any vector suitable for inducible expression (e.g., a galactose-inducible or doxycycline-inducible vector).
[0171] In some embodiments, the vector replicates autonomously in the cell. The vector may contain one or more endonuclease restriction sites that are cleaved by a restriction endonuclease, allowing the nucleic acid containing the gene described in this application to be inserted or ligated to produce a recombinant vector that can replicate in the cell. The vector may be composed of DNA or RNA. Cloning vectors include, but are not limited to, plasmids, fosmids, phagemids, viral genomes, and artificial chromosomes. The term "expression vector" or "expression construct" as used in this application refers to a nucleic acid construct that is recombinantly or synthetically produced with a set of specific nucleic acid elements that allow the transcription of a particular nucleic acid in a host cell, such as a yeast cell. In some embodiments, the nucleic acid sequence of the gene described in this application is inserted into the cloning vector so that it is operably linked to a regulatory sequence, and in some embodiments, is expressed as an RNA transcript. In some embodiments, the vector contains one or more markers, such as selectable markers, for identifying cells that have been transformed or transfected with the recombinant vector. In some embodiments, the nucleic acid sequence of the gene described in this application is codon-optimized. Codon optimization can increase production of a gene product by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%, relative to a non-codon optimized reference sequence, including all values in between.
[0172] A coding sequence and a regulatory sequence are said to be "operably connected" or "operably linked" when the coding sequence and the regulatory sequence are covalently linked and expression or transcription of the coding sequence is under the influence or control of the regulatory sequence. A coding sequence and a regulatory sequence are said to be operably connected or linked if induction of a promoter in the 5' regulatory sequence permits transcription of the coding sequence where the coding sequence is to be translated into a functional protein, and further if the nature of the linkage between the coding sequence and the regulatory sequence does not (1) cause the introduction of a frameshift mutation, (2) interfere with the ability of the promoter region to direct transcription of the coding sequence, or (3) interfere with the ability of the corresponding RNA transcript to be translated into a protein.
[0173] In some embodiments, the nucleic acid encoding any of the proteins described in the present application is under the control of a regulatory sequence (e.g., an enhancer sequence).In some embodiments, the nucleic acid is expressed under the control of a promoter.The promoter may be a natural promoter, for example, the promoter of the gene in its endogenous context, which provides the normal regulation of the expression of the gene.Alternatively, the promoter may be a promoter different from the natural promoter of the gene, for example, the promoter is different from the promoter of the gene in its endogenous context.
[0174] In some embodiments, the promoter is a eukaryotic promoter.Non-limiting examples of eukaryotic promoters include TDH3, PGK1, PKC1, PDC1, TEF1, TEF2, RPL18B, SSA1, TDH2, PYK1, TPI1GAL1, GAL10, GAL7, GAL3, GAL2, MET3, MET25, HXT3, HXT7, ACT1, ADH1, ADH2, CUP1-1, ENO2, and SOD1, as would be known to those skilled in the art (see, for example, Addgene website: blog.addgene.org / plasmids-101-the-promoter-region).In some embodiments, the promoter is a prokaryotic promoter (e.g., bacteriophage or bacterial promoter).Non-limiting examples of bacteriophage promoters include Pls1con, T3, T7, SP6, and PL. Non-limiting examples of bacterial promoters include Pbad, PmgrB, Ptrc2, Plac / ara, Ptac, and Pm.
[0175] In some embodiments, the promoter is an inducible promoter. As used in this application, an "inducible promoter" is a promoter that is controlled by the presence or absence of a molecule. Non-limiting examples of inducible promoters include chemically regulated promoters and physically regulated promoters. In the case of chemically regulated promoters, transcription activity can be regulated by one or more compounds, such as alcohol, tetracycline, galactose, steroids, metals, or other compounds. In the case of physically regulated promoters, transcription activity can be regulated by phenomena such as light or temperature. Non-limiting examples of tetracycline-regulated promoters include anhydrotetracycline (aTc)-responsive promoters and other tetracycline-responsive promoter systems, such as tetracycline repressor protein (tetR), tetracycline operator sequence (tetO) and tetracycline transactivator fusion protein (tTA). Non-limiting examples of steroid-regulated promoters include rat glucocorticoid receptor-based promoters, human estrogen receptor, gaecdysone receptor, and promoters from the steroid / retinoid / thyroid receptor superfamily. Non-limiting examples of metal-regulated promoters include promoters from metallothionein (proteins that bind and sequester metal ions) genes. Non-limiting examples of pathogenesis-regulated promoters include promoters induced by salicylic acid, ethylene, or benzothiadiazole (BTH). Non-limiting examples of temperature / heat-inducible promoters include heat shock promoters. Non-limiting examples of light-regulated promoters include light-responsive promoters from plant cells. In certain embodiments, the inducible promoter is a galactose-inducible promoter. In some embodiments, the inducible promoter is induced by one or more physiological conditions (e.g., pH, temperature, radiation, osmolality, saline concentration gradient, cell surface binding, or one or more exogenous or intrinsic inducer concentrations).Non-limiting examples of exogenous inducers or inducers include amino acids and amino acid analogs, sugars and polysaccharides, nucleic acids, protein transcriptional activators and repressors, cytokines, toxins, petroleum-based compounds, metal-containing compounds, salts, ions, enzyme substrate analogs, hormones, or any combination thereof.
[0176] In some embodiments, the promoter is a constitutive promoter. "Constitutive promoter" as used herein refers to an unregulated promoter that allows continuous transcription of a gene. Non-limiting examples of constitutive promoters include TDH3, PGK1, PKC1, PDC1, TEF1, TEF2, RPL18B, SSA1, TDH2, PYK1, TPI1, HXT3, HXT7, ACT1, ADH1, ADH2, ENO2 and SOD1.
[0177] Other inducible or constitutive promoters known to those of skill in the art are also contemplated.
[0178] Regulatory sequences required for gene expression may vary between species or cell types, but generally include 5' non-transcribed and 5' non-translated sequences involved in the initiation of transcription and translation, respectively, such as, for example, TATA box, capping sequence, CAAT sequence, etc., as necessary. In particular, such 5' non-transcribed regulatory sequences are expected to include a promoter region, including a promoter sequence for transcriptional control of an operably linked gene. Regulatory sequences may also include enhancer sequences or upstream activator sequences. Vectors may include 5' leader or signal sequences. Regulatory sequences may also include terminator sequences. In some embodiments, terminator sequences provide gene termination to DNA during transcription. The selection and design of one or more suitable vectors suitable for directing expression of one or more genes described in this application in a host cell is within the ability and discretion of one of ordinary skill in the art.
[0179] Expression vectors containing the necessary elements for expression are commercially available and known to those of skill in the art (see, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, Fourth Edition, Cold Spring Harbor Laboratory Press, 2012).
[0180] In some embodiments, introduction of a polynucleotide, such as a polynucleotide encoding a recombinant polypeptide, into a host cell results in genomic integration of the polynucleotide. In some embodiments, a host cell has in its genome at least 1 copy, at least 2 copies, at least 3 copies, at least 4 copies, at least 5 copies, at least 6 copies, at least 7 copies, at least 8 copies, at least 9 copies, at least 10 copies, at least 11 copies, at least 12 copies, at least 13 copies, at least 14 copies, at least 15 copies, at least 16 copies, at least 17 copies, at least 18 copies, at least 19 copies, at least 20 copies, at least 21 copies, at least 22 copies, at least 23 copies, at least 24 copies, at least 25 copies, at least 26 copies, at least 27 copies, at least 28 copies, at least 29 copies, at least 30 copies, at least 31 copies, at least 32 copies, at least 33 copies, at least 34 copies, at least 35 copies, at least 36 copies, at least 37 copies, at least 38 copies, at least 39 copies, at least 40 copies, at least 41 copies, at least 42 copies, at least 43 copies, at least 44 copies, at least 45 copies, at least 46 copies, at least 47 copies, at least 48 copies, at least 49 copies, at least 50 copies, at least 51 copies, at least 52 copies, at least 53 copies, at least 54 copies, at least 55 copies, at least 56 copies, at least 57 copies, at least 58 copies, at least 59 copies, at least 60 copies, at least 61 copies, at least 62 copies, at least 63 copies, at least 64 copies, at least 65 copies, at least 66 copies, at least 67 copies, at least 68 At least 25 copies, at least 26 copies, at least 27 copies, at least 28 copies, at least 29 copies, at least 30 copies, at least 31 copies, at least 32 copies, at least 33 copies, at least 34 copies, at least 35 copies, at least 36 copies, at least 37 copies, at least 38 copies, at least 39 copies, at least 40 copies, at least 41 copies, at least 42 copies, at least 43 copies, at least 44 copies, at least 45 copies, at least 46 copies, at least 47 copies, at least 48 copies, at least 49 copies, at least 50 copies, at least 60 copies, at least 70 copies, at least 80 copies, at least 90 copies, at least 100 copies, or more copies, which may be inserted at the same locus or at different loci of a recombinant host cell of the disclosure.
[0181] In some embodiments, the polynucleotide encoding the sesquiterpene synthase comprises a sequence that is at least 50% identical (e.g., at least 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, including all values therebetween) to any one of SEQ ID NOs: 40-76, 93-109 or 133-154. In certain embodiments, the polynucleotide encoding the sesquiterpene synthase comprises any one of SEQ ID NOs: 40-76, 93-109 or 133-154. In certain embodiments, the polynucleotide encoding the sesquiterpene synthase consists or consists essentially of any one of SEQ ID NOs: 40-76, 93-109, 133-154.
[0182] host cell Any of the proteins of the present disclosure can be expressed in a host cell. The term "host cell" as used in this application refers to a cell that can be used to express polynucleotides, such as polynucleotides encoding proteins used in the production of alpha-guayene and their precursors.
[0183] Any suitable host cell, including eukaryotic or prokaryotic cells, can be used to express any of the recombinant polypeptides, including sesquiterpene synthases and other proteins disclosed in the present application. Suitable host cells include, but are not limited to, fungal cells (e.g., yeast cells), bacterial cells (e.g., E. coli cells), algal cells, plant cells, insect cells, and animal cells, including mammalian cells.
[0184] Suitable yeast host cells include, but are not limited to, Candida, Hansenula, Saccharomyces, Schizosaccharomyces, Pichia, Kluyveromyces, and Yarrowia. In some embodiments, the yeast cell is selected from the group consisting of Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces carlsbergensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia finlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, and the like. membranaefaciens, Pichia opuntiae, Pichia pastoris, Pichia pseudopastoris, Pichia membranifaciens, Komagataella pseudopastoris, Komagataella pastoris, Komagataella kurtzmanii, Komagataella mondaviorum, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia piedperi pijperi, Pichia stipitis, Pichia methanolicaIn some embodiments, the yeast strain is an industrial polyploid yeast strain. Other non-limiting examples of fungal cells include cells obtained from Aspergillus species, Penicillium species, Fusarium species, Rhizopus species, Acremonium species, Neurospora species, Sordaria species, Magnaporthe species, Sphaerotheca species, Uproarium species, Botrytis species, and Trichoderma species.
[0185] In some embodiments, the host cell is a Saccharomyces cell, such as a S. cerevisiae cell.
[0186] In certain embodiments, the host cells are algal cells, such as Chlamydomonas (eg, C. reinhardtii) and Phormidium (P. sp. ATCC29409).
[0187] In other embodiments, the host cell is a prokaryotic cell. Suitable prokaryotic cells include gram-positive, gram-negative, and gram-variable bacterial cells. Host cells include, but are not limited to, Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter, Acidothermus, Arthrobacter, Azobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Campestris, Campylobacter ... Camplyobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter, Erwinia, Fusobacterium, Faecalis, Frankicera, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Iriobicus Ilyobacter, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Mycobacterium, Neisseria, Pantoea, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodopseudomonas, Roseburia, Rhodospirillum, Rhodococcus, Scenedesmus, Streptomyces, The bacteria may be of the genera Streptococcus, Synecoccus, Saccharomonospora, Saccharopolyspora, Staphylococcus, Serratia, Salmonella, Shigella, Thermoanaerobacterium, Tropherima, Tularensis, Temecula, Thermosynecococcus, Thermococcus, Ureaplasma, Xanthomonas, Xylella, Yersinia, and Zymomonas.
[0188] In some embodiments, the bacterial host cell is an Agrobacterium spp. (e.g., A. radiobacter, A. rhizogenes, A. rubi), Arthrobacter spp. (e.g., A. aurescens, A. citreus, A. globformis, A. hydrocarboglutamicus, A. hydr. ocarboglutamicus, A. mysorens, A. nicotianae, A. paraffineus, A. protophonniae, A. roseoparaffinus, A. sulfureus, A. ureafaciens], or Bacillus Bacillus species [e.g., B. thuringiensis, B. anthracis, B. megaterium, B. subtilis, B. lentus, B. circulans, B. pumilus, B. lautus, B. coagulans, B. brevis, B. brevis ... In certain embodiments, the host cell is an industrial Bacillus strain, including, but not limited to, Bacillus subtilis, B. pumilus, B. licheniformis, B. megaterium, B. clausii, B. stearothermophilus, B. halodurans, and B. amyloliquefaciens.In some embodiments, the host cell is an industrial Clostridium species (e.g., C. acetobutylicum, C. tetani E88, C. lituseburense, C. saccharobutylicum, C. perfringens, C. beijerinckii). In some embodiments, the host cell is an industrial Corynebacterium species (e.g., C. glutamicum, C. acetoacidophilum). In some embodiments, the host cell is an industrial Escherichia species (e.g., E. coli). In some embodiments, the host cell is an industrial Erwinia species (e.g., E. uredovora, E. carotovora, E. ananas, E. herbicola, E. punctata, E. terreus). In some embodiments, the host cell is an industrial Pantoea species (e.g., P. citrea, P. agglomerans). In some embodiments, the host cell is an industrial Pseudomonas species (e.g., P. putida, P. aeruginosa, P. mevalonii). In some embodiments, the host cell is an industrial Streptococcus species (eg, S. equisimiles, S. pyogenes, S. uberis).In some embodiments, the host cell is an industrial Streptomyces species (e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus, S. fungidicus, S. griseus, S. lividans). In some embodiments, the host cell is an industrial Zymomonas species (e.g., Z. mobilis, Z. lipolytica).
[0189] The present disclosure is also suitable for use with a variety of animal cell types, such as mammalian cells, such as human (e.g., 293, HeLa, WI38, PER.C6, and Bowes melanoma cells), mouse (e.g., 3T3, NS0, NS1, Sp2 / 0, etc.), hamster (CHO, BHK), monkey (COS, FRhL, Vero), and hybridoma cell lines.
[0190] The present disclosure is also suitable for use with a variety of plant cell types.
[0191] In various embodiments, strains that can be used in the practice of the present disclosure, including both prokaryotic and eukaryotic strains, are readily available to the public from a number of microbial collections, such as the American Type Culture Collection (ATCC), Deutsche Sammlung von Mikroorganismen and Zellkulturen GmbH (DSM), Centraalbureau Voor Schimmelcultures (CBS), and the Agricultural Research Service Patent Culture Collection, Northern Regional Research Center (NRRL).
[0192] The term "cell" as used in this application may refer to a single cell or a population of cells, such as a population of cells that belong to the same cell line or strain. Use of the singular term "cell" should not be construed as explicitly referring to a single cell rather than a population of cells.
[0193] A vector encoding any of the recombinant polypeptides described in this application can be introduced into a suitable host cell using any method known in the art. For example, a non-limiting example of a yeast transformation protocol is described in Gietz et al., Yeast transformation can be conducted by the LiAc / SS Carrier DNA / PEG method. Methods Mol Biol. 2006;313:107-20, which is incorporated by reference in its entirety. The host cell can be cultured under any suitable conditions, as would be understood by one skilled in the art. For example, any medium, temperature, and incubation condition known in the art can be used. In the case of a host cell carrying an inducible vector, the cell can be cultured with an appropriate inducer to promote expression.
[0194] Any of the cells disclosed in this application can be cultured in any type (rich or minimal) and any composition of medium before, during, and / or after contact and / or incorporation of nucleic acid. The term "medium" refers synonymously to culture medium or cultivation medium or fermentation medium. Typically, culture medium contains components essential or beneficial for the maintenance and / or growth of cells, such as carbon source or carbon substrate, nitrogen source, such as peptone, yeast extract, meat extract, malt extract, urea, ammonium sulfate, ammonium chloride, ammonium nitrate and ammonium phosphate; phosphorus source, such as monopotassium phosphate or dipotassium phosphate; trace elements (e.g., metal salts), such as magnesium salts, cobalt salts and / or manganese salts; and growth factors, such as amino acids, vitamins, growth promoters, etc. The term "carbon source" or "carbon substrate" or "carbon source", according to the present disclosure, means any source of carbon that can be used by one skilled in the art to support normal growth of cells, examples of which include hexoses (e.g., glucose, galactose or lactose), pentoses, monosaccharides, oligosaccharides, disaccharides (e.g., sucrose, cellobiose or maltose), molasses, starch or derivatives thereof, cellulose, hemicellulose and combinations thereof.
[0195] The conditions of culture or the culture process can be optimized through routine experimentation, as would be expected by one skilled in the art to understand.In some embodiments, the selected medium is supplemented with various components.In some embodiments, the concentration and amount of the supplemented components are optimized.In some embodiments, other aspects of the medium and growth conditions (e.g., pH, temperature, etc.) are optimized through routine experimentation.In some embodiments, the frequency of supplementing the medium with one or more supplemented components and the time of culturing cells are optimized.
[0196] The culturing of cells described in this application can be carried out in culture vessels known and used in the art. In some embodiments, an aerated reaction vessel (e.g., a stirred tank reactor) is used to culture the cells. In some embodiments, a bioreactor or fermenter is used to culture the cells. Thus, in some embodiments, fermenting cells are used. The terms "bioreactor" and "fermenter" as used in this application are used interchangeably and refer to an enclosure or partial enclosure in which biological, biochemical and / or chemical reactions involving organisms, parts of organisms, or purified proteins take place. A "large-scale bioreactor" or "industrial-scale bioreactor" is a bioreactor used to produce products on a commercial or semi-commercial scale. Large-scale bioreactors typically have volumes in the range of several liters, hundreds of liters, thousands of liters, or more.
[0197] Non-limiting examples of bioreactors include stirred tank fermenters, bioreactors agitated by a rotary mixing device, chemostats, bioreactors agitated by a shaking device, airlift fermenters, packed bed reactors, fixed bed reactors, fluidized bed bioreactors, bioreactors employing wave-induced agitation, centrifugal bioreactors, roller bottles, and hollow fiber bioreactors, roller apparatus (e.g., benchtop, cart-mounted, and / or automated varieties), vertically stacked plates, spinner flasks, stirred or rocking flasks, shaken multiwell plates, MD bottles, T-flasks, Roux bottles, tissue culture propagators with multiple surfaces, modified fermenters, and coated beads (e.g., beads coated with serum proteins, nitrocellulose, or carboxymethylcellulose to prevent cell attachment).
[0198] In some embodiments, the bioreactor includes a cell culture system in which the cells (e.g., yeast cells) are in contact with moving liquid and / or air bubbles. In some embodiments, the cells or cell cultures are grown in the form of a suspension. In other embodiments, the cells or cell cultures are attached to a solid support. Non-limiting examples of support systems include microcarriers (e.g., polymer spheres, microbeads, and microdisks, which may be porous or non-porous), crosslinked beads (e.g., dextran) charged with specific chemical groups (e.g., tertiary amine groups), 2D microcarriers containing cells trapped in non-porous polymer fibers, 3D supports (e.g., support fibers, hollow fibers, multi-cartridge reactors, and semi-permeable membranes that may include porous fibers), microcarriers with low ion exchange capacity, encapsulation cells, capillaries, and aggregates. In some embodiments, the supports are fabricated from materials such as dextran, gelatin, glass, or cellulose.
[0199] In some embodiments, the industrial scale process is operated in a continuous, semi-continuous or discontinuous mode. Non-limiting examples of operation modes are batch, fed-batch, expanded batch, repeated batch, withdrawal / replenishment, rotating wall, spinning flask, and / or perfusion mode of operation. In some embodiments, the bioreactor allows for continuous or semi-continuous replenishment of substrate stock, e.g., carbohydrate source, and / or continuous or semi-continuous separation of product from the bioreactor.
[0200] In some embodiments, the bioreactor or fermenter includes sensors and / or control systems for measuring and / or adjusting reaction parameters. Non-limiting examples of reaction parameters include biological parameters (e.g., growth rate, cell size, cell number, cell density, cell type, or cell status), chemical parameters (e.g., pH, redox potential, concentration of reaction substrates and / or products, concentration of dissolved gases, such as oxygen concentration and CO2, etc.).2 These parameters include: concentration, nutrient concentration, metabolite concentration, oligopeptide concentration, amino acid concentration, vitamin concentration, hormone concentration, additive concentration, serum concentration, ionic strength, ion concentration, relative humidity, molar concentration, osmolarity, other chemical concentrations such as buffers, adjuvants, or reaction by-products), physical / mechanical parameters (e.g., density, conductivity, degree of agitation, pressure, and flow rate, shear stress, shear rate, viscosity, color, turbidity, light absorption, mixing rate, conversion rate, as well as thermodynamic parameters such as temperature, light intensity / quality, etc.). Sensors for measuring the parameters described herein are well known to those skilled in the relevant mechanical and electronic arts. Control systems for adjusting parameters in a bioreactor based on input from the sensors described herein are well known to those skilled in the art of bioreactor engineering.
[0201] In some embodiments, the method includes batch fermentation (e.g., shake flask fermentation). General requirements for batch fermentation (e.g., shake flask fermentation) include oxygen and glucose levels. For example, batch fermentation (e.g., shake flask fermentation) may be oxygen or glucose limited, and therefore in some embodiments, the ability of a strain to perform in a well-designed fed-batch fermentation is underestimated. In addition, the end product (e.g., alpha-guayene) may show some differences from the substrate (e.g., farnesyl diphosphate) in terms of solubility, toxicity, cellular accumulation and secretion, and in some embodiments, may have different fermentation kinetics.
[0202] An embodiment of the present disclosure provides a method for increasing the production of a compound of interest, such as alpha-guayene, in a host cell by increasing sesquiterpene synthase activity by introducing one or more mutations described in this disclosure into the sesquiterpene synthase. The methods described in this application encompass the production of alpha-guayene using a host cell, a cell lysate, or an isolated recombinant polypeptide (e.g., sesquiterpene synthase, and any other protein related to this disclosure).
[0203] Alpha-guayene produced by any of the recombinant cells disclosed in the present application can be identified or extracted using any method known in the art. A non-limiting example of a method for identifying and / or quantifying the compound of interest is GC (e.g., GC-MS, GC-FID).
[0204] composition Further aspects of the present disclosure relate to compositions containing alpha-guayene. Cultivation of host cells related to the present disclosure can result in compositions containing sesquiterpene products such as alpha-guayene. In some embodiments, the compositions obtained by culturing host cells related to the present disclosure contain at least 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 108%, 109%, 109%, 102%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, This results in a composition that is 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% alpha-guayene.
[0205] In some embodiments, a composition obtained by culturing a host cell associated with the present disclosure results in a composition in which about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49% or 50% of the total sesquiterpene products in the composition are delta-guayene.
[0206] The composition related to the present disclosure may further comprise additional components as would be understood by a person skilled in the art.For example, it is understood that in some embodiments, the composition comprising alpha-guayene may comprise fermentation broth or cell culture supernatant of cell culture.In other embodiments, the composition may comprise alpha-guayene in purified form from fermentation broth or cell culture supernatant of cell culture.
[0207] The ratio of alpha-guayene to delta-guayene can be, for example, in the range of about 60:40 to about 99: 1. For example, the weight ratio of alpha-guayene to delta-guayene can be in the range of about 65:35 to about 99: 1, about 70:30 to about 99: 1, about 75:25 to about 99: 1, about 80:20 to about 99: 1, about 85:15 to about 99: 1, about 90:10 to about 99: 1, or about 95:5 to about 99: 1. For example, the weight ratio of alpha-guayene to delta-guayene can be in the range of about 65:35 to about 98: 2, about 70:30 to about 97: 3, about 75:25 to about 96: 4, about 80:20 to about 95: 5, or about 85:15 to about 90: 10.
[0208] In some embodiments, the cells related to the present invention are cultured in the presence of an organic solvent overlay. Organic solvent overlay, as used in this disclosure, refers to a layer containing one or more organic solvents added to a cell culture sample. The organic solvent overlay may partially or completely cover the cell culture sample. The use of an organic solvent overlay can help reduce or mitigate the toxicity of host cells caused by increased product concentration. In some embodiments, the composition comprising alpha-guayene further comprises one or more components of an organic solvent overlay (e.g., dodecane).
[0209] The phraseology and terminology used in this application are for the purpose of description and should not be regarded as limiting. The use in this application of terms such as "including," "comprising," "having," "containing," "involving," and / or variations thereof is meant to encompass additional items in addition to the items listed thereafter and equivalents thereof.
[0210] The present invention is further illustrated by the following examples, which should not be construed as further limiting in any way. The entire contents of all reference documents cited throughout this application (e.g., literature references, issued patents, published patent applications, and co-pending patent applications) are hereby expressly incorporated by reference. EXAMPLES
[0211] [Example 1] Identification of mutant sesquiterpene synthases producing increased alpha-guayene This example describes the identification of mutant sesquiterpene synthases capable of producing increased amounts of alpha-guayene relative to that produced by the parent delta-guayene synthase 2 protein from A. crassuna (UniProt Accession No. D0VMR7; SEQ ID NO:1).
[0212] Metagenomic libraries of putative sesquiterpene synthases were obtained from public and personal sequence databases. After codon optimization for yeast, sequences were screened for alpha-guayene production using a high-throughput GC-MS assay. Of the approximately 500 sesquiterpene synthases designed and tested, none produced alpha-guayene in greater than 15% abundance, as determined using extracted ion chromatogram (EIC) m / z fragments on sesquiterpene synthase products detected via GC-MS utilizing an authenticated alpha-guayene standard for identification.
[0213] Because the metagenomic screen did not identify any putative sesquiterpene synthases that exhibited increased alpha-guayene production, a second approach to identify sesquiterpene synthases capable of producing increased amounts of alpha-guayene was pursued, which involved generating sequence variants of the delta-guayene synthase 2 protein from A. crassuna (UniProt Accession No. D0VMR7; SEQ ID NO: 1).
[0214] Mutant sesquiterpene synthases were generated by PCR mutagenesis and synthesized in a yeast expression vector flanked by a GAL1 promoter and terminator. The plasmid containing the mutant sesquiterpene synthase DNA coding sequence was transformed into a screening strain for in vivo evaluation of the enzyme's performance. The screening strain was engineered to express higher concentrations of farnesyl diphosphate (FPP), which serves as a substrate for the enzyme being screened. If the screening strain can utilize galactose, within the screening strain, the sesquiterpene synthase will convert FPP. The GAL1 promoter is capable of responding to galactose, inducing mutant enzyme expression in the presence of galactose. After sufficient exposure to galactose, the strain was harvested and exposed to an organic solvent, which was used to isolate the reaction products. Once isolated, the reaction products were analyzed by GC to determine the relative occurrence, ratio, and identity of the products in the reaction product mixture.
[0215] Table 4 and Figure 2 show the top performing mutant sesquiterpene synthases identified from the GC-MS screen based on percent alpha-guayene produced. GC-MS assays were performed on a Thermo ISQ GC-MS with a mass scan range of 40-350 m / z and the ion source set at 250 °C. The front inlet was set at 250 °C with a He flow of 1.5 mL / min and a split flow ratio of 4. The GC ramp was started from 80 °C to 200 °C at a rate of 7.5 °C / min, followed by another 20 °C ramp to a final hold temperature of 280 °C. Mzmine2 and thermo xcalibur software were utilized for reported peak detection, spectra / product identification, and peak areas.
[0216] The values of percent alpha-guayene (%) and percent delta-guayene (%) represent the estimated product percentages of the total sesquiterpene synthase product profile detected by the assay. The percentage of each sesquiterpene detected was calculated using the extracted ion chromatogram (EIC) peak area based on the m / z fragment with the most similar peak area response (based on certified sesquiterpene standards and NIST spectral library). For example, WT (no mutation compared to SEQ ID NO: 1) produced peak areas of alpha-guayene (105 m / z), delta-guayene (107 m / z), aciphyllene* (105 m / z), beta-elemene (81 m / z), humulene* (93 m / z), and neointermedeol* (81 m / z) for the calculation of percent alpha-guayene (* denotes predicted product via NIST spectral library). The fold increase in alpha-guayene production was calculated as the amount of alpha-guayene product produced by the host cells expressing the mutant sesquiterpene synthase relative to the amount of alpha-guayene product produced by the host cells expressing a control sesquiterpene synthase comprising SEQ ID NO: 1. Thus, in Table 4, "productivity" refers to the fold increase in alpha-guayene production due to the mutant enzyme relative to the alpha-guayene production due to the wild-type sesquiterpene synthase, % alpha-guayene refers to the amount of alpha-guayene relative to the total sesquiterpenes produced by the host cells expressing the mutant enzyme, and % delta-guayene refers to the amount of delta-guayene relative to the total sesquiterpenes produced by the mutant enzyme.
[0217] [Table 4-1] [Table 4-2] [Table 4-3]
[0218] Surprisingly, as shown in Table 4 and FIG. 2, a mutant sesquiterpene synthase was identified that produced 50% more alpha-guayene, in contrast to the delta-guayene synthase 2 protein from A. crassuna (UniProt Accession No. D0VMR7; SEQ ID NO: 1), which was found to produce 14.6% alpha-guayene.
[0219] Following identification of mutant sesquiterpene synthases that produced increased amounts of alpha-guayene, a linear model was used to predict which specific amino acid substitution mutations contributed most to the apparent increase in alpha-guayene production. After fitting the linear model, positive and large coefficients in the model associated with specific amino acid substitution mutations were identified, signifying increased statistical correlation with alpha-guayene production. Table 5 shows the amino acid substitutions that are statistically correlated with alpha-guayene production based on this analysis.
[0220] [Table 5]
[0221] Without wishing to be bound by any theory, the shape of the active site pocket of a sesquiterpene synthase may affect the ratio of the various products produced by the sesquiterpene synthase. In the reaction catalyzed by the delta-guaiene synthase 2 protein from A. crassuna (UniProt accession number D0VMR7; SEQ ID NO: 1), the final step involves either proton abstraction from the pro-delta carbon, yielding delta-guaiene as the product, or proton abstraction from the pro-alpha carbon, yielding alpha-guaiene as the product. Some mutant sesquiterpene synthases are able to produce increased amounts of alpha-guaiene, at least in part, because they alter the substrate binding mode of the sesquiterpene synthase to allow easier access to the pro-alpha carbon from the catalytic residue Tyr-520. [Example 2]
[0222] Identification of additional mutant sesquiterpene synthases producing increased alpha-guayene To identify additional mutant sesquiterpene synthases capable of producing increased amounts of alpha-guayene, additional mutant versions of the delta-guayene synthase 2 protein from A. crassuna (UniProt Accession No. D0VMR7; SEQ ID NO: 1) were produced and screened.
[0223] Using the method described in Example 1, mutant sesquiterpene synthases were screened for alpha-guayene production in a primary screen. The primary screen employed GC-MS analysis using the method described in O'Maille et al. (2008) Nat. Chem. Biol. 4, 617-623 to identify top performing mutant sesquiterpene synthases. Approximately 1012 mutant sesquiterpene synthases were screened. Nineteen high performing mutant sesquiterpene synthases were identified in the primary screen, including the three mutant sesquiterpene synthases from Example 1. The nineteen mutant sesquiterpene synthases were further characterized in a secondary screen using a GC-MS assay based on the method described in O'Maille et al. (2008) Nat. Chem. Biol. 4, 617-623, but at a slower heating rate than that used in the primary screen. (Table 6 and Figure 3). For the secondary screening GC-MS method, a TG5-MS column (15 m x 0.250 mm x 0.25 μm) on a Thermo GC-MS-ISQ was used with a flow rate of 1.5 mL / min, respectively, and the inlet temperature was set at 250°C. The oven was set with an initial 0.1 min hold at 80.0°C, ramping to 200.0°C at 7.5°C / min, and ramping to 280°C at 20.0°C / min and holding for 1.00 min. The MS transfer line temperature was set at 280°C. The ion source temperature was set at 250°C. The mass range was set at 40-350 amu.
[0224] Tertiary screening was performed on a subset of strains using GC-FID, which provides a quantitative readout with higher accuracy than GC-MS, based on the method described in Greenhagen et al. (2006) Proc. Natl. Acad. Sci. 103 (26) 9826-9831. Results confirmed the identification of mutant sesquiterpene synthases that had increased alpha-guayene production compared to that produced by a positive control strain expressing delta-guayene synthase 2 protein from A. crassuna (UniProt accession number D0VMR7; SEQ ID NO: 1) (Table 6).
[0225] Table 6 shows the results from the primary, secondary, and tertiary screens. "% alpha-guayene (primary screen)" represents the percent of the amount of alpha-guayene product in the total sesquiterpene synthase product profile detected by the primary screening assay. "% alpha-guayene (secondary screen)" represents the percent of the amount of alpha-guayene product in the total sesquiterpene synthase product profile detected by the secondary screening assay. "% alpha-guayene (tertiary screen)" represents the percent of the amount of alpha-guayene product in the total sesquiterpene synthase product profile detected by the tertiary screening assay. "% other sesquiterpenes vs. tertiary screen" represents the percent of the amount of non-alpha-guayene products in the total sesquiterpene synthase product profile detected by the tertiary screening assay.
[0226] [Table 6-1] [Table 6-2]
[0227] As shown in Table 6, surprisingly, a mutant sesquiterpene synthase was identified that produced 60% more alpha-guayene, in contrast to the delta-guayene synthase 2 protein from A. crassuna (UniProt Accession No. D0VMR7; SEQ ID NO: 1), which was found to produce about 13% alpha-guayene. [Example 3]
[0228] Identification of additional mutant sesquiterpene synthases producing increased alpha-guayene To identify additional mutant sesquiterpene synthases capable of producing increased amounts of alpha-guayene, additional mutants of the delta-guayene synthase 2 protein from A. crassuna (UniProt accession number D0VMR7; SEQ ID NO:1) were generated.
[0229] Using the method described in Example 1, mutant sesquiterpene synthases were screened for alpha-guayene production in a primary screen as follows: A yeast expression plasmid vector containing mutant sesquiterpene synthase DNA coding sequence flanked by a GAL1 promoter and terminator was transformed into a screening strain for in vivo evaluation of enzyme performance. The screening strain was engineered to express higher concentrations of farnesyl diphosphate (FPP), which serves as a substrate for the enzyme being screened. In the screening strain, when galactose is available, the sesquiterpene synthase converts FPP. The GAL1 promoter responds to galactose and induces mutant enzyme expression in the presence of galactose. After sufficient exposure to galactose, the strain was harvested and exposed to an organic solvent, which was used to isolate the reaction products. Once isolated, the reaction products were analyzed by GC to determine the relative occurrence, ratio, and identity of the products in the reaction product mixture.
[0230] The primary screen employed GC-MS analysis to identify top performing mutant sesquiterpene synthases using the method described in O'Maille et al. (2008) Nat. Chem. Biol. 4, 617-623. Approximately 5000 mutant sesquiterpene synthases were screened through multiple rounds of screening. 23 high performing mutant sesquiterpene synthases were identified in the primary screen. The 23 mutant sesquiterpene synthases were further characterized in a secondary screen using a GC-MS assay based on the method described in O'Maille et al. (2008) Nat. Chem. Biol. 4, 617-623, but with a longer heating rate than that used in the primary screen. (Table 7 and Figure 4). For the secondary screening GC-MS method, a TG5-MS column (15 m x 0.250 mm x 0.25 μm) on a Thermo GC-MS-ISQ was used with a flow rate of 1.5 mL / min, respectively, and the inlet temperature was set at 250°C. The oven was set with an initial 0.1 min hold at 80.0°C, ramping to 200.0°C at 7.5°C / min, and ramping to 280°C at 20.0°C / min and holding for 1.00 min. The MS transfer line temperature was set at 280°C. The ion source temperature was set at 250°C. The mass range was set at 40-350 amu.
[0231] The results confirmed the identification of a mutant sesquiterpene synthase that had increased alpha-guayene production compared to that produced by a positive control strain expressing the delta-guayene synthase 2 protein from A. crassuna (UniProt accession number D0VMR7; SEQ ID NO:1) (Table 7).
[0232] Table 7 shows the results from the primary and secondary screens. "% alpha-guayene (primary screen)" represents the percent of the amount of alpha-guayene product in the total sesquiterpene synthase product profile detected by the primary screening assay. "% alpha-guayene (secondary screen)" represents the percent of the amount of alpha-guayene product in the total sesquiterpene synthase product profile detected by the secondary screening assay. "% alpha-guayene (tertiary screen)" represents the percent of the amount of alpha-guayene product in the total sesquiterpene synthase product profile detected by the tertiary screening assay. "% other sesquiterpenes vs. tertiary screen" represents the percent of the amount of non-alpha-guayene and non-acifylene products in the total sesquiterpene synthase product profile detected by the tertiary screening assay.
[0233] [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4]
[0234] As shown in Table 7, surprisingly, a mutant sesquiterpene synthase was identified that produced more than 70% alpha-guayene, in contrast to the delta-guayene synthase 2 protein from A. crassuna (UniProt Accession No. D0VMR7; SEQ ID NO: 1), which was found to produce about 15% alpha-guayene. In addition, surprisingly, a mutant sesquiterpene synthase was identified that produced aciphylene as the second most abundant sesquiterpene product.
[0235] Table 8 contains the protein and nucleic acid sequences of sesquiterpene synthase variants relevant to the present disclosure.
[0236] In some embodiments, the sesquiterpene synthase comprises a sequence that is at least 50% identical (e.g., at least 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, including all values in between) to any one of SEQ ID NOs: 3-39, 77-92 or 110-131, the protein sequences in Table 8, or conservatively substituted variants thereof. In certain embodiments, the sesquiterpene synthase comprises the sequence of any one of SEQ ID NOs: 3-39, 77-92, or 110-131, or a conservatively substituted version thereof. In certain embodiments, the sesquiterpene synthase consists of or consists essentially of the sequence of any one of SEQ ID NOs: 3-39, 77-92, 110-131, or a conservatively substituted variant thereof.
[0237] In some embodiments, the polynucleotide encoding the sesquiterpene synthase comprises a sequence that is at least 50% identical (e.g., at least 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, including all values therebetween) to any one of SEQ ID NOs: 40-76, 93-109, 133-154 or any one of the nucleic acid sequences in Table 8. In certain embodiments, the polynucleotide encoding the sesquiterpene synthase comprises any one of SEQ ID NOs: 40-76, 93-109, 133-154. In certain embodiments, the polynucleotide encoding the sesquiterpene synthase consists or consists essentially of any one of SEQ ID NOs: 40-76, 93-109, 133-154.
[0238] [Table 8-1] [Table 8-2]
[0239] Equivalent Those skilled in the art are expected to recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described in this disclosure. Such equivalents are intended to be encompassed by the scope of the following claims. All references, including patent documents, disclosed in this application are incorporated by reference in their entirety, particularly with respect to the disclosures referenced in this application.
[0240] It is to be understood that the sequences disclosed in this application may or may not contain a secretion signal. The sequences disclosed in this application include versions with or without a secretion signal. It should be understood that the amino acid sequences disclosed in this application may or may not be depicted with a start codon (M). The sequences disclosed in this application include versions with or without a start codon. Thus, in some cases, the amino acid numbering may correspond to an amino acid sequence that contains a secretion signal and / or a start codon, while in other instances, the amino acid numbering may correspond to an amino acid sequence that does not contain a secretion signal and / or a start codon. It should also be understood that the sequences disclosed in this application may or may not be depicted with a stop codon. The sequences disclosed in this application include versions with or without a stop codon.
Claims
1. A composition comprising: (a) sesquiterpenes, wherein at least 50% of the sesquiterpenes are alpha-guayene; and (b) one or more additional components comprising a fermentation medium, a cell culture supernatant, and / or a hydrophobic overlay.
2. 2. The composition of claim 1, wherein at least 15% of the sesquiterpenes are aciphyllene.
3. 10. The composition of claim 1, wherein the alpha-guayene is produced using a microbial host cell.
4. 10. The composition of claim 1, wherein alpha-guayene is produced using an in vitro or in vivo system.
5. 2. The composition of claim 1, wherein about 50% to about 90% of the sesquiterpenes are alpha-guayene.
6. The composition of claim 1 further comprising delta-guayene.
7. The composition of claim 1 further comprising acifylene.
8. 7. The composition of claim 6, wherein about 50% to about 100% of the sesquiterpenes are alpha-guayene.
9. 9. The composition of claim 8, wherein at least 15% of the sesquiterpenes are aciphyllene.
10. 10. The composition of claim 1, further comprising one or more non-terpene components or one or more additional terpene components.
11. The composition of claim 10, wherein the one or more non-terpene components comprises FPP.
12. 11. The composition of claim 10, wherein the one or more additional terpene components comprise delta-guayene, beta-guayene, gamma-guayene, germacrene A, acifyllene, and / or alpha-humulene, as determined by GC.
13. 13. A method for producing a composition described in any one of claims 1 to 12, comprising culturing a host cell containing a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sesquiterpene synthase comprises one or more amino acid substitutions compared to SEQ ID NO:1, and at least one of the amino acid substitutions is at a position corresponding to positions 44, 212, 217, 293, 295, 404, 448, 515, and / or 542 in SEQ ID NO:
1.
14. 14. The method of claim 13, wherein the sequence of the sesquiterpene synthase further comprises at least one amino acid substitution at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:
1.
15. 13. A method of making the composition of any one of claims 1 to 12, comprising culturing a host cell containing a heterologous polynucleotide encoding a sesquiterpene synthase, wherein the sequence of the sesquiterpene synthase comprises two or more amino acid substitutions compared to SEQ ID NO: 1; (i) at least one amino acid substitution is at a position corresponding to positions 44, 212, 217, 293, 295, 404, 515, and / or 542 in SEQ ID NO:1; (ii) the at least one amino acid substitution is at a position corresponding to positions 23, 72, 86, 111, 118, 134, 147, 188, 201, 224, 252, 255, 289, 290, 291, 292, 346, 381, 390, 406, 419, 433, 442, 443, 444, 448, 458, 467, 494, 499, 512, 516, and / or 519 in SEQ ID NO:1.