Engineered polypeptides and host cells for biosynthesis of thymohydroquinone (THQ)
Engineered recombinant polypeptides with specific amino acid and codon variations in host cells like Pichia pastoris and Saccharomyces cerevisiae significantly enhance THQ production, addressing the inefficiencies of existing systems by achieving higher yields and activity in THQ biosynthesis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-04-02
AI Technical Summary
Existing recombinant polypeptides and host cell systems are sub-optimal for the commercial bioproduction of thymohydroquinone (THQ) due to low activity and yield, necessitating improved CYP activity and biocatalytic methods for efficient conversion of thymol and carvacrol precursors.
Engineering recombinant polypeptides with specific amino acid and codon variations, integrated into host cells like Pichia pastoris and Saccharomyces cerevisiae, to enhance the conversion of thymol and carvacrol to THQ, achieving increased activity and yield through optimized expression systems and integration sites.
The engineered polypeptides and host cells demonstrate improved THQ production, with yields up to 2-fold or more compared to control systems, enhancing the commercial viability of THQ biosynthesis.
Smart Images

Figure US2025047773_02042026_PF_FP_ABST
Abstract
Description
ENGINEERED POLYPEPTIDES AND HOST CELLS FOR BIOSYNTHESIS OF THYMOHYDROQUINONE (THQ) CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority of US Provisional Application 63 / 698,893, filed September 25, 2024, the entirety of which is hereby incorporated by reference herein. FIELD
[0002] The present disclosure relates to engineered polypeptides with cytochrome P450 (CYP) activity capable of selectively catalyzing the conversion of precursor compounds carvacrol and / or thymol to the product compound, thymohydroquinone (THQ), host cells comprising heterologous nucleic acids encoding the engineered polypeptides, and processes using these host cells for improved fermentative bioproduction of THQ. REFERENCE TO SEQUENCE LISTING
[0003] The official copy of the Sequence Listing is submitted concurrently with the specification via USPTO Patent Center as an WIPO Standard ST.26 formatted XML file with file name “21896-023WO1.xml”, a creation date of September 23, 2025, and a size of 826,541 bytes. This Sequence Listing filed via USPTO Patent Center is part of the specification and is incorporated in its entirety by reference herein. BACKGROUND
[0004] Thymohydroquinone (THQ) is a phenolic terpene that is a minor component in many essential oils, such as oregano oil. THQ acts an antioxidant without disrupting the flavor profile of oregano oil. The THQ precursors, thymol and carvacrol, are abundant in oregano oil but have a strong taste and odor. We have developed a bioconversion process with a P450 / CPR enzyme pair to convert thymol and carvacrol in oregano oil to THQ. Two cytochrome P450 enzymes from oregano, CYP76S40, and CYP736A300, have been identified that can conduct this reaction when expressed heterologously in Nicotiana benthiana or Saccharomyces cerevisiae (see e.g., Krause, et al. “The Biosynthesis of Thymol, Carvacrol, and Thymohydroquinone in Lamiaceae Proceeds via Cytochrome P450s and a Short-Chain Dehydrogenase.” Proceedings of the National Academy of Sciences 118, no.52 (December 28, 2021). These P450 enzymes, however, are sub-optimal for commercial bioproduction of THQ in yeast.
[0005] PCT / US2024 / 020520, filed May 31, 2024, discloses certain recombinant polypeptides from natural sources that have cytochrome P450 (CYP) activity capable of selectively catalyzing the conversion of precursor compounds, carvacrol and / or thymol, to the product compound, thymohydroquinone (THQ). and recombinant host cells that heterologously express genes encoding these polypeptides in yeast. It also discloses ‐ 1 ‐ processes that use these recombinant polypeptides for the fermentative bioproduction of THQ.
[0006] There remains a need for more recombinant polypeptides with CYP activity and recombinant cell systems expressing these polypeptides, along with improved biocatalytic methods that allow for more commercially viable biosynthetic production of THQ. SUMMARY
[0007] The present disclosure relates generally to recombinant polypeptides with cytochrome P450 (CYP) activity capable of selectively catalyzing the conversion of the precursor compounds, thymol and carvacrol, to the product compound, THQ, and the engineering of these polypeptides for heterologous expression in recombinant host cells (e.g., yeast) allowing for improved fermentative bioproduction (e.g., increased activity and / or increased yield) of these product compounds. This summary is intended to introduce the subject matter of the present disclosure, but does not cover each and every embodiment, combination, or variation that is contemplated and described within the present disclosure. Further embodiments are contemplated and described by the disclosure of the detailed description, drawings, and claims.
[0008] In at least one embodiment, the present disclosure provides recombinant host cell comprising a heterologous nucleic acid encoding a polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, L105, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, G258, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, K420, Y457, D458, N472, I477, and T478. In at least one embodiment, the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, L105V, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, G258V, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, K420C, Y457C, D458V, N472D, I477M, I477V, and T478C.
[0009] In at least one embodiment of the recombinant host cell of the present disclosure the amino acid sequence comprises a combination of amino acid differences as compared to SEQ ID NO: 24 selected from: ‐ 2 ‐ T196A, M464G S27R, H51R, L99A, I102L, N119S, A176M, T196A, D415R, M464G R,‐ 3 ‐ S27R, G35R, H51R, L99A, I102L, N119S, A176M, T196A, S299A, D415R, K420C, M464G, N472D S27R G35R H51R L99A I102L N119S A176M T196A S299A D415R M464Ghe heterologous nucleic acid sequence has at least 80% identity to SEQ ID NO: 23, and a neutral codon difference as compared to SEQ ID NO: 23 at one or more positions encoding an amino acid residue selected from: S21, V26, L53, S111, V112, G143, A171, T174, S210, I211, P214, L246, D289, D296, S299, Q333, D339, P358, P365, A368, D371, P437, and D458. In at least one embodiment, the neutral codon differences are selected from: S21S (AGC), V26V (GTC), L53L (CTA), S111S (AGT), V112V (GTG), G143G (GGC), A171A (GCA), T174T (ACG), S210S (TCG), I211I (ATC), P214P (CCC), L246L (CTG), D289D (GAC), D296D (GAC), S299S (TCG), Q333Q (CAG), D339D (GAC), P358P (CCG), P365P (CCT), A368A (GCA), D371D (GAC), P437P (CCT), and D458D (GAC).
[0011] In at least one embodiment of the recombinant host cell of the present disclosure the heterologous nucleic acid sequence comprises a combination of neutral codon difference as compared to SEQ ID NO: 23 selected from: S111S (AGT), P358P (CCG), P437P (CCT) S21S (AGC), S111S (AGT), P358P (CCG), P437P (CCT)
[0012] In at least one embodiment of the recombinant host cell of the present disclosure: (a) the polypeptide with CYP activity comprises an amino acid sequence having at ‐ 4 ‐ least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206; or (b) the polypeptide with CYP activity comprises an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206.
[0013] In at least one embodiment of the recombinant host cell of the present disclosure: (a) the heterologous nucleic acid comprises a sequence of at least 80% identity, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205; or (b) the heterologous nucleic acid comprises a codon degenerate sequence of a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205.
[0014] In at least one embodiment of the recombinant host cell of the present disclosure, the heterologous nucleic acid further comprises a sequence encoding a polypeptide with CPR activity; optionally, wherein the polypeptide with CPR activity comprises an amino acid sequence of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to SEQ ID NO: 6.
[0015] In at least one embodiment of the recombinant host cell of the present disclosure, the heterologous nucleic acid: is under the control of a promoter system selected from pGal1 / 10, and pCAT1:pFDH. ‐ 5 ‐
[0016] In at least one embodiment of the recombinant host cell of the present disclosure, the source organism of the host cell is selected from Pichia pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, and Escherichia coli.
[0017] In at least one embodiment of the recombinant host cell of the present disclosure, heterologous nucleic acid is integrated at a site in the host cell genome.
[0018] In at least one embodiment of the recombinant host cell of the present disclosure, the recombinant host cell is Pichia pastoris and the heterologous nucleic acid is integrated at one or more sites in the genome selected from AOX1, Int6, Int15, and HIS4; optionally, the heterologous nucleic acid is integrated at two, three or more sites.
[0019] In at least one embodiment of the recombinant host cell of the present disclosure, the recombinant host cell is Saccharomyces cerevisiae and the heterologous nucleic acid is integrated at one or more sites in the genome selected from X-2, X-4, XI-2, XII-4, NDE1, XII- 5, Gal80, and ROQ1; optionally, wherein the heterologous nucleic acid is integrated at two, three, or more sites.
[0020] In at least one embodiment, the recombinant host cell of the present disclosure comprising a heterologous nucleic acid encoding a polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions, exhibits improved bioconversion of carvacrol and / or thymol to THQ relative to a control recombinant host cell comprising a heterologous nucleic acid encoding a polypeptide of SEQ ID NO: 24. In at least one embodiment, the improved bioconversion exhibited by the host cell comprises an increased yield of THQ relative to the control host cell of at least 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 4-fold, 5-fold, 10-fold, or more.
[0021] In another aspect, the present disclosure provides a recombinant polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, L105, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, G258, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, K420, Y457, D458, N472, I477, and T478. In at least one embodiment, the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, L105V, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, G258V, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, ‐ 6 ‐ N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, K420C, Y457C, D458V, N472D, I477M, I477V, and T478C.
[0022] In at least one embodiment of the recombinant polypeptide, the amino acid sequence comprises a combination of amino acid differences as compared to SEQ ID NO: 24 selected from: T196A, M464G S27R, H51R, L99A, I102L, N119S, A176M, T196A, D415R, M464G R,‐ 7 ‐ G35R, M132V R159W, S344P, N416D, S27R, H51R, L99A, I102L, N119S, A176M, T196A, D415R, M464G K59E Y60F VD74D L84P K151E S344P S396T I477V S27R H51R L99A I102Lre: (a) the polypeptide comprises an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206; or (b) the polypeptide comprises an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206.
[0024] In at least one embodiment of the recombinant polypeptide of the present disclosure the polypeptide is fused via a linker to a second polypeptide; optionally, wherein the second polypeptide has CPR activity. In at least one embodiment, the second polypeptide with CPR activity comprises an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from SEQ ID NO: 6; optionally, wherein the polypeptide comprises an amino acid sequence of SEQ ID NO: 6.
[0025] In another aspect, the present disclosure provides a recombinant polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and ‐ 8 ‐ an amino acid difference as compared to SEQ ID NO: 24 at one or more positions exhibits increased activity in the conversion of carvacrol and / or thymol to THQ relative to a control polypeptide of SEQ ID NO: 24. In at least one embodiment, the increased exhibited by the recombinant polypeptide comprises an increased activity in the conversion of carvacrol and / or thymol to THQ relative to the polypeptide of SEQ ID NO: 24 of at least 1.2-fold, 1.3- fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 4-fold, 5-fold, 10-fold, or more.
[0026] In another aspect, the present disclosure provides a polynucleotide encoding a recombinant polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ of the present disclosure. In at least one embodiment, the polynucleotide comprises a sequence of at least 80% identity, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205.
[0027] In another aspect, the present disclosure provides an expression vector comprising the polynucleotide comprising a sequence of at least 80% identity to a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205. In at least one embodiment, the expression vector further comprises a polynucleotide sequence of at least 80% identity, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to SEQ ID NO: 5. In at least one embodiment, the expression vector can further comprise a control sequence; optionally, wherein the control sequence comprises a promoter system selected from pGal1 / 10, and pCAT1:pFDH1.
[0028] In another aspect, the present disclosure provides composition, wherein the composition comprises: (a) a recombinant host cell of the present disclosure, or a recombinant polypeptide of the present disclosure; and (b) compound (2a) and / or compound (2b), or a derivative of compound (2a) and / or a derivative of compound (2b) ‐ 9 ‐
[0029] In at least one e ises a recombinant polypeptide with CPR activity; optionally, a polypeptide comprising an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 6. In at least one embodiment, the composition comprises a derivative of compound (2a) and / or a derivative of compound (2b), wherein the derivative is selected from compound (2c), compound (2d), compound (2e), compound (2f), compound (2g), compound (2h), compound (2i), compound (2j), compound (2k), compound (2l), compound (2m), compound (2n), compound (2o), compound (2p), and mixture thereof. In at least one embodiment, the composition is in aqueous solution.
[0030] In another aspect, the present disclosure provides a method for producing THQ comprising: (a) culturing a recombinant host cell of the present disclosure in a suitable medium comprising carvacrol and / or thymol; and (b) recovering the produced THQ.
[0031] In another aspect, the present disclosure provides a method for preparing compound (1a) or a derivative of compound (1a) comprising contacting under a compound (2a) and / orcompound (2b) or a derivative of compound (2a) and / or a derivative of compound (2b)with a recombinant or a polypeptide of the present disclosure. In at least one embodiment, the contacting with the recombinant ‐ 10 ‐ polypeptides with CYP activity and CPR activity occurs in the presence of host cell that expresses the polypeptides. In at least one embodiment, the contacting with the recombinant polypeptides with CYP activity and CPR activity occurs in a cell-free system.
[0032] In at least one embodiment of the method, a derivative of compound (1a) is prepared, wherein the derivative is selected from compound (1b), compound (1c), compound (1d), compound (1e), compound (1f), compound (1g), compound (1h), compound (1i), compound (1j), compound (1k), compound (1l), compound (1m), compound (1n), compound (1o), compound (1p), compound (1q), compound (1r), compound (1s), compound (1t), compound (1u), compound (1v), compound (1w), compound (1x), compound (1y), compound (1z), compound (1aa), compound (1ab), compound (1ac), compound (1ad), compound (1ae), compound (1af), compound (1ag), compound (1ah), and compound (1ai).
[0033] In at least one embodiment of the method, a derivative of compound (2a) and / or a derivative of compound (2b) is used to prepare a derivative of compound (1a), wherein the derivative of compound (2a) and / or a derivative of compound (2b), wherein the derivative is selected from compound (2c), compound (2d), compound (2e), compound (2f), compound (2g), compound (2h), compound (2i), compound (2j), compound (2k), compound (2l), compound (2m), compound (2n), compound (2o), compound (2p), and mixture thereof.
[0034] In at least one embodiment, the method can further comprise a chemical step to form a derivative of compound (1a); optionally, wherein the derivative of compound (1a) is selected from compound (1b), compound (1c), compound (1d), compound (1e), compound (1f), compound (1g), compound (1h), compound (1i), compound (1j), compound (1k), compound (1l), compound (1m), compound (1n), compound (1o), compound (1p), compound (1q), compound (1r), compound (1s), compound (1t), compound (1u), compound (1v), compound (1w), compound (1x), compound (1y), compound (1z), compound (1aa), compound (1ab), compound (1ac), compound (1ad), compound (1ae), compound (1af), compound (1ag), compound (1ah), and compound (1ai). BRIEF DESCRIPTION OF THE DRAWINGS
[0035] A better understanding of the novel features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0036] FIG.1A, 1B, and 1C show schematic depictions of the CYP and CPR gene insertion sites in the strain constructs described in Examples 1 and 2. FIG.1A depicts the X-4 insertion site of the OG002 strain containing m-Venus and the URA3 gene under the bidirectional pGal10 / 1 promoter. FIG.1B shows the X-4 insertion site of the OG003 strain ‐ 11 ‐ containing CPR006 (SEQ ID NO: 6) and the URA3 gene (which was used as a landing site to integrate the two CYP genes) under the bidirectional pGal10 / 1 promoter. FIG.1C shows the X-4 insertion site of the OG014 strain containing a non-functional CPR006 and the URA3 gene (which was used as a landing site to integrate various CYP homolog genes) under the bidirectional pGal10 / 1 promoter.
[0037] FIG.2 depicts exemplary UHPLC profiles obtained for the recombinant host strains OG004, OG005, OG008 and OG009, which were constructed and screened for THQ production as described in Example 1. DETAILED DESCRIPTION
[0038] For the descriptions herein and the appended claims, the singular forms “a”, and “an” include plural referents unless the context clearly indicates otherwise. Thus, for example, reference to “a protein” includes more than one protein, and reference to “a compound” refers to more than one compound. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation. The use of “comprise,” “comprises,” “comprising” “include,” “includes,” and “including” are interchangeable and not intended to be limiting. It is to be further understood that where descriptions of various embodiments use the term “comprising,” those skilled in the art would understand that in some specific instances, an embodiment can be alternatively described using language “consisting essentially of” or “consisting of.”
[0039] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening integer of the value, and each tenth of each intervening integer of the value, unless the context clearly dictates otherwise, between the upper and lower limit of that range, and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of these limits, ranges excluding (i) either or (ii) both of those included limits are also included in the invention. For example, “1 to 50,” includes “2 to 25,” “5 to 20,” “25 to 50,” “1 to 10,” etc.
[0040] Generally, the nomenclature used herein and the techniques and procedures described herein include those that are well understood and commonly employed by those of ordinary skill in the art, such as the common techniques and methodologies described in e.g., Green and Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Vols. 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y., 2012 (hereinafter ‐ 12 ‐ “Sambrook”); and Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., originally published in 1987 in book form by Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., and regularly supplemented through 2011, and now available in journal format online as Current Protocols in Molecular Biology, Vols.00 - 130, (1987-2020), published by Wiley & Sons, Inc. in the Wiley Online Library (hereinafter “Ausubel”).
[0041] All publications, patents, patent applications, and other documents referenced in this disclosure are hereby incorporated by reference in their entireties for all purposes to the same extent as if each individual publication, patent, patent application or other document were individually indicated to be incorporated by reference herein for all purposes.
[0042] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention pertains. It is to be understood that the terminology used herein is for describing particular embodiments only and is not intended to be limiting. For purposes of interpreting this disclosure, the following description of terms will apply and, where appropriate, a term used in the singular form will also include the plural form and vice versa.
[0043] Definitions
[0044] “THQ” refers to the compound thymohydroquinone having the chemical structure shown as compound (1a) in Table 1 (below).
[0045] “THQ derivatives” as referenced in the present disclosure include, but are not limited to, structural analogs of THQ, such as thymoquinone (TQ), dithymoquinone (DTQ), and the other exemplary compounds shown below in Table 1 (below).
[0046] TABLE 1: Exemplary THQ and THQ derivative compounds Compound Name Chemical Structure Th h d i‐ 13 ‐ Dithymoquinone (“DTQ”)‐ 14 ‐ 2-hydroxy-5-[2-(pyrrolidin-1- yl)ethoxy]-p-cymene hydrochloride‐ 15 ‐ 5-(2- Methylaminoethoxy)carvacrol hydrochloride‐ 16 ‐ thymoquinol 2-O-β- glucopyranoside‐ 17 ‐ 5-isopropyl-2-methyl-4- trifluoromethoxyphenol
[0047] “THQ precursor” as used herein refers to a compound capable of being converted into THQ, or a THQ derivative, by an enzyme capable producing THQ alone, or in combination with another enzyme or a non-enzymatic chemical reaction. THQ precursors as referenced in the present disclosure include, but are not limited to, the exemplary compounds summarized in Table 2 (below).
[0048] TABLE 2: Exemplary THQ precursor compounds Compound Name Chemical Structure l‐ 18 ‐ 5-[2-(pyrrolidin-1-yl)ethoxy]- p-cymene hydrochloride‐ 19 ‐ 5-isopropyl-2-methyl-4-(2- amino-4-chloropyrimidine)- benzene
[0049] CYP activity as used herein refers to the catalytic activity of a cytochrome P450 monooxygenase enzyme.
[0050] “CPR activity” as used herein refers to the catalytic activity of a cytochrome P450 reductase enzyme.
[0051] “Conversion” as used herein refers to the enzymatic conversion of a substrate(s) to a corresponding product(s). “Percent conversion” refers to the percent of the substrate that is converted to the product within a period of time under specified conditions. Thus, the ‐ 20 ‐ “enzymatic activity” or “activity” of an enzymatic conversion can be expressed as “percent conversion” of the substrate to the product.
[0052] “Substrate” as used herein in the context of an enzyme mediated process refers to the compound or molecule acted on by the enzyme.
[0053] “Product” as used herein in the context of an enzyme mediated process refers to the compound or molecule resulting from the activity of the enzyme.
[0054] “Host cell” as used herein refers to a cell capable of being functionally modified with recombinant nucleic acids and functioning to express recombinant products, including polypeptides and compounds produced by activity of the polypeptides.
[0055] “Nucleic acid,” or “polynucleotide” as used herein interchangeably to refer to two or more nucleosides that are covalently linked together. The nucleic acid may be wholly comprised ribonucleosides (e.g., RNA), wholly comprised of 2'-deoxyribonucleotides (e.g., DNA) or mixtures of ribo- and 2'-deoxyribonucleosides. The nucleoside units of the nucleic acid can be linked together via phosphodiester linkages (e.g., as in naturally occurring nucleic acids), or the nucleic acid can include one or more non-natural linkages (e.g., phosphorothioester linkage). Nucleic acid or polynucleotide is intended to include single- stranded or double-stranded molecules, or molecules having both single-stranded regions and double-stranded regions. Nucleic acid or polynucleotide is intended to include molecules composed of the naturally occurring nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), or molecules comprising that include one or more modified and / or synthetic nucleobases, such as, for example, inosine, xanthine, hypoxanthine, etc.
[0056] “Protein,” “polypeptide,” and “peptide” are used herein interchangeably to denote a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristilation, ubiquitination, etc.). As used herein “protein” or “polypeptide” or “peptide” polymer can include D- and L-amino acids, and mixtures of D- and L-amino acids.
[0057] “Naturally-occurring” or “wild-type” as used herein refers to the form as found in nature. For example, a naturally occurring nucleic acid sequence is the sequence present in an organism that can be isolated from a source in nature, and which has not been intentionally modified by human manipulation.
[0058] “Recombinant,” “engineered,” or “non-naturally occurring” when used herein with reference to, e.g., a cell, nucleic acid, or polypeptide, refers to a material, or a material corresponding to the natural or native form of the material, that has been modified in a manner that would not otherwise exist in nature, or is identical thereto but is produced or derived from synthetic materials and / or by manipulation using recombinant techniques. Non-limiting examples include, among others, recombinant cells expressing genes that are ‐ 21 ‐ not found within the native (non-recombinant) form of the cell or express native genes that are otherwise expressed at a different level.
[0059] “Nucleic acid derived from” as used herein refers to a nucleic acid having a sequence at least substantially identical to a sequence of found in naturally in an organism. For example, cDNA molecules prepared by reverse transcription of mRNA isolated from an organism, or nucleic acid molecules prepared synthetically to have a sequence at least substantially identical to, or which hybridizes to a sequence at least substantially identical to a nucleic sequence found in an organism.
[0060] “Coding sequence” refers to that portion of a nucleic acid (e.g., a gene) that encodes an amino acid sequence of a protein.
[0061] “Heterologous nucleic acid” as used herein refers to any polynucleotide that is introduced into a host cell by laboratory techniques and includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.
[0062] “Codon degenerate” describes a nucleotide sequence that has one or more different codons relative to the reference nucleotide sequence, but which encodes a polypeptide that is identical to the polypeptide encoded by a reference nucleotide sequence. The different codons between the nucleotide sequence and the reference nucleotide sequence are called “synonyms” or “synonymous” codons in that they use different triplets of nucleotides to encode the same amino acid in a polypeptide.
[0063] “Codon optimized” refers to changes in the codons of the polynucleotide encoding a protein to those preferentially used in a particular organism such that the encoded protein is efficiently expressed in the organism of interest. Although the genetic code is degenerate in that most amino acids are represented by several different “synonymous” codons, it is well known that codon usage by particular organisms is nonrandom and biased towards particular codon triplets. This codon usage bias may be higher in reference to a given gene, genes of common function or ancestral origin, highly expressed proteins versus low copy number proteins, and the aggregate protein coding regions of an organism's genome. In some embodiments, the polynucleotides encoding the imine reductase enzymes may be codon optimized for optimal production from the host organism selected for expression.
[0064] “Preferred, optimal, high codon usage bias codons” refers to codons that are used at higher frequency in the protein coding regions than other codons that code for the same amino acid. The preferred codons may be determined in relation to codon usage in a single gene, a set of genes of common function or origin, highly expressed genes, the codon frequency in the aggregate protein coding regions of the whole organism, codon frequency in the aggregate protein coding regions of related organisms, or combinations thereof. Codons whose frequency increases with the level of gene expression are typically optimal ‐ 22 ‐ codons for expression. A variety of methods are known for determining the codon frequency (e.g., codon usage, relative synonymous codon usage) and codon preference in specific organisms, including multivariate analysis, for example, using cluster analysis or correspondence analysis, and the effective number of codons used in a gene (see GCG CodonPreference, Genetics Computer Group Wisconsin Package; CodonW, John Peden, University of Nottingham; McInerney, J. O, 1998, Bioinformatics 14:372-73; Stenico et al., 1994, Nucleic Acids Res.222437-46; Wright, F., 1990, Gene 87:23-29). Codon usage tables are available for a growing list of organisms (see for example, Wada et al., 1992, Nucleic Acids Res.20:2111-2118; Nakamura et al., 2000, Nucl. Acids Res.28:292; Duret, et al., supra; Henaut and Danchin, "Escherichia coli and Salmonella," 1996, Neidhardt, et al. Eds., ASM Press, Washington D.C., p.2047-2066. The data source for obtaining codon usage may rely on any available nucleotide sequence capable of coding for a protein. These data sets include nucleic acid sequences actually known to encode expressed proteins (e.g., complete protein coding sequences-CDS), expressed sequence tags (ESTS), or predicted coding regions of genomic sequences (see for example, Mount, D., Bioinformatics: Sequence and Genome Analysis, Chapter 8, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 2001; Uberbacher, E. C., 1996, Methods Enzymol.266:259-281; Tiwari et al., 1997, Comput. Appl. Biosci.13:263-270).
[0065] “Control sequence” as used herein refers to all sequences, which are necessary or advantageous for the expression of a polynucleotide and / or polypeptide as used in the present disclosure. Each control sequence may be native or foreign to the nucleic acid sequence encoding a polypeptide. Such control sequences include, but are not limited to, a leader, a promoter, a polyadenylation sequence, a pro-peptide sequence, a signal peptide sequence, and a transcription terminator. At a minimum, control sequences typically include a promoter, and transcriptional and translational stop signals. The control sequences may be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the nucleic acid sequence encoding a polypeptide.
[0066] “Operably linked” as used herein refers to a configuration in which a control sequence is appropriately placed (e.g., in a functional relationship) at a position relative to a polynucleotide sequence or polypeptide sequence of interest such that the control sequence directs or regulates the expression of the sequence of interest.
[0067] “Promoter sequence” refers to a nucleic acid sequence that is recognized by a host cell for expression of a polynucleotide of interest, such as a coding sequence. The promoter sequence contains transcriptional control sequences, which mediate the expression of a polynucleotide of interest. The promoter may be any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid ‐ 23 ‐ promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.
[0068] “Percentage of sequence identity,” “percent sequence identity,” “percentage homology,” or “percent homology” are used interchangeably herein to refer to values quantifying comparisons of the sequences of polynucleotides or polypeptides, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (or gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage values may be calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity. Alternatively, the percentage may be calculated by determining the number of positions at which either the identical nucleic acid base or amino acid residue occurs in both sequences or a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity. Those of skill in the art appreciate that there are many established algorithms available to align two sequences. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math.2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol.48:443, by the search for similarity method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection (see generally, Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Examples of algorithms that are suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., 1990, J. Mol. Biol.215: 403-410 and Altschul et al., 1977, Nucleic Acids Res.3389-3402, respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website. This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as, the neighborhood word score threshold (Altschul et al, supra). These initial neighborhood word hits act as seeds for initiating searches to find ‐ 24 ‐ longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc Natl Acad Sci USA 89:10915). Exemplary determination of sequence alignment and % sequence identity can employ the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison Wis.), using default parameters provided.
[0069] “Reference sequence” refers to a defined sequence used as a basis for a sequence comparison. A reference sequence may be a subset of a larger sequence, for example, a segment of a full-length nucleic acid or polypeptide sequence. A reference sequence typically is at least 20 nucleotide or amino acid residue units in length but can also be the full length of the nucleic acid or polypeptide. Since two polynucleotides or polypeptides may each (1) comprise a sequence (i.e., a portion of the complete sequence) that is similar between the two sequences, and (2) may further comprise a sequence that is divergent between the two sequences, sequence comparisons between two (or more) polynucleotides or polypeptide are typically performed by comparing sequences of the two polynucleotides or polypeptides over a “comparison window” to identify and compare local regions of sequence similarity. “Comparison window” refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acids residues wherein a sequence may be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids and wherein the portion of the sequence in the comparison window may comprise additions or deletions (or gaps) of 20 percent or less as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences.
[0070] “Substantial identity” or “substantially identical” refers to a polynucleotide or polypeptide sequence that has at least 70% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95 % sequence identity, or at least 99% sequence identity, as compared to a reference sequence ‐ 25 ‐ over a comparison window of at least 20 nucleoside or amino acid residue positions, frequently over a window of at least 30-50 positions, wherein the percentage of sequence identity is calculated by comparing the reference sequence to a sequence that includes deletions or additions which total 20 percent or less of the reference sequence over the window of comparison.
[0071] “Corresponding to,” “reference to,” or “relative to” when used in the context of the numbering of a given amino acid or polynucleotide sequence refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue number or residue position of a given polymer is designated with respect to the reference sequence rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as that of an engineered CYP, can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, although the gaps are present, the numbering of the residue in the given amino acid or polynucleotide sequence is made with respect to the reference sequence to which it has been aligned.
[0072] “Isolated” as used herein in reference to a molecule means that the molecule (e.g., THQ, polynucleotide, polypeptide) is substantially separated from other compounds that naturally accompany it, e.g., protein, lipids, and polynucleotides. The term embraces nucleic acids which have been removed or purified from their naturally occurring environment or expression system (e.g., host cell).
[0073] “Substantially pure” refers to a composition in which a desired molecule is the predominant species present (i.e., on a molar or weight basis it is more abundant than any other individual macromolecular species in the composition) and is generally a substantially purified composition when the object species comprises at least about 50 percent of the macromolecular species present by mole or % weight.
[0074] “Recovered” as used herein in relation to an enzyme, protein, or compound (e.g., THQ), refers to a more or less pure form of the enzyme, protein, or compound.
[0075] Homologous Genes Encoding Recombinant Polypeptides with CYP Activity
[0076] The present disclosure provides genes encoding polypeptides with cytochrome P450 (CYP) monooxygenase activity capable of converting the substrate compounds, carvacrol and / or thymol to the product compound, THQ, as illustrated by the reaction of Scheme 1 (below). ‐ 26 ‐ Scheme 1
[0077] As shown in Scheme 1, the conversion of carvacrol (compound 2a) and / or thymol (compound 2b) to the product compound, THQ (compound 1a) can be catalyzed by the CYP polypeptide of the present disclosure in the presence of a cytochrome P450 reductase (CPR) enzyme, the cofactor NADPH, and oxygen. As described herein, CYP polypeptides with the activity of Scheme 1 were identified by sequence homology analysis of two cytochrome P450 enzymes from oregano (Origanum vulgare), CYP76S40, and CYP736A300. These two CYP polypeptides have been previously characterized as having this activity and heterologous genes encoding have been shown to be capable of expressing this activity in Nicotiana benthiana or Saccharomyces cerevisiae (see e.g., Krause, et al. “The Biosynthesis of Thymol, Carvacrol, and Thymohydroquinone in Lamiaceae Proceeds via Cytochrome P450s and a Short-Chain Dehydrogenase.” Proceedings of the National Academy of Sciences 118, no.52 (December 28, 2021). CYP76S40 has the following 491 amino acid sequences: MDFLTPCLVVASIAWICMLILRARKPSKLPPGPYGPPIIGNILHLGPKPHRSLADLARKYGPVMKLRL GSVTTVVISSPEAAKAVLQKHDSSFWNRPAPSSVRAVGHDEFSVAWLPVDKQWRKLRKIMKELMFSSP RLDAGQGMRRAKLQQLSDYVWGRCQAGRAVEVGEAAFTTSLNLMSATLFSTDFARFDSDSSQEMKEVV WGVMKCVGSPNLVDYFPVLKSLDPQGILKDAKFCFGKLFAIFDEILDERLKISRGEKQDLVEALIDLN QRDVPQLSRDDINHLLLDLFVAGSDTTSGTVEWAMTELIRHPEKMTKLRNEITSFVEENGPIEESDIS RLPYLQAVVKETFRLHPVAPFLLPHKASSDIEINGYTVPKNAQILVNIWASGRDPINWVDADKFVPER FLSENKGLNFIGQDFELIPFGAGRRICPGLPLANRMVHQMLVTFVGNFEWKLEGIKVEEMDMDENFGL TLQKAIPLRAIPTKL (SEQ ID NO: 2; CYP002)
[0078] CYP736A300 has the following 498 amino acid sequence: MEWFWAALSLIVFLSLLHQLLKKEKKREANLPPSPIALPVIGHLHLLGKNLPLKLHAIAERHGPIVFL RLGLVRALVVSTAAGAELVLKTHDLVFSGRVHHQASRYLGYDQKNIVFAPYGAYWRNMRRLCMVKLLN AAKINEFRPVRRAELEETVASMRRAAEERGVVDVSAVISGVIGDMNSLMVFGRKYVDRDLDEELGFKA VIDEMLHVGALPNLGDFFPFMAALDLQGLDRRMKELSKIFDGFLERIIDDHLLKKTENTKKGDFVDTM LAVMEAGEADFEFDRRHVKAVLLDMLIAGMDTSASTVEWALSELIRHPEITKKLQKELEQVVGMDQMV DESHLDKLDYLDSVLKETLRLYPPGRLLVHETMEECTVNGFHIPKGTWTFVNMWSIGRDPAMWHEPEK FVPERFAGENLDFLGQNFKFIPFGAGRRSCPGLQLGLTFVRLVLAQLVHCFDWELPNGMVPSDLDMNE KFGIVTSRDKHLMAIPTYRLNK (SEQ ID NO: 4; CYP003)
[0079] As described in the Examples and elsewhere herein, the CYP76S40 and CYP736A300 sequences were used to identify 128 unique homologs with at least 60% ‐ 27 ‐ sequence identity from 45 different plant species. These candidate homolog polypeptides were screened for heterologous expression of the desired activity of Scheme 1 in yeast systems. In a first screening system, a heterologous CPR gene from Arabidopsis thaliana encoding the 692 amino acid CPR polypeptide of SEQ ID NO: 6 (CPR006) was also integrated into the yeast genome to provide the necessary P450 reductase activity. MTSALYASDLFKQLKSIMGTDSLSDDVVLVIATTSLALVAGFVVLLWKKTTADRSGELKPLMIPKSLM AKDEDDDLDLGSGKTRVSIFFGTQTGTAEGFAKALSEEIKARYEKAAVKVIDLDDYAADDDQYEEKLK KETLAFFCVATYGDGEPTDNAARFYKWFTEENERDIKLQQLAYGVFALGNRQYEHFNKIGIVLDEELC KKGAKRLIEVGLGDDDQSIEDDFNAWKESLWSELDKLLKDEDDKSVATPYTAVIPEYRVVTHDPRFTT QKSMESNVANGNTTIDIHHPCRVDVAVQKELHTHESDRSCIHLEFDISRTGITYETGDHVGVYAENHV EIVEEAGKLLGHSLDLVFSIHADKEDGSPLESAVPPPFPGPCTLGTGLARYADLLNPPRKSALVALAA YATEPSEAEKLKHLTSPDGKDEYSQWIVASQRSLLEVMAAFPSAKPPLGVFFAAIAPRLQPRYYSISS SPRLAPSRVHVTSALVYGPTPTGRIHKGVCSTWMKNAVPAEKSHECSGAPIFIRASNFKLPSNPSTPI VMVGPGTGLAPFRGFLQERMALKEDGEELGSSLLFFGCRNRQMDFIYEDELNNFVDQGVISELIMAFS REGAQKEYVQHKMMEKAAQVWDLIKEEGYLYVCGDAKGMARDVHRTLHTIVQEQEGVSSSEAEAIVKK LQTEGRYLRDVW (SEQ ID NO: 6; CPR006).
[0080] In a second screening system, the heterologous CPR of SEQ ID NO: 6 was knocked out and the native yeast cytochrome P450 reductase activity was relied upon to provide the reductase activity for the reaction.
[0081] In at least one embodiment of the present disclosure, the CYP76S40 polypeptide of SEQ ID NO: 2 can be encoded for heterologous expression by the yeast codon-optimized nucleotide sequence of SEQ ID NO: 1. ATGGATTTCTTAACCCCCTGTCTCGTGGTAGCATCCATTGCCTGGATTTGTATGTTGATACTAAGAGC AAGAAAACCGTCTAAACTTCCTCCTGGCCCATATGGTCCGCCCATCATAGGTAATATACTACATTTAG GTCCTAAACCTCATAGATCTTTGGCCGACTTGGCCCGAAAATATGGGCCAGTCATGAAGCTTAGGTTG GGTTCGGTAACCACAGTTGTTATATCTTCCCCAGAAGCAGCTAAAGCTGTTTTACAAAAACACGATTC ATCATTTTGGAATAGACCTGCTCCTTCTTCTGTTAGGGCGGTTGGTCACGATGAATTCTCAGTCGCTT GGCTGCCTGTAGACAAACAATGGAGAAAACTTCGGAAAATAATGAAAGAATTAATGTTCTCATCGCCA AGATTAGATGCCGGCCAAGGTATGAGACGTGCAAAACTTCAACAATTAAGTGATTATGTCTGGGGTAG GTGCCAGGCTGGTAGAGCGGTCGAGGTTGGAGAGGCTGCATTCACGACTTCTCTGAATTTAATGTCAG CTACATTATTTTCTACTGATTTTGCTCGTTTCGATTCCGATAGCAGTCAGGAAATGAAGGAAGTGGTA TGGGGCGTTATGAAGTGTGTGGGTTCTCCAAACTTGGTAGATTACTTTCCAGTCTTGAAATCCCTCGA TCCGCAAGGGATCCTAAAAGACGCAAAGTTCTGCTTTGGGAAGTTGTTCGCGATCTTTGACGAAATTC TGGATGAACGCTTGAAGATCTCTCGTGGTGAAAAGCAGGATTTGGTAGAGGCACTAATCGATTTAAAT CAACGTGACGTTCCTCAACTAAGCCGCGACGATATTAACCACCTATTGTTAGACTTATTTGTTGCTGG TTCAGATACTACTAGTGGTACAGTCGAATGGGCAATGACAGAATTGATTAGGCATCCCGAAAAGATGA CTAAGTTGAGAAATGAGATTACTTCATTTGTTGAAGAGAACGGCCCAATTGAAGAAAGTGATATCAGC AGATTGCCATACTTACAAGCCGTCGTTAAAGAGACGTTTAGACTTCATCCTGTGGCCCCATTCCTACT TCCACATAAGGCTAGCTCGGACATTGAAATTAACGGTTACACGGTTCCTAAAAATGCACAGATCCTGG TTAACATTTGGGCCTCCGGAAGGGATCCGATTAACTGGGTGGATGCTGATAAGTTTGTGCCAGAAAGG TTTTTAAGTGAGAATAAAGGACTGAACTTTATAGGACAGGATTTTGAACTCATACCCTTTGGCGCTGG TAGAAGAATTTGTCCAGGGTTGCCCTTGGCTAATAGAATGGTCCATCAAATGCTGGTGACCTTCGTTG GAAATTTCGAGTGGAAGCTTGAAGGAATAAAGGTAGAGGAAATGGACATGGACGAAAATTTTGGCTTA ACCTTGCAAAAAGCAATCCCACTACGAGCGATTCCAACAAAATTATAA (SEQ ID NO: 1)
[0082] In at least one embodiment of the present disclosure, the CYP736A300 polypeptide of SEQ ID NO: 4 can be encoded for heterologous expression by the yeast codon-optimized nucleotide sequence of SEQ ID NO: 3. ‐ 28 ‐ ATGGAGTGGTTTTGGGCGGCTCTCTCATTGATAGTCTTTTTGTCGTTGTTACACCAATTACTAAAAAA GGAGAAGAAAAGGGAAGCTAACCTCCCACCGAGTCCTATTGCTCTTCCTGTAATCGGCCATTTACACC TATTGGGAAAAAACCTTCCCTTAAAGCTGCACGCAATTGCCGAACGTCATGGTCCTATAGTTTTTTTA AGATTAGGCCTTGTACGTGCCTTGGTGGTTAGCACAGCAGCTGGGGCGGAGTTAGTGTTAAAAACTCA TGATTTGGTTTTTAGTGGAAGAGTTCATCATCAAGCATCTCGCTACCTTGGTTATGACCAAAAGAATA TCGTATTTGCACCTTATGGGGCATATTGGCGAAACATGCGTAGATTATGTATGGTCAAATTGTTGAAT GCTGCCAAGATAAATGAATTTAGGCCCGTAAGAAGAGCTGAATTAGAAGAAACCGTTGCTTCCATGAG AAGGGCAGCTGAAGAAAGGGGAGTGGTGGACGTGTCAGCAGTAATTTCCGGTGTCATTGGAGATATGA ACTCGTTAATGGTTTTCGGAAGAAAATATGTGGACCGGGATCTGGATGAAGAGTTAGGTTTCAAGGCA GTCATAGATGAAATGTTGCACGTTGGTGCCTTGCCTAATCTAGGTGACTTTTTTCCATTCATGGCTGC ACTAGACTTGCAAGGTTTGGATAGAAGGATGAAAGAATTGAGCAAGATTTTTGACGGGTTCCTAGAAC GTATAATTGACGATCACTTGCTGAAAAAGACTGAGAATACCAAAAAGGGGGATTTCGTGGATACAATG CTGGCTGTTATGGAGGCCGGCGAAGCCGATTTTGAATTCGATCGTAGACATGTTAAAGCTGTTCTGCT CGATATGCTTATCGCTGGTATGGATACGTCTGCCAGCACAGTTGAATGGGCTCTGTCTGAGTTGATTA GACACCCTGAAATCACTAAGAAACTACAAAAAGAACTCGAGCAGGTTGTTGGTATGGACCAGATGGTA GATGAGTCTCATCTAGACAAACTAGATTACCTGGATTCCGTCCTTAAAGAGACCTTAAGATTATACCC ACCCGGTAGATTGCTTGTCCATGAAACGATGGAAGAATGTACAGTTAACGGTTTTCATATTCCAAAGG GTACATGGACTTTCGTCAATATGTGGAGTATAGGAAGAGATCCAGCGATGTGGCATGAACCAGAAAAG TTTGTACCTGAACGCTTTGCCGGTGAAAATTTAGATTTTCTAGGCCAAAATTTCAAATTCATTCCATT CGGTGCGGGCCGAAGGTCATGCCCGGGTTTACAGTTGGGCTTAACTTTTGTTAGGTTGGTGTTGGCAC AACTTGTACATTGTTTTGACTGGGAACTTCCAAACGGAATGGTTCCGTCAGACTTGGATATGAACGAG AAATTCGGTATCGTCACTTCTAGAGATAAGCATTTAATGGCTATTCCAACCTACAGATTAAATAAATA A (SEQ ID NO: 4)
[0083] In at least one embodiment of the present disclosure, the CPR006 polypeptide of SEQ ID NO: 6 can be encoded for heterologous expression by the yeast codon-optimized nucleotide sequence of SEQ ID NO: 5. ATGACCTCCGCACTCTACGCATCCGATTTGTTTAAGCAGTTGAAATCTATAATGGGAACCGACTCCTT GAGCGATGACGTGGTTTTAGTGATTGCGACTACTTCATTGGCCTTAGTAGCCGGATTCGTTGTACTTT TGTGGAAAAAAACAACTGCTGATCGCTCTGGTGAATTAAAACCACTAATGATCCCAAAGAGTCTGATG GCTAAAGATGAAGATGATGATCTAGACCTAGGTTCAGGTAAAACAAGAGTCTCTATATTTTTTGGCAC CCAGACAGGTACGGCTGAAGGCTTTGCTAAAGCATTGTCGGAAGAAATTAAGGCCAGATATGAAAAGG CAGCAGTCAAGGTTATAGACCTGGACGATTACGCTGCGGACGACGACCAATACGAGGAAAAACTAAAG AAGGAGACGTTGGCTTTTTTTTGTGTAGCTACATACGGCGATGGAGAACCGACAGACAATGCTGCTAG ATTCTACAAGTGGTTCACTGAAGAGAACGAACGTGACATAAAATTACAGCAACTTGCATACGGGGTGT TTGCCCTTGGCAACCGACAATACGAACACTTTAATAAGATCGGGATCGTTCTAGATGAGGAGCTGTGC AAAAAAGGAGCAAAAAGGCTTATTGAGGTGGGTCTCGGCGATGACGATCAAAGCATTGAGGATGATTT CAACGCGTGGAAAGAATCGTTGTGGAGTGAATTGGATAAACTTTTAAAGGACGAAGATGATAAGAGTG TGGCTACACCATATACTGCCGTTATTCCCGAGTATAGAGTCGTAACACACGATCCACGTTTCACAACT CAAAAGTCAATGGAATCAAATGTCGCCAATGGTAACACTACTATTGATATTCACCACCCTTGCAGAGT TGATGTAGCAGTACAAAAAGAGCTTCATACCCATGAATCTGATAGGAGCTGTATACATTTGGAGTTTG ATATCAGCAGAACGGGGATAACTTATGAGACAGGAGATCATGTTGGTGTCTATGCTGAGAATCATGTC GAGATTGTTGAAGAAGCTGGTAAACTACTGGGGCACTCCTTAGATTTAGTTTTCTCGATTCATGCTGA TAAAGAAGATGGCTCACCTTTGGAAAGTGCTGTACCACCACCATTTCCTGGGCCTTGTACTTTAGGAA CCGGTCTGGCTAGGTATGCAGATTTATTGAATCCGCCACGAAAATCAGCACTAGTAGCTTTAGCTGCA TATGCCACTGAACCTAGTGAAGCAGAAAAATTAAAACATCTCACCAGTCCTGATGGTAAAGACGAATA CTCACAGTGGATTGTTGCATCACAAAGATCATTATTAGAAGTGATGGCAGCGTTTCCCTCTGCCAAAC CTCCACTTGGAGTCTTTTTCGCTGCCATTGCCCCGAGATTGCAACCAAGATATTACAGTATTTCTTCT TCTCCAAGGTTGGCACCCAGCAGAGTTCATGTCACTTCAGCGTTGGTTTATGGTCCAACCCCAACAGG TAGGATCCATAAAGGTGTTTGTTCCACTTGGATGAAAAATGCAGTCCCAGCTGAAAAGTCTCACGAAT GTTCTGGTGCCCCTATCTTTATAAGAGCTTCTAATTTTAAGTTGCCCTCGAACCCGTCTACACCTATA GTTATGGTGGGTCCCGGAACCGGTCTAGCACCTTTCCGTGGTTTTTTGCAGGAAAGAATGGCGCTGAA AGAAGACGGTGAAGAATTAGGCTCCAGTTTACTGTTCTTTGGCTGTCGGAACAGGCAAATGGATTTCA TTTACGAAGACGAGCTTAACAATTTCGTCGATCAGGGCGTGATTTCCGAATTAATTATGGCCTTTTCG ‐ 29 ‐ AGAGAGGGAGCACAAAAGGAATATGTTCAGCATAAAATGATGGAGAAGGCTGCGCAAGTTTGGGACTT GATCAAGGAGGAAGGATATCTATATGTTTGCGGTGATGCCAAAGGTATGGCTAGAGATGTTCATCGTA CATTACATACGATAGTGCAAGAACAAGAAGGTGTATCTTCATCCGAAGCTGAAGCCATCGTTAAGAAG TTACAAACGGAAGGTCGCTATCTCAGAGACGTATGGTAA (SEQ ID NO: 5)
[0084] As a result of screening the 128 candidate homolog polypeptides for the desired CYP activity when expressed heterologously in two yeast screening systems, 15 exemplary polypeptide homologs were identified that exhibit the unexpected and surprising technical effect of THQ production when integrated in a recombinant host cell. These exemplary polypeptides and exemplary yeast optimized genes encoding them are summarized in Table 3A below (as well as in the following Examples and the accompanying Sequence Listing).
[0085] TABLE 3A: Homolog CYP polypeptides with carvacrol / thymol to THQ conversion activity NT AA SEQ SEQ :‐ 30 ‐ LDLQGLDRRMKKLSSIFDAFLDKIIDDHLQRKPEKTH NRDFVDTMLAVMDSGEAGFEFDRRHVKAVLLDILIGG MDTTLSTVEWAISELIRHPKVTKKLQSELERIVSLDE‐ 31 ‐ SAEFARFDSESSQEMKEAVCGVMKSVGSPNFADYFPL LKPADPQGLLKAAELSTGMLFVKLDEIIDEKLKSRGE KRDLVEALLEINQRDEAQLSRDDIRHLLLDLLVAGTD‐ 32 ‐ VGENVEVEESDISRLPYLQAVVKETFRLHPAAPFLVP HKVNADVEIKGFTVPKNAQVLVNVWASGRDPDTWPEA ESFSPERFMDGRQIDIRGKDFELIPFGSGRRICPGLP‐ 33 ‐ MEDCVVDGFHIPKNSRILVNVWAIGRDPNVWPDPEAF SPERFVGCNIDLRGRDFQLIPFGSGRRGCPGLQLGLT VVRLMVAQLVHCFDWKLPSGMKPNELDMSEHFGLVTStransformed with a heterologous nucleic acid encoding an exemplary homolog CYP polypeptide of the present disclosure (e.g., a CYP polypeptide of Table 3A), and the recombinant host cell is fed the substrate(s) carvacrol and / or thymol, the product molecule, THQ is produced by the host cell. Furthermore, the THQ production is in greater yield relative to a comparable recombinant host cell integrated with a gene encoding a recombinant CYP polypeptide from Origanum vulgare of SEQ ID NO: 2 or 4. Without intending to be bound by any particular theory or mechanism, the enhanced yield of the THQ biosynthetic product is correlated with the one or more amino acid residue differences in the homolog CYP polypeptides of the present disclosure, as compared to the amino acid sequences of the CYP polypeptides of SEQ ID NO: 2 or 4, both of which have sequence identity of 76% or less to any of the 15 exemplary homolog polypeptides of Table 3A.
[0087] Based on the correlation of recombinant polypeptide functional information provided herein with the sequence information provided in Table 3A, the accompanying Sequence Listing, and / or the Examples disclosed herein, one of ordinary skill can recognize that the present disclosure of homolog CYP polypeptides provides a range of recombinant polypeptides having CYP activity capable of converting carvacrol and / or thymol to THQ with improved activity (e.g., increased activity or increased yield of the desired product compound), wherein the polypeptide comprises an amino acid sequence comprising one or more of the amino acid differences or sets of amino acid differences (relative to SEQ ID NO: 2 or 4) disclosed in any one of polypeptide homolog sequences of SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, and 36, and otherwise have at least 80%, at least 85% at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, and 36.
[0088] Additionally, in at least one embodiment, a homolog CYP polypeptide of the present disclosure having activity capable of converting carvacrol and / or thymol to THQ can have an amino acid sequence comprising one or more of the amino acid differences or sets of amino acid differences relative to SEQ ID NO: 2 or 4 disclosed in any one of SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, and 36, and additionally have 1-2, 1-3, 1-4, 1-5, 1- 6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-14, 1-15, 1-16, 1-18, 1-20, 1-22, 1-24, 1-26, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, or 1-60 residue differences at other residue positions. In some ‐ 34 ‐ embodiments, the number of differences can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 20, 22, 24, 26, 30, 35, 40, 45, 50, 55, or 60 residue differences at the other residue positions.
[0089] Accordingly, in at least one embodiment, the present disclosure provides engineered CYP polypeptides derived from homolog CYP polypeptide CYP029 (SEQ ID NO: 24), wherein the engineered polypeptides have at least 80% sequence identity to SEQ ID NO: 24, and an amino acid difference as compared to SEQ ID NO: 24, and wherein the engineered polypeptide also has improved CYP activity (e.g., increased activity and / or product yield) in converting carvacrol and / or thymol to THQ relative to SEQ ID NO: 24. Exemplary CYP polypeptides derived from CYP029 that exhibit this improved activity are described in Table 3B below and demonstrated in the Examples.
[0090] TABLE 3B: Engineered CYP029 polypeptides with carvacrol / thymol to THQ conversion activity NT AA Neutral SEQ SEQ :‐ 35 ‐ PQKMVKLRNEIRSFMSESGKQQIEES DISSLSYLQAVVKENFRLHPAAPFLV PHKANCDVEINGYIIPKNAQIFVNVW‐ 36 ‐ RDLVEVLLEIHRRDEAQLSRDDIRHL LLDLLVAGTDTTSGTMEWAMTELIRN PQKMVVLRNEIRSFMSESGKQQIEES‐ 37 ‐ KSIGRPNLADYFPVLKAVDPQGIYKE TELNVGRLFVKLDEIIDEKLRSRGEK RDLVEVLLEIHRRDEAQLSRDDIRHL‐ 38 ‐ AERCERGTAVDVGEAAFTTALNLMSA SLFSAEFVQFGSETSQEMKEAISSVV KSIGRPNLADYFPVLKAVDPQGIYKE‐ 39 ‐ LRHHEFSVGWLPVGNQWRKLRKICKE QMFSAPRLDASEGLRREKLQKLQDYL AERCERGTAVDVGEAAFTTALNLMSA‐ 40 ‐ LAKLSRKYGPVMSLKLGSITTVVISS AETAKLVLQKHDSSFSNRTVLSAITA LRHHEFSVGWLPVGNQWRKLRKICKE‐ 41 ‐ OGP096 K201C,MDFLTSSVILASIAWICMLISRARKV125 126 N321M, SKLPPGPYGLPIIGNILQLGPKPHHS Q333S LAKLSRKYGPVMSLKLGSITTVVISS‐ 42 ‐ DRMVHQMLVTFVGNYDWKLENGKPEE MDMNENFGITLEKAIPLKAIPIKV libOGP0 S27R P365P MDFLTSSVILASIAWICMLISRARKV 131 132‐ 43 ‐ ASGRDSKIWKNPDEFLPERFLNENDN IDFKGRDFELIPFGSGRRMCPGLPLA DRMVHQMLVTFVGNYDWKLENGKPEE‐ 44 ‐ DISSLSYLQAVVKENFRLHPAAPFLV PHKANCDVEINGYIIPKNAQIFVNVW ASGRDSKIWKNPDEFLPERFLNENGN‐ 45 ‐ LLDLLVAGTDTTSGTMEWAMTELIRN PQKMVKLRNEIRSFMSESGKQQIEES DISSLSYLQAVVKENFRLHPAAPFLV‐ 46 ‐ TEHNVGRLFVKLDEIIDEKLRSRGEK RDLVEVLLEIHRRDEAQLSRDDIKHL LLDLLVAGTDTTSGTMEWAMTELIRN‐ 47 ‐ SLFSAEFVQFGSETSQEMKEAISSVV KSIGRPNLADYFPVLKAVDPQGIYKE TEHNMGRLFVKLDEIIDEKLRSRGEK‐ 48 ‐ S299A, QMFSAPRLDASEGLRREKLQKLQDYL M464G AESCERGTAVDVGEAAFTTMLNLMSA SLFSAEFVQFGSETSQEMKEAISSVV‐ 49 ‐ 9_AP3_ H51R, (AGT), AETAKLVLQKHDSSFSNRTVASALTA 50x:C10 L99A, P358P LRHHEFSVGWLPVGSQWRKLRKICKE I102L, (CCG), QMFSAPRLDASEGLRREKLQKLQDYL‐ 50 ‐ 2024071 S344P, S111S MDFLTSSVILASIAWICMLISRARKV 177 178 0_OG_li S27R, (AGT), RKLPPGPYGLPIIGNILQLGPKPHRS bOGP08 H51R, P358P LAKLSRKYGPVMSLKLGSITTVVISS‐ 51 ‐ DRMVHQMLVTFVGNYDWKLENGKPEE MDMDENFGITLEKAIPLKAIPIKV 2024071 N95S S111S MDFLTSSVILASIAWICMLISRARKV 183 184‐ 52 ‐ ASGRDSKIWKNPDEFLPERFLNENVN IDFKGRDFELIPFGSGRRMCPGLPLA DRMVHQMLVTFVGNYDWKLENGKPEE‐ 53 ‐ N119S, PQKMVKLRNEIRSFMSESGKQQIEES A176M, DISSLPYLQAVVKENFRLHPAAPFLV T196A, PHKANCDVEINGYIIPKNAQIFVNVW‐ 54 ‐ K420C, RDLVEVLLEIHRRDEAQLSRDDIRHL M464G, LLDLLVAGTDTTAGTMEWAMTELIRN N472D PQKMVKLRNEIRSFMSESGKQQIEES‐ 55 ‐ S299A, KSIGRPNLADYFPVLKAVDPQGIYKE D415R, TELNVGRLFVKLDEIIDEKLRSRGEK M464G, RDLVEVLLEIHRRDEAQLSRDDIRHLpolypeptides derived from CYP029 disclosed in Table 3B can further comprise other residue differences relative to the reference polypeptide of SEQ ID NO: 24 at other residue positions. Residue differences at these other residue positions can provide for additional variations in the amino acid sequence without adversely affecting the ability of the recombinant polypeptide to carry out the desired biocatalytic conversion of carvacrol (compound 2a) and / or thymol (compound 2b) to THQ (compound 1a). In some embodiments, the recombinant polypeptides can have additionally 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1- 11, 1-12, 1-14, 1-15, 1-16, 1-18, 1-20, 1-22, 1-24, 1-26, 1-30, 1-35, 1-40 residue differences at other amino acid residue positions as compared to SEQ ID NO: 2. In some embodiments, the number of differences can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 20, 22, 24, 26, 30, 35, and 40 residue differences at other residue positions. The residue difference at these other positions can include conservative changes or non-conservative changes. In some embodiments, the residue differences can comprise conservative substitutions and non-conservative substitutions as compared to the reference polypeptide of SEQ ID NO: 24. Exemplary amino acid residue differences as compared to SEQ ID NO: 24 useful in the engineered CYP polypeptides derived from CYP029 include differences at positions M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, L105, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, G258, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, K420, Y457, D458, N472, I477, and T478. The specific differences include, but are not limited to those exemplified in the polypeptides of Table 3B: M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, L105V, N119S, K126R, M132V, K151E, R159S, R159W, ‐ 56 ‐ A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, G258V, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, K420C, Y457C, D458V, N472D, I477M, I477V, and T478C.
[0092] Combinations of amino acid residue differences as compared to SEQ ID NO: 24 useful in the engineered CYP polypeptides, and the recombinant host cells expressing them, can include any of the following combinations found in the exemplary engineered polypeptides of Table 3B: T196A, M464G S27R, H51R, L99A, I102L, N119S, A176M, T196A, D415R, M464G‐ 57 ‐ G35R, N95S, M132V, R415V, S27R, H51R, L99A, I102L, N119S, A176M, T196A, D415R, M464G E259G S344P N375D Y457C S27R H51R L99A I102L N119S A176M T196A(SEQ ID NO: 24) of the present disclosure can be in the form of fusion polypeptides in which the polypeptides are fused to other polypeptides, such as, by way of example and not limitation, antibody tags (e.g., myc epitope), purification sequences (e.g., His tags for binding to metals), and cell localization signals (e.g., secretion signals). Thus, the recombinant polypeptides described herein can be used with or without fusions to other polypeptides. It is also contemplated that the recombinant polypeptides described herein are not restricted to the genetically encoded amino acids. In addition to the genetically encoded amino acids, the polypeptides described herein may be comprised, either in whole or in part, of naturally occurring and / or synthetic non-encoded amino acids.
[0094] In another aspect, the present disclosure provides polynucleotides encoding the engineered CYP polypeptides derived from CYP029 (SEQ ID NO: 24) with CYP activity capable of converting carvacrol and / or thymol to THQ with increased activity and / or yield as described herein (e.g., CYP polypeptides of Table 3B). In at least one embodiment, the polynucleotide encoding a recombinant CYP polypeptide comprises a polynucleotide sequence that is at least about 80% identity, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to SEQ ID NO: 23.
[0095] In at least one embodiment, the polynucleotide has a sequence encoding a recombinant polypeptide of the present disclosure in which the encoding polynucleotide sequence has one or more neutral codon differences relative to a polynucleotide sequence of SEQ ID NO: 23, which codon differences do not encode an amino acid difference but ‐ 58 ‐ result in improved properties, such as increased yield of the desired product compound (e.g., THQ) produced by a recombinant host cell in which the polynucleotide sequence is integrated.
[0096] It is also contemplated that the polynucleotides encoding the recombinant polypeptides having CYP activity and increased activity and / or yield as described herein, can include a combination of one or more neutral codon differences relative to SEQ ID NO: 23, wherein at least one of the codon differences encodes an amino acid difference as compared to the parent sequence and at least one codon difference is a neutral codon difference that does not encode an amino acid difference as compared to the parent sequence of SEQ ID NO: 24. Accordingly, in at least one embodiment, the present disclosure provides a polynucleotide sequence encoding a recombinant polypeptide having CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polynucleotide sequence comprises a combination of a codon differences encoding an amino acid difference and a neutral codon difference. Exemplary neutral codon differences as compared to SEQ ID NO: 23 include differences at the following encoded amino acid positions of SEQ ID NO: 24: S21, V26, L53, S111, V112, G143, A171, T174, S210, I211, P214, L246, D289, D296, S299, Q333, D339, P358, P365, A368, D371, P437, and D458. Specific neutral codon differences as compared to SEQ ID NO: 23 include, but are not limited to, the following neutral codon differences present in the engineered CYP polypeptides of Table 3B: S21S (AGC), V26V (GTC), L53L (CTA), S111S (AGT), V112V (GTG), G143G (GGC), A171A (GCA), T174T (ACG), S210S (TCG), I211I (ATC), P214P (CCC), L246L (CTG), D289D (GAC), D296D (GAC), S299S (TCG), Q333Q (CAG), D339D (GAC), P358P (CCG), P365P (CCT), A368A (GCA), D371D (GAC), P437P (CCT), and D458D (GAC). As noted in Table 3B, the amino acid residue and position encoded by the codon that does not change as compared to SEQ ID NO: 24 is indicated using standard notation (e.g., “S21S“) followed by the new triplet codon to which the codon is changed as compared to SEQ ID NO: 23 in parentheses (e.g., “(AGC)”).
[0097] Combinations of neutral codon differences as compared to SEQ ID NO: 23 useful in the engineered CYP polypeptides, and recombinant host cells expressing them, can include any of the following combinations exemplified in the polynucleotides encoding the exemplary engineered CYP polypeptides of Table 3B: S111S (AGT), P358P (CCG), P437P (CCT)‐ 59 ‐ P358P (CCG), P365P (CCT) S299S (TCG), P437P (CCT)exemplary recombinant polypeptide having CYP activity as disclosed in Table 3B and the accompanying Sequence Listing. In at least one embodiment, the polynucleotide comprises a sequence of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205. In at least one embodiment, the polynucleotide comprises a codon degenerate sequence of a sequence selected from the group consisting of SEQ ID NO: a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205.
[0099] The polynucleotide sequences encoding the recombinant polypeptides of the present disclosure may be operatively linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide capable of expressing the polypeptide. Expression constructs containing a heterologous polynucleotide encoding the recombinant polypeptide can be introduced into appropriate host cells to express the corresponding polypeptide. Because of the knowledge of the codons corresponding to the various amino acids, availability of a protein sequence provides a description of all the polynucleotides capable of encoding the subject. The degeneracy of the genetic code, where the same amino acids are encoded by alternative or synonymous codons allows an ‐ 60 ‐ extremely large number of nucleic acids to be made, all of which encode the enzymes with improved CYP activity disclosed herein. Thus, having identified a particular amino acid sequence, those skilled in the art could make any number of different nucleic acids by simply modifying the sequence of one or more codons in a way which does not change the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates each and every possible variation of polynucleotides that could be made by selecting combinations based on the possible codon choices, and all such variations are to be considered specifically disclosed for any polypeptide disclosed herein, including the amino acid sequences presented in Tables 3A and 3B and the accompanying Sequence Listing.
[0100] The codons can be selected to fit the host cell in which the protein is being produced. For example, preferred codons used in bacteria are used to express the gene in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells. It is contemplated that all codons need not be replaced to optimize the codon usage of the recombinant polypeptide since the natural sequence will comprise preferred codons and because use of preferred codons may not be required for all amino acid residues. Consequently, codon optimized polynucleotides encoding the recombinant polypeptide may contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of codon positions of the full-length coding region.
[0101] The present disclosure also provides an expression vector comprising a polynucleotide encoding a recombinant polypeptide having CYP activity capable of converting carvacrol and / or thymol to THQ, and one or more expression regulating regions such as a promoter, a terminator, a replication origin, or the like, depending on the type of hosts into which they are to be introduced. The various nucleic acid and control sequences described above may be joined together to produce a recombinant expression vector which may include one or more convenient restriction sites to allow for insertion or substitution of the nucleic acid sequence encoding the recombinant polypeptide at such sites. Alternatively, a polynucleotide sequence of the present disclosure may be expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression. In creating the expression vector, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression. The recombinant expression vector may be any vector (e.g., a plasmid or virus), which can be conveniently subjected to recombinant DNA procedures and can bring about the expression of the polynucleotide sequence. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vectors may be linear or closed circular plasmids. ‐ 61 ‐
[0102] The expression vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a mini- chromosome, or an artificial chromosome. The vector may contain any means for assuring self-replication. Alternatively, the vector may be one which, when introduced into the host cell, is integrated into the genome, and replicated together with the chromosome(s) into which it has been integrated. Furthermore, a single vector or plasmid or two or more vectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, or a transposon may be used. In at least one embodiment, the expression vector further comprises one or more selectable markers, which permit easy selection of transformed cells.
[0103] Use in Recombinant Host Cells
[0104] The recombinant genes of the present disclosure that encode recombinant engineered polypeptides derived from CYP029 (SEQ ID NO: 24) with CYP activity capable of converting carvacrol and / or thymol to THQ (e.g., exemplary polypeptides of Table 3B) can be incorporated in recombinant host cells to enable in vivo biosynthesis of the compound THQ and THQ derivative compounds. These recombinant host cells comprise a polynucleotide or expression vector that encodes the recombinant homolog polypeptide with CYP activity, wherein the polynucleotide is operatively linked to one or more control sequences for expression of the polypeptide in the host cell. Host cells for use in expressing recombinant genes encoding the polypeptides with CYP activity of the present disclosure are well known in the art and include but are not limited to, bacterial cells, such as E. coli, or fungal cells, such as Saccharomyces cerevisiae or Pichia pastoris, insect cells, such as Drosophila S2 and Spodoptera Sf9, animal cells, such as CHO, COS, BHK, 293, and plant cells. Appropriate mediums and growth conditions for culturing the recombinant host cells so that they express the polypeptide with CYP activity are well known in the art.
[0105] The recombinant host cells can comprise heterologous nucleic acids encoding not only polypeptides with CYP activity capable of converting carvacrol and / or thymol to THQ, and also other enzymes, such as CPR, and enzymes capable of producing other compounds, such as precursors for the substrates, carvacrol and / or thymol. As described elsewhere herein, nucleic acid sequences encoding pathway enzymes for producing such precursor compounds are known in the art and can readily be used in accordance with the present disclosure. Typically, the nucleic acid sequence encoding the enzymes which form a part of the pathway, further include one or more additional nucleic acid sequences, for example, a nucleic acid sequence controlling expression of the enzymes which form a part of the biosynthetic pathway, and these one or more additional nucleic acid sequences together with the nucleic acid sequence encoding the recombinant polypeptides with CYP ‐ 62 ‐ activity can be considered a heterologous nucleic acid sequence. A variety of techniques and methodologies are available and well known in the art for introducing heterologous nucleic acid sequences, such as nucleic acid sequences encoding the enzymes (e.g., CYP and CPR), into a host cell so as to attain expression the host cell. Such techniques are well known to the skilled artisan and can be found in, for example, Sambrook et al., Molecular Cloning, a Laboratory Manual, Cold Spring Harbor Laboratory Press, 2012, Fourth Ed.
[0106] For example, the introduction of the heterologous nucleic acids can include integration of the nucleic acids into specific loci in the genome of a host cell via CRISPR- Cas9 and other techniques, some of which are demonstrated in the Examples herein. Such techniques are well known to the skilled artisan and can, for example, be found in Sambrook and other well-known sources. The number of copies of heterologous genes and their locus of integration in a recombinant host cell’s genome can result in improved biosynthetic production of a desired product, such as THQ. For example, a heterologous nucleic acid encoding a polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ can be integrated into a host cell’s genome in 1, 2, 3, 4, or more copies. Accordingly, it is contemplated that in the recombinant host cells of the present disclosure, the heterologous nucleic acid encoding the recombinant polypeptide having CYP activity can be integrated in the host cell’s genome at one or more loci, including but not limited to the well- known genomic loci in Saccharomyces cerevisiae of X-2, X-4, XI-2, XII-4, NDE1, XII-5, Gal80, and ROQ1, or the loci in Pichia pastoris of AOX1, Int6, Int15, and HIS4.
[0107] One of ordinary skill will recognize that the heterologous nucleic acids encoding the recombinant enzymes with CYP activity capable of converting carvacrol and / or thymol to THQ, CPR activity, and any other pathway enzymes will further comprise transcriptional promoters capable of controlling expression of the enzymes in the recombinant host cell. Generally, the transcriptional promoters are selected to be compatible with the host cell, so that promoters obtained from bacterial cells are used when a bacterial host cell is selected in accordance herewith, while a fungal promoter is used when a fungal host cell is selected, a plant promoter is used when a plant cell is selected, and so on. Promoters useful in the recombinant host cells of the present disclosure may be constitutive or inducible, provided such promoters are operable in the host cells. Promoters that may be used to control expression in fungal host cells, such as Saccharomyces cerevisiae and Pichia pastoris, are well known in the art and include, but are not limited to, inducible promoters, such as a Gal1 promoter or Gal10 promoter, a constitutive promoter, such as an alcohol dehydrogenase (ADH) promoter, a glyceraldehyde-3-phosphate dehydrogenase (GPD) promoter, or an S. pombe Nmt, or ADH promoter. Exemplary promoters that may be used to control expression in bacterial cells can include the Escherichia coli promoters, lac, tac, trc, trp or the T7 promoter. Exemplary promoters that may be used to control expression in plant cells ‐ 63 ‐ include, for example, a Cauliflower Mosaic Virus 35S promoter (Odell et al. (1985) Nature 313:810-812), a ubiquitin promoter (U.S. Pat. No.5,510,474; Christensen et al. (1989)), or a rice actin promoter (McElroy et al. (1990) Plant Cell 2:163-171). Exemplary promoters that can be used in mammalian cells include, a viral promoter such as an SV40 promoter or a metallothionine promoter. All of these host cell promoters are well known by and readily available to one of ordinary skill in the art. Further nucleic acid control elements useful for controlling expression in a recombinant host cell can include transcriptional terminators, enhancers, and the like, all of which may be used with the heterologous nucleic acids incorporate in the recombinant host cells of the present disclosure.
[0108] A wide variety of techniques are well known in the art for linking transcriptional promoters and other control elements to heterologous nucleic acid sequences encoding pathway genes for biosynthesis of THQ. Such techniques are described in e.g., Sambrook et al., Molecular Cloning, a Laboratory Manual, Cold Spring Harbor Laboratory Press, 2012, Fourth Ed. Accordingly, in at least one embodiment, the heterologous nucleic acid sequences of the present disclosure comprise a promoter capable of controlling expression in a host cell, wherein the promoter is linked to a nucleic acid sequence encoding a recombinant polypeptide of the present disclosure having CYP activity capable of converting carvacrol and / or thymol to THQ, and as necessary, other enzymes constituting a pathway for production of a precursor, such as carvacrol and / or thymol, THQ, and / or THQ derivative. This heterologous nucleic acid sequence can be integrated into a recombinant expression vector which ensures good expression in the desired host cell, wherein the expression vector is suitable for expression in a host cell, meaning that the recombinant expression vector comprises the heterologous nucleic acid sequence linked to any genetic elements required to achieve expression in the host cell. Genetic elements that may be included in the expression vector in this regard include a transcriptional termination region, one or more nucleic acid sequences encoding marker genes, one or more origins of replication, and the like. In some embodiments, the expression vector further comprises genetic elements required for the integration of the vector or a portion thereof in the host cell's genome.
[0109] It is also contemplated that in some embodiments an expression vector comprising a heterologous nucleic acid of the present disclosure may further contain a marker gene. Marker genes useful in accordance with the present disclosure include any genes that allow the distinction of transformed cells from non-transformed cells, including all selectable and screenable marker genes. A marker gene may be a resistance marker such as an antibiotic resistance marker against, for example, kanamycin or ampicillin. Screenable markers that may be employed to identify transformants through visual inspection include β-glucuronidase (GUS) (U.S. Pat. Nos.5,268,463 and 5,599,670) and green fluorescent protein (GFP) (Niedz et al., 1995, Plant Cell Rep., 14: 403). ‐ 64 ‐
[0110] In at least one embodiment, the present disclosure also provides of a method for producing THQ, wherein a heterologous nucleic acid encoding a recombinant polypeptide having CYP activity capable of converting carvacrol and / or thymol to THQ (e.g., an exemplary engineered polypeptide of Table 3B) can be introduced into a recombinant host cell. The recombinant host cell can then be used for production of the polypeptide or incorporated in a biocatalytic process that utilized the CYP activity of the recombinant polypeptide expressed by the host cell for the catalytic conversion of a substrate, e.g., the conversion of carvacrol and / or thymol. In at one embodiment, the recombinant host cell can further comprise a pathway of enzymes capable of producing a compound precursor (e.g., carvacrol) which can act as a substrate for the recombinant polypeptides with CYP activity and CPR activity. It is contemplated that a recombinant host cell comprising a heterologous nucleic acid encoding a recombinant polypeptide having CYP activity of the present disclosure can provide improved biosynthesis of a desired product compound ( e.g., THQ or THQ derivative) in terms of titer, yield, and production rate, due to the improved characteristics of the expressed CYP activity in the cell associated with the amino acid and codon differences engineered in the gene.
[0111] Accordingly, in at least one embodiment, the present disclosure provides a method for producing THQ comprising: (a) culturing a recombinant host cell of the present disclosure in a suitable medium comprising carvacrol and / or thymol; and (b) recovering the produced THQ.
[0112] In at least one embodiment, it is contemplated the engineered polypeptides derived from CYP029 (SEQ ID NO: 24) with CYP activity of the present disclosure can be incorporated in any biosynthesis method requiring a CYP catalyzed biocatalytic step, whether in vivo or in vitro. For example, in at least one embodiment, the recombinant engineered polypeptides having CYP activity (e.g., exemplary engineered CYP polypeptides of Table 3B) can be used in a method for preparing compound (1a) or a derivative of compound (1a) comprising contacting undera substrate compound (2a) and / or a substrate compound (2b) or a derivative of compound (2a) and / or a derivative of compound (2b), ‐ 65 ‐ with a recombinant engineered polypeptide with CYP activity, or a recombinant host cell expressing the engineered polypeptide, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 24, and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, Y457, D458, N472, and I477; optionally, wherein the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, Y457C, D458V, N472D, and I477V.
[0113] The method comprises contacting a recombinant engineered polypeptide derived from CYP029 (SEQ ID NO: 24) having CYP activity of the present disclosure (e.g., an exemplary polypeptide of Table 3B), or a recombinant host cell expressing such a polypeptide, under suitable reactions conditions with conditions compound (2a) and / or compound (2b), or a derivative of compound (2a) and / or a derivative of compound (2b). Exemplary conversions of THQ precursor compounds, carvacrol, compound (2a) and / or thymol, compound (2b) to THQ, compound (1a), and related structural analogs or derivatives of compound (1a) that are catalyzed by the recombinant polypeptides having CYP activity of the present disclosure can include: (1) conversion of thymol to THQ; (2) the conversion of carvacrol to THQ; and (3) conversion of a mixture of carvacrol and thymol to THQ. Accordingly, in at least one embodiment of the biosynthesis method for conversion a THQ precursor compound to THQ or a THQ structural analog or THQ derivative compound, the THQ precursor compound is carvacrol and / or thymol and the resulting product compound THQ.
[0114] The present disclosure also contemplates that the methods for biocatalytic conversion of a THQ precursor compound to a THQ or a THQ structural analog or a THQ ‐ 66 ‐ derivative compound using a recombinant engineered polypeptide derived from CYP029 (SEQ ID NO: 24) having CYP activity of the present disclosure (e.g., an exemplary polypeptide of Table 3B) can further comprise chemical or biocatalytic steps carried out on the product compound of the bioconversion, including steps of product compound work-up, extraction, isolation, purification, and / or crystallization, each of which can be carried out under a range of conditions. For example, the biocatalytic conversion can further comprise a further biocatalytic or chemical step to form a derivative of compound (1a), such as compound (1b), compound (1c), compound (1d), compound (1e), compound (1f), compound (1g), compound (1h), compound (1i), compound (1j), compound (1k), compound (1l), compound (1m), compound (1n), compound (1o), compound (1p), compound (1q), compound (1r), compound (1s), compound (1t), compound (1u), compound (1v), compound (1w), compound (1x), compound (1y), compound (1z), compound (1aa), compound (1ab), compound (1ac), compound (1ad), compound (1ae), compound (1af), compound (1ag), compound (1ah), and compound (1ai).
[0115] In at least one embodiment of the method, a derivative of compound (2a) and / or a derivative of compound (2b) is used to prepare a derivative of compound (1a), wherein the derivative of compound (2a) and / or a derivative of compound (2b), wherein the derivative is selected from compound (2c), compound (2d), compound (2e), compound (2f), compound (2g), compound (2h), compound (2i), compound (2j), compound (2k), compound (2l), compound (2m), compound (2n), compound (2o), compound (2p), and mixture thereof.
[0116] Suitable reaction conditions for the biosynthesis of compounds such as THQ and THQ derivatives are known in the art and can be used with a recombinant engineered polypeptide derived from CYP029 (SEQ ID NO: 24) having CYP activity of the present disclosure (e.g., an exemplary polypeptide of Table 3B) or recombinant host cells expressing such polypeptides of the present disclosure. Additionally, suitable reaction conditions for the exemplary polypeptides of the present disclosure can be determined using routine techniques known in the art for optimizing biocatalytic reactions. It is contemplated that various ranges of suitable reaction conditions with the recombinant polypeptides of the present disclosure, including but not limited to ranges of pH, temperature, buffer, solvent system, substrate loading, polypeptide loading, co-substrate or co-factor loading, atmosphere, and reaction time. Suitable reaction conditions can be readily determined and optimized for particular reactions by routine experimentation that includes, but is not limited to, contacting the recombinant polypeptide and substrate under experimental reaction conditions of concentration, pH, temperature, solvent conditions, and detecting the production of the desired compound of structural formula (I). In at least one embodiment, the suitable reaction conditions comprise a reaction solution of ~pH 7-8, a temperature of 25C to 37C; optionally, the reaction conditions comprise a reaction solution of ~ pH 7 and a ‐ 67 ‐ temperature of ~30C. In at least one embodiment, the reaction solution is allowed to incubate at a temperature of 25C to 37C for a reaction time of at least 1, 6, 12, 24, or 48 hours, before the amount of reaction product is determined.
[0117] Associated with the above-described biocatalytic methods, the disclosure also provides enzymatic reaction mixture compositions useful for enzymatic synthesis of the compound (1a) or a derivative of compound (1a). For example, in at least one embodiment the disclosure provides a composition comprising: (a) a recombinant engineered polypeptide derived from CYP029 (SEQ ID NO: 24) having CYP activity of the present disclosure (e.g., an exemplary polypeptide of Table 3B) or a recombinant host cell expressing such polypeptides of the present disclosure; and (b) a compound (2a) and / or compound (2b), or a derivative of compound (2a) and / or a derivative of compound (2b). In at least one embodiment of the composition, the engineered CYP polypeptide comprises an amino acid sequence of at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 24, and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, Y457, D458, N472, and I477; optionally, wherein the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, Y457C, D458V, N472D, and I477V. In at least one embodiment, the composition further comprises a recombinant polypeptide with CPR activity; optionally, a polypeptide comprising an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 6.
[0118] It is also contemplated that the composition can comprise other compounds useful in further biocatalytic or chemical steps, such as reagents useful in the production of derivatives of compound (1a), such as compound (1b), compound (1c), compound (1d), compound (1e), compound (1f), compound (1g), compound (1h), compound (1i), compound (1j), compound (1k), compound (1l), compound (1m), compound (1n), compound (1o), compound (1p), compound (1q), compound (1r), compound (1s), compound (1t), compound (1u), compound (1v), compound (1w), compound (1x), compound (1y), compound (1z), compound (1aa), compound (1ab), compound (1ac), compound (1ad), compound (1ae), compound (1af), compound (1ag), compound (1ah), and compound (1ai). Accordingly, in at least one embodiment, the composition can comprise a derivative of compound (2a) and / or ‐ 68 ‐ a derivative of compound (2b), for example a derivative selected from compound (2c), compound (2d), compound (2e), compound (2f), compound (2g), compound (2h), compound (2i), compound (2j), compound (2k), compound (2l), compound (2m), compound (2n), compound (2o), compound (2p), and mixture thereof. EXAMPLES
[0119] Various features and embodiments of the disclosure are illustrated in the following representative examples, which are intended to be illustrative, and not limiting. Those skilled in the art will readily appreciate that the specific examples are only illustrative of the invention as described more fully in the claims which follow thereafter. Every embodiment and feature described in the application should be understood to be interchangeable and combinable with every embodiment contained within. Example 1: Establishment of CYP Activity for Conversion of Carvacrol and / or Thymol to Thymohydroquinone (THQ)
[0120] This example illustrates establishment of baseline activity of two CYP polypeptides from Origanum vulgare, CYP76S40 (CYP002; SEQ ID NO: 2) and CYP736A300 (CYP003; SEQ ID NO: 4) which have previously been shown to convert carvacrol and / or thymol to THQ in heterologous yeast systems (see e.g., Krause, et al. “The Biosynthesis of Thymol, Carvacrol, and Thymohydroquinone in Lamiaceae Proceeds via Cytochrome P450s and a Short-Chain Dehydrogenase.” Proceedings of the National Academy of Sciences 118, no. 52, December 28, 2021).
[0121] Material and Methods
[0122] A. Construction of control S. cerevisiae strains with integrated CYP genes previously identified.
[0123] Yeast codon-optimized genes encoding the enzymes CYP76S40 (SEQ ID NO: 1) and CYP736A300 (SEQ ID NO: 3) with previously reported activity to convert carvacrol and / or thymol to THQ, were synthesized by TWIST. The synthesized genes (also referred to herein as CYP002 and CYP003, respectively), were further amplified to contain homology sequences to the S. cerevisiae pGAL1 promoter and tPGK1 terminator with the wbOligos5915 and wbOligos5917 primers to create donor cassettes and further integrated into the X-4 locus of two distinct screening strains as described below.
[0124] A first screening strain (OG002) was constructed expressing two landing sites at the X-4 locus, m-Venus and URA3. This OG002 screening strain did not contain a heterologous CPR reductase to function as the redox partner required for activity of CYP enzymes, instead relying on the endogenous yeast reductases (PGA3, AIM33, MCR1, CYB5, NCP1) to provide this CPR activity. FIG.1A shows a schematic depiction of the X-4 insertion site of OG002 containing m-Venus and the URA3 gene under the bidirectional pGal10 / 1 promoter. ‐ 69 ‐ This strain was not capable of converting carvacrol and / or thymol to THQ and therefore was also used as the negative control.
[0125] CYP002 and CYP003 integrated into OG002 generated strains OG008 and OG009 respectively. A second screening strain (OG003) was constructed to include a heterologous cytochrome P450 reductase (CPR006, SEQ ID NO: 6) to function as the redox partner required to work in tandem with the recombinant CYP enzyme. CPR006 (SEQ ID NO: 6) is the ATR1 CPR polypeptide from Arabidopsis thaliana and is the CPR sequence reported to be functioning along with CYP76S40 and CYP736A300, in a strain capable of converting carvacrol and / or thymol to THQ (Krause, et al, 2021). FIG.1B shows a schematic depiction of the X-4 insertion site of OG003 containing CPR006 (SEQ ID NO: 6) and the URA3 gene (which was used as a landing site to integrate the two CYP genes) under the bidirectional pGal10 / 1 promoter. Using the Gal10 / Gal1 promoters, these screening systems could be induced with galactose and / or via glucose depletion.
[0126] CYP002 (SEQ ID NO: 2) and CYP003 (SEQ ID NO: 4) integrated into OG003 generated strains OG004 and OG005 respectively. Additional copies of the resulting ADH1t_CPR006-pGAL10-pGAL1_CYP002_PGK1t expression cassette were targeted to integration sites XI-2, XII-4, and X-2. Individual donors for integrating into each of these sites were generated by amplifying the CPR006-CYP002 cassette from OG004 then adding site- specific integration homology at both the 5’ and 3’ ends by overlap extension PCR. PCR fragments targeting the three different loci were amplified using the primers listed in Table 4 and the accompanying Sequence Listing.
[0127] TABLE 4 SEQ ID SEQ ID PCR Fragment Forward primer NO: Reverse Primer NO:
[0128] Primers used in overlap extension PCR are listed in Table 5 and the accompanying Sequence Listing.
[0129] TABLE 5 SEQ ID SEQ ID‐ 70 ‐
[0130] The resulting donors were transformed as a pool into OG004 which already contained a single copy of CPR006 and CYP002 at the X-4 locus as previously described. Individual clones were screened for THQ production using an HTP protocol (as described in Example 2C) and in shake flasks as described in section B below. A PCR confirmed two- copy strain, OGS034 (copies of CPR006 / CYP002 verified at X-4 and X-2) and a three-copy strain OGS035 (copies of CPR006 / CYP002 verified at X-4, XI-2, and XII-4) were identified with respective conversion shown in Table 7 (below).
[0131] B. Screening of initial control strains for THQ bioconversion in Shake Flasks
[0132] To screen initial control strains, a shake flask protocol was developed to detect and quantify the bioconversion of carvacrol to THQ from the recombinant S. cerevisiae strains. Individual colonies of recombinant S. cerevisiae strains were picked into 500 mL flasks containing 2 x YPD media (100 mL working volume). The flasks were incubated at 30oC for 48 h with shaking at 250 rpm at 85 % humidity. After this time, the strains were induced via the addition of galactose (1 % final concentration) from a 40 % stock solution. The flasks were incubated for a further 4 h before the biomass was harvested via centrifugation and the supernatant removed. A portion of the resulting biomass (2 g) was resuspended in 10 mL of bioconversion buffer (50 mg / L carvacrol, 0.1 M phosphate buffer, pH 7) in a 125 mL flask. Bioconversion was conducted at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, a 300 ^L aliquot was taken and extracted with MeOH (300 ^L) with shaking at 250 rpm at 30oC for a further 30 minutes. The biomass was then separated by centrifugation and the resulting supernatant diluted with MeOH (3 x total dilution). The samples were then loaded onto an Agilent 1290 Infinity II UHPLC equipped with a DAD and the compounds of interest were detected by monitoring at 274 nm and quantified relative to calibration curves which were prepared using serial dilutions of stock solutions of carvacrol and THQ.
[0133] UHPLC Instrumentation and parameters: UHPLC system: Agilent 1290 Infinity II UHPLC; Column: Agilent EclipsePlusC18 column 1.8 um, 3.0 x 50 mm; Column temperature: 40°C; Flow Rate: 1.1 mL / min; Mobile phase A: 100% water + 0.01% FA; Mobile phase B: 100% acetonitirile + 0.01% FA; Gradient: 0 - 0.5 min isocratic 30% B, 0.5 - 2.0 min gradient to 100% B, 2.0 - 2.5 min isocratic 100% B; UV detection wavelength: 274 nm, 4nm bandwidth.
[0134] Results
[0135] Exemplary UHPLC profiles for the strains OG004, OG005, OG008 and OG009 are shown in FIG.2. The results of screening assays in shake flasks for the four S. cerevisiae strain builds (OG004, OG005, OG008, and OG009), along with the OG002 negative control are summarized in Table 6 below. ‐ 71 ‐
[0136] TABLE 6 Substrate Strain CYP ID CPR ID Loading Conversion OG004 CYP76S40 / CYP002 CPR006 50 mg / L 9.3%summarized in Table 7 below.
[0138] TABLE 7 Carvacrol to THQ Strain Copies of CYP002 Carvacrol Loading (mg / L) % Conversionfor Conversion of Carvacrol and / or Thymol to Thymohydroquinone (THQ)
[0139] This example illustrates a study to identify and screen candidate genes from various organisms that encode polypeptides with CYP activity capable of converting carvacrol and / or thymol to THQ when expressed heterologously in yeast.
[0140] Material and Methods
[0141] A. Search for candidate CYP homologs
[0142] The amino acid sequences of CYP002 (SEQ ID NO: 2) and CYP003 (SEQ ID NO: 4) were used to conduct sequence searches with a 60% sequence identity cut-off for genes encoding homologous polypeptide sequences. A total of 128 unique homologs from 45 different plant species were identified from public databases (NCBI, OneKP), including 28 homologs from 14 species that had > 70% sequence homology.96 of these unique homologs were selected to be synthesized and further evaluated for activity to convert carvacrol and / or thymol to THQ.
[0143] B. Construction of S. cerevisiae strains with integrated CYP homolog genes
[0144] A third screening strain (OG014) was constructed to include a non-functional cytochrome P450 reductase (CPR006 with 1 base pair deletion causing a frameshift) unable to act as the redox partner, required to work in tandem with a CYP enzyme. FIG.1C shows a schematic depiction of the X-4 insertion site of OG014 containing a non-functional CPR006 ‐ 72 ‐ and the URA3 gene (which was used as a landing site to integrate various CYP homolog genes) under the bidirectional pGal10 / 1 promoter.
[0145] The 96 candidate homolog genes were first amplified using primers wbOligos5945 (SEQ ID NO: 77) and wbOligos5377 (SEQ ID NO: 57). To facilitate integration, a fragment containing the pGAL10 / 1 promoter and a fragment containing the tPGK1 terminator and a portion of the S. cerevisiae X4 locus were added on to the amplified fragments via overlap- extension PCR using primers wbOligos4668 (SEQ ID NO: 56) and wbOligos1135 (SEQ ID NO: 55) to create donor cassettes. The resulting donor cassettes were transformed into two different screening strains, OG003 and OG014 to identify CYP enzymes that had activity with and active CPR006 or inactive CPR006 (endogenous reductases) respectively.
[0146] C. HTP screening of strains for THQ bioconversion
[0147] A HTP screening assay was developed to detect and quantify the bioconversion of carvacrol to THQ from the recombinant S. cerevisiae strains. Individual colonies of recombinant S. cerevisiae strains were picked into 96-well plates containing 2 x YPD media (300 ^L per well) using a QPixTM420 colony picking system. The plates were incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, the strains were sub-cultured (40 ^L inoculation volume per strain) into a second 96-well plate containing 2 x YPD media (1.2 mL per well) using an Agilent Bravo automated liquid handling platform. The plates were incubated at 30oC for 48 h with shaking at 250 rpm at 85 % humidity. After this time, a solution of galactose (40 % stock solution) was added using the Bravo (1 % final concentration) and the plates were further incubated for 4 hours to induce protein production. After this time, the biomass was harvested via centrifugation and the supernatant removed. The resulting biomass was resuspended in 300 ^L of bioconversion buffer (50 mg / L carvacrol, 0.1 M phosphate buffer, pH 7) using the Agilent Bravo automated liquid handling platform and the plates were incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, MeOH (300 ^L) was added using the Bravo and the plates were incubated at 30oC with shaking at 250 rpm at 85 % humidity for a further 30 minutes. The biomass was then separated by centrifugation and the resulting supernatant diluted with MeOH (3 x total dilution) using the Agilent Bravo automated liquid handling platform. The samples were then loaded onto an Agilent 1290 Infinity II UHPLC equipped with DAD and the compounds of interest were detected by monitoring at 274 nm and quantified relative to calibration curves which were prepared using serial dilutions of stock solutions of carvacrol and THQ.
[0148] UHPLC Instrumentation and parameters: Identical to those described in Example 1 (above).
[0149] Results ‐ 73 ‐
[0150] The 96 CYP homologs were screened in S. cerevisiae (integrated to replace URA3 in OG003 and OG014) and compared to the control strain, OG004, which contains the gene encoding CYP002 (SEQ ID NO: 1) and CPR006 (SEQ ID NO: 5). 15 unique CYP homologs were identified with varying activity toward converting carvacrol and / or thymol to THQ. Of these 15 homologs the CYP identified from Salvia hispanica, CYP029 (SEQ ID NO: 24), was shown to fully convert 50mg / L of carvacrol substrate to THQ.
[0151] Results obtained with screening strain 1 (OG003) with an integrated functional CPR006 gene are shown in Table 5 below.
[0152] TABLE 5 AA Carvacrol to THQ Carvacrol to THQ SEQ ID % Conversion % Conversion ial CPR006 gene (endogenous reductases only) are shown in Table 6 below.
[0154] TABLE 6 AA Carvacrol to THQ Carvacrol to THQ SEQ ID % Conversion % ConversionExample 3: Construction of Pichia Strains with CYP Activity for Conversion of Carvacrol and / or Thymol to Thymohydroquinone (THQ) ‐ 74 ‐
[0155] This example illustrates a study to build strains of Pichia pastoris with heterologous genes encoding an enzyme with CYP activity capable of converting carvacrol and / or thymol to THQ.
[0156] Material and Methods
[0157] A. Construction of Pichia strains expressing CYP genes capable of converting carvacrol and / or thymol to THQ
[0158] To improve upon the bioconversion of carvacrol and / or thymol to THQ in S. cerevisiae, the genes encoding CYP002 (SEQ ID NO: 1) and CPR006 (SEQ ID NO: 6) were integrated into the P. pastoris genome under the bidirectional pCAT1:pFDH1 promoter system (SEQ ID NO: 83) and tDAS1 and tDAS2 terminators SEQ ID NOs: 84 and 85, respectively, to generate a single copy strain (OGP010) as follows. A single plasmid (wbplasmid162; SEQ ID NO: 79) containing a zeocin resistance marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 86), and a cassette comprising of tDAS1:CPR006:pCAT1:pFDH1:CYP002:tDAS2 was assembled. Sequences were amplified using PCR using the following primers: wboligos5866 (SEQ ID NO: 58), wboligos5867 (SEQ ID NO: 59), wboligos5868 (SEQ ID NO: 60), wboligos5869 (SEQ ID NO: 61), wboligos5870 (SEQ ID NO: 62), wboligos5871 (SEQ ID NO: 63), wboligos5872 (SEQ ID NO: 64), and wboligos5873 (SEQ ID NO: 65).
[0159] This was followed by agarose-gel purification and assembly using NEB’s HiFi DNA Assembly Master Mix (catalog no. E2621X) according to manufacturer’s instructions. Three microliters of the HiFi assembly were used to transform E. coli competent cells and plated on low-salt LB media containing zeocin. Plasmids were extracted from cultures inoculated by colonies using GeneJet Plasmid Miniprep Kit (ThermoFisher catalog no. K0503) and sequence verified. The sequence-confirmed plasmid was digested with BbvcI and cleaned using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The cut site, found in the HIS4 homologous sequence, linearized the plasmid, allowing for homologous integration at the HIS4 locus in Pichia competent cells, which were prepared from the ^aox1 BG-11 strain using Thermo’s Pichia EasyComp Transformation Kit (catalog no. K173001) according to manufacturer’s instructions. Briefly, 5 μL of linearized plasmid (5 μg total) was used to transform a 50 μL aliquot of BG-11 competent cells and plated on YPDS media plates containing zeocin. After 3 days of growth, colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each colony.
[0160] A recombinant Pichia host cell strain (OGP011) with one copy of CYP003 (SEQ ID NO: 3) and CPR006 (SEQ ID NO: 5) was also created as follows. A single plasmid (wbplasmid163; SEQ ID NO: 80) containing a zeocin resistance marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 86), and a cassette comprising of tDAS1:CPR006:pCAT1:pFDH1:CYP003:tDAS2 was assembled with the method described ‐ 75 ‐ above. Sequences were amplified using PCR using the following primers: wboligos5866, wboligos5867, wboligos5868, wboligos5869, wboligos5870, wboligos5871, wboligos5872, and wboligos5873.
[0161] Following plasmid recovery and sequence confirmation, a single plasmid was digested with BbvcI and cleaned using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The cut site, found in the HIS4 homologous sequence, linearized the plasmid, allowing for homologous integration at the HIS4 locus in Pichia competent cells, which were prepared from the ^aox1 BG-11 strain using Thermo’s Pichia EasyComp Transformation Kit (catalog no. K173001) according to manufacturer’s instructions. Briefly, 5 microliters of linearized plasmid (5 ug total) were used to transform a 50-microliter aliquot of BG-11 competent cells and plated on YPDS media plates containing zeocin. After 3 days of growth, colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each colony.
[0162] A recombinant Pichia host cell strain (OGP024) with one copy of a gene encoding CYP056 (SEQ ID NO: 28) and a gene encoding CPR006 (SEQ ID NO: 6) was also created as follows. A single plasmid (wbplasmid192; SEQ ID NO: 82) containing a zeocin resistance marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 86), and a cassette comprising of tDAS1:CPR006:pCAT1:pFDH1:CYP056:tDAS2 was assembled with the method described above. As wbplasmid192 (SEQ ID NO: 82) shares the same sequence as wbplasmid162 (SEQ ID NO: 79) except for the CYP002 CDS, inverse PCR was used to amplify the vector minus the CYP002 CDS using primers wboligos6113 (SEQ ID NO: 76) and wboligos6114 (SEQ ID NO: 77). CYP056 (SEQ ID NO: 28) was amplified from S. cerevisiae strain OG026 using primers wboligos6112 (SEQ ID NO: 75) and wboligos6115 (SEQ ID NO: 78). Following plasmid recovery and sequence confirmation, a single plasmid was digested with BbvcI and cleaned using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The cut site, found in the HIS4 homologous sequence, linearized the plasmid, allowing for homologous integration at the HIS4 locus in Pichia competent cells, which were prepared from the ^aox1 BG-11 strain using Thermo’s Pichia EasyComp Transformation Kit (catalog no. K173001) according to manufacturer’s instructions. Briefly, 5 microliters of linearized plasmid (5 ug total) were used to transform a 50 microliter aliquot of BG-11 competent cells and plated on YPDS media plates containing zeocin. After 3 days of growth, colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each colony.
[0163] A recombinant Pichia host cell strain (OGP025) with one copy of a gene encoding CYP029 (SEQ ID NO: 24) and a gene encoding CPR006 (SEQ ID NO: 6) was also created as follows. A single plasmid (wbplasmid191; SEQ ID NO: 81) containing a zeocin resistance ‐ 76 ‐ marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 86), and a cassette comprising of tDAS1:CPR006:pCAT1:pFDH1:CYP029:tDAS2 was assembled with the method described above. As wbplasmid191 shares the same sequence as wbplasmid162 except for the CYP002 CDS, inverse PCR was used to amplify the vector minus the CYP002 CDS using primers wboligos6103 and wboligos6100. The gene encoding CYP056 (SEQ ID NO: 28) was amplified from S. cerevisiae strain OG027 gDNA using primers wboligos6101 (SEQ ID NO: 72) and wboligos6102 (SEQ ID NO: 73). Following plasmid recovery and sequence confirmation, a single plasmid was digested with BbvcI and cleaned using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The cut site, found in the HIS4 homologous sequence, linearized the plasmid, allowing for homologous integration at the HIS4 locus in Pichia competent cells, which were prepared from the ^aox1 BG-11 strain using Thermo’s Pichia EasyComp Transformation Kit (catalog no. K173001) according to manufacturer’s instructions. Briefly, 5 μL of linearized plasmid (5 μg total) was used to transform a 50 μL aliquot of BG-11 competent cells and plated on YPDS media plates containing zeocin. After 3 days of growth, colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each colony.
[0164] B. Analysis of strains for THQ bioconversion
[0165] Screening of the recombinant Pichia strains for bioconversion of carvacrol and / or thymol to THQ was carried out according to the following assay: individual colonies of recombinant P. pastoris strains were picked into 500 mL baffled conical flasks containing BMGY media (100 mL working volume). The flasks were incubated at 30oC for 48 h with shaking at 250 rpm at 85% humidity and were supplemented with glycerol (2 % final concentration) after 24 hours. After this time, the biomass was harvested via centrifugation and the supernatant was discarded. A portion of the resulting biomass (0.5 g) was resuspended in 2.5 mL of bioconversion buffer (0.1 M phosphate buffer, 2% MeOH, pH 7) in a 20 mL scintillation vial and carvacrol (50-500 mg / L) was added. The vial was incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. The extraction, sample preparation and analytical methods are identical to those described in above examples.
[0166] Results
[0167] Results obtained from recombinant Pichia strains are shown below in Table 7.
[0168] TABLE 7 Strain ID CYP AA Carvacrol Loading Carvacrol to THQ CYP ID SEQ ID NO: (m / L) % C nv r i n‐ 77 ‐ OGP025 CYP029 24 100 100 OGP025 CYP029 24 250 95Host Cells by Site Saturation Mutagenesis (SSM)
[0169] This example illustrates the preparation of site saturation mutagenesis (SSM) libraries of engineered polypeptides derived from the parent CYP029 polypeptide of SEQ ID NO: 24 in the OGP074 screening strain. The resulting libraries were screened for improved activity in the bioconversion of carvacrol to THQ relative to the bioconversion of the parent strain OGP075 containing the parent CYP029 polypeptide of SEQ ID NO: 24 and the CPR006 polypeptide of SEQ ID NO: 6.
[0170] Material and Methods
[0171] A. Site Saturation Mutagenesis library construction
[0172] The heterologous nucleic acid sequence which encodes the CYP polypeptide of SEQ ID NO: 24, expressed under the pAOX1 promoter (SEQ ID NO: 217) and tDAS2 terminator (SEQ ID NO: 85), was used to generate SSM libraries at amino acid positions spanning the entire polypeptide sequence. To generate an appropriate screening strain to evaluate these libraries, the mVenus gene (SEQ ID NO.198) was integrated into the Intergenic Region 15 site of a Pichia strain under the pAOX1 promoter and the tDAS2 terminator, along with the codon optimized polynucleotide sequence of CPR006 (SEQ ID NO:6), under the pAOX1 promoter and tPMP20 terminator (SEQ ID NO: 219) integrated into Intergenic Region 21. The resulting strain, OGP074, was used as the screening host for Integration and screening of the resulting SSM libraries.
[0173] The SSM libraries were synthesized in-house using the Telesis BioXp® 9600 system with each variant (Fragment B) containing a single amino acid change relative to the parent CYP (SEQ ID No.24). For efficient integration of the SSM CYP variants, PCR amplicons with homologous 5’ and 3’ regions to the SSM library were generated (Fragment A and Fragment C respectively). Fragment A was amplified using the forward primer wboligos7178 (SEQ ID NO: 210) and reverse primer wboligos5732 (SEQ ID NO: 211) and Fragment C was amplified using the forward primer wboligos6862 (SEQ ID NO: 212) and the reverse primer wboligos7183 (SEQ ID NO: 213). Fragments A, B, and C were further assembled by NEBuilder® HiFi Assembly and amplified as a single linear DONOR using the following forward rescue primer wboligos7179 (SEQ ID NO: 214) and reverse primer wboligos7182 ‐ 78 ‐ (SEQ ID NO: 215). The assembled PCR products were then pooled together, and gel purified to provide a SSM library of linear donor DNA.
[0174] The sequences of the PCR primers described are listed Table 8 below and the accompanying Sequence Listing.
[0175] TABLE 8 SEQ ID Primer Name Primer Se uence NO: 012345integrated using CRISPR-Cas9 into the Intergenic Region 15 loci replacing the mVenus gene in OGP074.
[0177] B. HTP Screening of SSM library
[0178] A HTP screening assay was developed to detect and quantify the bioconversion of carvacrol to THQ from the recombinant Pichia pastoris strains. Individual colonies of recombinant P. pastoris strains were picked into 96-well plates containing 2x YPD media (300 ^L per well) using a QPixTM420 colony picking system. The plates were incubated at 30oC for 48 h with shaking at 250 rpm at 85 % humidity. After this time, the strains were sub-cultured (30 ^L inoculation volume per strain) into a second 96-well plate containing 1 x BMGY media (300 µL per well) using an Integra Viafill automated liquid handling platform. The plates were incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, the biomass was harvested via centrifugation and the supernatant removed. The resulting biomass was resuspended in 300 ^L of bioconversion buffer (500-700 mg / L carvacrol, 0.1 M phosphate buffer, pH 7, 2% MeOH, 10% IPM) using the Agilent Bravo automated liquid handling platform and the plates were incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, IPA (300 ^L) was added using the Bravo and the plates were incubated at 30oC with shaking at 250 rpm at 85 % humidity for a further 30 minutes. The biomass was then separated by centrifugation and the resulting supernatant diluted with MeOH (20 x total dilution) using the Agilent Bravo automated liquid handling platform before analysis. The samples were then loaded onto an Agilent 1290 ‐ 79 ‐ Infinity II UHPLC equipped with DAD and the compounds of interest were detected by monitoring at 274 nm and quantified relative to calibration curves which were prepared using serial dilutions of stock solutions of carvacrol and THQ.
[0179] Evaluation of SSM hits at SF scale
[0180] Evaluation of the SSM hits for bioconversion of carvacrol to THQ at shake flask scale was carried out according to the following assay: individual colonies of recombinant P. pastoris strains were picked into 500 mL baffled conical flasks containing BMGY media (100 mL working volume). The flasks were incubated at 30oC for 48 h with shaking at 250 rpm at 85% humidity and were supplemented with glycerol (2 % final concentration) after 24 hours. After this time, the biomass was harvested via centrifugation and the supernatant was discarded. A portion of the resulting biomass (0.5 g) was resuspended in 2.25 mL of bioconversion buffer (0.1 M phosphate buffer, 2% MeOH, pH 7) in a 20 mL scintillation vial and carvacrol (2 g / L) was added as a 10X IPM stock. The vial was incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, IPA (5 mL) was added and the vials were incubated at 30oC with shaking at 250 rpm at 85 % humidity for a further 30 minutes. The samples were then aliquoted into a 96-well plate before the biomass was separated by centrifugation and the resulting supernatant diluted with MeOH / water (3:1) (30 x total dilution) using the Agilent Bravo automated liquid handling platform before analysis using the UHPLC-DAD method described above.
[0181] Evaluation of SSM hits at Bioreactor scale
[0182] The recombinant Pichia pastoris strain was grown at 2 L scale in a glass jacketed fermenter according to the following protocol. The process consists of two consecutive shake flask cultures to expand a glycerol stock to a 100 mL culture to an OD600 of 30-50 to inoculate the growth tank. The Inoculum is added to a prepared sterile 2 L fermenter containing 1561 mL of the complex medium BMGY containing 4% glycerol as carbon source and Biotin supplement. The growth fermenter is controlled at 30 °C, 30 % dissolved oxygen and pH 5 throughout. There are 2 stages to the fermentation: Batch phase, until the initial 4% glycerol is consumed at approximately 16 hours and an additional fed-batch phase where 50% V / V glycerol is fed at 10 g / L / hr rate to increase biomass to the desired cell density. After glycerol feeding, the cells are harvested using centrifugation. Before setting up the bioconversion reaction, each 1L Dasgip bioreactor vessel was autoclaved with calibrated pH and DO% probes. After sterilization and calibration, the vessel was charged with 30% (W / V) biomass (90 g) obtained using the above-described fermentation protocol. The bioconversion buffer was then prepared and added to the vessel. The bioconversion buffer consists of 0.1 M phosphate buffer (pH 7) and MeOH (2% v / v). A total of 270 mL of bioconversion buffer was added to the vessel containing the biomass followed by carvacrol as a 10X stock in IPM (7 g / L final concentration). The vessel was then transferred to the ‐ 80 ‐ Dasgip stirring block where it was controlled and monitored. The reaction was stirred at 1000 RPM for 24 h at 30oC and monitoring %DO and pH profiles. Filtered clean dry air was sparged at 1.4 vvm for the duration. After MeOH in bioconversion buffer was consumed, signaled by an abrupt spike in DO% around 4 hours, pure MeOH was fed using DO% stat profile, adding MeOH when dissolved oxygen % rises above 30%. Aliquots (0.3 mL) were taken at regular time intervals for analysis using the UHPLC method described above.
[0183] UHPLC Instrumentation and parameters: Identical to those described in Example 1 (above).
[0184] Results
[0185] Data from HTP screening of SSM libraries in terms of fold-improvement in production of THQ from carvacrol relative to the control strain, OGP075, are summarized Table 9 (below).
[0186] TABLE 9 Carvacrol FIOPC NT AA and / or Carvacrol to‐ 81 ‐ libOGP074_SC05_115 116 T307Sn / a 65% 1.55B_20X:A3 libOGP074 SC06117 118 Y345Nn / a 64% 1.51d isproduction of THQ from carvacrol relative to the control strain, OGP075, are summarized Table 10 (below).
[0188] TABLE 10 Strain Carvacrol CodeP. pastoris Strain DescriptionLoading % C ion‐ 82 ‐
[0189] Data from bioreactor screening of SSM hits in terms of fold-improvement in production of THQ from carvacrol relative to the control strain, OGP075, are summarized Table 11 (below).
[0190] TABLE 11 Strain Carvacrol CodeP. pastoris Strain DescriptionLoading % (m / L) Conversiont P. pastoris Host Cells by Combinatorial Mutagenesis
[0191] This example illustrates the preparation of combinatorial libraries of engineered CYP genes based on parent strain OGP098 which carries a gene encoding the M464G mutant of CYP029 named “CYP301” (SEQ ID NO: 130) that was identified through SSM in Example 4 and shows improved bioconversion of carvacrol and / or thymol to THQ.
[0192] Material and Methods
[0193] A. Combinatorial library construction
[0194] The polynucleotide sequence of SEQ ID NO: 129 encoding the engineered CYP301 polypeptide (SEQ ID NO: 130) from OGP098 expressed under the pAOX1 promoter (SEQ ID NO: 217) and tDAS2 terminator (SEQ ID NO: 85), was used to design oligos to generate combinatorial libraries to randomly incorporate additional identified beneficial amino acid mutations identified from Example 4 into the CYP301 parent sequence. To generate an appropriate screening strain to evaluate these libraries, the ku70 gene was knocked out in the OGP074 strain via CRISPR-Cas9 with a gRNA targeting the ku70 gene. The resulting strain OGP084 was used as the screening host for CYP combinatorial libraries. The OGP098 strain comprising sequences encoding CYP301 and CPR006 was used as the parent control strain in order to determine fold-improvement in conversion of carvacrol to THQ as described below.
[0195] A semi-synthetic approach was used to construct the first set of combinatorial libraries. Genomic DNA from the OGP098 strain was used as the template to generate a full- length PCR product using primer pair of wboligo7751 (SEQ ID NO: 216) and wboligo7752 (SEQ ID NO: 217) while incorporating uracil using a dNTP mix comprising of the following deoxyribonucleotides: dATP, dGTP, dCTP, dTTP, dUTP. The resulting PCR product was column purified and digested with Uracil-DNA Glycosylase and Endonuclease IV at 37 C for ‐ 83 ‐ 2 hours, followed by enzyme denaturation at 94 C for two minutes, to generate a pool of fragments in the range of 50-100 bases. These fragments were further combined with differing ratios of pools of the synthesized oligonucleotide primers (each oligo up to 60 bases in length and encoding one or more amino acid change) in several individual assembly PCR reactions using forward primer wboligos6242 (SEQ ID NO: 218) and reverse primer wboligos6102 (SEQ ID NO: 219) to reassemble the full-length PCR product (Fragment B) and incorporate a combination of amino acid changes within each pool randomly. The sequences of the PCR primers described are listed Table 12 below and the accompanying Sequence Listing.
[0196] TABLE 12 SEQ ID i i :and their encoded mutations are listed in Table 13 below and the accompanying Sequence Listing.
[0198] TABLE 13 SEQ Name Mutation(s) Length Primer Sequence ID NO:‐ 84 ‐ _lib1 L99A60AGGACTGTAGCGTCTGCAATAACTGCACTACGC228 V112V CACCACGAATTCTCGGTGGGATGGCTT‐ 85 ‐ 30_lib1 S299A 43TTGCTGGCACCGATACCACCGCGGGAACTATGG249 AATGGGCCATg , p with homologous 5’ and 3’ regions to the CYP were generated (Fragment A and Fragment C respectively). Fragment A was amplified using the forward primer wboligos7178 (SEQ ID NO: 210) and reverse primer wboligos5732 (SEQ ID NO: 211) and Fragment C was amplified using the forward primer wboligos6862 (SEQ ID NO: 212) and the reverse primer wboligos7183 (SEQ ID NO: 213). Fragments A, B, and C were assembled by NEBuilder® HiFi Assembly and amplified as a single linear DONOR using the following forward primer wboligos7179 (SEQ ID NO: 214) and reverse primer wboligos7182 (SEQ ID NO: 215). The assembled PCR products were then pooled together, and gel purified to provide a combinatorial library of linear donor DNA.
[0200] The pooled linear donor DNA of combinatorial libraries were integrated using CRISPR-Cas9 into the Int15 locus in place of mVenus in OGP084.
[0201] B. HTP Screening of combinatorial library
[0202] Screening of the recombinant host cells for bioconversion of carvacrol to THQ was carried out using the same protocol described in Example 4 with the exception that the ‐ 86 ‐ carvacrol loading was increased from 500-700 mg / L to 1000 mg / L and the IPM was increased from 10% to 20%.
[0203] Screening of combinatorial library hits at SF scale
[0204] Screening of the combinatorial library hits for bioconversion of carvacrol to THQ at shake flask scale was carried out according to the following assay: individual colonies of recombinant P. pastoris strains were picked into 500 mL baffled conical flasks containing BMGY media (100 mL working volume). The flasks were incubated at 30oC for 48 h with shaking at 250 rpm at 85% humidity and were supplemented with glycerol (2 % final concentration) after 24 hours. After this time, the biomass was harvested via centrifugation and the supernatant was discarded. A portion of the resulting biomass (0.5 g) was resuspended in 2.25 mL of bioconversion buffer (0.1 M phosphate buffer, 2% MeOH, pH 7) in a 20 mL scintillation vial and carvacrol (2 g / L) was added as a 10X IPM stock. The vial was incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, IPA (5 mL) was added and the vials were incubated at 30oC with shaking at 250 rpm at 85 % humidity for a further 30 minutes. The samples were then aliquoted into a 96-well plate before the biomass was separated by centrifugation and the resulting supernatant diluted with MeOH / water (3:1) (30 x total dilution) using the Agilent Bravo automated liquid handling platform before analysis using the UHPLC-DAD method described above.
[0205] Screening of Round 1 combinatorial library hits at Bioreactor scale
[0206] The recombinant Pichia pastoris strain was grown at 2 L scale in a glass jacketed fermenter according to the following protocol. The process consists of two consecutive shake flask cultures to expand a glycerol stock to a 100 mL culture to an OD600 of 30-50 to inoculate the growth tank. The Inoculum is added to a prepared sterile 2 L fermenter containing 1561 mL of the complex medium BMGY containing 4% glycerol as carbon source and Biotin supplement. The growth fermenter is controlled at 30 °C, 30 % dissolved oxygen and pH 5 throughout. There are 2 stages to the fermentation: Batch phase, until the initial 4% glycerol is consumed at approximately 16 hours and an additional fed-batch phase where 50% V / V glycerol is fed at 10 g / L / hr rate to increase biomass to the desired cell density. After glycerol feeding, the cells are harvested using centrifugation. Before setting up the bioconversion reaction, each 1L Dasgip bioreactor vessel was autoclaved with calibrated pH and DO% probes. After sterilization and calibration, the vessel was charged with 30% (W / V) biomass (90 g) obtained using the above-described fermentation protocol. The bioconversion buffer was then prepared and added to the vessel. The bioconversion buffer consists of 0.1 M phosphate buffer (pH 7) and MeOH (2% v / v). A total of 270 mL of bioconversion buffer was added to the vessel containing the biomass followed by carvacrol as a 10X stock in IPM (7 g / L final concentration). The vessel was then transferred to the Dasgip stirring block where it was controlled and monitored. The reaction was stirred at 1000 ‐ 87 ‐ RPM for 24 h at 30oC and monitoring %DO and pH profiles. Filtered clean dry air was sparged at 1.4 vvm for the duration. After MeOH in bioconversion buffer was consumed, signaled by an abrupt spike in DO% around 4 hours, pure MeOH was fed using DO% stat profile, adding MeOH when dissolved oxygen % rises above 30%. Aliquots (0.3 mL) were taken at regular time intervals for analysis using the UHPLC method described above.
[0207] Results
[0208] Screening of the combinatorial libraries for fold-improvement in production of THQ from carvacrol relative to the control strain, OGP098, which expresses the M464G CYP polypeptide of CYP301 (SEQ ID NO: 130) and the CPR polypeptide of SEQ ID NO: 6, are summarized in Table 14 (below).
[0209] TABLE 14 Carvacrol NT AA and / or FIOPC l ol‐ 88 ‐ hlibOGP084 149 150 T196A, M464G V112V (GTG), 48% 1.02 _SC2_B_50x P437P (CCT) :H4 d is
[0210] Data from shake flask evaluation of combinatorial libraries hits in terms of fold- improvement in production of THQ from carvacrol relative to the control strain, OGP098, are summarized Table 15 (below).
[0211] TABLE 15 Strain Carvacrol Pt i t i D i tiL di % n‐ 89 ‐ OGP099 One copy of CYP301 combi mutant at Int15 locus 2000 30 and one copy of CPR006 at Int21 locus OGP100 One copy of CYP302 (CYP301 combi mutant) at 2000 56improvement in production of THQ from carvacrol relative to a previous high-performing control strain, OGP087, are summarized Table 16 (below).
[0213] TABLE 16 Strain Carvacrol Code Loading % nxampe : oun : ur er p mza on o uan enes n ecom nant P. pastoris Host Cells by Combinatorial Mutagenesis
[0214] This example illustrates the preparation of combinatorial libraries of CYP genes based on parent strain OGP100 which carries a gene encoding the CYP302 (SEQ ID NO: 156) which has the mutations S27R, H51R, L99A, I102L, S111S, N119S, A176M, T196A, P358P, D415R, and P437P relative to the CYP301 polypeptide of SEQ ID NO: 130) as identified through combinatorial library screening in Example 5 for improved bioconversion of carvacrol and / or thymol to THQ.
[0215] Material and Methods
[0216] A. Combinatorial library construction
[0217] The polynucleotide sequence of SEQ ID NO: 155 encoding the engineered CYP302 polypeptide (SEQ ID NO: 156) from OGP100 expressed under the pAOX1 promoter (SEQ ID NO: 217) and tDAS2 terminator (SEQ ID NO: 85), was used to design oligos to generate combinatorial libraries to randomly incorporate additional identified beneficial amino acid ‐ 90 ‐ mutations from Examples 4 and 5 into the CYP302 backbone sequence. The same screening strain was used to integrate the resulting libraries as in Example 5 (OGP084), however, the OGP100 strain that expresses CYP302 was used as the parent control strain in order to determine fold-improvement in conversion of carvacrol to THQ as described below.
[0218] A semi-synthetic approach was used to construct the first set of combinatorial libraries. Genomic DNA from the OGP100 strain was used as the template to generate a full- length PCR product using primer pair of wboligos7751 (SEQ ID NO: 216) and wboligos7752 (SEQ ID NO: 217) while incorporating uracil using a dNTP mix comprising of the following deoxyribonucleotides: dATP, dGTP, dCTP, dTTP, dUTP. The resulting PCR product was column purified and digested with Uracil-DNA Glycosylase and Endonuclease IV at 37 C for 2 hours, followed by enzyme denaturation at 94 C for two minutes, to generate a pool of fragments in the range of 50-100 bases. These fragments were further combined with differing ratios of pools of the synthesized oligonucleotide primers (each oligo up to 60 bases in length and encoding one or more amino acid change) in several individual assembly PCR reactions using forward primer wboligos6242 (SEQ ID NO: 218) and reverse primer wboligos6102 (SEQ ID NO: 219) to reassemble the full-length PCR product (Fragment B) and incorporate additional amino acid changes within each pool randomly. The synthesized oligonucleotide primers SEQ ID NOs: 261-324 used in the pools and their encoded mutations are listed in Table 17 below and the accompanying Sequence Listing.
[0219] TABLE 17 SEQ ID : 1 2 3 4 5 6 7 8 9 0‐ 91 ‐ _lib1 R257G 60GATGAAAAATTAAGATCCGGGGGTGAGAAGAGA271 GATTTAGTGGAAGTGTTGCTTGAAATT lib1 V264M 60GATGAAAAATTAAGATCCAGGGGTGAGAAGAGA272 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7‐ 92 ‐ _lib2 R159Y 43TACAGGATTATTTAGCTGAATATTGTGAAAGAG298 GTACAGCTGT lib2 N178C 60CAACAATGTTGTGCTTGATGAGCGCTTCTCTAT299 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1‐ 93 ‐ 36_lib2 D418R 58TGCCTGAAAGATTTCTGAACGAAAATCGGAACA322 TTCGCTTTAAAGGTAGAGACTTTGA 37 lib2 E413A58TGCCTGAAAGATTTCTGAACGCAAATCGGAACA323 4with homologous 5’ and 3’ regions to the CYP were generated (Fragment A and Fragment C respectively). Fragment A was amplified using the forward primer wboligos7178 (SEQ ID NO: 210) and reverse primer wboligos5732 (SEQ ID NO: 211) and Fragment C was amplified using the forward primer wboligos6862 (SEQ ID NO: 212) and the reverse primer wboligos7183 (SEQ ID NO: 213). Fragments A, B, and C were assembled by NEBuilder® HiFi Assembly and amplified as a single linear DONOR using the following forward rescue primer wboligos7179 (SEQ ID NO: 214) and reverse primer wboligos7182 (SEQ ID NO: 215). The assembled PCR products were then pooled together, and gel purified to provide a combinatorial library of linear donor DNA.
[0221] The pooled linear donor DNA of combinatorial libraries were integrated using CRISPR-Cas9 into the Int15 locus in place of mVenus in OGP084.
[0222] B. HTP Screening of combinatorial library
[0223] Screening of the recombinant host cells for bioconversion of carvacrol to THQ was carried out using the same protocol described in Example 5 except the MeOH addition was decreased from a final of 2% to 0.125% and the BMGY concentration was decreased from 1x to 0.5x.
[0224] Screening of Round 2 combinatorial library hits at SF scale
[0225] Screening of the Round 2 combinatorial library hits for bioconversion of carvacrol to THQ at shake flask scale was carried out according to the following assay: individual colonies of recombinant P. pastoris strains were picked into 500 mL baffled conical flasks containing BMGY media (100 mL working volume). The flasks were incubated at 30oC for 48 h with shaking at 250 rpm at 85% humidity and were supplemented with glycerol (2 % final concentration) after 24 hours. After this time, the biomass was harvested via centrifugation and the supernatant was discarded. A portion of the resulting biomass (0.5 g) was resuspended in 2.25 mL of bioconversion buffer (0.1 M phosphate buffer, 2% MeOH, pH 7) in a 20 mL scintillation vial and carvacrol (3 g / L) was added as a 10X IPM stock. The vial was incubated at 30oC for 24 h with shaking at 250 rpm at 85 % humidity. After this time, IPA (5 mL) was added and the vials were incubated at 30oC with shaking at 250 rpm at 85 % humidity for a further 30 minutes. The samples were then aliquoted into a 96-well plate before the biomass was separated by centrifugation and the resulting supernatant ‐ 94 ‐ diluted with MeOH / water (3:1) (30 x total dilution) using the Agilent Bravo automated liquid handling platform before analysis using the UHPLC-DAD method described above.
[0226] Results
[0227] Screening of the combinatorial libraries for fold-improvement in production of THQ from carvacrol relative to the control strain, OGP100, which expresses the CYP302 polypeptide of SEQ ID NO: 156 and the CPR polypeptide of SEQ ID NO:6, are summarized in Table 18 (below). As shown by the results in Table 18, strains expressing the engineered CYP variant polypeptides derived from CYP302 of even-numbered SEQ ID NOs: 168-196, were found to exhibit 1.2-fold to 1.5-fold improved activity in converting carvacrol to THQ relative to the parent control strain OGP100 that expresses the CYP302 polypeptide.
[0228] TABLE 18 Carvacrol FIOPC NT AA and / or Carvacrol o‐ 95 ‐ 20240710_O 177 178 S344P, S27R, S111S (AGT), 60% 1.53 G_libOGP08 H51R, L99A, P358P (CCG), 9 AP3 50x: I102L N119S P437P (CCT)‐ 96 ‐ K151E, S344P, P358P (CCG), S396T, I477V, P437P (CCT) S27R H51R simprovement in production of THQ from carvacrol relative to the control strain, OGP100, are summarized Table 19 (below). As shown in Table 19, the strains OGP109, OGP110, OGP111, OGP112, OGP113, and OGP114, each of which have multiple copies of CYP302 at the Int15 locus, each exhibited improved activity of between 23% and 28% higher conversion of carvacrol to THQ at 3000 mg / L carvacrol loading relative to the control strain OGP100.
[0230] TABLE 19 Strain Carvacrol P. pastoris Strain DescriptionLoading % CodenExample 7: Round 4: Further Optimization of Mutant CYP Genes in Recombinant P. pastoris Host Cells by Site Saturation Mutagenesis (SSM) ‐ 97 ‐
[0231] This example illustrates the preparation of site saturation mutagenesis (SSM) libraries of CYP genes based on parent strain OGP114 which carries a gene encoding the CYP312 (SEQ ID NO: 196) which has the amino acid mutations G35R, S299A, and N472D, relative to the CYP302 polypeptide of SEQ ID NO: 156) as identified through combinatorial library screening in Example 6 for improved bioconversion of carvacrol and / or thymol to THQ. The resulting libraries were screened for improved activity in the bioconversion of carvacrol to THQ relative to the bioconversion of the parent strain OGP114 containing the parent CYP polypeptide of SEQ ID NO: 196 (CYP312) and CPR006 (SEQ ID NO: 6).
[0232] Material and Methods
[0233] A. Site Saturation Mutagenesis library construction
[0234] The heterologous nucleic acid sequence of SEQ ID NO: 195 which encodes the CYP polypeptide of SEQ ID NO: 196, expressed under the pAOX1 promoter (SEQ ID NO: 217) and tDAS2 terminator (SEQ ID NO: 85), was used to generate SSM libraries at 192 amino acid positions of the polypeptide sequence. To generate an appropriate screening strain to evaluate these libraries, the mVenus gene (SEQ ID NO: 218) was integrated into the Intergenic Region 15 site of a Pichia strain under the pAOX1 promoter and the tDAS2 terminator, along with the codon optimized polynucleotide sequence of CPR006 (SEQ ID NO:6), under the pAOX1 promoter and tPMP20 terminator (SEQ ID NO: 219) integrated into Intergenic Region 21. Additionally, the ku70 gene was knocked out with the goal to increase DNA integration efficiency. The resulting strain, OGP084, was used as the screening host for Integration and screening of the resulting SSM libraries.
[0235] Genomic DNA from the OGP114 strain was used as the template to generate two PCR products: (1) a first PCR product (Fragment A), which does not harbor any degenerate codons, and (2) a second PCR product (Fragment B), which has sequence overlap with the Fragment A, and is amplified harboring one NNK degenerate codon only. Primers spanning 192 codons of the CYP polypeptide were designed according to standard site-saturation mutagenesis protocols and used for amplification of Fragments A and B and overlap extension.
[0236] Fragment A was amplified using a single forward primer wboligos7179 (SEQ ID NO: 325, AACAGAATTTTAGTATGTATAAAAGTTGCAAG) and a series of 192 reverse primers designed according to the location of the desired mutagenesis site in CYP. The 192 reverse primers for Fragment A listed in Table 20A (below) and are provided in the accompanying Sequence Listing as SEQ ID NOs: 326-517.
[0237] TABLE 20A SEQ ID i i‐ 98 ‐ Codon45_P2A2CTGTAGGATATTACCTATTATGGGTAATCT327Codon105_P2A3TGCAGTCAATGCAGACGCTACAGTCCTATT328d 2 2 P2A4 2‐ 99 ‐ Codon162_P2D12TTCACATCTTTCAGCTAAATAATCCTGTAG373Codon4_P2E1AAAATCCATTTTCAATAATTAGTTGTTTTT374d 4 7 P2E2 7‐ 100 ‐ Codon57_P2H10GAGTTTTGCTAAAGAACGATGGGGTTTAGG419Codon74_P2H11TGTCGTAATGGACCCCAACTTAAGAGACAT420d 1 P2H12 421‐ 101 ‐ Codon417_P1D8GTTCCGATTTTCGTTCAGAAATCTTTCAGG465Codon181_P1D9CATCAAATTCAACATTGTTGTAAAGGCGGC466d P1D1 4 7‐ 102 ‐ Codon112_P1H6ACTGAATTCGTGGTGGCGTAGTGCAGTCAA511Codon14_P1H7AATTGAAGCTAAAATCACTGAGCTAGTTAA512d P1H 1518, GAATGAAATCTTGATGGCCTGGGG) and a series of 192 forward primers that included a single NNK degenerate codon spanning across successive positions of the CYP312 encoding gene of SEQ ID NO: 196. The 192 forward primers for Fragment B are listed in Table 20B (below) and are provided in the accompanying Sequence Listing as SEQ ID NOs: 519-710.
[0239] TABLE 20B: Forward primers for SSM library Fragment B amplification SEQ ID : 9012345678901234567890123456789‐ 103 ‐ Codon420_P1C8TTTCTGAACGAAAATCGGAACATTGACTTTNNKGGTAGAGACTTT550Codon98_P1C9AAACATGACAGCTCATTCAGTAATAGGACTNNKGCGTCTGCATTG551Codon321 P1C10CGTAATCCACAGAAGATGGTCAAGTTGAGANNKGAGATCAGATCA55234567890123456789012345678901234567890123456789012‐ 104 ‐ Codon399_P1H1GTCTGGGCATCGGGGAGAGACTCTAAAATCNNKAAGAACCCAGAT603Codon460_P1H2TTGGTTACTTTCGTCGGAAATTACGATTGGNNKCTGGAAAACGGG604Codon492 P1H3GCGATTCCTTTGAAAGCTATTCCAATTAAANNKTAAGTAGATTTG60567890123456789012345678901234567890123456789012345‐ 105 ‐ Codon477_P1D6GAGGAGATGGATATGGACGAGAATTTCGGTNNKACACTGGAAAAA656Codon124_P1D7TGGCTTCCGGTCGGCAGTCAATGGAGGAAANNKAGAAAGATCTGT657Codon417 P1D8CCTGAAAGATTTCTGAACGAAAATCGGAACNNKGACTTTAAAGGT65890123456789012345678901234567890123456789012345678‐ 106 ‐ Codon207_P1H11AGCCAGGAGATGAAAGAGGCCATCAGTTCANNKGTAAAGTCAATA709Codon217_P1H12GTCGTAAAGTCAATAGGTCGTCCTAACCTCNNKGATTATTTTCCG710forward primer wboligos6153 (SEQ ID NO: 711, GATAATCGACCTAAATTCTCGTTGAAAGGC) and reverse primer wboligos7182 (SEQ ID NO: 712, CTTGTAGTTGAAAGATGTCTGACATATATC). The assembled OE-PCR products were then pooled together, and gel purified to provide a saturation mutagenesis library of linear donor DNA.
[0241] The pooled linear donor DNA of the site saturation mutagenesis library was integrated using CRISPR-Cas9 into the Intergenic Region 15 loci replacing the m-Venus gene in OGP084.
[0242] B. HTP Screening of SSM library
[0243] An HTP screening assay was developed to detect and quantify the bioconversion of carvacrol to THQ from the recombinant Pichia pastoris strains. Individual colonies of recombinant P. pastoris strains were picked into 96-well plates containing 2x YPD media (300 μL per well) using a QPixTM 420 colony picking system. The plates were incubated at 30oC for 48 h with shaking at 250 rpm, 85 % humidity with a 50-mm throw shaker. After this time, the strains were sub-cultured (10.5 μL inoculation volume per strain) into a second 96- well plate containing 1 x BMGY media (105 µL per well) using an Agilent Bravo automated liquid handling platform. The plates were incubated at 30oC for 24 h with shaking at 250 rpm, 85 % humidity with a 50-mm throw shaker. After this time, 45 μl of a MCT oil overlay containing 13.33 g / L carvacrol oil mix, 6.67% MeOH was added (final concentration at 4 g / L carvacrol oil mix and 2% MeOH) to the culture and incubated at 30oC for 24 h with shaking at 250 rpm, 85 % humidity with a 50-mm throw shaker. After this time, IPA (315 μL) was added using the Bravo and the plates were incubated at 30oC with shaking at 250 rpm, 85 % humidity with a 50-mm throw shaker for a further 30 minutes. The biomass was then separated by centrifugation at 4,000xg for 20 minutes and the resulting supernatant diluted with MeOH (100 x total dilution) using the Agilent Bravo automated liquid handling platform before analysis. The samples were then loaded onto an Agilent 1290 Infinity II UHPLC equipped with a Diode Array Detector and the compound of interest were detected by monitoring at 292 nm (THQ) and 274 nm (Carvacrol). Analytes were quantified relative to calibration curves which were prepared using 2-fold serial dilutions of stock solutions of carvacrol and THQ. UHPLC Instrumentation and parameters used were as described in Example 1.
[0244] C. Evaluation of SSM hits at Bioreactor scale ‐ 107 ‐
[0245] The recombinant Pichia pastoris strain was grown at a 1 L scale in a glass fermenter according to the following protocol. There are three stages in the fermentation process: a Batch phase, where the initial 4% glycerol is consumed at approximately 20 hours, a Fed- Batch phase where glycerol was fed to increase the biomass and a Biotransformation phase where MeOH is fed for induction of pathway genes for the biotransformation of oregano oil to THQ.
[0246] A high-performing control strain, OGP122, was constructed to evaluate expression at the bioreactor scale. The OGP122 strain includes the heterologous nucleic acid sequence encoding CYP312 expressed under the pAOX1 promoter and the tDAS2 terminator integrated at the Intergenic Region 15 of Pichia pastoris. Additionally, OGP122 includes the codon-optimized sequence for CPR006, also expressed under the pAOX1 promoter and terminated by tPMP20, integrated into the Intergenic Region 21.
[0247] The propagation consisted of one consecutive shake flask culture to expand a glycerol stock into an inoculum culture to a targeted OD600 > 25 to inoculate the growth tank. The inoculum was added to a prepared sterile 1 L fermenter containing about 440 mL of a phosphate citrate minimal media containing 4% glycerine as a carbon source (“PCMG”) and a modified Pichia Trace Metals mix (“PTM1”).
[0248] PCMG.1 liter preparation contains: citric acid (19.2 g); glycerine (40 g); 75% phosphoric acid (12.9 M); magnesium sulfate heptahydrate (2.4 g); mPTM1 (4.35 mL); 30% ammonium hydroxide (10.2 mL at 14.8 M) (1NA0109 Spectrum; VWR); 45% potassium hydroxide (26.5 mL at pH 5.6 - 5.75) (95031-518 Spectrum; VWR); and MCT oil (2 mL).
[0249] PTM1.1 liter preparation contains: trace metal master mix (94.289 g); copper sulfate pentahydrate (6 g) (12849-100G Fluka; VWR); potassium iodide (0.089 g); manganese(II) sulfate monohydrate (3 g) (MA164 Spectrum; VWR); zinc chloride (20 g); ferrous sulfate heptahydrate (65 g) (F1062 Spectrum; FisherScientific); biotin (Vitamin B7) (0.2 g) (BI115 Spectrum; VWR); citric acid or sulfuric acid (98%) 2 g 5 mL (BDH3068-500MLP; VWR).
[0250] The Batched phase growth fermenter started at a volume of 500 mL and an initial OD600 of 3, which was maintained at 30°C, 30% dissolved oxygen, and pH 5 using 28% ammonium hydroxide. This reaction was maintained at a DO of 30% using the following cascade parameters: agitation [0 %- 50%] 600 – 1200 and gassing [50% -100%] 15 sL / h – 66 sL / h. For the Fed-Batch phase, 50% (v / v) glycerol was fed at a 10 g / L / h to increase the biomass to a wet cell weight (WCW) of 29%. After the desired WCW was reached, culture was induced by adding 0.5% MeOH. When the MeOH bolus was added, the Biotransformation time started. An acclimation period was performed, the culture signaled ready for MeOH DO-stat feeding by an abrupt spike in DO%. Pure MeOH was fed using a DO% stat profile, adding MeOH when dissolved oxygen percentage rises above 30%. The vessel was charged with 24% MCT oil containing 5 g / L stock of Oregano oil. After 42 h, the ‐ 108 ‐ reaction was charged with another 6% of MCT oil containing 5 g / L stock of Oregano oil. The Biotransformation reaction time was 72 h, and the total Runtime was 110 h. Aliquots (0.3 mL) were taken at regular time intervals for analysis using the UHPLC method described above.
[0251] Results
[0252] The results of screening of the site saturation libraries for improved activity in production of THQ from carvacrol relative to the control strain, OGP114, which expresses the CYP312 polypeptide of SEQ ID NO: 196 and the CPR polypeptide of SEQ ID NO:6, are summarized in Table 21 (below). Strains expressing the CYP variants CYP379, CYP380, CYP381, CYP383, and CYP384 were found to exhibit at least 1.9-fold improved production (FIOPC) of THQ relative to parent strain OGP114.
[0253] TABLE 21 NT AA THQ SEQ ID SEQ ID Neutral Codon FIOPC L) wn
[0254] Data from bioreactor evaluation of SSM library hits in terms of % conversion of carvacrol to THQ (at 10000 mg / L loading) relative to the high-performing control strain, OGP122 are summarized Table 22 (below). As shown in Table 22, the OGP124, which has ‐ 109 ‐ copies of CYP379, CYP380, and CYP381 at Int15 locus, exhibited improved activity of 25% higher conversion to THQ at 10000 mg / L carvacrol loading relative to the control.
[0255] TABLE 22 Strain Carvacrol Code Loading % P. pastoris Strain Description (mg / L) Conversione detail by way of example and illustration for purposes of clarity and understanding, this disclosure including the examples, descriptions, and embodiments described herein are for illustrative purposes, are intended to be exemplary, and should not be construed as limiting the present disclosure. It will be clear to one skilled in the art that various modifications or changes to the examples, descriptions, and embodiments described herein can be made and are to be included within the spirit and purview of this disclosure and the appended claims. Further, one of skill in the art will recognize a number of equivalent methods and procedure to those described herein. All such equivalents are to be understood to be within the scope of the present disclosure and are covered by the appended claims.
[0257] Additional embodiments of the invention are set forth in the following claims.
[0258] The disclosures of all publications, patent applications, patents, or other documents mentioned herein are expressly incorporated by reference in their entirety for all purposes to the same extent as if each such individual publication, patent, patent application or other document were individually specifically indicated to be incorporated by reference herein in its entirety for all purposes and were set forth in its entirety herein. In case of conflict, the present specification, including specified terms, will control. ‐ 110 ‐
Claims
CLAIMS What is claimed is:
1. A recombinant host cell comprising a heterologous nucleic acid encoding a polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, L105, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, G258, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, K420, Y457, D458, N472, I477, and T478; optionally, wherein the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, L105V, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, G258V, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, K420C, Y457C, D458V, N472D, I477M, I477V, and T478C.
2. The cell of claim 1, wherein the amino acid sequence comprises a combination of amino acid differences as compared to SEQ ID NO: 24 selected from: T196A, M464G S27R H51R L99A I102L N119S A176M T196A D415R M464G‐ 111 ‐ L99A, R159S, M464G S12A, L99A, I378T, M464G L, 3.T e ce o any one o cams 1-2, w eren t e eteroogous nucec acd sequence as at least 80% identity to SEQ ID NO: 23, and at least one neutral codon difference as compared to SEQ ID NO: 23 at a position encoding an amino acid residue selected from: S21, V26, L53, S111, V112, G143, A171, T174, S210, I211, P214, L246, D289, D296, S299, Q333, D339, P358, P365, A368, D371, P437, and D458; optionally, wherein the neutral codon difference is selected from: S21S (AGC), V26V (GTC), L53L (CTA), S111S (AGT), V112V (GTG), G143G (GGC), A171A (GCA), T174T (ACG), S210S (TCG), I211I (ATC), P214P (CCC), L246L (CTG), D289D (GAC), D296D (GAC), S299S (TCG), Q333Q (CAG), D339D (GAC), P358P (CCG), P365P (CCT), A368A (GCA), D371D (GAC), P437P (CCT), and D458D (GAC). ‐ 112 ‐ 4. The cell of any one of claims 1-3, wherein the heterologous nucleic acid sequence comprises a combination of neutral codon difference as compared to SEQ ID NO: 23 selected from: S111S (AGT), P358P (CCG), P437P (CCT) S21S (AGC), S111S (AGT), P358P (CCG), P437P (CCT) 5.The cell of any one of claims 1-4, wherein the polypeptide comprises (a) an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206; or (b) to an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206.
6. The cell of any one of claims 1-5, wherein the heterologous nucleic acid comprises: (a) a sequence of at least 80% identity, at least 81%, at least 82%, at least 83%, at ‐ 113 ‐ least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205; or (b) a codon degenerate sequence of a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205.
7. The cell of any one of claims 1-6, wherein the heterologous nucleic acid encodes a polypeptide with CPR activity; optionally, wherein the polypeptide with CPR activity comprises an amino acid sequence of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to SEQ ID NO:
6.
8. The cell of any one of claims 1-7, wherein the heterologous nucleic acid is under the control of a promoter system selected from pGal1 / 10, and pCAT1:pFDH1.
9. The cell of any one of claims 1-8, wherein the source organism of the host cell is selected from Pichia pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, and Escherichia coli.
10. The cell of any one of claims 1-9, wherein the heterologous nucleic acid is integrated into a site in the host cell genome.
11. The cell of any one of claims 1-10, wherein the recombinant host cell is Pichia pastoris and the heterologous nucleic acid is integrated at one or more sites in the genome selected from AOX1, Int6, Int15, and HIS4; optionally, wherein the heterologous nucleic acid is integrated at three or more sites.
12. The cell of any one of claims 1-10, wherein the recombinant host cell is Saccharomyces cerevisiae and the heterologous nucleic acid is integrated at one or more sites in the ‐ 114 ‐ genome selected from X-2, X-4, XI-2, XII-4, NDE1, XII-5, Gal80, and ROQ1; optionally, wherein the heterologous nucleic acid is integrated at three or more sites.
13. The cell of any one of claims 1-12, wherein the cell exhibits increased bioconversion of carvacrol to THQ relative to a control host cell comprising a heterologous nucleic acid encoding a polypeptide of SEQ ID NO: 24; optionally, wherein the improved bioconversion exhibited by the host cell comprises an increased yield of THQ relative to the control host cell of at least 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8- fold, 1.9-fold, 2-fold, 4-fold, 5-fold, 10-fold, or more.
14. A recombinant polypeptide with CYP activity capable of converting carvacrol and / or thymol to THQ, wherein the polypeptide comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, L105, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, G258, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, K420, Y457, D458, N472, I477, and T478; optionally, wherein the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, L105V, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, G258V, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, K420C, Y457C, D458V, N472D, I477M, I477V, and T478C.
15. The polypeptide of claim 14, wherein the amino acid sequence comprises a combination of amino acid differences as compared to SEQ ID NO: 24 selected from: T196A, M464G 27R H 1R L A I1 2L N11 A17 M T1 A D41 R M4 4‐ 115 ‐ S27R, R284K, D415G, M464G G35K, E202K L,16. The polypeptide of any one of claims 14-15, wherein the polypeptide comprises (a) an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at ‐ 116 ‐ least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206; or (b) to an amino acid sequence selected from SEQ ID NO: 130, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, and 206.
17. The polypeptide of any one of claims 14-16, wherein the polypeptide is fused via a linker to a second polypeptide; optionally, wherein the second polypeptide has CPR activity.
18. The polypeptide of claim 17, wherein the second polypeptide with CPR activity comprises an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from SEQ ID NO: 6; optionally, wherein the polypeptide comprises an amino acid sequence of SEQ ID NO:
6.
19. The polypeptide of any one of claims 14-18, wherein the polypeptide has increased activity in the conversion of carvacrol to THQ relative to a control polypeptide of SEQ ID NO: 24; optionally, wherein the increased activity in the conversion of carvacrol to THQ relative to the polypeptide of SEQ ID NO: 24 of at least 1.2-fold, 1.3-fold, 1.4-fold, 1.5- fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 4-fold, 5-fold, 10-fold, or more.
20. A polynucleotide encoding the polypeptide of any one of claims 14-19.
21. A polynucleotide comprising a sequence of at least 80% identity, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 129, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, ‐ 117 ‐ 125, 127, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, and 205.
22. An expression vector comprising the polynucleotide of claim 20-21.
23. The expression vector of claim 22 further comprising a polynucleotide sequence of at least 80% identity, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to SEQ ID NO:
5.
24. The expression vector of any one of claims 22-23 further comprising a control sequence; optionally, wherein the control sequence comprises a promoter system selected from pGal1 / 10, and pCAT1:pFDH1.
25. A host cell comprising the polynucleotide claim 20 or the expression vector of any one of claims 22-24.
26. The host cell of claim 25, wherein the source organism of the host cell is selected from Pichia pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, and Escherichia coli.
27. A composition comprising: (a) a recombinant host cell of any one of claims 1-13 or a recombinant polypeptide of any one of claims 14-19; and (b) a compound (2a) and / or compound (2b) or a derivative of compound (2a) and / or a derivative of compound (2b)28. The composition of 27, a with CPR activity; optionally, a polypeptide comprising an amino acid sequence of at least 80% sequence identity to SEQ ID NO:
6. ‐ 118 ‐ 29. The composition of any one of claims 27-28, wherein the composition comprises a derivative of compound (2a) and / or a derivative of compound (2b) selected from compound (2c), compound (2d), compound (2e), compound (2f), compound (2g), compound (2h), compound (2i), compound (2j), compound (2k), compound (2l), compound (2m), compound (2n), compound (2o), compound (2p), and mixture thereof.
30. The composition of any one of claims 27-29, wherein the composition is in aqueous solution.
31. A method for preparing compound (1a) or a derivative of compound (1a) comprising contacting under compound (2a) and / orcompound (2b) or a derivative of a derivative of compound (2b)with a recombinant host cell of any one of claims 1-13, or a recombinant polypeptide of any one of claims 14-19.
32. A method for preparing compound (1a) or a derivative of compound (1a)‐ 119 ‐ comprising contacting under suitable reactions conditions compound (2a) and / or compound (2b) or a derivative of compound (2a) and / or a derivative of compound (2b) with a recombinantpolypeptide with CYP activity comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO: 24 and an amino acid difference as compared to SEQ ID NO: 24 at one or more positions selected from M464, S27, K28, G35, H51; L53, K59, Y60, L68; A74, T81, L84, N95, L99, I102, L105, N119, K126; M132, K151, R159, A176, T196, M200, K201, E202, R213, V226, L237, F243, G258, E259, I270, R284, S299, T307, R311, K318, N321, Q333, Q334, S344, Y345, L356, N375, S396, L406, P407, N412, E413, D415, N416, K420, Y457, D458, N472, I477, and T478; optionally, wherein the amino acid differences are selected from M464G, S27R, K28N, G35K, G35R, H51R; L53D, K59E, Y60F, L68F; A74D, T81A, L84P, N95S, L99A, I102L, L105V, N119S, K126R, M132V, K151E, R159S, R159W, A176M, T196A, M200T, K201C, E202K, R213S, V226A, L237H, L237R, F243L, G258V, E259G, I270L, R284K, S299A, T307S, R311S, K318V, N321M, Q333S, Q334L, S344P, Y345N, L356F, N375D, S396T, L406S, P407L, N412E, E413A, D415G, D415R, R415V, N416D, K420C, Y457C, D458V, N472D, I477M, I477V, and T478C.
33. The method of claim 32, wherein the method further comprises contacting with a recombinant polypeptide with CPR activity; optionally, wherein the recombinant polypeptide with CPR activity comprises an amino acid sequence of at least 80% sequence identity to SEQ ID NO:
6.
34. The method of any one of claims 32-33, wherein the contacting with the polypeptide with CYP activity occurs in the presence of host cell that expresses the polypeptide.
35. The method of any one of claims 32-34, wherein the contacting with the polypeptide with CYP activity occurs in a cell-free system. ‐ 120 ‐ 36. The method any one of claims 32-35, wherein a derivative of compound (1a) is prepared, wherein the derivative is selected from compound (1b), compound (1c), compound (1d), compound (1e), compound (1f), compound (1g), compound (1h), compound (1i), compound (1j), compound (1k), compound (1l), compound (1m), compound (1n), compound (1o), compound (1p), compound (1q), compound (1r), compound (1s), compound (1t), compound (1u), compound (1v), compound (1w), compound (1x), compound (1y), compound (1z), compound (1aa), compound (1ab), compound (1ac), compound (1ad), compound (1ae), compound (1af), compound (1ag), compound (1ah), and compound (1ai).
37. The method any one of claims 32-35, wherein a derivative of compound (2a) and / or a derivative of compound (2b) is used to prepare a derivative of compound (1a), wherein the derivative is selected from compound (2c), compound (2d), compound (2e), compound (2f), compound (2g), compound (2h), compound (2i), compound (2j), compound (2k), compound (2l), compound (2m), compound (2n), compound (2o), compound (2p), and mixture thereof.
38. The method of any one of claims 32-37, further comprising a chemical step to form a derivative of compound (1a).
39. The method of claim 38, wherein the derivative of compound (1a) is selected from compound (1b), compound (1c), compound (1d), compound (1e), compound (1f), compound (1g), compound (1h), compound (1i), compound (1j), compound (1k), compound (1l), compound (1m), compound (1n), compound (1o), compound (1p), compound (1q), compound (1r), compound (1s), compound (1t), compound (1u), compound (1v), compound (1w), compound (1x), compound (1y), compound (1z), compound (1aa), compound (1ab), compound (1ac), compound (1ad), compound (1ae), compound (1af), compound (1ag), compound (1ah), and compound (1ai).
40. A method for producing THQ comprising: (a) culturing a recombinant host cell of any one of claims 1-13 in a suitable medium comprising carvacrol and / or thymol; and (b) recovering the produced THQ. ‐ 121 ‐