Cannabinoid acid synthases

WO2026177620A1PCT designated stage Publication Date: 2026-08-27WAGENINGEN UNIVERSITEIT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/NL2026/050052
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2026-02-23
Publication Date
2026-08-27

Smart Images

  • Figure IMGF000026_0001_TABLE
    Figure IMGF000026_0001_TABLE
  • Figure IMGF000026_0002_TABLE
    Figure IMGF000026_0002_TABLE
  • Figure IMGF000027_0001_TABLE
    Figure IMGF000027_0001_TABLE
Patent Text Reader

Abstract

The invention provides an isolated or engineered polynucleotide (10) comprising a coding sequence (12), wherein the coding sequence (12) encodes a protein (20), wherein an amino acid sequence of the protein (20) has at least 85% sequence identity with a reference amino acid sequence in a sequence alignment between the amino acid sequence and the reference amino acid sequence, wherein the reference amino acid sequence is selected from the group consisting of SEQ ID NO:8 11 14 1720, and wherein the protein (20) is a cannabinoid acid synthase (25).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cannabinoid acid synthases

[0002] FIELD OF THE INVENTION

[0003] The invention relates to the field of synthetic biology. In general, the invention pertains to cannabinoid biosynthesis. In particular, the invention is directed to an isolated or engineered polynucleotide encoding a cannabinoid acid synthase. The invention is also directed to a protein. The invention is further directed to a method for producing a cannabinoid acid, to an engineered organism or part thereof, and to the use of the polynucleotide, the protein, and / or the organism in the manufacture of a cannabinoid acid.

[0004] BACKGROUND OF THE INVENTION

[0005] Approaches for producing cannabinoid acid synthases and cannabinoids are known in the art. For instance, LUO et al., (Nature, 2019, Vol. 567, pp. 123-126) describes the biosynthesis of cannabinoids in Saccharomyces cerevisiae via the introduction of a biosynthetic pathway, including the introduction of genes encoding cannabinoid acid synthases.

[0006] SUMMARY OF THE INVENTION

[0007] Cannabinoids are a class of chemical compounds, some of which have medicinal value. These compounds are primarily produced, in acid form, in the cannabis plant (Cannabis sativa L.) via metabolization of cannabigerolic acid (CBGA), catalyzed by enzymes known as cannabinoid oxidocyclases or cannabinoid acid synthases. Three genes from the cannabis plant have been identified that encode enzymes responsible for catalyzing the biosynthesis of three principal cannabinoid acids, namely tetrahydrocannabinolic acid (THCA), cannabidiolic acid (CBDA), and cannabichromenic acid (CBCA). Of the decarboxylated forms of these compounds, tetrahydrocannabinol (THC) is known for its therapeutic applications, including the control of nausea and vomiting, appetite stimulation, and pain relief. Cannabidiol (CBD) has gained attention for its potential in reducing epileptic seizures and is increasingly used in health and food supplement markets. Cannabichromene (CBC), though typically produced in lower quantities by the cannabis plant, has been recognized for its potential analgesic properties and non-psychoactive nature, making it a compound of emerging interest for pharmaceutical and therapeutic applications.

[0008] Given the well-documented medicinal value of certain cannabinoids, the demand for these compounds has been increasing and is expected to continue rising. However,the natural production of cannabinoids in wild-type cannabis plants may be insufficient to meet this growing demand.

[0009] The prior art describes the heterologous expression of cannabinoid acid synthase genes in microorganisms (e.g., yeast) for cannabinoid production. However, attempts to express CBDA synthase (CBDAS) and CBCA synthase (CBCAS) in heterologous systems have generally resulted in low expression levels.

[0010] Moreover, CBDAS exhibits relatively low enzymatic activity in vitro, leading to low CBDA production, which further complicates CBDA biosynthesis.

[0011] Additionally, wild-type cannabinoid acid synthases may exhibit poor product specificity. For instance, CBDAS can also catalyze the formation of CBCA and THCA, while THCA synthase (THCAS) may also catalyze the formation of CBCA.

[0012] The prior art describes rational engineering approaches to modify cannabinoid acid synthase activity and / or product specificity. Additionally, the prior art describes engineering approaches to reduce the number of disulfide bonds in cannabinoid acid synthases to facilitate heterologous expression in microbial hosts. However, engineered cannabinoid acid synthases typically still suffer from one or more limitations, including low enzymatic activity, poor product specificity, and low heterologous expression levels, particularly in the case of CBCAS and CBDAS.

[0013] Accordingly, it is an aim of the invention to provide an alternative cannabinoid acid synthase and a polynucleotide encoding the same, which cannabinoid acid synthase preferably at least partially mitigates one or more of the above-described drawbacks. The present invention further aims to overcome or ameliorate at least one of the disadvantages of the prior art, or to provide a useful alternative.

[0014] In a first aspect, the invention provides an isolated or engineered polynucleotide comprising a coding sequence. The coding sequence may encode a protein. The protein may have an amino acid sequence having at least 85% sequence identity with a reference amino acid sequence in a sequence alignment, especially in an amino acid sequence alignment. In particular, the sequence alignment may be an alignment between the amino acid sequence and the reference amino acid sequence. The reference amino acid sequence may especially be selected from the group consisting of SEQ ID NO:8 11 14 1720. In embodiments, the protein may be a cannabinoid acid synthase, especially one or more of a cannabichromenic acid synthase, a cannabidiolic acid synthase, and a tetrahydrocannabinolic acid synthase.

[0015] In specific embodiments, the isolated or engineered polynucleotide comprises a coding sequence, wherein the coding sequence encodes a protein, wherein an amino acidsequence of the protein has at least 80% sequence identity with a reference amino acid sequence in a sequence alignment between the amino acid sequence and the reference amino acid sequence, wherein the reference amino acid sequence is selected from the group consisting of SEQ ID NO:8 11 14 1720, and wherein the protein is a cannabinoid acid synthase.

[0016] In particular, the polynucleotide of the invention may encode a protein with a higher enzymatic activity, with an increased product specificity, and / or with an increased heterologous expression level relative to prior art polynucleotides and the therein encoded cannabinoid acid synthases. Especially, the encoded proteins may be particularly suitable for CBCA and / or CBDA synthesis. For example, a protein according to SEQ ID NO: 11 may (essentially) exclusively catalyze the production of CBCA. As a further example, a protein according to SEQ ID NO: 14 or according to SEQ ID NO:20 may have CBDA as predominant product and have a substantially higher enzymatic activity compared to wild-type CBDAS. Further, polynucleotides encoding proteins according to each of SEQ ID NO:8 11 14 17 20 have been found to be particularly suitable for heterologous expression.

[0017] The sequences described herein were obtained using a combination of ancestral gene reconstruction and rational enzyme engineering (see below).

[0018] In particular, SEQ ID NO:1 corresponds to a reconstructed (common) ancestral polynucleotide sequence of (i) THCAS, CBCAS, and CBDAS genes, and (ii) the closest homolog from hop (Humulus lupulus L.) a sister species of the cannabis lineage. Enzyme characterization revealed that the encoded protein (SEQ ID NO:2) was unable to metabolize CBGA, and, consequently, could not catalyze cannabinoid acid production. The protein of SEQ ID NO:2 may hereinafter be referred to as ‘HCa’.

[0019] SEQ ID NO:4 corresponds to a reconstructed ancestral polynucleotide sequence of THCAS, CBCAS, and CBDAS genes. Enzyme characterizations revealed that the encoded protein (SEQ ID NO:5) could metabolize CBGA into THCA, CBDA, and CBCA. The protein of SEQ ID NO:5 may hereinafter be referred to as ‘Ca’.

[0020] The amino acid sequence of HCa was engineered by introducing the active site of the Ca protein, resulting in SEQ ID NO: 16 for the polynucleotide and SEQ ID NO: 17 for the encoded protein. The protein according to SEQ ID NO: 17 was found to exhibit high enzymatic activity and was particularly well-suited for heterologous expression. The protein of SEQ ID NO: 17 may hereinafter be referred to as ‘HCa Ca’.

[0021] The amino acid sequence of HCa Ca was further engineered by removing four codons encoding part of a loop structure, resulting in SEQ ID NO: 10 for the polynucleotide and SEQ ID NO: 11 for the encoded protein. Surprisingly, enzyme characterization of theprotein according to SEQ ID NO: 11 revealed that the enzyme converts CBGA (essentially) exclusively into CBCA, thus exhibiting a significantly improved product specificity for CBCA. The protein of SEQ ID NO: 11 may hereinafter be referred to as ‘HCa Ca del’. In addition, the protein of SEQ ID NO: 11 may be suitable for expression in heterologous systems, whereas the expression of wild-type CBC synthases in such systems may present challenges.

[0022] The amino acid sequence of HCa Ca was further engineered by introducing the flavin adenine dinucleotide (FAD) binding pocket of the Ca protein, resulting in SEQ ID NO:7 for the polynucleotide and SEQ ID NO: 8 for the encoded protein. Enzyme characterization revealed that the protein according to SEQ ID NO: 8 can also metabolize CBGA into THCA, CBDA, and CBCA. Surprisingly, the protein according to SEQ ID NO:8 exhibited particularly high enzymatic activity and was well-suited for heterologous expression. The protein of SEQ ID NO:8 may hereinafter be referred to as ‘HCa_Ca-FAD’.

[0023] Separately, the amino acid sequence of Ca was engineered by introducing the active site from a CBDAS, resulting in SEQ ID NO: 19 for the polynucleotide and SEQ ID NO:20 for the encoded protein, respectively. Enzyme characterization of the protein according to SEQ ID NO:20 revealed high product specificity for CBDA and, surprisingly, substantially higher enzymatic activity than the wild-type CBDAS. The protein of SEQ ID NO:20 may hereinafter be referred to as ‘Ca CBDAS’.

[0024] The amino acid sequence of Ca CBDAS was further engineered by introducing the FAD binding site from a CBDAS, resulting in SEQ ID NO: 13 for the polynucleotide and SEQ ID NO: 14 for the encoded protein. Similar as for SEQ ID NO: 20, enzyme characterization of the protein according to SEQ ID NO: 14 revealed high product specificity for CBDA and, surprisingly, substantially higher enzymatic activity than the wild-type CBDAS. The protein of SEQ ID NO: 14 may hereinafter be referred to as ‘Ca_CBDAS-FAD’. In addition, the protein of SEQ ID NO: 14 may be suitable for expression in heterologous systems, whereas wild-type CBDAS may exhibit low expression levels in such systems.

[0025] Surprisingly, the generated ancestral sequences and the engineered variants derived therefrom were observed to be substantially easier to express and produce than wild-type (WT) THCAS and CBDAS.

[0026] As described above, the invention may provide an isolated or engineered polynucleotide, especially an isolated or engineered DNA molecule.

[0027] The term “isolated” with regards to a polynucleotide may herein refer to a polynucleotide that has been separated from its natural environment. In embodiments, the polynucleotide may especially be an isolated polynucleotide.The term “engineered” with regards to a polynucleotide may herein refer to a polynucleotide that relates to or contains genetically engineered DNA, such as a polynucleotide produced via genetic engineering. The polynucleotide may be (artificially) synthesized or engineered to deviate from a naturally occurring polynucleotide, particularly in its nucleotide sequence. In embodiments, the polynucleotide may be an engineered polynucleotide, such as a recombinant polynucleotide.

[0028] In further embodiments, the polynucleotide, especially the coding sequence, may have a non-naturally-occurring nucleotide sequence. In further embodiments, the polynucleotide may encode a non-naturally-occurring amino acid sequence.

[0029] The polynucleotide may comprise a coding sequence encoding a protein. In particular, the coding sequence may encode a protein with cannabinoid acid synthase activity, i.e., a cannabinoid acid synthase. In further embodiments, the cannabinoid acid synthase may be selected from the group consisting of CBCAS, CBDAS, and THCAS. In embodiments, the encoded protein may be a CBCAS. In further embodiments, the encoded protein may be a CBDAS (EC 1.21.3.8). In further embodiments, the encoded protein may be a THCAS (EC 1.21.3.7).

[0030] The term “cannabinoid acid synthase” may herein refer to an enzyme configured to catalyze the conversion of CBGA into a cannabinoid acid, such as CBCA, CBDA, and / or THCA. The resulting cannabinoid acid may subsequently undergo (spontaneous) non-enzymatic decarboxylation, thereby converting into the corresponding cannabinoid. In particular, CBCA may convert into CBC, CBDA may convert into CBD, and THCA may convert into THC, all via non-enzymatic decarboxylation.

[0031] Hence, in embodiments, the protein may comprise a CBCAS configured to catalyze the conversion of CBGA into CBCA, wherein CBCA may spontaneously convert into CBC via non-enzymatic decarboxylation. Particularly, the protein may exhibit a product specificity PCBCA for CBCA, typically wherein PCBCA > 0.5, such as > 0.6, especially > 0.8. In further embodiments, PCBCA > 0.9, especially > 0.95, such as > 0.99, including 1.

[0032] Similarly, in embodiments, the protein may comprise a CBDAS configured to catalyze the conversion of CBGA into CBDA, wherein CBDA may spontaneously convert into CBD via non-enzymatic decarboxylation. Particularly, the protein may exhibit a product specificity PCBDA for CBDA, typically wherein PCBDA > 0.5, such as > 0.6, especially > 0.75. In further embodiments, PCBDA > 0.8, especially > 0.85, such as > 0.9, including 1.

[0033] In further embodiments, the protein may comprise a THCAS configured to catalyze the conversion of CBGA into THCA, wherein THCA may spontaneously convert intoTHC via non-enzymatic decarboxylation. Particularly, the protein may have a product specificity PTHCA for THCA, typically wherein PTHCA > 0.5, such as > 0.6, especially > 0.75. In further embodiments, PTHCA > 0.8, especially > 0.85, such as > 0.9, including 1.

[0034] The term “product specificity” of an enzyme with respect to a cannabinoid acid may herein refer to the fraction of that cannabinoid acid produced from CBGA in the presence of the enzyme relative to the total amount of cannabinoid acids produced (see further below). For instance, if exposure of CBGA to a cannabinoid acid synthase results in 80% CBCA (and / or CBC), 15% CBDA (and / or CBD), and 5% THCA (and / or THC), the cannabinoid acid synthase has a product specificity PCBCA of 0.8, PCBDA of 0.15, and PTHCA of 0.05.

[0035] In particular, the coding sequence of the polynucleotide may encode an amino acid sequence of the protein.

[0036] In a sequence alignment between the amino acid sequence and a reference amino acid sequence, the sequence identity between the amino acid sequence and the reference amino acid sequence may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0037] In general, proteins with (highly) similar amino acid sequences are more likely to exhibit the same or similar function. This relationship between amino acid sequence similarity and protein function may, for example, be utilized to predict the function of a protein based on its sequence identity with proteins of known function, such as via annotation by sequence homology-based inference.

[0038] The term “sequence identity” herein refers to the percentage of identical characters (e.g., letters corresponding to amino acids in an amino acid sequence) between two aligned sequences (see further below). The higher the sequence identity between two proteins, the greater the likelihood that they exhibit the same or similar function.

[0039] Amino acid sequence alignments may especially be obtained using BLASTp (Protein BLAST) on the website of the National Center for Biotechnology Information (NCBI). Two sequences may be aligned via BLASTp, especially using default algorithm parameters, such as a BLOSUM62 matrix and a gap cost of 11 : 1 (existence extension).

[0040] Accordingly, in embodiments, the sequence alignment between the amino acid sequence and the reference amino acid sequence may be obtained using a BLOSUM62 matrix, with an existence gap cost of 11 and an extension gap cost of 1.In embodiments, the reference amino acid sequence may be selected from the group consisting of SEQ ID NO:8 11 14 1720, particularly from the group consisting of SEQ ID NO:8 11 14, i.e., especially from the group consisting of SEQ ID NO:8, SEQ ID NO:11, and SEQ ID NO: 14.

[0041] For instance, in embodiments, the reference amino acid sequence may be SEQ ID NO: 11. In such embodiments, the protein may especially be a CBCAS. Especially, the protein may exhibit a product specificity PCBCA for CBCA, typically wherein PCBCA > 0.8, such as > 0.9, especially > 0.95, such as > 0.99, including 1.

[0042] In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO: 11 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0043] As described above, the protein of SEQ ID NO: 11 includes a deletion of four amino acids relative to the protein of SEQ ID NO: 8. Accordingly, in further embodiments, the amino acid sequence comprises (or ‘has’) one or more gaps aligned to (amino acids at) positions 331-334 of SEQ ID NO:8 in a second amino acid sequence alignment. In particular, the second amino acid sequence alignment may be between the amino acid sequence and SEQ ID NO:8. Such embodiments may result in a particularly high product specificity PCBCA for CBCA.

[0044] In further embodiments, the reference amino acid sequence may be SEQ ID NO: 14. In such embodiments, the protein may be a CBDAS. Especially, the protein may exhibit a product specificity PCBDA for CBDA of > 0.6, such as > 0.7, especially > 0.75, such as > 0.8, especially > 0.85. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO: 14 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0045] In further embodiments, the reference amino acid sequence may be SEQ ID NO: 17. In such embodiments, the protein may be a CBCAS. Especially, the protein may exhibit a product specificity PCBCA for CBCA of > 0.5, especially > 0.6, such as > 0.7, especially > 0.75, such as > 0.8. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO: 17 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially atleast 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0046] In further embodiments, the reference amino acid sequence may be SEQ ID NO:20. In such embodiments, the protein may be a CBDAS. Especially, the protein may exhibit a product specificity PCBDA for CBDA of > 0.55, especially > 0.6, such as > 0.7, especially > 0.75, such as > 0.8. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO:20 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0047] In further embodiments, the reference amino acid sequence may be SEQ ID NO:8. In such embodiments, the protein may be particularly suitable for heterologous expression and / or exhibit particularly high enzymatic activity. Further, in such embodiments, the protein may be a CBCAS. Especially, the protein may exhibit a product specificity PCBCA for CBCA of > 0.45, such as > 0.5, especially > 0.55, such as > 0.6. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO:8 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0048] It will be clear to the person skilled in the art that the amino acid sequence of a protein does not necessarily need to have the same length as a reference amino acid sequence in order for the protein to retain the same function as the reference protein. In particular, the amino acid sequence of the protein may be substantially shorter or longer than the reference amino acid sequence. Furthermore, it will be clear to the person skilled in the art that a sequence alignment does not necessarily need to cover the full length of either the protein or the reference protein. However, in general, a larger sequence alignment with a reference sequence may be more informative than a smaller sequence alignment. Accordingly, while a sequence alignment does not necessarily need to cover the full length of either the protein or the reference protein, such sequence alignment may typically be over the entire length of the reference amino acid sequence. The amino acid sequence may have a length corresponding to the length of the reference amino acid sequence, or may be longer or shorter than the reference amino acid sequence.Hence, in embodiments, the sequence alignment may have a sequence alignment length of at least 70% of the total sequence length of the reference amino acid sequence, such as at least 80%, especially at least 90%, such as at least 95%, including 100%. Accordingly, the query coverage in BLASTp (or BLASTn), using the reference sequence as the query, may be at least 70%, such as at least 80%, especially at least 90%, such as at least 95%, including 100%.

[0049] In further embodiments, the sequence alignment between the amino acid sequence and the reference amino acid sequence may have a corresponding query coverage (with the reference amino acid sequence as the query) and a sequence identity, wherein the product of the query coverage (in %) and the sequence identity (in %) is at least 60%, especially at least 70%, such as at least 80%, especially at least 90%. In further embodiments, the product of the query coverage and the sequence identity is at least 95%, such as at least 97%, especially at least 98%, including 100%. For instance, in such embodiments, the reference amino acid sequence may be selected from the group consisting of SEQ ID NO:8, SEQ ID NO: 11, SEQ ID NO: 14, SEQ ID NO: 17, and SEQ ID NO:20, such as especially SEQ ID NO:8, or especially SEQ ID NO:11, or especially SEQ ID NO: 14, or especially SEQ ID NO: 17, or especially SEQ ID NO:20.

[0050] It will be clear to the person skilled in the art that multiple different coding sequences may encode the same protein due to codon redundancy. Further, it will be clear that different organisms may use different codon tables, meaning that the mapping of nucleotide triplets (‘codons’) to amino acids may vary between organisms. For explanatory purposes, specific polynucleotide sequences encoding the aforementioned proteins have been provided herein. Specifically, SEQ ID NO:1 corresponds to a coding sequence encoding SEQ ID NO:2. Similarly, SEQ ID NO:4 encodes SEQ ID NO:5, SEQ ID NO:7 encodes SEQ ID NO:8; SEQ ID NOTO encodes SEQ ID NO:11; SEQ ID NO:13 encodes SEQ ID NO:14; SEQ ID NO:16 encodes SEQ ID NO: 17; and SEQ ID NO: 19 encodes SEQ ID NO:20. However, it will be clear to the person skilled in the art that the polynucleotides of the invention are not limited to the specific polynucleotide sequences disclosed herein.

[0051] The phrase “a coding sequence encoding a protein” and similar phrases used herein, may herein refer to a nucleotide sequence, specifically a sequence of codons, that provides the genetic instructions for assembling a specific sequence of amino acids in the encoded protein through the biological processes of transcription and translation. To determine which amino acid sequence is encoded by a given coding sequence, the codon table of theintended expression host may be applied. In the absence of a designated expression host, the canonical codon table may be used as a reference instead.

[0052] In specific embodiments, the invention may provide an isolated or engineered polynucleotide comprising a coding sequence. The coding sequence may have a sequence identity of at least 70%, such as of at least 80%, especially at least 90%, with a reference DNA sequence, especially in a (nucleotide) sequence alignment between the coding sequence and the reference DNA sequence. In embodiments, the reference DNA sequence may be selected from the group consisting of SEQ ID NO:7 10 13 16 19, especially from the group consisting of SEQ ID NO:7 10 13, such as SEQ ID NOTO or SEQ ID NO: 13. The coding sequence may particularly encode a cannabinoid acid synthase, especially a CBCAS, a CBDAS, or a THCAS.

[0053] In further embodiments, the nucleotide sequence alignment may have a sequence alignment length of at least 70% of the total sequence length of the reference DNA sequence, such as at least 80%, especially at least 90%, such as at least 95%, including 100%.

[0054] Nucleotide sequence alignments may especially be obtained using BLASTn (Nucleotide BLAST) on the website of the National Center for Biotechnology Information (NCBI). In particular, two sequences may be aligned via BLASTn, especially using default algorithm parameters, such as megablast with 1,-2 Match / Mismatch scores and Linear Gap Costs. Hence, in embodiments, the sequence alignment of the coding sequence and the reference DNA sequence may be obtained using megablast, applying 1,-2 Match / Mismatch scores and Linear Gap Costs.

[0055] In further embodiments, the sequence alignment between the coding sequence and the reference DNA sequence may have a corresponding query coverage (with the reference DNA sequence as the query) and a sequence identity, wherein the product of the query coverage (in %) and the sequence identity (in %) is at least 70%, such as at least 80%, especially at least 90%. In further embodiments, the product of the query coverage and the sequence identity is at least 95%, such as at least 97%, especially at least 98%, including 100%. For instance, in such embodiments, the reference DNA sequence may be selected from the group consisting of SEQ ID NOT, SEQ ID NOTO, SEQ ID NO:13, SEQ ID NO:16, and SEQ ID NO:19.

[0056] In embodiments, the polynucleotide may comprise an expression cassette that includes the coding sequence encoding the protein. The expression cassette may further comprise a promoter arranged upstream of the coding sequence, and / or a terminator arranged downstream of the coding sequence, preferably the expression cassette comprises both the promoter and the terminator. In further embodiments, the promoter may be a constitutive promoter or an inducible promoter.The term “constitutive promoter” may herein refer to a promoter configured to facilitate continual transcription of a coding sequence downstream of the promoter, i.e., the promoter may essentially be unregulated. The term “inducible promoter” may herein refer to a promoter that (only) facilitates transcription of a coding sequence downstream of the promoter in dependence of an external stimulus, i.e., the promoter may be regulated in dependence of an external stimulus. The external stimulus may, for instance, be the presence of a specific compound in the extracellular environment. The external stimulus may further be an environmental temperature or light, especially light of a specific wavelength.

[0057] The terms “upstream” and “downstream” may refer to the relative arrangement of genetic elements, particularly with respect to the reading direction of a coding sequence, or especially with respect to the corresponding encoded amino acid sequence. Hence, a genetic element positioned at the 5’ end relative to another genetic element (especially with respect to the reading direction of a gene) may be referred to as “upstream”, such as a promotor upstream of the respective gene. Conversely, a genetic element positioned at the 3 ’ end relative to another genetic element, such as a terminator downstream of the respective gene, may be referred to as “downstream”.

[0058] Also provided herein is an expression vector comprising the polynucleotide of the invention. Typically, the polynucleotide is operably linked to regulatory elements suitable for expression in, for example, a host cell, such as promoters, enhancers, signal peptides, and the like, which may be endogenous or heterologous, and may be constitutive, inducible, or tissue-specific, depending on the intended expression system. The expression vector may be a plasmid in particular, including any plasmid described herein, a viral vector, a cosmid, a bacterial artificial chromosome, or a yeast artificial chromosome.

[0059] In a second aspect, the invention provides a protein, especially an isolated protein, or especially an engineered protein. In particular, an amino acid sequence of the protein may have at least 70% sequence identity, such as at least 80%, especially at least 90%, with a reference amino acid sequence, especially in a sequence alignment between the amino acid sequence and the reference amino acid sequence. The reference amino acid sequence may be selected from the group consisting of SEQ ID NO:8 11 14 17 20. The protein may especially be a cannabinoid acid synthase.

[0060] In further embodiments, the sequence alignment has a sequence alignment length of at least 80% of the reference amino acid sequence length.

[0061] As described above, the protein of the invention may have a higher enzymatic activity and / or increased product specificity relative to prior art cannabinoid acid synthasesand / or wild-type cannabinoid acid synthases. The protein may thus be a cannabinoid acid synthase, especially one or more of a CBCAS, a CBDAS, and a THCAS.

[0062] Accordingly, the invention may provide a protein, especially an engineered protein. The term “engineered” with regards to the protein may herein refer to a non-naturally-occurring protein, such as a protein encoded by an engineered polynucleotide (with an engineered coding sequence) or an (artificially) synthesized protein.

[0063] In particular, the amino acid sequence of the protein may exhibit at least 60% sequence identity to a reference amino acid sequence in an amino acid sequence alignment. The reference amino acid sequence may be selected from the group consisting of SEQ ID NO:8 11 14 17 20, especially from the group consisting of SEQ ID NO:8 11 14, such as SEQ ID NO: 11 or SEQ ID NO: 14.

[0064] In an amino acid sequence alignment between the amino acid sequence and the reference amino acid sequence, the sequence identity may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0065] In embodiments, the reference amino acid sequence may be SEQ ID NO: 11. Especially, the sequence identity between the amino acid sequence and SEQ ID NO:11 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%. In such embodiments, the protein may be a CBCAS. Particularly, the protein may exhibit a product specificity PCBCA for CBCA, especially wherein PCBCA > 0.8, such as > 0.9, especially > 0.95, such as > 0.99, including 1.

[0066] In further embodiments, the reference amino acid sequence may be SEQ ID NO: 11, and the amino acid sequence may further comprise one or more deletions relative to SEQ ID NO: 8. In such embodiments, the amino acid sequence may comprise one or more deletions aligned to (amino acids at) positions 331-334 of SEQ ID NO:8 in a second amino acid sequence alignment. The one or more deletions may especially comprise at least two deletions, such as at least three deletions, especially four deletions. Such embodiments may result in a particularly high product specificity PCBCA for CBCA.

[0067] In further embodiments, the reference amino acid sequence may be SEQ ID NO: 14. In such embodiments, the protein may be a CBDAS. Particularly, the protein mayexhibit a product specificity PCBDA for CBDA of > 0.6, such as > 0.7, especially > 0.75, such as > 0.8, especially > 0.85. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO: 14 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0068] In further embodiments, the reference amino acid sequence may be SEQ ID NO: 17. In such embodiments, the protein may be a CBCAS. Particularly, the protein may exhibit a product specificity PcBCAfor CBCA of > 0.5, especially > 0.6, such as > 0.7, especially > 0.75, such as > 0.8. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO: 17 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0069] In further embodiments, the reference amino acid sequence may be SEQ ID NO:20. In such embodiments, the protein may be a CBDAS. Particularly, the protein may exhibit a product specificity PCBDA for CBDA of > 0.55, especially > 0.6, such as > 0.7, especially > 0.75, such as > 0.8. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO:20 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.

[0070] In further embodiments, the reference amino acid sequence may be SEQ ID NO:8. In such embodiments, the protein may be particularly suitable for heterologous expression and / or exhibit a particularly high enzymatic activity. Further, in such embodiments, the protein may be a CBCAS. Especially, the protein may exhibit a product specificity PCBCA of > 0.45, such as > 0.5, especially > 0.55, such as > 0.6. In further embodiments, the sequence identity between the amino acid sequence and SEQ ID NO: 8 may be at least 60%, such as at least 70%, especially at least 80%. In further embodiments, the sequence identity may be at least 85%, such as at least 90%, especially at least 94%, such as at least 95%. In further embodiments, the sequence identity may be at least 97%, such as at least 98%, especially at least 99%, including 100%.Accordingly, in embodiments, the sequence alignment may have a length of at least 70% of the reference amino acid sequence length, such as at least 80%, especially at least 90%, such as at least 95%, including 100%.

[0071] In embodiments, the sequence alignment length may be at least 70% of the reference amino acid sequence length, wherein the sequence identity between the amino acid sequence and the reference amino acid sequence is at least 80%, such as at least 90%, especially at least 95%, such as at least 98%, including 100%.

[0072] In further embodiments, the sequence alignment length may be at least 80% of the reference amino acid sequence length, wherein the sequence identity between the amino acid sequence and the reference amino acid sequence is at least 80%, such as at least 90%, especially at least 95%, such as at least 98%, including 100%.

[0073] In further embodiments, the sequence alignment length may be at least 90% of the reference amino acid sequence length, wherein the sequence identity between the amino acid sequence and the reference amino acid sequence is at least 80%, such as at least 90%, especially at least 95%, such as at least 98%, including 100%.

[0074] In further embodiments, the sequence alignment length may be at least 95% of the reference amino acid sequence length, wherein the sequence identity between the amino acid sequence and the reference amino acid sequence is at least 80%, such as at least 90%, especially at least 95%, such as at least 98%, including 100%.

[0075] In a further aspect, the invention provides a method for producing (or ‘preparing’) a cannabinoid acid. In particular, the method may comprise exposing the protein of the invention to cannabigerolic acid to produce the cannabinoid acid. The cannabinoid acid may spontaneously convert into its corresponding cannabinoid, particularly via decarboxylation (as described herein). The method may further comprise (facilitating) converting the produced cannabinoid acid into its corresponding cannabinoid, for example, by applying heat and / or UV light.

[0076] In embodiments, the method may comprise expressing the (coding sequence of the) polynucleotide of the invention to produce the protein and subsequently exposing the protein to CBGA to convert the CBGA into a cannabinoid acid. In particular, the method may comprise expressing the polynucleotide in an expression host or in an in vitro expression system (or “cell-free expression system”), especially in an expression host, or especially in an in vitro expression system.

[0077] It will be clear to the person skilled in the art that the occurrence of the conversion of CBGA to a cannabinoid acid may further depend on the availability of co-factorsand co-substrates. Hence, in embodiments, the method may comprise exposing CBGA to the protein in the presence of suitable co-factors and / or co-substrates to facilitate the conversion. In particular, the conversion of CBGA to a cannabinoid acid may typically involve the reduction of flavin adenine dinucleotide (FAD) to the hydroquinone form of FAD (FADH2). Hence, in embodiments, the method may comprise exposing the CBGA to the protein in the presence of FAD. Further, FADH2 may be oxidized back to FAD, which can be reused for continued conversion of cannabigerolic acid into cannabinoid acid. FADH2 oxidation may, for instance, be facilitated by molecular oxygen (O2). Accordingly, in further embodiments, the method may comprise exposing CBGA to the protein in the presence of FAD and O2 (optionally dissolved in an aqueous medium).

[0078] It will further be clear to the person skilled in the art that the performance and the stability of proteins may depend on operational conditions, such as temperature, salinity, pH, etc. The person skilled in the art will be capable of selecting suitable operational conditions for CBGA conversion in the presence of the protein.

[0079] For instance, in embodiments, the method may comprise exposing CBGA to the protein at a temperature selected from the range of 20-70 °C, such as from the range of 30-60 °C, especially from the range of 40-60 °C, such as from the range of 40-50 °C.

[0080] Further, pH may affect the product specificity of cannabinoid acid synthases, e.g., a higher pH may favor the production of CBCA over CBDA and THCA.

[0081] In particular, in embodiments, the method may comprise exposing CBGA to the protein at a pH selected from the range of 6 - 8. Such pH may favor CBCA production.

[0082] In further embodiments, the method may comprise exposing CBGA to the protein at a pH selected from the range of 4-6. Such pH may favor CBDA and / or THCA production.

[0083] In embodiments, the amino acid sequence of the protein may have at least 80%, such as at least 90%, sequence identity with SEQ ID NO: 11 in a sequence alignment with the reference amino acid sequence. Additionally, the amino acid sequence may comprise one or more deletions aligned to positions 331-334 of SEQ ID NO:8 in a second amino acid sequence alignment (with SEQ ID NO: 8). The one or more deletions may especially comprise at least two deletions, such as at least three deletions, especially four deletions. In such embodiments, the cannabinoid acid produced by the protein may especially comprise CBCA.

[0084] In further embodiments, the amino acid sequence of the protein may have at least 80% sequence identity, such as at least 95%, with SEQ ID NO: 14. In such embodiments, the cannabinoid acid produced by the protein may especially comprise CBDA.In further embodiments, the cannabinoid acid may comprise THCA.

[0085] It will be clear to the person skilled in the art that cannabinoid acid synthases typically catalyze the conversion of CBGA into one or more of CBCA, CBDA, and THCA. Hence, depending on the protein, the method of the invention may also result in the production of one or more of CBCA, CBDA, and THCA and, by extent, one or more of CBC, CBD and THC.

[0086] The present invention is not limited to the specific amino acid sequences disclosed herein but further encompasses proteins with similar amino acid sequences, as defined in the appended claims. Further, the invention further encompasses any functional variants resulting from post-translational modifications of the protein. These modifications may include, but are not limited to, phosphorylation, glycosylation, acetylation, methylation, ubiquitination, and any other covalent or non-covalent modifications that alter the activity, stability, or function of the protein. The invention includes all such modified forms of the protein, as long as they retain the cannabinoid acid synthase activity described herein.

[0087] In a further aspect, the invention provides an engineered organism or part thereof, particularly a recombinant organism or part thereof. The engineered organism or part thereof may comprise a genome modification relative to a parent organism or part thereof, wherein the genome modification comprises a genomic insertion of the polynucleotide according to the invention. Alternatively or additionally, the engineered organism or part thereof may comprise a plasmid comprising the polynucleotide according to the invention. In particular, the engineered organism or part thereof may be configured to express the coding sequence of the polynucleotide, especially thereby providing the protein according to the invention.

[0088] The engineered organism or part thereof may be particularly suitable for the production of one or more cannabinoid acids, or their corresponding cannabinoid, as described herein. For instance, the engineered organism or part thereof may be used for cannabinoid production, wherein the cannabinoid is either extracted from the organism or part thereof, or obtained from its environment (due to secretion of the cannabinoid and / or the protein). Alternatively or additionally, the engineered organism or part thereof may be used for the production of the protein, wherein the produced protein may typically be used for cannabinoid production in a cell-free system.

[0089] For instance, in embodiments, the engineered organism or part thereof is configured to produce and secrete the protein. In further embodiments, the engineered organism or part thereof is configured to express the protein with a signal peptide, wherein the signalpeptide comprises a secretion tag (specific to the engineered organism or part thereof). For example, the polynucleotide may further encode the signal peptide. Such embodiments may be particularly suitable for separating the protein from the engineered organism or part thereof for subsequent use in, e.g., in vitro cannabinoid production.

[0090] In further embodiments, the engineered organism or part thereof is configured to express the protein with a purification tag, especially wherein the purification tag comprises a His-tag. For example, the polynucleotide may further encode the purification tag. Such embodiments may be particularly suitable for isolating the protein from the engineered organism or part thereof for subsequent use in, e.g., in vitro cannabinoid production.

[0091] Hence, the invention may further provide an engineered organism or part thereof, particularly a recombinant organism or part thereof. The term “recombinant organism” may herein refer to an organism comprising an engineered polynucleotide (as described herein), especially engineered DNA.

[0092] In embodiments, the engineered organism or part thereof comprises a genome modification relative to a parent organism or part thereof, for instance a wild-type organism or part thereof, or for instance a (pre-engineered) host organism or part thereof, or chassis organism or part thereof, wherein the genome modification comprises a genomic insertion of the polynucleotide according to the invention.

[0093] In further embodiments, the engineered organism or part thereof comprises a plasmid comprising the polynucleotide according to the invention,

[0094] The engineered organism or part thereof may especially be configured to produce the protein according to the invention. In particular, the engineered organism or part thereof may be configured to express the coding sequence of the polynucleotide. The polynucleotide may, for instance, comprise a promoter suitable for gene expression in the engineered organism or part thereof. Alternatively, the polynucleotide may be integrated into a genomic locus comprising a (native) promoter suitable for expression in the engineered organism or part thereof. Alternatively, the polynucleotide may be integrated into a plasmid locus comprising a promoter suitable for expression in the engineered organism or part thereof. For instance, in embodiments, the polynucleotide, especially the coding sequence, may be configured under control of a constitutive promoter.

[0095] In embodiments, the engineered organism or part thereof may be a (recombinant) plant or plant part, such as a cannabis plant or part thereof, i.e., a plant belonging to the genus Cannabis. In further embodiments, the engineered organism may especially belongto a subspecies selected from the group comprising C. sativa sativa (including formerly named C. ruderalis) and C. sativa indica.

[0096] A phrase such as ‘the engineered organism may belong to species X’ may herein especially refer to the engineered organism being derived from a (parent) organism belonging to species X.

[0097] Hence, in embodiments, the engineered organism or part thereof may be derived from a parent organism or part thereof, wherein the parent organism is a plant, typically belonging to the genus Cannabis, especially a subspecies selected from the group comprising C. sativa sativa (including formerly named C. ruderalis) and C. sativa indica. For instance, such engineered organism or part thereof may be provided by genetically modifying plant cells, tissues or seeds obtained from a Cannabis plant, such as using Jgrohactenwm-mediated transformation and / or using CRISPR-based gene editing.

[0098] In further embodiments, the engineered organism may be a (recombinant) filamentous fungus. Especially, the engineered organism may belong to a genus selected from the group consisting of Aspergillus, Trichoderma, Penicillium, and Myceliophthora.

[0099] In embodiments, the engineered organism may be a single-celled organism. The engineered organism may thus comprise an engineered cell, such as an engineered yeast cell, an engineered bacterial cell, or an engineered algal cell.

[0100] In further embodiments, the engineered organism may be a (recombinant) yeast. Especially, the engineered organism, such as the engineered cell, may belong to a genus selected from the group consisting of Saccharomyces, Kluyveromyces, Komagataella, Hansenula, and Yarrowia. For instance, in specific embodiments, the engineered cell may belong to the species Saccharomyces cerevisiae. Such an engineered organism may be particularly suitable for the production of the protein of the invention as such.

[0101] In further embodiments, the engineered organism may be a (recombinant) bacterium. Especially, the engineered organism, such as the engineered cell, may belong to a genus selected from the group consisting of Escherichia, Bacillus, Pseudomonas, and Streptomyces.

[0102] In further embodiments, the engineered organism may be a (recombinant) alga, especially a microalga. Especially, the engineered organism, such as the engineered cell, may belong to the group consisting of Chiorophytes, Charophyta, and Rhodophyta. In further embodiments, the engineered organism, such as the engineered cell, may be selected from the group consisting of green algae, red algae, Euglenoids, Chromista, Dinoflagellates and Cyanobacteria. Especially, the engineered organism, such as the engineered cell, may belongto a genus selected from the group consisting of Chlamydomonas, Chlorella, Phaeodactylum, Nannochlor opsis, Cyanobacterium, and Tetraselmis.

[0103] Alternatively, the engineered organism or part thereof may comprise a cell derived from a multi-celled organism, such as an insect cell.

[0104] In further embodiments, the engineered organism may be an insect cell. Especially, the engineered organism, such as the engineered cell, may be derived from a (recombinant) insect cell line. For instance, the engineered organism, especially the engineered cell, may belong to a genus selected from the group consisting of Spodoptera, Trichoplusia, and Drosophila.

[0105] Also provided herein is a method of producing an engineered organism, especially the engineered organism of the invention. Typically, the method comprises genomically modifying an organism, especially wherein the modifying comprises genomically inserting the polynucleotide of the invention and / or an expression vector as described herein that typically comprises the polynucleotide of the invention, such as a plasmid comprising the polynucleotide of the invention. It is to be understood that the organism typically is as described herein.

[0106] Further provided herein is a method of increasing or upregulating cannabinoid acid formation in an organism or part thereof, as well as the resulting (engineered, or recombinant) organism or part thereof obtainable, or obtained, by the method. In particular, the cannabinoid acid may be cannabichromenic acid (CBCA) and / or cannabidiolic acid (CBDA). It is to be understood that the organism or part thereof typically is as described herein. The method typically comprises expressing the (coding sequence of the) polynucleotide of the invention in an organism or part thereof, thereby increasing or upregulating cannabinoid acid formation when compared to its parent organism or part thereof.

[0107] In embodiments, the organism or part thereof may be a plant or plant part, such as a cannabis plant or part thereof, i.e., a plant belonging to the genus Cannabis, e.g., C. sativa sativa (including formerly named C. ruderalis) or C. sativa indica. In further embodiments, the organism or part thereof may be a filamentous fungus, such as belonging to a genus selected from the group consisting of Aspergillus, Trichoderma, Penicillium, and Myceliophthora. In yet further embodiments, the organism or part thereof may be a single-celled organism, such as a yeast cell, for example belonging to a genus selected from the group consisting of Saccharomyces, Kluyveromyces, Komagataella, Hansenula, and Yarrowia, particularly the species Saccharomyces cerevisiae, a bacterial cell, for example belonging to a genus selected from the group consisting of Escherichia, Bacillus, Pseudomonas, and Streptomyces, or analgal cell, especially a microalgal, for example belonging to the group consisting of Chiorophytes, Charophyta, Rhodophyta, green algae, red algae, Euglenoids, Chromista, Dinoflagellates and Cyanobacteria, especially belonging to a genus selected from the group consisting of Chlamydomonas, Chlorella, Phaeodactylum, Nannochloropsis, Cyanobacterium, and Tetraselmis. In still further embodiments, the organism or part thereof may be, or comprise, a cell derived from a multi-celled organism, such as an insect cell, for example belonging to a genus selected from the group consisting of Spodoptera, Trichoplusia, and Drosophila.

[0108] In still a further aspect, the invention provides the use of the isolated or engineered polynucleotide of the invention, the protein of the invention, and / or the engineered organism or part thereof of the invention, in the manufacture of a cannabinoid acid. In particular, the cannabinoid acid may be cannabichromenic acid (CBCA) and / or cannabidiolic acid (CBDA).

[0109] The embodiments described herein are not limited to a single aspect of the invention. For example, an embodiment describing the protein may, for example, further relate to the polynucleotide or to the method. Similarly, an embodiment of the polynucleotide may further relate to embodiments of the protein and of the engineered organism.

[0110] BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings in which corresponding reference symbols indicate corresponding parts, and in which: Fig. 1 schematically depicts embodiments of the polynucleotide, the protein, and the method; Figs. 2A-B schematically depict embodiments of the method and of the engineered organism; Figs. 3A-B schematically depict embodiments of the proteins. The drawings are not necessarily on scale.

[0112] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0113] Fig. 1 schematically depicts an embodiment of the polynucleotide 10 comprising a coding sequence 12 that encodes a protein 20. In particular, in the depicted embodiment, the expression of the coding sequence 12 is depicted to be under control of a promoter 11. The protein 20 may have at least 80% sequence identity with a reference amino acid sequence in a sequence alignment between its amino acid sequence and the reference amino acid sequence. In embodiments, the reference amino acid sequence may be selected from the group consisting of SEQ ID NO:8 11 14 1720.

[0114] Fig. 1 further schematically depicts an embodiment of the protein 20 as such.The protein 20 may be a cannabinoid acid synthase 25, i.e., an enzyme capable of catalyzing the conversion of CBGA 30 into a cannabinoid acid 40, which cannabinoid acid 40 may subsequently decarboxylate into a cannabinoid. As schematically depicted in Fig. 1, the CBGA 30 can be converted into CBDA 46 by a CBDAS 26, into THCA 47 by a THCAS 27, or into CBCA 48 by a CBCAS 28.

[0115] Fig. 1 further schematically depicts an embodiment of the method for producing (or ‘preparing’) a cannabinoid acid 40 using the protein 20 of the invention. In the depicted embodiment, the method comprises exposing cannabigerolic acid 30 to the protein 20 to produce the cannabinoid acid 40.

[0116] Fig. 2A schematically depicts another embodiment of the method. In the depicted embodiment, the method comprises exposing CBGA 30 to the protein 20 in the presence of a suitable cofactor 35, specifically in the presence of FAD 36, to provide the cannabinoid acid 40.

[0117] In particular, Fig. 2 A schematically depicts the transformation of a parent organism 50, specifically a parent yeast cell 52, by introducing a plasmid 15 comprising the polynucleotide 10 according to the invention. This transformation results in an engineered organism 60, particularly an engineered cell 61, and more specifically an engineered yeast cell 62. The engineered yeast cell 62 is cultivated in a suitable growth medium 160 such that the engineered yeast cell 62 produces the protein 20. Hence, the engineered organism 60 is configured to produce and secrete protein 20. Secretion may, for instance, result from the inclusion of a signal tag for protein secretion in the coding sequence 12 of the protein 20.

[0118] Fig. 2A further schematically depicts an embodiment of the engineered organism 60, particularly an engineered cell 61. In this embodiment, the engineered organism 60 comprises a plasmid 15 comprising the polynucleotide 10 according to the invention. Additionally, as depicted, the engineered organism 60 is configured to express the coding sequence 12 of the polynucleotide 10, thereby yielding protein 20.

[0119] In embodiments, the engineered yeast cell 62 may belong to a genus selected from the group comprising Saccharomyces, Kluyveromyces, Komagataella, Hanseniila, and Yarrowia.

[0120] Fig. 2B schematically depicts a further embodiment of the engineered organism 60. In this embodiment, the engineered organism 60 is an engineered cannabis plant 63, i.e., a plant belonging to the genus Cannabis, such as a plant of a subspecies selected from the group consisting of C. sativa sativa and C. sativa indica. In this embodiment, the engineered organism 60 may comprise (a) a genome modification relative to the parent organism 50, such asespecially a parent cannabis plant 53, wherein the genome modification comprises a genomic insertion of the polynucleotide 10, e.g., introduced into the genome using CRISPR genome editing or another targeted genetic modification technique. In the depicted embodiment, a cannabinoid acid 40 is obtained from the engineered organism 60. In further embodiments, the protein 20 may be obtained from the engineered organism 60.

[0121] Materials and methods

[0122] Unless specified otherwise, the experiments described herein were performed using the following materials and methods.

[0123] Phylogenetic tree inference

[0124] A berberine bridge-like (BBL) dataset was assembled, comprising 248 nucleotide sequences from Cannabis sativa, Humulus lupulus, and Trema orientale. To improve data quality, redundant and ambiguous sequences e.g., potential chimeras) were removed. Additionally, sequences with near-perfect identity to others, differing only by a few recent evolutionary mutations, were excluded, as their contribution to ancestral reconstruction was negligible. The resulting representative 77 sequences covered all major subclades, from C. sativa cannabinoid acid synthases to their closest T. orientale relatives. For subsequent analyses, frameshift-containing pseudogenes were adjusted by inserting “N” at indel sites. BBL nucleotide sequences were aligned with GENEIOUS (Geneious prime 2019), relying on Translation Alignment based on MAFFT v.7.450. The alignment was then manually curated, partitioned (CodonPosl CodonPos2, CodonPos3) in MESQUITE v.3.6, and subsequently used to construct a Bayesian gene tree in MRBAYES v.3.2.7. The resulting phylogenetic tree was used to identify clades and ancestors of interest.

[0125] Ancestral Sequence Reconstruction

[0126] Each ancestral sequence was reconstructed using both Bayesian statistics and Maximum Likelihood approaches. Bayesian inference was performed with MrBayes v.3.2.7, using abovementioned BBL alignment and the following parameters: nst = mixed, rates = gamma, shape = (all), 10 million generations, temperature heating 0.05, partitioning: CodonPosl CodonPos2, CodonPos3. Maximum Likelihood reconstruction was conducted in PAML v.4.9, using abovementioned BBL alignment and an unrooted version of the BBL gene tree. Three models were tested, based on nucleotide, codon, and amino acid analyses. The best-fit models were selected via likelihood ratio tests, with parameters as follows. Nucleotide: baseml, model = 7, Mgene = 4, alpha = 1, ncatG = 5. Codon: codeml, model = 0, alpha = 0.5, ncatG = 20, CodonFreq = 1, aaDist = 1, NSsites = 0. Amino acids: codeml, model = 3,alpha = 0.5, ncatG = 10, aaRatefile = jones. For each ancestor to reconstruct, sequences inferred by four Bayesian and Maximum Likelihood methods were analyzed. The following refinement steps were performed. First, as ancestral inference places a nucleotide at each alignment position, even if primarily consisting of gaps, the position of ancestral gaps was manually predicted, and associated nucleotides or amino acids were deleted. Next, the sequences obtained with the different models were compared, and Bayesian sequences were selected for further steps. The posterior probabilities of the Bayesian sequences were analyzed for each alignment position, and nucleotides with a posterior probability above 0.6 were retained. Positions with a posterior probability below 0.6 were classified as ambiguous. For the ambiguous positions, the results from Bayesian and Maximum Likelihood models were compared, and alternative nucleotides were considered (e.g., silent or missense mutation, amino acid properties, residue location). This procedure resulted in the replacement of two and one codon(s) of the Bayesian HCa and Ca sequences, respectively, by their second-best alternative in terms of Bayesian posterior probability.

[0127] Homology modeling and chimera design

[0128] Three-dimensional (3D) homology models of wild-type (WT) and ancestral BBE sequences were generated based on the THCAS crystal structure (PDB ID: 3VTE), using the SWISS-MODEL protein homology modeling server as available at https: / / swissmodel.expasy.org / . The BBE N-terminal signal peptide was excluded from the modelling. The 3D homology models were visualized, analyzed, and compared using PyMOL™ v.2.5.4 (Schrodinger, LLC) to investigate residue spatial positioning — i.e., proximity to the active site, the cofactor, or the enzyme surface — to aid in chimera design. Fig. 3 A depicts the relationship between two initial (ancestral) proteins 120, 120a, 120b and three chimeric proteins 20, 20a, 20b, 20c derived therefrom. In particular, a first initial protein 120a corresponds to HCa and a second initial protein 120b corresponds to Ca. The initial proteins thus both correspond to reconstructed ancestral sequences. Protein 20a corresponds to HCa Ca del, protein 20b corresponds to HCa Ca, and protein 20c corresponds to HCa_Ca-FAD. The initial proteins 120, 120a, 120b were compared to identify common amino acids, as well as different amino acids 125, that are either specific for HCa (reference 125a), or specific for Ca (reference 125 b). The sequence identity between initial proteins 120a and 120b was found to be 77%, the remainder corresponding to 116 mutations and an insertion 122 of four amino acids in Ca relative to HCa. Further, for each initial protein 120, 120a, 120b a substrate binding region 121, 121a, 121b and a FAD binding site 123, 123a, 123b were identified.In particular, Fig. 3A specifically depicts the parts of the substrate binding regions 121, 121a, 121b and the FAD binding sites 123, 123a, 123b that differ between HCa and Ca.

[0129] HCa Ca del 20a was generated by taking the backbone of HCa 120a, replacing the HCa substrate binding region 121a by the Ca substrate binding region 121b without the insert 122.

[0130] HCa Ca 20b was generated by taking the backbone of HCa 120a, replacing the HCa substrate binding region 121a by the Ca substrate binding region 121b with the insert 122.

[0131] HCa_Ca-FAD 20c was generated by taking the backbone of HCa 120a, replacing the HCa substrate binding region 121a and the HCa FAD binding site 123 a by the Ca substrate binding region 121b with the insert 122 and the CA FAD binding site 123b, respectively.

[0132] Fig. 3B depicts the relationship between two initial proteins 120, 120b, 120c and two chimeric proteins 20d,20e derived therefrom. In particular, initial protein 120b corresponds to Ca and initial protein 120c corresponds to C CBDAS. Protein 20d corresponds to Ca CBDAS, and protein 20e corresponds to Ca_CBDAS-FAD. The initial proteins 120, 120b, 120c were compared to identify common amino acids, as well as different amino acids 125, that are either specific for Ca (reference 125b2), or specific for C CBDAS (reference 125c). The sequence identity between initial proteins 120b and 120c was found to be 91%, the remainder corresponding to 45 mutations and one deletion. Further, for each initial protein 120, 120b, 120c a substrate binding region 121, 121b2, 121c and a FAD binding site 123, 123b2, 123c were identified. In particular, Fig. 3B specifically depicts the parts of the substrate binding regions 121, 121b2, 121c and the FAD binding sites 123, 123b2, 123c that differ between Ca and C CBDAS.

[0133] Ca CBDAS 20d was generated by taking the backbone of Ca, replacing the Ca substrate binding region 121b by the C CBDAS substrate binding region 121c.

[0134] Ca CBDAS FAD 20e was generated by taking the backbone of Ca, replacing the Ca substrate binding region 121b and the HCa FAD binding site 123b by the C CBDAS substrate binding region 121c and the C CBDAS FAD binding site 123c, respectively.

[0135] Expression and purification of proteins

[0136] Proteins were heterologously expressed in yeast and purified following the protocol described in VILLARD et al., “Natural gene variation in Cannabis sativa unveils a key region of cannabinoid synthase enzymes”, biorxiv (2023), DOI: https: / / doi.Org / 10.l 101 / 2023.08.30.555511, which is hereby herein incorporated by reference. In short: His-tagged coding sequences of THCAS (GenBank accessionAB057805.1), CBDAS (GenBank accession NM 001397936.1), Hop-BBE (HopBase accession 000840F.g23.tl), along with associated ancestors and chimeras, were domesticated by removal of restriction sites and plant signal peptides, synthesized by GenScript® (Leiden, The Netherlands), and subcloned into the pPICZaA expression vector (Invitrogen™, Thermo Fisher Scientific). Recombinant plasmids were transformed into the Komagataella phaffii strain X-33 (Invitrogen™, Thermo Fisher Scientific). Heterologous expression was achieved by culturing the recombinant K. phaffii in methanol-containing induction medium for two days. Resulting His-tagged enzymes were harvested, purified by affinity, and exchanged into assay buffer (sodium citrate, 100 mM, pH 5.0). Purified enzymes were used immediately for activity assay or were kept on ice until the activity assay. The presence of His-tagged proteins in the purified solution was confirmed with Western blot analysis using a 6x-His Tag Monoclonal Antibody conjugated to alkaline phosphatase (Thermo Fisher Scientific). Protein concentration was quantified by Bradford test, using bovine gamma globulin as the protein standard.

[0137] Enzyme in vitro assays

[0138] To determine which enzymes could metabolize CBGA, qualitative activity assays were performed by incubating 80-800 pg / mL of freshly produced enzymes (depending on enzyme production levels) in 80 pL of sodium citrate buffer (100 mM, pH 5.0) containing 100 pM of CBGA. After 60 min at 45°C, 700 rpm in a Thermomixer, reactions were stopped by adding 20 pL of 100% acetonitrile. To quantify enzymatic activity of the enzymes that could metabolize CBGA, standardized assays were performed using 75 pM CBGA, providing good solubility in the assay buffer and operating within the linear range of the Michaelis-Menten curve of these enzymes. Standardized assays were conducted by incubating 30 pg / mL of fresh enzymes in 80 pL of sodium citrate buffer (100 mM, pH 5.0) containing 75 pM CBGA. After 30 min at 30°C, 700 rpm, reactions were stopped by adding 0.25 volume 100% acetonitrile. All reactions were performed in triplicate.

[0139] Identification and quantification of cannabinoids

[0140] Reaction mixtures were filtered at 0.2 pm and analyzed by ultra-high performance liquid chromatography (UHPLC), following the procedure described in VILLARD et al., “Natural gene variation in Cannabis sativa unveils a key region of cannabinoid synthase enzymes”, biorxiv (2023), DOI: https: / / doi.org / 10.1101 / 2023.08.

[0141] 30.555511, which is hereby herein incorporated by reference. The mobile gradient phase was modified as follows (A / B; v / v): 60 : 40 (0-1 min), 25 : 75 (4 min), 10 : 90 (19 min), 0 : 100 (20 min), and 60 : 40 (25-30 min). Compounds were detected by UV absorbance scans at 270,254, 220, and 305 nm. Reaction products were identified by comparing their retention times to those of a standard for each (relevant) cannabinoid (acid).

[0142] Experiments

[0143] Yeast cells were transformed with plasmids containing different versions of cannabinoid acid synthases. Specifically, the yeast Komagataella phaffii was used, and the following sequences were tested:

[0144]

[0145] For each yeast cell line, the yeast cells were cultivated, after which the produced protein was purified and tested as described above. The following results were obtained:

[0146]

[0147]

[0148] C THCAS and C CBDAS are known cannabinoid acid synthases (see references in the table above) and were included as comparative examples. These specific sequences were selected because they encode commonly used reference enzymes, with well-documented properties known among experts in the field.

[0149] The expression levels represent the observed heterologous expression in K. phaffii. In particular, C_THCAS serves as the reference, with other sequences scored relative to C THCAS, where indicates about 50% lower expression, ‘+ / -‘ indicates a similar expression, ‘+’ indicates 50-100% higher expression, and ‘++’ indicates 100-300% higher expression.

[0150] The relative (enzymatic) activities correspond to the observed turnover numbers for CBGA conversion into a cannabinoid acid. These values are normalized relative to that of C THCAS.

[0151] N.A. indicates that the corresponding product was either not formed or below the detection limit.

[0152] As shown in the table above, HCa demonstrated particularly high suitability for heterologous expression. However, HCa exhibited no detectable cannabinoid acid synthase activity.

[0153] All other engineered proteins tested in this study displayed a substantially higher activity than C CBDAS. Notably, HCa_Ca-FAD also exhibited a superior activity compared to C THCAS.

[0154] Each of Ca, HCa_Ca-FAD, HCa Ca del and HCa Ca was found to exhibit particularly high product specificities PCBCA for CBCA compared to the reference enzymes. Notably, no reference cannabichromenic acid synthase was included in the analysis due to issues with heterologous expression of CBCAS. In particular, HCa_Ca-FAD and HCa Ca demonstrated a substantially higher or comparable enzymatic activity relative to C THCAS, while exhibiting a significantly higher product specificity PCBCA. Further, HCa_Ca-FAD was observed to outperform HCa Ca in terms of relative activity, i.e., the replacement of the HCa FAD binding site 123a by the Ca FAD binding site 123b appears beneficial for cannabinoid acid activity.Accordingly, in embodiments, the amino acid sequence of the protein may have at least 80% sequence identity with SEQ ID NO:8 (in a sequence alignment), wherein (in a second sequence alignment) relative to SEQ ID NO: 17 the amino acid sequence comprises one or more substitutions selected from the group consisting of V91M, M104L, A208E, and L214V, especially two or more of the substitutions, such as three or more of the substitutions, more especially all four of the (indicated) substitutions.

[0155] Furthermore, HCa Ca del was observed to convert CBGA exclusively into CBCA, with no detectable conversion into CBDA or THCA. Hence, HCa Ca del may outperform HCa_Ca-FAD as a cannabichromenic acid synthesis in terms of PCBCA, i.e., the omission of the insert 122 from Ca (see above) appears beneficial for PCBCA.

[0156] Accordingly, in embodiments, the amino acid sequence of the protein may have at least 80% sequence identity with SEQ ID NO: 11 (in a sequence alignment), wherein (in a second sequence alignment) relative to SEQ ID NO: 8 the amino acid sequence comprises a deletion at one or more of positions 331-334 (in SEQ ID NO:8), especially at two or more of the positions, such as three or more of the positions, especially at all of positions 331-334.Ca_CBDAS-FAD and Ca CBDAS exhibited high product specificities PCBDA for CBDA, while also demonstrating substantially higher enzymatic activities compared to C CBDAS. Notably, Ca_CBDAS-FAD retained a product specificity PCBDA that is only 10% lower than that of C CBDAS, yet displayed an enzymatic activity more than two times that of C CBDAS. It further appears that Ca_CBDAS-FAD outperforms Ca CBDAS as a cannabidiolic acid synthase in terms of both activity and PCBDA, i.e., the replacement of the Ca FAD binding site 123b by the C CBDAS FAD binding site 123c appears beneficial for CBDAS activity.

[0157] Accordingly, in embodiments, the amino acid sequence of the protein may have at least 80% sequence identity with SEQ ID NO: 14 (in a sequence alignment), wherein (in a second sequence alignment) relative to SEQ ID NO:20 the amino acid sequence comprises one or more substitutions selected from the group consisting of G152C, S158G, I202L, and G207A, especially two or more of the substitutions, such as three or more of the substitutions, more especially all four of the (indicated) substitutions.

[0158] The term “plurality” refers to two or more. Furthermore, the terms “a plurality of’ and “a number of’ may be used interchangeably.

[0159] The terms “substantially” or “essentially”, and similar terms, will be understood by the person skilled in the art. These terms may also include “entirely”, “completely”, “all”, etc. Hence, in embodiments, the adjectives “substantially” or “essentially” may be removedwithout altering the scope. Where applicable, “substantially” or “essentially” may refer to 90% or higher, such as 95% or higher, especially 99% or higher, even more especially 99.5% or higher, including 100%. Similarly, the terms ’’about” and “approximately” may denote 90% or higher, such as 95% or higher, especially 99% or higher, even more especially 99.5% or higher, including 100%. For numerical values, it is to be understood that the terms “substantially”, “essentially”, “about”, and “approximately” may indicate a range of 90-110%, such as 95-105%, especially 99-101%, relative to the referenced value.

[0160] The term “comprise” also includes embodiments wherein “comprises” means “consists of’. The term "comprising" may in an embodiment refer to "consisting of' but may in another embodiment also refer to "containing at least the defined species and optionally one or more other species".

[0161] The term “and / or” especially refers to one or more of the elements mentioned before and after “and / or”. For instance, the phrase “item 1 and / or item 2” and similar phrases may refer to one or more of item 1 and item 2.

[0162] Furthermore, the terms “first”, “second”, “third” and the like in the description and claims are used for differentiation between similar elements and do not necessarily imply a specific sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments of the invention may operate in other sequences than described or illustrated herein.

[0163] The term “further embodiment”, and similar terms, may refer to an embodiment comprising the features of a previously described embodiment, but may also refer to an alternative embodiment.

[0164] It should be noted that the embodiments are intended to illustrate, rather than limit, the invention, and that those skilled in the art will be able to design alternative embodiments without departing from the scope of the appended claims.

[0165] In the claims, any reference signs in parentheses shall not be construed as limiting the claim.

[0166] Use of the verb "to comprise" and its conjugations does not exclude the presence of additional elements or steps beyond those explicitly stated in a claim. Unless the context clearly requires otherwise, throughout the description and the claims, the terms “comprise”, “comprising”, “include”, “including”, “contain”, “containing” and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; meaning “including, but not limited to”.The use of "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.

[0167] The various aspects discussed in this patent can be combined in order to provide additional advantages. Further, the person skilled in the art will understand that embodiments can be combined, and that also more than two embodiments can be combined. Furthermore, some of the features can form the basis for one or more divisional applications.

Claims

CLAIMS:

1. An isolated or engineered polynucleotide (10) comprising a coding sequence (12), wherein the coding sequence (12) encodes a protein (20), wherein an amino acid sequence of the protein (20) has at least 85% sequence identity with a reference amino acid sequence in a sequence alignment between the amino acid sequence and the reference amino acid sequence, wherein the reference amino acid sequence is selected from the group consisting of SEQ ID NO:8 11 14 1720, and wherein the protein (20) is a cannabinoid acid synthase (25).

2. The polynucleotide (10) according to claim 1, wherein the reference amino acid sequence is SEQ ID NO:11, and wherein the protein (20) is a cannabichromenic acid synthase (28).

3. The polynucleotide (10) according to claim 2, wherein the amino acid sequence comprises one or more deletions aligned to positions 331-334 of SEQ ID NO:8 in a second amino acid sequence alignment, and wherein the protein (20) has a product specificity PCBCA for cannabichromenic acid (48), PCBCA being defined as the fraction of cannabichromenic acid produced from cannabigerolic acid relative to the total amount of cannabinoid acids produced, and wherein PCBC > 0.9.

4. The polynucleotide (10) according to claim 1, wherein the reference amino acid sequence is SEQ ID NO: 14, wherein the amino acid sequence of the protein (20) has at least a 95% sequence identity with the reference amino acid sequence in the sequence alignment, and wherein the protein (20) is a cannabidiolic acid synthase (26).

5. The polynucleotide (10) according to any one of the preceding claims, wherein the sequence identity between the amino acid sequence and the reference amino acid sequence in the sequence alignment is at least 95%.

6. The polynucleotide (10) according to any one of the preceding claims, wherein the sequence alignment has a sequence alignment length, wherein the sequence alignment length is at least 80% of a sequence length of the reference amino acid sequence.

7. A protein (20), wherein an amino acid sequence of the protein (20) has at least 85% sequence identity with a reference amino acid sequence in a sequence alignment between the amino acid sequence and the reference amino acid sequence, wherein the reference amino acid sequence is selected from the group consisting of SEQ ID NO:8 11 14 1720, and wherein the protein (20) is a cannabinoid acid synthase (25).

8. The protein (20) according to claim 7, wherein the reference amino acid sequence is SEQ ID NO:11, wherein the amino acid sequence comprises one or more deletions aligned to positions 331-334 of SEQ ID NO:8 in a second amino acid sequence alignment, wherein the protein (20) is a cannabichromenic acid synthase (28).

9. The protein (20) according to claim 7, wherein the reference amino acid sequence is SEQ ID NO: 14, wherein the amino acid sequence of the protein (20) has at least a 95% sequence identity with the reference amino acid sequence in the sequence alignment, and wherein the protein (20) is a cannabidiolic acid synthase (26).

10. The protein (20) according to any one of the preceding claims 7-9, wherein the sequence identity between the amino acid sequence and the reference amino acid sequence in the sequence alignment is at least 95%.

11. A method for producing a cannabinoid acid (40) using the protein (20) according to any one of the preceding claims 7-10, wherein the method comprises exposing cannabigerolic acid (30) to the protein (20) to produce the cannabinoid acid (40).

12. The method according to claim 11, wherein the cannabinoid acid (40) comprises cannabichromenic acid (48), and wherein the protein (20) is according to any one of claims 7, 8, and 10.

13. The method according to claim 11, wherein the cannabinoid acid (40) comprises cannabidiolic acid (46), and wherein the protein (20) is according to any one of claims 7, 9, and 10.

14. An engineered organism (60) or part thereof, wherein the engineered organism (60) or part thereof comprises: (a) a genome modification relative to a parent organism (50) orpart thereof, wherein the genome modification comprises a genomic insertion of the polynucleotide (10) according to any one of the preceding claims 1-6; and / or (b) a plasmid (15) comprising the polynucleotide (10) according to any one of the preceding claims 1-6.

15. The engineered organism (60) or part thereof according to claim 14, wherein the engineered organism (60) is a yeast (62).

16. The engineered organism (60) or part thereof according to claim 14, wherein the engineered organism (60) is a yeast belonging to a genus selected from the group consisting of Saccharomyces, Kluyveromyces, Komagataella, Hansemila, and Yarrowia.

17. The engineered organism (60) or part thereof according to claim 14, wherein the engineered organism (60) is a plant.

18. The engineered organism (60) or part thereof according to claim 14, wherein the engineered organism (60) is a cannabis plant (63).

19. The engineered organism (60) or part thereof according to claim 14, wherein the engineered organism (60) is an alga.

20. Use of the isolated or engineered polynucleotide of any one of claims 1-6, the protein of any one of claims 7-10, and / or the engineered organism or part thereof of any one of claims 14-19 in the manufacture of a cannabinoid acid, preferably cannabichromenic acid and / or cannabidiolic acid.