Enzymes for luciferin biosynthesis and their use
By identifying and characterizing fungal luciferin biosynthesis proteins, the challenges of low bioluminescence intensity and toxicity in existing bioluminescent systems are addressed, enabling stable and cost-effective luciferin synthesis for enhanced bioluminescence in eukaryotic cells.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LIGHT BIO INC
- Filing Date
- 2026-02-19
- Publication Date
- 2026-05-26
AI Technical Summary
The existing bioluminescent systems, particularly those involving fungal luciferases, face limitations due to the lack of understanding of enzymes that promote luciferin synthesis and the reduction of oxyluciferin to luciferin, leading to low bioluminescence intensity and toxicity issues in eukaryotic cells, limiting their application in organismic medical studies and other bioluminescent applications.
Identification and characterization of fungal luciferin biosynthesis proteins, including hispidin hydroxylase, hispidin synthase, and caffeylpyruvate hydrolase, along with their nucleic acid sequences, which catalyze the conversion of precursor compounds to luciferin and preluciferin, enabling stable and non-toxic bioluminescent systems in eukaryotic cells.
Enables simple and inexpensive methods for luciferin synthesis, allowing for autonomous bioluminescent systems in eukaryotic cells with enhanced bioluminescence intensity, overcoming toxicity issues and expanding the applicability of bioluminescent systems in organismic medical studies and other applications.
Smart Images

Figure 2026086812000106 
Figure 2026086812000107 
Figure 2026086812000108
Abstract
Description
Technical Field
[0001] The group of the present invention relates to the fields of organism engineering and genetic engineering. In particular, the present invention relates to enzymes of the bioluminescence system of fungi.
Background Art
[0002] An enzyme capable of catalyzing the oxidation of a low molecular weight compound of luciferin with light emission or bioluminescence is called "luciferase". When luciferin is oxidized, oxyluciferin is released from the complex with the luciferase enzyme.
[0003] Luciferase is widely used as a reporter gene in many organismic medical applications and organism engineering. For example, in carcinogenicity studies in animal models, in methods for detecting toxic substances in microorganisms or media, luciferase is used as an indicator for determining the concentration of various substances, for determining the viability of cells and the activity of promoters or other components of biological systems, and for visualizing the passage of signal transduction cascades, etc. [Non-Patent Document 1; Non-Patent Document 2; Non-Patent Document 3]. Many applications of luciferase are described in the references [Non-Patent Document 4; Non-Patent Document 5; Non-Patent Document 6]. All the main applications of luciferase are based on the detection of the light emitted according to the phenomenon or signal being studied. Such detection is, in principle, carried out using a luminometer or a modified optical microscope.
[0004] Thousands of species capable of bioluminescence are known, and among them, about a dozen luciferins with various structures and dozens of corresponding luciferase enzymes have been reported. It has been found that the bioluminescence system has occurred independently in various organisms in the process of evolution more than 40 times [Non-Patent Document 7; Non-Patent Document 8].
[0005] A group of insect luciferases that catalyze the oxidation of D-luciferin have been reported [Non-Patent Literature 9; Non-Patent Literature 10]. A group of luciferases that catalyze the oxidation of coelenterazine have been reported [Non-Patent Literature 11]. The bioluminescent system of ostracods of the genus *Odontolabis* is known, characterized by high luciferin chemical activity and high luciferase stability [Non-Patent Literature 12]. The bioluminescent systems of dinoflagellates and krill are also known. Currently, genes encoding three types of luciferases have been cloned from this group [Non-Patent Literature 11]. However, this system is still not well studied, and in particular, the complete luciferase sequence has not yet been established.
[0006] In recent years, a group of luciferases and luciferins from the fungal bioluminescence system have been reported. Fungal bioluminescence has been known for several centuries, but fungal luciferin was only identified in 2015: it was found to be 3-hydroxyhispidin, a metabolite capable of penetrating the cell membrane [Non-Patent Literature 13]. The same publication confirmed the existence of an enzyme capable of hydroxylating hispidin in fungal lysates to form luciferin, but this enzyme was not identified. Patent Literature 1 reports several fungal luciferase genes containing luciferin in the form of 3-hydroxyhispidin having the following structure.
[0007] [ka]
[0008] It has been found that fungal luciferases can also catalyze the photo-emission oxidation of other compounds having the structures shown in Table 1 [Non-Patent Literature 14]. All of these compounds, which are fungal luciferins containing 3-hydroxyhispidin, belong to the group 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one and have the following general formula: where R is aryl or heteroaryl.
[0009] [ka]
[0010] [Table 1]
[0011] [Table 2]
[0012] [Table 3]
[0013] The enzymes that promote either the synthesis of luciferin in vivo or the reduction of oxyluciferin to luciferin are unknown in the vast majority of cases. Therefore, most bioluminescent applications of luciferin involve introducing exogenous luciferase-containing luciferin (e.g., cell culture medium or organism) into the system. As a result, the use of bioluminescent systems remains limited for many reasons, particularly the low ability of many luciferins to penetrate cell membranes, the chemical instability of luciferin, and the complex, multi-step, and costly process of luciferin synthesis.
[0014] In the only bioluminescent system reported in marine bacteria, an enzyme that promotes luciferin synthesis has been identified. However, this system differs significantly from other bioluminescent systems. Bacterial luciferin (myristicaldehyde) is oxidized during the reaction but does not emit light [Non-Patent Literature 11]. In addition to luciferin, other major components of the bioluminescence reaction include NAD (nicotinamide adenine dinucleotide) and FMN-H2 (flavin mononucleotide). The oxidized derivative of FMN-H2 acts as the true light source. To date, the bioluminescent system of marine bacteria is the only system that is fully encoded in a heterologous expression system and is considered to be the closest to the prior art of the present invention. However, this system is generally only applicable to prokaryotic organisms. To obtain autonomous bioluminescence, luciferase (luxA and luxB heterodimers) that acts as a bioluminescent substrate and the luxCDABE operon encoding the luxCDE luciferin biosynthesis protein are used (Non-Patent Literature 15). In 2010, autonomous bioluminescence in human cells was achieved using this system. However, the bioluminescence intensity level was low, only 12 times higher than the signal emitted from non-bioluminescent cells, making it virtually impossible to apply the developed system to solve the given problem [Non-Patent Literature 16]. Attempts to increase the luminescence intensity failed because the bacterial components were toxic to eukaryotic cells [Non-Patent Literature 17].
[0015] From this perspective, identifying enzymes that promote the stable and / or abundant synthesis of luciferin from precursor compounds present in cells, and the reduction of oxyluciferin to luciferin, is an urgent task. Identifying such enzymes would enable simple and inexpensive methods for luciferin synthesis and pave the way for the creation of autonomous bioluminescent systems. Among these, bioluminescent systems that are non-toxic to eukaryotic cells are of particular interest. [Prior art documents] [Patent Documents]
[0016] [Patent Document 1] Russian Patent Application Publication No. 2017102986 Specification
Non-licensed literature
[0017]
Non-licensed literature 1
Non-licensed Document 2
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed Document 8
Non-licensed literature 9
[0018] The applicants deciphered the steps of luciferin biosynthesis in the bioluminescent system of fungi and identified the enzymes involved in the cyclic cycle of fungal luciferin, as well as the nucleic acid sequences encoding those enzymes.
[0019] The following flowchart illustrates the stages of fungal luciferin turnover.
[0020] [ka] [Means for solving the problem]
[0021] Therefore, the present invention first provides isolated fungal luciferin biosynthesis proteins and nucleic acids encoding them.
[0022] In preferred embodiments, the present invention provides hispidin hydroxylase characterized by amino acid sequences selected from the following sequence numbers: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, as well as essentially similar proteins, homologs, variants, and derivatives of these hispidin hydroxylases.
[0023] In some embodiments, the hispidine hydroxylase of the present invention is characterized by an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, within the range of at least 350 amino acids, to an amino acid sequence selected from the following sequence number group: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0024] In some embodiments, the amino acid sequence of the hispidine hydroxylase of the present invention is characterized by the presence of several consensus sequences delimited by non-conservative amino acid insertion segments, characterized by the following sequence numbers: 29-33.
[0025] The hispidine hydroxylase of the present invention catalyzes the conversion of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0026] [ka]
[0027] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0028] [ka]
[0029] The present invention also provides hispidin synthases characterized by amino acid sequences selected from the following sequence numbers: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, as well as essentially similar proteins, homologs, variants, and derivatives of these hispidin synthases.
[0030] In some embodiments, the amino acid sequence of the hispidin synthase of the present invention is characterized by the presence of several consensus sequences delimited by non-conservative amino acid insertion segments, characterized by the following SEQ ID NOs: 56-63.
[0031] In some embodiments, the hispidin synthase of the present invention is characterized by an amino acid sequence having at least 40% identity, for example, at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0032] The hispidin synthase of the present invention catalyzes the conversion of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula, where R is selected from the aryl or heteroaryl group.
[0033] [ka]
[0034] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0035] [ka]
[0036] The present invention also provides caffeylpyruvate hydrolases characterized by amino acid sequences selected from the following sequence numbers: 65, 67, 69, 71, 73, and 75, as well as essentially similar proteins, homologs, variants, and derivatives of these caffeylpyruvate hydrolases.
[0037] In some embodiments, the amino acid sequence of the caffeylpyruvate hydrolase of the present invention is characterized by the presence of several consensus sequences delimited by non-conservative amino acid insertion segments, characterized by the following sequence numbers: 76-78.
[0038] In some embodiments, the caffeylpyruvate hydrolase of the present invention is characterized by an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity) of an amino acid sequence selected from the following SEQ ID NOs: 65, 67, 69, 71, 73, 75.
[0039] The caffeylpyruvate hydrolase of the present invention catalyzes the conversion of 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid to 3-arylacrylic acid. The 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid has the following structural formula, where R is aryl or heteroaryl.
[0040] [ka]
[0041] The 3-arylacrylic acid in question has the following structural formula.
[0042] [ka]
[0043] In a preferred embodiment, the hispidin hydroxylase of the present invention catalyzes the conversion of fungal luciferin to preluciferin, for example, the conversion of hispidin to 3-hydroxyhispidin.
[0044] In a preferred embodiment, the hispidin synthase of the present invention catalyzes the conversion of a preluciferin precursor to preluciferin, for example, the conversion of caffeic acid to hispidin.
[0045] In preferred embodiments, the caffeylpyruvate hydrolase of the present invention catalyzes the conversion of fungal oxyluciferin to a preluciferin precursor, for example, the conversion of caffeylpyruvate to caffeic acid.
[0046] The present invention also provides applications for proteins. The protein has an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, within the range of at least 350 amino acids, to an amino acid sequence selected from the following sequence number group: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity), and / or a consensus sequence having sequence numbers 29-33 separated by non-conserved amino acid insertion segments as hispidin hydroxylase. The hispidine hydroxylase catalyzes the in vitro or in vivo conversion of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0047] [ka]
[0048] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0049] [ka]
[0050] The present invention also provides applications for proteins. The protein has an amino acid sequence having at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity), and / or a consensus sequence having SEQ ID NOs. 56-63 separated by non-conserved amino acid insertion segments as a hispidin synthase. The hispidin synthase catalyzes the in vitro or in vivo conversion of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula.
[0051] [ka]
[0052] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0053] [ka]
[0054] The present invention also provides applications for proteins. The proteins have an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity) to an amino acid sequence selected from the following sequence numbers: 65, 67, 69, 71, 73, 75, and / or a consensus sequence having sequence numbers 76-78 separated by non-conserved amino acid insertion segments as caffeyylpyruvate hydrolase. The caffeyylpyruvate hydrolase catalyzes the in vitro or in vivo conversion of 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid to 3-arylacrylic acid. The 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid has the following structural formula, where R is aryl or heteroaryl.
[0055] [ka]
[0056] The 3-arylacrylic acid in question has the following structural formula.
[0057] [ka]
[0058] The present invention also provides nucleic acids encoding the hispidine hydroxylase, hispidine synthase, and caffeylpyruvate hydrolase.
[0059] In some embodiments, the nucleic acid encoding hispidine hydroxylase has an amino acid sequence selected from the following group: (a) The following sequence numbers: amino acid sequences presented as 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28; (b) an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, within the range of at least 350 amino acids, selected from the following sequence numbers: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity); (c) Amino acid sequences including the consensus sequences presented as sequence numbers 29-33 below.
[0060] In some embodiments, the nucleic acid encoding hispidin synthase has an amino acid sequence selected from the following group: (a) The amino acid sequences presented as the following sequence numbers: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55; (b) Amino acid sequences having at least 40% identity with an amino acid sequence selected from the following sequence numbers: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, for example, at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity); (c) Amino acid sequences including the consensus sequences presented as sequence numbers 56-63 below.
[0061] In some embodiments, the nucleic acid encoding caffeylpyruvate hydrolase has an amino acid sequence selected from the following group: (a) The amino acid sequences presented as sequence numbers 65, 67, 69, 71, 73, and 75 below; (b) Amino acid sequences having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity); (c) Amino acid sequences containing consensus sequences having sequence numbers 76-78 separated by non-conservative amino acid insertion segments.
[0062] The present invention also provides applications of nucleic acids encoding proteins. The protein has an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, within the range of at least 350 amino acids, to an amino acid sequence selected from the following sequence number group: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, for example, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity), and / or a consensus sequence having sequence numbers 29-33 separated by non-conservative amino acid insertion segments that generate hispidine hydroxylase in vitro or in vivo. The hispidine hydroxylase catalyzes the conversion of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0063] [ka]
[0064] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0065] [ka]
[0066] The present invention also provides applications of nucleic acids encoding proteins. The proteins have an amino acid sequence having at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 96%, 97%, 98%, 98%, or 99%) of an amino acid sequence selected from the following sequence numbers: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, and / or a consensus sequence having sequence numbers 56-63 separated by non-conservative amino acid insertion segments that generate hispidin synthase in vitro or in vivo. The hispidin synthase catalyzes the conversion of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula.
[0067] [ka]
[0068] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0069] [ka]
[0070] The present invention also provides applications of nucleic acids encoding proteins. The protein has an amino acid sequence having at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 96%, 97%, 98%, 98%, or 99% identity) to an amino acid sequence selected from the following sequence numbers: 65, 67, 69, 71, 73, 75, and / or a consensus sequence having sequence numbers 76-78 separated by non-conservative amino acid insertion segments that generate caffeylpyruvate hydrolase in vivo or in vitro. The caffeylpyruvate hydrolase catalyzes the conversion of 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid to 3-arylacrylic acid, and 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid has the following structural formula, where R is aryl or heteroaryl.
[0071] [ka]
[0072] The 3-arylacrylic acid in question has the following structural formula.
[0073] [ka]
[0074] The present invention also provides a fusion protein comprising at least one hispidin hydroxylase and / or at least one hispidin synthase and / or at least one caffeypyruvate hydrolase of the present invention, which are operationally, directly, or crosslinked via an amino acid linker, and a luciferase capable of oxidizing fungal luciferin accompanied by an intracellular localization signal and / or a signal peptide and / or luminescence.
[0075] Luciferases capable of oxidizing fungal luciferin with photoemission are known in the art. In preferred embodiments, the luciferase has an amino acid sequence substantially similar to or identical to an amino acid sequence selected from the following sequence number group: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98. For example, the luciferase may have an amino acid sequence that is at least 40% identical, for example, at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 70%, or at least 75%, or at least 80%, or at least 85% identical to an amino acid sequence selected from the following sequence number group: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98. In many embodiments, the amino acid sequence of the luciferase has at least 90% identity, or at least 95% identity (for example, at least 96%, 97%, 98%, 98%, or 99% identity) with an amino acid sequence selected from the following sequence number group: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98.
[0076] In some embodiments, the fusion protein has an amino acid sequence having SEQ ID NO: 101.
[0077] The present invention also provides nucleic acids encoding the aforementioned fusion protein.
[0078] The present invention also provides an expression cassette, which comprises (a) a transcription initiation domain functional in a host cell; (b) a nucleic acid encoding a fungal luciferin biosynthesis enzyme, i.e., hispidin synthase, hispidin hydroxylase, or caffeylpyruvate hydrolase, or a fusion protein according to the present invention; and (c) a transcription termination domain functional in a host cell.
[0079] The present invention also provides a vector for transferring nucleic acids to host cells. The vector comprises a fungal luciferin biosynthesis enzyme of the present invention, namely hispidin synthase, hispidin hydroxylase, or caffeylpyruvate hydrolase, or a nucleic acid encoding a fusion protein of the present invention.
[0080] The present invention also provides a host cell, which comprises an expression cassette containing nucleic acids encoding the hispidin synthase and / or hispidin hydroxylase and / or caffeypyruvate hydrolase of the present invention, either as part of an extrachromosomal element or incorporated into the cell's genome as a result of introducing the cassette into the cell. Such a cell produces at least one of the fungal luciferin biosynthesis enzymes by expression of the introduced nucleic acid.
[0081] The present invention also provides antibodies obtained using the protein of the present invention.
[0082] The present invention also provides a method for producing fungal luciferin having the chemical formula 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one and the following structural formula, either in vitro or in vivo, where R is aryl or heteroaryl.
[0083] [ka]
[0084] The above method involves combining, under physiological conditions, at least one molecule of the hispidine hydroxylase of the present invention with at least one molecule of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one having the following structural formula, at least one NAD(P)H molecule, and at least one oxygen molecule.
[0085] [ka]
[0086] The present invention also provides a method for producing fungal preluciferin having the chemical formula 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one and the following structural formula, either in vitro or in vivo, where R is aryl or heteroaryl.
[0087] [ka]
[0088] The above method involves combining, under physiological conditions, at least one molecule of 3-arylacrylic acid having the following structural formula with at least one molecule of the hispidin synthase of the present invention, at least one molecule of coenzyme A (CoA), at least one molecule of ATP, and at least two molecules of malnyl-CoA.
[0089] [ka]
[0090] The present invention also provides a method for producing fungal luciferin in vitro or in vivo. The method comprises combining, under physiological conditions, at least one molecule of the hispidin hydroxylase of the present invention with at least one molecule of 3-arylacrylic acid, at least one molecule of the hispidin synthase of the present invention, at least one molecule of coenzyme A, at least one molecule of ATP, at least two molecules of malonyl-CoA, at least one molecule of NAD(P)H, and at least one molecule of molecular oxygen.
[0091] Methods for producing fungal luciferin and preluciferin can be carried out in cells or organisms. In this case, the method involves introducing a cellular nucleic acid encoding the corresponding luciferin biosynthesis enzyme (hispidin synthase and / or hispidin hydroxylase) capable of expressing the enzyme into a cell or organism. In preferred embodiments, the nucleic acid is introduced into the cell or organism as part of an expression cassette or vector of the present invention.
[0092] In some embodiments, a nucleic acid encoding a 4'-phosphopantotheinyl transferase capable of transferring 4-phosphopantotheinyl from coenzyme A to serine within the acyl transfer domain of a polyketide synthase is additionally introduced into the cell or organism. In some embodiments, the 4'-phosphopantotheinyl transferase has an amino acid sequence substantially similar to or identical to that of SEQ ID NO: 105.
[0093] The present invention also provides applications of polyketide synthases (PKS) having amino acid sequences that are at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% identical to sequences selected from the following sequence numbers: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, and 139, which generate hispidin in vitro or in vivo.
[0094] In some embodiments, a method for preparing hispidin includes combining at least one molecule of PKS with at least two molecules of malonyl-CoA and at least one molecule of caffeyl-CoA under physiological conditions. In some embodiments, the method includes combining at least one molecule of PKS with at least two molecules of malonyl-CoA, at least one molecule of caffeic acid, at least one molecule of coenzyme A, at least one molecule of coumarate-CoA ligase, and at least one molecule of ATP under physiological conditions.
[0095] For the purposes of the present invention, any coumarate-CoA ligase that catalyzes the conversion of caffeic acid to caffeyl-CoA can be used. For example, the coumarate-CoA ligase may have an amino acid sequence that matches the sequence of SEQ ID NO: 141 by at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%.
[0096] The reaction described above can be carried out by any of the above methods, instead of the reaction for producing fungal preluciferin from a preluciferin precursor using the hispidin synthase of the present invention. For example, the reaction can be carried out in a cell or organism by introducing an expression cassette having a nucleic acid encoding PKS into the cell or organism. If necessary, a nucleic acid encoding coumarate-CoA ligase can also be additionally introduced into the cell or organism.
[0097] In some embodiments, nucleic acids encoding 3-arylacrylic acid biosynthetic enzymes are further introduced into the same cells or organisms. For example, the 3-arylacrylic acid biosynthetic enzyme may be a nucleic acid encoding a tyrosine ammonia lyase having an amino acid sequence substantially similar to, or identical to, that of Rhodobacter capsulatus tyrosine ammonia lyase having sequence number 107, or a nucleic acid encoding the HpaB and HpaC components of 4-hydroxyphenylacetate 3-monooxygenase reductase having an amino acid sequence substantially similar to that of E. coli 4-hydroxyphenylacetate 3-monooxygenase reductase having sequences of HpaB and HpaC components having sequences 109 and 111. In some embodiments, nucleic acids encoding phenylalanine ammonia lyase having an amino acid sequence substantially similar to that of sequence number 117 are used.
[0098] The present invention also provides a method for generating transgenic bioluminescent cells or organisms consisting of plant, animal, bacterial, or fungal cells or organisms.
[0099] In a preferred embodiment, a method for generating a transgenic bioluminescent cell or organism comprises introducing at least one nucleic acid of the present invention into a cell or organism together with a nucleic acid encoding a luciferase capable of oxidizing fungal luciferin with photoemission. The nucleic acid is introduced into the cell or organism in a form that enables its expression and the production of functional protein products. For example, the nucleic acid may be contained in an expression cassette. The nucleic acid can be generated intracellularly as part of an extrachromosomal element or in a form integrated into the cell's genome by insertion of an expression cassette into the cell.
[0100] In a preferred embodiment, a method for generating a transgenic bioluminescent cell or organism comprises introducing a nucleic acid encoding the hispidine hydroxylase of the present invention and a nucleic acid encoding a luciferase capable of oxidizing fungal luciferin with photoemission into a cell or organism. As a result, the cell or organism acquires bioluminescent ability in the presence of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one and fungal preluciferin having the following structural formula, where R is aryl or heteroaryl.
[0101] [ka]
[0102] In some embodiments, instead of nucleic acids encoding hispidin synthase and luciferase, nucleic acids encoding a hispidin hydroxylase and luciferase fusion protein are introduced into the cells.
[0103] In some embodiments, a method for generating a transgenic bioluminescent cell or organism also includes introducing a nucleic acid encoding the hispidin synthase of the present invention into a cell or organism. The cell or organism acquires bioluminescent ability in the presence of a fungal preluciferin precursor in the form of a 3-arylacrylic acid having the following structural formula, where R is aryl or heteroaryl.
[0104] [ka]
[0105] In some embodiments, nucleic acids encoding PKS are introduced into the cells instead of nucleic acids encoding hispidin synthase.
[0106] In some embodiments, a method for generating transgenic bioluminescent cells or organisms also includes introducing the nucleic acid encoding the caffeylpyruvate hydrolase of the present invention into cells or organisms to enhance the intensity of bioluminescence.
[0107] In some embodiments, a method for generating transgenic bioluminescent cells or organisms also includes introducing a nucleic acid encoding 4'-phosphopantotheinyltransferase into cells or organisms.
[0108] In some embodiments, a method for generating transgenic bioluminescent cells or organisms also includes introducing nucleic acids encoding coumarate-CoA ligase into cells or organisms.
[0109] In some embodiments, a method for generating transgenic bioluminescent cells or organisms also includes introducing nucleic acids encoding 3-arylacrylate biosynthesis enzymes into cells or organisms.
[0110] The present invention also provides transgenic bioluminescent cells and organisms obtained by the method described above, comprising one or more nucleic acids of the present invention as part of an extrachromosomal element or in a form incorporated into the genome of a cell.
[0111] In some embodiments, the transgenic bioluminescent cells and organisms of the present invention are capable of autonomous bioluminescence without the exogenous addition of luciferin, preluciferin, and preluciferin precursors.
[0112] The present invention also provides combinations of proteins and nucleic acids, as well as products and kits comprising the proteins and nucleic acids of the present invention. For example, the nucleic acid combinations are provided for generating autonomously luminescent cells, cell lines, or transgenic organisms; and for assaying promoter activity to label cells.
[0113] In some embodiments, a kit for producing fungal luciferin and / or fungal preluciferin is provided to consist of the hispidin hydroxylase and / or hispidin synthase and / or PKS, or nucleic acids encoding them.
[0114] In some embodiments, a kit is provided for generating bioluminescent cells or bioluminescent transgenic organisms comprising a nucleic acid encoding hispidin hydroxylase and a nucleic acid encoding luciferase capable of oxidizing fungal luciferin with photoemission. The kit also includes caffeylpyruvate hydrolase. The kit may also include a nucleic acid encoding hispidin synthase or PKS. The kit may also include a nucleic acid encoding 4'-phosphopantotheinyltransferase and / or a nucleic acid encoding coumarate-CoA ligase and / or a nucleic acid encoding 3-arylacrylate biosynthesis enzyme. The kit may also include additional components such as buffers, antibodies, fungal luciferin, fungal preluciferin, and a precursor of fungal preluciferin. The kit may also include a kit application guide. In some embodiments, the nucleic acids are provided in an expression cassette or vector for introduction into cells or organisms.
[0115] In preferred embodiments, the cells or transgenic organisms of the present invention can produce fungal luciferin from a precursor. In some embodiments, the cells and transgenic organisms of the present invention can bioluminesce in the presence of a fungal luciferin precursor. In some embodiments, the cells or transgenic organisms of the present invention are capable of autonomous bioluminescence.
[0116] Preferred embodiments of the methods and applications disclosed above include the following: (E)-6-(3,4-dihydroxystyryl)-4-hydroxy-2H-pyran-2-one (hispidine), (E)-4-dihydroxy-6-styryl-2H-pyran-2-one, (E)-4-hydroxy-6-(4-hydroxystyryl)-2H-pyran-2-one (bisnoriangonin), (E)-4-hydroxy-6-(2-hydroxystyryl)-2H-pyran-2-one, (E)-4-hydroxy-6-(2,4-dihydroxystyryl)-2H-pyran-2-one, (E)-4-hydroxy-6-(4-hydroxy-3,5-dimethoxystyryl)-2H-pyran-2-one, (E)-4-hydroxy-6-(4-hydroxy-3-methoxystyryl)-2H-pyran-2-one, (E)-4-hydroxy-6-(2-(6-hydroxynaphthalene-2-yl)vinyl)-2H-pyran-2-one, (E)-6-(4-aminostyryl)-4-hydroxy-2H-pyran-2-one, (E)-6-(4-(diethylamino)styryl)-4-hydroxy-2H-pyran-2-one, (E)-6-(2-(1H-indole-3-yl)vinyl)-4-hydroxy-2H-pyran-2-one, (E)-4-hydroxy-6-(2,3,6,7-tetrahydro-1H,5H-pyrido[3,2,1-ij]quinoline-9-yl)vinyl)-2H-pyran-2-one Preluciferin having the chemical formula 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one is used, selected from the following.
[0117] In preferred embodiments, 3-arylacrylic acid selected from the group consisting of caffeic acid, cinnamic acid, paracoumaric acid, coumaric acid, umbellic acid, sinapic acid, and ferulic acid is suitable for the purposes of the present invention.
[0118] In preferred embodiments, 3-hydroxyhispidin is used as luciferin, hispidin as preluciferin, and caffeic acid as a precursor of preluciferin.
[0119] One of the objectives of the present invention is to provide an effective method for generating visible-emitting autonomous bioluminescent systems, including eukaryotic non-luminescent cells and organism-based systems.
[0120] Another object of the present invention is to provide a novel and effective method for synthesizing hispidin or its functional analogues.
[0121] Another object of the present invention is to provide a novel and effective method for synthesizing fungal luciferin or its functional analogues.
[0122] Another object of the present invention is to provide a cell or organism that emits light autonomously.
[0123] The objective of this invention is achieved by identifying the steps of luciferin conversion in bioluminescent fungi and identifying the amino acid and nucleotide sequences of proteins involved in luciferin biosynthesis. The function of all proteins has been demonstrated for the first time. [Brief explanation of the drawing]
[0124] [Figure 1] The arrangement of multiple amino acid sequences of hispidin hydroxylase is shown. The FAD / NAD(P) ligation domain is underlined. The consensus sequence is shown below the arrangement. [Figure 2] The arrangement of multiple amino acid sequences of hispidin synthase is shown. The consensus sequence is shown below the arrangement. [Figure 3] The arrangement of multiple amino acid sequences of caffeylpyruvate hydrolase is shown. The consensus sequence is shown below the arrangement. [Figure 4] The graphs show the luminescence intensity of hispidin hydroxylase-expressing Pichia pastoris cells and luciferase (A), or luciferase alone (B), and the luminescence intensity of wild-type yeast (C) when colonies are sprayed with 3-hydroxyhispidin (luciferin, left plot) or hispidin (preluciferin, right plot). [Figure 5] The luminescence intensity of HEK293NT cells expressing both hispidin hydroxylase and luciferase is shown in comparison to the luminescence intensity of HEK293NT cells expressing only luciferase when hispidin is added. [Figure 6](1) Hispidin hydroxylase and luciferase genes when hispidin is added separately; (2) Hispidin hydroxylase and luciferase chimeric protein genes when hispidin is added; (3) Emission curves of HEK293T cells expressing hispidin hydroxylase and luciferase chimeric protein genes when 3-hydroxyhispidin is added. [Figure 7] The autonomous bioluminescence ability of transfected Pichia pastris cells is shown compared to that of wild-type cells. Left: Cells in a petri dish under daylight; Right: Cells in the dark. [Figure 8] This shows the luminescence of transfected Pichia pastris cell culture medium in the dark. [Figure 9] This image shows Nicotiana benthamiana, a transgenic plant that exhibits autonomous bioluminescence. The left image was taken in ambient light, while the right image was taken in darkness. [Modes for carrying out the invention]
[0125] definition Various terms related to the subject matter of the present invention are used above, as well as in the specification and claims. In the specification of the present invention, the terms "consisting of" and "comprising of" are to be interpreted as "consisting of, but not limited to," and are not intended to be interpreted as "consisting of only."
[0126] For the purposes of this invention, the terms "luminescence" and "bioluminescence" are mutually interchangeable and refer to the phenomenon of light emission during a chemical reaction catalyzed by the enzyme luciferase.
[0127] Terms related to protein activity, such as "reactive" and "reaction-promoting," mean that the protein is an enzyme that catalyzes the indicated reaction.
[0128] For the purposes of this invention, the term "luciferase" means a protein that has the ability to catalyze the oxidation of a compound (luciferin) with molecular oxygen, causing the oxidation reaction to be accompanied by photoluminescence (luminescence or bioluminescence) and the formation of oxidized luciferin.
[0129] For the purposes of this invention, the term "fungal luciferin" means a compound selected from the group of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one having the following structural formula, where R is aryl or heteroaryl.
[0130] [ka]
[0131] Fungal luciferin is oxidized by a group of luciferases referred to below as "luciferases capable of oxidizing fungal luciferin with photoemission," etc. Such luciferases have been found in bioluminescent fungi and are described, for example, in Russian Patent No. 2017102986 / 10 (005203), filed on January 30, 2017. The amino acid sequences of luciferases useful in the methods and combinations of the present invention are substantially similar to or identical to amino acid sequences selected from the following sequence number group: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98. In many embodiments of the present invention, luciferases useful for the purposes of the present invention are characterized by an amino acid sequence that is at least 40% identical, for example, at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, to an amino acid sequence selected from the following sequence number group: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98. In many cases, the luciferase is characterized by an amino acid sequence that has at least 90% identity (for example, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity) to an amino acid sequence selected from the following sequence number group: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98.
[0132] The oxidation of fungal luciferin produces "fungal oxyluciferin," a product having the chemical formula 6-aryl-2-hydroxy-4-oxohexa-2,5-dienonic acid and the following structural formula.
[0133] [ka]
[0134] The term “fungal preluciferin” or simply “preluciferin” is used herein to refer to a compound selected from the group of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one having the following structural formula, where R is aryl or heteroaryl. Preluciferin is converted to fungal luciferin by a chemical reaction catalyzed by the enzyme of the present invention.
[0135] [ka]
[0136] The term “preluciferin precursor” is used herein to refer to a group of compounds belonging to the group of 3-arylacrylic acids having the following structural formula, where R is aryl or heteroaryl. Preluciferin is formed from 3-arylacrylic acid in the course of a chemical reaction catalyzed by the enzyme of the present invention.
[0137] Examples of fungal luciferins are shown in Table 1. Examples of fungal luciferin-related preluciferins, oxyluciferins, and preluciferins are shown in Table 2.
[0138] [Table 4]
[0139] [Table 5]
[0140] [Table 6]
[0141] [Table 7]
[0142] [Table 8]
[0143] [Table 9]
[0144] The term "aryl" or "aryl substituent" refers to an aromatic radical in a single or fusion carbocyclic ring system containing 5 to 14 ring members. In preferred embodiments, the ring system contains 6 to 10 ring members. One or more hydrogen atoms may be substituted with substituents selected from acyl groups, acylamino groups, acyloxy groups, alkenyl groups, alkoxy groups, alkyl groups, alkynyl groups, amino groups, aryl groups, aryloxy groups, azide groups, carbamoyl groups, carboalkoxy groups, carboxyl groups, carboxyamide groups, carboxyamino groups, cyano groups, disubstituted amino groups, formyl groups, guanidino groups, halogen groups, heteroaryl groups, heterocyclyl groups, hydroxyl groups, iminoamino groups, monosubstituted amino groups, nitro groups, oxo groups, phosphoamino groups, sulfinyl groups, sulfonamino groups, sulfonyl groups, thio groups, thioacylamino groups, thioureido groups, or ureido groups. Examples of aryl groups include, but are not limited to, phenyl, naphthyl, biphenyl, and terphenyl. Furthermore, as used herein, the term "aryl" refers to a group in which an aromatic ring is linked to one or more non-aromatic rings.
[0145] The terms "heterocyclic aromatic substituent," "heteroaryl substituent," or "heteroaryl" refer to an aromatic radical containing 1 to 4 heteroatoms or heterogroups selected from O, N, S, or SO in a single or fusion heterocyclic system containing 5 to 15 ring members. In preferred embodiments, the heteroaryl ring system contains 6 to 10 ring members. Furthermore, one or more hydrogen atoms can be substituted with substituents selected from acyl groups, acylamino groups, acyloxy groups, alkenyl groups, alkoxy groups, alkyl groups, alkynyl groups, amino groups, aryl groups, aryloxy groups, carbamoyl groups, carboalkoxy groups, carboxyl groups, carboamide groups, carboxyamino groups, cyano groups, disubstituted amino groups, formyl groups, guanidino groups, halogen groups, heteroaryl groups, heterocyclyl groups, hydroxyl groups, iminoamino groups, monosubstituted amino groups, nitro groups, oxo groups, phosphoamino groups, sulfinyl groups, sulfonamino groups, sulfonyl groups, thio groups, thioacylamino groups, thioureido groups, or ureido groups. Examples of heteroaryl groups include, but are not limited to, pyridinyl groups, thiazolyl groups, thiadiazolyl groups, isoquinolinyl groups, pyrazolyl groups, oxazolyl groups, oxadiazoyl groups, triazolyl groups, and pyrrolyl groups. Furthermore, as used herein, the term "heteroaryl" refers to a group in which a heteroaromatic ring is linked to one or more nonaromatic rings.
[0146] Compound names in this invention will be used in accordance with the International Union of International Purposes (IUPAC) nomenclature. Common names will also be provided (if any).
[0147] The terms "luciferin biosynthesis enzyme" or "enzyme involved in the periodic turnover of luciferin conversion" are used to refer to enzymes that catalyze the conversion from preluciferin precursor to preluciferin, and / or from preluciferin to fungal luciferin, and / or from oxyluciferin to preluciferin precursor in vitro and / or in vivo systems. The term "fungal luciferin biosynthesis enzyme" does not refer to luciferase unless otherwise specified.
[0148] The term "hispidin hydroxylase" is used herein to describe an enzyme that catalyzes the reaction that converts preluciferin to fungal luciferin, for example, the reaction that synthesizes 3-hydroxyhispidin from hispidin.
[0149] The term "hispidin synthase" is used herein to describe an enzyme that can catalyze the synthesis of fungal preluciferin from preluciferin precursors, for example, the synthesis of hispidin from caffeic acid.
[0150] The term "PKS" is used herein to describe enzymes belonging to a group of type III polyketide synthases capable of catalyzing the synthesis of hispidin from caffeyl-CoA.
[0151] The term "caffeylpyruvate hydrolase" is used herein to describe an enzyme that can catalyze the breakdown of fungal oxyluciferin into simpler compounds in order to form a precursor of preluciferin. For example, it can catalyze the conversion of caffeylpyruvate to caffeic acid.
[0152] The term "functional analogue" is used herein to describe proteins that perform the same function, and / or compounds or proteins that can be used for the same purpose. For example, all fungal luciferins listed in Table 1 are functionally analogues of one another.
[0153] The term "ATP" refers to adenosine triphosphate, the primary energy carrier within cells, which has the following structural formula.
[0154] [ka]
[0155] In this specification, the term "NAD(P)H" refers to the reduced nicotinamide adenine dinucleotide phosphate (NADPH) moiety or the nicotinamide adenine dinucleotide (NADH) moiety. The term "NAD(P)" refers to the oxidized form of nicotinamide adenine dinucleotide phosphate (NADP) or nicotinamide adenine dinucleotide (NAD). Nicotinamide adenine dinucleotide is represented by the following formula.
[0156] [ka]
[0157] Nicotinamide adenine dinucleotide phosphate is represented by the following formula.
[0158] [ka]
[0159] Nicotinamide adenine dinucleotide and nicotinamide adenine dinucleotide phosphate are dinucleotides constructed from nicotinamide and adenine linked by a chain consisting of two D-ribose residues and two phosphate residues. NADP differs from NAD due to the presence of a phosphate residue attached to the hydroxyl group of the D-ribose residue. Both compounds are widely present in nature, participate in many redox reactions, and function as carriers of electrons and hydrogen received from oxidized substances. The reduced form transfers the received electrons and hydrogen to other substances.
[0160] The term "coenzyme A" or "CoA" refers to a coenzyme well known from prior art, which is involved in the oxidation or synthesis of fatty acids, the biosynthesis of fats, and the oxidative transformation of carbohydrate degradation products, and has the following structural formula.
[0161] [ka]
[0162] The term "malonyl-CoA" refers to a derivative of coenzyme A containing a malonic acid residue, formed during fatty acid synthesis, and is represented by the following formula.
[0163] [ka]
[0164] The term "coumaroyl-CoA" refers to the thioester and coumaric acid of coenzyme A, and is represented by the following formula.
[0165] [ka]
[0166] The term "caffeyl-CoA" refers to the thioester and caffeic acid of coenzyme A, and is represented by the following formula.
[0167] [ka]
[0168] As used herein, the terms “mutant” or “derivative” refer to the proteins disclosed herein, in which one or more amino acids are added to, / or substituted in, / or deleted from, and / or inserted into the N-terminal and / or C-terminal sequences and / or native amino acid sequences of the proteins of the present invention. As used herein, the term “mutant” refers to the nucleic acid region encoding the mutant protein. Furthermore, as used herein, the term “mutant” refers to any variant that is shorter or longer than the proteins or nucleic acids disclosed herein.
[0169] The term "homology" is used to describe the relationship between nucleotide sequences and amino acid sequences, which is determined by the degree of identity and / or similarity between the sequences being compared.
[0170] As used herein, an amino acid or nucleotide sequence is "nearly identical" or "nearly the same" as a reference sequence if it has at least 40% identity with a sequence selected within the reference domain. Thus, nearly similar sequences include, for example, sequences having at least 40% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 62% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity). Two sequences that are identical to each other are also nearly similar. For the purposes of this invention, the length of the comparison sequence must be at least 100 amino acids, preferably at least 200 amino acids, for example, 300 amino acids. In particular, it is possible to compare the full-length amino acid sequences of proteins. In the case of nucleic acids, the length of the comparison sequence must be at least 300 nucleotides; preferably at least 600 nucleotides, for example, 900 nucleotides.
[0171] An example of an algorithm suitable for determining sequence identity and sequence similarity is the BLAST algorithm described by Altschul et al., Journal of Molecular Biology, 1990, No. 215; pp. 403-410. Software for performing BLAST analysis is available from the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). This algorithm first searches for high-scoring segment pairs (HSPs) by identifying whether a short string of length W in the test sequence, when sequenced with strings of the same length in the database sequence, perfectly matches or satisfies a certain positive threshold score T. T represents the neighbor string score threshold (Altschul et al., 1990). These initial neighbor string hits serve as seed values to initiate the search for longer HSPs containing that string. These string hits are then expanded bidirectionally along each sequence as long as the cumulative marshalling score can be increased. For nucleotide sequences, the cumulative score is calculated using parameters M (reward score set for matching residue pairs; always > 0) and N (penalty score set for mismatched residues; always < 0). A scoring matrix is used to calculate the cumulative value for amino acid sequences. Expansion of string hits in each direction stops when the cumulative marshalling score drops by X from its maximum achieved value; or when the end of either sequence is reached. The parameters W, T, and X of the BLAST algorithm determine the sensitivity and speed of marshalling. In the BLASTN program (for nucleotide sequences), the default string length (W) is 11, the expected value (E) is 10, the drop-off (cut-off) value is 100, M=5 and N=-4, and comparisons are performed on both strands.The BLASTP program (for amino acid sequences) has a default string length (W) of 3, an expected value (E) of 10, and uses the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proceedings of the National Academy of Sciences, (USA), 1989, No. 89: p. 10915).
[0172] In addition to calculating the sequence identity rate, the BLAST algorithm also performs a statistical similarity analysis between two sequences (see, for example, Karlin and Altschul, Proceedings of the National Academy of Sciences, (USA), 1993, No. 90: pp. 5873-5787). One of the parameters provided by the BLAST algorithm for determining similarity is the minimum cumulative probability (P(N)), which indicates the probability of random agreement between two nucleotide or amino acid sequences. For example, if the minimum cumulative probability when comparing a test nucleic acid sequence with a reference nucleic acid sequence is less than 0.1, more preferably less than 0.01, and most preferably less than 0.001, the test nucleic acid sequence is considered similar to the reference nucleic acid sequence.
[0173] The term "consensus sequence" refers to a typical amino acid sequence used as a reference for comparing all variants of a particular protein or sequence. Consensus sequences and methods for determining them are well known to those skilled in the art. For example, a consensus sequence can be determined by identifying the amino acid that most frequently occurs at a given position within a given set of related sequences, through comparisons of multiple known homologous proteins.
[0174] The term "conserved sequence" is used to specify a nucleotide sequence in a nucleic acid or an amino acid sequence in a polypeptide chain that remains completely or nearly unchanged throughout the evolution of different organisms. Conversely, a "non-conserved sequence" refers to a sequence that differs significantly between the organisms being compared.
[0175] The term "amino acid insertion segment" refers to one or more amino acids within a polypeptide chain between protein fragments (protein domains, linkers, consensus sequences) under consideration. Those skilled in the art should understand that it is possible to operationally link these amino acid insertion segments and fragments to form a single polypeptide chain.
[0176] The domain structure of a protein can be determined using any suitable software known in the art. For example, the Simple Modular Architecture Research Tool (SMART) software, available online at http: / / smart.embl-heidelberg.de, can be used for this purpose [Schultz et al., Proceedings of the National Academy of Sciences (PNAS), 1998; No. 95: pp. 5857-5864; Letunic I, Doerks T, and Bork P, Nucleic Acids Research, 2014; doi:10.1093 / nar / gku949].
[0177] In the description of fusion proteins, terms such as "operationally linked" refer to polypeptide sequences resulting from the physical and functional relationships between them. In the most preferred embodiment, the function of the polypeptide component of the chimeric molecule is unchanged compared to the functional properties of the isolated polypeptide component. For example, the hispidin hydroxylase of the present invention can be operationally linked to a target fusion partner, such as luciferase. In this case, the target polypeptide retains its original biological activity, such as the ability to oxidize luciferin with photoemission, while the fusion protein retains the properties of hispidin hydroxylase. In some embodiments of the present invention, the activity of the fusion partner may be reduced compared to the activity of the isolated protein. Such fusion proteins also find applications within the scope of the present invention.
[0178] In the description of nucleic acids, terms such as "operationally linked" mean that nucleic acids are covalently linked in such a way that no malfunction or stop signal occurs in the reading frame at their junction. As will be apparent to those skilled in the art, a nucleotide sequence encoding a fusion protein with "operationally linked" components (proteins, polypeptides, linker sequences, amino acid insertion segments, protein domains, etc.) consists of fragments encoding the said components, and these fragments are covalently linked in such a way that a full-length fusion protein is generated during the transcription and translation of the nucleotide sequence.
[0179] In describing the relationship between regulatory coding sequences (promoters, enhancers, transcriptional terminators) and nucleic acids, the term "operationally linked" means that the sequences are positioned and linked in such a way that the regulatory sequence affects the expression level of the coding nucleic acid or nucleic acid sequence.
[0180] In the context of this invention, "linking" of nucleic acids means linking two or more nucleic acids together using any means known in the art. As a non-limiting example, multiple nucleic acids can be linked together using DNA ligase or polymerase chain reaction (PCR) during annealing. Multiple nucleic acids can also be linked by nucleic acid chemosynthesis using a single sequence from two or more distinct nucleic acids.
[0181] The term "regulatory element" or "regulatory sequence" refers to sequences involved in regulating the expression of coding nucleic acids. Regulatory elements include promoters, terminal signals, and other sequences that affect nucleic acid expression. They also typically consist of sequences necessary for the proper translation of nucleotide sequences.
[0182] The term "promoter" is used to describe the untranslated and untranscribed DNA sequences upstream of the coding region, which includes the RNA polymerase binding site and the transcription initiation DNA binding site. The promoter region may also consist of other gene expression regulatory elements.
[0183] As used herein, the term “functional” refers to a nucleotide or amino acid sequence capable of performing a role in a particular test or task. When used to describe luciferase, the term “functional” means that the protein has the ability to produce a luciferin oxidation reaction accompanied by luminescence. When used to describe hispidin hydroxylase, the same term “functional” means that the protein has the ability to catalyze a reaction that converts at least one of the preluciferins shown in Table 2 to the corresponding luciferin. When used to describe hispidin synthase, the same term “functional” means that the protein has the ability to catalyze a reaction that converts at least one of the preluciferin precursors to preluciferin, for example, the reaction that converts caffeylpyruvate to hispidin. When used to describe caffeyylpyruvate hydrolase, the same term “functionality” means that the protein has the ability to catalyze a reaction that converts at least one of the oxyluciferins to a preluciferin precursor (for example, the conversion of caffeyylpyruvate to caffeylpyruvate).
[0184] As used herein, the term "enzyme properties" refers to the ability of a protein to catalyze a given chemical reaction.
[0185] As used herein, the term "biochemical properties" refers to protein folding, and also includes maturation rate, half-life, catalytic efficiency, pH and temperature stability, and other similar properties.
[0186] As used herein, the term "spectral properties" refers to spectra, quantum yields, emission intensities, and other similar properties.
[0187] A reference to a polypeptide-coding nucleotide sequence means that the polypeptide is produced during mRNA transcription and translation according to this nucleotide sequence. In this case, it is possible to refer to both the coding strand, which is identical to the mRNA and commonly used in sequence listings, and the complementary strand, which is used as a transcription template. As will be apparent to those skilled in the art, this term also encompasses any degenerate nucleotide sequence that codes for the same amino acid sequence. A polypeptide-coding nucleotide sequence consists of a sequence containing introns.
[0188] The term “expression cassette” or “expression cassette” is used herein to mean a nucleic acid sequence capable of controlling the expression of a specific nucleotide sequence in a suitable host cell. In principle, an “expression cassette” comprises a heterogeneous nucleic acid encoding a protein or a functional fragment thereof, operationally linked to a promoter and terminal signaling pathway. Typically, an “expression cassette” also includes sequences necessary for the proper translation of a significant nucleotide sequence. Expression cassettes may be naturally occurring (including in host cells), but are produced in recombinant forms useful for the expression of heterogeneous nucleic acids. However, in many cases, an “expression cassette” is heterogeneous with respect to the host; that is, the specific nucleic acid sequence of this expression cassette does not naturally exist in the host cell and must be introduced into the host cell or a host cell precursor by transformation. The expression of this nucleotide sequence is controllable by a constituent or inducible promoter that initiates transcription only when the host cell is vulnerable to a specific external stimulus. In the case of multicellular organisms, promoters may also be specific to a particular tissue, organ, or developmental stage.
[0189] "Exogenous" or "heterogeneous" nucleic acids refer to nucleic acids that are never present in wild-type host cells.
[0190] The term "endogenous" refers to a native protein or nucleic acid that is located in its natural position within the genome of an organism.
[0191] As used herein, the term “specifically hybridize” refers to the association of two single-stranded nucleic acid molecules or sufficiently complementary (sometimes the term “nearly complementary” is used) sequences that enable hybridization under certain conditions commonly used in the art.
[0192] "Isolated" nucleic acid sites or isolated proteins are nucleic acid sites or proteins that have been separated from the natural environment by human activity and are therefore not products of nature. Isolated nucleic acid molecules or isolated proteins can be generated in a purified form or in non-natural environments such as (but not limited to) recombinant prokaryotic cells, plant cells, animal cells, non-bioluminescent fungal cells, transgenic organisms (fungi, plants, animals), etc.
[0193] "Transformation" is the process of introducing a different nucleic acid into a host cell or organism. In particular, "transformation" means the stable integration of a DNA region into the genome of a target organism.
[0194] The terms "transformed / transgenic / recombinant" refer to host organisms such as bacteria, plants, fungi, or animals that have been modified by introducing a heterologous nucleic acid site. This nucleic acid site may be stably integrated into the host genome or may exist as an extrachromosomal location. Such extrachromosomal locations may be capable of self-replication. It should be understood that transgenic or stably transformed cells, tissues, or organisms include both end products of the transformation process, but also transgenic offspring. The terms "untransformed," "non-transgenic," "non-recombinant," or "wild-type" refer to natural host organisms or host cells that do not contain heterologous nucleic acid sites, such as bacteria or plants.
[0195] The term "autonomously bioluminescent" or "autonomously bioluminescent" refers to a transgenic organism or host cell capable of bioluminescence without the exogenous addition of luciferin, preluciferin, or a preluciferin precursor.
[0196] The term "4'-phosphopantotheinyltransferase" is used herein to refer to the enzyme that transfers 4-phosphopantotheinyl from coenzyme A to serine within the acyl transfer domain of a polyketide synthase. 4'-phosphopantotheinyltransferase is naturally expressed in many plants and fungi and is well known in the art [Gao Menghao et al., Microbial Cell Factories, 2013, No. 12: p. 77]. It will be obvious to those skilled in the art that any functional variant of 4'-phosphopantotheinyltransferase can be used for the purposes of this invention. For example, the NpgA4'-phosphopantotheinyltransferase of Aspergillus nidulans (SEQ ID NOs. 104, 105) described in [Gao Menghao et al., Microbial Cell Factories, 2013, No. 12: p. 77], or its homologs or mutants, i.e., proteins having amino acid sequences that are nearly similar to or identical to the sequence containing SEQ ID NO. 105. Another example is a 4'-phosphopantotheinyltransferase having at least 40% identity with the sequence characterized by sequence number 105, for example, at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 62% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, or at least 85% identity, or at least 90% identity (for example, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0197] Nucleotides are specified according to their base using the following standard abbreviations: adenine (A), cytosine (C), thymine (T), and guanine (G). Similarly, amino acids are designated by the following standard abbreviations: alanine (Ala;A), arginine (Arg;R), asparagine (Asn;N), aspartic acid (Asp;D), cysteine (Cys;C), glutamine (Gln;Q), glutamic acid (Glu;E), glycine (Gly;G), histidine (His;H), isoleucine (He;1), leucine (Leu;L), lysine (Lys;K), methionine (Met;M), phenylalanine (Phe;F), proline (Pro;P), serine (Ser;S), threonine (Thr;T), tryptophan (Trp;W), tyrosine (Tyr;Y), and valine (Val;V).
[0198] The present invention aims to identify novel fungal luciferin biosynthetic enzymes, nucleic acids capable of encoding these enzymes, and proteins capable of catalyzing specific steps in fungal luciferin biosynthesis. The present invention also provides applications of nucleic acids for generating the enzymes in cells or organisms. Methods for preparing compounds consistent with fungal luciferin and preluciferin in vitro or in vivo are also provided. Vectors comprising the nucleic acids described in the present invention are also provided. The present invention also provides expression cassettes comprising the nucleic acids of the present invention and regulatory elements necessary for nucleic acid expression in selected host cells. Furthermore, cells, stable cell lines, and transgenic organisms (e.g., plants, animals, fungi, or microorganisms) containing the nucleic acids, vectors, or expression cassettes of the present invention are also provided. The present invention also provides combinations of nucleic acids for obtaining autoluminescent cells, cell lines, or transgenic organisms. In preferred embodiments, cells or transgenic organisms are capable of generating fungal luciferin from precursors. In some embodiments, cells or transgenic organisms are capable of generating fungal preluciferin from precursors. In some embodiments, cells or transgenic organisms can bioluminesce in the presence of fungal luciferin precursors. In some embodiments, cells or transgenic organisms can bioluminesce autonomously. Combinations of proteins for generating luciferin or its precursors from simpler compounds are also provided. The present invention also provides kits comprising the nucleic acids, vectors, or expression cassettes of the present invention for generating luminescent cells, cell lines, or transgenic organisms.
[0199] protein As described above, the present invention provides a protein that is involved as an enzyme in fungal luciferin biosynthesis (cyclic system of transformation).
[0200] The proteins of the present invention can be obtained from natural sources or by recombinant technology. For example, wild-type proteins can be isolated from bioluminescent fungi, such as basidiomycetes, mainly the basidiomycete class, particularly fungi of the Agaricales order. For example, wild-type proteins can be isolated from fungi such as Neonothopanus nambi, Armillaria fuscipes, Armillaria mellea, Guyanagaster necrorhiza, Mycena citricolor, Neonothopanus gardneri, Omphalotus olearius, Panellus stipticus, Armillaria gallica, Armillaria ostoyae, and Mycena chlorophos. The proteins of the present invention can also be obtained by expressing recombinant nucleic acids and coding protein sequences in their respective hosts or cell-free expression systems, as described in the "Nucleic Acids" section. In some embodiments, the protein is used in a host cell encoding the protein by introducing an expressible nucleic acid.
[0201] In preferred embodiments, the claimed protein folds rapidly after expression in host cells. It is understood that “rapid folding” means that the protein reaches a tertiary structure that ensures its enzymatic properties within a short period of time. In these embodiments, the protein folds within a period of time generally not exceeding about 3 days, usually not exceeding about 2 days, and typically not exceeding about 12–24 hours.
[0202] In some embodiments, the protein is used in an isolated form. Any of the general techniques described in the Guide to Protein Purification (edited by Deuthser, Academic Press, 1990) can be used for protein purification. For example, the lysate can be prepared from the initial source and purified using HPLC, substitution chromatography, gel electrophoresis, affinity chromatography, etc.
[0203] When the protein of the present invention is in an isolated form, it means that the protein is substantially free from other proteins or other naturally occurring biological molecules, such as oligosaccharides, nucleic acids and their fragments. In this context, the term "substantially free" means that less than 70%, usually less than 60%, and typically less than 50%, of the composition comprising the isolated protein are other naturally occurring biological molecules. In some embodiments, the protein is in a substantially purified form. The term "substantially purified form" means a purity equal to at least 95%, usually at least 97%, and typically at least 99%.
[0204] The protein of the present invention maintains its activity at temperatures below 50°C, typically up to 45°C, meaning it maintains its activity at temperatures between 20 and 42°C and can be used in heterologous expression systems in vitro and in vivo.
[0205] The requested protein has pN stability in the range of 4–10, typically in the range of 6.5–9.5. The optimal pN stability for the requested protein is in the range of 6.8–8.5, for example, in the range of 7.3–8.3.
[0206] The claimed protein is active under physiological conditions. The term "physiological conditions" in the present invention is intended to refer to a temperature within the range of 20 to 42 °C, a pH within the range of 6.8 to 8.5, physiological saline, and a medium having an osmotic pressure of 300 to 400 mOsm / l. In particular, the term "physiological conditions" includes intracellular media, cell-free preparations, and liquids extracted from living organisms such as plasma. "Physiological conditions" can be artificially formed. For example, a reaction mixture ensuring "physiological conditions" can be formed by combining known compounds. Such methods of forming media are well known from the prior art. Non-limiting examples include the following.
[0207] 1) Ringer's solution, which is isotonic with mammalian plasma Ringer's solution consists of 6.5 g of NaCl, 0.42 g of KCl, and 0.25 g of CaCl2 dissolved in 1 liter of double-distilled water. When preparing the solution, the salts are added sequentially. Each subsequent salt should not be added until the previously added salt has dissolved. To prevent calcium carbonate precipitation, it is recommended to pass carbon dioxide gas through the sodium bicarbonate solution. The solution is prepared with fresh distilled water.
[0208] 2) Basen solution Basen solution is a mixture of EDTA and inorganic salts, dissolved in distilled water or water for injection, and sterilized by membrane filtration using a filter with a final pore size of 0.22 μm. 1 liter of Basen solution consists of 8.0 g of NaCl, 0.2 g of KCl, 1.45 g of disodium phosphate dodecahydrate, 0.2 g of potassium dihydrogen phosphate, 0.2 g of palkelate, and double-distilled water for making up to 1 liter. The buffering capacity of Basen solution needs to be 1.4 ml. The chloride ion content is 4.4 to 5.4 g / l, and the amount of EDTA is at least 0.6 mmol / l.
[0209] 3) Phosphate-buffered saline (PBS, Na-phosphate buffer) Sodium phosphate buffer consists of 137 mM NaCl, 10 mM Na2HPO4, and 1.76 mM KH2PO4. The buffer can also contain KCl at concentrations up to 2.7 mM. To prepare 1 liter of normal-strength sodium phosphate buffer, use the following: 8.00 g NaCl, 1.44 g Na2HPO4, 0.24 g KH2PO4, and 0.20 g KCl (optional). Dissolve in 800 ml of distilled water. Adjust to the desired pH using hydrochloric acid or sodium hydroxide. Then add distilled water to a total volume of 1 liter.
[0210] The specific proteins under consideration are enzymes involved in the biosynthesis of luciferin in cyclic fungi, their variants, homologs, and derivatives. Each of these specific types of polypeptide structures will be analyzed individually and in further detail.
[0211] Hispidin hydroxylase The hispidin hydroxylase of the present invention is a protein capable of catalyzing the synthesis of luciferin from preluciferin. In other words, hispidin hydroxylase is an enzyme that catalyzes the reaction in which 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one is converted to 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one. The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0212] [ka]
[0213] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0214] [ka]
[0215] The reaction is carried out under in vitro and in vivo physiological conditions in the presence of at least one molecule of NAD(P)H and at least one molecule of molecular oxygen (O2) per molecule of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one.
[0216] [ka]
[0217] The hispidin-hydroxylases in question include proteins derived from the bioluminescent fungi *Lycoperdon perlatum*, *Armillaria mellea*, *Armillaria mellea*, *Guyanagastar necrohiza*, *Misena citricolor*, *Neonotepanus gardneri*, *Omphalotus olearius*, *Armillaria mellea*, *Armillaria gracilis*, and *Mycena chlorophos*. Their amino acid sequences are shown in SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, as well as their functional variants, homologs, and derivatives.
[0218] In preferred embodiments, the hispidin hydroxylase of the present invention is characterized by the presence of a FAD / NAD(P) binding domain IPR002938 - code in the InterPro public database, available on the Internet at the website (http: / / www.ebi.ac.uk / interpro). The domain is involved in the binding of flavin adenine dinucleotide (FAD) to nicotinamide adenine dinucleotide (NAD) in multiple enzymes, the addition of hydroxyl groups to substrates, and multiple organisms found in metabolic pathways. The hispidin hydroxylase of the present invention consists of the domain having a length of 350–385 amino acids, typically 360–380 amino acids, e.g., 364–377 amino acids, and non-conserved N-terminal and C-terminal amino acid sequences into which loxP is introduced with a low degree of identity with each other. The location of the FAD / NAD binding domain in the claimed hispidin hydroxylase is illustrated in Figure 1 in a multiple alignment of the individual protein amino acid sequences.
[0219] Homogenetics or variants of hispidine hydroxylase are also provided. Their sequences differ from the specific amino acid sequences claimed in this invention, i.e., SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, and 28. The homolog or mutant in question has at least 40% identity, for example, at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity). In particular, the homologous or mutant in question relates to the amino acid sequence that provides the functional site of the protein, namely the sequence of the FAD / NAD binding domain, which is part of hispidin hydroxylase.
[0220] In preferred embodiments, the amino acid sequences of the hispidine hydroxylase of the present invention are characterized by the presence of several conserved amino acid motifs (consensus sequences) that are typical only to this group of enzymes. These consensus sequences are shown in SEQ ID NOs: 29-33. The consensus sites within the hispidine hydroxylase amino acid sequences are operationally bound to the lower insertion sites via amino acid inserts.
[0221] Hispidin Synthase The hispidin synthase of the present invention is a protein capable of catalyzing the synthesis of preluciferin from its precursor. In other words, the hispidin synthase is an enzyme that catalyzes the transformation of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula, where R is aryl or heteroaryl.
[0222] [ka]
[0223] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula. In the formula, R is aryl or heteroaryl.
[0224] [Chemical formula]
[0225] Examples of 3-arylacrylic acid, which is a precursor of preleuciferin, are shown in Table 2.
[0226] The reaction is carried out under in vitro and in vivo physiological conditions in the presence of at least 1 molecule of coenzyme A, at least 1 molecule of ATP, and at least 2 molecules of malonyl-CoA.
[0227] [Chemical formula]
[0228] The target hispidin-synthase includes proteins derived from the bioluminescent fungi Shirohikari take, Asigronara take, Narata take, Guyanagaster necrohiza, Mycena citricolor, Neonotepanus gardneri, Omphalotus olearius, Wasabitake, Watagenara take, Oninarata take, and Yakoutake. These amino acid sequences are shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, as well as their functional variants, homologs, and derivatives.
[0229] In a preferred embodiment, the amino acid sequence of the hispidin-synthase of the present invention is characterized by the presence of several conserved amino acid motifs (consensus sequences) typical only of this group of enzymes. These consensus sequences are shown in SEQ ID NOs: 56 to 63. The consensus sites within the hispidin-synthase amino acid sequence are operably linked to the lower insertion part via an amino acid insert.
[0230] In many embodiments of the present invention, the relevant amino acid sequences of specific hispidin-synthase homologs and mutants are characterized by being substantially identical to the sequences shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, which have, for example, at least 40% identity in the whole protein amino acid sequence, e.g., at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, e.g., at least 80% identity, at least 85% identity, or at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0231] In preferred embodiments, the hispidin synthase of the present invention is a polydomain protein related to the polyketide synthase superfamily. In preferred embodiments, the hispidin synthase of the present invention undergoes posttranslational modification. Specifically, the transfer of 4-phosphopantetheinyl from coenzyme A to serine in the acyl carrier domain of the polyketide synthase is required for the maturation of the hispidin synthase. The enzyme 4'-phosphopantetheinyltransferase, which performs such modification, is known from the prior art [Gao Menghao et al., Microbial Cell Factories, 2013, No. 12: p. 77]. 4'-phosphopantetheinyltransferase is naturally expressed in many plants and fungi, and in their cells, the functional hispidin synthase of the present invention matures without the introduction of additional enzymes or nucleic acids that encode it. Simultaneously, the introduction of sequences encoding 4'-phosphopantheteinyltransferase into host cells is necessary for the maturation of hispidin synthases in several lower fungal (e.g., yeast) and animal cells. It will be obvious to those skilled in the art that any functional variant of 4'-phosphopantheteinyltransferase known from the prior art can be used for the purposes of the present invention. For example, 4'-phosphopantheteinyltransferase NpgA derived from Aspergillus nidurans (SEQ ID NOs. 104, 105) described in [Gao Menghao et al., Microbial Cell Factories, 2013, No. 12: p. 77], and its homologs or variants whose activity has been confirmed can be used.
[0232] Caffeoyl pyruvate hydrolase The caffeioylpyruvate hydrolase of the present invention is a protein capable of catalyzing the transformation of 3-arylacrylic acid into oxyluciferin, which is 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid. The 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid has the following structural formula, where R is aryl or heteroaryl.
[0233] [ka]
[0234] The 3-arylacrylic acid in question has the following structural formula, where R is aryl or heteroaryl.
[0235] [ka]
[0236] Table 2 shows an example of oxyluciferin.
[0237] The reactions will be carried out under the following in vitro and in vivo physiological conditions.
[0238] [ka]
[0239] In a preferred embodiment, the caffeylpyruvate hydrolase of the present invention transforms caffeylpyruvate into caffeic acid. In a preferred embodiment, the caffeylpyruvate hydrolase transforms oxyluciferin shown in Table 2 into a preluciferin precursor.
[0240] The target caffeoylpyruvate hydrolases include proteins derived from the bioluminescent fungi *Lycoperdon perlatum*, *Armillaria mellea*, *Armillaria mellea*, *Guyanagastar necrohiza*, *Misena citricolor*, *Neonotepanus gardneri*, *Omphalotus olearius*, *Armillaria mellea*, *Armillaria gracilis*, and *Mycena chlorophos*. Their amino acid sequences are shown in SEQ ID NOs: 65, 67, 69, 71, 73, 75, and their functional variants, homologs, and derivatives.
[0241] In preferred embodiments, the caffeypyruvate hydrolase amino acid sequence of the present invention (including the homologs and variants of the present invention) is characterized by the presence of several conserved amino acid motifs (consensus sequences) that are typical only to this group of enzymes. These consensus sequences are shown in SEQ ID NOs: 76-78. The consensus sites within the caffeypyruvate hydrolase amino acid sequence are operationally bound to the lower insertion site via amino acid inserts.
[0242] In many embodiments of the present invention, the relevant amino acid sequence of caffeylpyruvate hydrolase is characterized by being substantially identical to the sequence shown in SEQ ID NOs: 65, 67, 69, 71, 73, 75, which has, for example, at least 40% identity in the whole protein amino acid sequence, for example, at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0243] Homogenetics of the above-mentioned specific proteins (i.e., proteins with amino acid sequence numbers: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 65, 67, 69, 71, 73, 75) can be isolated from natural sources. Homogenetics can be found in many organisms (fungi, plants, microorganisms, animals). In particular, homologouss are found in various types of bioluminescent fungi, such as basidiomycetes, mainly the basidiomycete class, especially in the Agaricales order.Furthermore, non-bioluminescent fungi and plants that produce hispidin have attracted particular attention as sources of the protein homologs of the present invention. Examples include Pteris ensiformis [Yung-Husan Chen et al., "Identification of phenolic antioxidants from Sword Brake fern (Pteris ensiformis Burm.)", Food Chemistry, 2007, Vol. 105, No. 1, pp. 48-56], Inonotus xeranticus [In-Kyoung Lee et al., "Hispidin Derivatives from the Mushroom Inonotus xeranticus and Their Antioxidant Activity", Journal of Natural Products, 2006, No. 69(2), pp. 299-301], and Phellinus sp. [In-Kyoung Lee et al., "Highly oxygenated and unsaturated metabolites providing a diversity of hispidin class antioxidants in the medicinal Examples include "Mushrooms Inonotus and Phellinus," Bioorganic & Medicinal Chemistry, No. 15(10): pp. 3309-14, and horsetail (Equisetum arvense) [Markus Herderich et al., "Establishing styrylpyrone synthase activity in cell-free extracts obtained from gametophytes of Equisetum arvense L. by high performance liquid chromatography-tandem mass spectrometry," Phytochemical Analysis, No. 8: pp. 194-197].
[0244] Proteins that are derivatives or variants of the aforementioned naturally occurring proteins are also provided. The variants and derivatives may retain the biological properties of the wild-type protein (e.g., naturally occurring) or may have different biological properties. Examples of mutations include the substitution of one or more amino acids, the deletion or insertion of one or more amino acids, the substitution, truncation, or extension of the N-terminus, and the substitution, truncation, or extension of the C-terminus. The variants and derivatives are obtained using standard molecular biology methods, as detailed in the "Nucleic Acids" section. The variants are substantially identical to the wild-type protein, i.e., they have at least 40% identity with the wild-type protein in selected regions for comparison. Therefore, nearly similar sequences will have, for example, at least 40% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 62% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, or at least 85% identity, or at least 90% identity (e.g., 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99%) within the region selected for comparison. In many embodiments, the homologues in question have a much higher degree of identity with respect to the amino acid sequence that provides the protein functional domain, for example, 70%, 75%, 80%, 85%, 90% (e.g., 92%, 93%, 94%) or more, such as 95%, 96%, 97%, 98%, 99%, 99.5%.
[0245] Derivatives can be obtained using standard methods, including those involving RNA-mediated changes, chemical modifications, post-translational modifications, and post-transcriptional modifications. For example, derivatives can be obtained by methods such as modified phosphorylation, glycosylation, acetylation, and lipidation, or by heterologous separation during maturation.
[0246] Methods well known to those skilled in the art are used to search for functional variants, homologs, and derivatives. For example, functional screening of expression libraries consisting of variants (e.g., protein variant morphologies, homologous proteins, or protein derivatives) is performed. Expression libraries are obtained by cloning nucleic acids encoding test variants of a protein into expression vectors and introducing them into suitable host cells. Procedures using nucleic acids are described in detail in the "Nucleic Acids" section. To identify the functional enzymes of the present invention, appropriate substrates are added to cells expressing the test nucleic acid. The expected product formation of reactions catalyzed by the functional enzymes can be detected by HPLC using synthetic variants of the expected reaction products as standards. For example, hispidin or other preluciferins shown in Table 2 can be used as substrates for identifying functional hispidin hydroxylase. The expected reaction product is fungal luciferin. Preluciferin precursors (e.g., caffeic acid) can be used as substrates for identifying hispidin synthase, and the corresponding fungal preluciferin is the reaction product. It should be noted that to screen for functional hispidin synthases, host cells express 4'-phosphopantheteinyltransferase, which promotes post-translational modification of proteins.
[0247] Oxylciferin (Table 2) was used as a substrate to search for functional caffeylpyruvate hydrolase, and the test reaction product was preluciferin precursor-3-arylacrylic acid.
[0248] In many embodiments of the present invention, bioluminescence reactions can be used to explore the functional enzymes of the present invention. In this case, for the purpose of preparing an expression library, cells that produce luciferase capable of oxidizing fungal luciferin by luminescence, and functional enzymes that promote the production of fungal luciferin from the products of an enzymatic reaction carried out by the test protein are used.
[0249] Thus, to screen for functional hispidin hydroxylase, host cells that produce functional luciferase using fungal luciferin as a substrate are used. When preluciferin is added to cells consisting of functional variants of hispidin hydroxylase, fungal luciferin is formed, and luminescence appears due to the oxidation of fungal luciferin by luciferase.
[0250] To screen for functional hispidin synthases, host cells that additionally produce functional luciferase and functional hispidin hydroxylase using fungal luciferin as a substrate are used. When a preluciferin precursor is added to such cells, fungal luciferin is formed, and luminescence appears upon oxidation of fungal luciferin by luciferase.
[0251] To screen for functional caffeylpyruvate hydrolase, host cells that produce functional luciferase, functional hispidin hydroxylase, and functional hispidin synthase using fungal luciferin as a substrate are used. When oxyluciferin is added to such cells, fungal luciferin is formed, and luminescence appears upon oxidation of fungal luciferin by luciferase.
[0252] Any luciferase capable of oxidizing luminescent luciferin, selected from the group of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one having the following general formula, can be used for screening. In the formula, R is either aryl or heteroaryl.
[0253] [ka]
[0254] Table 1 shows non-limiting examples of luciferin. Non-limiting examples of suitable luciferases are described in the "Applications, Combinations, and Methods of Use" section below.
[0255] Luminescence was detected during luciferin oxidation by luciferase. The light emitted during oxidation can be detected using standard methods (e.g., visual observation, observation with night vision devices, spectrophotometer, fluorometer, photographic recording, special equipment for luminescence and fluorescence detection, e.g., IVIS Spectrum In Vivo Imaging System (Perkin Elmer)). The recorded luminescence can be emitted in an intensity range from one photon to luminescence easily perceived by the eye, e.g., an intensity of 1 cd, or bright luminescence with an intensity of 100 cd or more. The light emitted during the oxidation of 3-hydroxyhispidine is in the range of 400-700 nm, preferably 450-650 nm, with a maximum emission value of 520-590 nm. The light emitted during the oxidation of other 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one may have a maximum emission value shift (Table 3).
[0256] [Table 10]
[0257] Examples of functional screening using bioluminescence are described in the experimental section below.
[0258] The present invention also covers fusion proteins comprising the protein of the present invention. Fusion proteins are homologs or variants, including truncated or elongated forms. The protein of the present invention can be operationally fused with intracellular localization signals (e.g., nuclear localization signals, localization signals in mitochondria or peroxisomes or lysosomes or Golgi apparatus or other organelles), signal peptides that facilitate protein isolation into intracellular space, transmembrane domains, or any protein or polypeptide of interest (fusion partner). Fusion proteins may, for example, include operationally crosslinked hispidin hydroxylase and / or hispidin synthase and / or caffeylpyruvate hydrolase as claimed in the present invention, which have a fusion partner linked to the C-terminus or N-terminus. Non-limiting examples of fusion partners include the protein of the present invention having other enzymatic functions, antibodies or their linked fragments, ligands or receptors, and luciferases that can use fungal luciferin as a substrate in bioluminescence reactions. In some embodiments, the fusion partner and the protein of the present invention are operationally crosslinked via linking sequences (peptide linkers) that facilitate the folding and function of independent fusion proteins. Methods for producing fusion proteins are well known to those skilled in the art.
[0259] In some embodiments, the fusion protein comprises the hispidine hydroxylase of the present invention and a luciferase capable of oxidizing the luminescent fungal luciferin, which are operationally crosslinked via a short-chain peptide linker. Such fusion proteins can be used to obtain bioluminescence in vitro and in vivo in the presence of preluciferin (e.g., in the presence of hispidine). It will be obvious to those skilled in the art that any of the functional hispidine hydroxylases described above can be used together with any functional luciferase to produce the fusion protein. Specific examples of fusion proteins are described in the experimental section below. Examples of luciferases that can be used in producing the fusion protein are described in the "Applications, Combinations, and Methods of Use" section below.
[0260] nucleic acid The present invention provides nucleic acids encoding enzymes for fungal luciferin biosynthesis, including truncated and elongated forms, variants and homologs of the protein thereof.
[0261] The nucleic acids used herein are isolated DNA molecules such as genomic DNA molecules or cDNA molecules, or RNA molecules such as mRNA molecules. In particular, the nucleic acids are cDNA molecules having an open reading frame encoding the luciferin biosynthesis enzyme of the present invention, and it is possible to ensure the expression of the enzyme of the present invention under appropriate conditions.
[0262] The term "cDNA" is used to describe mature mRNA, which is a nucleic acid that reflects the naturally occurring arrangement of sequence elements, and whose sequence elements are exons and 5' and 3' non-coding regions. Immature mRNA may have exons separated by intervening introns, which, if present, are removed during post-translational RNA spicing, resulting in mature mRNA with an open reading frame.
[0263] The target genome sequence may include nucleic acids located between the start and end codons determined in the sequence, and these nucleic acids may include all introns normally present in native chromosomes. The target genome sequence may additionally include the 5' and 3' untranslated regions of mature mRNA, as well as specific transcription and translation regulatory sequences such as promoters and enhancers, and may include adjacent genomic DNA of approximately 1 kbp in size, and possibly larger, at the 5' or 3' end of the transcription region.
[0264] The present invention also covers nucleic acids that are homologous, substantially similar, or identical derivatives or mimics of the nucleic acids encoding the proteins of the present invention.
[0265] The claimed nucleic acid exists in an environment different from the medium in which it naturally occurs, for example, in isolation, in concentrated quantities, or present or expressed in vitro or within cells, or in an organism different from the environment in which it naturally occurs.
[0266] The specific nucleic acids of interest include nucleic acids encoding hispidin hydroxylase, hispidin synthase, or caffeylpyruvate hydrolase, as described in the "Proteins" section above. Each of these specific nucleic acids of interest is disclosed individually and in more detail.
[0267] Nucleic acid encoding hispidin hydroxylase In preferred embodiments, the nucleic acid of the present invention encodes a protein capable of catalyzing the transformation reaction of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one (fungal luciferin) with 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one (preluciferin). The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0268] [ka]
[0269] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0270] [ka]
[0271] In a preferred embodiment, the nucleic acid encodes hispidine hydroxylase, and the amino acid sequence of hispidine hydroxylase is characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 29-33.
[0272] Specific examples of nucleic acids include nucleic acids encoding hispidin hydroxylase having the amino acid sequences shown in SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, and 28. Examples of nucleic acids encoding the aforementioned proteins are shown in SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, and 27. Functional variants, homologs, and derivatives of the aforementioned specific nucleic acids are also included.
[0273] In a preferred embodiment, the nucleic acid of the present invention encodes a protein having an amino acid sequence that matches at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% of the sequence shown in SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, and 28, at least 350 amino acids.
[0274] Nucleic acid encoding hispidin synthase In preferred embodiments, the nucleic acid of the present invention encodes a protein capable of catalyzing the transformation reaction of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula, where R is aryl or heteroaryl.
[0275] [ka]
[0276] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0277] [ka]
[0278] In a preferred embodiment, the nucleic acid encodes a hispidin synthase, and the amino acid sequence of the hispidin synthase is characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 56-63.
[0279] Specific examples of nucleic acids include nucleic acids encoding the hispidin-synthase of the present invention, having the amino acid sequences shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55. Examples of nucleic acids encoding the protein are shown in SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, and 54.
[0280] Furthermore, functional variants, homologs, and derivatives of the specific nucleic acids mentioned above are also included.
[0281] In a preferred embodiment, the nucleic acid of the present invention encodes a protein having an amino acid sequence that matches at least 45%, typically at least 50%, for example, at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% of the sequence shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55 in the entire protein polypeptide chain.
[0282] Nucleic acid encoding caffeylpyruvate hydrolase In a preferred embodiment, the nucleic acid of the present invention encodes a protein capable of catalyzing the transformation reaction of oxyluciferin to 3-arylacrylic acid. The oxyluciferin has the following structural formula, where R is aryl or heteroaryl.
[0283] [ka]
[0284] The 3-arylacrylic acid in question has the following structural formula, where R is selected from aryl and heteroaryl groups.
[0285] [ka]
[0286] In preferred embodiments, the nucleic acid encodes caffeypyruvate hydrolase, and the amino acid sequence of caffeypyruvate hydrolase is characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 76-78. Specific examples of nucleic acids include those encoding caffeypyruvate hydrolase, the amino acid sequence of which is shown in SEQ ID NOs: 65, 67, 69, 71, 73, and 75. Examples of nucleic acids encoding the protein are shown in SEQ ID NOs: 64, 66, 68, 70, 72, and 74.
[0287] Furthermore, nucleic acids encoding functional variants, homologs, and derivatives of the aforementioned proteins are also included.
[0288] In a preferred embodiment, the nucleic acid of the present invention codes for a protein whose amino acid sequence matches, by at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, the sequence shown in SEQ ID NOs: 65, 67, 69, 71, 73, and 75 of the entire protein polypeptide chain.
[0289] The target nucleic acids (for example, nucleic acids encoding protein homologs characterized by the amino acid sequences shown in SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 65, 67, 69, 71, 73, 75) can be isolated from any organism (fungi, plants, microorganisms, animals), and are particularly diverse. Bioluminescent fungi can be isolated from various types of bioluminescent fungi, such as basidiomycetes, mainly the basidiomycete class, especially the Agaricales order, for example, bioluminescent fungi such as *Lycoperdon perlatum*, *Armillaria mellea*, *Armillaria mellea*, *Guyanagastar necrohiza*, *Misena citricolor*, *Neonotepanus gardneri*, *Omphalotus olearius*, *Armillaria mellea*, *Armillaria gracilis*, and *Mycena chlorophos*.TAlso, non-bioluminescent fungi and plants that produce hispidin, such as Pteris ensiformis, are specially targeted as sources of nucleic acids encoding the homologous proteins of the present invention [Yung-Husan Chen et al., "Identification of phenolic antioxidants from Sword Brake fern (Pteris ensiformis Burm.)", Food Chemistry, 2007, Vol. 105, No. 1, pp. 48-56], Inonotus xeranticus [In-Kyoung Lee et al., "Hispidin Derivatives from the Mushroom Inonotus xeranticus and Their Antioxidant Activity", Journal of Natural Products, 2006, No. 69(2), pp. 299-301], and Phellinus [In-Kyoung Lee et al., "Highly oxygenated and unsaturated metabolites providing a diversity of hispidin class antioxidants in the medicinal mushrooms Inonotus and Phellinus", Bioorganic & Medicinal]. Chemistry, No. 15(10): p.3309-14], horsetail [Markus Herderich et al., "Establishing styrylpyrone synthase activity in cell free extracts obtained from gametophytes of Equisetum arvense L. by high performance liquid chromatography-tandem mass spectrometry", Phytochemical Analysis, No. 8: p.194-197].
[0290] Homogenies are identified using one of several methods. The cDNA fragments of the present invention can be used as a comparison between a hybridization probe and a cDNA library from a target organism, utilizing low stringency conditions. The probe may be a large fragment, a single degenerate primer, or a short degenerate primer. Nucleic acids with sequence similarity are detected by hybridization under low stringency conditions, e.g., 50°C and 6×SSC (0.9M sodium chloride / 0.09M sodium citrate), followed by washing at 50×C in 1×SSC (0.15M sodium chloride / 0.015M sodium citrate). Sequence identity can be determined by hybridization under high stringency conditions, e.g., 50°C or higher and 0.1×SSC (15mM sodium chloride / 1.5mM sodium citrate). A nucleic acid containing a region nearly identical to the presented sequence, such as an allelic variant or a genetically modified variant of the nucleic acid, is bound to the presented sequence under high stringency hybridization conditions. Using probes, particularly DNA sequence-labeled probes, makes it possible to recover homologous or similar genes.
[0291] Homogenies can be identified by performing polymerase chain reactions from a genomic library or cDNA library. Oligonucleotide primers representing known sequence fragments of specific nucleic acids can be used as PCR primers. In a preferred embodiment, the oligonucleotide primers have a degenerate structure and correspond to nucleic acid fragments encoding a conserved region of the protein amino acid sequence; for example, consensus sequences are shown in SEQ ID NOs: 29-33, 56-63, and 76-78. The full-length coding sequence can then be detected using 3'- and 5'-RACE methods, which are well known from the prior art.
[0292] By comparing the amino acid sequence predicted based on sequencing with the sequence numbers 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 65, 67, 69, 71, 73, and 75 of the amino acid sequence, homologs can also be identified in the results of whole-genome sequencing of an organism. Sequence identity is determined based on the reference sequence. Algorithms for sequence analysis are publicly known in the art, such as BLAST described by Altschul et al., Journal of Molecular Biology, 1990, No. 215; pp. 403-410. For the purposes of this invention, it is possible to use a comparison of nucleotide sequences and amino acid sequences performed by the Blast software package provided by the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / blast) using notched alignment with standard parameters to determine the level of identity and similarity between nucleotide sequences and amino acid sequences.
[0293] We also provide nucleic acids hybridized with the above nucleic acids under stringency conditions, preferably high stringency conditions (i.e., nucleic acids complementary to the above nucleic acids). Examples of hybridization under high stringency conditions include hybridization at 50°C or higher and in 0.1×SSC (15 mM sodium chloride / 1.5 mM sodium citrate). Another example of hybridization under high stringency conditions is incubation overnight at 42°C in 50% formamide, 5×SSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH7.6), 5× Denhardt's solution, 10% dextran sulfate, and 20 μg / ml salmon spermatogenic cleavage DNA, followed by pre-washing with 0.1×SSC at approximately 65°C. Other high stringency conditions for hybridization are known in the art and can also be used for the identification of nucleic acids of the present invention.
[0294] The present invention also provides nucleic acids encoding variants, mutants, or derivatives of the protein. Variants of the nucleic acid template can be obtained by modifying, deleting, or adding one or more nucleotides to a template sequence or combination thereof, thereby obtaining mutants or derivatives from a nucleic acid template selected from the above-mentioned nucleic acids. Modification, addition, or deletion can be carried out by any method known in the art (e.g., Gustin et al., Biotechniques, 1993, No. 14: p. 22; Barany, Gene, 1985, No. 37: pp. 111-123; Colicelli et al., Molecular Genetics and Genomics, 1985, No. 199: pp. 537-539; Sambrook et al., Molecular Cloning: A Laboratory Manual, CSH Press, 1989, pp. 15.3-15.108). The methods include error-prone PCR, shuffling, oligonucleotide-directed mutagenesis, assembly PCR, paired PCR mutagenesis, in vivo mutagenesis, cassette mutagenesis, inductive ensemble mutagenesis, exponential ensemble mutagenesis, oligonucleotide-directed mutagenesis, random mutagenesis, genetic reconstruction, site-saturated mutagenesis (GSSM), synthetic ligation reconstruction (SLR), or combinations thereof. Modification, addition, or deletion can be performed by methods including recombination, inductive sequence recombination, phosphorothioate-modified DNA mutagenesis, uracil template mutagenesis, double-skip mutagenesis, point reducing mismatch mutagenesis, recovered deletion mutant mutagenesis, chemical mutagenesis, radioactive mutagenesis, deletion mutagenesis, restriction-selective mutagenesis, restriction mutagenesis with purification, artificial gene synthesis, multiple mutagenesis, creation of chimeric multiple nucleic acids, or combinations thereof. Nucleic acids encoding truncated and elongated variants of the luciferase are also within the scope of the present invention. As used herein, these protein variants consist of amino acid sequences in which the C-terminus, N-terminus, or both ends of a polypeptide chain are modified.
[0295] In preferred embodiments, the homologs and mutants discussed are functional enzymes capable of fungal luciferin biosynthesis, such as fungal luciferin. The homologs and mutants of interest may have modified properties such as maturation rate in host cells, aggregation or dimerization, half-life, or other biochemical properties including substrate binding constant, thermal stability, pH stability, optimal activity temperature conditions, optimal activity pH conditions, Michaelis-Menten constant, substrate specificity, and secondary problem ranges. In some embodiments, the homologs and mutants have the same properties as the claimed protein.
[0296] Nucleic acids encoding functional homologs and variants of the present invention can be identified during functional testing, for example, in the expression library functional screening described in the "Proteins" section.
[0297] Furthermore, degenerate variants of nucleic acids encoding the proteins of the present invention are also provided. Degenerate variants of nucleic acids include substituted nucleic acid codons in which the codons are replaced with other codons encoding the same amino acid. In particular, degenerate variants of nucleic acids are created to increase expression in host cells. In this embodiment, undesirable or less desirable nucleic acid codons in host cell genes are replaced with codons that are excessively presented in the coding sequence of the host cell gene, where the substituted codons encode the same amino acid. In particular, humanized versions of the nucleic acids of the present invention are also covered. As used herein, the term “humanization” refers to substitutions made in nucleic acid sequences to optimize codons for protein expression in mammalian cells (Yang et al., Nucleic Acids Research, 1996, No. 24: pp. 4592-4593). See also U.S. Patent No. 5,795,737, which describes humanization of proteins. This disclosure is incorporated herein by reference. Variants of nucleic acids optimized for expression in plant cells are of particular interest. Examples of nucleic acids encoding the proteins of the present invention are shown in SEQ ID NOs: 103, 113, and 114.
[0298] The requested nucleic acid is isolated and obtained in a nearly purified form. Primarily, the purified form means that the nucleic acid has a purity of at least about 50%, and usually at least about 90%, and is typically a "recombinant," i.e., one or more nucleotides into which loxP has been introduced. This nucleic acid does not typically bind to chromosomes naturally present in its native host organism.
[0299] The claimed nucleic acid can be synthesized artificially. Methods for generating nucleic acids are well known from the prior art. For example, if information on amino acid sequences or nucleotide sequences is available, isolated molecules of the nucleic acid of the present invention can be produced by oligonucleotide synthesis. If information on amino acid sequences is available, several distinct nucleic acids can be synthesized by degeneracy of the gene code. Methods for selecting the necessary host codon variants are well known in the art.
[0300] Synthetic oligonucleotides can be produced by the phosphoramidite method, and the resulting constructs can be purified by methods well known in the art, such as high-performance liquid chromatography (HPLC), or by other methods. Other methods are described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, New York, 1989, 2nd edition, and also in accordance with the instructions provided in the guidelines of the United States Department of Health and Human Services and the National Institutes of Health (NIH) for recombinant DNA research. The long double-stranded DNA molecule of the present invention can be synthesized as follows: Several small fragments containing suitable ends capable of aggregation with adjacent fragments can be synthesized with the necessary complementarity. Adjacent fragments can be crosslinked by DNA ligase, a recombinant-based method, or a PCR-based method.
[0301] We also provide nucleic acids encoding fusion proteins, including the protein of the present invention. Examples of such proteins are described in the "Proteins" section above. Nucleic acids encoding fusion proteins can be synthesized artificially as described above.
[0302] In particular, the present invention also provides expression cassettes or systems for obtaining fusion proteins based on or for replicating the claimed proteins (i.e., hispidin hydroxylase, hispidin synthase, and caffeylpyruvate hydrolase) or the claimed nucleic acid molecules. The expression cassettes may exist as extrachromosomal elements or may be included in a cell genome obtained by introducing the expression cassettes into cells. When the expression cassettes are introduced into cells, protein products encoded by the nucleic acids of the present invention are formed; in this case, the protein can be said to be "produced" or "expressed" by the cells. For example, any expression system including bacterial, yeast, plant, insect, amphibian, or mammalian cells is applicable. The target nucleic acid in the expression cassette is operationally bound to a regulatory sequence which may include a promoter, enhancer, terminator sequence, operator, repressor, and inductor. Generally, the expression cassettes consist of at least (a) a transcription start region that functions in a host cell; (b) the nucleic acid of the present invention; and (c) a transcription termination region that functions in a host cell. Methods for obtaining an expression cassette or system capable of expressing a desired product are known to those skilled in the art.
[0303] Vectors comprising the claimed nucleic acids and other nucleic acid structures are also provided. Suitable vectors include viral and nonviral vectors, plasmids, cosmids, phages, etc., and are used for cloning, amplification, expression, transfer, etc., of the nucleic acid sequence of the present invention to a suitable host. The selection of a suitable vector is obvious to those skilled in the art. Generally, a full-length nucleic acid or a portion thereof is introduced into the vector by a DNA ligase linked to the site of division by restriction enzymes in the vector. Alternatively, the desired nucleotide sequence can be inserted by in vivo homologous recombination, usually by linking a homologous region to the vector adjacent to the desired nucleotide sequence. The homologous region is added as part of the desired nucleotide sequence, for example, by oligonucleotide ligation or polymerase chain reaction using a primer containing the homologous region. In principle, the vector has a replication origin that promotes replication in the host cell as a result of being introduced into the cell as an extrachromosomal element. The vector also consists of regulatory elements that promote the expression of nucleic acids in the host cell to obtain recombinant functional proteins. In an expression vector, the nucleic acid is functionally bound to a regulatory sequence which may include a promoter, enhancer, terminator, operator, repressor, silencer, insulator, and inductor. For the purpose of expressing a functional protein or a truncated form thereof, the coding nucleic acid is operationally crosslinked to at least the regulatory sequence and the nucleic acid constituting the transcription start site. These nucleic acids may consist of a histidine tag (6His tag), a signal peptide, or a sequence encoding the domain of the functional protein. In many embodiments, the vector facilitates the integration of the operationally bound nucleic acid with the regulatory elements into the host cell genome. The vector may consist of an expression cassette for a selectable marker such as a fluorescent protein (e.g., GFP), an antibiotic resistance gene (e.g., resistance genes for ampicillin, kanamycin, neomycin, or hygromycin, etc.), a gene that modulates herbicide resistance such as a gene that modulates resistance to phosphinotricin and sulfonamide herbicides, or other selectable markers known from the prior art.
[0304] The vector can consist of an additional expression cassette containing 4'-phosphopantetheinyltransferase, nucleic acids encoding 3-arylacrylic acid synthesis protein (e.g., as described in the "Applications, Combinations, and Uses" section), luciferase, etc.
[0305] The expression systems described above can be used in prokaryotic or eukaryotic hosts. To obtain the protein, host cells other than Escherichia coli, Bacillus subtilis, S. cerevisiae, insect cells, or human embryonic cells, such as yeast, plants (e.g., Arabidopsis thaliana, Nicotiana benthamiana, Physcomitrella patens), and vertebrates (e.g., COS7 cells, HEK293, CNO, Xenopus oocytes, etc.) can be used.
[0306] Cell lines that reliably produce the protein of the present invention can be selected by methods known in the art (e.g., cotransfection with selectable markers such as dhfr, gpt, or antibiotic resistance genes (e.g., ampicillin, kanamycin, neomycin, or hygromycin)), and this method makes it possible to identify and isolate transfected cells consisting of genes that are included in the genome or incorporated into extrachromosomal elements.
[0307] When using the aforementioned host cells or other host cells or organisms suitable for replication and / or expression of the nucleic acids of the present invention, the resulting replicated nucleic acids, expressed proteins, or polypeptides fall within the scope of the present invention as host cell or organism products. The products can be isolated by preferred methods known in the art.
[0308] In many embodiments of the present invention, cells are co-transfected with multiple expression cassettes consisting of nucleic acids of the present invention encoding various enzymes of fungal luciferin biosynthesis. In some embodiments, an expression cassette consisting of a nucleic acid encoding a luciferase capable of oxidizing fungal luciferin with luminescence is additionally introduced into the cells. In some cases, the expression cassette is bound to a single vector used for cell transformation. In some embodiments, a nucleic acid encoding 4'-phosphopantetheinyltransferase and / or 3-arylacrylic acid synthesis protein is additionally introduced into the cells.
[0309] We also provide short DNA fragments of the claimed nucleic acid for use as PCR primers, rolling circle amplifiers, hybridization screening probes, etc. As described above, the long DNA fragments are used to obtain encoded polypeptides. However, in geometric amplification reactions such as PCR, a pair of short DNA fragments, i.e., primers, are used. The exact primer sequence is not important to the present invention, but in most applications, as is known in the art, the primers hybridize with the claimed sequence under stringent conditions. It is preferable to select a pair of primers, which can yield amplification products from at least about 50 nucleotides, preferably at least about 100 nucleotides, and extend to the entire nucleic acid sequence. Algorithms for primer sequence selection are known and are available in commercially available software packages. The amplification primers hybridize with complementary DNA strands and serve as the basis for amplification counter-reactions.
[0310] The nucleic acid molecules of the present invention can also be used to determine gene expression in biological specimens. Methods for examining cells for the presence of specific nucleotide sequences, such as genomic DNA or RNA, are well known in the art. Specifically, DNA or mRNA is isolated from a cell specimen. A complementary DNA strand can be formed using reverse transcriptase, and then the mRNA can be amplified by polymerase chain reaction using primers specific to the claimed DNA sequence. Alternatively, the mRNA specimen can be isolated by gel electrophoresis, transferred to a suitable carrier, such as nitrocellulose or nylon, and then tested with the claimed DNA fragment as a probe. Other methods can also be used, such as oligonucleotide ligation analysis, hybridization with insights, and hybridization with DNA probes immobilized on hard arrays. If mRNA hybridizing with the claimed sequence is detected, it indicates that the gene was expressed in the specimen.
[0311] transgenic organisms The present invention also provides transgenic organisms, transgenic cells, and transgenic cell lines expressing the nucleic acids of the present invention. The transgenic cells of the present invention contain one or more nucleic acids studied in the present invention as introduced genes. For the purposes of the present invention, any suitable host cell can be used, including prokaryotic host cells (e.g., Escherichia coli, Streptomyces sp., Bacillus subtilis, Lactobacillus acidophilus, etc.) or eukaryotic host cells other than human embryonic cells. The transgenic organisms of the present invention can be prokaryotic or eukaryotic organisms including bacteria, cyanobacteria, fungi, plants, and animals, where one or more organismal cells consisting of heterologous nucleic acids of the present invention are introduced, for example, by artificial manipulation in accordance with transgenic techniques known in the art.
[0312] In one embodiment of the present invention, the transgenic organism can be a prokaryotic organism. Methods for transforming prokaryotic host cells are well known in the art (see, for example, Sambrook et al. (Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, 1989, 2nd edition) and Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc., 1995).
[0313] In other embodiments of the present invention, the transgenic organism may be a fungus, such as yeast. Yeast is widely used as a carrier for heterologous gene expression (see, for example, Goodey et al., Yeast biotechnology, DR. Berry et al., Allen and Unwin, London, 1987, pp. 401-429; and Kong et al., Molecular and Cell Biology of Yeasts, EFWalton and GTYarronton, Blackie, Glasgow, 1989, pp. 107-133). Several yeast vectors are available, including embedded vectors that require recombination with the host genome for their maintenance, and plasmid vectors that replicate autonomously.
[0314] Other host organisms include animal organisms. Transgenic animals are well known in the art and can be obtained using transgenic techniques described in standard manuals (e.g., Pinkert, Transgenic Animal Technology: A Laboratory Handbook, Academic Press, San Diego, 2003, 2nd edition; Gersenstein and Vinterstein, Manipulating the Mouse Embryo: A Laboratory Manual, 2002, 3rd edition; Nagy A. (ed.), Cold Spring Harbor Laboratory; Blau et al., Laboratory Animal Medicine, 2002, 2nd edition; Fox JG, Anderson LC, Loew FM, Quimby FW (eds.), American Medical Association, American Psychological Association; Alexandra L. Joyner (ed.), Gene Targeting: A Practical Approach, Oxford University Press; 2000, 2nd edition). For example, transgenic animals can be obtained by homologous recombination within a framework in which endogenous loci have been altered, or by randomly incorporating nucleic acid structures into the genome. Suitable vectors for stable incorporation include plasmids, retroviruses and other animal viruses, and YACs.
[0315] Nucleic acids can be introduced into cells directly or indirectly by introduction into cell precursors using careful genetic manipulation such as microinjection, recombinant viral infection or the use of recombinant viral vectors, transfusion, transformation, gene gun delivery, or transconjugation. Techniques for such nucleic acid (e.g., DNA) molecular transfer into organisms are well-known and described in standard manuals, such as Sambrook et al. (Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, New York, 2001, 3rd edition).
[0316] The term "genetic manipulation" refers not to classical crossbreeding or in vitro fertilization, but rather to the introduction of nucleic acid recombinant molecules. These nucleic acid molecules may be incorporated into chromosomes or to extrachromosomal replicated DNA.
[0317] The DNA structure for homologous recombination comprises at least a portion of the nucleic acid of the present invention, which is operationally linked to a homologous region targeting a gene locus. In random incorporation, it is not necessary to include the homologous region in the DNA structure to facilitate recombination. Positive and negative selection markers can also be included. Methods for obtaining cells containing targeted gene modifications by homologous recombination are known in the art. Different techniques for transfection of mammalian cells are described, for example, in the paper Keown et al., Methods in Enzymology, 1990, No. 185: pp. 527-537).
[0318] In the case of embryonic stem cells (ES), ES cell lines can be used, or fresh embryonic cells can be obtained from host organisms such as mice, rats, or guinea pigs. Such cells proliferate on the corresponding fibroblast nurse layer or in the presence of leukemia suppressor (LIF). Transformed ES cells or embryonic cells can be used to create transgenic animals using related techniques known in the art.
[0319] Transgenic animals can be any animal other than humans, including mammals (e.g., mice or rats), birds, amphibians, etc., and are used for functional testing, drug screening, etc.
[0320] It is also possible to obtain transgenic plants. Methods for obtaining transgenic plant cells are described in U.S. Patents 5,767,367, 5,750,870, 5,739,409, 5,689,049, 5,689,045, 5,674,731, 5,656,466, 5,633,155, 5,629,470, 5,595,896, 5,576,198, 5,538,879, and 5,484,956, which are referenced in this invention. Methods for obtaining transgenic plants are summarized in the following reviews: Lea and Leegood (eds.), Plant Biochemistry and Molecular Biology, John Wiley & Sons, 1993, pp. 275–295; and Oksman, Caldentey and Barz (eds.), Plant Biotechnology and Transgenic Plants, 2002, p. 719.
[0321] To obtain a transgenic host organism, for example, embryogenetic explants consisting of somatic cells can be used. After harvesting the cells or tissue, the target exogenous DNA is introduced into the plant cells. Many different techniques are available for such introduction. The availability of isolated protoplasts allows for introduction using DNA-mediated gene transfer protocols. Such protocols involve incubating protoplasts with deproteinized DNA, such as plasmids containing the target exogenous coding sequence, in the presence of a polyvalent cation (e.g., PEG or poly-L-ornithine) or in the presence of native DNA containing the target exogenous sequence, according to protoplast electroporation. Protoplasts that have successfully incorporated the exogenous DNA are then selected and grown to callus formation, and finally, transgenic plants are obtained by contacting them with enhancing factors such as auxin and cytokinin incorporated in the relevant amounts and ratios.
[0322] The plants are obtained by methods based on "gene guns" known to those skilled in the art, or by other suitable methods such as Agrobacterium-mediated transformation.
[0323] antibody In this specification, the term "antibody" refers to a polypeptide or group of polypeptides containing at least one antibody active site (antigen-binding site). The term "antigen-binding site" refers to a spatial structure whose surface parameters and charge distribution are complementary to the antigen epitope: the antigen-binding site facilitates antibody binding to the relevant antigen. The term "antibody" includes, for example, vertebrate antibodies, chimeric antibodies, hybrid antibodies, humanized antibodies, modified antibodies, monovalent antibodies, Fab fragments, and single-domain antibodies.
[0324] The antibodies specific to the proteins of the present invention are applicable to affinity chromatography, immunological screening, and the detection and identification of the proteins of the present invention (hispidin hydroxylase, hisspidin synthase, caffeylpyruvate hydrolase). The target antibodies bind to the antigen polypeptides or proteins or protein fragments described in the "Proteins" section. The antibodies of the present invention can be immobilized on a carrier and used in immunological screening or affinity chromatography columns to detect and / or separate polypeptides, proteins or protein fragments, or cells containing such polypeptides, proteins or protein fragments. Alternatively, such polypeptides, proteins or protein fragments can be immobilized in a manner that allows for the detection of antibodies that can specifically bind to them.
[0325] Antibodies specific to the protein of the present invention can be obtained as polyclonal and monoclonal antibodies using standard methods. Generally, suitable mammals, preferably mice, rats, rabbits, or goats, are first immunized with the protein. Rabbits and goats are preferred subjects for obtaining polyclonal serum because a considerable amount of serum can be obtained, and significant anti-rabbit and anti-goat antibodies can be obtained. Immunization is usually carried out by mixing or emulsifying the specific protein with an adjuvant, preferably Freund's adjuvant, in physiological saline, and then introducing the resulting mixture or emulsion parenterally (usually by subcutaneous or intramuscular injection). A sufficient dose is usually 50-200 μg per injection.
[0326] In various embodiments of the present invention, recombinant proteins or native proteins are used for immunization in their native or denatured forms. Protein fragments or synthetic polypeptides consisting of a portion of the protein amino acid sequence of the present invention can also be used for immunization.
[0327] Immunization is typically performed in physiological saline, preferably with an incomplete Freund's adjuvant, with one or several protein supplemental injections for 2 to 6 weeks. Alternatively, antibodies can be obtained by in vitro immunization using methods known in the art, which, from the viewpoint of the objectives of the present invention, is equivalent to in vitro immunization. Polyclonal antiserum is obtained by collecting blood from immunized animals into a glass or plastic container, then incubating the blood at 25°C for up to 1 hour, and then incubating at 4°C for 2 to 18 hours. The serum is extracted by centrifugation (e.g., 1000 g for up to 10 minutes). 20 to 50 ml of blood can be obtained from rabbits at a time.
[0328] Monoclonal antibodies are obtained using the standard Kohler-Milstein technique (Kohler & Milstein, Nature, 1975, No. 256, pp. 495-496) or a modified thereof. Typically, mice or rats are immunized according to the information above. However, in contrast to blood collection from animals to obtain serum, this technique involves splenectomy (and, if not necessary, removal of several large lymph nodes) and tissue maceration to isolate individual cells. If necessary, splenocytes can be screened (after extracting unspecified adherent cells) by adding them to a cell suspension on a plate or to another plate well coated with protein-antigen. B lymphocytes expressing membrane-bound immunoglobulins specific to the test antigen are bound on the plate so that the B lymphocytes are not washed off the plate with suspension residue. Next, the resulting B lymphocytes or macerated whole spleen cells fuse with myeloma cells that result in hybridoma formation; then they are incubated in a selective medium (e.g., HAT medium consisting of hypoxanthine, aminopterin, and thymidine). The resulting hybridomas are seeded in a limited culture medium and tested for the responsiveness of antibodies that are specifically bound to the antigen used for immunization (but not bound to exogenous agents). Then, selected hybridomas that secrete monoclonal antibodies (mAbs) are incubated either in vitro (e.g., in a fermenter in the form of hollow fiber bundles, or in a glass container for tissue culture) or in vivo (in mouse ascites).
[0329] Antibodies (as monoclonal or polyclonal) can be tagged using standard methods. Preferred tags include fluorescent agents, chromogenic agents, radionuclides (particularly 32P and 125I), high electron-density reagents, enzymes, and ligands, for which specific binding partners are known. Enzymes are usually detected by their catalytic activity. For example, wasabi peroxidase is generally detected by its ability to convert 3,3',5,5'-tramethylbenzidine (TMB) into a blue pigment and is quantitatively evaluated by spectrophotometer. The term “specific binding partner” refers to a protein capable of binding a molecule-ligand to a molecular-ligand with a high level of specificity, such as in the case of an antigen and its specific monoclonal antibody. Other examples of specific binding partners include biotin and avidin (or streptavidin), immunoglobulin-G and protein-A, and several pairs of receptors and their ligands known in the art. Other variations and capabilities are obvious to those skilled in the art and are considered equivalent within the scope of the present invention.
[0330] The antigens, immunogens, polypeptides, proteins, or protein fragments of the present invention form specific binding partners—antibodies. The antigens, immunogens, polypeptides, proteins, or protein fragments of the present invention comprise the immunogenic compositions of the present invention. Such immunogenic compositions may further consist of, or contain, adjuvants, carriers, or other compositions that stimulate, enhance, or stabilize the antigens, polypeptides, proteins, or protein fragments of the present invention. Such adjuvants and carriers are obvious to those skilled in the art.
[0331] Applications, combinations, and usage methods The present invention provides the application of a fungal luciferin biosynthesis protein (i.e., 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one (fungal luciferin)) as an enzyme that catalyzes the reaction (1) of luciferin synthesis from preluciferin (i.e., 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one). The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0332] [ka]
[0333] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0334] [ka]
[0335] Alternatively, the above enzyme catalyzes the synthesis of preluciferin from 3-arylacrylic acid (preluciferin precursor). This 3-arylacrylic acid has the following structural formula, where R is selected from aryl or heteroaryl groups.
[0336] [ka]
[0337] Alternatively, the above enzyme catalyzes the synthesis of 3-arylacrylic acid from 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid (oxyluciferin). The 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid has the following structural formula.
[0338] [ka]
[0339] Fungal luciferin biosynthesis proteins are used in many embodiments of the present invention, and non-limiting examples of such embodiments are shown in the following chapters.
[0340] The fungal luciferin biosynthesis proteins reliably applied by the present invention can be obtained from various natural sources or by recombinant technology. For example, wild-type proteins can be isolated from bioluminescent fungi, such as basidiomycetes, mainly the basidiomycete class, especially fungi of the Agaricales order. For example, wild-type proteins can be isolated from bioluminescent fungi such as Luciferum sarmentosum, Armillaria mellea, Armillaria mellea, Guyanagastar necrohiza, Misena citricolor, Neonotepanus gardneri, Omphalotus olearius, Amomum erythrorhizon, Armillaria mellea, Armillaria gracilis, and Mycena chlorophos. These can be obtained by expressing recombinant nucleic acids encoding the protein sequence in their respective hosts, or in cell-free expression systems.
[0341] In some embodiments, the protein is applied and expressed within host cells, where periodic transformation of fungal luciferin occurs. In other embodiments, an extract consisting of isolated recombinant protein, native protein, or applied protein is used.
[0342] Fungal luciferin biosynthesis proteins exhibit activity under physiological conditions.
[0343] In some embodiments, the protein hispidin hydroxylase is applied in vitro and in vivo to obtain luciferin oxidized by bioluminescent fungal luciferases, their homologs, and luminescent mutants. Thus, the present invention provides an application of the hispidin hydroxylase of the present invention to catalyze the conversion of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one (preluciferin) to 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one (fungal luciferin). The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0344] [ka]
[0345] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0346] [ka]
[0347] A method for obtaining fungal luciferin from preluciferin comprises combining at least one molecule of hispidin hydroxylase, at least one molecule of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one, at least one molecule of NAD(P)H, and at least one molecule of molecular oxygen (O2). The reaction is carried out in vitro and in vivo under physiological conditions at a temperature of 20-42°C, and the reaction can be carried out in cells, tissues, and host organisms expressing hispidin hydroxylase. In preferred embodiments, the cells, tissues, and organisms consist of sufficient amounts of NAD(P)H and molecular oxygen to carry out the reaction. Exogenously induced 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one, or endogenous 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one produced in cells, tissues, and organisms, can be used in the reaction.
[0348] In preferred embodiments, the hispidin hydroxylase of the present invention synthesizes 3-hydroxyhispidin from hispidin. In preferred embodiments, the hispidin hydroxylase synthesizes at least one functional analogue of 3-hydroxyhispidin from the corresponding preluciferin shown in Table 2. In some embodiments, the hispidin hydroxylase of the present invention synthesizes 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one from the corresponding 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one having the following structural formula, where R is aryl or heteroaryl.
[0349] [ka]
[0350] The obtained 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one is used to induce luminescence in in vitro and in vivo systems consisting of functional luciferases, and fungal luciferin is identified as the substrate.
[0351] For the purposes of the present invention, proteins having the amino acid sequences shown in SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28 and their variants, homologs, and derivatives can be applied as hispidin hydroxylase. For example, functional hispidin hydroxylase 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28 can be used, having at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% identity in at least 350 amino acids. For example, a functional hispidin hydroxylase can be identical to the entire protein polypeptide chain by at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%.
[0352] In preferred embodiments, for the purposes of the present invention, proteins having amino acid sequences characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 29-33 can be applied as hispidin hydroxylase. The consensus sites within the amino acid sequence of the hispidin hydroxylase are operationally linked to the lower insertion site via amino acid inserts (Figure 1).
[0353] In some embodiments, the hispidin-synthase protein is applied in vitro and in vivo to produce fungal luciferin from its precursor. Specifically, it is applied to catalyze the transformation of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula, where R is an aryl or heteroaryl group.
[0354] [ka]
[0355] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0356] [ka]
[0357] A method for obtaining preluciferin involves combining at least one molecule of hispidin synthase, at least one molecule of 3-arylacrylic acid, at least one molecule of coenzyme A, at least one molecule of AMP, and at least two molecules of malonyl-CoA.
[0358] The reaction is carried out under physiological conditions at a temperature of 20-42°C, and can be carried out in cells, tissues, and host organisms that express hispidin synthase. In preferred embodiments, the cells, tissues, and organisms consist of sufficient amounts of coenzyme A, malonyl-CoA, and AMP to carry out the reaction.
[0359] Exogenously induced 3-arylacrylic acid, or 3-arylacrylic acid produced in cells, tissues, and organisms, can be used in the reaction.
[0360] For example, the hispidin synthase of the present invention can be used to produce hispidin from caffeic acid. In a preferred embodiment, the hispidin synthase synthesizes a functional analog of hispidin (6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one) from 3-arylacrylic acid shown in Table 2.
[0361] The obtained 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one is used in the production of fungal luciferin in the presence of the hispidin hydroxylase of the present invention. Hispidin and its functional analogues have been applied in the medical field due to their antioxidant and antitumor properties; there are also several arguments that hispidin can prevent obesity [Be Tu et al., Drug Discoveries & Therapeutics, June 2015; No. 9(3): pp. 197-204; Nguyen et al., Drug Discoveries & Therapeutics, December 2014; No. 8(6): pp. 238-44; Yousfi et al., Phytotherapy Research, September 2009; No. 23(9): pp. 1237-42].
[0362] For the purposes of the present invention, proteins having the amino acid sequences shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55, as well as variants, homologs, and derivatives of these proteins, are applicable as hispidin synthases. For example, functional hispidin synthases having amino acid sequences that are at least 40%, typically at least 45%, usually at least 50%, for example, at least 55%, at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% identical to sequences selected from the group SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55.
[0363] In a preferred embodiment for the purposes of the present invention, a protein having an amino acid sequence characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 56-63 can be used as a hispidin synthase. The consensus sites within the amino acid sequence of the hispidin synthase are operationally linked to the lower insertion site via amino acid inserts (Figure 2).
[0364] In some embodiments, the caffeylpyruvate hydrolase protein is applied in vitro and in vivo to produce 3-arylacrylic acid from 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid. The 3-arylacrylic acid has the following structural formula, where R is selected from an aryl or heteroaryl group.
[0365] [ka]
[0366] The 6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid has the following structural formula, where R is aryl or heteroaryl.
[0367] [ka]
[0368] This reaction is carried out under physiological conditions both in vitro and in vivo. The caffeoylpyruvate hydrolase of the present invention is applied to the autonomous bioluminescent system described in detail below.
[0369] For the purposes of the present invention, proteins having the amino acid sequences shown in SEQ ID NOs: 65, 67, 69, 71, 73, and 75, as well as variants, homologs, and derivatives of these proteins, can be applied as caffeylpyruvate hydrolases. For example, functional caffeylpyruvate hydrolases having amino acid sequences that are at least 60%, at least 65%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sequences selected from the group of SEQ ID NOs: 65, 67, 69, 71, 73, and 75, are available.
[0370] In a preferred embodiment for the purposes of the present invention, a protein having an amino acid sequence characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 76-78 can be applied as caffeylpyruvate hydrolase. The consensus sites within the amino acid sequence of caffeylpyruvate hydrolase are operationally linked to the lower insertion site via amino acid inserts (Figure 3).
[0371] The present invention also provides a combination of proteins applicable to the method of the present invention. In a preferred embodiment, the combination includes a functional hispidin hydroxylase and a functional hispidin synthase. This combination is applied to produce 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one from 3-arylacrylic acid having the following structural formula, where R is aryl or heteroaryl.
[0372] [ka]
[0373] For example, this combination can be used to produce hydroxyhispidin caffeate. This reaction is carried out under physiological conditions in the presence of at least one molecule of hispidin hydroxylase, at least one molecule of hispidin synthase, at least one molecule of 3-arylacrylic acid, at least one molecule of coenzyme A, at least one molecule of AMP, at least two molecules of malonyl-CoA, at least one molecule of NAD(P)H, and at least one molecule of molecular oxygen (O2).
[0374] In some embodiments, the combination may also include a luciferase capable of using 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one having the following structural formula, where R is aryl or heteroaryl, as in luciferin.
[0375] [ka]
[0376] The oxidation of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one by such luciferase is accompanied by bioluminescence and the formation of oxyluciferin (6-aryl-2-hydroxy-4-oxohexa-2,5-dienoic acid).
[0377] Any protein exhibiting the above-mentioned activity can be used as a luciferase. For example, known luciferases derived from bioluminescent fungi include the luciferase described in Russian Patent Application No. 2017102986 / 10(005203), filed on January 30, 2017, and its homologs, mutants, and fusion proteins having luciferase activity.
[0378] In many embodiments of the present invention, the luciferase applicable to the purposes of the present invention is characterized by an amino acid sequence that is at least 40% identical, for example, at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, to an amino acid sequence selected from the group of SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, and 98. Luciferases are frequently characterized by amino acid sequences that have the following identity with respect to an amino acid sequence selected from the group of SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98: at least 90% identity (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity, or 100% identity).
[0379] Mutants may retain the biological properties of the wild-type luciferase from which they were obtained, or they may have different biological properties from the wild-type protein. The term "biological properties" in the context of luciferase as used in this invention refers, without limitation, to the ability to oxidize various luciferins; biochemical properties such as in vivo and / or in vitro stability (e.g., half-life); maturation rate; tendency to aggregate or oligomerize; and other similar properties. Mutations include changes in one or more amino acids, deletions or insertions, substitutions or truncations of one or more amino acids, or N-terminal truncation or extension, C-terminal truncation or extension, etc.
[0380] In some embodiments of the present invention, the luciferase is used in an isolated form, that is, the luciferase is substantially free from other proteins or other naturally occurring biological molecules such as oligosaccharides, nucleic acids and their fragments. Herein, the term “substantially free” means that less than 70%, usually less than 60%, and typically less than 50%, of the composition consisting of the isolated protein are other naturally occurring biological molecules. In some embodiments, the protein is in a substantially purified form, where the term “substantially purified form” means that the purity is equal to at least 95%, usually at least 97%, and typically at least 99%.
[0381] In some embodiments, the luciferase is used as part of an extract obtained from a bioluminescent fungus or host cell, comprising nucleic acids encoding recombinant luciferase.
[0382] In many embodiments, the luciferase is present (in the cells or organisms of the present invention) in a heterologous expression system consisting of nucleic acids encoding recombinant luciferase.
[0383] Methods for generating recombinant proteins, particularly luciferases, in isolated form, as part of an extract, or in a heterologous expression system are well known in the art and are described in the "Nucleic Acids" section. Methods for protein purification are described in the "Proteins" section.
[0384] In preferred embodiments, the luciferase retains activity at temperatures below 50°C, typically up to 45°C, i.e., retains activity at temperatures between 20 and 42°C, and can be used in heterologous expression systems in vitro and in vivo. Typically, the pH stability of the described luciferase is in the range of 4 to 10, typically between 6.5 and 9.5. The optimal pH stability of the claimed protein is in the range of 7.0 to 8.5, for example, between 7.3 and 8.0. In preferred embodiments, the luciferase is active under physiological conditions.
[0385] The combination of hispidin hydroxylase and luciferase that oxidizes fungal luciferin with luminescence is applied to a method for identifying hispidin and its functional analogs in biological subjects: cells, tissues, or organisms. This method involves contacting the combination of isolated hispidin hydroxylase and luciferase with the biological subject or an extract obtained from the subject in a suitable reaction buffer consisting of components necessary to create physiological conditions and carry out the reaction. Those skilled in the art can prepare a variety of reaction buffers that satisfy these conditions. Non-limiting examples of reaction buffers include 0.2 M sodium phosphate buffer (pH 7.0-8.0) with 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM), and 1 mM NADPH added.
[0386] The presence of hispidin or its functional analogues is determined by the occurrence of detectable bioluminescence. The method for detecting detectable bioluminescence is described above in the "Proteins" section when explaining functional screening methods.
[0387] A combination of hispidin hydroxylase, hispidin synthase, and luciferase that oxidizes fungal luciferin with luminescence is applied to a method for identifying 3-arylacrylic acid having the following structural formula, where R is aryl or heteroaryl.
[0388] [ka]
[0389] The method includes isolated hispidin hydroxylase, hispidin synthase, and The reaction involves contacting a test biological object or an extract obtained from such object with a combination of luciferases consisting of components necessary to create physiological conditions and carry out the reaction. Those skilled in the art can prepare a variety of reaction buffers that satisfy these conditions. Non-limiting examples of reaction buffers include 0.2 M sodium phosphate buffer (pH 7.0-8.0) with 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM), 1 mM NADPH, 10 mM ATP, 1 mM CoA, and 1 mM malonyl-CoA added.
[0390] The presence of 3-arylacrylic acid is determined by the occurrence of detectable bioluminescence. The method for detecting detectable bioluminescence is described above in the "Proteins" section when explaining functional screening methods.
[0391] In some embodiments, a fusion protein, as described in the "Proteins" section above, can be used instead of the combination of hispidin hydroxylase and luciferase that oxidizes fungal luciferin with luminescence. The fusion protein simultaneously exhibits hispidin hydroxylase activity and luciferase activity and can be used in any way in place of the aforementioned enzyme combination.
[0392] In some embodiments, instead of the above-mentioned hispidin synthase, a type III polyketide synthase featuring the same amino acid sequence as an amino acid sequence selected from the group of SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139 can be used. For the purposes of the present invention, type III polyketide synthases having amino acid sequences that are at least 40%, typically at least 45%, usually at least 50%, for example at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% identical to sequences selected from the group of SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, and 139.
[0393] Typical polyketide synthases (PKS) have been identified in many plant organisms; the mutagenicity of these polyketide synthases and / or polyketide synthases catalyzing bisnoryangonin synthesis from coumar-CoA is well known in the art [Lim et al., Molecules, June 22, 2016; 21(6)]. The applicants have demonstrated that these enzymes can catalyze hispidin synthesis from caffeyl-CoA in vitro and in vivo as follows.
[0394] [ka]
[0395] Therefore, the application of the aforementioned protein for hispidin synthesis is also within the scope of the present invention.
[0396] In some embodiments of the present invention, PKS is used in an isolated form. That is, the PKS is substantially free from other proteins or other naturally occurring biological molecules such as oligosaccharides, nucleic acids and their fragments. In this context, the term “substantially free” means that less than 70%, usually less than 60%, and typically less than 50%, of the composition consisting of the isolated protein are other naturally occurring biological molecules. In some embodiments, the protein is in a substantially purified form. The term “substantially purified form” means a purity equal to at least 95%, usually at least 97%, and typically at least 99%.
[0397] In many embodiments, PKS is present in a heterologous expression system (in the cells or organisms of the present invention) consisting of nucleic acids encoding recombinant enzymes.
[0398] Methods for generating recombinant proteins in isolation, as part of an extract, or in a heterologous expression system are well known in the art and are described in the "Nucleic Acids" section. Methods for protein purification are described in the "Proteins" section.
[0399] In preferred embodiments, the PKS retains activity at temperatures below 50°C, typically up to 45°C, i.e., at temperatures between 20 and 42°C, and can be used in heterologous expression systems in vitro and in vivo. Typically, the pH stability of the described PKS is in the range of 4 to 10, typically between 6.0 and 9.0. The optimal pH stability of the claimed protein is in the range of 6.5 to 8.5, for example, between 7.0 and 7.5. In preferred embodiments, the PKS is active under physiological conditions.
[0400] A method for obtaining hispidin involves combining at least one molecule of the above-mentioned type III polyketide synthase, at least one molecule of 3-arylacrylic acid, at least two molecules of malonyl-CoA, and at least one molecule of caffeyl-CoA.
[0401] In some embodiments, the method involves generating caffeyl-CoA from caffeic acid during an enzymatic reaction catalyzed by coumarate-CoA ligase. In this case, the method involves combining the type III polyketide synthase described above with at least one molecule of caffeic acid, at least one molecule of coenzyme A, at least one molecule of coumarate-CoA ligase, at least one molecule of ATP, and at least two molecules of malonyl-CoA.
[0402] For the purposes of the present invention, any coumarate-CoA ligase enzyme known in the art that carries out the reaction of adding coenzyme A to caffeic acid and forming caffeyl-CoA, as shown in the following formula, can be used.
[0403] [ka]
[0404] In particular, coumarate-CoA ligase 1 derived from Arabidopsis thaliana, having the amino acid sequence and nucleic acid sequence shown in SEQ ID NO: 141, as well as its functional variants and homologs, can be used. For example, for the purposes of the present invention, a functional coumarate-CoA ligase having an amino acid sequence having at least 40% identity with the amino acid sequence shown in SEQ ID NO: 141, for example, at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, for example, at least 80%, at least 85%, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity) can be applied.
[0405] All of the above reactions are carried out under physiological conditions at temperatures of 20-50°C, and these reactions can be performed in cells, tissues, and host organisms that express functional enzymes.
[0406] The PKS and coumarate-CoA ligase combined with the hispidin hydroxylase of the present invention can be used to produce 3-hydroxyhispidin from caffeic acid. This reaction is carried out under physiological conditions in the presence of at least one molecule of hispidin hydroxylase, at least one molecule of PKS, at least one molecule of coumarate-CoA ligase, at least one molecule of caffeic acid or caffeyl-CoA, at least one molecule of coenzyme A, at least one molecule of ATP, at least one molecule of NAD(P)H, at least one molecule of oxygen, and at least two molecules of malonyl-CoA.
[0407] Furthermore, the present invention provides applications of nucleic acids encoding enzymes for fungal luciferin biosynthesis, variants and homologs of these proteins including truncated and elongated forms, and fusion proteins, for obtaining enzymes involved in fungal luciferin biosynthesis in vitro and in vivo.
[0408] In preferred embodiments, the present invention provides nucleic acids encoding hispidine hydroxylase, i.e., proteins characterized by the amino acid sequences shown in SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, and 28, as well as applications of their functional homologs, variants, and derivatives. In a preferred embodiment, the nucleic acid encodes a protein having an amino acid sequence that matches at least 40%, typically at least 45%, usually at least 50%, for example at least 55%, at least 60%, or at least 65%, or at least 70%, at least 80%, at least 85%, at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, the sequence shown in SEQ ID NOs: 29-33.
[0409] The present invention also provides applications of nucleic acids encoding hispidin synthases, i.e., proteins characterized by the amino acid sequences shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55, as well as functional homologs, variants, and derivatives thereof. In preferred embodiments, the nucleic acids of the present invention encode proteins having amino acid sequences that are at least 40%, typically at least 45%, usually at least 50%, for example at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% identical to the sequences shown in SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55 of the entire protein polypeptide chain. In a preferred embodiment, the nucleic acid encodes a protein having an amino acid sequence characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 56-63.
[0410] The present invention also provides applications of nucleic acids encoding caffeylpyruvate hydrolase, i.e., proteins characterized by the amino acid sequences shown in SEQ ID NOs: 65, 67, 69, 71, 73, and 75, as well as functional homologs, variants, and derivatives thereof. In preferred embodiments, the nucleic acids of the present invention encode proteins having amino acid sequences that are at least 40%, typically at least 45%, usually at least 50%, for example at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% identical to the sequences shown in SEQ ID NOs: 65, 67, 69, 71, 73, and 75 of the entire protein polypeptide chain. In a preferred embodiment, the nucleic acid encodes a protein having an amino acid sequence characterized by the presence of several conserved amino acid motifs (consensus sequences) shown in SEQ ID NOs: 76-78.
[0411] The above-mentioned nucleic acid group will be used for the generation of recombinant proteins of hispidin hydroxylase, hispidin synthase, and caffeylpyruvate hydrolase, as well as for the expression of these proteins in heterologous expression systems.
[0412] In particular, nucleic acids encoding hispidin hydroxylase are used to obtain precursor cells of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one from exogenous or endogenous 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula.
[0413] [ka]
[0414] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0415] [ka]
[0416] Nucleic acids encoding caffeylpyruvate hydrolase are applied to obtain cells and organisms that can transform oxyluciferin into preluciferin precursors.
[0417] Nucleic acids encoding hispidin synthase are used to obtain producer cells for the above-mentioned 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one from the corresponding 3-arylacrylic acid. For example, cells expressing hispidin synthase are used to produce hispidin from caffeic acid.
[0418] In some embodiments, nucleic acids encoding hispidin synthase are used to produce hispidin from tyrosine. In these embodiments, nucleic acids encoding an enzyme that promotes the synthesis of caffeic acid from tyrosine are additionally introduced into the cell. Such enzymes are known in the art. For example, as described in [Lin and Yan., Microbial Cell Factories, 2012, Vol. 4; No. 11: p. 42], a combination of nucleic acids encoding tyrosine ammonia lyase from Rhodobacter capsulatus and components HpaB and HpaC of Escherichia coli 4-hydroxyphenylacetate 3-monooxygenase reductase can be used. It will be obvious to those skilled in the art that other enzymes known in the art, such as enzymes that convert tyrosine to caffeic acid, for example, enzymes having an amino acid sequence substantially identical to that of Rhodobacter capsulata tyrosine-ammonia-lyase, as well as components HpaB and HpaC of Escherichia coli 4-hydroxyphenylacetate 3-monooxygenase-reductase shown in SEQ ID NOs: 107, 109, and 111, can be used as substitutes. For example, the enzyme can have an amino acid sequence that has at least 40% identity with the amino acid sequences shown in SEQ ID NOs: 107, 109, and 111, respectively, for example, at least 45% identity, or at least 50% identity, or at least 55% identity, or at least 60% identity, or at least 65% identity, or at least 70% identity, or at least 75% identity, for example, at least 80% identity, at least 85% identity, or at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0419] In some embodiments, nucleic acids encoding hispidin synthase are used to obtain producer cells for hispidin functional analogs from aromatic compounds, including aromatic amino acids and their derivatives. In these embodiments, nucleic acids encoding enzymes that promote the synthesis of 3-arylacrylic acid, from which the hispidin functional analogs are biosynthesized, are additionally introduced into the cells. Such enzymes are known in the art. For example, as described in [Bang, HB, Lee, YH, Kim, SC et al., Microbial Cell Factories, 2016, No. 15: p. 16, https: / / doi.org / 10.1186 / s12934-016-0415-9], nucleic acids encoding phenylalanine-ammonia-lyase from Streptomyces maritimus can be used for cinnamic acid biosynthesis. It will be obvious to those skilled in the art that other enzymes known in the art that convert aromatic amino acids and other aromatic compounds to 3-arylacrylic acid can be used as alternatives. For example, in cinnamic acid biosynthesis, any functional phenylalanine-ammonia-lyase can be used, and its amino acid sequence is nearly identical to the sequence shown in SEQ ID NO: 117, for example, having at least 40% identity with the sequence of SEQ ID NO: 117, for example, a minimum of 45%, or a minimum of 50%, or a minimum of 55%, or a minimum of 60%, or a minimum of 65%, or a minimum of 70%, or a minimum of 75%, for example, a minimum of 80%, a minimum of 85%, or a minimum of 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0420] In some embodiments for obtaining host cells expressing functional hispidin-synthase, it is necessary to cotransfect the host cells with the nucleic acid encoding the hispidin-synthase of the present invention and with the nucleic acid encoding a 4'-phosphopantetheinyltransferase capable of transferring 4'-phosphopantetheinyl from coenzyme A to serine in the acyl carrier domain of polyketide synthase. In other embodiments, selected host cells, such as plant cells or cells of several lower fungi (e.g., Aspergillus), consist of endogenous 4'-phosphopantetheinyltransferase, and cotransfection is not required.
[0421] The present invention also provides applications of nucleic acid combinations. Thus, a combination of nucleic acids encoding hispidin hydroxylase and hispidin synthase can be applied, for example, to produce 3-hydroxyhispidin from caffeic acid and / or tyrosine, in order to obtain producer cells of 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one from 3-arylacrylic acid. In other embodiments, the nucleic acid combination includes a nucleic acid encoding 4'-phosphopantetheinyltransferase. In some embodiments, the nucleic acid combination includes a nucleic acid encoding an enzyme that promotes the synthesis of 3-arylacrylic acid from cellular metabolites, for example, an enzyme that promotes the synthesis of caffeic acid from tyrosine or cinnamic acid from phenylalanine.
[0422] In some embodiments, a combination of nucleic acid, coding PKS, and coumarate-CoA ligase is used to obtain hispidin producer cells from caffeic acid. For the purposes of the present invention, nucleic acids encoding functional PKS are applicable, the amino acid sequence of which is substantially similar to or identical to sequences selected from the group of SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139; for example, the amino acid sequence of PKS is at least 40% typical of sequences selected from the group of SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139. In terms of the components, at least 45%, usually at least 50%, for example, at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% are identical. Furthermore, nucleic acids encoding functional coumarate-CoA ligases that catalyze the reaction of adding coenzyme A to caffeic acid to form caffeyl-CoA are also applicable to the purposes of the present invention. For example, nucleic acids encoding a functional coumarate-CoA ligase having an amino acid sequence identical to the sequence shown in SEQ ID NO: 141, or having a minimum of 40% identity, for example, a minimum of 45% identity, or a minimum of 50% identity, or a minimum of 55% identity, or a minimum of 60% identity, or a minimum of 65% identity, or a minimum of 70% identity, or a minimum of 75% identity, for example, a minimum of 80% identity, a minimum of 85% identity, or a minimum of 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, or 99% identity).
[0423] In some embodiments, a combination of nucleic acids encoding hispidine hydroxylase and PKS is used. In preferred embodiments, the combination also includes nucleic acids encoding coumarate-CoA ligase. The combination is applied to obtain 3-hydroxyhispidin producer cells from caffeic acid and / or caffeyl-CoA.
[0424] In some embodiments, the nucleic acid combination includes a nucleic acid that encodes an enzyme that promotes the synthesis of caffeic acid from tyrosine.
[0425] The nucleic acid combinations of the present invention, used with nucleic acids encoding luciferase capable of oxidizing fungal luciferin with luminescence, are of particular interest. The luciferase-encoding nucleic acid molecules for the purposes of the present invention are obtained from biological sources, for example, by cloning from basidiomycete types, primarily basidiomycete classes, particularly the Agaricales order, or by genetic engineering techniques. Luciferase mutants possessing luciferase activity are obtained using standard molecular biology techniques, as detailed later in the "Nucleic Acids" section. Mutations include changes in one or more amino acids, deletions or insertions, substitutions or truncations of one or more amino acids, or N-terminal truncation or extension, C-terminal truncation or extension, etc. In preferred embodiments, these nucleic acids encode a luciferase having an amino acid sequence that is at least 40% identical, for example, at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, with an amino acid sequence selected from the group SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98. For example, these nucleic acids can have an amino acid sequence that is at least 90% identical (for example, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical) to an amino acid sequence selected from the group SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98. Non-exclusive examples of nucleic acids encoding luciferase are shown in SEQ ID NOs: 79, 81, 83, 85, 87, 89, 91, 93, and 95.
[0426] In some embodiments, a combination of the nucleic acid encoding hispidine hydroxylase and the nucleic acid encoding the luciferase described above is used. This combination is widely applicable to labeling organisms, tissues, cells, organelles, or proteins by bioluminescence. Methods for labeling organisms, tissues, cells, organelles, or proteins with luciferase are well known in the art and, for example, presuppose the introduction of the nucleic acid encoding luciferase into a host cell that is part of an expression cassette that promotes luciferase expression in the cell, tissue, or organism. When a suitable luciferin is added to the cell, tissue, or organism expressing luciferase, detectable luminescence is produced. When labeling organelles or proteins, the nucleic acid encoding luciferase operationally binds to the nucleic acid encoding the localization signal in the target organelle or protein, respectively. In the intracellular co-expression of luciferase and hispidin-synthase of the present invention, biological targets (cells, tissues, organisms, organelles, or proteins) acquire the ability to luminescent not only in the presence of fungal luciferin but also in the presence of preluciferin (the latter being more stable in the presence of ambient oxygen in most cases).
[0427] Furthermore, this nucleic acid combination can be applied to studies on the activity dependence of two promoters in heterologous expression systems. In this case, nucleic acids operationally bound to promoter A encoding luciferase and nucleic acids operationally bound to promoter B encoding hispidin hydroxylase are introduced into host cells. When luciferin, preluciferin, or a mixture of preluciferin and luciferase is added to aliquots of cells (or cell extracts), the activity of one promoter A (luminescence is detected in the presence of luciferin alone), the activity of the other promoter B (luminescence is detected in the presence of a mixture of preluciferin and luciferase), or the activity of both promoters (luminescence is detected in all cases) can be detected by the occurrence of luminescence.
[0428] In some embodiments, the combination also comprises a nucleic acid encoding hispidin synthase. In some embodiments, the combination further comprises a nucleic acid encoding 4'-phosphopantheteinyltransferase.
[0429] In some embodiments, the combination consists of a nucleic acid encoding hispidin-synthase, a nucleic acid encoding luciferase, a nucleic acid encoding PKS, and a nucleic acid encoding coumarate-CoA ligase.
[0430] When labeling organisms, tissues, cells, organelles, or proteins by bioluminescence, these combinations are widely applicable. In this embodiment, to obtain luminescence, a suitable preluciferin precursor, such as caffeic acid or coumaric acid, is added to a biological target expressing hispidin hydroxylase, luciferase, and hispidin synthase, or hispidin hydroxylase, luciferase, PKS, and coumarate-CoA ligase.
[0431] These combinations can also be applied to methods for studying the activity dependence of three promoters in heterologous expression systems. This method assumes the introduction of nucleic acids encoding luciferase under promoter A control, hispidin hydroxylase under promoter B control, and hispidin synthase (or PKS) under promoter B control into host cells. If co-expression of 4'-phosphopantetheinyltransferase is required for the maturation of functional hispidin-synthase, 4'-phosphopantetheinyltransferase is also introduced into the cells under the control of any suitable constituent or inducing promoter. Upon addition of a suitable preluciferin precursor to the cells (or an extract thereof), detectable luminescence appears, indicating simultaneous activation of all three promoters.
[0432] This combination can also be applied to the generation of transgenic bioluminescent organisms. In a preferred embodiment, the transgenic organism is obtained from an organism that is not bioluminescent in its wild type. The nucleic acid encoding the target protein is introduced into the transgenic organism as part of an expression cassette or vector present in the organism as an extrachromosomal element, or is incorporated into the organism genome as described in the "Transgenic Organisms" section above to promote the expression of the target protein. The transgenic organisms of the present invention are unique in that they express at least hispidin hydroxylase, with the exception of luciferase whose substrate is fungal luciferin. In a preferred embodiment, these transgenic organisms also express hispidin synthase. In another preferred embodiment, these transgenic organisms also express PKS. In another preferred embodiment, these transgenic organisms also express PKS. In some embodiments, these transgenic organisms also express coumarate-CoA ligase. Endogenous coumarate-CoA ligase is known to be present in many plant organisms; therefore, if endogenous coumarate-CoA ligase is not present, it should be introduced.
[0433] In some embodiments, these transgenic organisms also express caffeylpyruvate hydrolase. In contrast to organisms that express only luciferase, transgenic organisms obtained using the nucleic acids of the present invention acquire the ability to emit light in the presence of preluciferin and / or preluciferin precursors—3-arylacrylic acid (typically caffeic acid). 3-arylacrylic acid is the cheapest and most stable substrate for obtaining bioluminescence and can be added to water for plant watering, or microbial cultures, or feed, or to animal (e.g., fish) habitats. Bioluminescent transgenic organisms (plants, animals, or fungi) can be applied as light sources and are also used for decorative purposes. Bioluminescent transgenic organisms, cells, and cell cultures can also be used for different screenings in which bioluminescence intensity changes in response to external influences, for example, to analyze the effects of various factors on the activity of promoters that control the expression of exogenous nucleic acids.
[0434] Autonomous bioluminescent transgenic organisms, as provided in this invention, are a special subject.
[0435] In some embodiments, the organism has at least one 3-arylacrylic acid having the following structural formula as a metabolite, where R is aryl or heteroaryl.
[0436] [ka]
[0437] Non-limiting examples include superordinate and subordinate plants, such as flowering plants and mosses. To obtain autonomously bioluminescent transgenic plants, nucleic acids encoding hispidin hydroxylase, hispidin synthase, and luciferase, which are capable of oxidizing fungal luciferin with bioluminescence and expressing the corresponding enzymes, are introduced into these plants. Since plants normally consist of endogenous 4'-phosphopantetheinyltransferase, the additional introduction of nucleic acids encoding this enzyme is generally not necessary to obtain autonomously bioluminescent plants.
[0438] In some embodiments, organisms that do not naturally produce 3-arylacrylic acid are used to obtain autonomous bioluminescent transgenic organisms. Examples of such organisms include animals and various microorganisms, such as yeast and bacteria. In this case, an expressible nucleic acid encoding an enzyme that promotes the biosynthesis of 3-arylacrylic acid from cellular metabolites, such as caffeic acid derived from tyrosine, is additionally introduced into the organism to obtain autonomous bioluminescence. If necessary, a nucleic acid encoding 4'-phosphopantheteinyltransferase is also introduced into the organism.
[0439] In some embodiments for obtaining autonomously luminescent organisms, nucleic acids encoding PKS, hispidin hydroxylase, and luciferase, capable of expressing the corresponding enzyme and oxidizing fungal luciferin with luminescence, are introduced into these organisms. In preferred embodiments, the cells, tissues, and organisms consist of sufficient amounts of caffeyl-CoA and malonyl-CoA to carry out hispidin synthesis.
[0440] If the transgenic organism does not produce a sufficient amount of caffeyl-CoA through its normal metabolic processes, a nucleic acid encoding coumarate-CoA ligase, and, if necessary, an enzyme that biosynthesizes caffeic acid from tyrosine, are also introduced into the cell or organism.
[0441] In a preferred embodiment, the nucleic acid combination for obtaining autonomous bioluminescent cells or transgenic organisms also comprises a nucleic acid encoding caffeyylpyruvate hydrolase. As described in the experimental section below, the expression of caffeyylpyruvate hydrolase consequently increases the bioluminescence intensity of autonomously bioluminescent cells or transgenic organisms. In a preferred embodiment, the bioluminescence intensity increases by at least 1.5 times, typically at least 2 times, usually at least 5 times, for example 7 to 9 times, for example 8 times or more.
[0442] Autonomous bioluminescent transgenic organisms (plants, animals, or fungi), as well as cells and cell structures, differ from transgenic organisms, cells, and cell cultures known in the art that express only luciferase, in that their luminescence does not require the exogenous addition of luciferin or its precursor.
[0443] In some embodiments, instead of a combination of nucleic acids encoding hispidin hydroxylase and luciferase, a nucleic acid encoding a fusion protein of these two enzymes is used. It will be obvious to those skilled in the art that the fusion protein and the combination of nucleic acids encoding hispidin hydroxylase and luciferase are interchangeable in all uses. It will also be obvious that other fusion proteins retaining the properties of the fusion partner can be generated based on the nucleic acids of the present invention; and such fusion proteins and the nucleic acids encoding them can be used without limitation as substitutes for the individual protein and nucleic acid combinations.
[0444] In all of the above applications and methods, nucleic acids can also be in the form of expression cassettes that can be used to promote coding sequence expression in host cells. Nucleic acids can be introduced into host cells as part of a vector for expression in suitable host cells, or in a form not included in a vector, for example, by being incorporated into liposomes or viral particles. Alternatively, purified nucleic acid molecules can be directly incorporated into host cells by suitable means, for example, by direct endocytosis. Genetic components can be directly introduced into host organism cells (e.g., plants) by transfection, infection, microinjection, cell fusion, protoplast fusion, or by using a microparticle bombardment, or by using a "gene gun" (a gun that fires microparticles carrying genetic components).
[0445] Applications of the polyclonal and monoclonal antibodies of the present invention are also provided. These antibodies are used to stain tissues, cells, or organisms, localizing expressed or naturally occurring hispidin hydroxylase, hispidin synthase, and caffeylpyruvate hydrolase of the present invention. Methods of staining using specific antibodies are well known in the art and are described, for example, in [VLBykov, Cytology and General Histology]. Direct immunohistochemical methods are based on the direct reaction of a specifically bound labeled antibody with a detectable substance, while indirect immunohistochemical methods are based on the detection of an unlabeled primary antibody by a secondary labeled antibody after it has bound to the detectable substance, provided that the primary antibody is the antigen of the secondary antibody. Antibodies can also be used to halt enzymatic reactions. Contacting an antibody with a specific binding partner inhibits the enzymatic reaction. Antibodies can also be used in methods for purifying recombinant proteins and natural proteins of the present invention by affinity chromatography. Affinity chromatography techniques are well known in the art, for example, as described in Ninfa et al., Fundamental Laboratory Approaches for Biochemistry and Biotechnology, Wiley, 2009, 2nd edition, p. 133; Cuatrecasasas, JBC, 1970, reprinted November 22, 2017.
[0446] Sets and Products A further embodiment of the present invention is a product comprising the above-mentioned hispidin hydroxylase, or hispidin synthase, or caffeylpyruvate hydrolase, or nucleic acid encoding the enzyme, having an expression vector or cassette comprising an element for promoting the expression of a target protein in a host cell, for example, a nucleic acid encoding the target protein. Alternatively, the nucleic acid may consist of an adjacent sequence for incorporation into the target vector. The nucleic acid may be included in an unpromotered vector for the purpose of simple cloning of the target regulatory element. The recombinant protein may be lyophilized or dissolved in a buffer. The nucleic acid may be lyophilized, precipitated in an alcohol solution, or dissolved in water or a buffer.
[0447] In some embodiments, the product includes cells expressing one or more of the above-mentioned nucleic acids.
[0448] In some embodiments, the product comprises a transgenic organism expressing one or more of the above nucleic acids.
[0449] In some embodiments, the product includes an antibody for staining and / or inhibiting and / or affinity chromatography of the enzyme.
[0450] The product is a container with a label and instructions for use attached to the label. Acceptable containers include, for example, bottles, ampoules, glass tubes, syringes, cell plates, petri dishes, etc. Containers can be made from different materials such as glass or polymer materials. The selection of a suitable container will be obvious to those skilled in the art.
[0451] Furthermore, the product may include other products required from a commercial or consumer perspective, such as: reaction buffers or components for their preparation; buffers or components for their preparation for diluents and / or solutions and / or storage solutions of proteins and nucleic acids; deionized water; secondary antibodies against specific antibodies of the present invention; cell culture media or components for their preparation; and nutrients for transgenic organisms.
[0452] The product also includes instructions for implementing the proposed method. The instructions can be attached to the product in one or more forms, and can take various forms, provided, for example, that the instructions are available as an electronic file and / or on paper.
[0453] The present invention also relates to a kit applicable to a variety of purposes. The kit may preferably include a combination of proteins or nucleic acids of the present invention, using an expression vector or cassette consisting of elements for promoting the expression of a target protein in a host cell, such as a nucleic acid encoding the target protein. In some embodiments, the kit may also include a nucleic acid encoding a luciferase that can oxidize fungal luciferin with luminescence. In some embodiments, the kit may also include a nucleic acid encoding an enzyme involved in the biosynthesis of caffeic acid from tyrosine. In some embodiments, the kit may also include a nucleic acid encoding 4'-phosphopantetheinyltransferase. In some embodiments, the kit may also include a nucleic acid encoding PKS. In some embodiments, the kit may also include a nucleic acid encoding coumarate-CoA ligase.
[0454] In some embodiments, the kit may also include antibodies for purifying recombinant proteins or staining proteins expressed in host cells. In some embodiments, the kit may also include primers complementary to the region of the nucleic acid for amplifying the nucleic acid or a cross-section thereof. In some embodiments, the kit may also include one or more fungal luciferins and / or preluciferins and / or preluciferin precursors. The compounds may be in the form of a dry powder, an organic solvent solution, or an aqueous solution. In some embodiments, the kit may include cells comprising one or more of the above nucleic acids. In some embodiments, the kit may include the transgenic organism of the present invention, such as a producer strain or a transgenic autonomous bioluminescent plant. All components of the kit are placed in a suitable container. Generally, the kit also includes instructions for use.
[0455] The following examples are provided to better illustrate the present invention. These examples are provided for illustrative purposes only and should not be construed as limiting the scope of the present invention in any way.
[0456] All publications, patents, and patent applications described herein are incorporated herein by reference. Although the above inventions have been described in considerable detail by illustration and illustrative purposes for clarity, it will be obvious to those skilled in the art, based on the ideas disclosed herein, that several changes and modifications can be introduced without departing from the spirit and scope of the proposed embodiments of the invention. [Examples]
[0457] Example 1. Isolation of hispidin hydroxylase sequence Total RNA was isolated from the mycelium of *Lycoperdon perlatum* according to the method described in [Chomczynski and Sacchi, Analytical Biochemistry, 1987, No. 162, pp. 156-159]. cDNA was amplified using the SMART PCR cDNA Synthesis Kit (Clontech, USA) according to the manufacturer's protocol. The obtained cDNA was used to amplify the luciferase coding sequence. The nucleotide and amino acid sequences are shown in SEQ ID NOs: 79 and 80. The coding sequence was cloned into a pGAPZ vector (Invitrogen, USA) according to the manufacturer's protocol and transformed into competent *E. coli* XL1 Blue strain cells. The bacteria were cultured in a petri dish in the presence of the antibiotic zeosin. After 16 hours, the colonies were rinsed from the petri dish, thoroughly mixed, and plasmid DNA was isolated from the colonies using a plasmid DNA isolation kit (Evrogen, Russia). The isolated plasmid DNA was linearized using the restriction site AvrII and used to transform Pichia pastris GS115 cells. Electroporation was performed according to the method using lithium acetate and dithiothreitol described in [Wu and Letchworth, Biotechniques, 2004, No. 36: pp. 152-4]. The electroporated cells were dispersed in a petri dish containing RDB medium consisting of 1 M sorbitol, 2% (weight / volume) glucose, 1.34% (weight / volume) yeast nitrogen base (YNB), 0.005% (weight / volume) amino acid mixture, 0.00004% (weight / volume) biotin, and 2% (weight / volume) agar. The resulting colonies were sprayed with 3-hydroxyhispidine solution, and the presence of intracellular luciferase was detected by the generation of luminescence. The luminescence emitted by the colonies was detected using IVIS Spectrum CT (PerkinElmer, USA). Colonies that showed luminescence in response to the addition of 3-hydroxyhispidine were selected for further study.
[0458] Next, the total cDNA amplified from *Lycoperdon perforatum* was cloned into a pGAPZ vector and transformed into competent *E. coli* XL1 Blue strain cells. The bacteria were cultured in a petri dish in the presence of the antibiotic zeosin. After 16 hours, the colonies were rinsed from the petri dish, thoroughly mixed, and plasmid DNA was isolated from the colonies using a plasmid DNA isolation kit (Evrogen, Russia). The isolated plasmid DNA was linearized using the restriction site AvrII and used for the transformation of *Pichia pastris* GS115 yeast cells. Transformation was performed by electroporation as described above. The cells were dispersed in a petri dish containing RDB medium consisting of 1 M sorbitol, 2% (weight / volume) glucose, 1.34% (weight / volume) yeast nitrogen bases (YNB), 0.005% (weight / volume) amino acid mixture, 0.00004% (weight / volume) biotin, and 2% (weight / volume) agar. The diversity of the obtained *Lycoperdon perlatum* cDNA library in yeast was approximately 1 million clones. Hispidin solution was sprayed onto the obtained colonies, and the presence of hispidin hydroxylase in the cells was detected by the occurrence of luminescence. The luminescence emitted by the colonies was detected using IVIS Spectrum CT (PerkinElmer, USA). Cells expressing only luciferase and wild yeast cells were used as negative controls. During library screening, colonies showing luminescence were selected and used as a matrix for PCR with standard plasmid primers. The PCR products were sequenced using the Sanger method to determine the sequence of the expressed gene. The sequence of the obtained hispidin hydroxylase nucleic acid is shown in SEQ ID NO: 1. The amino acid sequence encoded by that nucleic acid is shown in SEQ ID NO: 2.
[0459] Figure 4 shows the luminescence of Pichia pastris cells expressing hispidin hydroxylase and luciferase, or Pichia pastris cells expressing only luciferase, or wild yeast, when colonies were sprayed with 3-hydroxyhispidine (luciferin) and hispidin (preluciferin). The data demonstrate that luciferin was produced intracellularly in the presence of hispidin hydroxylase.
[0460] In the next step, genomic DNA was isolated from the fungi Armillaria mellea, Armillaria japonica, Armillaria tinctoria, Armillaria mellea, Guyanagastar necrohiza, Misena citricolor, Mycena chlorophos, Pyrocotyle fuciformis, Neonotepanus gardneri, Omphalotus olearius, and Wasabi mushrooms. Whole genome sequencing was performed using Illumina HiSeq technology (Illumina, USA) according to the manufacturer's recommendations. The sequencing results were used to predict the amino acid sequence of a hypothetical protein and to search for hispidin hydroxylase homologs from Pyrocotyle fuciformis. Homophila searches were performed using software provided by the National Center for Biotechnology Information (NCBI). Amino acid sequences were searched in the fungal genome sequencing data of the NCBI Genbank database. The standard search parameter blastp was used for the search. As a result, sequences of hispidine hydroxylase homologs derived from *Lycoperdon perlatum*, *Armillaria mellea*, *Armillaria mellea*, *Guyanagastar necrohiza*, *Misena citricolor*, *Neonotepanus gardneri*, *Omphalotus olearius*, *Armillaria mellea*, *Armillaria gracilis*, and *Mycena chlorophos* were identified.
[0461] The nucleotide and amino acid sequences of the hispidin-synthase homologs of *Lycoperdon perlatum* are shown in Sequence IDs 3-28.
[0462] All the identified enzymes are nearly identical to each other. The degree of amino acid sequence identity is shown in Table 4.
[0463] [Table 11]
[0464] Several highly homologous hispidin hydroxylase amino acid sequences, characterized by single amino acid substitutions, were isolated from *Pholiota rupestris* and *Misena citricolor*. The nucleotide and amino acid sequences of these hispidin hydroxylase amino acid sequences are shown in SEQ ID NOs. 7-13 (*Pholiota rupestris*) and SEQ ID NOs. 15-18 (*Misena citricolor*). Further examination of the protein properties revealed no effect of these substitutions on the enzyme properties.
[0465] The coding sequences of the detected homologs (SEQ ID NOs: 3-28) were cloned into Pichia pastris GS115 cells constitutively expressing *Pichia leucocarpa* luciferase according to the protocol described above, and the cells were transformed. The resulting colonies were sprayed with a hispidin solution, and the presence of intracellular hispidin hydroxylase was detected by the generation of luminescence. The luminescence emitted by the colonies was detected using IVIS Spectrum CT (PerkinElmer, USA). All colonies expressing the test genes (SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27) showed luminescence 1,000 to 100 million times greater than that of control cells when sprayed with a hispidin solution, confirming the ability of the enzyme encoded by the test genes to catalyze the transformation of hispidin into 3-hydroxyhispidin (fungal luciferin).
[0466] Structural analysis of the amino acid sequence of the detected enzyme was performed. Analysis using SMART (Simple Modular Architecture Research Tool), software available online at the website http: / / smart.embl-heidelberg.de [Schultz et al., Proceedings of the National Academy of Sciences (PNAS), 1998; No. 95: pp. 5857-5864; Letunic I, Doerks T, and Bork P, Nucleic Acids Research, 2014; doi: 10.1093 / nar / gku949] revealed that the entire detected protein consists of a FAD / NAD(P) binding domain, IPR002938 — a code from the InterPro public database available online at the website http: / / www.ebi.ac.uk / interpro —. This domain is involved in the binding of FAD and NAD in several enzymes, particularly monooxygenases — a representative of a large family of enzymes that add hydroxyl groups to substrates — and in several organisms found in metabolic pathways. Excluding the FAD / NAD(P) binding domain, the detected hispidin hydroxylase consists of N-terminal and C-terminal amino acid sequences that are operationally bound to the hispidin hydroxylase (Figure 1). Multiple alignment and comparison of the detected hispidin hydroxylase amino acid sequences (Figure 1) revealed that the amino acid sequences of these hispidin hydroxylases consist of several conserved amino acid motifs (consensus sequences) (SEQ ID NOs: 29-33) that are characteristic of this enzyme group alone. The consensus sites within the amino acid sequences are operationally bound via amino acid inserts. [Examples]
[0467] Example 2. Expression of hispidin hydroxylase and fungal luciferase in mammalian cells, and their combined use for cell labeling. The coding sequences of hispidin hydroxylase and luciferase derived from *Lycoperdon perlatum*, obtained according to Example 1, were optimized (humanized) for expression in mammalian cells. The optimized nucleic acids (SEQ ID NOs: 99 and 100) were synthetically obtained. The coding sequence of hispidin hydroxylase was cloned into a pmKate2-keratin vector (Evrogen, Russia) using restriction sites NheI and NotI instead of the sequence encoding the fusion protein mKate2-keratin. The luciferase sequence was amplified by PCR, treated with restriction endonucleases NheI and EcoRV (New England Biolabs, Ipswich, Massachusetts), and ligated into the lentiviral vector pRRLSIN.cPPT.EF1. Rasmid DNA was purified using a plasmid DNA purification kit (Evrogen). Using plasmid DNA consisting of the luciferase gene, a stable expression strain HEK293NT was developed. Vector particles were obtained by calcium-phosphate transfection of HELK293T cells (Invitrogen, Carlsbad, California) according to the protocol provided on the manufacturer's website. 1,500,000 cells were placed in a 60 mm culture dish 24 hours prior to transfection. For transfection, approximately 4 μg and 1.2 μg of packaging plasmids pR8.91 and pMD.G, and 5 μg of a transfer plasmid consisting of the luciferase sequence were used. Virus particles were collected 24 hours after transfection, enriched 10-fold, and used for transduction of HEK293NT cells. Approximately 100% of HEK293NT cells stably expressed *Luciferum erythrorhizon* luciferase.
[0468] The obtained cells were re-transfected with a vector consisting of the hispidin hydroxylase coding sequence using the transfection reagent FuGENE HD (Promega, USA) according to the manufacturer's protocol. 24 hours after transfection, hispidin at a concentration of 800 μg / ml was added to the culture medium, and cell luminescence was detected using IVIS Spectrum CT (PerkinElmer). The obtained cells emitted luminescence at an intensity more than two orders of magnitude higher than the signal emitted from untransfected control cells (Figure 5).
[0469] Cells were visualized using transmitted light in a green emission detection channel. When human cells were expressed with *Hispidin hydroxylase*, a clear emission signal appeared in the green spectrum in the presence of hispidin, allowing us to distinguish between transformed and untransformed cells. [Examples]
[0470] Example 3. Use of hispidin hydroxylase and hispidin analog in cell lysate. HEK293NT cells expressing luciferase and hispidin hydroxylase of *Hikaritake* obtained according to Example 2 were rinsed from the petri dish 24 hours after transfection with Barzen's solution supplemented with 0.025% trypsin, the medium was replaced with pH 8.0 phosphate-buffered saline by centrifugation, the cells were resuspended, and lysed by sonication at 0°C for a maximum of 7 minutes using a Bioruptor (Diagenode, Belgium) under manufacturer-recommended conditions. Then, 1 mM NADPH (Sigma-Aldrich, USA) and one of hispidin or its analogs were added at a concentration of 660 μg / ml as follows: The following compounds were added to the culture medium: (E)-4-hydroxy-6-(4-hydroxystyryl)-2H-pyran-2-one, (E)-6-(2-(1H-indole-3-yl)vinyl)-4-hydroxy-2H-pyran-2-one, (E)-6-(2-(1,2,3,5,6,7-hexahydropyrido[3,2,1-ij]quinoline-9-yl)vinyl)-4-hydroxy-2H-pyran-2-one, (E)-6-(4-(diethylamino)styryl)-4-hydroxy-2H-pyran-2-one, or (E)-4-hydroxy-6-(2-(6-hydroxynaphthalene-2-yl)vinyl)-2H-pyran-2-one. The bioluminescence spectra were detected using a Varian Cary Eclipse spectrophotometer. Luminescence was observed in the lysate when all of the above hispidin functional analogs were added. Depending on the type of luciferin used, a displacement of the expected emission peak was observed. [Examples]
[0471] Example 4. Obtaining recombinant hispidine hydroxylase A polyhistidine sequence (His tag) was operationally added to the 5' end of the nucleic acids encoding hispidin-3-hydroxylase and luciferase derived from *Hikaritake* mushrooms, obtained according to Examples 1 and 2. The resulting structures were cloned into the pET-23 vector using restriction endonucleases BamHI and HindIII. This vector was used to transform *E. coli* cells of strain BL21-DE3. These cells were dispersed in a petri dish containing LB medium consisting of 1.5% agar and 100 μg / ml ampicillin and incubated overnight at 37°C. The *E. coli* colonies were then transferred to 4 ml of liquid LB medium containing ampicillin and incubated overnight at approximately 37°C. 1 ml of the culture incubated overnight was transferred to 100 ml of Overnight Express Autoinduction medium (Novagen) pre-added with ampicillin. The culture medium was incubated at 37°C for up to 2.5 hours until an optical density of 0.6 OE was reached at 600 nm, followed by incubation at room temperature for up to 16 hours. The cells were then pelleted in an Eppendorf 5810R centrifuge at 4500 rpm for up to 20 minutes and resuspended in 35 ml of buffer (50 mM Tris HCl ρH8.0, 150 mM NaCl). The cells were sonicated and pelleted again. TALON resin metal affinity chromatography (Clontech, USA) was used to purify recombinant proteins. The presence of the predicted recombinant product was confirmed by electrophoresis.
[0472] Aliquotes of isolated recombinant hispidin hydroxylase were used to test functionality and stability. To determine functionality, 15 μl of the isolated recombinant protein solution was placed in a glass tube containing 100 μl of buffer (0.2 M Na-phosphate buffer, 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM, pH 8.0), 0.5 μl purified recombinant luciferase from *Lycoperdon perlatum*, 1 mM NADPH, and 0.2 μM hispidin). The glass tube was placed in a luminometer. The activity of the isolated recombinant protein yielded luminescence in the presence of *Lycoperdon perlatum* luciferase in combination with hispidin and its analogs as described in Example 3. In all examples, the luminescence intensity was highest when using *Lycoperdon perlatum* hispidin hydroxylase and lowest when using *Armillaria mellea* hispidin hydroxylase. [Examples]
[0473] Example 5. Acquisition of 3-hydroxyhispidin, (E)-3,4-dihydroxy-6-styryl-2H-pyran-2-one and (E)-3,4-dihydroxy-6-(4-hydroxystyryl)-2H-pyran-2-one using recombinant hispidin hydroxylase. Recombinant hispidin hydroxylase isolated from *Lycoperdon perlatum* obtained according to Example 4 was added to a reaction mixture containing 1 mM NADPH, 0.2 μM hispidin, and (E)-4-hydroxy-6-styryl-2H-pyran-2-one or (E)-4-hydroxy-6-(4-hydroxystyryl)-2H-pyran-2-one in 100 μl of buffer (0.2 M Na-phosphate buffer, 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM) pH 8.0). After 30 minutes, the reaction mixture was analyzed by HPLC using synthetic luciferin as a standard. Chromatography demonstrated the generation of peaks corresponding to the 3-hydroxylated derivatives: 3-hydroxyhispidine, (E)-3,4-dihydroxy-6-styryl-2H-pyran-2-one, and (E)-3,4-dihydroxy-6-(4-hydroxystyryl)-2H-pyran-2-one. [Examples]
[0474] Example 6. Detection of bioluminescence by hispidin hydroxylase and luciferase fusion protein The humanized DNA sequences encoding hispidine hydroxylase and luciferase from *Hikaritake* obtained according to Example 2 were operably crosslinked with a flexible short-chain peptide linker having the amino acid sequence GGSGGGS (SEQ ID NO: 115). The nucleotide and amino acid sequences of the resulting fusion protein are shown in SEQ ID NOs: 101 and 102. The nucleic acid encoding the fusion protein was cloned into the pEGFP-N1 vector (Clontech, USA) in place of the EGFP gene under the control of a cytomegalovirus promoter. The resulting structure was transfected into HEK293T cells. Analog vectors consisting of the individual hispidine hydroxylase and luciferase genes were also co-transfected. Transfection was performed using the transfection agent FuGENE HD (Promega, USA) according to the manufacturer's protocol. Twenty-four hours after transfection, one million cells were resuspended in 0.5 ml of PBS, and luminescence was recorded using a luminometer with and without hispidin addition (10 μg per million cells). Hispidin addition resulted in cellular luminescence in the green spectrum (Figure 6). Bioluminescence signals were also induced with the addition of 3-hydroxyhispidin. Because the expression of a hispidin hydroxylase and luciferase fusion protein allows the use of a more stable luciferin precursor (hispidin, bisnoriangonin, etc.) instead of a single luciferase for cellular bioluminescence labeling, co-transfection of cells with two nucleic acids is unnecessary. [Examples]
[0475] Example 7. Preparation of polyclonal antibodies The coding sequences for hispidine hydroxylase from *Lycoperdon perlatum* (SEQ ID NO: 1) and *Armillaria mellea* (SEQ ID NO: 19) were synthetically obtained in the form of linear double-stranded DNA, and cloned into the expression vector pQE-30 (Qiagen, Germany) so that the recombinant protein contained a histidine tag at the N-terminus. After expression in *E. coli*, the recombinant protein was purified under denaturing conditions using the metal-affinity resin TALON (Clontech, USA). The purified protein product emulsified with Freund's adjuvant was used for rabbit immunization four times at one-month intervals. Blood was collected from the rabbits on day 10 or 11 post-immunization. The activity of the obtained polyclonal antiserum was tested by ELISA and Western immunoblotting in a panel of purified recombinant hispidine hydroxylase obtained according to Example 4.
[0476] Antibodies obtained by immunizing rabbits with proteins derived from *Lycoperdon perlatum* showed activity against denatured and undenatured hispidin hydroxylase of *Lycoperdon perlatum* and denatured hispidin hydroxylase of *Neonotepanus gardneri*. Furthermore, antibodies obtained by immunizing rabbits with Armillaria mellea protein showed activity against denatured and undenatured hispidin hydroxylase from Armillaria mellea, Armillaria gracilis, Armillaria tinctoria, and Armillaria mellea. [Examples]
[0477] Example 8. Obtaining transgenic plants expressing *Hikaritake* hispidin hydroxylase and luciferase. The coding sequences for *Physcomitrella patens* hispidin hydroxylase and luciferase were optimized for expression in moss cells of *Physcomitrella patens*. Next, an expression cassette was created in silico consisting of the rice aktI gene promoter, the human cytomegalovirus 5'-untranslated region encoding a hispidin hydroxylase sequence optimized for expression in plant cells (SEQ ID NO: 103), a stop codon, the Agrobacterium osc gene terminator sequence, the rice ubiquitin promoter, the coding sequence for *Physcomitrella patens* luciferase optimized for expression in moss cells (SEQ ID NO: 112), and the Agrobacterium tumefaciens nos gene terminator.
[0478] The obtained sequences were synthesized so that all the fragments appeared to be operationally cross-linked with one another, and then cloned into the expression vector pLand#1 (Institut Jean-Pierre Bourgin, France) located between DNA fragments corresponding to the locus of the moss genome DNA of Physcomitrella patens between the highly expressed moss genes Pp3c16_6440V3.1 and Pp3c16_6460V3.1 using Gibson assembly technique [Gibson et al., Nature Methods, 2009, Vol. 6: pp. 343-5]. Vector pLand#1 also contained a Cas9 nuclease guide RNA (sgRNA) sequence complementary to the same DNA locus region.
[0479] Following the polyethylene glycol transformation protocol described in [Cove et al., Cold Spring Harbor Protocols, 2009, No. 2], plasmid DNA products were co-transformed into moss protoplasts of Physcomitrella patens with an expression vector consisting of a Cas9 nuclease sequence under the Arabidopsis thaliana ubiquitin promoter. The protoplasts were then cultured in BCD medium at approximately 50 rpm under dark conditions for 2 days to regenerate the cell wall. Subsequently, the protoplasts were transferred to a petri dish consisting of agar and BCD medium and cultured under 16 hours of illumination for up to 1 week. Transformed moss colonies were screened by PCR using external genome primers to determine the progress of genome integration of gene constructs, transferred to fresh petri dishes, and cultured under the same illumination conditions for up to 30 days.
[0480] The obtained moss gametophytes were immersed in BCD medium containing hispidin at a concentration of 900 μg / ml and analyzed using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). All transgenic plants analyzed showed bioluminescence intensity at least two orders of magnitude higher than that of control plants expressing only luciferase incubated in the same solution as hispidin. [Examples]
[0481] Example 9. Identification of hispidin synthase and caffeylpyruvate hydrolase Fungal luciferin precursors such as hispidin belong to a large group of compounds—polyketide derivatives. These compounds are theoretically derived from 3-arylacrylic acid, where the aryl or heteroaryl aromatic substituents are located at the 3-position. It is known in the art that the enzymes involved in the synthesis of polyketides and their derivatives are multidomain complexes associated with the polyketide synthase protein superfamily. At the same time, no polyketide synthase capable of catalyzing the transformation of 3-arylacrylic acid to substituted 4-hydroxy-2H-pyran-2-one is known in the art. Target polyketide synthases were searched for using a screening of the *Hikaritake* cDNA library.
[0482] It is known that to obtain functional polyketide synthase in heterologous yeast systems, it is necessary to additionally introduce a gene expressing 4'-phosphopantetheinyltransferase, an enzyme that transfers 4-phosphopantetheinyl from coenzyme A to serine within the acyl carrier domain of polyketide synthase, into the culture medium [Gao Menghao et al., Microbial Cell Factories, 2013, Vol. 12: p. 77]. The NpgA gene for 4'-phosphopantetheinyltransferase from Aspergillus nidurans (SEQ ID NOs. 104, 105), which is known in the art, was synthetically obtained and cloned into a pGAPZ vector. The plasmid was linearized at restriction site AvrII and used to transform Pichia pastrius GS115 yeast strain constitutively expressing the white-flowered hyacinth cDNA obtained according to Example 1. The diversity of the obtained library of white-flowered hyacinth cDNA in yeast was approximately 1 million clones.
[0483] A Pichia pastris cDNA library expressed in the Pichia pastris yeast strain was obtained according to the protocol given in Example 1 and used for the identification of hispidin synthase and caffeylpyruvate hydrolase. Cells were dispersed in a petri dish containing RDB medium consisting of 1 M sorbitol, 2% (weight / volume) glucose, 1.34% (weight / volume) yeast nitrogen base (YNB), 0.005% (weight / volume) amino acid mixture, 0.00004% (weight / volume) biotin, and 2% (weight / volume) agar.
[0484] The obtained colonies were sprayed with a caffeic acid solution (potentially a hispidin precursor), and the presence of intracellular hispidin-synthase was detected by the occurrence of luminescence. The luminescence emitted by the colonies was detected using IVIS Spectrum CT (PerkinElmer, USA). Cells expressing only luciferase and hispidin hydroxylase, as well as wild yeast cells, were used as negative controls. When screening the library, colonies in which luminescence was detected were selected and used as a matrix for PCR with standard plasmid primers. The PCR products were sequenced by Sanger sequencing to determine the sequences of the expressed genes. The sequence of the obtained hispidin-synthase nucleic acid is shown in SEQ ID NO: 34. The amino acid sequence encoded by that nucleic acid is shown in SEQ ID NO: 35.
[0485] Next, using the Pichia pastris yeast strain obtained, consisting of the gene for Pichia pastris luciferase, hispidin hydroxylase, and hispidin synthase integrated into the genome, and the NpgA gene of 4'-phosphopantetheinyltransferase from Aspergillus nidurans, we identified the enzymes that catalyze the transformation of oxyluciferin ((2E,5E)-6-(3,4-dihydroxyphenyl)-2-hydroxy-4-oxohexa-2,5-dienoic acid) to caffeate. The cell line was re-transformed with a linearized plasmid library of Pichia pastris genes obtained in the first step of the process. The resulting colonies were sprayed with caffeoylpyruvic acid solution, and the presence of the target enzyme in the cells was detected by the generation of luminescence. The luminescence emitted by the colonies was detected using IVIS Spectrum CT (PerkinElmer, USA). Cells expressing only luciferase and hispidin hydroxylase, as well as wild yeast cells, were used as negative controls. During library screening, colonies exhibiting luminescence were selected and used as a matrix for PCR with standard plasmid primers. The PCR products were sequenced using the Sanger assay to determine the sequence of the expressed gene. The sequence of the isolated enzyme nucleic acid is shown in SEQ ID NO: 64. The amino acid sequence encoded by this nucleic acid is shown in SEQ ID NO: 65. The identified enzyme was named caffeylpyruvate hydrolase. [Examples]
[0486] Example 10. Identification of hispidin synthase and caffeylpyruvate hydrolase homologs of *Lycoperdon perlatum*. Using whole-genome sequencing data from bioluminescent fungi obtained according to Example 1, homologs of *Hipyrus erythrosora* hispidin synthase and caffeylpyruvate hydrolase were searched. Homogenetic searches were performed using software provided by the National Center for Biotechnology Information (NCBI). Amino acid sequences were searched in fungal genome sequencing data in the NCBI Genbank database. The standard search parameter blastp was used for the search.
[0487] The sequences of hispidin-synthase homologs derived from *Lycoperdon perlatum*, *Armillaria mellea*, *Armillaria mellea*, *Guyanagastar necrohiza*, *Misena citricolor*, *Neonotepanus gardneri*, *Omphalotus olearius*, *Armillaria mellea*, *Armillaria gracilis*, and *Mycena chlorophos* were identified. These nucleotide and amino acid sequences are shown in Sequence IDs 36-55. All identified enzymes were nearly identical to each other. The degree of amino acid sequence identity is shown in Table 5.
[0488] [Table 12]
[0489] Two highly homologous hispidin-synthase amino acid sequences, characterized by single amino acid substitutions, were isolated from the wasabi mushroom. The nucleotide and amino acid sequences of these hispidin-synthase amino acid sequences are shown in Sequence IDs 36-39.
[0490] The identified enzymes were tested for their ability to transform caffeic acid into hispidin using the technique described in Example 9.
[0491] By multiple alignment of the identified protein amino acid sequences, several highly homologous fragments of typical amino acid sequences for this enzyme group were identified. The consensus sequences of these fragments are shown in sequence numbers 70-77. As shown in Figure 2, these sequences are separated by long-chain amino acid sequences.
[0492] The caffeylpyruvate hydrolase homologous sequences of *Lycoperdon perlatum* were identified in *Neonotepanus gardneri*, *Armillaria mellea*, *Armillaria holosteoides*, *Armillaria mellea*, and *Armillaria gracilis*. The nucleotide and amino acid sequences of the identified homologs are shown in SEQ ID NOs: 66-75. The identified enzymes were tested for their ability to transform caffeoylpyruvate into caffeic acid using the technique described in Example 9.
[0493] All identified enzymes are nearly identical to each other and have amino acid lengths of 280–320. The degree of amino acid sequence identity is shown in Table 6.
[0494] [Table 13]
[0495] Analysis using the SMART (Simple Modular Architecture Research Tool) software, available online at the website http: / / smart.embl-heidelberg.de [Schultz et al., Proceedings of the National Academy of Sciences, 1998; No. 95: pp. 5857-5864; Letunic I, Doerks T, and Bork P, Nucleic Acids Research, 2014; doi: 10.1093 / nar / gku949] revealed that the detected whole protein consists of a fumarylacetase domain (EC3.7.1.2) approximately 200 amino acids long located near the C-terminus. However, the conserved region, according to the amino acid numbering of caffeylpyruvate hydrolase in *Hikaritake*, begins at approximately 8 amino acids. Multiple alignment made it possible to identify typical consensus sequences (SEQ ID NOs: 76-78) for this group of proteins separated by amino acid inserts with low identity. The locations of the consensus sequences are shown in Figure 3. [Examples]
[0496] Example 11. Obtaining recombinant hispidin synthase and caffeylpyruvate hydrolase, and using them to obtain bioluminescence. According to Example 9, a polyhistidine (His tag) coding sequence was operationally added to the 5' end of the nucleic acids encoding hispidin synthase and caffeylpyruvate hydrolase of *Hilippines*, and the resulting structure was cloned into the pET-23 vector using restriction endonucleases NotI and SacI. This vector was used to transform *E. coli* cells of the BL21-DE3-codon+ strain by electroporation. These transformed cells were dispersed in a petri dish containing LB medium consisting of 1.5% agar and 100 μg / ml ampicillin and incubated overnight at 37°C. Next, the *E. coli* colonies were transferred to 4 ml of liquid LB medium containing 100 μg / ml ampicillin and incubated overnight at approximately 37°C. 1 ml of the culture incubated overnight was transferred to 200 ml of Overnight Express Autoinduction medium (Novagen) with ampicillin pre-added. The culture medium was incubated at 37°C for up to 3 hours until it reached an optical density of 0.6 OE at 600 nm, and then incubated at room temperature for up to 16 hours. The cells were then pelleted in an Eppendorf 5810R centrifuge at 4500 rpm for up to 20 minutes, resuspended in 20 ml of buffer (50 mM Tris HCl ρH8.0, 150 mM NaCl), dissolved by sonication in a Bioruptor (Diagenode, Belgium) at 0°C for up to 7 minutes under manufacturer-recommended conditions, and pelleted again. Proteins were obtained from the lysate by Talon resin affinity chromatography (Clontech, USA). Since bands of the predicted length were available, the presence of the predicted recombinant product was confirmed by electrophoresis.
[0497] Aliquots of isolated recombinant proteins were used to test their functionality and stability.
[0498] To determine the functionality of hispidin synthase, 30 μl of isolated recombinant protein solution was placed in a glass tube containing 100 μl of buffer (0.2 M Na-phosphate buffer, 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM) pH 8.0, all components from Sigma-Aldrich, USA), 0.5 μl of purified recombinant luciferase from *Lycoperdon perlatum* obtained according to Example 4, 1 mM NADPH (Sigma-Aldrich, USA), 15 μl of purified recombinant hispidin hydroxylase from *Lycoperdon perlatum* obtained according to Example 4, 10 mM ATP (ThermoFisher Scientific, USA), 1 mM CoA (Sigma-Aldrich, USA), and 1 mM malonyl-CoA (Sigma-Aldrich, USA). The glass tube was placed in a GloMax 20 / 20 luminometer (Promega, USA). When 20 μM caffeic acid was added to a solution (Sigma-Aldrich, USA), the reaction mixture exhibited bioluminescence. The maximum wavelength of the emission was 520–535 nm.
[0499] To determine the functionality of caffeylpyruvate hydrolase, 10 μl of isolated recombinant protein solution was placed in a glass tube containing 100 μl of buffer (0.2 M Na-phosphate buffer, 0.5 M Na2SO4, 0.1% dodecylmaltoside (DDM) pH 8.0), 0.5 μl of *Lycoperdon perlatum* luciferase, 1 mM NADPH (Sigma-Aldrich, USA), 15 μl of hispidin hydroxylase, 10 mM ATP (ThermoFisher Scientific, USA), 1 mM CoA (Sigma-Aldrich, USA), 1 mM malonyl-CoA (Sigma-Aldrich, USA), and 30 μl of purified recombinant hispidin synthase. The glass tube was then placed in a GloMax 20 / 20 luminometer (Promega, USA). When 25 μM caffeoylpyruvic acid was added to the solution, bioluminescence was detected in the reaction mixture, serving as an indicator of the ability of the test enzyme to decompose caffeoylpyruvic acid into caffeic acid. The maximum wavelength of the emission was 520–535 nm.
[0500] The enzyme obtained was used to obtain luminescence (bioluminescence) in the reaction of the white-flowered mushroom lysiferase and hispidin hydroxylase obtained according to Example 4. 5 μl of each isolated recombinant protein solution was placed in a glass tube containing 100 μl of buffer (0.2 M Na-phosphate buffer, 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM) pH 8.0), 1 mM NADPH (Sigma-Aldrich, USA), 10 mM ATP (ThermoFisher Scientific, USA), 1 mM CoA (Sigma-Aldrich, USA), 1 mM malonyl-CoA (Sigma-Aldrich, USA), and 0.2 μM of one type of 3-arylacrylic acid: paracoumaric acid (Sigma-Aldrich, USA), coumaric acid (Sigma-Aldrich, USA), or ferulic acid (Abcam, USA). In other experiments, instead of 3-arylacrylic acid, analogues of fungal oxyluciferin—(2E,5E)-2-hydroxy-6-(4-hydroxyphenyl)-4-oxohexa-2,5-diene, (2E,5E)-2-hydroxy-4-oxo-6-phenylhexa-2,5-diene, or (2E,5E)-2-hydroxy-6-(4-hydroxy-3-methoxyphenyl)-4-oxohexa-2,5-dienoic acid—were also placed in glass tubes at a concentration of 0.2 μM. The glass tubes were placed in a luminometer. The activity of the isolated recombinant protein resulted in luminescence in each of the described reactions. [Examples]
[0501] Example 13. Obtaining hispidin from caffeic acid An expression cassette consisting of nucleic acids encoding *Hikaritake* hispidin-synthase (SEQ ID NOs: 34, 35) under the control of the J23100 promoter, and an expression cassette consisting of the NpgA gene (SEQ ID NOs: 104, 105) of *Aspergillus nidurans* 4'-phosphopantetheinyltransferase, under the control of the araBAD promoter, with loxP introduced via the homologous region of the SS9 site, were synthetically obtained and cloned into a bacterial expression vector consisting of a zeosin-resistant cassette. The obtained structures were used for transformation and integration into the *E. coli* BW25113 genome by lambda bacteriophage protein-mediated recombination as described by Bassalo et al. [ACS Synthetic Biology, July 15, 2016; Vol. 5(7): p. 561-8], utilizing zeosin resistance selection. After confirming the integration of the full-length structure by PCR using primers specific to the SS9 homologous region, the appropriateness of the integrated structure was verified by sequencing of the PCR product of genomic DNA using the Sanger method.
[0502] The obtained E. coli strain was used for hispidin production. In the first step, the bacteria were incubated in LB medium at approximately 200 rpm at 37°C for a maximum of 10 hours in five 50 ml plastic tubes. 250 ml of the obtained culture was added to 3.3 liters of fermentation medium so that the initial culture optical density at 600 nm was approximately 0.35, and the mixture was placed in a Biostat B5 fermenter (Braun, Germany). The fermentation medium contained 10 g / l peptone, 5 g / l caffeic acid, 5 g / l yeast extract, 10 g / l NaCl, 25 g / l glucose, 15 g / l (NH4)2SO4, 2 g / l KH2PO4, 2 g / l MgSO4·7H2O, 14.7 mg / l CaCl2, 0.1 mg / l thiamine, and 1.8 mg / l and 0.1% solutions consisting of the following: 8 mg / The solution contained 1 mg / l EDTA, 2.5 mg / l CoCl2·6H2O, 15 mg / l MnCl2·4H2O, 1.5 mg / l CuCl2·2H2O, 3 mg / l H3BO3, 2.5 mg / l Na2MoO4·2H2O, 13 mg / l Zn(CH3COO)2·2H2O, 100 mg / l ferric citrate, and 4.5 mg / l thiamine hydrochloride. Fermentation was carried out at 37°C with aeration at 3 L / min and mixing at 200 rpm. After 25 hours of incubation, arabinose was added to the culture medium until the final concentration was 0.1 mM. NH4OH was added to automatically control the pH and lower it to 7.0. A solution consisting of 500 g / l glucose, 5 g / l caffeic acid, 2 g / l arabinose, 25 g / l tryptone, 50 g / l yeast extract, 17.2 g / l MgSO4·7H2O, 7.5 g / l (NH4)SO4, and 18 g / l ascorbic acid was added to a fermenter, and the glucose level was maintained each time the pH rose to 7.1. After 56 hours of incubation, the hispidin concentration in the culture medium was 1.23 g / l. Hispidin in the fermentation medium and hispidin purified from the medium by HPLC were active in bioluminescence reactions with *Hipylissericea* hispidin hydroxylase and luciferase. [Examples]
[0503] Example 13. Acquisition of 3-hydroxyhispidine from caffeic acid An expression cassette consisting of nucleic acids encoding *Lyophyllum sarmentosum* hispidin hydroxylase (SEQ ID NOs: 1, 2), controlled by the J23100 promoter, was synthetically obtained and cloned into a bacterial expression vector consisting of a spectinomycin resistance gene. The obtained vector was transformed into *E. coli* cells expressing *Lyophyllum sarmentosum* hispidin synthase, zeosin resistance gene, and NpgA gene, as obtained according to Example 12. Using the obtained bacterial cells, 3-hydroxyhispidin was produced by fermentation according to the protocol described in Example 12, with spectinomycin added to all culture media at a concentration of 50 mg / ml. After 48 hours of incubation, the concentration of 3-hydroxyhispidin in the culture medium was 2.3 g / l. 3-hydroxyhispidin from the fermentation medium and purified from the medium by HPLC was active in the bioluminescence reaction with *Lyophyllum sarmentosum* luciferase. [Examples]
[0504] Example 14. Obtaining hispidin from cell metabolites and tyrosine To synthesize hispidin from tyrosine, we obtained an E. coli strain that effectively produces tyrosine and caffeic acid. The E. coli strain was obtained as described in [Lin and Yan., Microbial Cell Factories, April 4, 2012; No. 11: p. 42]. The E. coli strain BW25113, in which a mutant gene for acY (lacY A177C) permease was incorporated into the attB site, which enables uniform consumption of arabinose by bacterial cells, was used as the basis for strain development. Expression cassettes consisting of the Rhodobacter capsulata tyrosine-ammonia-lyase genes (SEQ ID NOs: 106, 107) and the coding sequences of the E. coli 4-hydroxyphenyl acetate 3-monooxygenase reductase components HpaB and HpaC (SEQ ID NOs: 108-111) were synthetically obtained under the control of the constitutive J23100 promoter and incorporated into the genome of the E. coli strain as described in Example 12. In the next step, a plasmid consisting of the coding sequence of *Hikaritake* hispidin-synthase, obtained according to Example 12 and under the control of the constitutive J23100 promoter, a zeosin-resistant cassette derived from a pGAP-Z vector, and the NpgA gene were incorporated into the *E. coli* genome. Integration into the *E. coli* genome was performed by lambda bacteriophage protein-mediated recombination according to the method described in [Bassalo et al., ACS Synth Biol., 2016; Vol. 5(7): p. 561-568]. After confirming the integration of the full-length structure by PCR using primers specific to the SS9 homologous regions (5'-CGGAGCATTTTGCATG-3' and 5'-TGTAGGATCAAGCTCAG-3'), the appropriateness of the integrated structure was verified by sequencing of the PCR product of the genomic DNA using the Sanger method. The resulting strain was used for biosynthetic hispidin production in a fermenter.
[0505] As described in Example 12, bacteria were cultured in a fermenter, the only difference being that caffeic acid was not added to the bacterial culture medium. Biosynthetic hispidin was isolated from the medium by HPLC. The obtained strain was able to produce 1.20 mg / l of hispidin per 50 hours of fermentation. The purity of the obtained product was 97.3%. Furthermore, by adding tyrosine at a concentration of 10 g / ml to the culture medium, the amount of hispidin produced could be increased to 108.3 mg / ml. [Examples]
[0506] Example 15. Growth of the autonomous bioluminescent yeast Pichia pastris To promote the growth of the autonomous bioluminescent yeast Pichia pastris, expression cassettes were synthesized under the control of the GAP promoter and tAOX1 terminator, consisting of coding sequences for Pichia nidurans cyferase (SEQ ID NOs: 79, 80), Pichia nidurans hispidin hydroxylase (SEQ ID NOs: 1, 2), Pichia nidurans hispidin synthase (SEQ ID NOs: 34, 35), Pichia nidurans caffeypyruvate hydrolase (SEQ ID NOs: 64, 65), Aspergillus nidurans NpgA protein (SEQ ID NOs: 104, 105), Rhodobacter capsulata tyrosine ammonia lyase (SEQ ID NOs: 106, 107), and HpaB and HpaC (SEQ ID NOs: 108-111), components of Escherichia coli 4-hydroxyphenyl acetate 3-monooxygenase reductase. Each expression cassette was modified with loxP via a BsmBI restriction enzyme recognition sequence. Homologous regions to the MET6 Pichia pastris gene (Uniprot, F2QTU9) with loxP introduced at the BsmBI restriction site were also synthetically obtained. The synthetic DNA was treated with BsmBI restriction enzyme and ligated to a single plasmid according to the Golden Gate cloning protocol described in [Iverson et al., ACS Synthetic Biology, January 15, 2016; Vol. 5(1): p. 99-103]. 10 fmol of each DNA fragment was mixed into a total volume of 10 μl reaction solution consisting of standard strength buffer for DNA ligase (Promega, USA), 20 units of DNA ligase activity (Promega, USA), and 10 units of DNA restriction endonuclease activity. The resulting reaction mixture was placed in an amplifier and incubated at 16°C and 37°C according to the following protocol: 25 cycles of incubation at 37°C for up to 1.5 minutes and 16°C for 3 minutes, followed by a single incubation at 50°C for up to 5 minutes, followed by a single incubation at 80°C for up to 10 minutes. 5 μl of the reaction mixture was used to transform chemically competent E. coli cells. The appropriateness of the plasmid DNA assembly was confirmed by Sanger assay, and the purified plasmid DNA product was used for electroporation-induced transformation of Pichia pastris GS11 cells.Electroporation was performed according to the method using lithium acetate and dithiothreitol described in [Wu and Letchworth, Biotechniques, 2004, No. 36: pp. 152-4]. Electroporated cells were dispersed in a petri dish containing RDB medium consisting of 1 M sorbitol, 2% (weight / volume) glucose, 1.34% (weight / volume) yeast nitrogen base (YNB), 0.005% (weight / volume) amino acid mixture, 0.00004% (weight / volume) biotin, and 2% (weight / volume) agar. Integration of the gene cassette into the genome was confirmed by PCR from primers annealed in homologous regions. The resulting yeast strains containing the correct genome insert were able to luminescent autonomously, in contrast to wild yeast strains (Figures 7 and 8). [Examples]
[0507] Example 16. Growth of an autonomous bioluminescent flowering plant To promote the growth of autonomous bioluminescent flowering plants using the pBI121 vector (Chlontech, USA), a binary vector for Agrobacterium transformation was constructed, consisting of the coding sequences of Agrobacterium cyferase (SEQ ID NO: 112), Agrobacterium cyferase (SEQ ID NO: 103), Agrobacterium cyferase (SEQ ID NO: 113), Agrobacterium caffeypyruvate hydrolase (SEQ ID NO: 114), and a kanamycin resistance gene, all optimized for expression in plants. Each gene is under the control of a 35S promoter derived from cauliflower mosaic virus. The expression cassette assembly sequences were obtained synthetically. The vector was assembled according to the Golden Gate cloning protocol described in [Iverson et al., ACS Synthetic Biology, January 15, 2016; Vol. 5(1):p.99-103].
[0508] Arabidopsis thaliana was transformed by co-culturing plant tissue with the AGL0 strain of Agrobacterium tumefaciens, consisting of a created binary vector [Lazo et al., Biotechnology, October 1991; No. 9(10): p.963-7]. Transformation was performed using co-culturing of Arabidopsis thaliana root segments (C24 ecotype), as described in [Valvekens et al., Proceedings of the National Academy of Sciences, (USA), 1988, No. 85, p.5536-5540]. Arabidopsis thaliana roots were cultured for up to 3 days in agar-treated Gamborg medium B-5 containing 20 g / l glucose, 0.5 g / l 2,4-dichlorophenoxyacetic acid, and 0.05 g / l kinetin. Subsequently, the roots were cut to a length of 0.5 cm and transferred to 10 ml of liquid Gamborg medium B-5 containing 20 g / l glucose, 0.5 g / l 2,4-dichlorophenoxyacetic acid, and 0.05 g / l kinetin, and 1.0 ml of Agrobacterium overnight culture solution was added. The explants and Agrobacterium were co-cultured for up to 2-3 minutes. Then, the explants were placed on a sterile filter in a petri dish containing agar medium of the same composition. After incubation in a 25°C incubator for 48 hours, the explants were transferred to fresh medium containing 500 mg / l cefotaxime and 50 mg / l kanamycin. After 3 weeks, plant regeneration was started in a selective medium consisting of 50 mg / l kanamycin. The transgenic plants developed roots and were transferred to germination medium or soil. Bioluminescence was visualized using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). Over 90% of the transgenic plants emitted luminescence at least two orders of magnitude higher than the signals from wild-type plants.
[0509] Tobacco benthamiana was transformed by co-culturing plant tissue with Agrobacterium tumefaciens strain AGL0, which consisted of a prepared binary vector [Lazo et al., Biotechnology, October 1991; No. 9(10): p.963-7]. Transformation was performed using co-culturing of Tobacco benthamiana segments. Next, leaves were cut to a length of 0.5 cm and transferred to 10 ml of liquid Gamborg medium B-5 containing 20 g / l glucose, 0.5 g / l 2,4-dichlorophenoxyacetic acid, and 0.05 g / l kinetin, and 1.0 ml of Agrobacterium overnight culture solution was added. The explants and Agrobacterium were co-culturned for a maximum of 2-3 minutes. Then, the explants were placed on a sterile filter in a petri dish containing agar medium of the same composition. After incubation in a 25°C incubator for 48 hours, the explants were transferred to fresh culture medium containing 500 mg / l cefotaxime and 50 mg / l kanamycin. Three weeks later, plant regeneration was initiated in a selective medium consisting of 50 mg / l kanamycin. The transgenic plants developed roots and were transferred to germination medium or soil. Bioluminescence was visualized using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). Over 90% of the transgenic plants emitted luminescence at least two orders of magnitude higher than the signal from wild-type plants. A photograph of the autoluminescent tobacco plant is shown in Figure 9.
[0510] For the purpose of cultivating the bioluminescent Agrostis stolonifera L., the coding sequences of fungal luciferin metabolism cascade genes optimized for expression in plants and into which loxP was introduced via a BsaI restriction endonuclease site—specifically, Agrostis stolonifera luciferase (SEQ ID NO: 126), Agrostis stolonifera hispidin hydroxylase (SEQ ID NO: 117), Agrostis stolonifera hispidin (SEQ ID NO: 127), Agrostis stolonifera caffeyylpyruvate hydrolase (SEQ ID NO: 128), and the herbicide glyphosate resistance gene (bar gene)—were cloned into the pBI121 vector (Chlontech, USA). Each sequence was under the control of the CmYLCV promoter [Stavolone et al., Plant Molecular Biology, November 2003; No. 53(5): p.663-73]. The sequences were synthesized according to standard techniques. Vectors were assembled according to the Golden Gate cloning protocol. Transformation was performed using the embryonic callus Agrobacterium transformation method. The overnight culture solution of Agrobacterium tumefaciens strain AGL0, consisting of the prepared binary vector [Lazo et al., Biotechnology, October 1991; No. 9(10): p.963-7] was added to a liquid medium. After co-culturing for 2 days in agar-based Murashige and Skoog medium, the plants were transferred to a fresh medium containing 500 mg / l cefotaxime and 10 mg / l phosphinothricin. Plant regeneration was initiated after 3 weeks. Transgenic plants were transplanted into a medium containing half the salinity of Murashige and Skoog and supplemented with 8 mg / l phosphinothricin to induce rooting. The rooted plants were then placed in a greenhouse. Approximately 25% of the plants that were properly and completely integrated into the resulting metabolic cascade genome emitted bioluminescence exceeding that of the control wild-type plants.
[0511] Currently, organisms capable of bioluminescence in specific tissues or at specific times are of special interest. Such organisms consume the resources necessary for bioluminescence more efficiently. To cultivate autonomous bioluminescent roses that emit light only in their petals, several varieties of roses with white petals were selected. Based on the pBI121 vector (Clontech, USA), two binary vectors for Agrobacterium transformation were created, consisting of metabolic cascades derived from the coding sequences of *Luciferum erythrorhizon* luciferase, *Luciferum erythrorhizon* hispidin hydroxylase, *Luciferum erythrorhizon* hispidin synthase, *Luciferum erythrorhizon* caffeyylpyruvate hydrolase, and the neomycin resistance gene, optimized for expression in plants. All genes except the luciferase gene were placed under the control of the cauliflower mosaic virus 35S promoter. In one vector, the luciferase gene was placed under the control of the rose chalcone synthase promoter, and in the other vector, it was placed under the control of the chrysanthemum chalcone UEP1 promoter. Using synthetic nucleic acids necessary for vector assembly, loxP was introduced at the BsaI restriction enzyme recognition site, and the vector was assembled according to the Golden Gate cloning protocol. By co-culturing Agrobacterium tumefaciens strain AGL0, consisting of the above binary vector, with embryonic callus, a transgenic plant of Rosa hybrida L. cv. Tinike was obtained [Lazo et al., Biotechnology, October 1991; No. 9(10): p.963-7]. Culturing was carried out for up to 40 minutes in a liquid medium consisting of macro and micro salts of Murashige-Skoog, supplemented with 1-2 mg / l kinetin, 3 mg / l 2,4-dichlorophenoxyacetic acid, and 1 mg / l 6-benzylaminopurine. The callus was transferred to an agar medium of the same composition. Two days later, the explants were transferred to fresh Murashige-Skoog medium containing 500 mg / l cefotaxime and 50 mg / l kanamycin. Bud formation and regeneration appeared after 5–8 weeks. The buds were transferred to propagation medium or rooting medium. Rooted buds were placed in a peat mixture in a greenhouse. Flowering was observed after 8 weeks.Plants bearing mature flowers were visualized using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). All test plants for each test structure autonomously emitted luminescence at least three orders of magnitude higher than the signal from wild-type plants. The luminescence was emitted only from the petal tissue, confirming the tissue-specific function of the promoter.
[0512] To cultivate self-bioluminescent plants whose bioluminescence is controlled by circadian rhythms and activated at night, a binary vector for Agrobacterium transformation consisting of pre-obtained coding sequences for *Agrobacterium sibiricum* luciferase, *Agrobacterium sibiricum* hispidin hydroxylase, *Agrobacterium sibiricum* hispidin synthase, *Agrobacterium sibiricum* caffeyylpyruvate hydrolase, and a neomycin resistance gene was used. Each gene is under the control of a 35S promoter derived from cauliflower mosaic virus. The promoter for *Agrobacterium sibiricum* luciferase expression was replaced with the promoter of the *Arabidopsis thaliana*-derived CAT3 gene. Transcription from the CAT3 gene promoter is controlled by circadian rhythms and activated at night. The CAT3 promoter sequence is publicly known in the art [Michael and McClung, Plant Physiology, October 2002; No. 130(2): p. 627-38]. Arabidopsis thaliana was transformed by co-culturing plant tissue with Agrobacterium tumefaciens strain AGL0, consisting of a created binary vector [Lazo et al., Biotechnology, October 1991; No. 9(10): p.963-7]. Transformation was performed using the co-culturing of Arabidopsis thaliana root segments (C24 ecotype) described in [Valvekens et al., Proceedings of the National Academy of Sciences, USA, 1988, No. 85, p.5536-5540]. Roots were cut to a length of 0.5 cm and transferred to 10 ml of liquid Gamborg medium B-5 containing 20 g / l glucose, 0.5 g / l 2,4-dichlorophenoxyacetic acid, and 0.05 g / l kinetin, with 1.0 ml of Agrobacterium overnight culture solution added. Explants and Agrobacterium were co-cultured for up to 2-3 minutes. The explants were then placed on a sterile filter in a petri dish containing agar medium of the same composition. After incubation at 25°C for 48 hours, the explants were transferred to fresh medium containing 500 mg / l cefotaxime and 50 mg / l kanamycin. After 3 weeks, plant regeneration was initiated in a selective medium consisting of 50 mg / l kanamycin. The transgenic plants developed roots, were transferred to germination medium, and grown under natural daylight cycle conditions.Bioluminescence was visualized using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). Plants were placed in the system for 24 hours, and bioluminescence intensity was recorded every 30 minutes. The plants emitted light within 24 hours, but the bioluminescence intensity was greatly modulated by the circadian rhythm: in 85% of the tested plants, the integrated luminescence intensity at night was more than 1000 times greater than the integrated luminescence intensity during the day. [Examples]
[0513] Example 17. Growth of transgenic autonomous bioluminescent hypoplants Autonomously bioluminescent moss *Physcomitrella patens* was grown by protoplast co-transformation with a plasmid using the method described in Example 8. An expression cassette containing the coding sequences for *Physcomitrella patens* luciferase (SEQ ID NO: 112), *Physcomitrella patens* hispidin hydroxylase (SEQ ID NO: 103), *Physcomitrella patens* hispidin synthase (SEQ ID NO: 113), *Physcomitrella patens* caffeypyruvate hydrolase (SEQ ID NO: 114), and the kanamycin resistance gene, optimized for expression in plants, was synthetically obtained. Each enzyme is under the control of the rice actin 2 promoter. The expression cassette was operationally crosslinked to the pBI121 vector (Clontech, CSHA) so that loxP was introduced into the structure containing the entire metabolic cascade and the kanamycin resistance gene at a sequence matching the moss genome target locus. The vector was assembled according to the Golden Gate cloning protocol [Iverson et al., ACS Synthetic Biology, January 15, 2016; Vol. 5(1): p.99-103]. A guide RNA gene complementary to the target region in the moss genome was also cloned into the vector. The plasmid containing the identified gene was co-transformed with a plasmid for constitutive expression of Cas9 nuclease according to the polyethylene glycol transformation protocol described in [Cove et al., Cold Spring Harbor Protocols, 2009, Vol. 2]. The resulting transformed protoplasts were incubated in BG-11 medium for up to 24 hours under dark conditions and transferred to a petri dish containing BG-11 medium and 8.5% agar. One month after growth in the petri dish, visualization was performed using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer) under continuous illumination. 70% of the test plants emitted luminescence at least an order of magnitude higher than the signal from wild-type plants. [Examples]
[0514] Example 18. Growth of transgenic bioluminescent animals A transgenic zebrafish (Danio rerio), consisting of the gene for *Hispidin hydroxylase*, was created according to the technique described in [Hisano et al., Scientific Reports, 2015, Vol. 5: p. 8841]. This technique involves the expression of guide RNA and the expression of Cas9 nuclease to create a break point in a region homologous to the guide RNA sequence. For the purpose of growing transgenic animals, a synthetic DNA fragment consisting of a pX330 plasmid, a guide RNA sequence derived from Addgene#42230, and Cas9 nuclease mRNA was prepared under the control of the bacteriophage polymerase T7 promoter. The resulting fragment was used for in vitro transcription using reagents from the MAXIscript T7 kit (Life Technologies, USA), and the synthesized RNA was purified using a DNA isolation kit (Evrogen, Russia).
[0515] The coding sequence for *Hisano et al., Scientific Reports, 2015, No. 5: p. 8841*, into which loxP was introduced, was synthetically obtained using a 50-nucleotide sequence derived from the krtt1c19e zebrafish gene, and cloned into a pEGFP / C1 plasmid base consisting of a pUC replication origin and a kanamycin-resistant cassette. The obtained vector, Cas9 nuclease mRNA, and guide RNA were dissolved in injection buffer (40 mM HEPES (pH 7.4), 240 mM KCl, with 0.5% phenol red added), and injected at a volume of approximately 1-2 nl into 1-2 cell embryos of a pre-existing zebrafish strain that stably expresses *Hisano et al., Scientific Reports, 2015, No. 5: p. 8841). Of the 48 embryos, approximately 12 survived the injection and showed normal development on day 4 after fertilization.
[0516] Zebrafish larvae were intravenously injected with a hispidin solution to record bioluminescence signals according to the technique described in [Cosentino et al., Journal of Visualized Experiments, 2010; (42): p.2079]. Bioluminescence was recorded using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). After recording, genomic DNA was isolated from the larvae, and the integration of hispidin hydroxylase into the genome was confirmed. All larvae in which the *Hippodium verum* hispidin hydroxylase gene was correctly integrated into the genome showed bioluminescence intensity at least two orders of magnitude higher than the signal emitted from wild-type fish after hispidin solution injection. [Examples]
[0517] Example 19. Investigation of the effect of caffeyl vinate hydrolase on the luminescence of autonomous bioluminescent organisms. To investigate the effect of caffeylpyruvate hydrolase on the luminescence of autonomous bioluminescent organisms, a binary vector for Agrobacterium transformation was used, consisting of the coding sequences for *Lyophyllum sarmentosum* luciferase, *Lyophyllum sarmentosum* hispidin hydroxylase, *Lyophyllum sarmentosum* hispidin synthase, *Lyophyllum sarmentosum* caffeylpyruvate hydrolase, and a kanamycin resistance gene. Each gene was under the control of a 35S promoter derived from cauliflower mosaic virus obtained according to Example 16, and the control vector was characterized by the removal of the caffeylpyruvate hydrolase sequence from the test vector. Using these vectors, *Arabidopsis thaliana* was transformed under the same conditions according to the protocol described in Example 16. Bioluminescence was visualized using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). When comparing the bioluminescence intensity of plants expressing all four genes of the *Lycoperdon perlatum* bioluminescence system with plants expressing only luciferase, hispidin hydroxylase, and hispidin synthase, it was found that plants additionally expressing caffeylpyruvate hydrolase exhibited bioluminescence that was, on average, 8.3 times brighter. The provided data suggests that caffeylpyruvate hydrolase enables an increase in the efficiency of the bioluminescence cascade, thereby increasing the intensity of the luminescence emitted by plants. [Examples]
[0518] Example 20. Effect of external addition of caffeic acid on bioluminescence of transgenic organisms Autonomous bioluminescent transgenic plants, *Benthamia japonica*, obtained according to Example 16, were transplanted into soil and cultured for up to 8 weeks. The plant stems were then cut and placed in water for 2 hours, after which the bioluminescence intensity was measured using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). The plants were then transferred to one of five aqueous solutions with caffeic acid concentrations of 0.4 g / l, 0.8 g / l, 1.6 g / l, 3.2 g / l, and 6.4 g / l, while control plants were placed in water. After a further 2 hours of incubation in either the caffeic acid solution or water, the bioluminescence intensity was measured again. In all cases, the bioluminescence intensity of plants incubated in caffeic acid solution increased compared to the intensity before placement in the caffeic acid solution, with the greatest change observed in plants incubated in the 6.4 g / l solution. Control plants incubated in water did not show a significant change in bioluminescence intensity within 4 hours of the start of incubation. [Examples]
[0519] Example 21. Promoter activity assay and use of fungal bioluminescent system genes for intracellular logical integration of external signals. The coding sequences for *Lycoperdon perforatum* hispidin hydroxylase, hispidin synthase, and luciferase were used to monitor the simultaneous activation of several promoters. The synthetic expression cassette consisted of the coding sequences for *Lycoperdon perforatum* hispidin synthase (SEQ ID NOs: 34, 35) under the control of the *E. coli* araBAD promoter induced by arabinose, the coding sequences for hispidin hydroxylase (SEQ ID NOs: 1, 2) under the control of the T7 / lacO promoter induced by IPTG, the luciferase gene (SEQ ID NOs: 79, 80) under the control of the pRha promoter induced by rhamnose, and the NpgA gene (SEQ ID NOs: 104, 105) under the control of the constitutive J23100 promoter (Registry of Standard Biological parts, Part: BBa_J23100). The obtained synthetic nucleic acid was cloned into the MoClo_Level2 vector [Weber et al., PLoS One, February 18, 2011, Vol. 6(2):p.e16768] using BpiI restriction endonuclease instead of an insert consisting of the LacZ gene. The resulting vector was transformed into competent E. coli BL21 (NEB, USA) strain cells consisting of a genome copy of T7 bacteriophage polymerase.
[0520] To determine the possibility of recording the simultaneous activation of several promoters, the cells obtained in the preceding step were grown overnight in a flask containing 100 ml of LB medium supplemented with 100 mg / l ampicillin. The following day, cell culture aliquots were placed in one of the following media compositions at 24°C and 200 rpm for 120 minutes. LB medium supplemented with 1.1% arabinose, LB medium supplemented with 2.0.2% rhamnose, LB medium supplemented with 3.0.5% IPTG, LB medium supplemented with 4.1% arabinose and 0.2% rhamnose, LB medium supplemented with 5.1% arabinose and 0.5% IPTG, LB medium supplemented with 6.0.2% rhamnose and 0.5% IPTG, LB medium supplemented with 7.1% arabinose, 0.2% rhamnose, and 0.5% IPTG. 8. LB medium (control).
[0521] After incubation, the cells were pelleted and the culture medium was replaced with pH 7.4 phosphate-buffered saline (Sigma-Aldrich, USA) supplemented with 1 g / l caffeic acid (Sigma-Aldrich, USA), and the cells were resuspended by pipetting. After 30 minutes, cellular bioluminescence was analyzed using a GloMax 20 / 20 luminometer (Promega, USA). The experiment was repeated three times. In eight test samples, the bioluminescence intensity of bacteria incubated in medium No. 7 (LB medium supplemented with 1% arabinose, 0.2% rhamnose, and 0.5% IPTG) differed significantly from that of test bacteria incubated only in LB medium (medium No. 8). Therefore, bacterial luminescence served as an indicator that three different promoters are reliably activated simultaneously when bacteria are introduced into the culture medium. In this experiment, bacterial cells incorporated information about the presence of substances, induced promoter activity in the external medium, and transmitted signals via luminescence only when all three substances were simultaneously present in the medium, performing an "AND" logic operation within the cell.
[0522] A synthetic expression cassette consisting of (1) the coding sequences of hispidin hydroxylase (SEQ ID NOs: 1, 2) under the control of the Odf2 promoter according to [Pletz et al., Biochimica et Biophysica Acta, June 2013; No. 1833(6): p.1338-46], (2) the coding sequences of hispidin synthase (SEQ ID NOs: 34, 35) under the control of the cyclin-dependent kinase CDK7 promoter, and (3) luciferase (SEQ ID NOs: 79, 80) under the control of the CCNH gene promoter was cloned into a pmKate2-keratin vector (Evrogen, Russia) instead of the cytomegalovirus promoter abd mKate2-keratin insert sequence. In addition, the coding sequences of the NpgA gene (SEQ ID NOs: 104, 105) were cloned into the pmKate2-keratin vector instead of the mKate2-keratin insert sequence. All obtained vectors were co-transfected into HEK293T cells using the transfection agent FuGENE HD (Promega, USA) according to the manufacturer's protocol. 24 hours after transfection, caffeate at a concentration of 5 mg / ml was added to the culture medium, and cell luminescence was detected using a Leica TCS SP8 microscope. The luminescence allowed for the identification of simultaneous activation of the Odf2, CCNH, and CDK7 promoters, and the luminescence intensity was related to the cell cycle stage.
[0523] The data obtained suggests that fungal bioluminescence system genes can be used to monitor the simultaneous activation of several promoters, detect the presence and combinations of different substances in the culture medium, and integrate external signals into the intracellular logic. [Examples]
[0524] Example 22. Identification of hispidin in plant extracts The coding sequences of *Hispidin hydroxylase* and luciferase obtained according to Example 1 were cloned into the pET23 vector under the control of the T7 promoter. The purified plasmid DNA product was used for in vitro protein transcription and translation using the PURExpress In Vitro Protein Synthesis Kit (NEB, USA). By adding 2 μl of plant lysate to 100 μl of the reaction mixture and recording the luminescence intensity with a GloMax luminometer (Promega, USA), approximately 19 plant species were identified: chrysanthemum (Chrysanthemum sp.), pineapple (Ananas comosus), petunia atkinsiana, European spruce (Picea abies), European nettle (Urtica dioica), tomato (Solanum lycopersicum), benthamiana tobacco (Nicotiana benthamiana), tobacco (Nicotiana tobacum), Arabidopsis thaliana, Rosa glauca, Rosa rubiginosa, horsetail (Equisetum arvense), and Equisetum thermaeia. The obtained reaction mixtures were used to analyze the presence and concentration of hispidin and its functional analogues in the lysis solutions of horsetail (Equisetum telmateia), Polygala sabulosa, Rosa rugosa, Clematis tashiroi, Kalanchoe sp., Triticum aestivum, and carnation (Dianthus caryophyllus). The highest concentrations of hispidin and its functional analogues were determined to be present in the lysis solutions of horsetail and Equisetum telmateia. Hispidin or its functional analogues were also identified in Polygala sabulosa, Rosa rugosa, and Clematis tashiroi. [Examples]
[0525] Example 23. Identification of PKS capable of catalyzing hispidin synthesis, and use of PKS for in vitro and in vivo hispidin production. Fungal luciferin precursors such as hispidin are related to the polyketide derivative group. Enzymes involved in polyketide synthesis in plants belong to the polyketide synthase protein superfamily, and it is known in the art that plant polyketide synthases, in contrast to fungal polyketide synthases, are relatively small proteins using CoA ethers of acids including 3-arylacrylic acid. Furthermore, although no polyketide synthase capable of catalyzing the transformation of hispidin with caffeic acid CoA ether is known in the art, hispidin is present in many plant organisms.
[0526] Using bioinformatics analysis, we selected 11 polyketide synthases from the following sources that have the potential to catalyze hispidin synthase. Aquilaria sinensis (2 types of enzymes) Hydrangea macrophylla, Arabidopsis thaliana, Physcomitrella patens Japanese knotweed (Polygonum cuspidatum), Rheum palmatum Rheum tataricum, Wachendorfia thyrsiflora, Kava (Piper methysticum) (containing two types of enzymes).
[0527] We optimized selected nucleotide sequences for expression in Pichia pastris yeast cells and Tobacco benthamiana plant cells. The resulting nucleic acids were synthesized, cloned into pGAPZ vectors, and used to verify the hispidin synthesis ability of the expressed proteins.
[0528] For this purpose, in the genome of Pichia pastris GS115 yeast strains constitutively expressing Pichia leucosurase and hispidin hydroxylase obtained according to Example 1, a pGAPZ plasmid consisting of the gene for Arabidopsis coumarate-CoA ligase 1 (its nucleotide and amino acid sequences are shown in SEQ ID NOs: 140 and 141) was also introduced during oligonucleotide synthesis. This plasmid was linearized at restriction site AvrII and used for transformation into Pichia pastris GS115 cells.
[0529] Yeast cells constitutively expressing *Hypholoma fasciculare* lysiferase, hispidin hydroxylase, and *Arabidopsis thaliana* coumarate-CoA ligase 1 were linearized using a plasmid consisting of a PKS coding sequence and dispersed in a petri dish containing RDB medium consisting of 1 M sorbitol, 2% (weight / volume) glucose, 1.34% (weight / volume) yeast nitrogen base (YNB), 0.005% (weight / volume) amino acid mixture, 0.00004% (weight / volume) biotin, and 2% (weight / volume) agar. To identify enzymes with hispidin synthase activity, caffeic acid solution was sprayed onto the resulting colonies, and the presence of intracellular hispidin synthase was detected by the generation of luminescence. The luminescence emitted by the colonies was detected using IVIS Spectrum CT (PerkinElmer, USA). Yeast strains constitutively expressing luciferase, hispidin hydroxylase, and coumarate-CoA ligase 1, as well as wild yeast cells, were used as negative controls. Of the tested genes, 11 enzymes possessed hispidin synthase activity, and their sequences are shown at SEQ ID NOs: 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, and 138. Next, the encoded amino acid sequences are shown at SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, and 139, respectively. The enzymes derived from PKS1 and PKS2 obtained from Aquilaria sinensis (SEQ ID NOs: 119, 121), Arabidopsis thaliana (SEQ ID NO: 123), and Hydrangea macrophylla (SEQ ID NO: 125) showed the highest activity.
[0530] Recombinant proteins were generated using nucleic acids encoding PKS derived from hydrangea (SEQ ID NOs: 124, 125) according to the technique described in Example 4. Since bands of the predicted length were available, the presence of the predicted recombinant product was confirmed by electrophoresis. Aliquots of isolated recombinant proteins were used for functional validation: 30 μl of the isolated recombinant protein solution was placed in a glass tube containing 100 μl of buffer (0.2 M Na-phosphate buffer, 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM) pH 8.0, all components from Sigma-Aldrich, USA), 0.5 μl of purified recombinant luciferase from *Lycoperdon perlatum* obtained according to Example 4, 1 mM NADPH (Sigma-Aldrich, USA), 15 μl of purified recombinant hispidin hydroxylase from *Lycoperdon perlatum* obtained according to Example 4, 10 mM ATP (ThermoFisher Scientific, USA), 1 mM CoA (Sigma-Aldrich, USA), and 1 mM malonyl-CoA (Sigma-Aldrich, USA). The glass tube was placed in a GloMax 20 / 20 luminometer (Promega, USA). When 20 μM caffeyl-CoA was added to the solution, the reaction mixture exhibited bioluminescence. The maximum wavelength of the emission was 520–535 nm.
[0531] Nucleic acids encoding PKS2 from Aquilaria sinensis (SEQ ID NOs: 120, 121) were used to generate hispidin producer strains. An expression cassette consisting of nucleic acid SEQ ID NO: 120 under the control of the constitutive J23100 promoter, and an expression cassette consisting of nucleic acid SEQ ID NO: 140 encoding 4-coumarate-CoA ligase 1 from Arabidopsis thaliana under the control of the araBAD promoter were synthesized; loxP was introduced into both expression cassettes via a homologous region at the SS9 site. These expression cassettes were cloned into a bacterial expression vector consisting of a zeosin-resistant cassette and used for transformation and integration into the E. coli BW25113 genome by lambda bacteriophage protein-mediated recombination, as described in Bassalo et al. [ACS Synthetic Biology, July 15, 2016; Vol. 5(7): p. 561-8], utilizing zeosin resistance selection. After confirming the incorporation of the full-length structure by PCR using primers specific to the SS9 homologous region, the appropriateness of the incorporated structure was verified by sequencing of the PCR product of genomic DNA using the Sanger method.
[0532] The obtained E. coli strain was used for hispidin production. In the first step, the bacteria were incubated in LB medium at approximately 200 rpm at 37°C for a maximum of 10 hours in five 50 ml plastic tubes. 250 ml of the obtained culture was added to 3.3 liters of fermentation medium so that the initial culture optical density at 600 nm was approximately 0.35, and the mixture was placed in a Biostat B5 fermenter (Braun, Germany). The fermentation medium contained 10 g / l peptone, 5 g / l caffeic acid, 5 g / l yeast extract, 10 g / l NaCl, 25 g / l glucose, 15 g / l (NH4)2SO4, 2 g / l KH2PO4, 2 g / l MgSO4·7H2O, 14.7 mg / l CaCl2, 0.1 mg / l thiamine, and 1.8 mg / l and 0.1% solutions consisting of the following: 8 mg / The mixture contained 1 mg / l EDTA, 2.5 mg / l CoCl2·6H2O, 15 mg / l MnCl2·4H2O, 1.5 mg / l CuCl2·2H2O, 3 mg / l H3BO3, 2.5 mg / l Na2MoO4·2H2O, 13 mg / l Zn(CH3COO)2·2H2O, 100 mg / l ferric citrate, and 4.5 mg / l thiamine hydrochloride. Fermentation was carried out at 37°C with aeration at 3 l / min and mixing at 200 rpm. After 25 hours of incubation, arabinose was added to the culture medium until the final concentration was 0.1 mM. NH4OH was added to automatically control the pH and lower it to 7.0. A solution consisting of 500 g / l glucose, 5 g / l caffeic acid, 2 g / l arabinose, 25 g / l tryptone, 50 g / l yeast extract, 17.2 g / l MgSO4·7H2O, 7.5 g / l (NH4)SO4, and 18 g / l ascorbic acid was added to a fermenter, and the glucose level was maintained each time the pH rose to 7.1. After 56 hours of incubation, the hispidin concentration in the culture medium was 3.48 g / l. The fermentation medium and the hispidin purified from the medium by HPLC were active in the bioluminescent reaction with the hispidin hydroxylase and luciferase obtained according to Example 4.
[0533] For the purpose of growing the autonomous bioluminescent yeast Pichia pastris, an expression cassette consisting of coding sequences for Pichia leucocephala cyferase (SEQ ID NOs: 79, 80), Pichia leucocephala hispidin hydroxylase (SEQ ID NOs: 1, 2), Pichia leucocephala hispidin synthase (SEQ ID NOs: 34, 35), Pichia leucocephala caffeypyruvate hydrolase (SEQ ID NOs: 64, 65), Rhodobacter capsulata tyrosine ammonia lyase (SEQ ID NOs: 106, 107), and HpaB and HpaC (SEQ ID NOs: 108-111), components of Escherichia coli 4-hydroxyphenyl acetate 3-monooxygenase reductase obtained according to Example 15, was used under the control of the GAP promoter and tAOX1 terminator. Similar expression cassettes were synthesized consisting of the coding sequences of Arabidopsis thaliana-derived 4-coumarate-CoA ligase 1 (SEQ ID NOs: 140, 141), and three PKSs: PKS from Aquilaria sinensis (SEQ ID NOs: 120, 121), PKS from Arabidopsis thaliana (SEQ ID NOs: 122, 123), and PKS from Hydrangea (SEQ ID NOs: 124, 125). Each expression cassette was modified by introducing loxP via a BsmBI restriction enzyme recognition sequence. Homologous regions to the MET6 Pichia pastris gene (Uniprot, F2QTU9) with loxP introduced at the BsmBI restriction enzyme site were also synthetically obtained. The synthetic DNA was treated with BsmBI restriction enzyme and ligated into a single plasmid according to the Golden Gate cloning protocol described in [Iverson et al., ACS Synthetic Biology, January 15, 2016; Vol. 5(1): p. 99-103]. Three plasmids with different PKS compositions were generated. The obtained plasmids were used to generate transgenic yeast Pichia pastris according to the technique described in Example 15. Incorporation of the gene cassette into the genome was confirmed by PCR from primers annealed in homologous regions. All three resulting yeast strains, consisting of appropriate genomic inserts, were able to luminescent autonomously, in contrast to the wild yeast strain.
[0534] To promote the growth of autonomous bioluminescent flowering plants based on the pBI121 vector (Chlontech, USA), a set of binary vectors for Agrobacterium transformation was formed, consisting of the coding sequences of *Lyophyllum sarmentosum* luciferase (SEQ ID NO: 112), *Lyophyllum sarmentosum* hispidine hydroxylase (SEQ ID NO: 103), *Lyophyllum sarmentosum* caffeypyruvate hydrolase (SEQ ID NO: 114), a kanamycin resistance gene, and PKS (SEQ ID NOs: 122, 123), all optimized for expression in plants. Each gene is under the control of a 35S promoter derived from cauliflower mosaic virus. The sequences of the expression cassette assemblies were obtained synthetically. The vectors were assembled according to the Golden Gate cloning protocol described in [Iverson et al., ACS Synthetic Biology, January 15, 2016; Vol. 5(1):p.99-103]. Tobacco (Nicotiana tabacum) was transformed by co-culturing plant tissue with the AGL0 strain of Agrobacterium tumefaciens [Lazo et al., Biotechnology, October 1991; No. 9(10): p.963-7] consisting of a prepared binary vector. Transformation was performed by co-culturing tobacco leaf segments. The leaves were then cut into 0.5 cm long fragments and transferred to 10 ml of liquid Gamborg medium B-5 containing 20 g / l glucose, 0.5 g / l 2,4-dichlorophenoxyacetic acid, and 0.05 g / l kinetin, with 1.0 ml of Agrobacterium overnight culture solution added. The explants and Agrobacterium were co-cultured for a maximum of 2-3 minutes. The explants were then placed on a sterile filter in a petri dish containing agar medium of the same composition. After incubation in a 25°C incubator for 48 hours, the explants were transferred to fresh culture medium containing 500 mg / l cefotaxime and 50 mg / l kanamycin. Three weeks later, plant regeneration was initiated in a selective medium consisting of 50 mg / l kanamycin. The transgenic plants developed roots and were transferred to germination medium or soil. Bioluminescence was visualized using the IVIS Spectrum In Vivo Imaging System (Perkin Elmer). Over 10% of the transgenic plants emitted luminescence at least three orders of magnitude higher than the signal from wild-type plants. [Examples]
[0535] Example 24. Combinations of nucleic acids Combination 1: Composition: (a) a nucleic acid encoding hispidine hydroxylase having an amino acid sequence selected from the group SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28; and (b) a nucleic acid encoding luciferase having an amino acid sequence selected from the group SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98.
[0536] This combination can be used to obtain bioluminescence in an in vitro or in vivo expression system in the presence of a substance selected from the group of 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one having the following structural formula.
[0537] [ka]
[0538] In the formula, the substituent at position 6 is a 2-aryl vinyl or 2-heteroaryl vinyl substituent (R-CH=CH-) including 2-(3,4-dihydroxystyryl), 2-(4-hydroxystyryl), 2-(4-(diethylamino)styryl), 2-(2-(1H-indole-3-yl)vinyl), 2-(2-(1,2,3,5,6,7-hexahydropyrido[3,2,1-ij]quinoline-9-yl)vinyl), and 2-(6-hydroxynaphthalene-2-yl)vinyl.
[0539] This combination can also be used to study the dependence of two promoters in heterologous expression systems.
[0540] This combination can also be used to identify hispidin and its analogues in biological objects.
[0541] This combination can also be used for cell labeling by bioluminescence that occurs in the presence of hispidin and its functional analogues.
[0542] Combination 2: Composition: (a) a nucleic acid encoding hispidin hydroxylase having an amino acid sequence selected from the group SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28; and (b) a nucleic acid encoding hispidin synthase having an amino acid sequence selected from the group SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55.
[0543] This combination can be used to produce fungal luciferin in an in vitro or in vivo expression system from a substance selected from substituted acrylic acids having the following structural formula, where R is an aryl or heteroaryl (e.g., obtained from caffeic acid).
[0544] [ka]
[0545] Combination 3: It contains all the components specified in combination 2, and nucleic acids encoding luciferase having an amino acid sequence selected from the group of SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, and 98.
[0546] This combination can be used to obtain bioluminescence in an in vitro or in vivo expression system in the presence of a substance selected from substituted acrylic acids having the following structural formula, where R is aryl or heteroaryl.
[0547] [ka]
[0548] This combination can be used to generate bioluminescent cells and transgenic organisms. The combination can also be used to study the dependence of three promoters in heterologous expression systems.
[0549] Combination 4: It contains all the components specified in combination 3, and nucleic acids encoding caffeylpyruvate hydrolase having an amino acid sequence selected from the group of SEQ ID NOs: 65, 67, 69, 71, 73, and 75.
[0550] This combination can be used to generate bioluminescent cells and transgenic organisms.
[0551] Combination 5: Composition: (a) a nucleic acid encoding a hispidin synthase having an amino acid sequence selected from the group of SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55; and (b) a nucleic acid encoding a gene for 4'-phosphopantetheinyltransferase having the amino acid sequence shown in SEQ ID NO: 105.
[0552] This combination can be used to produce hispidin from caffeic acid in in vitro and in vivo expression systems.
[0553] Combination 6: Composition: (a) a nucleic acid encoding hispidin synthase having an amino acid sequence selected from the group of SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55; (b) a nucleic acid encoding the gene for 4'-phosphopantetheinyltransferase having the amino acid sequence shown in SEQ ID NO: 105; and (c) a nucleic acid encoding an enzyme for 3-arylacrylic acid biosynthesis having the following structural formula. In the formula, R is an aryl or heteroaryl obtained from a cellular metabolite (e.g., a nucleic acid encoding tyrosine ammonia lyase, and nucleic acids encoding HpaB and HpaC, components of 4-hydroxyphenylacetate 3-monooxygenase reductase).
[0554] [ka]
[0555] This combination can be used to produce hispidin from tyrosine in in vitro and in vivo expression systems.
[0556] Combinations 2-4 may also contain the coding sequences for the 4'-phosphopantheteinyltransferase NpgA gene (SEQ ID NOs: 104, 105) or other enzymes exhibiting the same activity.
[0557] Combination 7: Composition: (a) nucleic acids encoding hispidine hydroxylase having an amino acid sequence selected from the group SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28; and (b) nucleic acids encoding PKS having an amino acid sequence selected from the group SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139.
[0558] This combination can be used to produce 3-hydroxyhispidin from caffeyl-CoA in an in vitro or in vivo expression system.
[0559] Combination 8: The combination includes all the components specified in combination 7, and a nucleic acid encoding a luciferase having an amino acid sequence selected from the group of SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, and 98. This combination can be used to obtain bioluminescence in vitro or in vivo in the presence of caffeyl-CoA.
[0560] Combination 9: It contains all the components specified in combination 8, and nucleic acids encoding caffeylpyruvate hydrolase having an amino acid sequence selected from the group of SEQ ID NOs: 65, 67, 69, 71, 73, and 75.
[0561] This combination can be used to generate bioluminescent cells and transgenic organisms.
[0562] Combination 10: Composition: (a) Nucleic acid encoding PKS having an amino acid sequence selected from the group of SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139; and (b) Nucleic acid encoding Arabidopsis thaliana-derived 4-coumarate-CoA ligase 1 having the amino acid sequence shown in SEQ ID NO: 141.
[0563] This combination can be used to produce hispidin from caffeate in in vitro and in vivo expression systems.
[0564] Combination 11: It includes all the components specified in combination 10, and nucleic acids encoding enzymes for caffeic acid biosynthesis (for example, nucleic acids encoding tyrosine ammonia lyase and components HpaB and HpaC of 4-hydroxyphenyl acetate 3-monooxygenase reductase).
[0565] This combination can be used to produce hispidin from tyrosine in in vitro and in vivo expression systems. [Examples]
[0566] Example 25. Combinations of recombinant proteins Combination 1: Composition: (a) Hispidin hydroxylase having an amino acid sequence selected from the group of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28; and (b) Hispidin synthase having an amino acid sequence selected from the group of SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55.
[0567] This combination can be used to produce fungal luciferin from a substance selected from 3-arylacrylic acids having the following structural formula, where R is an aryl or heteroaryl (e.g., obtained from caffeic acid).
[0568] [ka]
[0569] Combination 2: The combination includes the components specified in Combination 1, and a luciferase having an amino acid sequence selected from the group of Sequence IDs: 80, 82, 84, 86, 88, 90, 92, 94, 96, and 98.
[0570] This combination can be used to detect a sample of 3-arylacrylic acid having the following structural formula, where R is an aryl or heteroaryl (e.g., obtained from caffeic acid).
[0571] [ka]
[0572] Combination 3: Composition: (a) hispidin hydroxylase having an amino acid sequence selected from the group of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28; (b) PKS having an amino acid sequence selected from the group of SEQ ID NOs: 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139; and (c) Arabidopsis thaliana-derived 4-coumarate-CoA ligase 1 having the amino acid sequence shown in SEQ ID NO: 141. This combination can be used to produce fungal luciferin from caffeic acid. [Examples]
[0573] Example 25. Kit In the following examples, nucleic acids are contained in an expression cassette or vector and can be operationally crosslinked to regulatory elements for expression in host cells. Alternatively, nucleic acids can consist of adjacent sequences for integration into a target vector. Nucleic acids can also be included in unpromotered vectors intended for easy cloning of target regulatory elements.
[0574] Reagent Kit No. 1 contains a purified hispidin synthase of the present invention and can be used to produce hispidin from caffeic acid. This kit can also be used to produce another 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one from the corresponding 3-arylacrylic acid. This other 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0575] [ka]
[0576] The 3-arylacrylic acid in question has the following structural formula, where R is aryl or heteroaryl.
[0577] [ka]
[0578] The reagent kit may also include a reaction buffer. For example, 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM), 1 mM NADPH, 10 mM ATP, 1 mM CoA, 1 mM malonyl-CoA, or 0.2 M sodium phosphate buffer (pH 8.0) with added components for reaction buffer preparation.
[0579] The reagent kit can also include deionized water.
[0580] Reagent kits can also include instructions for use.
[0581] Reagent Kit No. 2 comprises a purified hispidin synthase and a purified hispidin hydroxylase of the present invention, and can be used to produce fungal luciferin from a substance selected from 3-arylacrylic acids having the following structural formula: where R is an aryl or heteroaryl (e.g., obtained from caffeic acid).
[0582] [ka]
[0583] The reagent kit may also include 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM), 1 mM NADPH, 10 mM ATP, 1 mM CoA, 1 mM malonyl-CoA, or 0.2 M sodium phosphate buffer (pH 8.0) with added components for reaction buffer preparation.
[0584] The reagent kit can also include deionized water.
[0585] Reagent kits can also include instructions for use.
[0586] Reagent Kit No. 3 contains a purified hispidine hydroxylase of the present invention and can be used to produce fungal luciferin from 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one having the following structural formula, where R is aryl or heteroaryl.
[0587] [ka]
[0588] For example, the kit can be used to produce 3-hydroxyhispidin from hispidin.
[0589] The reagent kit may also include a reaction buffer. For example, 0.2 M sodium phosphate buffer (pH 8.0) with 0.5 M Na2SO4, 0.1% dodecyl maltoside (DDM), and 1 mM NADPH added.
[0590] The reagent kit can also include deionized water.
[0591] Reagent kits can also include instructions for use.
[0592] Reagent kits No. 4 and No. 5 differ from kits No. 2 and No. 3 in that they consist of purified luciferase having a substrate which is 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one having the following structural formula, where R is aryl or heteroaryl.
[0593] [ka]
[0594] This kit can be used to identify 3-arylacrylic acid and / or 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one in biological specimens, such as plant extracts, fungal extracts, and microorganisms. 3-arylacrylic acid has the following structural formula, where R is an aryl or heteroaryl compound (e.g., obtained from caffeic acid).
[0595] [ka]
[0596] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula, where R is an aryl or heteroaryl (e.g., hispidin).
[0597] [ka]
[0598] The reagent kit may also include a reaction buffer for carrying out the reaction (see the description for kits 2 and 3), or components for preparing the reaction buffer.
[0599] The reagent kit can also include deionized water.
[0600] Reagent kits can also include instructions for use.
[0601] The reagent kit may also include caffeic acid. For example, an aqueous solution of caffeic acid or the residue for dissolving it in water.
[0602] The reagent kit can also include hispidin.
[0603] Kit applications To identify the presence of caffeic acid in the test sample, add 5 μl of enzyme mixture to 95 μl of ice-cold reaction buffer in a cuvette, mix carefully, add 5 μl of the test sample, mix carefully again, and place in a luminometer. Integrate the bioluminescence signal at a maximum of 30°C for a maximum of 2 minutes. Perform a control measurement under the same conditions by adding 5 μl of caffeic acid solution or 5 μl of water instead of an aliquot of the test sample. If the luminescence emitted by the sample exceeds the background signal recorded from the sample with water, it can be said that caffeic acid is present in the sample in a detectable amount.
[0604] Sensitivity: This kit can determine if the caffeic acid concentration in the culture medium exceeds 1 nM.
[0605] Storage conditions: All components of the kit should be stored at a temperature of -20°C or below.
[0606] To identify the presence of hispidin in the test sample, add 5 μl of enzyme mixture to 95 μl of ice-cold reaction buffer in a cuvette, mix carefully, add 5 μl of the test sample, mix carefully again, and place in a luminometer. Integrate the bioluminescent signal at a maximum of 30°C for a maximum of 2 minutes. Perform a control measurement under the same conditions by adding 5 μl of hispidin solution or 5 μl of water instead of an aliquot of the test sample. If the luminescence emitted by the sample exceeds the background signal recorded from the sample with water, it can be said that hispidin is present in the sample in a detectable amount.
[0607] Sensitivity: This kit can determine if the hispidin concentration in the culture medium exceeds 100 pM.
[0608] Storage conditions: All components of the kit should be stored at a temperature of -20°C or below.
[0609] Reagent kit No. 6 contains a nucleic acid encoding the hispitidine hydroxylase of the present invention. For example, a hispitidine hydroxylase having an amino acid sequence selected from the group of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28.
[0610] Reagent kits can also include instructions for using nucleic acids.
[0611] The reagent kit may also include deionized water or buffer solution for dissolving lyophilized nucleic acids and / or for diluting nucleic acid solutions.
[0612] The reagent kit may also include primers complementary to the region of the nucleic acid for amplifying the nucleic acid or a fragment thereof.
[0613] The reagent kit can be used to generate the recombinant hispidin hydroxylase of the present invention or for hispidin hydroxylase expression in cells and / or cell lines and / or organisms. Following nucleic acid expression in cells, cell lines and / or organisms, these cells, cell lines and / or organisms acquire the ability to catalyze the transformation of exogenous or endogenous 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one to 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one. The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0614] [ka]
[0615] The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0616] [ka]
[0617] These cells, cell lines, and / or organisms acquire the ability to catalyze the transformation of hispidin to 3-hydroxyhispidin.
[0618] Reagent kit No. 7 contains a nucleic acid encoding the hispitidine synthase of the present invention. For example, a hispitidine synthase having an amino acid sequence selected from the group of SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55.
[0619] Reagent kits can also include instructions for using nucleic acids.
[0620] The reagent kit may also include deionized water or buffer solution for dissolving lyophilized nucleic acids and / or for diluting nucleic acid solutions.
[0621] The reagent kit may also include primers complementary to the region of the nucleic acid for amplifying the nucleic acid or a fragment thereof.
[0622] The reagent kit may also include nucleic acids encoding 4'-phosphopantheinyltransferase, for example, a 4'-phosphopantheinyltransferase having the amino acid sequence shown in SEQ ID NO: 105.
[0623] The reagent kit can be used to generate the recombinant hispidin synthase of the present invention, or for the expression of hispidin hydroxylase in cells and / or cell lines and / or organisms.
[0624] Following nucleic acid expression in cells, cell lines, and / or organisms, these cells, cell lines, and / or organisms acquire the ability to catalyze the transformation of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one. The 3-arylacrylic acid has the following structural formula, where R is aryl or heteroaryl.
[0625] [ka]
[0626] The 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one has the following structural formula.
[0627] [ka]
[0628] For example, these cells, cell lines and / or organisms acquire the ability to catalyze the transformation of caffeic acid to hispidin, and / or the conversion of cinnamic acid to (E)-4-hydroxy-6-styryl-2H-pyran-2-one, and / or the transformation of paracoumaric acid to bisnoriangonin, and / or the transformation of propenoic acid to (E)-4-hydroxy-6-(2-(6-hydroxynaphthalen-2-yl)vinyl)-2H-pyran-2-one with (E)-3-(6-hydroxynaphthalen-2-yl), and / or the transformation of propenonic acid to (E)-6-(2-(1H-indole-3-yl)vinyl)-4-hydroxy-2H-pyran-2-one with (E)-3-(1H-indole-3-yl).
[0629] The reagent kit may also include the nucleic acid encoding tyrosine-ammonia-lyase, as well as the components HpaB and HpaC of 4-hydroxyphenylacetate 3-monooxygenase-reductase. Kits with such compositions can be used to produce hispidin from tyrosine in in vitro and in vivo expression systems.
[0630] Reagent Kit No. 8 contains a nucleic acid encoding the hispitidine synthase of the present invention, and a nucleic acid encoding the hispitidine hydroxylase of the present invention. For example, an amino acid sequence selected from the group SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55; and a hispitidine hydroxylase having an amino acid sequence selected from the group SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28.
[0631] Reagent kits can also include instructions for using nucleic acids.
[0632] The reagent kit may also include deionized water or buffer solution for dissolving lyophilized nucleic acids and / or for diluting nucleic acid solutions.
[0633] The reagent kit may also include primers complementary to the nucleic acid region included in the kit for amplifying nucleic acids or fragments thereof.
[0634] The reagent kit may also include nucleic acids encoding 4'-phosphopantheinyltransferase, for example, a 4'-phosphopantheinyltransferase having the amino acid sequence shown in SEQ ID NO: 105.
[0635] The reagent kit may also include nucleic acids encoding enzymes for 3-arylacrylic acid biosynthesis from cellular metabolites, such as tyrosine ammonia lyase and nucleic acids encoding components HpaB and HpaC of 4-hydroxyphenyl acetate 3-monooxygenase reductase.
[0636] This kit can be used for all the purposes described for kits 6 and 7. This kit can be used for the expression of hispidin hydroxylase and hispidin synthase in cells and / or cell lines and / or organisms. After nucleic acid expression in cells, cell lines and / or organisms, these cells, cell lines and / or organisms acquire the ability to produce 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one from the corresponding 3-arylacrylic acid. The 6-(2-arylvinyl)-3,4-dihydroxy-2H-pyran-2-one has the following structural formula, where R is aryl or heteroaryl.
[0637] [ka]
[0638] The 3-arylacrylic acid in question has the following structural formula.
[0639] [ka]
[0640] This kit can be used to express hispidin hydroxylase and hispidin synthase in cells and / or cell lines and / or organisms, along with the components HpaB and HpaC of tyrosine-ammonia-lyase and 4-hydroxyphenyl acetate 3-monooxygenase-reductase. After nucleic acid expression in cells, cell lines and / or organisms, these cells, cell lines and / or organisms acquire the ability to produce hispidin from tyrosine and cellular metabolites.
[0641] Reagent kit No. 9 contains nucleic acids encoding the hispitidine hydroxylase of the present invention. For example, a hispitidine hydroxylase having an amino acid sequence selected from the group of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, and a nucleic acid encoding a luciferase capable of oxidizing at least one fungal luciferin with luminescence. For example, a luciferase having an amino acid sequence selected from the group of SEQ ID NOs: 80, 82, 84, 86, 88, 90, 92, 94, 96, 98 can be selected.
[0642] Reagent kits can also include instructions for using nucleic acids.
[0643] The reagent kit may also include deionized water or buffer solution for dissolving lyophilized nucleic acids and / or for diluting nucleic acid solutions.
[0644] The reagent kit may also include primers complementary to the nucleic acid region included in the kit for amplifying nucleic acids or fragments thereof.
[0645] The kit can be used to label cells and / or cell lines and / or organisms. Here, as a result of the expression of the nucleic acid, the cells, cell lines and / or organisms acquire bioluminescence in the presence of exogenous or endogenous fungal preluciferin. For example, the cells, cell lines and / or organisms acquire bioluminescence in the presence of hispidin.
[0646] This kit can also be used to study the simultaneous activation of target gene promoters.
[0647] The kit may also include hispidin-synthases having amino acid sequences selected from the group of SEQ ID NOs: 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, and 55. In this case, the kit can be used to generate cells, cell lines, and transgenic organisms capable of bioluminescence in the presence of exogenous or endogenous 3-arylacrylic acid. 3-arylacrylic acid has the following structural formula, where R is aryl or heteroaryl.
[0648] [ka]
[0649] For example, in the presence of a 3-arylacrylic acid selected from the following group: caffeic acid or cinnamic acid, or paracoumaric acid, or coumaric acid, or umberic acid, or sinapic acid, or ferulic acid. In particular, the kit can be used to generate autonomous bioluminescent transgenic organisms (e.g., plants or fungi).
[0650] The kit may also include nucleic acids encoding 4'-phosphopantheteinyltransferases, such as a 4'-phosphopantheteinyltransferase having the amino acid sequence shown in SEQ ID NO: 105.
[0651] The kit may also include nucleic acids encoding enzymes for 3-arylacrylic acid biosynthesis derived from cell metabolites.
[0652] The kit may also include nucleic acids encoding caffeylpyruvate hydrolase of the present invention, for example, caffeylpyruvate hydrolase having an amino acid sequence selected from the group of SEQ ID NOs: 65, 67, 69, 71, 73, and 75.
[0653] This kit can be used for all the purposes described in kits 6 and 8.
[0654] This kit can be used to generate cell lines that enable the identification...
Claims
1. A hispidin synthase protein comprising an amino acid sequence having at least 90% identity with an amino acid sequence selected from SEQ ID NOs: 37, 39, 41, 43, 45, 47, 49, 53, 55, wherein the hispidin synthase catalyzes the conversion of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one.
2. The nucleic acid encoding the hispidin synthase protein according to claim 1.
3. A vector comprising the nucleic acid according to claim 2.
4. A nonfungal cell comprising the nucleic acid described in claim 2, wherein the nonfungal cell expresses the hispidin synthase protein.
5. A nonfungal cell according to claim 4, further comprising a nucleic acid encoding 4'-phosphopantotheinyltransferase.
6. A nonfungal cell according to claim 5, wherein the 4'-phosphopantotheinyltransferase has at least 90% identity with the amino acid sequence of SEQ ID NO:
105.
7. A method for producing 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one, comprising culturing the cells described in any one of claims 4 to 6 under conditions sufficient to produce 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one.
8. A method for catalyzing the conversion of 3-arylacrylic acid to 6-(2-arylvinyl)-4-hydroxy-2H-pyran-2-one, comprising contacting 3-arylacrylic acid with the hispidin synthase protein described in claim 1 under conditions sufficient for the conversion to occur.
9. A method according to claim 8, wherein the 3-arylacrylic acid is selected from the group consisting of caffeic acid, cinnamic acid, paracoumaric acid, coumaric acid, umbellic acid, sinapic acid, and ferulic acid.
10. A method for producing hispidin in vitro or in vivo, wherein the in vivo system is a non-human system, and comprises conjugating the hispidin synthase protein described in claim 1 with 3-arylacrylic acid under physiological conditions.
11. A method according to claim 10, wherein the 3-arylacrylic acid is selected from the group consisting of caffeic acid, cinnamic acid, paracoumaric acid, coumaric acid, umbellic acid, sinapic acid, and ferulic acid.
12. A transgenic non-human organism comprising the nucleic acid described in claim 2, wherein the transgenic non-human organism expresses the hispidin synthase protein.