Newly designed luciferase

JP2025502272A5Pending Publication Date: 2026-01-16UNIV OF WASHINGTON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024542021
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-01
Filing Date
2023-01-13
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

The development of luciferase as a molecular probe has been hindered by the lack of suitable luciferases that recognize synthetic luciferin, require multiple disulfide bonds for stability in mammalian cells, and have low substrate specificity, limiting multiplexed imaging capabilities.

Method used

Design of a protein with a specific secondary structure configuration and amino acid residues that enhance luciferase activity, including helical and loop domains, to create a luciferase with high specificity for the synthetic luciferin substrate diphenyltetrazolium chloride (DTZ), using a deep learning-based approach to generate a protein scaffold with optimized binding pockets.

Benefits of technology

The engineered luciferase exhibits high substrate specificity, enabling multiplexed bioassays with improved sensitivity and specificity, surpassing natural luciferases in activity and stability, allowing for efficient luminescence in biological samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A protein having luciferase activity is disclosed, the protein having a secondary structure arrangement H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, where "H" is a helical domain, "L" is a loop domain, and "E" is a β-strand domain; (a) the H1 domain is at least 18 or 19 amino acids in length; residue 14 of the H1 domain is Y, D, or E, and residue 18 of the H1 domain is D or E; (b) the E3 domain is at least 6, 7, 8, 9, or 10 amino acids in length; residue 2 of the E3 domain is R; and (c) the E5 domain is at least 10, 11, 12, 13, or 14 amino acids in length; and residue 9 of the E5 domain is H or N.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 300,171, filed January 17, 2022, and U.S. Provisional Patent Application No. 63 / 381,922, filed November 1, 2022. These provisional applications are incorporated herein by reference in their entireties. Federal Government Support Statement This invention was made with Government support under Grant No. K99EB031913 awarded by the National Institutes of Health. The United States Government has certain rights in the invention. Sequence Listing A computer readable form of the sequence listing is being filed with this application by electronic submission and is hereby incorporated by reference in its entirety. The sequence listing is contained in a file entitled "21-1622-WO.xml", created on January 8, 2023, and is 638kb in size. [Background technology]

[0002] Bioluminescence, generated by the enzymatic oxidation of a luciferin substrate, is widely used for bioassays and imaging in biomedical research. Because no excitation light source is required, luminescence photons are generated in the dark, which results in higher sensitivity than fluorescence imaging in living animal models and biological samples where autofluorescence or phototoxicity is a concern. However, the development of luciferases as molecular probes has lagged behind the well-developed fluorescent protein toolkit for several reasons: (i) only very few natural luciferases have been identified; (ii) many of those identified require multiple disulfide bonds to stabilize their structure and are therefore prone to misfolding in mammalian cells; (iii) most natural luciferases do not recognize synthetic luciferins that have more desirable photophysical properties; and (iv) multiplexed imaging to track multiple processes in parallel using mutually orthogonal luciferase-luciferin pairs is limited by the low substrate specificity of natural luciferases. Summary of the Invention

[0003] In one aspect, the disclosure provides a protein having luciferase activity, the protein comprising a secondary structure arrangement H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, where "H" is a helical domain, "L" is a loop domain, and "E" is a β-strand domain; (a) the H1 domain is at least 18 or 19 amino acids in length; residue 14 of the H1 domain is Y, D, or E, and residue 18 of the H1 domain is D or E; (b) the E3 domain is at least 6, 7, 8, 9, or 10 amino acids in length; residue 2 of the E3 domain is R; and (c) the E5 domain is at least 10, 11, 12, 13, or 14 amino acids in length; and residue 9 of the E5 domain is H or N.

[0004] In various embodiments, residue 7 of the E5 domain is M; the E6 domain is at least 9, 10, 11, 12, or 13 amino acids in length, and residue 5 of E6 is V; residue 1 of the L5 domain is S; residue 7 of the E5 domain is M and residue 5 of the E6 domain is V; and / or residue 7 of the E5 domain is M, residue 5 of the E6 domain is V, and residue 1 of the L5 domain is S.

[0005] In other embodiments, the H2 domain is at least 5, 6, or 7 amino acids in length, the H3 domain is at least 9, 10, 11, 12, 13, or 14 amino acids in length, the E1 domain is at least 3 or 4 amino acids in length, the E2 domain is at least 3 or 4 amino acids in length, and / or the E4 domain is at least 8, 9, 10, 11, or 12 amino acids in length.

[0006] In one embodiment, one, two, three, four, or all five of the following are true: (a) Residue 13 of domain H1 is F; (b) residue 1 of domain L3 is W; (c) residue 5 of domain E5 is V or another hydrophobic residue; (d) residue 8 of domain E5 is A or L or another hydrophobic residue; and / or (e) Residue 11 of domain E5 is W.

[0007] In another embodiment, one, two, three, four, five, or all six of the following are true: (a) Residue 2 of domain E1 is I or another hydrophobic residue; (b) residue 4 of domain H3 is F; (c) residue 6 of domain E4 is V or another hydrophobic residue; (d) residue 8 of domain E4 is L or another hydrophobic residue; (e) residue 5 of domain E6 is M or V or another hydrophobic residue; and / or (f) Residue 7 of domain E6 is V or another hydrophobic residue.

[0008] In further embodiments, the protein comprises an amino acid sequence at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-181, or SEQ ID NOs: 1-3. In another embodiment, the protein comprises the amino acid sequence of SEQ ID NO:4.

[0009] In one aspect, the disclosure provides a protein having luciferase activity, the protein comprising an amino acid sequence at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO:1; Residue 14 is Y, D, or E and residue 98 is H or N; and Residue 18 is D or E and residue 65 is R.

[0010] In various embodiments, the protein comprises one or both of an A96M and an M110V substitution relative to SEQ ID NO: 1; both an A96M and an M110V substitution relative to SEQ ID NO: 1; and / or an R60S substitution relative to SEQ ID NO: 1; an R60S, A96M, and M110V substitution relative to SEQ ID NO: 1. In other embodiments, the protein comprises an amino acid sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-3, or SEQ ID NOs: 1-181.

[0011] In another aspect, the disclosure provides a protein comprising the formula X1-Z1-X2-Z2-X3-Z3-X4-Z4-X5-Z5-X6-Z6-X7-Z7-X8-Z8, X1 is, [ka] wherein residue 14 is Y, D, or E, and residue 18 is D or E; X2 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of ADTAASLF (SEQ ID NO: 183); X3 has an amino acid sequence at least 50%, 75%, or 100% identical to the amino acid sequence of TIHL (SEQ ID NO: 184); X4 has an amino acid sequence at least 33%, 66%, or 100% identical to the amino acid sequence of VTF; X5 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of EEFR EWFERLFST (SEQ ID NO: 185); The X6 is [ka] and wherein residue 2 is R; X7 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of VEVH VQLHATH (SEQ ID NO: 187); The X8 is [ka] and wherein residue 8 is H or N; X9 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of VTEM RVHINPTG (SEQ ID NO: 189); and Z1, Z2, Z3, Z4, Z5, Z6, Z7, and Z8 are independently present or absent and, if present, may comprise any amino acid sequence.

[0012] In another embodiment, the present disclosure provides a self-complementary multipartite protein having luciferase activity, the protein comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked, and wherein collectively the at least first polypeptide component and the second polypeptide component comprise the domains X1-Z1-X2-Z2-X3-Z3-X4-Z4-X5-Z5-X6-Z6-X7-Z7-X8-Z8-X9, each domain being as defined herein; (a) each X domain is present entirely within one of at least the first and second polypeptide components, and (b) none of the at least first and second polypeptide components comprises each of X1, X2, X3, X4, X5, X6, X7, X8, and X9.

[0013] In a further embodiment, the present disclosure provides a self-complementary multipartite protein having luciferase activity, the protein comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked, and wherein collectively the at least first polypeptide component and the second polypeptide component comprise the secondary structure sequence H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, each domain as defined herein.

[0014] In a further embodiment, the disclosure provides a protein having luciferase activity, the protein comprising a secondary structure arrangement H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, where "H" is a helical domain, "L" is a loop domain, and "E" is a β-strand domain; (a) the H1 domain is at least 18 or 19 amino acids in length; residue 14 of the H1 domain is Y, D, or E, and residue 18 of the H1 domain is D or E; (b) the E3 domain is at least 6, 7, 8, 9, or 10 amino acids in length; residue 2 of the E3 domain is R; and (c) the E5 domain is at least 10, 11, 12, 13, or 14 amino acids in length; and residue 9 of the E5 domain is H or N.

[0015] The present disclosure also provides (a) a protein or polypeptide component of any embodiment; and (b) one or more additional functional domains The present invention provides a fusion protein comprising the

[0016] The present disclosure further provides nucleic acids encoding the proteins, polypeptide components, or fusion proteins of the present disclosure, expression vectors comprising the nucleic acids operably linked to suitable control sequences, host cells comprising the proteins, polypeptide components, fusion proteins, nucleic acids, and / or expression vectors of the present disclosure; and kits comprising the proteins, polypeptide components, fusion proteins, nucleic acids, expression vectors, and / or host cells of the present disclosure; and instructions for their use. The present disclosure also provides methods for the use of the proteins, polypeptide components, fusion proteins, nucleic acids, expression vectors, host cells, and / or kits of the present disclosure. [Brief description of the drawings]

[0017] [Figure 1]Generation of ideal scaffolds and computational designs for novel luciferases. (a) Family-wide hallucination. Sequences encoding proteins with desired topology are optimized by Monte Carlo sampling with a multicomponent loss function. Structurally conserved regions are evaluated based on the consistency of input residue-residue distance and orientation distributions from 85 experimental structures of NTF2-like proteins, while variable non-ideal regions are evaluated based on the reliability of predicted inter-residue shapes calculated as the KL information between the network prediction and background distribution. Sequence-space MCMC sampling incorporates both sequence changes and insertions / deletions (see Methods) to guide hallucination sequences to encode structures with the desired fold. Hydrogen-bonding networks are incorporated into the designed structures to enhance structural specificity. (b-d) Luciferase active site design. (b) Generation of DTZ conformers using AIMNet. (c) Generation of a rotamer interaction field (RIF) to stabilize anionic DTZ and form hydrophobic packing interactions around DTZ conformers. (d) Docking of RIF into the hallucination scaffold and optimization of substrate-scaffold interactions using position specific score matrix (PSSM)-biased sequence design. (e) Selection of NTF2 topology. RIF was docked into 4000 natural small molecule binding proteins, excluding proteins that bind luciferin substrates with more than five loop residues. Most of the top hits were from the NTF2-like protein superfamily. Using a family-wide hallucination scaffold generation protocol, we generated 1615 scaffolds and found that these yielded better predicted RIF binding energies than the natural proteins. (f) scaffolds generated using family-wide hallucination generate more samples within the native structure space than scaffolds generated with previous blueprints, and (g) have stronger sequence to structure correlations than either native NTF2 scaffolds or blueprinted novel NTF2 scaffolds. [Diagram 2] Biophysical characterization of LuxSit. (a) Coomassie stained SDS-PAGE of purified recombinant LuxSit from E. coli. (b) Size exclusion chromatography of purified LuxSit suggested monodispersed monomeric properties. (c) Far-UV CD spectra at 25°C, 95°C, and cooled back to 25°C. Inset: CD melting curve of LuxSit at 220 nm. (d) Luminescence emission spectra of DTZ in the presence and absence of LuxSit. (e) Structural alignment of the designed model and the AlphaFold2 predicted model, which are nearly identical at both the main chain (left) and side chain level (right). (fi) Site-saturation mutagenesis of residues interacting with the substrate. Zoomed-in view (left) of the design and AlphaFold2 model at the side-chain level illustrating the designed enzyme-substrate interactions of (f) Tyr14-His98 core HBNet, (g) Asp18-Arg65 dyad, (h) π-stacking, and (i) hydrophobic packing residues. The sequence profile (right) is scaled by the activity of the different sequence variants: (activity for the indicated amino acid) / (sum of activities of all tested amino acids at that position). Substitutions with increased activity (Ala96 and Met110) are highlighted. [Diagram 3] Characterization of luciferase activity in vitro and in human cells. (a) Substrate concentration dependence of LuxSit, LuxSit-f, and LuxSit-i activity. Numbers indicate signal / background ratio at Vmax. (b) Fluorescence and luminescence imaging of live HEK293T cells transiently expressing LuxSit-i-mTagBFP2; LuxSit-i activity is detectable at single-cell resolution. Left: Fluorescence channel representing mTagBFP2 signal. Right: Total luminescence photons were collected during a 10 s exposure. Inset: Negative control, non-transfected cells containing DTZ. Luminescence images were acquired immediately after addition of 25 μM DTZ without excitation light. Scale bar: 20 μm. 40X. [Figure 4]The high substrate specificity of the engineered luciferases enables multiplexed bioassays. (a) Chemical structures of coelenterazine substrate analogs. (b) Activity of LuxSit-i against selected luciferin substrates. Luminescence images (top) and signal quantification (bottom) of the indicated substrates in the presence of 100 nM LuxSit-i. LuxSit-i has high specificity for its designed target substrate, DTZ. (c) Heatmap visualization of the substrate specificity of LuxSit-i, Renilla luciferase (RLuc), Gaussia luciferase (Gluc), and engineered NLuc from Oplophorus luciferase. The heatmap shows the luminescence of each enzyme against each substrate; values ​​are normalized for each enzyme to the highest signal of that enzyme against all substrates. (d) Luminescence emission spectra of LuxSit-i / DTZ and RLuc / PP-CTZ can be spectrally resolved by 528 / 20 and 390 / 35 filters (indicated by dotted bars) to recognize only their cognate substrates. (e) Schematic of multiplexed luciferase assay. HEK293T cells transiently transfected with CRE-RLuc, NFκB-LuxSit-i, and CMV-CyOFP plasmids were treated with either forskolin or human tumor necrosis factor alpha (TNFα) to induce expression of labeled luciferase. (fg) Luminescence signals from cells can be measured by either substrate resolution or spectral resolution methods using a plate reader. (f) For substrate resolution method, luminescence intensity was recorded without filters after addition of either PP-CTZ or DTZ. (g) For the spectral decomposition method, both PP-CTZ and DTZ were added and signals were acquired using 528 / 20 and 390 / 35 filters simultaneously. In (f) and (g), the lower panel shows the addition of forskolin or TNFα. Luminescence signals were acquired from lysates of 15,000 cells in CelLytic™ M reagent, while CyOFP fluorescence signals were used to normalize cell number and transfection efficiency. All data were normalized to the corresponding unstimulated control. Data are shown as mean ± SD (n=3). [Diagram 5]Proposed catalytic mechanism of coelenterazine-utilizing luciferase. Density functional theory (DFT) calculations suggested that the formation of an anionic state is an essential electron source for the activation of triplet oxygen (3O2). Supported by both theoretical26,27 and experimental evidence28,29, the subsequent oxygenation process may be via a single electron transfer (SET) mechanism, in which the surrounding reaction field can significantly affect the change in Gibbs free energy (ΔGSET). Finally, the thermal decomposition of the dioxetane light emitter intermediate can produce photons via progressive reversible charge transfer induced luminescence (GRCTIL), which is generally energetic. All previous examples of evidence are based on calculations in virtual solvents or chemiluminescence in ideal organic solvents. The detailed mechanism of the luciferase-catalyzed luminescence reaction remains unclear. We proposed that the key steps of the enzyme are to promote the formation of an anionic state and to create a favorable environment to promote efficient SET. Therefore, the goal of this work is to engineer an enzymatic reaction field around the substrate to stabilize the anionic substrate state and alter the local proton activity, solvent polarity, and hydrophobicity for efficient activation of 3O2. [Figure 6]Schematic diagram depicting colony-based luciferase screening. Computationally designed DNA sequences were provided in an oligo array, fragments were amplified by PCR, assembled, and ligated into the pBAD bacterial expression vector. The plasmid library was used to transform DH10B cells. Each colony grown on an LB agar plate represented one luciferase design. Plates were sprayed with DTZ solution and imaged to identify active colonies using a ChemiDoc™ imager. All active colonies were inoculated into 96-well plates, expressed, and purified to confirm individual luciferase activity. Selected plasmids can then be sequenced to demonstrate active design models that provide insight into design principles and enzyme function, or can be subjected to random mutagenesis for further development. Inset: Three luciferases were identified from this screen. We call the most active, DTZ-specific luciferase "LuxSit". [Figure 7]Expression, purification, and structural characterization of LuxSit mutants. (ac) Recombinant expression of (a) LuxSit, (b) LuxSit-i, and (c) LuxSit-f in E. coli. Lanes are annotated as follows: 1: pre-IPTG; 2: post-IPTG; 3: soluble lysate; 4: flow-through; 5: wash; 6: eluate; 7: post-TEV cleavage; 8: post-SEC. (df) Size-exclusion chromatography of purified (d) LuxSit; (e) LuxSit-i; and (f) LuxSit-f monomers. (gi) Deconvoluted mass spectra of (g) LuxSit, (h) LuxSit-i, and (i) LuxSit-f. (jk) Far-UV circular dichroism (CD) spectra (left panels) and CD melting curves at 220 nm (right panels) of (j) LuxSit-i; and (k) LuxSit-f at 25°C, 95°C, and cooled back to 25°C. (l) A dimer SEC peak was observed when LuxSit-i was concentrated to high concentrations (~50 μM) in Tris pH 8.0 buffer. Both dimer and monomer n SEC fractions showed the predicted size on SDS PAGE, and both peaks were catalytically active and emitted luminescence in the presence of 25 μM DTZ. [Figure 8]Screening of randomized NNK libraries at 60, 96, and 110 positions and sequence alignment between LuxSit and its variants. We generated fully randomized libraries at 60, 96, and 110 positions to thoroughly screen all possible combinations. After colony-based screening, we identified many colonies with strong luciferase activity against DTZ. Each colony was individually expressed (1 mL culture) in each well of a 96-well plate and purified accordingly (see Methods). (a) The individual luminescence activity of each selected variant was plotted and compared with the parental LuxSit. Luminescence activity was measured in the presence of 25 μM DTZ. Luminescence activity (RLU) was shown as the integrated signal over the first 15 min. Statistical analysis of amino acid frequency for luciferase activity at residues (b) 60, (c) 96, and (d) 110. Among all the selected mutants, Arg60 was identified as mutagenic because it may be structurally poorly defined since it arises from a loop and has no hydrogen-bonding partners. Ala96 prefers larger side chains (Leu, Ile, Met, and Cys), and Met110 prefers hydrophobic residues (Val, Ile, and Ala). The newly discovered mutant (R60S / A96L / M110V) had a 100-fold higher photon flux relative to LuxSit and was designated LuxSit-i due to its high brightness. [Figure 9] Sequence alignment of LuxSit (SEQ ID NO: 1), LuxSit-i (SEQ ID NO: 2), and LuxSit-f (SEQ ID NO: 3). The sequence alignment highlights the mutations. The conserved catalytic dyads of Asp18-Arg65 and Tyr14-His98 are shown. [Figure 10]Additional characterization of LuxSit variants. (a) Normalized luminescence kinetics of 15,000 untreated HeLa cells expressing LuxSit-i, 100 nM purified LuxSit-i, or 100 nM purified LuxSit-f in the presence of 50 μM DTZ. The more extended emission kinetics in HeLa cells is likely due to the diffusion rate of DTZ through the cell membrane. (b) Normalized luminescence decay curves of LuxSit-i in various pH buffers revealed a pH-dependent catalytic mechanism. (c) Luminescence quantum yields were estimated from the integrated luminescence signal up to the complete conversion of 125 pmol substrate into photons in the presence of 50 nM of the corresponding luciferase (see Methods). All data points were plotted as the average of three replicate measurements. [Figure 11]Expression, localization, and luminescence activity of LuxSit-i in live HEK293T and HeLa cells. (ab) Fluorescence imaging of live (a) HEK293T and (b) HeLa cells expressing LuxSit-i-mTagBFP2, either untargeted or localized to the nuclear (histone 2B), plasma membrane (KRasCAAX), or mitochondrial (DAKAP) subcellular compartments. Scale bar: 10 μm. (cd) Luminescence signal was measured in 15,000 untreated (c) HEK293T or (d) HeLa cells in the presence of 25 μM DTZ in DPBS. Gene transfer efficiency ranges from 60–70% for HEK293T cells and 5–10% for HeLa cells. (e) Luminescence emission spectrum obtained from LuxSit-i expressing HEK293T cells matches that of recombinant LuxSit-i purified from E. coli. (fg) Luminescence signal was measured in 15,000 (f) untreated LuxSit-i expressing HEK293T cells or (g) cell lysates in the presence of 25 μM of the indicated substrates. Luminescence intensity was normalized to the DTZ signal, demonstrating high DTZ specificity over other substrates in cell-based assays. Data are shown as total luminescence signal over the first 20 min measured by technical triplicates. (h) Outline of normalized luminescence intensity profile across different cells (n=10) of the luminescence image in main Fig. 3b; grey lines represent non-transfected cells. Error bars represent ±SEM. [Figure 12]The substrate specificity of LuxSit-i and the spectrally resolved luciferase-luciferin pair allow for multiplexed bioassays. (a) Orthogonal relationship between LuxSit-i-DTZ and RLuc-PP-CTZ (Prolume Purple, methoxy e-coelenterazine) luminescence pair. The indicated amounts of each luciferase were mixed in different ratios that totaled 100%. After addition of both DTZ and PP-CTZ substrates at 25 μM, the filtered light from 528 / 20 and 390 / 35 was measured simultaneously. Data are shown as mean ± SD (n=3). Heatmaps show the luminescence signals for individual luciferases (100 nM) or 1:1 mixtures in the presence of cognate or non-cognate (DTZ or PP-CTZ or both) substrates. Response signals were acquired simultaneously with Neo2™ with 528 / 20 and 390 / 35 filters. (c) Multiplexed luciferase assay in live HEK293T after co-transfection of Cre-RLuc, NFkB-LuxSit-i, and CMV-CyOFP plasmids and stimulation with forskolin (FSK) or human tumor necrosis factor alpha (TNFα). (d, e) (d) 15,000 untreated cells were assayed after addition of DTZ, PP-CTZ, or both DTZ and PP-CTZ in DPBS without cell lysis in substrate-resolved or (e) spectral-resolved mode (see Methods). Area scans of CyOFP fluorescence signals were used to estimate cell number and transfection efficiency. Reported units were RLU / au; Ex. / Em.=relative light unit / fluorescence intensity measurements at 480 / 580 nm. All data were normalized to the corresponding unstimulated controls. Data are shown as mean ± SD (n=3). [Figure 13] The secondary structure is shown mapped onto an exemplary protein of the disclosure (SEQ ID NO:1). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] All documents cited are incorporated herein by reference in their entirety.

[0019] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.

[0020] As used herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine ​​(Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).

[0021] In all embodiments of the polypeptides disclosed herein, all N-terminal methionine residues are optional (ie, N-terminal methionine residues may be present or omitted).

[0022] All embodiments of any aspect of this disclosure can be used in combination unless the context clearly dictates otherwise.

[0023] Unless the context clearly dictates otherwise, throughout the specification and claims, the terms "comprise" or "comprising" are to be construed in the inclusive sense, i.e., "including, but not limited to," rather than in the exclusive or exhaustive sense. Terms using the singular or plural also include the plural and singular, respectively. Additionally, the terms "herein," "above," and "below," and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application.

[0024] In a first aspect, the present disclosure provides a protein having luciferase activity comprising a secondary structure arrangement H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, where "H" is a helical domain, "L" is a loop domain, and "E" is a β-strand domain; (a) the H1 domain is at least 18 or 19 amino acids in length; residue 14 of the H1 domain is Y, D, or E, and residue 18 of the H1 domain is D or E; (b) the E3 domain is at least 6, 7, 8, 9, or 10 amino acids in length; residue 2 of the E3 domain is R; and (c) the E5 domain is at least 10, 11, 12, 13, or 14 amino acids in length; and residue 9 of the E5 domain is H or N.

[0025] As disclosed in the following examples, the proteins of the present disclosure are of non-natural origin, have luciferase activity, and share this described secondary structure arrangement. The arrangement is shown for the amino acid sequence of SEQ ID NO: 1 in FIG. 13. The inventors have conducted extensive studies to evaluate the important residues in the polypeptide for retaining luciferase activity, and have created a number of modified polypeptides, detailed in SEQ ID NO: 4 and in the following examples. The above required amino acids are those involved in the catalytic dyad, as described below and in the examples. Except as noted above, the different domains may be of any suitable length.

[0026] In one embodiment, residue 7 of the E5 domain is M. In another embodiment, the E6 domain is at least 9, 10, 11, 12, or 13 amino acids long, and residue 5 of the E6 domain is V. In a further embodiment, residue 1 of the L5 domain is S. In one embodiment, residue 7 of the E5 domain is M and residue 5 of the E6 domain is V. In a further domain, residue 7 of the E5 domain is M, residue 5 of the E6 domain is V, and residue 1 of the L5 domain is S. In another embodiment, the H2 domain is at least 5, 6, or 7 amino acids long, the H3 domain is at least 9, 10, 11, 12, 13, or 14 amino acids long, the E1 domain is at least 3 or 4 amino acids long, the E2 domain is at least 3 or 4 amino acids long, and the E4 domain is at least 8, 9, 10, 11, or 12 amino acids long. In all these embodiments, one or more of the described domains may independently comprise an amino acid residue. In one embodiment, one or more of the descriptive domains may independently comprise an additional 1, 2, 3, 4, or 5 residues.

[0027] In one embodiment, The H1 domain is 19 amino acids long; The H2 domain is 7 amino acids long; The E1 domain is 4 amino acids long; The E2 domain is 4 amino acids long; The H3 domain is 14 amino acids long; The E3 domain is 10 amino acids long; The E4 domain is 12 amino acids long; The E5 domain is 14 amino acids in length; and The E6 domain is 12 or 13 amino acids long.

[0028] The loop domains may be of any length and may include the insertion of any residues or functional domains deemed appropriate relative to the sequences exemplified herein, including, but not limited to, metal binding domains, drug binding domains, GPCR receptors, protein switches, and small molecule binding domains.

[0029] In another embodiment, the protein comprises an amino acid sequence that is at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-3, where the residues in brackets are optional and may be present or deleted.

[0030] SEQ ID NO:1 is the LuxSit construct disclosed herein. Figure 13 shows the domain structure mapped onto the SEQ ID NO:1 amino acid sequence. [ka] SEQ ID NO:2 is the LuxSit-i construct disclosed herein. [ka] SEQ ID NO:3 is the LuxSit-f construct disclosed herein. [ka]

[0031] In each of the annotated sequences shown as SEQ ID NOs: 1-3: (a) Positions in bold and underlined and in expanded font size are dyad 1 (catalytic residues) Y14 (H1 domain residue 14) + H98 (E5 domain residue 9); (b) Positions in bold and enlarged font size are dyad 2 (catalytic residues) D18 (H1 domain residue 9) + R65 (E3 domain residue 2); (c) Expanded font and non-bold positions indicate core packing (recognition residues) F13 (residue 13 in domain H1), I35 (residue 2 in domain E1), W38 (residue 1 in domain L3), F49 (residue 4 in domain H3), V81 (residue 6 in E4 domain), L83 (residue 8 in domain E4), V94 (residue 5 in domain E5), A / L97 (residue 8 in E5 domain), W100 (residue 11 in domain E5), M / V110 (residue 5 in domain E6), V112 (residue 7 in domain E6); and (d) The underlined and non-bolded positions are regions (loop domains or immediately adjacent) for splitting the enzyme or inserting other functional domains. In some embodiments of the protein, one, two, three, four, or all five of the following are true: (a) Residue 13 of domain H1 is F; (b) residue 1 of domain L3 is W; (c) residue 5 of domain E5 is V or another hydrophobic residue; (d) residue 8 of domain E5 is A or L or another hydrophobic residue; and / or (e) Residue 11 of domain E5 is W.

[0032] In other embodiments of the protein, the H2 domain is 7 amino acids long, the H3 domain is 14 amino acids long, the E1 domain is 4 amino acids long, the E2 domain is 4 amino acids long, and / or the E4 domain is 12 amino acids long. In further embodiments, one, two, three, four, five, or all six of the following are true: (a) Residue 2 of domain E1 is I or another hydrophobic residue; (b) residue 4 of domain H3 is F; (c) residue 6 of domain E4 is V or another hydrophobic residue; (d) residue 8 of domain E4 is L or another hydrophobic residue; (e) residue 5 of domain E6 is M or V or another hydrophobic residue; and / or (f) Residue 7 of domain E6 is V or another hydrophobic residue.

[0033] In another embodiment, the protein comprises the amino acid sequence of SEQ ID NO:4. [Table 1] TIFF2025502272000009.tif242151TIFF2025502272000010.tif145159

[0034] In another embodiment, the protein comprises an amino acid sequence at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-181, as shown in Table 1. SEQ ID NOs: 5-181 in Table 1 are redesigned amino acid sequences based on LuxSit-i (SEQ ID NO: 2) and exhibited their luciferase activity. [Table 2] TIFF2025502272000012.tif229159TIFF2025502272000013.tif219159TIFF2025502272000014.tif228159TIFF2025502272000015.tif229159TIFF2025502272000016.tif228159TIFF2025502272000017.tif229159TIFF2025502272000018.tif229159TIFF2025502272000019.tif219159TIFF2025502272000020.tif228159TIFF2025502272000021.tif228159TIFF2025502272000022.tif230159TIFF2025502272000023.tif229159TIFF2025502272000024.tif227159TIFF2025502272000025.tif218159TIFF2025502272000026.tif228159TIFF2025502272000027.tif228159TIFF2025502272000028.tif229159TIFF2025502272000029.tif229159TIFF2025502272000030.tif228159TIFF2025502272000031.tif219159TIFF2025502272000032.tif228159TIFF2025502272000033.tif229159TIFF2025502272000034.tif229159TIFF2025502272000035.tif229159TIFF2025502272000036.tif229159TIFF2025502272000037.tif219159TIFF2025502272000038.tif227159TIFF2025502272000039.tif229159TIFF2025502272000040.tif228159TIFF2025502272000041.tif228159TIFF2025502272000042.tif229159TIFF2025502272000043.tif219159TIFF2025502272000044.tif229159TIFF2025502272000045.tif228159TIFF2025502272000046.tif229159TIFF2025502272000047.tif228159TIFF2025502272000048.tif169159.

[0035] In another aspect, the disclosure provides a protein having luciferase activity comprising an amino acid sequence at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:1; Residue 14 is Y, D, or E and residue 98 is H or N; Residue 18 is D or E and residue 65 is R.

[0036] The protein of this aspect is of non-natural origin. In one embodiment, percent identity to a reference sequence is performed by sequence alignment by the Needleman-Wunsch algorithm, which is a common sequence alignment tool for those skilled in the art and allows for insertions and deletions.

[0037] In one embodiment, the protein comprises one or both of A96M and M110V substitutions relative to SEQ ID NO:1. In another embodiment, the protein comprises an R60S substitution relative to SEQ ID NO:1. In a further embodiment, the protein comprises an R60S, A96M, and M110V substitution relative to SEQ ID NO:1. In another embodiment, all substitutions relative to SEQ ID NO:1 at residues F12, I35, W38, F49, V81, L83, V94, A97, W100, M110, V112 are conservative amino acid substitutions. In one embodiment, the protein comprises an amino acid sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from SEQ ID NOs:1-3. In another embodiment, the protein comprises an amino acid sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from SEQ ID NOs: 1-181.

[0038] In another embodiment, the protein comprises the formula X1-Z1-X2-Z2-X3-Z3-X4-Z4-X5-Z5-X6-Z6-X7-Z7-X8-Z8, X1 is, [ka] wherein residue 14 is Y, D, or E, and residue 18 is D or E; X2 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of ADTAASLF (SEQ ID NO: 183); X3 has an amino acid sequence at least 50%, 75%, or 100% identical to the amino acid sequence of TIHL (SEQ ID NO: 184); X4 has an amino acid sequence at least 33%, 66%, or 100% identical to the amino acid sequence of VTF; X5 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of EEFREWFERLFST (SEQ ID NO: 185); The X6 is [ka] and wherein residue 2 is R; X7 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of VEVHVQLHATH (SEQ ID NO: 187); The X8 is [ka] and wherein residue 8 is H or N; X9 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of VTEMRVHINPTG (SEQ ID NO: 189); and Z1, Z2, Z3, Z4, Z5, Z6, Z7, and Z8 are independently present or absent and, if present, may comprise any amino acid sequence. In one embodiment, one, two, three, four, five, six, seven, or all eight of the following are true: Z1 includes SGD; Z2 contains HPGV (SEQ ID NO: 190); Z3 includes WDG; Z4 includes TSR; Z5 contains RKDA (SEQ ID NO: 191); Z6 includes GDT; Z7 includes NGQ; Z8 includes GNR; and Zero, one, two, three, four, five, six, seven, or all eight of Z1, Z2, Z3, Z4, Z5, Z6, Z7, and Z8 further comprise an additional polypeptide domain.

[0039] In one embodiment of all proteins of the present disclosure, the amino acid substitutions relative to the reference protein are conservative amino acid substitutions. As used herein, "conservative amino acid substitution" means that a given amino acid can be replaced by a residue with similar physicochemical properties, for example, replacing one aliphatic residue with another (such as Ile, Val, Leu, or Ala with each other), or replacing one polar residue with another (such as between Lys and Arg; Glu and Asp; or Gln and Asn). Other such conservative substitutions are known, for example, full-region substitutions with similar hydrophobic properties. Proteins containing conservative amino acid substitutions can be tested in any one of the assays described herein to confirm that the desired activity is retained. Amino acids can be grouped according to the similarity of their side chains (A. L. Lehninger, in Biochemistry, second ed., pp. 73-75, Worth Publishers, New York (1975)): (1) nonpolar: Ala (A), Val (V), Leu (L), Ile (I), Pro (P), Phe (F), Trp (W), Met (M); (2) uncharged polar: Gly (G), Ser (S), Thr (T), Cys (C), Tyr (Y), Asn (N), Gln (Q); (3) acidic: Asp (D), Glu (E); (4) basic: Lys (K), Arg (R), His (H). Alternatively, naturally occurring residues can be divided into groups based on common side chain properties: (1) hydrophobic: norleucine, Met, Ala, Val, Leu, Ile; (2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues that affect chain orientation: Gly, Pro; (6) aromatic: Trp, Tyr, Phe. Non-conservative substitutions entail the exchange of a member of one of these classes for another.Particular conservative substitutions include, for example; Ala to Gly or Ser; Arg to Lys; Asn to Gln or His; Asp to Glu; Cys to Ser; Gln to Asn; Glu to Asp; Gly to Ala or Pro; His to Asn or Gln; Ile to Leu or Val; Leu to Ile or Val; Lys to Arg, Gln, or Glu; Met to Leu, Tyr or Ile; Phe to Met, Leu or Tyr; Ser to Thr; Thr to Ser; Trp to Tyr; Tyr to Trp; and / or Phe to Val, Ile, or Leu.

[0040] In another embodiment, the present disclosure provides a self-complementary multipartite protein having luciferase activity comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked, and wherein collectively the at least first polypeptide component and the second polypeptide component comprise the domains X1-Z1-X2-Z2-X3-Z3-X4-Z4-X5-Z5-X6-Z6-X7-Z7-X8-Z8-X9, each domain being as defined above; and (a) each X domain is present entirely within one of at least the first and second polypeptide components, and (b) none of the at least first and second polypeptide components comprises each of X1, X2, X3, X4, X5, X6, X7, X8, and X9.

[0041] Split proteins are of non-natural origin. Split proteins comprise at least a first polypeptide component and a second polypeptide component, in which the X domain is conserved, while the split point is taken only in the Z domain. In other words, each X strand or (X1, X2, X3, X4, X5, X6, X7, X8, and X9) is entirely present in at least one of the first and second polypeptide components, while the protein is split into separate components at the Z domain (Z1, Z2, Z3, Z4, Z5, Z6, Z7, and Z8), and the Z domain where the split occurs may be absent or partially present in one or both of the first and second polypeptide components. As a non-limiting example, in various embodiments of split luciferase proteins, the first and second polypeptide components may comprise the components illustrated in Table 2. [Table 3]

[0042] In various embodiments, the split can occur at Z4, Z5, Z6, or Z7. In another embodiment, the present disclosure provides a self-complementary multipartite protein having luciferase activity comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked, and wherein collectively the at least first polypeptide component and the second polypeptide component comprise the secondary structure arrangement H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, each domain being as defined above; (a) each H and E domain is present entirely within at least one of the first and second polypeptide components, and (b) neither of the at least first and second polypeptide components contains all of the H and E domains.

[0043] In this embodiment, the split protein comprises at least a first and a second polypeptide component, in which the H and E domains are conserved, while the split point is taken only in the L domain, in various embodiments, the split occurs at L4, L5, L6, L7, or L8.

[0044] The split-proteins of these embodiments are only active when they are brought together and are thus conditionally active.

[0045] In another embodiment, the disclosure provides a fusion protein comprising: (a) a protein or polypeptide component of any embodiment or combination of embodiments herein; and (b) one or more additional functional domains.

[0046] As used herein, a "functional domain" is any polypeptide that can be usefully fused to a luciferase protein or split-protein component of the present disclosure. By way of non-limiting example, the one or more additional functional domains can include diagnostic polypeptides, any protein that one wishes to localize within a cell, tissue, or organism, and the like.

[0047] In another aspect, the present disclosure provides nucleic acids encoding proteins, protein components, or fusion proteins of any embodiment or combination of embodiments of the present disclosure. The nucleic acid sequences may include single- or double-stranded RNA (e.g., mRNA) or DNA in genomic or cDNA form, or DNA-RNA hybrids, each of which may contain chemically or biochemically modified non-natural or derivatized nucleotide bases. Such nucleic acid sequences may include additional sequences useful for facilitating expression and / or purification of the encoded polypeptide, including, but not limited to, polyA sequences, modified Kozak sequences, and sequences encoding epitope tags, export signals, and secretory signals, nuclear transport signals, and cell membrane localization signals. Based on the teachings herein, it will be clear to one of skill in the art what nucleic acid sequences encode the polypeptides of the present disclosure. In various non-limiting embodiments, the nucleic acid may include the nucleotide sequence of any one of SEQ ID NOs: 200-380, where the residues in brackets are optional and may be present or absent.

[0048] In a further aspect, the disclosure provides an expression vector comprising the nucleic acid of any aspect of the disclosure operably linked to a suitable control sequence. An "expression vector" includes a vector that operably links a nucleic acid coding region or gene to any control sequence that can affect the expression of the gene product. A "control sequence" operably linked to a nucleic acid sequence of the disclosure is a nucleic acid sequence that can affect the expression of a nucleic acid molecule. Control sequences need not be contiguous with the nucleic acid sequence, so long as they function to direct their expression. Thus, for example, intervening non-translated but transcribed sequences may be present between the promoter sequence and the nucleic acid sequence, and the promoter sequence can still be considered to be "operably linked" to the coding sequence. Other such control sequences include, but are not limited to, polyadenylation signals, termination signals, and ribosome binding sites. Such expression vectors can be of any type, including, but not limited to, plasmids and viral-based expression vectors. The control sequences used to drive expression of the disclosed nucleic acid sequences in mammalian systems can be constitutive (driven by any of a variety of promoters, including but not limited to, CMV, SV40, RSV, ACTIN, EF) or inducible (driven by any of a number of inducible promoters, including but not limited to, tetracycline, ecdysone, steroid responsive). The expression vector must be replicable in the host organism, either as an episome or by integration into the host chromosomal DNA. In various embodiments, the expression vector may comprise a plasmid, a virus-based vector, or any other suitable expression vector.

[0049] In another aspect, the disclosure provides a host cell comprising a nucleic acid, expression vector (i.e., episomally or chromosomally integrated), non-naturally occurring polypeptide, fusion protein, or composition disclosed herein, the host cell can be prokaryotic or eukaryotic. The cell can be transiently or stably modified to incorporate a nucleic acid or expression vector of the disclosure using techniques including, but not limited to, bacterial transformation, calcium phosphate co-precipitation, electroporation, or liposome-mediated DEAE-dextran-mediated, polycation-mediated, or viral-mediated gene transfer. The present disclosure also provides (a) a protein, polypeptide component, fusion protein, nucleic acid, expression vector, and / or host cell of any embodiment or combination of embodiments herein; and (b) Instructions for their use A kit comprising:

[0050] In one embodiment, the kit further comprises diphenylterazine (DTZ). In another aspect, the disclosure provides methods for the use of any preceding claimed protein, polypeptide component, fusion protein, nucleic acid, expression vector, host cell, and / or kit for any suitable purpose, including, but not limited to, use for luminescence reporting assays, diagnostic assays, cellular localization of targets of interest, cell imaging, gene editing, live animal imaging, cancer labeling, CART cell reporting, secretion assays, gene delivery, tissue engineering, etc. Additional details can be found in the Examples.

[0051] In another aspect, the present disclosure provides a method for making luciferase, including de novo design using the method of any embodiment disclosed herein, starting from a protein comprising the amino acid sequence of SEQ ID NO: 381. The examples provide detailed methods for de novo design of luciferase for DTZ. As described in the examples, the method includes designing a shape-complementary catalytic site that stabilizes the anionic state of DTZ and lowers the SET energy barrier, assuming that the downstream dioxetane light emitter pyrolysis step is spontaneous. To stabilize the anionic species of DTZ, we focused on the placement of the positively charged guanidinium group of the arginine residue that interacts with the anionic imidazopyrazinon core. To computationally design such an active site, we first used AIMNet to generate a set of anionic DTZ conformers (Figure 1b). Next, around each conformer, we enumerated a rotamer interaction field (RIF) on a 3D grid consisting of millions of amino acid side chain configurations that form hydrogen bonds and nonpolar interactions with DTZ using the RIFgen method (Fig. 1c). In addition, we included an arginine guamidium group near the deprotonation site of the nitrogen (N1 atom) of the imidazopyrazinones in the RIFs. RIFdock was then used to dock each DTZ conformer and the associated RIF in the central cavity of each scaffold to maximize protein-DTZ interactions. On average, eight side chain rotamers containing arginine to stabilize the anionic imidazopyrazinones core were placed in each pocket (Fig. 1d). For the top 50,000 dockeds with the most favorable side chain-DTZ interactions, we optimized the remainder of the sequence for high affinity binding to DTZ, with a bias against sequence changes observed in nature to ensure foldability (Fig. 1d). During the design process, predefined hydrogen bond networks (HBNets) in the scaffold were kept open for structural specificity and stability, and interactions of these HBNet side chains with DTZ were explicitly required in the RIFdock step to ensure pre-construction of residues essential for catalysis.In the first sequence design phase, the identities of all RIF and HBNet residues were kept fixed and the surrounding residues were optimized to hold the side chain-DTZ interactions in place and maintain structural specificity. In the second sequence design step, the identities of RIF residues (other than arginine) were also varied to identify apolar and aromatic packing interactions that were missed by RIF due to binning effects. During sequence design, the scaffold backbone, side chains, and DTZ substrate were relaxed in Cartesian space.

[0052] Working Example summary De novo enzyme design has sought to introduce active sites and substrate-binding pockets predicted to catalyze reactions of interest into geometrically flexible natural scaffolds. 1,2 However, they have been limited by a lack of suitable protein structures and the complexity of natural protein sequence-structure relationships. Here, we describe a deep learning-based "family-wide hallucination" approach that generates a large number of ideal protein structures containing diverse pocket shapes and the designed sequences that encode them. We use these scaffolds to bind a synthetic luciferin substrate, diphenylterazine (DTZ), by positioning the arginine anion group adjacent to the anionic species generated during the reaction in a highly complementary binding pocket. 3 We design artificial luciferases that selectively catalyze the oxidative chemiluminescence of . For both luciferin substrates, we obtain engineered luciferases with high selectivity; the most active of these is small (13.9 kDa) and thermostable (T M >95°C) enzyme with catalytic efficiency (K cat / K M =10 6 M -1 s -1 ) but with much higher substrate specificity. The design from scratch of highly active and specific biocatalysts with broad applications in biomedicine is a significant milestone for computational enzyme design, and our approach should enable the design of a wide range of novel and useful luciferases and other enzymes.

[0053] the study Bioluminescence, generated by the enzymatic oxidation of a luciferin substrate, is widely used for bioassays and imaging in biomedical research. Because no excitation light source is required, luminescence photons are generated in the dark, which makes fluorescence imaging in living animal models and biological samples a concern for autofluorescence or phototoxicity. 4,5 However, the development of luciferases as molecular probes has lagged behind the development of a well-developed fluorescent protein toolkit for a number of reasons: (i) very few natural luciferases have been identified; (ii) many of those identified require multiple disulfide bonds to stabilize their structure and are therefore prone to misfolding in mammalian cells; (iii) most natural luciferases do not recognize synthetic luciferins, which have more desirable photophysical properties; and (iv) multiplexed imaging to simultaneously follow multiple processes using mutually orthogonal luciferase-luciferin pairs is limited by the poor substrate specificity of natural luciferases.

[0054] We explored the use of novel protein design to generate novel luciferases that are small, highly stable, well expressed in cells, specific for one substrate, and do not require cofactors to function. Because we are agnostic to natural luciferase substrates, we have the advantage of being able to exploit their good quantum yield, red-shifted emission, and 3 , favorable in vivo pharmacokinetics 14,15Due to the lack of a cofactor required for light emission, and the large number of folds required for DTZ binding, we select diphenylterazine (DTZ), a synthetic luciferin, as the target substrate. Previous computational enzyme design studies have primarily repurposed natural protein scaffolds in the PDB, but few natural structures have a suitable binding pocket for DTZ, and the effect of sequence changes on natural proteins is unpredictable. To circumvent these constraints, we set out to generate a large number of ideal protein scaffolds with a pocket of appropriate size and shape for DTZ and with a well-defined sequence-structure relationship that facilitates the incorporation of the subsequent active site. To identify protein folds that can accommodate such pockets, we first docked DTZ into 4000 natural small molecule binding proteins. We found that many NTF2 (nuclear transport factor 2)-like folds have binding pockets with the appropriate shape-complementary size for DTZ placement (Figure 1e), and therefore selected the NTF2-like superfamily as the target topology.

[0055] Family Wide Hallucination Natural NTF2 structures have a variety of pocket sizes and shapes, but also contain non-ideal features such as long loops that compromise stability. To generate a large number of ideal NTF2-like structures, we used unconstrained de novo design. 19,20 and fixed chain sequence design method 21 We developed a deep learning-based "family-wide hallucination" method that integrates multiple methods to enable the generation of proteins with a virtually unlimited number of desired folds (Fig. 1a). The family-wide hallucination approach uses unconstrained protein hallucination for loops and variable regions. 19,20 We exploit the de novo sequence and structure discovery capabilities of the and structure-guided sequence optimization for core regions. 22We employed a novel method for identifying novel designed proteins that work experimentally and that is effective in hallucinating novel globular proteins of diverse topologies. Starting with the sequences and predicted structures of 2,000 naturally occurring NTf2s, we optimized the amino sequences of the conserved core and variable loop regions using trRosetta™. Protein core idealization was performed using topology-specific loss functions for core residues versus shape (see Methods) and variable loop optimization by optimizing sequence length and identity to maximize the neural network's confidence in the predicted structure. To further encode structural specificity, we incorporated a buried long-range hydrogen bond network. The resulting 1615 family-wide hallucinating NTF2 scaffolds yielded shape-complementary binding pockets for more DTZs than natural small molecule protein-binding proteins (Fig. 1e). This approach sampled protein backbones closer to natural NTF2-like proteins (Fig. 1f) and was more robust than previous non-deep learning approaches. 23 has better scaffold quality metrics than (Fig. 1g).

[0056] We selected the NTF2 scaffold of SEQ ID NO: 381, from which we engineer a luciferase for DTZ, as described in detail below. >2692_0_0.35_5_19_1806_2_0.45_5_Y14_H98_W100d1o7nb__clean_0001_D18_R65.rd1.pdb MSEEEIRQFLRRFYEAFDKGDVDTFASLFHPGVTIHVWQGITFTSREELREWVERFLRNFKDMQREILSLEVRGDTVEVHVQVHTTHNGQKYTFDVTHHWHFRGHRVTEIRVHVNPT (SEQ ID NO: 381)

[0057] De novo design of luciferase for DTZ Standard computational enzyme design typically begins with an idealized active site or theozyme, consisting of protein functional groups around the reaction transition state, and is then extrapolated to a set of existing scaffolds. 1,2 However, the detailed mechanism of native marine luciferase is not fully understood, since only a few apo structures have been characterized, and the holo structure has not been characterized. 24,25 (Excluding calcium-regulated photoproteins). Quantum chemical calculations 26,27 and experimental data 28,29 Both of these indicate that the chemiluminescence reaction proceeds via an anionic species, and that the polarity of the surroundings is such that the subsequent triplet molecular oxygen ( 3 These results suggest that the free energy of the single electron transfer (SET) process with DTZ (O2) can be substantially altered. Guided by these data (Figure 5), we explored the design of a shape-complementary catalytic site that would stabilize the anionic state of DTZ and lower the SET energy barrier, assuming that the downstream dioxetane light emitter thermolysis step is spontaneous. To stabilize the anionic species of DTZ, we focused on the placement of a positively charged guanidinium group of an arginine residue that interacts with the anionic imidazopyrazinon core.

[0058] To computationally design such active sites into multiple hallucination NTF2 scaffolds, we first used the AIMNet 30 We used the RIFgen method to generate a set of anionic DTZ conformers (Figure 1b). Then, around each conformer, we 31,32Using RIFdock, we enumerated rotamer interaction fields (RIFs) on a 3D grid consisting of millions of amino acid side chain configurations that form hydrogen bonds and nonpolar interactions with DTZ (Fig. 1c). In addition, we included an arginine guamidium group near the deprotonation site of the nitrogen (N1 atom) of the imidazopyrazinones in the RIFs. RIFdock was then used to dock each DTZ conformer and the associated RIF in the central cavity of each scaffold to maximize protein-DTZ interactions. On average, eight side chain rotamers, including arginines to stabilize the anionic imidazopyrazinones core, were placed in each pocket (Fig. 1d). For the top 50,000 docked bodies with the most favorable side chain-DTZ interactions, we used RosettaDesign™ (Fig. 1d) to optimize the remainder of the sequence for high affinity binding to DTZ, with a bias against sequence changes observed in nature to ensure foldability. During the design process, the predefined hydrogen bond networks (HBNet) in the scaffold were kept raw for structural specificity and stability, and the interactions of these HBNet side chains with DTZ were explicitly required in the RIFdock step to ensure pre-construction of residues essential for catalysis. In the first sequence design stage, the identities of all RIF and HBNet residues were kept fixed, and the surrounding residues were optimized to hold the side chain-DTZ interactions in place and maintain structural specificity. In the second sequence design step, the identities of RIF residues (other than arginine) were also varied to identify apolar and aromatic packing interactions that were missed by RIF due to binning effects. During sequence design, the scaffold backbone, side chains, and DTZ substrate were relaxed in Cartesian space. After sequence optimization, the designs were screened based on ligand binding energy, protein-ligand hydrogen bonds, shape complementarity, and contact molecular surface, and 7982 designs were selected and organized as pool oligos for experimental screening.

[0059] Screening and characterization of DTZ-specific luciferases Oligonucleotides encoding the halves of each design were assembled into full-length genes and cloned into an E. coli expression vector (see Methods). A colony-based screening method was used to directly image active luciferase colonies from the library, and activity of selected clones was confirmed using 96-well plate expression (Figure 6). Three active designs were identified; we call the most active of these LuxSit (Latin: make light exist); LuxSit is the smallest known luciferase with 117 residues (13.9 kDa). Biochemical analyses including SDS-PAGE and size-exclusion chromatography (Figure 2ab and Figure 7) showed that LuxSit is a soluble monomer that is highly expressed from E. coli. Circular dichroism (CD) spectroscopy showed strong far-UV CD characteristics, suggesting an organized α-β structure. CD melting experiments showed that the protein was not completely unfolded at 95°C, and that its intact structure was restored when the temperature was reduced (Figure 2c). Incubation of LuxSit with DTZ resulted in luminescence with an emission peak at approximately 480 nm (Figure 2d), consistent with the chemiluminescence spectrum of DTZ. We were unable to determine the crystal structure of LuxSit, but AlphaFold2 33 The predicted structure is quite close to the designed model at the main chain level (RMSD = 1.3 Å) and with respect to the side chains interacting with the substrate (Fig. 2e). The designed LuxSit active site contains the Tyr14-His98 and Asp18-Arg65 dyads; the imidazole nitrogen atom of His98 forms hydrogen-bonding interactions with Tyr14 and the O1 atom of DTZ (Fig. 2f). The center of the Arg65 guanidinium cation is 4.2 Å from the N1 atom of DTZ, and Asp18 forms a bidentate hydrogen bond to the guanidinium group and the main chain NH of Arg65 (Fig. 2g).

[0060] Activity Optimization To better understand the contributions to catalysis of our most active luciferase design, LuxSit, we constructed a site-saturation mutagenesis (SSM) library in which all mutations were made at all pocket residues, one by one (see Methods). Figure 2f-i illustrates the amino acid selection at key positions. Arg65 is highly conserved and its dyad partner Asp18 can only be mutated to Glu, which reduces activity, suggesting that the carboxylate-Arg65 hydrogen bond is important for luciferase activity. In the Tyr14-His98 dyad, Tyr14 can be substituted with Asp and Glu, while His98 can be substituted with Asn. Since all active mutants have hydrogen bond donors and acceptors at these positions, the dyad can assist in the transfer of electrons and protons required for luminescence. Hydrophobic (Figure 2h) and π-stacked (Figure 2i) residues at the binding interface generally favor amino acids in the original design, which tolerates other aromatic or aliphatic substitutions and is consistent with model-based affinity predictions of mutational effects. The A96M and M110V mutants enhance activity 16- and 19-fold, respectively, compared to LuxSit (Table 4). Optimization guided by these results yielded LuxSit-f (A96M / M110V) with strong initial flash emission and LuxSit-i (R60S / A96L / M110V) with 100-fold higher photon flux than LuxSit (Figure 9). Overall, the results of active site saturation mutagenesis corroborate the design model, with Tyr14-His98 and Asp18-Arg65 dyads playing key roles in catalysis and the substrate binding pocket being largely conserved.

[0061] The most active catalysts, LuxSit-f and LuxSit-i, are both soluble expressed at high levels in E. coli and are monomeric (some dimerization was observed at high protein concentrations, Fig. 7i) and thermostable (Fig. 7j-k). Similar to the native CTZ-utilizing luciferase, the apparent Michaelis constants K MThe activity is in the low μM range (Figure 3a) and the luminescence signal decays over time due to rapid catalytic turnover (Figure 10a). LuxSit-i is a highly efficient enzyme, with 6 M -1 s -1 K cat / K M The luminescence signal is easily visible to the naked eye and has a photon flux (photons s -1 ) is 38% greater than that of native Renilla luciferase (RLuc) (Table 3). The DTZ luminescence reaction catalyzed by LuxSit-i is pH dependent (FIG. 10b), consistent with the proposed mechanism.

[0062] Cellular imaging and multiplexed bioassays Since luciferase is a commonly used gene tag and reporter for the study of cellular functions, we evaluated the expression and function of LuxSit-i in living mammalian cells. LuxSit-i-mTagBFP2-expressing HEK293T cells had DTZ-specific luminescence (Figure 3b), which was maintained after targeting of LuxSit-i to the nucleus, membrane, and mitochondria (Figure 11). Natural and previously engineered luciferases, likely due to their large opening pocket (luciferases with high specificity for one luciferin substrate have been difficult to control, even with extensive directed evolution). 35 ) are fairly promiscuous, with activity towards many luciferin substrates (Figure 4ac). In contrast, LuxSit-i showed exceptional specificity towards its target luciferin, with a 50-fold preference for DTZ over bis-CTZ (which differ only at the benzyl carbon). Overall, the specificity of our engineered luciferases was comparable to that of natural luciferases. 36,37 or previously modified luciferase 38 Much larger than that.

[0063] The high substrate specificity of LuxSit-i may enable multiplexing of luminescence reporters via substrate-specific or spectrally resolved luminescence signals (Fig. 4d and Fig. 12ab). To investigate this possibility, we followed two independent signaling pathways (cAMP / pKa and NF-κB) by placing expression of RLuc or LuxSit-i downstream of NF-κB or cAMP response element promoters, respectively (Fig. 4e). Imaging one by one in the presence of substrates for the two luciferases (PP-CTZ for RLuc and DTZ for LuxSit-i) could clearly distinguish known activators of the two pathways. Since the luminescence of the two reactions occurs at different wavelengths, we could also simultaneously assess the activation of two signaling pathways in the same sample in untreated HEK293T cells or cell lysates (Fig. 12c-e) by supplying both substrates together and monitoring the luminescence at different wavelengths (Fig. 4g).

[0064] conclusion To date, computational enzyme design has been constrained by the number of available scaffolds, which limits the catalytic configuration and the degree to which enzyme-substrate shape complementation can be achieved. 16-18 Our use of deep learning to generate a large number of de novo designed scaffolds removes this limitation; in the near future we will develop more accurate RoseTTAfold™ 39 and AlphaFold2™ 33This should enable the generation of even more effective protein scaffolds by leveraging family-wide hallucination capabilities. Diversity in scaffold pocket shapes and sizes allows for exploration of numerous catalytic shapes and maximization of substrate-enzyme shape complementarity; to our knowledge, no natural luciferase has a similar fold to LuxSit, and the two enzymes have high specificity for the non-naturally occurring, fully synthetic luciferin substrate. With the incorporation of 2-3 substitutions that provide more complementary pockets to stabilize the transition state, LuxSit-i has higher activity than any previous de novo designed enzyme; 6 M -1 s -1 K cat / K M is within the range of natural luciferase. This is a remarkable advance for computational enzyme design, as dozens of rounds of directed evolution were required to obtain catalytic performance in this range for a designed retroaldolase, and the structure was extensively reconstructed. 40 in contrast, the predicted differences in ligand-side chain interactions between LuxSit and LuxSit-i are quite small. Achieving such high activity from the computer remains an open goal for computational enzyme design. The small size of LuxSit makes it well suited as a gene tag for volume-limited viral vectors, biosensor development, and fusion to proteins of interest. On the basic science side, the small size, simplicity, and high activity make LuxSit-i an excellent model system for computational and experimental studies aimed at improving understanding of luciferase catalytic mechanisms. Extension of the approach used here to generate novel luciferases similarly specific for synthetic luciferin substrates beyond DTZ and h-CTZ is shown in Figure 4, or by using a microscopic phasor. 41This will greatly expand multiplexing opportunities, resulting in a broad range of useful multiplexed luminescence toolkits. More generally, our family-wide hallucination method opens up the possibility of a nearly unlimited number of novel scaffolds for substrate binding and catalytic residue arrangements, which is particularly important when the reaction mechanism, and how it is facilitated, is not fully understood and many structural and catalytic hypotheses can be readily enumerated with different catalytic residue arrangements in the binding pocket with shape and chemical complementarity. Although luciferases are unique in catalyzing the emission of light, the chemical conversion of substrates to products is common to all enzymes, and the approach developed here should be readily extendable to a wide variety of chemical reactions. [Table 4]

[0065] method: 1. Materials and General Methods Synthetic genes and oligonucleotides were purchased from Integrated DNA Technologies or GenScript. Synthetic genes were inserted between the NdeI and XhoI sites of the pET29b+ vector, which contains an N-terminal hexahistidine tag followed by a TEV protease cleavage site and a C-terminal stop codon. Restriction endonucleases, Q5 PCR polymerase, and T4 ligase were purchased from NEB. Plasmid DNA, PCR products, or digested fragments were purified by Qiagen DNA purification kits. DNA sequences were analyzed by Genewiz. Coelenterazine (CTZ) was purchased from Gold Biotechnology. Diphenylterazine (DTZ), pyridyldiphenylterazine (8pyDTZ), and furimazine (FRZ) were purchased from MedChemExpress. All other coelenterazine analogs (bis-CTZ: bisdeoxycoelenterazine; f-CTZ: f-coelenterazine; e-CTZ: e-coelenterazine-F; PP-CTZ: methoxy e-coelenterazine; v-CTZ: v-coelenterazine). All other reagents were purchased from Sigma-Aldrich or Fisher Scientific and used without further purification. To identify the molecular weight of each protein, raw mass spectra were obtained via reversed-phase LC / MS with an AdvanceBio RP-desalting column on an Agilent 6230B TOF, and subsequently deconvoluted by Bioconfirm software using a total entropy algorithm. An AKTA pure M (GE Healthcare) with UNICORN 6.3.2 workstation control combined with a Superdex™ 75 Increase 10 / 300 GL column was used for size exclusion chromatography. DNA and protein concentrations were measured by a NanoDrop™ small-volume 8-channel UV / vis spectrometer. CD spectra and CD melting experiments were performed on a J-1500 circular dichroism spectroscopy spectropolarimeter (Jasco) with default settings. All luminescence measurements were acquired on a BioTek Synergy Neo2™ Multi-Mode Plate Reader.To convert relative arbitrary units (RLU) to photon counts, use a Neo2 plate reader as previously described. 45 The SDS PAGE and luminescence images were captured on a Bio-Rad ChemiDoc™ XRS+. Images were analyzed using Fiji image analysis software.

[0066] 2. General Procedure for Protein Production and Purification Lemo21(DE3) strain was used for transformation with pET29b+ plasmids encoding genes of interest. Transformed cells were grown for 12 hours in TB medium supplemented with kanamycin. Cells were inoculated in 100 mL fresh TB medium at a 1:50 ratio and grown for 4 hours at 37°C, then induced with IPTG for another 18 hours at 16°C. Cells were harvested by centrifugation at 4,000g for 10 minutes and resuspended in 30 mL of lysis buffer (20 mM Tris-HCl pH 8.0, 300 mM NaCl, 30 mM imidazole, and Pierce™ Protease Inhibitor Tablet). Cell resuspension was lysed by sonication for 5 minutes (10 seconds per cycle). Lysates were clarified by centrifugation at 24,000g for 40 minutes at 12°C and pre-equilibrated with 1 mL of Ni-NTA nickel agarose for 1 hour at 4°C. The resin was washed twice with 10 mL of wash buffer and then eluted with 1 mL of elution buffer (20 mM Tris-HCl pH 8.0, 300 mM NaCl, 300 mM imidazole). The eluted protein was purified by size-exclusion chromatography in PBS. Fractions were collected based on A280 recording, flash frozen in liquid nitrogen, and stored at -80°C.

[0067] 3. Computational design of ideal scaffolds The generation of our ideal NTF2 scaffold can be divided into four parts: (3.1) generation of seed structure, (3.2) optimization of main-chain geometry using trRosetta™-based hallucination, (3.3) generation of structure-constrained sequence models for biased design, and (3.4) design and selection.

[0068] 3.1 Seed structure generation We use trRosetta(TM) 22 We sought to increase the set of NTF2 structures by complementing the experimentally solved structures from the PDB with highly accurate models generated by . To achieve this, we first collected 85 NTF2-like protein structures from the PDB based on SCOPe annotation (d.17.4 SCOPe v2.05). The corresponding sequences were then used as queries to collect sequence homologs from UniProt™ by performing eight iterations of hhblits against the uniclust30_2018_08 database with a 1e-20 e-value cutoff; the default sorting cutoff was relaxed to maximize the number of output hits (-maxfilt 100000000 -neffmax 20 -nodiff -realign_max 10000000). All hits were matched using cd-hit with a 60% sequence identity cutoff. 46 The redundancy was reduced using, yielding a set of 7,573 candidates for modeling.

[0069] To generate inputs for structural modeling with trRosetta™, we constructed multiple sequence alignments (MSAs) for each of the 7,573 selected sequences with hhblits using a more conservative e-value cutoff of 1e-50; the resulting MSAs were also complemented with hits from hmmsearch against uniref100 (release-2019_11) using a bit-score threshold of 115 (i.e., approximately 1 bit per position). After concatenating the two sets of alignments and filtering them at 90% sequence identity and 75% coverage cutoff, only sequences with more than 50 homologs in the corresponding MSAs were retained for modeling (2,005 sequences). The filtered MSAs, together with information about the top 25 putative structural homologs identified by hhsearch, were then filtered using trRosetta™. 47 The distances and orientations of residue pairs were predicted against the PDB100 database of templates as input to a template-aware version of . The network predictions were then analyzed using a Rosetta™-based folding protocol as previously described. 22 was used to reconstruct the all-atom 3D structural model.

[0070] 3.2 Ideal hallucination of NTF2 Exploring idealization of natural structure seeds, we determined that trRosetta™, a convolutional residual neural network that predicts residue-residue orientation and distance from sequence, could serve as a key component in a protein idealizer. Previously, this network was used to generate a variety of proteins that resembled de novo designed proteins of "ideal" structure by altering protein sequences to optimize the contrast (KL disparity) between the shapes of predicted and randomly generated sequences. 19 .

[0071] For our purposes, the target folding space was not diverse, but instead focused on NTF2-like topologies. To ensure the generation of ideal structures within this folding space, we implemented a novel fold-specific loss function that was biased towards hallucination based on shapes observed in native crystal structures. Since many experimentally characterized NTF2s contain non-ideal regions, we started by creating a truncated set of non-ideal NTF2s (Χ) by manually removing non-ideal structural elements such as twisted helices and long or rarely observed loops. For each seed structure, we then found equivalent positions between the seed structure and Χ using a structure-based sequence alignment method (see 3.3). A residue pair was considered to be in a conserved tertiary motif (TERM) if it had 5 or more equivalent positions in Χ. A smooth probability distribution based on the shapes observed in Χ was then calculated. For distances, we used a Gaussian distribution with a mean equal to the true distance, denoted by D, and a standard deviation, denoted by σ, equal to 0.5 Å. The probability density function for distance d is given by:

number

[0072] Using this density function, we can construct a categorical distribution for the binned distances by evaluating this function at the bin centers and then normalizing by the sum of all the values ​​in the different bins. Similarly, we can write the von Mises distribution as

number

number

number

number

number

[0073] We used a Markov Chain Monte Carlo (MCMC) method to search for sequences that trRosetta™ predicted would fold into a structure that minimized this loss function. We allowed measures with four types of different sampling probabilities: mutations (p=0.55), insertions (p=0.15), deletions (p=0.15), and moving segments (p=0.15). Mutations randomly changed one amino acid to another, with equal transition probability for all 20 amino acids. Insertions inserted new amino acids at random positions that underwent KL information loss (all to the same degree). Deletions removed random residues from the same positions. Finally, we also allowed the "segments" to move, cutting and pasting themselves from one part of the sequence to another while maintaining the same overall segment order. Here, the "segments" are contiguous runs of amino acids that all underwent folding-specific loss, often consisting of single strands or helices. Starting with a random sequence of initial length (usually 120 amino acids), we use the standard Metropolis criteria:

number

[0074] 3.3 Structure-constrained multiple sequence alignment Given the complexity of NTF2-like protein folding, we hypothesized that it would be necessary to impose sequence design rules that reject alternative states (negative designs). To this end, we calculated structure-constrained multiple sequence alignments based on NTF2-like proteins. In particular, we used TMalign to overlay each of the 2005 predicted native structures (from 3.1) onto each hallucination backbone (from 3.2). 48Next, to find the structurally corresponding positions, we used the Needleman-Wunsch algorithm. 49 We implemented a structure-based dynamic programming algorithm similar to that of . However, instead of using amino acid similarity as the scoring metric, we used an adjustable structure-based score function. After aligning the two structures, we scored the structural similarity of any two residues by several empirically weighted metrics: (1) the distance between the Ca atoms, (2) the difference between the main-chain torsion angles (φ and ψ), and (3) the angle (°) between the vectors pointing from Cα to Cβ in each residue. To calculate an unweighted score for each component, we normalized each by the maximum possible value (180° for angles and 10 Å for distances) and included a "set point" that roughly delineated where we judged the metric to indicate that the two residues were more similar. Values ​​above this set point were positive, indicating that the two residues were similar, and values ​​below the set point indicated that the two residues were dissimilar.

number

[0075] Each value was scaled by its normalized weight and summed to give an overall similarity score between any two amino acids. These similarity scores were used as the similarity metric in our dynamic programming algorithm instead of the typical BLOSUM62 similarity metric. We used a gap penalty of 0.1 and an extension penalty of 0.0. Finally, after concatenating all the structure-conditionally aligned sequences, we used the PSI-BLAST-exB 50,51 was used to calculate a redundancy-weighted log-odds score for each amino acid at each position (position-specific scoring matrix, PSSM).

[0076] 3.4 Design To design the resulting backbones, we explored further specifying the backbone structure and functionalizing the pocket by incorporating a complete hydrogen-bonding network from a native NTF2-like protein, in addition to the sequence patterns captured in PSSM (3.3). We compiled two sets of hydrogen-bonding networks: a set of 85 networks that included the cavity, and another set of networks that connected the C-terminal region of the first helix with the third β-strand, containing 25 networks. In 20 independent trials for each backbone, we randomly grafted one network from each set, fixed the identity of the hydrogen-bonding residues, and designed sequences for all other positions under the PSSM constraints. The resulting models were screened for various backbone quality metrics and for preservation of the hydrogen-bonding network in the absence of constraints, yielding 1615 ideal scaffolds.

[0077] 4.RIFdock adjustment file RIFdock's hierarchical search framework is a powerful method for searching 6D rigid-body orientations. Although originally designed to handle physics-based force fields, its scoring mechanism can be easily modified to do other things. A system called "tuning files" has been added that allows tuning the energy of RIFdock by "requiring" specific interactions. Specific interactions can range from specific hydrogen-bonding interactions, to specific bidentate interactions, and even specific hydrophobic interactions. The details are that during the RIFgen stage, each pooled rotamer is compared against a list of definitions in the tuning file. If a rotamer meets the definitions, it is saved in the RIF with a "requirement number". At later stages of RIFdock, these requirement numbers are available during scoring, and the presence or absence of specific rotamer interactions can be used to penalize or even completely discard the docking solution. In this work, tuning files were used to require specific hydrogen-bonding interactions between arginine and a secondary amine in the pyrazine ring of a coelenterazine-like substrate.

[0078] 5. Design of Theozyme Structure in a Novel NTF2 Scaffold The de novo design of luciferase can be divided into three major steps: scaffold construction, placement of substrates with required interactions, and sequence design. Obtaining an ideal NTF2-like scaffold, we selected five diverse rotamers from AIMNet and used the Rotamer Interaction Field (RIF) docking method to comprehensively search a large space of interacting side chains for the anionic form of DTZ. 31 Chemically, deprotonation of the N1 hydrogen is the first step to form an anionic species (Figure 5). We first used the RifGen ion-guiding configuration in the protein scaffold. 31 We generated the RIF using the RT-PCR tool, where we required the placement of a positively charged arginine side chain to stabilize the formation of the negatively charged N1 atom, where deprotonation occurs first, and enumerated a number of possible side chain interactions with the rest of the DTZ. HBOND_DEFINITION N1 1 ARG END_HBOND_DEFINITION REQUIREMENT_DEFINITION 1 HBOND N1 1 END_REQUIREMENT_DEFINITION

[0079] RIFdock was then used to hierarchically search for the best combination of RIFs to place on the input backbone. Although the negative charge can be transferred to another electronegative atom O1 via resonance of the imidazopyrazinon core, it is not clear which anionic species is more important for luciferase-catalyzed luminescence emission. Therefore, we let RIFdock place polar rotamers relative to the hydrogen bond placement to O1 and non-polar rotamers relative to DTZ without any specific requirement. In the next docking step, we parsed the -scaffold_res argument with a list of residue numbers as positions of the scaffold backbone annotated as pocket residues to allow hierarchical search of RIF placements. We only allowed RIF placements in pocket residues and left the predefined hydrogen bond network (HBNet) unprocessed. After RIFdock, we continued with Rosetta™ sequence design, where we changed the score function from higher buried_unsat_penalty to higher buried_unsat_penalty. 52 The amino acid selection was biased by feeding a pre-generated PSSM file via the SeqprofConsensus task operation, which was weighted again for the core region known to be beneficial as a catalytic pocket. 53 This can minimize unfilled buried residues and increase pre-organized structures in the first round. Two rounds of FastDesign calculations were included: we limited the RIF rotamers and core HBNet to repacking in the first round, while we allowed other residues for PSSM-based redesign during the Monte Carlo simulated annealing method. After the surrounding residues were optimized to retain RIF interactions, we allowed redesign of the RIF rotamers to find efficient aromatic and hydrophobic packing around the DTZ, while still restricting the catalytic residues (N1 requirement) to repacking only. The final set of designs was obtained after screening by ligand binding boundary energy, shape complementarity, contact molecular surface, HbondsToResidue, and the presence of N1_hbond.

[0080] 6. LuxSit structure prediction by AlphaFold2 and comparison with design model To computationally evaluate the accuracy of the LuxSit designed models, we performed single sequence structure prediction using AlphaFold2. All models were run with 12 recycles and the generated models were compared using AMBER. 54 The model with the best pLDDT was used for comparison with the Rosetta™ design model, and structural superposition was performed using the Theseus alignment tool to align the design model with the AlphaFold2 model. 33 The main chain RMSD between was measured.

[0081] 8. Computational SSM experiments to estimate mutant binding free energies Rosetta(TM) Cartesian_ddg Application 55,56 The enzyme and substrate binding free energies were computationally estimated using the LuxSit design model. The LuxSit design model was previously relaxed in the substrate-bound Cartesian space. For the 21 positions experimentally screened for the effect of single mutations on luciferase activity, each residue was computationally mutated to other amino acid types and packings, and Cartesian relaxation was performed to evaluate the final score in REU. The method was applied three times in parallel to both the substrate-bound and apo states. The average of the triplicate results was used to estimate the relative binding free energy (ddG bind ) was used to calculate

[0082] 9. Construction and Screening of Engineered Luciferase Libraries The construction of the assembled gene library was previously described in detail. 57Briefly, the amino acid sequences of all designed luciferases were first reverse-translated in E. coli. All DNA sequences were sorted into multiple subpools according to gene length (approximately 500 designs per subpool). Each gene was then split into two fragments (fragment A and fragment B) and outer and inner primer sequences were added at the 5' and 3' ends (e.g., Outer_oligoA_5primer + design_half_A + Inner_oligoA_3primer and Inner_oligoB_5primer + design_half_B + Outer_oligoB_3primer). All oligos were organized into one Twist 250nt Oligo Pool. To construct the libraries, individual fragments A or B were amplified from each subpool using polymerase chain reaction (PCR) with oligoA_5primer / oligoA_3primer or oligoB_5primer / oligoB_3primer oligonucleotide pairs. Pool-specific sequences were removed with Uracil Specific Excision Reagent (USER) followed by NEB end repair kit. The outer primers (oligoA_5primer and oligoB_3primer) were then used to assemble and amplify fragment A and fragment B. The assembled full-length fragments were digested with XhoI / HindIII and ligated into predigested pBAD / His B vector. All ligation products were used to transform ElectroMAX™ DH10B Cells, which were then plated onto 150mmx15mm LB agar plates supplemented with carbenicillin and L-arabinose. We sequenced 30 random colonies, and 11 sequences were present in our designed library. Plates (approximately 2000 colonies per plate) were incubated overnight at 37°C to allow bacterial colonies to form and left at 4°C for an additional 24 hours.To image luminescence activity directly from bacterial colonies, we sprayed each agar plate with 30 μM DTZ in PBS for 2 min, and luminescence images were acquired and processed with a Bio-Rad ChemiDoc XRS+. After screening 15 plates, active colonies were collected for sequence analysis, protein expression, and other downstream characterization, and LuxSit was selected from three active designs that showed catalytic signals in the above background.

[0083] 10. Construction and Evaluation of LuxSit Site-saturation Mutagenesis Libraries A mixture of forward oligos with degenerate codons (NDT, VHG, and TGG = 1:1:0.1 ratio) and overlapping reverse oligos were used to amplify the LuxSit plasmid to generate libraries of each single amino acid substitution at residues 13, 14, 17, 18, 35, 37, 38, 49, 52, 53, 56, 60, 65, 81, 83, 94, 96, 98, 100, 110, and 112. The resulting PCR products were circularized by the Gibson Assembly protocol and subsequently used to transform ElectroMAX™ DH10B Cells. Cells were plated on 150mmx15mm LB agar plates supplemented with carbenicillin and L-arabinose, incubated overnight at 37°C, and left at 4°C for an additional 24 hours. Colony-based screening by spraying with DTZ solution was used to identify active colonies as described in Screening of the Luciferase Library. Inactive colonies were also randomly selected. As a result, a total of 32 colonies were selected for each residue library. 32x21 individual colonies were grown in 1mL of TB supplemented with carbenicillin and L-arabinose in 96-well deep-well culture plates. Plates were shaken at 1,100 rpm on a 96-well plate shaker at 37°C overnight (approximately 16-18 hours). Cells were pelleted by centrifugation at 4,000g for 15 minutes in a tabletop centrifuge. The medium was discarded and the cell pellet was resuspended in 0.2mL of BugBuster HT protein extraction buffer. Plates were returned to the 96-well plate shaker and incubated at 1,100 rpm for an additional 30 minutes. Cell debris was pelleted again by centrifugation at 4,000g for 15 minutes and the soluble lysate was transferred to a new semi-deep 96-well plate and incubated with 10μL of magnetic Ni-NTA beads for 30 minutes to allow binding. Using a magnetic extractor, the beads were first transferred from the binding plate to a wash plate containing 200 μL of IMAC wash buffer per well, and then the beads were transferred to an elution plate containing 30 μL of IMAC wash buffer per well. The concentration of total protein in each well was measured by direct Bradford assay.The elution solution in each well was used to make 25 μL of protein solution at the indicated concentration and mixed with 25 μL of 50 μM DTZ in PBS. Luminescence signals were acquired for 15 min during which the actual point mutations were identified by sequencing. Thus, mutation-versus-activity relationships could be mapped. To assess whether these beneficial mutations were synergistic, we compiled individual mutants with combined mutations at residues 14, 60, 96, 98, and 110 (see Table 4) and expressed and purified these LuxSit mutants for kinetics, emission spectra, and luminescence intensity. We identified four mutants that produced 47-77 times more photons than the parent LuxSit. We designated one of these as LuxSit-f(A96M / M110V) due to its strong initial flash emission. Because mutations at residues 96 and 110 are robust and residue 60 is versatile, we generated a fully randomized library at positions 60, 96, and 110 to exhaustively explore all possible combinations. After colony-based screening, we identified many colonies with strong luciferase activity against DTZ (Figure 9). Among all selected mutants, Arg60 was confirmed to be mutagenic, Ala96 prefers larger hydrophobic side chains (Leu, Ile, Met, and Cys), and Met110 prefers hydrophobic residues (Val, Ile, and Ala). The newly discovered mutant R60S / A96L / M110V, which has 100-fold higher photon flux versus LuxSit, was designated LuxSit-i due to its high brightness.

[0084] 11. In Vitro Characterization of Photoluminescence Properties For Michaelis-Menten kinetic measurements, 25 μL of serially diluted DTZ substrate in Tris pH 8.0 buffer was added into wells of a white 96-well half-area microplate containing 25 μL of purified luciferase (final enzyme concentration: 100 nM; substrate concentration: 0.78–50 μM). Measurements were taken every minute (0.1 s integration at each interval, 10 s shaking) for a total of 20 min. Initial kinetics were estimated as the average of the light intensity from the first three data points fitted to the Michaelis-Menten equation. All relative arbitrary units (RLU) per second were calculated using the luminol-H2O2-HRP calibration method. 45 was converted to photons / s using the following formula: max =LQYxk cat x[E], I max is the maximum photon flux (photons s -1 ), [E] is the total enzyme concentration, and V max is the maximum photon flux / molecule (photons s) from fitting the Michaelis-Menten equation -1 molecule -1 ). To measure the luminescence quantum yield, 25 μL of 5 μM of the individual substrates in PBS were injected into 25 μL of PBS containing 100 nM of the corresponding luciferase. DTZ was used for all LuxSit variants, while CTZ was used as the native RLuc substrate. The luminescence signal was monitored until the reaction was complete (0.1 s integration and measurements were performed every 5 s for a total of 40 min). The sum of the luminescence photon counts was normalized to the total photon counts of the RLuc / CTZ pair (LQY=5.3±0.1%). 30 The relative luminescence quantum yields of the LuxSit mutants were derived (Fig. 10c). cat The value is calculated by the formula: cat =V maxCalculated using / LQY. To record emission spectra, 50 μM DTZ in 25 μL PBS was injected into 25 μL 200 nM pure luciferase and emission spectra were collected with 0.1 s integration and 2 nm increments between 300 and 700 nm. In vitro luminescence activity measurements of HEK293T or HeLa cells expressing LuxSit-i were performed similarly when 15,000 untreated cells or lysates were used in the assay instead of purified luciferase. To assess substrate specificity, 50 μM of substrate analogs in 25 μL PBS were added to 25 μL 200 nM of the indicated luciferase and the signal was recorded for 20 min. Data were presented as total luminescence signal over the first 10 min. We normalized the data by setting the most luminescent substrate to 100%.

[0085] 12. Circular dichroism spectroscopy (CD) Purified protein samples were prepared at 15 μM in 10 mM phosphate buffer, pH 7.4. Spectra from 190 nm to 260 nm were recorded at 25° C., 50° C., 75° C., 95° C., and upon cooling back to 25° C. Thermal denaturation was monitored at 220 nm from 25° C. to 95° C. (increments of 1° C. per minute). Values ​​were not reported as there was no clear inflection point in the melting curve.

[0086] 13. Mammalian Cell Culture and Transfection HEK293T and HeLa cell lines were maintained at 37°C in a humidified 5% CO2 atmosphere and cultured in Dulbecco's Modified Eagle's Medium (DMEM, GIBDO) supplemented with 10% fetal bovine serum (FBS, Sigma). Cells were transfected with Turbofectin™ 8.0 (Origene) containing 500 μg of plasmid DNA. After 24 hours at 37°C in a CO2 incubator, the medium was removed and cells were harvested and resuspended in Dulbecco's Phosphate Buffered Saline (DPBS).

[0087] 14. Fluorescence Microscopy and Image Analysis Cells were washed twice with HBSS and subsequently imaged in HBSS at 37°C in the dark. Immediately prior to imaging, cells were incubated with 25 μM DTZ. Epifluorescence imaging was performed on a Yokogawa CSU-X1 microscope equipped with a Hamamatsu ORCA-Fusion scientific CMOS camera and a Lumencor Celesta light engine. The objectives used were: 10x, NA 0.45, WD 4.0 mm, 20x, NA 1.4, WD 0.13 mm, and 40x, NA 0.95, WD 0.17-0.25 mm, with correction collars (Plan Apochromat Lambda) for cover glass thickness (0.11 mm-0.23 mm). Imaging of BFP used a 408 nm laser, 436 / 36 nm dichroic, and 440 / 40 nm emission filter (Semrock). Exposure times were 200 ms for BFP and 10 s for luminescence. All epifluorescence experiments were then analyzed using NIS Elements software.

[0088] 15. Multiplexed dual luciferase reporter assay for cAMP / pKa and NF-κB pathways HEK293T cells were grown in tissue culture grade white 96-well plates and transfected with the indicated Cre-RLuc, NFkB-LuxSit-i, and CMV-CyOFP plasmids. 24 hours after transfection, medium was replaced with 2 μM forskolin (FSK) or 300 ng / mL human tumor necrosis factor alpha (TNFα) in regular cell medium. 23 hours after stimulation, cells were resuspended in DPBS by pipette mixing. 30,000 untreated cells in 25 μL DPBS were mixed with 25 μL CelLytic M for 15 min to generate cell lysates. For untreated cell assays, 15,000 untreated cells in 25 μL DPBS were mixed with 25 μL PP-CTZ (2 μM) and / or DTZ (10 μM) in DPBS. For cell lysate assays, 25 μL of cell lysate was added to 25 μL of PP-CTZ (2 μM) and / or DTZ (10 μM) to initiate the luminescence reaction. Signals were recorded every minute for a total of 10 min. Optical signals were collected in substrate decomposition mode without filters and in spectral decomposition mode with 528 / 20 and 390 / 35 filters. Area scanning of the fluorescence intensity of CyOFP at 480 nm (excitation wavelength) and 580 nm (emission wavelength) was used to estimate the total cell number and transfection efficiency. The reported units were the average of the first 10 min of luminescence (RLU), higher than the relative fluorescence unit (au). To derive fold activation, all data were normalized to the corresponding unstimulated control.

[0089] statistical analysis No statistical methods were used to predetermine sample sizes. No samples were excluded from data analysis. Results were reproduced with different batches of pure protein on different days. Data are presented as mean ± sd unless otherwise indicated, and error bars in figures represent the sd of technical triplicates. Data were analyzed and plotted using GraphPad Prism8, seaborn, and matplotlib.

[0090] Supplementary Information

Table 5

Table 6

[0091] References 1. Jiang, L., et al. De novo computational design of retro-aldol enzymes. Science 319, 1387 - 1391 (2008). 2. Rothlisberger, D., et al. Kemp elimination catalysts by computational enzyme design. Nature 453, 190 - 195 (2008). 3. Yeh, H. W., et al. Red-shifted luciferase-luciferin pairs for enhanced bioluminescence imaging. Nat. Methods 14, 971 - 974 (2017). 4. Love, A. C. & Prescher, J. A. Seeing (and Using) the Light: Recent Developments in Bioluminescence Technology. Cell Chemical Biology 27, 904 - 920 (2020). 5. Syed, A. J. & Anderson, J. C. Applications of bioluminescence in biotechnology and beyond. Chem. Soc. Rev. 50, 5668 - 5705 (2021). 6.Yeh,H.-W.& Ai,H.-W.Development and Applications of Bioluminescent and Chemiluminescent Reporters and Biosensors.Annu.Rev.Anal.Chem.12,129-150(2019). 7.Zambito,G.,Chawda,C.& Mezzanotte,L.Emerging tools for bioluminescence imaging.Curr.Opin.Chem.Biol.63,86-94(2021). 8.Markova,S.V.,Larionova,M.D.& Vysotski,E.S.Shining Light on the Secreted Luciferases of Marine Copepods:Current Knowledge and Applications.Photochem.Photobiol.95,705-721(2019). 9.Wu,N.et al.Solution structure of Gaussia Luciferase with five disulfide bonds and identification of a putative coelenterazine binding cavity by heteronuclear NMR.Sci.Rep.10,(2020). 10.Jiang,T.Y.,Du,L.P.& Li,M.Y.Lighting up bioluminescence with coelenterazine:strategies and applications.Photochem.Photobiol.Sci.15,466-480(2016). 11.Shakhmin,A.et al.Coelenterazine analogues emit red-shifted bioluminescence with NanoLuc.Org.Biomol.Chem.15,8559-8567(2017). 12.Michelini,E.et al.Spectral-resolved gene technology for multiplexed bioluminescence and high-content screening.Anal.Chem.80,260-267(2008). 13.Rathbun,C.M.et al.Parallel screening for rapid identification of orthogonal bioluminescent tools.ACS Cent.Sci.3,1254-1261(2017). 14.Yeh,H.-W.,Wu,T.,Chen,M.& Ai,H.-W.Identification of Factors Complicating Bioluminescence Imaging.Biochemistry 58,1689-1697(2019). 15.Su,Y.C.et al.Novel NanoLuc substrates enable bright two-population bioluminescence imaging in animals.Nat.Methods 17,852-860(2020). 16.Lombardi,A.,Pirro,F.,Maglio,O.,Chino,M.& DeGrado,W.F.De Novo design of four-helix bundle metalloproteins:One scaffold,diverse reactivities.Acc.Chem.Res.52,1148-1159(2019). 17.Chino,M.et al.Artificial diiron enzymes with a DE Novo designed four-helix bundle structure.Eur.J.Inorg.Chem.2015,3352-3352(2015). 18.Basler,S.et al.Efficient Lewis acid catalysis of an abiological reaction in a de novo protein scaffold.Nat.Chem.13,231-235(2021). 19.Anishchenko,I.et al.De novo protein design by deep network hallucination.Nature(2021)doi:10.1038 / s41586-021-04184-w. 20.Wang,J.et al.Scaffolding protein functional sites using deep learning.Science 377,387-394(2022). 21.Norn,C.et al.Protein sequence design by conformational landscape optimization.Proc.Natl.Acad.Sci.U.S.A.118,(2021). 22.Yang,J.Y.et al.Improved protein structure prediction using predicted interresidue orientations.Proc.Natl.Acad.Sci.U.S.A.117,1496-1503(2020). 23.Basanta,B.et al.An enumerative algorithm for de novo design of proteins with diverse pocket structures.Proc.Natl.Acad.Sci.U.S.A.117,22135-22145(2020). 24.Loening,A.M.,Fenn,T.D.& Gambhir,S.S.Crystal structures of the luciferase and green fluorescent protein from Renilla reniformis.J.Mol.Biol.374,1017-1028(2007). 25.Tomabechi,Y.et al.Crystal structure of nanoKAZ:The mutated 19 kDa component of Oplophorus luciferase catalyzing the bioluminescent reaction with coelenterazine.Biochem.Biophys.Res.Commun.470,88-93(2016). 26.Ding,B.W.& Liu,Y.J.Bioluminescence of Firefly Squid via Mechanism of Single Electron-Transfer Oxygenation and Charge-Transfer-Induced Luminescence.J.Am.Chem.Soc.139,1106-1119(2017). 27.Isobe,H.,Yamanaka,S.,Kuramitsu,S.& Yamaguchi,K.Regulation mechanism of spin-orbit coupling in charge-transfer-induced luminescence of imidazopyrazinone derivatives.J.Am.Chem.Soc.130,132-149(2008). 28.Kondo,H.et al.Substituent effects on the kinetics for the chemiluminescence reaction of 6-arylimidazo[1,2-a]pyrazin-3(7H)-ones(Cypridina luciferin analogues):support for the single electron transfer(SET)-oxygenation mechanism with triplet molecular oxygen.Tetrahedron Lett.46,7701-7704(2005). 29.Branchini,B.R.et al.Experimental Support for a Single Electron-Transfer Oxidation Mechanism in Firefly Bioluminescence.J.Am.Chem.Soc.137,7592-7595(2015). 30.Zubatyuk,R.,Smith,J.S.,Leszczynski,J.& Isayev,O.Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network.Science Advances 5,(2019). 31.Dou,J.Y.et al.De novo design of a fluorescence-activating beta-barrel.Nature 561,485-491(2018). 32.Cao,L.et al.Design of protein-binding proteins from the target structure alone.Nature 605,551-560(2022). 33.Jumper,J.et al.Highly accurate protein structure prediction with AlphaFold.Nature 596,583-+(2021). 34.Dauparas,J.et al.Robust deep learning based protein sequence design using ProteinMPNN.bioRxiv(2022)doi:10.1101 / 2022.06.03.494563. 35.Yeh,H.-W.et al.ATP-Independent Bioluminescent Reporter Variants To Improve in Vivo Imaging.ACS Chem.Biol.14,959-965(2019). 36.Bhaumik,S.& Gambhir,S.S.Optical imaging of Renilla luciferase reporter gene expression in living mice.Proc.Natl.Acad.Sci.U.S.A.99,377-382(2002). 37.Szent-Gyorgyi,C.,Ballou,B.T.,Dagnal,E.& Bryan,B.Cloning and characterization of new bioluminescent proteins.in Biomedical Imaging:Reporters,Dyes,and Instrumentation(eds.Bornhop,D.J.,Contag,C.H.& Sevick-Muraca,E.M.)(SPIE,1999).doi:10.1117 / 12.351015. 38.Hall,M.P.et al.Engineered luciferase reporter from a deep sea shrimp utilizing a novel imidazopyrazinone substrate.ACS Chem.Biol.7,1848-1857(2012). 39.Baek,M.et al.Accurate prediction of protein structures and interactions using a three-track neural network.Science 373,871-+(2021). 40.Giger,L.et al.Evolution of a designed retro-aldolase leads to complete active site remodeling.Nat.Chem.Biol.9,494-498(2013). 41.Yao,Z.et al.Multiplexed bioluminescence microscopy via phasor analysis.Nat.Methods 19,893-898(2022). 42.Loening,A.M.,Dragulescu-Andrasi,A.& Gambhir,S.S.A red-shifted Renilla luciferase for transient reporter-gene expression.Nat.Methods 7,5-6(2010). 43.Dijkema,F.M.et al.Flash properties of Gaussia luciferase are the result of covalent inhibition after a limited number of cycles.Protein Sci.30,638-649(2021). 44.Schenkmayerova,A.et al.Engineering the protein dynamics of an ancestral luciferase.Nat.Commun.12,(2021). 45.Ando,Y.et al.Development of a quantitative bio / chemiluminescence spectrometer determining quantum yields:Re-examination of the aqueous luminol chemiluminescence standard.Photochem.Photobiol.83,1205-1210(2007). 46.Li,W.& Godzik,A.Cd-hit:a fast program for clustering and comparing large sets of protein or nucleotide sequences.Bioinformatics 22,1658-1659(2006). 47.Farrell,D.P.et al.Deep learning enables the atomic structure determination of the Fanconi Anemia core complex from cryoEM.IUCrJ 7,881-892(2020). 48.Zhang,Y.& Skolnick,J.TM-align:a protein structure alignment algorithm based on the TM-score.Nucleic Acids Res.33,2302-2309(2005). 49.Needleman,S.B.& Wunsch,C.D.A general method applicable to the search for similarities in the amino acid sequence of two proteins.J.Mol.Biol.48,443-453(1970). 50.Oda,T.,Lim,K.& Tomii,K.Simple adjustment of the sequence weight algorithm remarkably enhances PSI-BLAST performance.BMC Bioinformatics 18,288(2017). 51.Altschul,S.F.et al.Gapped BLAST and PSI-BLAST:a new generation of protein database search programs.Nucleic Acids Res.25,3389-3402(1997). 52.Coventry,B.& Baker,D.Protein sequence optimization with a pairwise decomposable penalty for buried unsatisfied hydrogen bonds.PLoS Comput.Biol.17,(2021). 53.Smith,A.J.T.et al.Structural Reorganization and Preorganization in Enzyme Active Sites:Comparisons of Experimental and Theoretically Ideal Active Site Geometries in the Multistep Serine Esterase Reaction Cycle.J.Am.Chem.Soc.130,15361-15373(2008). 54.Salomon-Ferrer,R.,Case,D.A.& Walker,R.C.An overview of the Amber biomolecular simulation package.Wiley Interdiscip.Rev.Comput.Mol.Sci.3,198-210(2013). 55.Kellogg,E.H.,Leaver-Fay,A.& Baker,D.Role of conformational sampling in computing mutation-induced changes in protein structure and stability.Proteins 79,830-838(2011). 56.Park,H.et al.Simultaneous optimization of biomolecular energy functions on features from small molecules and macromolecules.J.Chem.Theory Comput.12,6201-6212(2016). 57.Klein,J.C.et al.Multiplex pairwise assembly of array-derived DNA oligonucleotides.Nucleic Acids Res.44,(2016). 58.Loening,A.M.,Wu,A.M.& Gambhir,S.S.Red-shifted Renilla reniformis luciferase variants for imaging in living subjects.Nat.Methods 4,641-643(2007). 59.Liang,J.,Feng,X.,Hait,D.& Head-Gordon,M.Revisiting the performance of time-dependent density functional theory for electronic excitations:Assessment of 43 popular and recently developed functionals from rungs one to four.J.Chem.Theory Comput.18,3460-3473(2022). 60.Chai,J.-D.& Head-Gordon,M.Long-range corrected hybrid density functionals with damped atom-atom dispersion corrections.Phys.Chem.Chem.Phys.10,6615-6620(2008). 61.Ditchfield,R.,Hehre,W.J.& Pople,J.A.Self-consistent molecular-orbital methods.IX.An extended Gaussian-type basis for molecular-orbital studies of organic molecules.J.Chem.Phys.54,724-728(1971). 62.Grimme,S.Exploration of chemical compound,conformer,and reaction space with meta-dynamics simulations based on tight-binding quantum chemical calculations.J.Chem.Theory Comput.15,2847-2862(2019). 63.Pracht,P.,Bohle,F.& Grimme,S.Automated exploration of the low-energy chemical space with fast quantum chemical methods.Phys.Chem.Chem.Phys.22,7169-7192(2020). 64.Luchini,G.,Alegre-Requena,J.V.,Funes-Ardoiz,I.& Paton,R.S.GoodVibes:automated thermochemistry for heterogeneous computational chemistry data.F1000Res.9,291(2020). 65.Li,Y.-P.,Gomes,J.,Mallikarjun Sharada,S.,Bell,A.T.& Head-Gordon,M.Improved force-field parameters for QM / MM simulations of the energies of adsorption for molecules in zeolites and a free rotor correction to the rigid rotor harmonic oscillator model for adsorption enthalpies.J.Phys.Chem.C Nanomater.Interfaces 119,1840-1850(2015). 66.Goetz,A.W.et al.Routine microsecond molecular dynamics simulations with AMBER on GPUs.1.Generalized Born.J.Chem.Theory Comput.8,1542-1555(2012). 67.Becke,A.D.Density-functional thermochemistry.III.The role of exact exchange.J.Chem.Phys.98,5648-5652(1993). 68.Grimme,S.,Antony,J.,Ehrlich,S.& Krieg,H.A consistent and accurate ab initio parametrization of density functional dispersion correction(DFT-D)for the 94 elements H-Pu.J.Chem.Phys.132,154104(2010). 69.Grimme,S.,Ehrlich,S.& Goerigk,L.Effect of the damping function in dispersion corrected density functional theory.J.Comput.Chem.32,1456-1465(2011). 70.Meiler,J.& Baker,D.ROSETTALIGAND:Protein-small molecule docking with full side-chain flexibility.Proteins 65,538-548(2006). 71.Davis,I.W.& Baker,D.RosettaLigand docking with full ligand and receptor flexibility.J.Mol.Biol.385,381-392(2009). 72.Davis,I.W.,Raha,K.,Head,M.S.& Baker,D.Blind docking of pharmaceutically relevant compounds using RosettaLigand.Protein Sci.18,1998-2002(2009). 73.Wang,J.,Wolf,R.M.,Caldwell,J.W.,Kollman,P.A.& Case,D.A.Development and testing of a general amber force field.J.Comput.Chem.25,1157-1174(2004). 74.Bayly,C.I.,Cieplak,P.,Cornell,W.& Kollman,P.A.A well-behaved electrostatic potential based method using charge restraints for deriving atomic charges:the RESP model.J.Phys.Chem.97,10269-10280(1993). 75.Besler,B.H.,Merz,K.M.& Kollman,P.A.Atomic charges derived from semiempirical methods.J.Comput.Chem.11,431-439(1990). 76.Singh,U.C.& Kollman,P.A.An approach to computing electrostatic charges for molecules.J.Comput.Chem.5,129-145(1984). 77.Jorgensen,W.L.,Chandrasekhar,J.,Madura,J.D.,Impey,R.W.& Klein,M.L.Comparison of simple potential functions for simulating liquid water.J.Chem.Phys.79,926-935(1983). 78.Maier,J.A.et al.Ff14SB:Improving the accuracy of protein side chain and backbone parameters from ff99SB.J.Chem.Theory Comput.11,3696-3713(2015). 79.Darden,T.,York,D.& Pedersen,L.Particle mesh Ewald:AnN·log(N)method for Ewald sums in large systems.J.Chem.Phys.98,10089-10092(1993). 80.Roe,D.R.& Cheatham,T.E.,III.PTRAJ and CPPTRAJ:Software for processing and analysis of molecular dynamics trajectory data.J.Chem.Theory Comput.9,3084-3095(2013).

Claims

1. A protein having luciferase activity comprising the secondary structure configuration H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, where "H" is a helical domain, "L" is a loop domain, and "E" is a β-strand domain; (a) the H1 domain is at least 18 or 19 amino acids in length; residue 14 of the H1 domain is Y, D, or E and residue 18 of said H1 domain is D or E; (b) the E3 domain is at least 6, 7, 8, 9, or 10 amino acids in length and residue 2 of the E3 domain is R; (c) the E5 domain is at least 10, 11, 12, 13, or 14 amino acids in length, and residue 9 of the E5 domain is H or N; protein.

2. The protein described in claim 1, wherein residue 7 of the E5 domain is M.

3. The protein described in claim 1, wherein the E6 domain is at least 9, 10, 11, 12, or 13 amino acids in length and residue 5 of the E6 domain is V.

4. The protein described in claim 1, wherein residue 1 of the L5 domain is S.

5. The protein described in claim 3, wherein residue 7 of the E5 domain is M and residue 5 of the E6 domain is V.

6. The protein described in claim 3, wherein residue 7 of the E5 domain is M, residue 5 of the E6 domain is V, and residue 1 of the L5 domain is S.

7. The protein of claim 1, wherein the H2 domain is at least 5, 6, or 7 amino acids in length, the H3 domain is at least 9, 10, 11, 12, 13, or 14 amino acids in length, the E1 domain is at least 3 or 4 amino acids in length, the E2 domain is at least 3 or 4 amino acids in length, and / or the E4 domain is at least 8, 9, 10, 11, or 12 amino acids in length.

8. The H1 domain is 19 amino acids in length; The H2 domain is 7 amino acids long; The E1 domain is 4 amino acids long; The E2 domain is 4 amino acids long; The H3 domain is 14 amino acids long; The E3 domain is 10 amino acids long; The E4 domain is 12 amino acids long; The E5 domain is 14 amino acids in length; and The E6 domain is 12 or 13 amino acids long. The protein of claim 1.

9. 10. The protein of claim 1, wherein one, two, three, four, or all five of the following are true: (a) residue 13 of domain H1 is F; (b) residue 1 of domain L3 is W; (c) residue 5 of domain E5 is V or another hydrophobic residue; (d) residue 8 of domain E5 is A or L or another hydrophobic residue; and / or (e) Residue 11 of domain E5 is W.

10. 9. The protein of claim 8, wherein one, two, three, four, five, or all six of the following are true: (a) residue 2 of domain E1 is I or another hydrophobic residue; (b) residue 4 of domain H3 is F; (c) residue 6 of domain E4 is V or another hydrophobic residue; (d) residue 8 of domain E4 is L or another hydrophobic residue; (e) residue 5 of domain E6 is M or V or another hydrophobic residue; and / or (f) Residue 7 of domain E6 is V or another hydrophobic residue.

11. 2. The protein of claim 1, comprising an amino acid sequence that is at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-181, or SEQ ID NOs: 1-3.

12. The protein of claim 11, comprising the amino acid sequence of SEQ ID NO:

4.

13. 1. A protein having luciferase activity, comprising an amino acid sequence that is at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO:1, Residue 14 is Y, D, or E and residue 98 is H or N; and Residue 18 is D or E and residue 65 is R; protein.

14. 14. The protein of claim 13, comprising one or both of the A96M and M110V substitutions relative to SEQ ID NO:

1.

15. 14. The protein of claim 13, comprising both an A96M and an M110V substitution relative to SEQ ID NO:

1.

16. 14. The protein of claim 13, comprising an R60S substitution relative to SEQ ID NO:

1.

17. 17. The protein of claim 16, comprising R60S, A96M, and M110V substitutions relative to SEQ ID NO:

1.

18. 14. The protein of claim 13, wherein any substitutions relative to SEQ ID NO: 1 at residues F12, 135, W38, F49, V81, L83, V94, A97, W100, M110, V112 are conservative amino acid substitutions.

19. 14. The protein of claim 13, comprising an amino acid sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-3, or SEQ ID NOs: 1-181.

20. 14. The protein of claim 13, wherein any substitutions relative to the reference sequence are conservative amino acid substitutions.

21. A protein comprising the formula X1-Z1-X2-Z2-X3-Z3-X4-Z4-X5-Z5-X6-Z6-X7-Z7-X8-Z8, X1 is, 【Chemistry 1】 wherein residue 14 is Y, D, or E and residue 18 is D or E; X2 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of ADTAASLF (SEQ ID NO: 183); X3 has an amino acid sequence at least 50%, 75%, or 100% identical to the amino acid sequence of TIHL (SEQ ID NO: 184); X4 has an amino acid sequence at least 33%, 66%, or 100% identical to the amino acid sequence of VTF; X5 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of EEFREWFERLFST (SEQ ID NO: 185); The X6 is 【Chemistry 2】 and wherein residue 2 is R; X7 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of VEVHVQLHATH (SEQ ID NO: 187); The X8 is 【Transformation 3】 wherein residue 8 is H or N; X9 has an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of VTEMRVHINPTG (SEQ ID NO: 189); and Z1, Z2, Z3, Z4, Z5, Z6, Z7, and Z8 are independently present or absent, and when present may comprise any amino acid sequence; protein.

22. One, two, three, four, five, six, seven, or all eight of the following are true: Z1 comprises SGD; Z2 contains HPGV (SEQ ID NO: 190); Z3 comprises WDG; Z4 includes TSR; Z5 comprises RKDA (SEQ ID NO: 191); Z6 includes GDT; Z7 contains NGQ; and / or Z8 comprises a GNR; and 22. The protein of claim 21 , wherein 0, 1, 2, 3, 4, 5, 6, 7, or all 8 of Z1, Z2, Z3, Z4, Z5, Z6, Z7, and Z8 further comprise an additional polypeptide domain.

23. 22. A self-complementary multipartite protein having luciferase activity comprising at least a first polypeptide component and a second polypeptide component, wherein at least said first polypeptide component and said second polypeptide component are not covalently linked, and wherein collectively at least said first polypeptide component and said second polypeptide component comprise the domains X1-Z1-X2-Z2-X3-Z3-X4-Z4-X5-Z5-X6-Z6-X7-Z7-X8-Z8-X9, each domain being as defined in claim 21; (a) each X domain is present entirely within at least one of the first and second polypeptide components; and (b) none of the first and second polypeptide components comprises each of X1, X2, X3, X4, X5, X6, X7, X8, and X9. Self-complementary multipartite proteins.

24. The self-complementary multi-split protein of claim 23, wherein the split occurs at Z4, Z5, Z6, or Z7.

25. 1. A self-complementary multipartite protein having luciferase activity comprising at least a first polypeptide component and a second polypeptide component, wherein at least said first polypeptide component and said second polypeptide component are not covalently linked, and wherein collectively at least said first polypeptide component and said second polypeptide component comprise the secondary structural arrangement H1-L1-H2-L2-E1-L3-E2-L4-H3-L5-E3-L6-E4-L7-E5-L8-E6, wherein each domain is as defined in claim 1; (a) each H and E domain is present entirely within at least one of the first and second polypeptide components; and (b) neither the first nor the second polypeptide component contains all of the H and E domains. Self-complementary multipartite proteins.

26. 26. The self-complementary multi-split protein of claim 25, wherein the split occurs at L4, L5, L6, L7, or L8.

27. (a) a protein or polypeptide component according to any one of claims 1 to 26; and (b) one or more additional functional domains; A fusion protein comprising:

28. A protein or polypeptide component according to any one of claims 1 to 26; or a fusion protein comprising said protein or polypeptide component and one or more additional functional domains; A nucleic acid encoding

29. 29. The nucleic acid of claim 28, comprising the nucleotide sequence of any one of SEQ ID NOs: 200 to 380.

30. 29. An expression vector comprising the nucleic acid of claim 28 operably linked to suitable control sequences.

31. A protein or polypeptide component according to any one of claims 1 to 26; a fusion protein comprising said protein or polypeptide component and one or more additional functional domains; a nucleic acid encoding said protein or polypeptide component, or said fusion protein; or an expression vector comprising said nucleic acid operably linked to a suitable control sequence; A recombinant host cell comprising:

32. (a) a protein or polypeptide component according to any one of claims 1 to 26; a fusion protein comprising said protein or polypeptide component and one or more additional functional domains; a nucleic acid encoding said protein or polypeptide component or said fusion protein; an expression vector comprising said nucleic acid operably linked to suitable control sequences; and / or a host cell comprising said protein or polypeptide component, said fusion protein, said nucleic acid, or said expression vector; and (b) instructions for their use; Kit including:

33. 33. The kit of claim 32, further comprising diphenylterazine (DTZ).

34. 27. A method for using a protein or polypeptide component according to any one of claims 1 to 26; a fusion protein comprising said protein or polypeptide component and one or more additional functional domains; a nucleic acid encoding said protein or polypeptide component or said fusion protein; an expression vector comprising said nucleic acid operably linked to a suitable control sequence; a host cell comprising said protein or polypeptide component, said fusion protein, said nucleic acid, or said expression vector; and / or a kit comprising said protein or polypeptide component, said fusion protein, said nucleic acid, said expression vector, or said host cell; for any suitable purpose, including but not limited to, luminescence reporting assays, diagnostic assays, cellular localization of a target of interest, cell imaging, gene editing, live animal imaging, cancer labeling, CAR T cell reporting, secretion assays, gene delivery, and tissue engineering.

35. A method for making a luciferase, including de novo design using the method of any embodiment disclosed herein, starting with a protein comprising the amino acid sequence of SEQ ID NO:381.