De novo designed bright and multi-color luciferases
De novo designed luciferases with improved properties address limitations of natural luciferases, enabling robust and multiplexed bioluminescence imaging in biomedical research.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Existing luciferases derived from natural sources have limitations in terms of compact size, stability, cofactor dependence, cellular expression efficiency, catalytic efficiency, and substrate orthogonality, which constrain the potential of bioluminescence technology in biomedical research.
Development of de novo designed luciferases with improved properties, including compact size, robust stability, cofactor independence, efficient cellular expression, and higher catalytic efficiency, along with luciferase-fluorescent protein FRET fusions for multi-parametric imaging.
The new luciferases enable reliable, versatile, and multiplexed bioluminescence imaging, facilitating sensitive, real-time imaging in living organisms with enhanced sensitivity and specificity.
Smart Images

Figure US2025046932_26032026_PF_FP_ABST
Abstract
Description
Atty. Docket: UCSC-412WODE NOVO DESIGNED BRIGHT AND MULTI-COLOR LUCIFERASESSTATEMENT OF GOVERNMENT SUPPORT
[0001] This invention was made with government support under EB031913 awarded by the National Institutes of Health. The government has certain rights in the invention.CROSS-REFERENCE
[0002] This application claims benefit of U.S. Provisional Application No. 63 / 697,339 filed September 20, 2024, which application is incorporated herein by reference in its entirety.INCORPORATION BY REFERENCE OF SEQUENCE LISTING XML
[0003] A Sequence Listing is provided herewith as a Sequence Listing XML, “UCSC- 412WO_SEQ_LIST_SEP_2025” created on September 17, 2025, and having a size of 202,046 bytes. The contents of the Sequence Listing XML are incorporated by reference herein in their entirety.INTRODUCTION
[0004] Bioluminescence technology stands as a powerful tool in biomedical research, offering highly sensitive, non-invasive, and real-time imaging in living organisms without the need for external excitation. However, traditional luciferases derived from natural sources have constrained the full potential of luminescence technology.
[0005] Work has been performed to engineer native luciferases to improve their use as molecular probes. However, satisfactory luciferases based on native luciferases have not been generated. A synthetic luciferase, named LuxSit, has been developed by the Baker lab at University of Washington (Nature 614, 774-780 (2023)).
[0006] There is a need for improved luciferases that have a compact size, robust stability, cofactor independence, efficient cellular expression, higher catalytic efficiency, and / or unique substrate orthogonality. There is also a need for luciferases that can be used in multi-parametric imaging.SUMMARY
[0007] The present disclosure provides luciferases that have a compact size, robust stability, cofactor independence, efficient cellular expression, higher catalytic efficiency, and / orAtty. Docket: UCSC-412WO unique substrate orthogonality. The present disclosure provides luciferase-fluorescent protein FRET fusions that can be used in multi-parametric imaging, e.g., in cellulo, in vivo, or both.
[0008] In certain aspects, the present disclosure provides a polypeptide having luciferase activity and comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the amino acid sequences shown in Fig. 14A (SEQ ID NOs:2-3, 1 , and 4-7), Figs. 15A-15C (SEQ ID NOs:151- 154, 1 , and 155-171 ), Figs. 16A-16C (SEQ ID NOs:172, 2, and 173-184), or Figs. 17A-17C (SEQ ID NOs:185-188, 3, and 189-196). In some case, the polypeptide having luciferase activity may comprise an amino acid sequence that has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the amino acid sequences shown in Fig. 14A (SEQ ID NOs:2-3, 1 , and 4-7), Figs. 15A-15C (SEQ ID NOs:151-154, 1 , and 155-171), Figs. 16A-16C (SEQ ID NOs:172, 2, and 173-184), or Figs. 17A-17C (SEQ ID NOs:185- 188, 3, and 189-196), and differs from these amino acid sequences by conservative amino acid substitutions.
[0009] In certain aspects, the present disclosure provides a luciferase-fluorescent protein FRET fusions comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the amino acid sequences of a luciferase-fluorescent protein FRET fusion protein provided herein. In some cases, the luciferase-fluorescent protein FRET fusions comprise an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the amino acid sequences of a luciferase-fluorescent protein FRET fusion protein provided herein and differs from these amino acid sequences by conservative amino acid substitutions.
[0010] In certain aspects, the present disclosure provides a self-complementing multipartite protein having luciferase activity, comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked or are covalently linked via a cleavable linker.
[0011] Additional aspects of the present disclosure include nucleic acids and vectors encoding the disclosed polypeptides and kits for expressing the disclosed polypeptides and kits comprising assay buffer for measuring the luciferase activity of the disclosed polypeptides.Atty. Docket: UCSC-412WOBRIEF DESCRIPTION OF THE DRAWINGS
[0012] FIGS. 1 A-1 B. Computational design and experimental characterization of the second-generation de novo luciferases.
[0013] FIGS. 2A-2B. Characterization of activity and protein folding of newly designed luciferase sequences generated by ProteinMPNN (FIG. 2B, SEQ ID NO:149).
[0014] FIGS. 3A-3C. Evaluation of de novo luciferase sequences designed by RFjoint Inpainting (FIG. 3A, SEQ ID NOs:48, 2, 150, and 4-7, from top to bottom).
[0015] FIGS. 4A-4D. Additional expression, purification, activity, and structural characterization of neoLux series, native, and engineered luciferases.
[0016] FIGS. 5A-5C. The design of neoLux-based FRET pairs for multiplexed imaging.
[0017] FIGS. 6A-6C. Computational modeling and experimental characterization of neoLux-FP FRET pairs.
[0018] FIG. 7. The signal and noise comparison between HeLa cells expressing luxNeon or Antares2 at single-cell resolution microscopic imaging.
[0019] FIGS. 8A-8D. Neoluminescent probes facilitate reliable, versatile, and multiplexed luciferase bioassays.
[0020] FIGS. 9A-9C. Multiplexed neoluminescence imaging of tumor xenografts / n vivo.
[0021] FIG. 10. Spectral unmixing and multiplexed imaging of FRET pairs by conventional filters.
[0022] FIGS. 11 A-11C. Luminescence and fluorescence analysis of B16F10 and HeLa cells expressing various FRET probes after lentiviral transduction.
[0023] FIGS. 12A-12B. Unmixed images of heterogeneous tumors in vivo and endpoint analysis of Tumor 4 by fluorescence-activated cell sorter (FACS).
[0024] FIGS. 13A-13B. Additional xenograft tumor images to showcase the inherent complexity of tumor heterogeneity in vivo.
[0025] FIGS. 14A-14B. FIG. 14A provides an alignment of exemplary polypeptides of the present disclosure (SEQ ID NOs:2-3,1 , and 4-7, from top to bottom). FIG.14B provides an identity matrix for these polypeptides.
[0026] FIGS. 15A-15D. FIGS. 15A-15C provide an alignment of variants of 1c1 (neoLuxI) (SEQ ID NOs:151 -154,1 , and 155-171 , from top to bottom). FIGS.15D-15E provides an identity matrix for these variants.Atty. Docket: UCSC-412WO
[0027] FIGS. 16A-16D. FIGS. 16A-16C provide an alignment of variants of 2b9 (SEQ ID NOs:172,2, and 173-184, from top to bottom). FIG.16D provides an identity matrix for these variants.
[0028] FIGS. 17A-17D. FIGS. 17A-17C provide an alignment of variants of 2c7 (SEQ ID NOs:185-188,3, and 189-196, from top to bottom). FIG.17D provides an identity matrix for these variants.DETAILED DESCRIPTION
[0029] The present disclosure provides luciferases that have a compact size, robust stability, cofactor independence, efficient cellular expression, higher catalytic efficiency, and / or unique substrate orthogonality. The present disclosure provides luciferase-fluorescent protein FRET fusions that can be used in multi-parametric imaging, e.g., in cellulo, in vivo, or both.
[0030] In certain aspects, the present disclosure provides a polypeptide having luciferase activity and comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the amino acid sequences shown in Fig. 14A, Figs. 15A-15C, Figs. 16A-16C, or Figs. 17A-17C.
[0031] In certain aspects, the present disclosure provides luciferase-fluorescent protein FRET fusions comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the amino acid sequences of luciferase-fluorescent protein FRET fusion proteins provided herein.
[0032] In certain aspects, the present disclosure provides a self-complementing multipartite protein having luciferase activity, comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked or are covalently linked via a cleavable linker.
[0033] Additional aspects of the present disclosure include nucleic acids and vectors encoding the disclosed polypeptides and kits for expressing the disclosed polypeptides and kits comprising assay buffer for measuring the luciferase activity of the disclosed polypeptides.
[0034] Before the present invention is described in greater detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particularAtty. Docket: UCSC-412WO embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0035] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0036] Certain ranges are presented herein with numerical values being preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.
[0037] Unless defined otherwise, all technical and scientific terms used herein have the same meanin as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.
[0038] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission thatthe present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.
[0039] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement isAtty. Docket: UCSC-412WO intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0040] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.
[0041] While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in anyway by the construction of "means" or "steps" limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112.DEFINITIONS
[0042] “Derived from” in the context of an amino acid sequence or polynucleotide sequence is meant to indicate that the polypeptide or nucleic acid has a sequence that is based on that of a reference polypeptide or nucleic acid, and is not meant to be limiting as to the source or method in which the protein or nucleic acid is made.
[0043] The terms "polypeptide", and "protein" are used interchangeably herein to designate a linear series of amino acid residues connected one to the other by peptide bonds between the alpha-amino and carboxy groups of adjacent residues. The amino acid residues are usually in the natural "L" isomeric form. However, residues in the "D" isomeric form can be substituted for any L-amino acid residue, as long as the desired functional property is retained by the polypeptide. In addition, the amino acids, in addition to the 20 "standard" amino acids, include modified and unusual amino acids, which include, but are not limited to those listed in 37 CFR (§1 .822(b)(4)). Furthermore, it should be noted that a dash at the beginning or end of an amino acid residue sequence indicates either a peptide bond to a further sequence of one or more amino acid residues ora covalent bond to a carboxyl or hydroxyl end group. However, the absence of a dash should not be taken to mean that such peptide bonds or covalent bond to a carboxyl or hydroxylAtty. Docket: UCSC-412WO end group is not present, as it is conventional in representation of amino acid sequences to omit such. The term "peptide” also refers to a linear series of amino acid residues connected one to the other by peptide bonds between the alpha-amino and carboxy groups of adjacent residues but is generally shorterthan a protein or a polypeptide, e.g., less than 50 amino acids long, e.g., 2-50 amino acids in length. The terms protein, polypeptide, and peptide may be used interchangeably.
[0044] As used herein, the term "binding" refers to the non-covalent interactions of the type which occur between two molecules. The strength or affinity of binding interactions can be expressed in terms of the dissociation constant (KD) of the interaction, wherein a smaller KDrepresents a greater affinity. Binding properties of selected polypeptides can be quantified using methods well known in the art.
[0045] “Isolated” refers to an entity of interest that is in an environment different from that in which the entity may naturally occur or is initially produced in. An “isolated” compound (e.g., an “isolated” polypeptide) is separated from all or some of the components that accompany it and may be substantially enriched, e.g., may be purified so that the compound is at least about 70% pure, at least about 80% pure, at least about 90% pure, at least about 95% pure, at least about 98% pure, at least about 99%, or greater than 99% pure, or free of impurities, contaminants, and / or components other than the compound. “Isolated” also refers to the state of a compound separated from all or some of the components that accompany it during manufacture (e.g., chemical synthesis, recombinant expression, culture medium, and the like).
[0046] As used herein, the amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gin; Q), glycine (Gly; G), histidine (His; H), isoleucine (lie; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Vai; V).
[0047] In all embodiments of polypeptides disclosed herein, any N-terminal methionine residues are optional (i.e., the N-terminal methionine residue may be present or absent). In all embodiments of polypeptides disclosed herein, any C-terminal glycine residues are optional (i.e., the C-terminal glycine residue may be present or absent).
[0048] The term "conservative substitution" is used in reference to proteins to reflect amino acid substitutions that do not substantially alter the activity (specificity or binding affinity) of the molecule. Typically, conservative amino acid substitutions involve substituting one amino acid for another amino acid with similar chemical properties (e.g., charge or hydrophobicity). TheAtty. Docket: UCSC-412WO following six groups each contain amino acids that are typical conservative substitutions for one another: 1 ) Alanine (A), Serine (S), Threonine (T); 2) Aspartic acid (D), Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); and 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W). The polypeptides encompassed by the present disclosure include those that have one or more conservative substitutions relative to the amino acid sequences provided here.
[0049] Percent identity between a pair of sequences may be calculated by multiplying the number of matches in the pair by 100 and dividing by the length of the aligned region, including gaps. Identity scoring only counts perfect matches and does not consider the degree of similarity of amino acids to one another. Only internal gaps are included in the length, not gaps at the sequence ends. Percent Identity = (Matches x 100) / Length of aligned region (with gaps).
[0050] Numeric ranges are inclusive of the numbers definingthe range.POLYPEPTIDES
[0051] The present disclosure provides polypeptides having luciferase activity and comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the following amino acid sequences and / or having one or more conservative substitutions relative to:
[0052] 1 c1 (neoLuxI ):MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIES VEIKDGEAWKVTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:1);
[0053] 2b9:MSAAEVRDFVDRFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARR EITAFEVSDGVAHVDWLRAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL (SEQ ID NO:2);
[0054] 2c7:MSPEEKRVFVERFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARRE ITAFEVSDGVAHVDWLIAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL (SEQ ID NO:3);
[0055] 1 c3:MSPEEIRDFVKRFYEALDAGDAETAAQLLWDAGCRRIELWDGTVFEGPDVRDQFVAWFRALQASVTGAKR EILKVEVKDGTVAWEVRLTATYKATGKTFWRLTHVFTFDPETGELVEVKVTLTPL (SEQ ID NO:4);Atty. Docket: UCSC-412WO
[0056] 1 b12:MSEEIREFVDRFYAALDAGDADTAADLLFSSGCKKIHLWDGTVFDGDKEAFKAWFEDLFAKSEGATRRVTSF AVDLDGLPRADVEVELTTTIDGKEVRVRLRHTFYFDAEGRLVEVWERLPL (SEQ ID NO:5);
[0057] 1f3:MSEEEKREFVERFYAALDKGGEEGAEEAADLLFSSGCKEIHLWDGRVFTSKEEFKAWFVELWASLGEKGAR REVTAFEVNEDGTAWDWLTAEWKDGTVRWRLRHVFHFEDGKLVRVEVERLPL (SEQ ID NO:6); and
[0058] 2a 1 :MSEEEMREFVERFYAALDAGDAETASSLLFDSGCKKIHLWDGRVFTSKEEFKDWFRHLHEDVLEGAVRKVT SFEVDPEKGVAWDWLTARVKATGEEVQVRLRHTFYFEEGKLVEVWERLPL (SEQ ID NO:7).
[0059] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and / or having one or more conservative amino acid substitutions relative to SEQ ID NO:1 .
[0060] The specification provides amino acid sequences that are different from the amino acid sequence of SEQ ID NO:1 at the listed position(s). The numbering of these positions is based on the numbering of amino acid positions in SEQ ID NO:1 , where the M at position 1 in SEQ ID NO:1 is position 1 and T at position 4 is counted as position 2, i.e., the SG residue between M and T are not counted. Thus, in the numbering used herein M in SEQ ID NO:1 is position 1 and residues startingfrom the 4thresidue, T, are counted from position 2. Thus, for example, D at position 18, V at position 83, and L at position 100, numbered based on the numbering of amino acid positions in SEQ ID NO:1 counted as set forth above can also be referred to as to D at position 20, V at position 85, and L at position 102.
[0061] Table 1 summarizes the alternatively counted positions:Atty. Docket: UCSC-412WO
[0062] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and comprises one or more of the following amino acids: D at position 18, V at position 83, and L at position 100, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
[0063] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and comprises one or more of the following amino acids: E at position 18, L at position 83, and I at position 100, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
[0064] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and comprises one or more of the following amino acids: L at position 17, D at position 18, and V at position 83, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
[0065] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and comprises one or more of the following amino acids: I at position 17, E at position 18, and L at position 83, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1 .
[0066] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises E at position 18, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
[0067] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises L at position 17 and V at position 83, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
[0068] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises I at position 17 and L at position 83, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .Atty. Docket: UCSC-412WO
[0069] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and comprises one or more of the following amino acids: D at position 18, V at position 83, and V at position 98, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
[0070] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence having at least at least 85% identity to SEQ ID NO:1 and comprises one or more of the following amino acids: E at position 18, L at position 83, and L at position 98, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
[0071] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises E at position 18, V at position 37, and I at position 85, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
[0072] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises E at position 18 and V at position 37, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
[0073] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises E at position 18 and L at position 98, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
[0074] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises E at position 18 and L at position 98, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
[0075] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence havingat least at least 85% identity to SEQ ID NO:1 and comprises L at position 83.
[0076]
[0077] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:1 or wherein theAtty. Docket: UCSC-412WO amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequences set forth in Figs. 15A-15C.
[0078] The present disclosure provides a luciferase, referred to as neoLuxI or 1 c1 , comprising the amino acid sequence set forth in SEQ ID NO:1 . The present disclosure also provides variants of neoLuxI , where the variant comprises the following substitutions relative to neoLuxI :
[0079] >neoLux1 -D18E / V83L / L100l:MSGTDEEIAEFVKAFYEALEAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWEIEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:8)
[0080] >neoLux1 -L17l / D18E / V83L:MSGTDEEIAEFVKAFYEAIEAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:9)
[0081] >neoLux1 -D18E:MSGTDEEIAEFVKAFYEALEAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKVTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NQ:10)
[0082] >neoLux1 -L17l / V83L:MSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:197)
[0083] >neoLux1 -D18E / V83L / V98L:MSGTDEEIAEFVKAFYEALEAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFVLELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:11 )
[0084] >neoLux1-D18E / l37V / V85l:MSGTDEEIAEFVKAFYEALEAGDAETAADLLFGAGCKEVHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIE SVEIKDGEAWK1TLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:12)
[0085] >neoLux1 -D18E / V85l / V98L:MSGTDEEIAEFVKAFYEALEAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKITLTATYKATGKKFVLELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:13)
[0086] >neoLux1 -V83L (neoLuxI .2):MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFVVELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:14)Atty. Docket: UCSC-412WO
[0087] The substituting amino acids are indicated in bold and are underlined in the sequences listed above.
[0088] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:2 or wherein the amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequences set forth in Figs. 16A-16C.
[0089] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:3 or wherein the amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequences set forth in Figs. 17A-17C.
[0090] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:4.
[0091] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:5.
[0092] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:6.
[0093] In certain cases, the polypeptide having luciferase activity comprises an amino acid sequence that is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to:Atty. Docket: UCSC-412WO
[0094] >1 a10:MSGTPEEKREFVERFYAALDKGAEGAEEAAELLERDGAEIYLWDGTVFKPETIKEDFIAWFKALNSTLSYAKREIVGFEVDEEGTAHVEWLTAQFKGQDVTVRLRHRFYFEDGRLVRVEVEAEPL (SEQ ID NO:15)
[0095] >1 c6:MSGSPEEKKTFVDRFYAALDAGDAKTAADLLFGDDGKCRIRLWDGREFVDDKEAFERWFEGLLSLTEPGTAKREWAFEVDENGRAHVDWLTARVKGSADEFFVRLHHTFYFEDGKLVEVDVEAEPL (SEQ ID NO:16)
[0096] >1 h10:MSGTPAQMRAFVEEFYAALDRGDADTASDLLFSSGCRTIHLWDGRVFTSREGFKAWFRKLYAETVGARREVTAFAVDLEGPQRADVEVTLTARVKSLDYREVTVRLRHTFYFDAAGRLVEVWEVLPP (SEQ ID NO:17)
[0097] >2b12:
[0098] MSGSKEEMKAFVKEFYEALDAEGEESAKKAAGLLFKDGAKIKLWDGTEFTSKEDFEKWFLKLRAETIGGARREILKLEVDEERQEAIVEVLLTARFKSLGGEEVKVRLVHRFKFKDGKLVEVEVRAEPV (SEQ IDNO:18)
[0099] >2c5:MSGTPEEMREFVEKFYAALDAGDADTATALLFEHGAEIYLWNGKVFRPDEAEAFRAWFEALLAAVEGGAKRRVTSFEVNEDRTAIVEVELTAHWKGDPRERWRLRHVFYFASPEDKRLVSVWEVLPN (SEQ ID NO:19)
[0100] >2c7:
[0101] MSGSPEEKRVFVERFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARREITAFEVSDGVAHVDWLIAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL (SEQ ID NO:20)
[0102] >mplux_1 A8:
[0103] MSGSAEAHRRFVDRFYAALDAGDADTASALFPDGTEIHLWDGRTFRTRAEFRTWFRELRARSDNARREWAFEVDGDTAHVEWLRASIDGEERWRLRHTFYFEGDRLVRVEVEIEP (SEQ ID NO:21 )
[0104] >mplux_1 A7:MSGSEEEQREFVDRFYAALDAGDAETASALFPDGTKIYLWDGKVFTTREEFRAWFEKLYSTSENAKRHWSFKVDGNKADVEWLHANINGEKKTVRLRHVFYFEGDKLVEVKVEIKPL (SEQ ID NO:22)
[0105] >mplux_1 A12:MSGSEEEIREFVRRFYEALDAGDAATASALFPDGTEIHLWDGTTFRTRAQFRAWFERLRAQSANARREIVDL KVEGDRAKVEVILRASFDGEEKWNLTHEFLFEGDRLVRVSVTIPRWVPNSVAVARTCTSKDP (SEQ ID NO:23)
[0106] >mplux_1 B6:Atty. Docket: UCSC-412WO
[0107] MSGSEEEQREFVDRFYAALDAGDAETASALFPDGTKIYLWDGKVFTTREEFRAWFEKLYSTSENAKRHWSFKVDGNKADVEWLHANINGEKKAVRLRHVFYFEGDKLVEVKVEIRPL (SEQ ID NO:24)
[0108] >mplux_1 B11 :
[0109] MSGSSDAQRAFVDRLYRALDAGDAETASALFPDGTRIHLWDGTTFTTREEFRAWFVDLRSRSENAAREWSFDVDGDVAHVEWLKAVIEGEEVWRLRHVFEWEGDRLVEVYVEIDPL (SEQ ID NO:25)
[0110] >mplux_1 D4:
[0111] MSGSAEQQREFVKHFYEALDAGDADTASALFPDGTEIHLWDGTTFRTRAEFRAWFEEQYSTSENASREVTSFSVDGDVADVEWLRANLGGEDRTVSLRHVFHFAGDRMMRVEVSIRPL (SEQ ID NO:26)
[0112] >mplux_1 D7:
[0113] MSGSEEEIREFVRRFYEALDAGDAVTASALFPDGTEIHLWDGTTFRTQAQFRAWFERLRAQSANARREIVDLKVEGDRAIVEWLRASFDGEEKWNLTHEFLFEGDRLVRVSVTITPL (SEQ ID NO:27)
[0114] >mplux_2C2:
[0115] MSGSEEEQREFVDRFYAALDAGDAETASALFPDGTKIYLWDGKVFTTREEIRAWFEKLYSTSENAKRHWSFKVDGNKADVEWLHANINGEKKTVRLRHVFYFEGDKLVEVKVEIKPL (SEQ ID NO:28).
[0116] The polypeptides of the present disclosure in addition to comprising an amino acid sequence as disclosed herein may include catalytic dyads of (i) D residue at position 20 and R residue at position 69; and (ii) Y residue at position 16 and H residue at position 104, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1.
[0117] The polypeptides of the present disclosure in addition to comprising an amino acid sequence and catalytic dyads as disclosed herein may have the secondary structure arrangement H1 -L1 -H2-L2-B1 -L3-B2-L4-H3-L5-B3-L6-B4-L7-B5-L8-B6-L9, wherein “H” is a helical domain, “L” is a loop domain, and “B” is a beta strand domain.
[0118] In certain cases, the H1 domain is at least or up to 20 amino acids in length; the L1 domain is at least or up to 4-5 amino acids in length; the H2 domain is at least or up to 9 amino acids in length; the L2 domain is at least or up to 5-6 amino acids in length; the B1 domain is at least or up to 2 amino acids in length; the L3 domain is at least or up to 3 amino acids in length; the B2 domain is at least or up to 4 amino acids in length; the L4 domain is at least or up to 4-5 amino acids in length; the H3 domain is at least or up to 12-16 amino acids in length; the L5 domain is at least or up to 5-8 amino acids in length; the B3 domain is at least or up to 8-10 amino acids in length; the L6 domain is at least or up to 3-7 amino acids in length; the B4 domain is at least or upAtty. Docket: UCSC-412WO to 10-12 amino acids in length; the L7 domain is at least or up to 4-6 amino acids in length; the B5 domain is at least or up to 12-14 amino acids in length; the L8 domain is at least or up to 4-7 amino acids in length; the B6 domain is at least or up to 6-9 amino acids in length; and the L9 domain is at least or up to 3 amino acids in length.
[0119] The secondary structure arrangement H1-L1 -H2-L2-B1 -L3-B2-L4-H3-L5-B3-L6-B4-L7-B5-L8-B6-L9 for exemplary luciferases is illustrated below:
[0120] 1 c1 (neoLuxI ):MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVIGAKRKIESVEJKDGEAWKVTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:1 )
[0121] 2b9:MSAAEVRDFVDRFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARR EITAFEVSDGVAHVDWLRAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL (SEQ ID NO:2)
[0122] 1c3:MS PEE I RD FVKRFYEALDAGDAETAAQ LLWDAGCRRI ELWDGTVFEGPDVRDQ FVAWFRALQASVTGAKR EILKVEVKDGTVAWEVRLTATYKATGKTFWRLTHVFTFDPETGELVEVKVTLTPL (SEQ ID NO:4)
[0123] 1 b12:MSEEIREFVDRFYAALDAGDADTAADLLFSSGCKKIHLWDGTVFDGDKEAFKAWFEDLFAKSEGATRRVTSFAVDLDGLPRADVEVELTTTIDGKEVRVRLRHTFYFDAEGRLVEVWERLPL (SEQ ID NO:5)
[0124] 1f3:MSEEEKREFVERFYAALDKGGEEGAEEAADLLFSSGCKEIHLWDGRVFTSKEEFKAWFVELWASLGEKGARREVTAFEVNEDGTAWDWLTAEWKDGTVRWRLRHVFHFEDGKLVRVEVERLPL (SEQ ID NO:6)
[0125] 2a1 :MSEEEMREFVERFYAALDAGDAETASSLLFDSGCKKIHLWDGRVFTSKEEFKDWFRHLHEDVLEGAVRKVTSFEVDPEKGVAWDWLTARVKATGEEVQVRLRHTFYFEEGKLVEVWERLPL (SEQ ID NO:7)
[0126] 2c7:MSPEEKRVFVERFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARRE ITAFEVSDGVAHVDWLIAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL(SEQ ID NO:3)
[0127] The loop regions are underlined.
[0128] In certain cases, the H1 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: MSGTDEEIAEFVKAFYEAID (SEQ ID NO:29);MSAAEVRDFVDRFYAALD (SEQ ID NO:30); or MSPEEKRVFVERFYAALD (SEQ ID NO:31 ).Atty. Docket: UCSC-412WO
[0129] In certain cases, the H2 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: ETAADLLFG (SEQ ID NO:32), AAASGLL (SEQ ID NO:33), or AAASGL (SEQ ID NO:34).
[0130] In certain cases, the B1 domain comprises an amino acid sequence having at least 50% or 100% identity to the amino acid sequence: IH or IE.
[0131] In certain cases, the B2 domain comprises an amino acid sequence having at least 50% or 100% identity to the amino acid sequence: GTV or GT.
[0132] In certain cases, the H3 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: EGFETWFKKLQS (SEQ ID NO:35), VLEQFETWFNGLLST (SEQ ID NO:36), or DDVLEQFETWFNGLLS (SEQ ID NO:37).
[0133] In certain cases, the B3 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 100% identity to the amino acid sequence: RKIESVEI (SEQ ID NO:38), RREITAFEV (SEQ ID NO:39), or RREITAFEV (SEQ ID NO:39).
[0134] In certain cases, the B4 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: AWKLTLTAT (SEQ ID NO:40), VAHVDWLRAK (SEQ ID NO:41 ), or AHVDWLIAK (SEQ ID NO:42).
[0135] In certain cases, the B5 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: KFWELEHTFTF (SEQ ID NO:43), AEVTVRLHHTFFFR (SEQ ID NO:44), or VTVRLHHTFFF (SEQ ID NO:45).
[0136] In certain cases, the B6 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: ELVEVKVTA (SEQ ID NO:46), RLVRVEVEV (SEQ ID NO:47), or RLVRVEVEV (SEQ ID NO:47).
[0137] In certain cases, the L1 , L2, L3, L4, L5, L6, L7, L8, and L9 domains are at least 1 , 2, 3, 4, 5, 6, 7, or 8 amino acids in length and comprise any amino acid and optionally are up to 2, 3, 4, 5, 6, 7, or 8 amino acids in length.Atty. Docket: UCSC-412WO
[0138] The polypeptides of the present disclosure possess one or more improved properties as compared to LuxSit-i. LuxSit-i is described in W02024097640 and has the following amino acid sequence:
[0139] MSEEQIRQFLRRFYEALDSGDADTAASLFHPGVTIHLWDGVTFTSREEFREWFERLFSTSK DAQREIKSLEVRGDTVEVHVQLHATHNGQKHTVDLTHHWHFRGNRVTEVRVHINPT (SEQ ID NO:48).
[0140] In certain aspects, these polypeptides have improved activity compared to LuxSit-i. For example, these polypeptides have a luciferase activity that is at least 10% higher than LuxSit-i luciferase activity, e.g., at least 20% higher, at least 30% higher, at least 40% higher, at least 50% higher, at least 60% higher, at least 70% higher, at least 80% higher, at least 90% higher, at least 100% higher, at least 150% higher, or up to 150% higher, or up to 180% higher, or up to 200% higher than LuxSit-i luciferase activity. The luciferase activity may be measured using any suitable assay, including assays provided herein. The luciferase activity may be measured using a luciferin substrate, e.g., DTZ, coelenterazine, furimazine, coelenterazine-n, coelenterazine-f, coelenterazine-h, coelenterazine-hcp, coelenterazine-cp, coelenterazine-c, coelenterazine-e, coelenterazine-fcp, bis-deoxycoelenterazine, coelenterazine-i, coelenterazine-icp, coelenterazine- v, and 2-methyl coelenterazine, or another luciferin substrate, or an analog thereof.
[0141] In certain aspects, these polypeptides have improved stability at high temperatures as compared to LuxSit-i. For example, these polypeptides are stable at higher temperatures as compared to LuxSit-i. Stability may be measured by enzymatic activity and / or protein misfolding measured over a period of time. In certain embodiments, stability may be measured using static light scattering (SLS). In certain embodiments, the polypeptides disclosed herein are more stable than LuxSit-i at a temperature higher than 37. C. In certain embodiments, the polypeptides disclosed herein are more stable than LuxSit-i at a temperature higher than 37. C, as measured by SLS.
[0142] In certain aspects, these polypeptides have improved specificity as compared to LuxSit-i. For example, it may have 2X, 3, X, 5X, 10X higher specificity for a luciferin substrate, e.g., DTZ, as compared to LuxSit-i.
[0143] In certain aspects, these polypeptides have improved yield compared to LuxSit-i. For example, these polypeptides are expressed at higher levels and / or with lower levels of aggregated or misfolded proteins as compared to LuxSit-i when expressed in standard expression systems such as E. Coli, yeast, mammalian cell lines, and the like.Atty. Docket: UCSC-412WOConjugated Proteins
[0144] The luciferases disclosed herein may be conjugated to another moiety. The moiety may be a small molecule, peptide, polypeptide, nucleic acid, or lipid. The moiety may be a peptide or polypeptide for localization of the luciferase to a cellular compartment, cell membrane, orfor secretion. The luciferases disclosed herein can be used as biosensors by conjugating a moiety to the N-terminus, the C-terminus, or in between the N- and the C-terminus. The moiety may be conjugated directly to the luciferase, e.g., via a peptide bond to the N-terminus and / or the C- terminus and / or to an amino acid side chain or may be conjugated to the luciferases via a linker. The linker may be a polymer, e.g., an amino acid linker or a sugar linker.
[0145] A variety of linkers may be used and may include alkyl groups, methylene carbon chains, ether, polyether, alkyl amide linker, a peptide linker, a modified peptide linker, a Polyethylene glycol) (PEG) linker, a streptavidin-biotin or avidin-biotin linker, polyaminoacids (e.g., polylysine), functionalized PEG, polysaccharides, glycosaminoglycans, oligonucleotide linker, phospholipid derivatives, alkenyl chains, alkynyl chains, disulfide, or a combination thereof. In some embodiments, the linker is cleavable (e.g., enzymatically (e.g., TEV protease site), chemically, photoinduced cleavage, etc.).
[0146] In certain aspects, the moiety may be a heterologous amino acid sequence. In certain aspects, the moiety is conjugated to the luciferase post-translationally. In certain aspects, the moiety is conjugated to the luciferase during translation, e. g. , a nucleic acid may encode a fusion protein comprising the luciferase and the moiety.
[0147] In certain aspects, the heterologous amino acid sequence includes a protein bindingdomain, such as one that binds IL-17RA, e.g., IL-17A, orthe IL-17A binding domain of IL- 17RA, Jun binding domain of Erg, or the EG binding domain of Jun; a potassium channel voltage sensingdomain, e.g., one useful to detect protein conformational changes, the GTPase binding domain of a Cdc42 or rac target, or other GTPase binding domains, domains associated with kinase or phosphotase activity, e.g., regulatory myosin light chain, PKC5, pleckstrin containing PH and DEP domains, other phosphorylation recognition domains and substrates; glucose binding protein domains, glutamate / aspartate binding protein domains, PKA or a cAMP-dependent binding substrate, lnsP3 receptors, GKI, PDE, estrogen receptor ligand binding domains, apoK1 -er, or calmodulin binding domains.
[0148] In certain aspects, a fusion protein comprising a luciferase fused to a heterologous amino acid sequence may be a biosensor. The biosensor is useful to detect a GTPase, e.g., bindingAtty. Docket: UCSC-412WO of Cdc42 or Rac to a EBFP, EGFP PAK fragment, Raichu-Rac, Raichu-Cdc42, integrin alphavbeta3, IBB of importin-a, DMCA or NBD-Ras of CRafl (for Ras activation), binding domain of Ras / Rap Rai RBD with Ras prenylation sequence. In one embodiment, the biosensor detects PI(4,5)P2 (e.g., using PH-PCLdelta1 , PH-GRP1 ), PI(4,5)P2 or PI (4) P (e.g., PH-OSBP), PI(3,4,5)P3 (e.g., using PH- ARNO, or PH-BTK, or PH-Cytohesin1 ), PI(3,4,5)P3 or PI(3,4)P2 (e.g., using PH Akt), PI(3)P (e.g., using FYVE-EEA1 ), or Ca2+ (cytosolic) (e.g., using calmodulin, or C2 domain of PKC.
[0149] In one aspect, a fusion protein comprising a luciferase is fused to a protein domain. In one embodiment, the domain is one with a phosphorylated tyrosine (e.g., in Src, Ab1 and EGFR), that detects phosphorylation of ErbB2, phosphorylation of tyrosine in Src, Ab1 and EGFR, activation of MKA2 (e.g., using MK2), cAMP induced phosphorylation, activation of PKA, e.g., using KID of CREG, phosphorylation of Crkll, e.g., using SH2 domain pTyr peptide, binding of bZIP transcription factors and REL proteins, e.g., bFos and bJun ATF2 and Jun, or p65 NFkappaB, or microtubule binding, e.g., using kinesin.
[0150] In some cases, a luciferase of the present disclosure is fused to a fluorescent polypeptide. The fluorescent polypeptide may produce an optical signal that is different from the optical signal produced by the luciferase. In certain cases, when both polypeptides simultaneously produce optical signals, the combination of the optical signals may be different from the individual optical signals.
[0151] In some cases, the fusion protein includes a luciferase and a fluorescent polypeptide that form a FRET pair which produces light that is distinct from and brighter than the fluorescent signal of the fluorescent polypeptide and the luminescent signal of the luciferase. In some cases, the luciferase is the FRET donor and the fluorescent polypeptide is the FRET acceptor.
[0152] In some cases, a fusion protein is provided which comprises a luciferase of the present disclosure and a fluorescent polypeptide fused via a linker. In some cases, a part of the N- terminal region of the fluorescent polypeptide may be deleted to increase the FRET. In some case, a part of the N-terminal region of the fluorescent polypeptide may be deleted and a linker may be added between the C-terminus of the luciferase and the N-terminus of the fluorescent polypeptide to increase the FRET. The linker length may be determined empirically. The linker length may be 1- 20 amino acids, 1 -18 amino acids, 2-18 amino acids, 2 -16 amino acids, 4 -14 amino acids, 5-12 amino acids, e.g., 1 , 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. The part of the N-terminal region of the fluorescent polypeptide that is deleted may be the first amino acid, the first 2 amino acids, the first 3 amino acids, the first 4 amino acids, the first 5 amino acids, the first 6 amino acids, the first 7Atty. Docket: UCSC-412WO amino acids, the first 8 amino acids, the first 9 amino acids, or the first 10 amino acids of the N- terminal region of the corresponding full-length fluorescent protein.
[0153] In some cases, a fusion protein is provided which comprises a luciferase of the present disclosure and a fluorescent polypeptide, where the polypeptide is mNeonGreen, mGold, mKok, mKate, CyOFP, or Neon-PEST.
[0154] In some cases, a mNeonGreen fluorescent polypeptide may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to:
[0155] MVSKGEEDNMASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFS PWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKG TGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPM YVFRKTELKHSKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:49).
[0156] In some cases, a mGold fluorescent polypeptide may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to:
[0157] MVSKGEELFTGWPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKE DGNILGHKLEYNYNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSY QSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NG:50).
[0158] In some cases, a mKok fluorescent polypeptide may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to:
[0159] MVSVIKPEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPAD GPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRK TEGNITEQVEDAVAHS (SEQ ID NO:51).
[0160] In some cases, a CyOFP fluorescent polypeptide may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to:
[0161] MVSKGEELIKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPAAtty. Docket: UCSC-412WONGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLE RIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:52).
[0162] In some cases, a mKate fluorescent polypeptide may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to:
[0163] MVSELIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGP VMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERI KEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:53).
[0164] In some cases, a Neon-PEST fluorescent polypeptide may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to:
[0165] MVSKGEEDNMASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKG TGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPM YVFRKTELKHSKTELNFKEWQKAFTDVMGMDELYKNSHGFPPEVEEQAAGTLPMSCAQESGMDRHPAAC ASARINV (SEQ ID NO:54).
[0166] In some case, the fluorescent polypeptide is mNeonGreen and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0167] MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLNSLPATHELHIF GSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDG SGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMGMDE LYK (SEQ ID NO:55);
[0168] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLFDNMASLPAT HELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQA AMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPAtty. Docket: UCSC-412WONDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDV MGMDELYK (SEQ ID NO:56);
[0169] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPRDDNMASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDV MGMDELYK (SEQ ID NO:57);
[0170] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPTGMASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMG MDELYK (SEQ ID NO:58);
[0171] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAVVKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPISSLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMGMD ELYK (SEQ ID NO:59);
[0172] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLNSLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMGMD ELYK (SEQ ID NQ:60); or
[0173] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPPYSLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMGMD ELYK (SEQ ID NO:61),Atty. Docket: UCSC-412WO
[0174] In some case, the fusion protein comprises a fusion of a luciferase of the present disclosure and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:SLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMS PFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRS KKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKA FTDVMGMDELYK (SEQ ID NO:62), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -5, 1 -8, or 1 -11 amino acids at the N-terminus of the mNeonGreen fluorescent protein and / or (ii) the C-terminus of any one of the luciferase of the present disclosure is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, 1-6, 1 -8, 1 -12, 1 -14, 1-16, or 1 -18 amino acids.
[0175] In some case, the fluorescent polypeptide is mGold and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0176] MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYFTGWPILVELD GDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSAMP EGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKNGIK ANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGM DELYK (SEQ ID NO:63);
[0177] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATGSELFTGWPILV ELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSA MPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKN GIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITL GMDELYK (SEQ ID NO:64);
[0178] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAVVKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATHGELFTGWPIL VELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSAtty. Docket: UCSC-412WOAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKN GIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITL GMDELYK (SEQ ID NO:65);
[0179] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPWLLFTGWPIL VELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKS AMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKN GIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:66);
[0180] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPQLLFTGWPILV ELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSA MPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKN GIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:67);
[0181] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPWDTGWPILVE LDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSA MPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKN GIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:68); or
[0182] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYFTGWPILVEL DGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMKQHDFFKSAM PEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYITADKQKNGI KANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:69),
[0183] In some cases, the fusion protein comprises a fusion of any one of the luciferases of the present disclosure and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to::Atty. Docket: UCSC-412WOFTGWPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMK QHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYI TADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEF VTAAGITLGMDELYK (SEQ ID NG:70), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -5, 1 -8, or 1 -9 amino acids at the N-terminus of the Gold fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N- terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -3, 1 -4, 1-6, 1 -8, 1 -13, or 1 -15 amino acids.
[0184] In some cases, the fluorescent polypeptide is mKok and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0185] MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATIQAEMKMRYYMD GSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFP EGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKG DVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:71 );
[0186] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATITPEMKMRYYMD GSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFP EGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKG DVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:72);
[0187] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATIAKEMKMRYYM DGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAF PEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLK GDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:73);
[0188] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATIQAEMKMRYYM DGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAF PEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKAtty. Docket: UCSC-412WOGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:74);
[0189] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATLTHEMKMRYYM DGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAF PEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLK GDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:75),
[0190] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPNPIKPEMKMRY YMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFK QAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDG VLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:76),
[0191] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAVVKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPSPIKPEMKMRY YMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFK QAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:77),
[0192] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPPPKPEMKMRY YMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFK QAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDG VLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:78),
[0193] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPAIKPEMKMRYY MDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQ AFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGV LKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:79),Atty. Docket: UCSC-412WO
[0194] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPILPEMKMRYYM DGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAF PEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLK GDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:80), or
[0195] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIQPEMKMRYY MDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQ AFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGV LKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:81),
[0196] In some cases, the fusion protein comprises a fusion of a luciferase of the present disclosure and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:PEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYP EEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPST EKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:82), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -4, 1 -5, or 1 -6 amino acids at the N-terminus of the mKok fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1-3, 1 -4, 1 -5, 1 -6, 1 -8, or 1 -1 1 amino acids.
[0197] In some examples, the fluorescent polypeptide is cyOFP and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0198] MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYIKENMRSKLYLE GSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGF TWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRAtty. Docket: UCSC-412WOCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:83);
[0199] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATDSIKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:84);
[0200] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYIKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNE1YVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:85);
[0201] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPFVKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:86);
[0202] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPSKKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:87);
[0203] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPILENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:88);Atty. Docket: UCSC-412WO
[0204] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIRENMRSKLYL EGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEG FTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGR CDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDE LYK (SEQ ID NO:89),
[0205] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTETPIKENMRSKLYL EGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEG FTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGR CDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDE LYK (SEQ ID NO:90);
[0206] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPHIKENMRSKLY LEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPE GFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLE GRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:91 ); or
[0207] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATSPIKENMRSKLY LEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPE GFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLE GRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGM DELYK (SEQ ID NO:92);
[0208] In some cases, the fusion protein comprises a fusion of a luciferase of the present disclosure and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:IKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADL PDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTET MYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVAAtty. Docket: UCSC-412WORYSNLGGGMDELYK (SEQ ID NO:93), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -5, 1 -8, 1 -9, 1 -10, or 1 -11 amino acids at the N-terminus of the cyOFP fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1-2, 1 -4, 1 -6, 1 - 8, 1 -13, or 1-15 amino acids.
[0209] In some cases, the fluorescent polypeptide is mKate and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0210] MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIRENMHMKLYM EGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGF TWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRAD MALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:94);
[0211] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPCLIKENMHMK LYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSF PEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLE GRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSK LGHR (SEQ ID NO:95);
[0212] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPPLIKENMHMKL YMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFP EGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEG RADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKL GHR (SEQ ID NO:96);
[0213] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPILENMHMKLY MEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPE GFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRAtty. Docket: UCSC-412WOADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLG HR (SEQ ID NO:97);
[0214] MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIRENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLG HR (SEQ ID NO:98),
[0215] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYPLIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLXDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:99),
[0216] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAVVKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPNPLIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NQ:100),
[0217] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYYIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:101),
[0218] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPMFIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSK LGHR (SEQ ID NO:102),Atty. Docket: UCSC-412WO
[0219] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYIKENMHMKLY MEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPE GFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGR ADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLG HR (SEQ ID NO:103), or
[0220] MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLIKENMHMKLY MEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPE GFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGR ADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLG HR (SEQ ID NO:104);
[0221] In some cases, the fusion protein comprises a fusion of any one of the luciferases of the present disclosure and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:ENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIP DFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYP ADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVAR YCDLPSKLGHR (SEQ ID NO:105), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -4, 1 -5, 1 -6, or 1 -7 amino acids at the N-terminus of the mKate fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N- terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, or 1-6 amino acids.
[0222] In some cases, the fusion protein comprises a fusion of any one of the luciferases of the present disclosure and the fluorescent polypeptide is Neon-PEST and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0223] MTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSA VTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLNSLPATHELHIFGSAtty. Docket: UCSC-412WOINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSG YQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTF KWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMGMDELYK NSHGFPPEVEEQAAGTLPMSCAQESGMDRHPAACASARINV (SEQ ID NO:106) or
[0224] In some cases, the fusion protein comprises a fusion of any one of the luciferases of the present disclosure and the fusion protein comprises a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:
[0225] SLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGF HQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMT NSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSK TELNFKEWQKAFTDVMGMDELYKNSHGFPPEVEEQAAGTLPMSCAQESGMDRHPAACASARINV (SEQ ID NO:107), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -6, 1-8, 1 -10, or 1 -11 amino acids at the N-terminus of the Neon-PEST fluorescent protein and / or (ii) wherein the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, or 1 -6 amino acids.
[0226] In certain cases, the luciferases, the fusion proteins, the first and second polypeptide components of a self-complementing multipartite protein of the present disclosure may be expressed as a fusion with a purification tag, e.g., a His-tag. For example, the His-tag, HHHHHH (SEQ ID NO: 108), may be fused to the N-terminus or the C-terminus of the luciferase. The purification tag may be fused to the luciferase directly or via a linker, e.g., (GS)n, (G)n, or (GSG)nlinker, where n=1 , 2, 3, 4, 5, or 6.SELF-COMPLEMENTING MUTIPARTITE PROTEIN HAVING LUCIFERASE ACTIVITY
[0227] A self-complementing multipartite protein having luciferase activity is provided. In certain aspects, the multipartite protein may have two self-complementing components or three self-complementing components.
[0228] Self-complementing refers to the characteristic of two or more polypeptides being able to form a complex with each other to regain enzymatic activity absent or substantially absent when the two or more polypeptides are not associated. Complementary polypeptides may require assistance to form a stable complex (e.g., from interaction elements), for example, to place theAtty. Docket: UCSC-412WO polypeptides in the proper conformation for complementarity, to co-localize complementary polypeptides, to lower interaction energy for polypeptides, etc.
[0229] Multipartite protein refers to a protein complex in which the polypeptide components of the multipartite protein are in direct and / or indirect contact with one another. In one aspect, direct contact means two or more molecules are close enough so that attractive noncovalent interactions between the molecules, such as Van der Waal forces, hydrogen bonding, ionic and hydrophobic interactions, and the like, influence the interaction of the molecules. An example of direct contact can include a multipartite protein comprising from N-terminus to C- terminus, a first polypeptide component, a linker, and a second polypeptide component, where the first polypeptide component and the second polypeptide component associate and have luciferase activity and upon cleavage of the linker are separated and lack or have substantially reduced cleavage activity. In one aspect, indirect contact means two or more molecules interact when bridging moieties conjugated to the two or more molecules bring the two or more molecules close together in a stable complex so that attractive noncovalent interactions between the molecules, such as Van der Waal forces, hydrogen bonding, ionic and hydrophobic interactions, and the like, influence the interaction of the molecules.Self-complementing multipartite protein havingtwo or more components
[0230] In certain aspects, the self-complementing multipartite protein includes at least a first polypeptide component and a second polypeptide component, where the at least first polypeptide component and the second polypeptide component are not covalently linked or are covalently linked via a linker (e.g., a cleavable linker), where in total the first polypeptide component and the second polypeptide component comprise the secondary structure arrangement H1 -L1 -H2-L2-B1 -L3-B2-L4-H3-L5-B3-L6-B4-L7-B5-L8-B6, where each domain is as described herein and (a) each H and B domain is fully present within one polypeptide component of eitherthe first polypeptide component orthe second polypeptide component, (b) the first polypeptide component and the second polypeptide component do not include all of the H and B domains, (c) the relative order of the H, L, and B domains when present in either of the first polypeptide component or the second polypeptide component is unchanged with reference to the order set forth in the secondary structure arrangement H1-L1-H2-L2-B1-L3-B2-L4-H3-L5-B3-L6-B4- L7-B5-L8-B6, and (d) the first component and the second component when not present in the selfcomplementing multipartite protein do not possess detectable luciferase activity or have luciferase activity lower than the luciferase activity of the self-complementing multipartite protein.Atty. Docket: UCSC-412WO
[0231] In certain aspects, the first polypeptide component and the second polypeptide component comprise the secondary structure arrangement as set forth in Table 1 :Table 1 :Atty. Docket: UCSC-412WO
[0232] In Table 1 , the L domain in parenthesis is (i) present in one but not both of the first and second polypeptide components, (ii) is split between the first and second polypeptide components, or (iii) absent.
[0233] In certain aspects, one or both of the first component and the second component includes an additional domain. The additional domain may be covalently linked to one or both of the first component and the second component. The domain may be a small molecule, a peptide, a polypeptide, nucleic acid, lipid, an aptamer, etc. In some cases, one of the first component and the second component may be fused to a fluorescent protein, such as, the fluorescent proteins provided herein. In some cases, the first component is smaller than the second component and the first component is fused to a fluorescent protein. In some cases, the first component may additionally be fused to another moiety, e.g., conjugated to a nucleic acid and the second component may be fused to another moiety, e.g., conjugated to a nucleic acid.
[0234] In some aspects, the first component and the second component have high affinity for each other and form a high-affinity two-component protein having luciferase activity by direct interaction. In other words, the two components form a stable complex having luciferase activity when present in close vicinity, e.g., in a polypeptide, in a cell, in a cell lysate, in a cell free solution, etc.
[0235] In some aspects, the first component and the second component have low affinity for each other and form a two-component protein having luciferase activity by indirect interaction mediated by a binding pair. In other words, the two components form a stable complex having luciferase activity when each is conjugated to a member of a binding pairand the interaction between the binding pair members allow formation of a two-component protein having luciferase activity.
[0236] Binding pairs can be a ligand and a receptor; an antigen and an antibody; selfcomplementing enzyme fragments, such as, beta-galactosidase; biotin-avidin; two complementary nucleic acids; two polypeptides capable of dimerization (e.g., homodimer, heterodimer, etc.); and the like.
[0237] In some embodiments, the self-complementing multipartite protein comprises from N-terminus to C-terminus: a first polypeptide component, a linker, and a second polypeptide component or a second polypeptide component, a linker, and a first polypeptide component and has luciferase activity. The linker may be cleavable linker. For example, the linker may include a cleavage site for a protease. In the presence of the protease, the linker is cleaved resulting inAtty. Docket: UCSC-412WO separation of the first and second polypeptide components and loss or significant reduction of the luciferase activity as compared to the luciferase activity of the self-complementing multipartite protein. In certain embodiments, the protease may be a neurotoxin and the self-complementing multipartite protein may be used to detect presence of the protease. In certain embodiments, the self-complementing multipartite protein may include spacer regions between the linker and the first and / or the second component. In certain embodiments, the neurotoxin cleavage site comprises a Clostridium botulinum neurotoxin (BoNT) or a Tetanus neurotoxin cleavage site.
[0238] In some embodiments, the self-complementing multipartite protein comprises a first polypeptide and a second polypeptide, where the first polypeptide comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to
[0239] MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKVTLTATYKATGKKFWELEHTFTFDRE (SEQ ID NO:109);
[0240] MSEEEQREFVDRFYAALDAGDAETASALFPDGTKIYLWDGKVFTTREEFRAWFEKLYSTSENAKRHVVSFKVDGNKADVEVVLHANINGEKKTVRLRHVFYFEG (SEQ ID NO:1 10);
[0241] MSAEQQREFVKRFYEALDAGDADTASALFPDGTEIHLWDGTTFRTRAEFRAWFEELYSTSENASREVTSFSVDGDVADVEWLRANLGGEDRTVSLRHVFHFAG (SEQ ID NO:111);
[0242] MSSDAQRAFVDRFYRALDAGDAETASALFPDGTRIHLWDGTTFTTREEFRAWFVDLRSRSENAAREWSFDVDGDVAHVEWLKAVIEGEEVWRLRHVFEWEGD (SEQ ID NO:112);
[0243] MSAEAQRRFVDRFYAALDAGDADTASALFPDGTEIHLWDGRTFRTRAEFRAWFRELRARSDNARREWAFEVDGDTAHVEWLRASIDGEERWRLRHTFYFEG (SEQ ID NO:113);
[0244] MSEEEIREFVRRFYEALDAGDAATASALFPDGTEIHLWDGTTFRTQAQFRAWFERLRAQSANARREIVDLKVEGDRAKVEVILRASFDGEEKWNLTHEFLFEGD (SEQ ID NO:114);
[0245] MSAAEVRDFVDRFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARREITAFEVSDGVAHVDWLRAKLAGGAEVTVRLHHTFFFRPD (SEQ ID NO:1 15);
[0246] MTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEVHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTITATYKATGKKFWEIENTFTFDRE (SEQ ID NO:116);
[0247] MSAEAHRRFVDRFYAALDAGDADTASALFPDGTEIHLWDGRTFRTRAEFRTWFRELRARSDNARREVVAFEVDGDTAHVEVVLRASIDGEERVVRLRHTFYFEG (SEQ ID NO:117);Atty. Docket: UCSC-412WO
[0248] MSEEIREFVDRFYAALDAGDADTAADLLFSSGCKKIHLWDGTVFDGDKEAFKAWFEDLFA KSEGATRRVTSFAVDLDGLPRADVEVELTTTIDGKEVRVRLRHTFYFDAEG (SEQ ID NO:118);
[0249] MSPEEIRDFVKRFYEALDAGDAETAAQLLWDAGCRRIELWDGTVFEGPDVRDQFVAWFRALQASVTGAKREILKVEVKDGTVAWEVRLTATYKATGKTFWRLTHVFTFDPETG (SEQ ID NO:119);
[0250] MSPEEKKTFVDRFYAALDAGDAKTAADLLFGDDGKCRIRLWDGREFVDDKEAFERWFEGLLSLTEPGTAKREWAFEVDENGRAHVDWLTARVKGSADEFFVRLHHTFYFEDG (SEQ ID NO:120);
[0251] MSEEEKREFVERFYAALDKGGEEGAEEAADLLFSSGCKEIHLWDGRVFTSKEEFKAWFVELWASLGEKGARREVTAFEVNEDGTAWDWLTAEWKDGTVRWRLRHVFHFEDG (SEQ ID NO:121);
[0252] MSEEEMREFVERFYAALDAGDAETASSLLFDSGCKKIHLWDGRVFTSKEEFKDWFRHLHEDVLEGAVRKVTSFEVDPEKGVAWDWLTARVKATGEEVQVRLRHTFYFEEG (SEQ ID NO:122);
[0253] MSEEEQREFVARFYAALDAGDAETASALFPDGTEIHLWDGKTFTTRAEFRAWFEKLHSLSDNASRHVTSFKVDGNVAEVEWLHADFKGKKLTVKLRHRYQFEG (SEQ ID NO:123);
[0254] MSKEEQEEFVKQFYEALDAGDAETASALFPDGTVIHLWDGKTFHTQAEFRAWFEELKSTSENAKREVTKFEVDGDVADVEWLKANINGEEKWNLKHKFKFEG (SEQ ID NO:124);
[0255] MSEEEIKEFVKRFYEALDAGDAETASALFPDGTRIYLWDGRVFRTRAEFRAWFVELHSTSEDAKREVIELKVEGNVAKVKWLHANINGEKKTVLLEHYFEFEG (SEQ ID NO:125);
[0256] MSEEEIREFVRRFYEALDAGDAETASALFKDGTKIYLWDGTVFETREEFRAWFVELYSKSENARRRWSFKVDGNVAEVEWLHASFQGEDKWRLKHRFKFEG (SEQ ID NO:126);
[0257] MSEESQREFWKRFYAALDAGDAETASALFPDGTEIHLWDGTVFRTRAEFRAWFVDLHSKSDNASREITSFKVEGNKALVEWLHASFKGEERTVKLTHVFEFEG (SEQ ID NO:127); or
[0258] MSPEEKRVFVERFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGPGSARREITAFEVSDGVAHVDWLIAKLAGGAEVTVRLHHTFFFRPDEN (SEQ ID NO:128), and
[0259] the second polypeptide component comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to
[0260] RNELVEVKVTATPL (SEQ ID NO:129);
[0261] DKLVEVKVEIKPL (SEQ ID NO:130);
[0262] DRLVRVEVSIRPL (SEQ ID NO:131 );
[0263] RLVEVYVEIDPL (SEQ ID NO:132);
[0264] DRLVRVEVEIEPL (SEQ ID NO:133);Atty. Docket: UCSC-412WO
[0265] RLVRVSVTITPL (SEQ ID NO:134);
[0266] ENRLVRVEVEVEPL (SEQ ID NO:135);
[0267] RNELVEMKATATPL (SEQ ID NO:136);
[0268] DRLVRVEVEIEPG (SEQ ID NO:137);
[0269] RLVEVWERLPL (SEQ ID NO:138);
[0270] ELVEVKVTLTPL (SEQ ID NO:139);
[0271] KLVEVDVEAEPL (SEQ ID NO:140);
[0272] KLVRVEVERLPL (SEQ ID NO:141);
[0273] KLVEVWERLPL (SEQ ID NO:142);
[0274] DKWEVWVEIEPL (SEQ ID NO:143);
[0275] DRLVRVDVEIFPL (SEQ ID NO:144);
[0276] DRLVEVRVEIKPL (SEQ ID NO:145);
[0277] DEWEVEVDIEPL (SEQ ID NO:146); or
[0278] DRLVRVEVEIKPL (SEQ ID NO:147).
[0279] The first and second polypeptide component of the above lists can combine to provide a luciferase whose activity can be lower, higher, or same as the luciferase or luciferases from which these first and second polypeptide components originate.
[0280] In some cases, the first polypeptide component and the second polypeptide component originate from the same luciferase. Exemplary pairs of such first and second polypeptide components are set forth below:Atty. Docket: UCSC-412WOAtty. Docket: UCSC-412WONUCLEIC ACIDS
[0281] In some aspects, where the polypeptide is relatively short, e.g., include one or a few H or B domains, such polypeptides may be synthesized using synthetic chemistry. In other aspects, the present disclosure provides nucleic acids comprising nucleotide sequences encoding the polypeptides described herein. These nucleic acids may be used for cell-free transcription and translation. A nucleotide sequence encoding a subject polypeptide can be operably linked to one or more regulatory elements, such as a promoter and enhancer, that allow expression of the nucleotide sequence in a recombinant cell that is genetically modified to produce the polypeptide.
[0282] Suitable promoter and enhancer elements are known in the art. For expression in a bacterial cell, suitable promoters include, but are not limited to, lacl, lacZ, T3, T7, gpt, lambda P and trc. Forexpression in a eukaryotic cell, suitable promoters include, but are not limited to, cytomegalovirus immediate early promoter; herpes simplex virus thymidine kinase promoter; early and late SV40 promoters; promoter present in long terminal repeats from a retrovirus; mouse metallothionein-l promoter; and the like.Atty. Docket: UCSC-412WO
[0283] A nucleotide sequence encoding a subject polypeptide can be present in an expression vector and / or a cloning vector. An expression vector can include a selectable marker, an origin of replication, and other features that provide for replication and / or maintenance of the vector. Large numbers of suitable vectors and promoters are known to those of skill in the art; many are commercially available for generating a subject recombinant construct. The following vectors are provided byway of example. Bacterial: pBs, phagescript, PsiX174, pBluescript SK, pBs KS, pNH8a, pNH16a, pNH18a, pNH46a (Stratagene, La Jolla, Calif., USA); pTrc99A, pKK223-3, pKK233-3, pDR540, and pRIT5 (Pharmacia, Uppsala, Sweden). Eukaryotic: pWLneo, pSV2cat, pOG44, PXR1, pSG (Stratagene) pSVK3, pBPV, pMSG and pSVL (Pharmacia). Expression vectors generally have convenient restriction sites located near the promoter sequence to provide for the insertion of nucleic acid sequences encoding polypeptides. A selectable marker operative in the expression host cell may be present.
[0284] Nucleic acids, e.g., as described herein, may, in some instances, be introduced into a cell, e.g., by contacting the cell with the nucleic acid. Cells with introduced nucleic acids will generally be referred to herein as genetically modified cells. Various methods of nucleic acid delivery may be employed including but not limited to e.g., naked nucleic acid delivery, viral delivery, chemical transfection, biolistics, and the like.
[0285] The nucleic acids of the present disclosure may be provided in a kit. The kit may include additional components such as reconstitution buffer for resuspending the nucleic acid provided in the kit in a lyophilized form.HOST CELLS
[0286] The present disclosure provides isolated genetically modified cells (e.g., in vitro cells, ex vivo cells, cultured cells, etc.) that are genetically modified with a subject nucleic acid. In some aspects, a subject isolated genetically modified cell can produce a subject polypeptide. In some instances, a genetically modified cell may be used in the screening, and / or discovery of protein-protein interaction; protein-drug interactions; protein-nucleic acid interaction, etc.
[0287] Suitable cells include eukaryotic cells, such as a mammalian cell, an insect cell, a yeast cell; and prokaryotic cells, such as a bacterial cell. Introduction of a subject nucleic acid into the host cell can be affected, for example by calcium phosphate precipitation, DEAE dextran mediated transfection, liposome-mediated transfection, electroporation, or other known methods.Atty. Docket: UCSC-412WOKITS
[0288] Aspects of the present disclosure include kits for measuring luciferase activity of a luciferase and / or for measuring FRET from a fusion protein of the present disclosure. The kit may include components for measuring activity of a luciferase. The components may be present in separate compartments, e.g., in separate vials. Aspects of the present disclosure include kits for measuring luciferase activity of a polypeptide having luciferase activity and / or a selfcomplementing multipartite protein having luciferase activity. In certain aspects, the kit may include one or more of the polypeptides, the first polypeptide component, the second polypeptide component, and / orthe fusion proteins as disclosed herein.
[0289] Aspects of the present disclosure include kits comprising one or more nucleic acids encodingthe polypeptides, the first polypeptide component, the second polypeptide component, and / or the fusion proteins.
[0290] In certain aspects, the kit may include an assay buffer suitable for measuring luciferase activity. The kit may include one or more container means such as vials, tubes, and the like, each of the container means comprising the different polypeptides, substrates, assay buffer, etc., to be used in a method for measuring luciferase activity. For example, one of the containers may include a polypeptide having luciferase activity or a polynucleotide (e.g., in the form of a vector) encodingthe polypeptide. A second container may contain a substrate for the polypeptide. The assay buffer may be any suitable buffer such as a solution described in the present disclosure.
[0291] The kit may include a luciferin substrate, such as, DTZ, coelenterazine, fu rimazine, coelenterazine-n, coelenterazine-f, coelenterazine-h, coelenterazine-hcp, coelenterazine-cp, coelenterazine-c, coelenterazine-e, coelenterazine-fcp, bis-deoxycoelenterazine, coelenterazine-i, coelenterazine-icp, coelenterazine-v, and 2-methyl coelenterazine, or another luciferin substrate, or an analog thereof.
[0292] The kit may also include one or more buffers, such as the assay buffer disclosed herein. The kit may include instructions to enable a userto perform assays such as those disclosed herein. In certain aspects, the kit includes instructions fora method for detecting luminescence in a cell comprises contacting a cell with a luciferin substrate. In certain aspects, the cell is a live cell. In certain aspects, the cell is in vivo, ex vivo, or in vitro.
[0293] A kit of the present disclosure can include two or more of a luciferase, a first fusion protein comprising a luciferase and a first fluorescent protein, a second fusion protein comprising a luciferase and a second fluorescent protein, a third fusion protein comprising a luciferase and aAtty. Docket: UCSC-412WO third fluorescent protein. A kit of the present disclosure can include two or more of a nucleic acid encoding a luciferase, a nucleic acid encoding a first fusion protein comprising a luciferase and a first fluorescent protein, a nucleic acid encoding a second fusion protein comprising a luciferase and a second fluorescent protein, a nucleic acid encoding a third fusion protein comprising a luciferase and a third fluorescent protein, etc. The kit may also include information regarding cloning sites included in the nucleic acids, substrate, assay buffer, excitation light for the fluorescent protein, etc. The kit may also include one or both of a substrate for the luciferase and an assay buffer.Methods
[0294] The polypeptides of the disclosure may be used in any way that luciferases and fluorescent proteins have been used. For example, they may be used in a bioluminogenic method which employs a luciferin substrate to detect one or more molecules in a sample, e.g., an enzyme, a cofactor for an enzymatic reaction, an enzyme substrate, an enzyme inhibitor, an enzyme activator, or OH radicals, or one or more conditions, e.g., redox conditions. The sample may include an animal (e.g., a vertebrate), a plant, a fungus, physiological fluid (e.g., blood, plasma, urine, mucous secretions), a cell, a cell lysate, a cell supernatant, or a purified fraction of a cell (e.g., a subcellularfraction). The presence, amount, spectral distribution, emission kinetics, or specific activity of such a molecule may be detected or quantified. The molecule may be detected orquantified in solution, including multiphasic solutions (e.g., emulsions or suspensions), oron solid supports (e.g., particles, capillaries, or assay vessels).
[0295] In certain aspects, the polypeptides can be used for detecting luminescence or FRET in live cells. In some aspects, a luciferase or FRET fusion protein can be expressed in cells (as a reporter or otherwise), and the cells treated with a luciferin substrate which will permeate the cells, react with the luciferase and generate luminescence. In still other aspects, a sample (including cells, tissues, animals, etc.) containing a luciferase and a substrate may be assayed using various microscopy and imaging techniques.
[0296] The luciferases and / or FRET fusion proteins can be expressed in a model animal, e.g., mouse models. The luciferases and / or FRET fusion proteins can be injected into a live mammal for imaging purposes. For example, the luciferases and / or FRET fusion proteins can be injected into a tumor in a mammal to assist in imaging of the tumor.
[0297] Methods disclosed herein include use of a polypeptide having luciferase activity, as disclosed herein, for imaging cells expressing the polypeptide. In certain aspects, the polypeptideAtty. Docket: UCSC-412WO having luciferase activity may be expressed as a fusion protein for imaging cells expressing a protein of interest fused to the polypeptide. The polypeptides having luciferase activity, as disclosed herein, may be used for imaging live mammalian cells, e.g., a mammal.
[0298] Methods disclosed herein include use of the self-complementing multipartite protein having luciferase activity to assay for the detection of molecular interactions (e.g., transient association, stable association, complex formation, etc.) between a first moiety and a second moiety. The first and second moieties can be a peptide, a protein, a nucleic acid, a small molecule, etc. The first moiety may be conjugated to a first polypeptide component of the selfcomplementing multipartite and the second moiety conjugating the second polypeptide component to the other moiety, where the two components do not stably associate and produce no signal (e.g., substantially no signal) in the absence of the molecular interaction between the first and second moieties, but stably associate to form the self-complementing multipartite protein and produce a detectable (e.g., bioluminescent) signal upon interaction of the first and second moieties. In such embodiments, assembly of the self-complementing multipartite protein is operated by the molecular interaction of the first and second moieties. If the first and second moieties engage in a sufficiently stable interaction, the self-complementing multipartite protein having luciferase activity forms, and a bioluminescent signal is generated. If the first and second moieties fail to engage in a sufficiently stable interaction, the self-complementing multipartite protein having luciferase activity does not form, or only weakly forms, and a bioluminescent signal is not generated or is substantially reduced (e.g., substantially undetectable, essentially not detectable, differentially detectable as compared to a stable control signal, etc.). In some embodiments, the magnitude of the detectable bioluminescentsignal is proportional (e.g., directly proportional) to the amount, strength, favorability, and / or stability of the molecular interactions between the first and second moieties. In certain aspects, the first moiety may be a protein and the second moiety may be a small molecule or vice versa. In certain aspects, the first moiety is a protein and is conjugated to the first component where the first component is larger than the second component and the second component is conjugated to a second moiety that is a small molecule or vice versa.
[0299] Methods disclosed herein include use of the self-complementing multipartite protein having luciferase activity to assay for the detection of molecular interactions (e.g., transient association, stable association, complex formation, etc.) between a first moiety and a second moiety. The method may involve use of a first polypeptide component and a second polypeptideAtty. Docket: UCSC-412WO component that can associate to form the self-complementing multipartite protein having luciferase activity, when either the first or the second or both components are conjugated to a moiety and do not associate when the moiety(ies) are bound to another moiety. For example, the first polypeptide component may be fused to a first moiety and may associate with the second polypeptide component to form the self-complementing multipartite protein having luciferase activity. However, when a moiety, e.g., a ligand interacts with the first moiety, the first and second components can no longer associate to form the self-complementing multipartite protein having luciferase activity.
[0300] In some aspects, the interaction is detected in living cells, in vivo or in vitro, by detecting the bioluminescence signal emitted by the cells. In some embodiments, the interaction is detected outside a living cell, where the first and second components are secreted by the cell. In some embodiments, the interaction is detected in living organism, either inside the cells or inside tissues of the living organism.
[0301] In some aspects, an alteration in the interaction resultingfrom an alteration of the environment of the cells is detected by detecting a difference in the emitted bioluminescent signal relative to control cells absent the altered environment. In some embodiments, the altered environment is the result of adding or removing a molecule from the culture medium (e.g., a drug).
[0302] The polypeptides having luciferase activity as described herein, e.g., FRET fusion proteins are useful for many purposes including, but not limited to, detecting the amount or presence of a particular molecule (a biosensor), isolating a particular molecule, detecting conformational changes in a particular molecule, e.g., due to binding, phosphorylation or ionization, facilitating high or low throughput screening, detecting protein-protein, protein-DNAor other protein-based interactions, or selecting or evolving biosensors. For instance, a polypeptides having luciferase activity or a fusion thereof, is useful to detect, e.g., in an in vitro or cell-based assay, the amount, presence or activity of a particular kinase (for example, by inserting a kinase site into the protein), RNAi (e.g., by inserting a sequence suspected of being recognized by RNAi into a coding sequence for the protein, then monitoring reporter activity after addition of RNAi), or protease, such as one to detect the presence of a particular viral protease, which in turn is indicator of the presence of the virus, or an antibody; to screen for inhibitors, e.g., protease inhibitors; to identify recognition sites or to detect substrate specificity, e.g., using a luciferase with a selected recognition sequence or a library of polypeptides having luciferase activity having a plurality of different sequences with a single molecule of interest or a plurality (for instance, aAtty. Docket: UCSC-412WO library) of molecules; to select or evolve biosensors or molecules of interest, e.g., proteases; or to detect protein-protein interactions via complementation or binding, e.g., in an in vitro or cell-based approach. In one aspect, a polypeptide having luciferase activity which includes an inserted amino acid sequence is contacted with a random library or mutated library of molecules, and molecules identified which interact with the inserted amino acid sequence. In another aspect, a library of polypeptides having luciferase activity having a plurality of insertions is contacted with a molecule, and polypeptides having luciferase activity which interact with the molecule identified. In one embodiment, a polypeptide having luciferase activity or fusion thereof, is useful to detect, e.g., in an in vitro or cell-based assay, the amount or presence of cAMP or cGMP (for example, by inserting a cAMP or cGMP binding site into the polypeptide having luciferase activity), to screen for inhibitors or activators of, e.g., cAMP or cGMP, inhibitors or activators of cAMP binding to a cAMP binding site or inhibitors or activators of G protein coupled receptors (GPCR), to identify recognition sites or to detect substrate specificity, e.g., using a polypeptide having luciferase activity with a selected recognition sequence ora library of polypeptides having luciferase activity having a plurality of different sequences with a single molecule of interest or a plurality (for instance, a library) of molecules, to select or evolve cAMP or cGMP binding sites, or in whole animal imaging.
[0303] Also encompassed herein are methods to monitor the expression, location and / or trafficking of molecules in a cell, as well as to monitor changes in microenvironments within a cell, using a polypeptide having luciferase activity or a fusion protein thereof. In one aspect, a polypeptide having luciferase activity comprises a recognition site for a molecule, and when the molecule interacts with the recognition site, that results in an increase in activity, and thus can be employed to detect or determine the presence or amount of the molecule. For example, in one aspect, a polypeptide having luciferase activity comprises an internal insertion containing two domains which interact with each other under certain conditions. In one embodiment, one domain in the insertion contains an amino acid which can be phosphorylated and the other domain is a phosphoamino acid binding domain. In the presence of the appropriate kinase or phosphatase, the two domains in the insertion interact and change the conformation of the polypeptide having luciferase activity resulting in an alteration in the detectable activity of the modified luciferase. In another embodiment, a modified luciferase comprises a recognition site for a molecule, and when the molecule interacts with the recognition site, results in an increase in activity, and so can be employed to detect or determine the presence of amount or the other molecule.Atty. Docket: UCSC-412WO
[0304] In certain aspects, a method for detecting luminescence in a cell further comprises contacting the cell with a polypeptide having luciferase activity. In certain aspects, the polypeptide is fused to a targeting moiety that specifically binds to the cell. In certain aspects, the targeting moiety is a peptide, lipid, protein, or a small molecule. In certain aspects, the targeting moiety is an antibody or an antigen binding fragment thereof, a receptor, a ligand, or a substrate.
[0305] In certain aspects, the cell is contacted with a luciferin analog after contacting the cell with the polypeptide having luciferase activity, wherein the luciferin analog is a compound described herein or a stereoisomer, a tautomer or a salt thereof. In certain aspects, the cell is in a tissue sample. In certain aspects, the cell is in vivo in a subject.
[0306] In certain aspects, the method comprises contacting the tissue with the polypeptide having luciferase activity and fused to a targeting moiety and contacting the tissue with a luciferin substrate; and detecting localization of the polypeptide in the tissue.
[0307] In certain aspects, the method comprises administering to the subject the polypeptide having luciferase activityand fused to a targeting moietyfor localizingthe polypeptide to the cell, administering to the subject a luciferin substrate and detecting luminescence to determine localization of the polypeptide to the cell. In certain aspects, the subject is a mammal, a primate, or a human.
[0308] In certain aspects, a method for detecting luminescence in a transgenic animal comprises administering a luciferin substrate to a transgenic animal that expresses a polypeptide having luciferase activity.EXAMPLES
[0309] The following examples are offered to illustrate, but not to limit any embodiments provided by the present disclosure.Example 1 : Computational design of new de novo luciferases
[0310] A deep learning-based hallucination approach14was used to generate thousands of de novo protein scaffolds and internal pockets were designed to host a hypothetical luciferase catalytic site for diphenylterazine (DTZ)15, a synthetic luciferin. Experimental validation through site-saturation mutagenesis of LuxSit-i confirmed the designed catalytic Tyr-His and Asp-Arg dyads, and verified that substrate-binding pocket residues remained largely conserved.Atty. Docket: UCSC-412WO
[0311] To diversify the sequence space of LuxSit-i, the amino acid identities in proximity to DTZ within the pocket were fixed and ProteinMPNN16was utilized to design the remaining protein sequence (Fig. 1 a). Of the 10,000 new amino acid sequences being generated, the structures were predicted with AlphaFold217, and candidates were filtered by the predicted local distance difference test score (pLDDT > 82), backbone Ca (< 1 ,4A) and pocket cp (< 1 ,3A) root-mean-square deviation (RMSD) relative to the input structure (the design model of LuxSit-i), to ensure that the newly designed sequences encode the same fold and pocket geometry of LuxSit-i.
[0312] 191 sequences were selected, expressed in E coli, and their luciferase activity in the presence of DTZ was measured. 134 out of 191 (~70%) exhibited luciferase activity above the background (Fig. 2a). The top 20 with higher activity were chosen for scale-up expression and purification in which all 20 successfully expressed, 16 out of 20 folded mostly into monomers as per size exclusion chromatography (SEC) analysis, and the remainders displayed a mix of monomers, dimers, and soluble aggregates (Fig. 2b). Interestingly, several sequences showed enhanced luciferase activity compared to the original LuxSit-i. These sequences, with only 48-56% sequence identity compared to LuxSit-i (13.7% of which was kept fixed by design), signify the ability to explore diverse sequence spaces while preserving the same underlying structure and function (Fig. 2c and 2d). This outcome further supports that the design model of LuxSit-i is highly confident and designable, in which the sequence-structure relationship and structure-activity correlation are clear to enable aggressive sequence design.
[0313] The conformation space of loop regions were sampled, considering that enzyme activity is often sensitive to subtle changes in backbone conformation18. First, the flexible regions of LuxSit-i were predicted using ENCoM19. It was unsurprising to find that the overall fold of LuxSit-i is rigid, while Loop regions exhibited some degree of flexibility (Fig. 3a). Next, RFioint inpainting20was used to remodel loops, sequence design with ProteinMPNN was performed while the pocket residues were fixed, predicted the structures, and filtered based on pocket residue alignment (Fig. 1 a and Fig. 3b). Subsequently, six designed sequences were tested and two that exhibited over 10- fold higher activity than LuxSit-i were identified. One of these designs showed a monodisperse and monomeric SEC trace. This newly designed sequence is referred to as neoLuxI (Fig. 3c and 3d). This work showcases the successful implementation of computational protein design algorithms to explore both sequence and conformation spaces for designing new luciferases with one order of magnitude improved activity. Importantly, this approach eliminates the need for conventional labor-intensive directed evolution to improve luciferase activity.Atty. Docket: UCSC-412WO
[0314] Fig. 1 : Computational design and experimental characterization of the second- generation de novo luciferases, a, Exploration of the designed sequence space using ProteinMPNN while preserving luciferase function by fixing residues close to the luciferin substrate, leading to new amino acid sequences that fold similarly to the input structure. Loop regions were remodeled using RFjoint inpainting to further sample the conformation space of input structures, b, Coomassie-stained SDS-PAGE of purified recombinant LuxSit-i, neoLuxI , neoLuxI .2, NLuc, RLuc, and GLuc (from left to right), c, Size-exclusion chromatography of purified neoLuxI .2 suggests monodispersed and monomeric folding properties, d, Michaelis-Menten kinetics of LuxSit-i, neoLuxI , and neoLuxI .2 with substrate titrations, indicating 10 and 15-fold improved photon flux over LuxSit-i, respectively. Data presented as mean ± SEM (n = 3). e, Luciferase activity of neoLuxI .2 (10 nM) in the presence of various luciferin analogs (25 pM), highlighting its highly specific for DTZ. f, Temperature-dependent luciferase activity for neoLuxI .2 (cyan), GLuc (green), NLuc (blue), and RLuc (purple), with each curve depicting normalized luciferase activities as a function of temperature. The midpoint temperatures (Tm) represent the transition inflection point where luciferase activity reaches 50% of its maximum value, g, Emission kinetics of neoLuxI .2, GLuc, NLuc, and RLuc (1 nM) in the presence of their respective luciferin substrates (50 pM), illustrating the prolonged emission profile of neoLuxI .2.
[0315] Fig. 2. Characterization of activity and protein folding of newly designed luciferase sequences generated by ProteinMPNN. a, Luminescence intensity of each newly designed luciferase sequence, normalized to LuxSit-i R65A (an inactive mutant as 65Arg is the key catalytic residue). Sequences displaying luminescence at leastthree times above the baseline (blue dotted line) were deemed positive. The top 20 sequences, surpassing the threshold indicated by the green dotted line, were selected for detailed biochemical analysis, b, Size-exclusion chromatography (SEC) analysis of top 20 recombinant luciferase sequences, revealing their respective folding states, c, Sequence logo plot representing the large sequence diversity explored among the 20 selected sequences. 16 out of 117 amino acid identities (letters in blue) in the pocket (13.7%) were fixed during ProteinMPNN design, d, Histogram plot shows the distribution of sequence identity for the top 20 sequences relative to the original LuxSit-i sequence, ranging from 48% to 56%.
[0316] Fig. 3. Evaluation of de novo luciferase sequences designed by RFjoint Inpainting, a, The colored representation of LuxSit-i structure is based on ENCoM-predicted B-factors, which indicate regions of predicted flexibility within the protein structure. Warmer colors represent higher predicted flexibility, while cooler colors denote rigidity, b, Inpainted regions were highlighted inAtty. Docket: UCSC-412WO which the corresponding loops were remodeled to sample conformation space (left). Sequence alignment displays the inpainted areas of the proteins (right), c, Luminescence intensity for each inpainted protein in the presence of 50 pM DTZ was plotted where some showed enhanced activity levels compared to the original LuxSit-i. d, Size-exclusion chromatography (SEC) traces of the inpainted proteins. The variant 1c1 was selected and named neoLuxI due to its improved activity and monomeric properties, e, Assessment of mutations within the enzyme pocket to identify the V83L mutant (neoLuxI .2). The panel presents the results from mutagenesis experiments to evaluate the impact of these mutations in the pocket on luciferase activities, f, Zoom-in views of AlphaFold2-predicted neoLuxI .2 model (blue) and LuxSit-i (grey) designed models in the pocket. The designed interaction residues are similar at the side-chain level and the hypothetical catalytic Tyr-His (green) and Asp-Arg (orange) dyads were kept intact. The V83L mutation in neoLuxI .2 is shown in red.Example 2: Improvement, characterization, and benchmarking of neoLux series
[0317] To investigate the transferability of mutations identified in the previous study to enhance luciferase activity14, those mutations were introduced combinatorically into the pocket of neoLuxI . The result indicated that a V83L mutation in the pocket of neoLuxI increased luciferase activity by 47%, which is designated neoLuxI .2 while most of the core residues remained the same as designed (Fig. 3e and 3f). This approach demonstrates the synergy between computational sampling and empirical optimization to improve the overall brightness of designer luciferases.
[0318] Biochemically, both neoLuxI and neoLuxI .2 are 13.7kDa proteins (123 residues), the smallest among commonly used luciferases, which are 62%, 28%, and 25% smaller than RLuc, NLuc, and Glue, respectively (Fig. 1 b). Both are highly soluble and expressed well in E. coli, combined with monodispersed and monomeric folding as indicated by size-exclusion chromatography (SEC) analysis (Fig. 1c and Fig. 4a). neoLuxI and neoLuxI .2 exhibited ~10-fold and ~15-fold improvement in maximum brightness than LuxSit-i, and showed low micromolar Kmvalues close to LuxSit-i, reflecting the strategy of preserving the pocket residues through design (Fig. 1 d and Table 1). Importantly, neoLuxI .2 maintains the exquisite substrate specificity for DTZ (Fig. 1e), unlike the promiscuous substrate recognition observed in native or engineered luciferases (Fig. 4c). From here, we demonstrated the tailored design of orthogonal luciferase-luciferin pairs and enhance the luciferase activity significantly without sacrificing substrate selectivity, which is adept at producing a distinct signal for multiplexing assays or imaging purposes.Atty. Docket: UCSC-412WO
[0319] The principle of de novo protein design aims to find the lowest-energy sequence for their structures21, thereby de novo proteins generally have exceptionally high folding stability22. Circular dichroism (CD) analyses revealed that neoLuxI .2 retains structural integrity even at 95 °C and is capable of reversible refolding upon cooling to 25 °C, while other native or engineered luciferases, such as RLuc, NLuc, and GLuc, irreversibly unfolded during temperature ramping (Fig. 4d). Furthermore, neoLuxI .2 maintained its activity across a broad range of temperature and retained ~83% activity even after exposure to 100 °C for one hour (Fig. 1 f), unlike RLuc, NLuc, and GLuc which experienced a 50% reduction in activity at 52 °C, 66 °C, and 91 °C, respectively (Table 1). These results collectively highlighted the exceptional thermostability of neoLuxI .2, surpassing even GLuc, which features five covalently linked disulfide bonds for its high structural stability.
[0320] An overlooked aspect of luciferase functionality is the kinetics of signal decay. While researchers often use light intensity as a proxy for luciferase expression in reporter assays for biological evaluation23, it's vital to recognize that the light intensity is a time-dependent decay process, due to the consumption of luciferin substrate or irreversible modifications that permanently deactivate the catalytic center of luciferases. Given that RLuc, NLuc, and GLuc are inherently flash-type luciferases, their signal decay half-lives usually range from a few minutes (Fig. 1g). This can cause signal decay artifacts and inaccuracies when the experimental conditions are not carefully optimized, including instrumental setup (e.g., injector and well-to-well time-delay), luciferase / luciferin concentrations, insufficient time point measurements, etc. In contrast, neoLuxI .2 offers stable glow-type emission kinetics with an extended decay half-life (—43 minutes) in a simple and commercially available HBS buffer (Fig. 1g), making our probes ideal for high- throughput screening by obviating the need for injectors and mitigatingthe impact of signal decay artifacts. Our designer luciferase exhibits a balance between brightness and decay half-life in comparison to luciferases with high initial brightness but rapid decay2425. In addition, following treatment with luciferin substrates, our mass spectrometry analysis revealed the presence of heterogenous +16Da oxidation species with NLuc, but not with neoLuxI .2, (Fig. 4e-f). This observation suggested that neoLuxI .2 possesses a catalytic center of greater stability, and the rapid signal decay observed with NLuc does not solely stem from substrate depletion but also from the covalent modifications that render enzyme deactivation, similar to the previous findings for GLuc26.
[0321] Fig. 4. Additional expression, purification, activity, and structural characterization of neoLux series, native, and engineered luciferases, a, SEC traces for recombinant neoLuxI , NLuc,Atty. Docket: UCSC-412WORLuc, and GLuc (from left to right), b, Enzyme kinetic of GLuc, NLuc, and RLuc with titrated CTZ substrate, fitted to the Michaelis-Menten equation to determine Vmax and Km values. Quantitative values were listed in Table 1 . c, Substrate specificity comparison among neoLuxI , NLuc, RLuc, and GLuc (from left to right) in the presence of 25 pM of the indicated substrates. Luminescence intensity for each luciferase was normalized to the highest signal among all tested luciferin substrates, with neoLuxI showing high DTZ specificity. Data were shown as the maximum luminescence signal ± SEM (n = 3). d, Far-ultraviolet circular dichroism (CD) spectra of neoLuxI .2, NLuc, RLuc, and GLuc (from left to right) at 25 °C (black), 95 °C (red), and cooled back to 25 °C (blue). MRE denotes molar residue ellipticity. The overlapping of the black and blue curves suggests the reversible folding of neoLuxI .2, unlike NLuc, RLuc, and GLuc, which do not exhibit this reversibility. Deconvoluted mass spectra for e, neoLuxI .2, and f, NLuc indicated the correct molecular mass (left panel). After exposure to DTZ substrate, NLuc exhibited a series of laddered +16Da species (right panel), indicating the occurrence of covalently heterogeneous oxygenation of NLuc after catalyzing light emission. In contrast, the active site of neoLuxI .2 remained intact and exhibited resistance to such modifications.Example 3: Designing efficient luciferase-fluorescent protein fusions
[0322] In the pursuit of expanding the color repertoire of neoluminescent proteins to facilitate the monitoring of multiple biological events simultaneously, we reasoned that our compact de novo luciferases could serve as efficient energy donors in Forster Resonance Energy Transfer (FRET) systems. Accordingto the FRET equation (Fig. 5a), the FRET efficiency is inversely proportional to the sixth power of the distance between the donor and acceptor. To model this crucial factor, the distance (r), we utilized AlphaFold2 structure prediction to model the optimal spatial arrangement and the linker lengths between FRET donor and acceptor, ensuring minimal distance while preserving the overall structural integrity (Fig. 5a). As a result, we found that the predictions can capture the shortest distance between neoLuxI .2 (FRET donor) and various fluorescent proteins (FRET acceptor) while suggesting suitable lengths for linker truncation (Fig. 6a and 6b).
[0323] After surveying the FPbase with the extensive repertoire of FPs27, we selected mNeonGreen28, mGold29, mKok30, CyOFPI31, and mKate232as the FRET acceptors based on their absorption spectra, quantum yields, and emission wavelengths. Based on the prediction results, we constructed and screened a strategically designed linker library to fine-tune the orientationAtty. Docket: UCSC-412WO factor (K2) for each FP (Fig. 6c). The screening process resulted in the identification of five FRET pairs (Fig. 5b and 5c) - luxNeon, luxGold, luxKok, luxOFP, and luxKate -with efficiencies surpassing previously reported pairs using native or engineered luciferases15,33,34. For instance, luxOFP emits at 588 nm, similar to Antares2 and ReNL, but achieved a much higher 91% energy transfer efficiency than Antares2 and ReNL (65%, Fig. 6d). Additionally, luxOFP is 45% smaller than that of Antares2 and ReNL (Fig. 6e). From our biochemical evaluation, we observed all FRET pairs are brighter than or equivalent to neoLuxI .2 itself (Fig. 6f), potentially due to the improved quantum yields via FRET. Moreover, all of our FRET pairs maintain similar Kmvalues and signal decay half-lives compared to neoLuxI .2, indicating that the FRET fusions did not disrupt either the FP folding (Fig. 6g) nor the luciferase activity. In brief, our strategy, aided by structural prediction, circumvents the need for extensive library screening and yields significantly more efficient FRET pairs, which is especially pronounced in the far-red spectral range (e.g. 81 % efficiency and Aem= 635 nm for luxKate), where energy transfer efficiencies have historically been suboptimal with native or engineered luciferases35.
[0324] Fig. 5: The design of neoLux-based FRET pairs for multiplexed imaging, a, AlphaFold2 was used to predict the distances between the FRET donor and acceptor, ensuring the integrity of fusion protein folding while identifying the minimal distances between neoLux and the fluorescent proteins (FPs). Spectra overlap ( / ), quantum yields (ct>), orientation factor (K2), and refraction index (n4) contribute to the Forster distance (Ro6). The FRET efficiency is majorly influenced by the distance (r) between the donor and acceptor, following an exponent of six. b, Photograph captured with a Canon 5D camera showcased distinct emissions from neoLuxI .2, luxNeon, luxGold, luxKok, luxOFP, and luxKate (tubes from left to right), c, Normalized emission spectra for neoLuxI .2, luxNeon, luxGold, luxKok, luxOFP, and luxKate in the presence of DTZ, indicating high FRET efficiencies, d, Luminescence microscopy of HeLa cells expressing luxNeon, luxGold, luxKok, luxOFP, luxKate, or Antares2 (top to bottom) under various channels, e, Excitation- free multiplexed luminescence image of subcellular structures using luxNeon for mitochondria; luxGold for plasma membrane (Lyn); and luxOFP for nucleus (H2B) localizations. Signals were resolved using linear unmixing. Scale bar=10 pm.
[0325] Fig. 6. Computational modeling and experimental characterization of neoLux-FP FRET pairs, a, AlphaFold predictions illustrating the structural model of luciferase-fluorescent fusion proteins, with distances calculated between the FRET donor and acceptor to determine the optimal linker length. The truncation site was chosen to minimize these distances, resulting in theAtty. Docket: UCSC-412WO shortest calculated distance while preserving structural integrity (e.g., removal of the first 7 residues of mKate2). b, Distance graphs for additional fluorescent proteins including mNeonGreen, mGold, mKOk, and CyOFPI (from left to right). The arrows indicate the selected position for constructing neoLux-FPs libraries, c, Experimental evaluation of randomized linker libraries (plus 2xNNK codons at the junction) by truncating nine (d9), eleven (d11), and thirteen residues (d 13) from the N-terminal of CyOFP. The shortest predicted truncation number (9 for CyOFP) revealed a transition point where fluorescence colony numbers were reduced. This suggested improper folding of the fluorescent protein when the linker is below the predicted length, d, Normalized luminescence emission spectra of Antares2, ReNL, and Nanolantern (from left to right), e, SDS- PAGE gel analysis of all FRET pairs used in this study (from left to right: neoLuxI .2, luxNeon, luxGold, luxKOk, luxOFP, luxKate, Antares2, ReNL, and Nanolantern). Samples were not denatured by heat, f, DTZ titration curves for each FRET pair, fitted to the Michaelis-Menten equation to calculate Km and Vmax parameters. All FRET pairs have similar or higher Vmax compared to neoLuxI .2. Quantitative values were listed in Table 3. g, Emission kinetics of FRET pairs in the presence of their respective luciferin substrates to calculate the decay half-life listed in Table 3. h, Excitation and emission spectra of the fluorescent proteins (FPs) from five FRET pairs show signatures consistent with the original full-length FPs, confirming the formation of the intended chromophores.Example 4: Efficient FRET pairs enable multiplexed cellular imaging
[0326] As our FRET pairs retain the same photophysical characteristics as the original FPs (Fig. 6g), users can perform microscopic imaging under eithertraditional fluorescence or excitation-free luminescence modalities. Unmixing multiple luminescence signals should be, in principle, simpler than fluorescence due to the absence of excitation spectral bleed-through under luminescence imaging mode36. Given the narrower emission spectra of our FRET pairs, we reason that we can differentiate multiple luminescent signals using standard optical filters. To assess the multiplexing capabilities of our probes, we expressed these designer FRET pairs in HeLa cells and successfully acquired live cell imaging at single-cell resolution (Fig. 5d). Due to their high FRET efficiencies, each emission can be effectively separated using conventional filters; in contrast, Antares2 with lower FRET efficiency exhibited severe crosstalk between each channel (Fig. 5d), which limits its multiplexing capability. Moreover, even though the spectral separation from luxNeon and luxGold is only 12 nm, we showed that the luxNeon / luxGold / luxOFP combination canAtty. Docket: UCSC-412WO enable the labeling of multiple sub-cellular organelles within the same cell after spectral unmixing (Fig. 5e), offering a new dual-modality neoluminescent toolkit for monitoring complex cellular dynamics without phototoxicity and photobleaching. Notably, although luxNeon exhibits 5-fold lower Vmax than Antares2 at the purified protein level, we observed a higher signal-to-noise ratio from luxNeon-expressing cell compared to Antares2-expressing cell (Fig. 6), likely due to superior cellular expression and folding of luxNeon in human cells.Example 5: Efficient FRET pairs enable real-time dual- and triple-luciferase bioassay
[0327] Luciferase assay is widely used as a reporter in molecular biology for monitoring gene expression3738. Traditionally, one luciferase gene (typically from the firefly or Renilla) is used as the experimental reporter, and the other as an internal control in dual-luciferase assays. Due to the intrinsic differences in light generation mechanisms of FLuc and RLuc, the traditional dualluciferase assay involves the sequential addition of two luciferin substrates and quenching steps. Our FRET pairs have transformed the configurations possible with luciferase assay, as these highly efficient FRET pairs emit distinguishable signals that can be readily separated using optical filters. Here, we showed that our FRET pairs can enable dual luciferase assays to acquire both reporter (luxNeon or luxGold) and reference (luxOFP) signals simultaneously without cumbersome stop- and-go procedures (Fig. 7a). Since our probes utilize the same DTZ luciferin substrate, discrepancies arising from using two different luciferase systems - such as varied signal intensity, buffer conditions, and emission kinetics - can be avoided39. The integration of our FRET pairs in reporter assays allows the measuring of ratiometric signals, which helps circumvent artifacts associated with time-sensitive decay. The assay format is designed to remove experimental variables such as cell numbers, cell types, plasmid amounts, and transfection efficiency, which enhanced the reliability and consistency of the data obtained from luciferase reporter assays (Fig 8b).
[0328] Furthermore, after constructing our neoluminescent reporters downstream of commonly used regulatory promoters, we observed dose-dependent expression curves in which the response curve EC5o are similar to commercialized luciferase kits (Fig. 7c). Since the signals from luxNeon and luxGold can be spectrally unmixed using 508 / 20 nm and 540 / 25 nm filters (Fig. 8a), we further expanded the assay’s capabilities to a triple-luciferase assay format. This allows for the concurrent quantification of two cellular pathways alongside an internal reference in which we successfully quantified the activation of cAMP-PKA and NFKB pathways from the same transfectedAtty. Docket: UCSC-412WO cells through spectrally unmixable triple-luciferase assay format (Fig. 7d). Additionally, the cell- permeable nature of the DTZ substrate1540enables continuous monitoring of gene expression in live cells - providing time course information beyond the reach of traditional lysis-based end-point assays. For instance, we showed a faithful recording of dynamic changes in the cAMP-PKA signaling with PEST-tagged luxNeon / CMV-luxOFP, in which the pathway activation was observed over extended durations after the addition of forskolin, and the luciferase level dropped after treating with KG-501 , a CRE signaling inhibitor (Fig 7e). Overall, the expanded flexibility and multiplexing capacity provided by our neoluminescence technology pave the way for sophisticated experimental designs and provide opportunities for deeper exploration into cellular mechanisms.
[0329] Fig. 7. The signal and noise comparison between HeLa cells expressing luxNeon or Antares2 at single-cell resolution microscopic imaging. Each violin plot represents data collected from 100 points (pixels) sampled from the 14 cells with the highest signal of the respective image, along with 100 pixels from the background for noise measurement. The distribution indicated that luxNeon exhibited a higher signal-to-noise ratio compared to Antares2 in HeLa cells.
[0330] Fig. 8. Neoluminescent probes facilitate reliable, versatile, and multiplexed luciferase bioassays, a, Schematic diagram illustrated the experimental setup for assessing quantitative relationships between multiple responsive luciferases in a single emission recording experiment. HEK293 cells were co-transfected with plasmids encoding transcriptionally responsive luxNeon and / or luxGold, along with a constitutive CMV promoter containing luxOFP, allowing simultaneous dual and triple luciferase assays, b, Signals from luxNeon and luxOFP can be simultaneously acquired in two channels (508 / 20 nm and 620 / 40 nm) to calculate ratiometric readouts, ensuring reproducible and reliable bioassays. Incorporation of luxOFP reference eliminated the effect of various experimental factors such as cell numbers, cell types, transfection efficiency, plasmid quantity, and time-sensitive decay artifacts, c, The discosed designer luciferases work broadly in reporting different commonly used response elements, including TGFp / activin, AP1 , Wnt, cAMP-PKA, and NFKB signaling pathways, d, Triple luciferase assay enabled simultaneous monitoring of three signals (508 / 20 nm, 540 / 25 nm, and 620 / 40 nm): two for the activities of labeled pathways by incorporating specific NFKB and cAMP-PKA regulatory elements upstream of luxNeon and luxGold, along with luxOFP as the reference under a constitutive CMV promoter, enabling two ratiometric measurements simultaneously after spectral unmixing. The unmixed ratios clearly show the individual (FSK orTNFa) and combined (FSK + TNFa) effects of these treatments on the respective signaling pathways, e, Real-time live-cell assaysAtty. Docket: UCSC-412WO enable kinetic measurements of luciferase expression within the cAMP-PKA pathway, facilitating the study of dynamic transcriptional responses overtime in the same sample. The first arrow indicated the addition of 100 pM forskolin, which increased luminescence after 2-4 h. The second arrow indicates the addition of the KG-501 inhibitor, which blocks CREB signaling.Example 6: FRET pairs enable multiplexed in vivo imaging
[0331] Current bioluminescence tools are predominantly employed for tracking a single parameter in vivo. Achieving multiplexed imaging in vivo poses a significant challenge, often requiring either sequential delivery of different substrates36or the use of two luciferases with spectrally distinct outputs41. Notably, all current dual luciferase in vivo imaging studies rely on at least one ATP-dependent luciferase42-44. Yet, the inherent differences in the mechanisms of light emission of different luciferases and substrate biodistribution among different luciferin substrates inevitably introduce imaging biases.
[0332] Our neoluminescent probes provide an innovative solution to address the multiplexing challenge. First, our designer probes are specific to one synthetic luciferin, DTZ, preventing substrate cross-reactivity. This unique feature enables the use of two ATP-independent luciferases for imaging within a single in vivo object, thereby mitigating fundamental imaging biases. To demonstrate this, we first produced lentivirus-transduced HeLa cell lines that stably express either Nanolatern45(an RLuc-Venus fusion), luxNeon, luxGold, luxKok, luxOFP, or luxKate. We verified their luminescence intensity, emission spectra, and expression in cells (Fig. 11 ) and subsequently used these cell lines for subcutaneous xenograft experiments in mice. Due to the high substrate specificity, luxNeon, luxGold, and luxOFP can be selectively illuminated only after the administration of DTZ while Nanolatern is reactive only with CTZ (Fig. 9a, left panel). This substrate-specific feature offers users the chemo-selectivity to image a particular labeled event at the desired time point. Secondly, the photons emitted from these efficient FRET probes can be spectrally resolved by standard band-pass filters (Fig. 10b). This capability was further confirmed in xenograft models, where the signals were simultaneously distinguished using MS filters (Fig. 9a, right panel) without the need for multiple sequential substrate administrations that complicate the experimental setup, data analysis, and can impose time constraints and additional stress on the animals. We also demonstrated that the luminescence signals from two other combinations - luxNeon / luxKok / luxOFP and luxNeon / luxOFP / luxKate - can also be distinguished in vivo usingAtty. Docket: UCSC-412WO standard IVIS filters and spectral unmixing (Fig. 9b), which highlights the high FRET efficiencies for easy spectral unmixing and versatility of our neoluminescent probes.
[0333] Cancer heterogeneity is a major factor in the complexity of cancer treatment and research, including how cancers develop, metastasize, and respond to treatments46. However, effective, non-invasive, and real-time methods to image tumor heterogeneity in vivo are lacking. To assess the effectiveness of our FRET probes in monitoring heterogeneous environments in living animals, we established subcutaneous xenografts of two-population (Fig. 9c) and three-population (Fig. 9d) mixed HeLa cells with different proportions in which each population is labeled with either luxNeon, luxOFP, or luxKate. In this model, we were able to distinguish heterogeneous tumors with varying cell populations in vivo after spectral unmixing. We also continuously acquired the unmixed images for 14 days post heterogeneous tumor implantation (Fig. 9e). The results showed that we can monitor the growth of each population over time, indicating that these cell mixtures can be simultaneously spectrally resolved at the same location during a single imaging event. This enables real-time and non-invasive observation of temporal dynamic changes in tumor composition.
[0334] We further applied this approach to image the mixtures of HeLa and B16F10 melanoma heterogeneous tumors using these distinct FRET probes. After unmixing, we successfully differentiated the signals, representing the growth of each labeled cell type in vivo overtime. We observed that the tumor with a sub-population of B16F10 melanoma cells (Fig. 9f, right panel) proliferated more than other sub-populations and suppressed the growth of the luxKate-labeled HeLa sub-population. This suppression was not seen in tumors composed entirely of HeLa cells (Fig. 9f, left panel), underscoring the unpredictable nature of cancer heterogeneity in live animals while our method offers dynamic, on-site information without the need to sacrifice the mice. To confirm our in vivo analysis, we harvested and sectioned the tumors at the study end point, followed by fluorescence microscopy to quantify the relative contribution of each cell population percentage, in which the in vitro validation agreed with our observation of distinct tumor cell populations in vivo (Fig. 12). These findings underscore the capacity of our neoluminescent toolkit to deliver precise spatiotemporal resolution in vivo, allowing for the simultaneous detection of multiple probes with a single substrate. The high specificity and multiplexing ability of our probes represent a leap forward in the in vivo analysis of complex biological systems in realtime.
[0335] Fig. 9: Multiplexed neoluminescence imaging of tumor xenografts in vivo, a, 5x105of various luciferase-expressing HeLa cells were injected into the left (luxNeon)Zright (Nano-lantern) dorsolateral trapezius regions, and the left (luxGold)Zright (luxOFP) dorsolateral thoracolumbarAtty. Docket: UCSC-412WO regions of NGS mice. Due to the high specificity of neoLuxI .2, luminescence signals from luxNeon, luxGold, and luxOFP were exclusively detectable following intravenous injection with DTZ, but not CTZ. Conversely, Nano-lantern emitted light in the presence of CTZ, allowing substrate-resolved orthogonal imaging using all ATP-independent luciferases. Signals from luxNeon, luxGold, and luxOFP were effectively resolved using IVIS filters, enabling simultaneous multi-color imaging in the same subject, b, Spectral-unmixing of luxNeon / luxKok / luxOFP (top row) and luxNeon / luxOFP / luxKate (bottom row) combinations in mice. The composite image merges these signals, showing their spatial distribution in vivo, c, HeLa cells expressing either luxNeon, luxOFP, or luxKate, were mixed in various pairwise proportions (pie charts) and engrafted in NGS mice at different subcutaneous sites. Images were spectrally unmixed to derive signals from each channel and merged as the composite image, d, Mixtures of three-population HeLa cells expressing luxNeon, luxOFP, and luxKate with different mixed percentages were prepared to simulate cancer heterogeneity. Spectrally unmixed images were used to distinguish the proportions of each population at different sites in vivo, e, Tracking three unmixed signals overtime allows real-time monitoring of three HeLa cell population growth in vivo. Pie charts show the proportion of each labeled cell. The graphs display unmixed signal intensities at each channel over 14 days postimplantation. Images of solid tumors harvested on day 15 were shown, f, Comparison of mixing equal percentages of luxNeon, luxOFP, and luxKate HeLa cells with a mixture where luxNeon- expressing HeLa cells are substituted with luxNeon-expressing B16F10 melanoma cells. The change in the micro-environment, as a result of substituting cell types, affected the growth of the heterogeneous tumor.
[0336] Fig. 10: Spectral unmixing and multiplexed imaging of FRET pairs by conventional filters, a, Graphs show linear spectral unmixing to determine 508 / 620 ratios from luxNeon and 540 / 620 ratios from luxGold. Different ratios totaling 100% of either purified luxNeon or luxGold with a constant amount of luxOFP in solutions (left) and in intact cells co-expressing luxNeon / luxOFP or luxGold / luxOFP (right) were mixed as indicated. After adding DTZ, filtered light from the 508 / 20 nm and 540 / 25 nm channels were successfully unmixed to show the respective emission ratios, b, Comparison of the normalized emission of specified FRET reporters under individual IVIS Spectrum emission filters. Our findings reveal that our FRET probes emit predominantly within their designated channels, showcasing high FRET efficiency. In contrast, Antares2 exhibits bleed-through across multiple channels. Triplicates were done for each probe on the same plate.Atty. Docket: UCSC-412WO
[0337] Fig. 11 . Luminescence and fluorescence analysis of B16F10 and HeLa cells expressing various FRET probes after lentiviral transduction, a, Time-dependent luminescence emission from 10,000 B16F10 or HeLa cells expressing indicated FRET probes in the presence of DTZ (left panel). The bar plot compares the highestthree data points of each stable cell line (right panel). Data reflects total relative light units (RLU) without considering PMT sensitivity across different wavelengths, b, Luminescence emission spectra of luxNeon in B16F10, luxNeon in HeLa, luxGold in HeLa, luxKOk in HeLa, luxOFP in HeLa, and luxKate in HeLa cells (from left to right) stably expressing these probes, indicating proper chromophore maturation, c, Microscopic analysis of HeLa and B16F10 cells post-lentiviral transduction, expressing corresponding FRET probes. Images include brightfield (top), fluorescence channel (middle), and merged views (bottom), illustrating nearly 100% transduction efficiency and probe expression.
[0338] Fig. 12. Unmixed images of heterogeneous tumors in vivo and endpoint analysis of Tumor 4 by fluorescence-activated cell sorter (FACS), a, Unmixed images at 520 nm, 570 nm, and 670 channels of tumor 1 , 2, 3, and 4 over a 14-day imaging period. Image quantitative plots are shown in Fig. 4e and 4f. These images show the temporal and spatial growth patterns of each cell population within the tumors, allowing for the observation of heterogeneous tumor development in living subjects without sacrificing animals, b, FACS histograms displaying fluorescence intensity counts for an equally mixed population of cultured HeLa cells expressing luxNeon, luxOFP, and luxKate as positive control. This suggests the ability to separately quantify individual populations using 488ex / 525em (left), 488ex / 610em (middle), and 561ex / 675em (right) channels, c, FACS histograms displaying the fluorescence intensity of heterogeneous cells dissociated from a biopsy of tumor 4 at the endpoint. Each histogram corresponds to a different excitation / emission channel. The colored peaks in the histograms represent detected signals under that channel, where both luxNeon-expressing and luxOFP-expressing cells were identified, but the population of luxKate- expressing cells was not detected. The ex vivo FACS results align with the unmixed luminescence signals observed in vivo from the same tumor.
[0339] Fig. 13. Additional xenograft tumor images to showcase the inherent complexity of tumor heterogeneity in vivo, a, Tumor growth dynamics for four different tumors (Tumor 5, 6, 7, and 8) with various initial proportions of HeLa or B16F10 cells expressing luxNeon, luxOFP, and luxKate FRET probes, as shown by the pie charts, were imaged over 14 days post-implantation. Unmixed signal intensities at three wavelengths (520 nm, 570 nm, and 670 nm) represent each population’s growth curve. These three-population tumors were implanted in different locations on another NGSAtty. Docket: UCSC-412WO mouse where we observed varied growth curves for each population in vivo. The average and standard deviation were plotted for the three highest intensities recorded during each imaging session, b, Unmixed images at 520 nm, 570 nm, and 670 channels of tumor 5, 6, 7, and 8 across 14 days of imaging period.
[0340] Discussion
[0341] Currently, there is a lack of genetically encoded probes capable of imaging and tracking biological events across molecular, cellular, and individual levels. Our strategy in creating designer luciferases represents a significant advancement in the development of ideal luminescent toolkits, overcoming the limitations imposed by native luciferases. For the first time, a fully artificial enzyme illuminates biological processes across micro to macro scales — from individual proteins and cellular cultures to in vivo heterogeneous tumor models — which provide a versatile and non-invasive platform for real-time, multiplexed studies of biological events.
[0342] The historical reliance on ATP-dependent firefly luciferases for in vivo imaging, imposing a notable metabolic burden particularly critical in pre-clinical studies spanning extended durations40,47,48, underscores the need for a transition to mechanistically superior ATP-independent luciferases. These suite of designer luciferases, characterized by ATP-independency, small size, robust folding, extreme stability, and excellent substrate specificity, positions them as preferable probes for bioimaging applications, particularly ideal for integration into viral gene delivery systems. The simultaneous dual- and triple-luciferase assay setups demonstrated here ensure robust and more reproducible bioassay results for biomedical research and drug discovery.References1 . Wang, M., Da, Y. & Tian, Y. Fluorescent proteins and genetically encoded biosensors. Chem. Soc. Rev. 52, 1189-1214 (2023).2. Frei, M. S., Mehta, S. & Zhang, J. Next-generation genetically encoded fluorescent biosensors illuminate cell signaling and metabolism. Annu. Rev. Biophys. 53, (2024).3. Townsend, K. M. & Prescher, J. A. Recent advances in bioluminescent probes for neurobiology. Neurophotonics 11, (2024).4. Yeh, H. W. & Ai, H. W. Development and Applications of Bioluminescent and Chemiluminescent Reporters and Biosensors. Annual Review of Analytical Chemistry, Vol 12 12, 129-150 (2019).5. Liu, S., Su, Y., Lin, M. Z. & Ronald, J. A. Brightening up Biology: Advances in Luciferase Systems for in Vivo Imaging. 4CS Chem. Biol. (2021 ) doi: 10.1021 / acschembio.1 c00549.Atty. Docket: UCSC-412WO Wu, N. etal. Solution structure of Gaussia Luciferase with five disulfide bonds and identification of a putative coelenterazine binding cavity by heteronuclear NMR. Sci. Rep. 10, (2020). Loening, A. M., Wu, A. M. & Gambhir, S. S. Red-shifted Renilla reniformis luciferase variants for imaging in living subjects. Nat. Methods 4, 641-643 (2007). Hall, M. P. etal. Engineered luciferase reporter from a deep sea shrimp utilizing a novel imidazopyrazinone substrate. ACS Chem. Biol. 7, 1848-1857 (2012). Yeh, H.-W. etal. ATP-lndependent Bioluminescent Reporter Variants To Improve in Vivo Imaging. ACS Chem. Biol. 14, 959-965 (2019). Syed, A. J. & Anderson, J. C. Applications of bioluminescence in biotechnology and beyond. Chem. Soc. Rev. 50, 5668-5705 (2021). Xiong, Y. etal. Engineered Amber-emitting nano Luciferase and its use for immunobioluminescence imaging in vivo. J. Am. Chem. Soc. 144, 14101-14111 (2022). Iwano, S. etal. Single-cell bioluminescence imaging of deep tissue in freely moving animals. Science 359, 935-939 (2018). Delroisse, J., Duchatelet, L., Flammang, P. & Mallefet, J. Leaving the dark side? Insights into the evolution of luciferases. Front. Mar. Sci. 8, (2021). Yeh, A. H.-W. etal. De novo design of luciferases using deep learning. Nature 614, 774-780 (2023). Yeh, H. W. etal. Red-shifted luciferase-luciferin pairs for enhanced bioluminescence imaging. Nat. Methods 14, 971-974 (2017). Dauparas, J. etal. Robust deep learning-based protein sequence design using ProteinMPNN. Science 378, 49-56 (2022). Jumper, J. etal. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-+ (2021). Corbella, M., Pinto, G. P. & Kamerlin, S. C. L. Loop dynamics and the evolution of enzyme activity. Nat. Rev. Chem. 7, 536-547 (2023). Frappier, V. & Najmanovich, R. J. A coarse-grained elastic network atom contact model and its use in the simulation of protein dynamics and the prediction of the effect of mutations. PLoS Comput. Biol. 10, e1003569 (2014). Wang, J. etal. Scaffolding protein functional sites using deep learning. Science 377, 387-394 (2022). Kortemme, T. De novo protein design — From new structures to programmable functions. Cell 187, 526-544 (2024). Cao, L. etal. De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science 370, 426-431 (2020). Sarrion-Perdigones, A. etal. Examining multiple cellular pathways at once using multiplex hextuple luciferase assaying. Nat. Commun. 10, 5710 (2019). Jiang, T. Y., Du, L. P. & Li, M. Y. Lighting up bioluminescence with coelenterazine: strategies and applications. Photochem. Photobiol. Sci. 15, 466-480 (2016). Markova, S. V., Larionova, M. D. & Vysotski, E. S. Shining Light on the Secreted Luciferases of Marine Copepods: Current Knowledge and Applications. Photochem. Photobiol. 95, 705-721 (2019).Atty. Docket: UCSC-412WO Dijkema, F. M. et al. Flash properties of Gaussia Luciferase are the result of covalent inhibition after a limited number of cycles. Protein Sci. 30, 638-649 (2021 ). Lambert, T. J. FPbase: a community-editable fluorescent protein database. Nat. Methods 16, 277-278 (2019). Shaner, N. C. et al. A bright monomeric green fluorescent protein derived from Branchiostoma lanceolatum. Nat. Methods 10, 407-409 (2013). Lee, J. etal. Versatile phenotype-activated cell sorting. Sci. Adv. 6, eabb7438 (2020). Tsutsui, H., Karasawa, S., Okamura, Y. & Miyawaki, A. Improving membrane voltage measurements using FRET with new fluorescent proteins. Nat. Methods 5, 683-685 (2008). Chu, J. etal. A bright cyan-excitable orange fluorescent protein facilitates dualemission microscopy and enhances bioluminescence imaging in vivo. Nat. Biotechnol. 34, 760-767 (2016). Shcherbo, D. etal. Far-red fluorescent tags for protein imaging in livingtissues. Biochem. J. 418, 567-574 (2009). Schaub, F. X. etal. Fluorophore-NanoLuc BRET reporters enable sensitive in vivo optical imaging and flow cytometry for monitoring tumorigenesis. Cancer Res. 75, 5023-5033 (2015). Suzuki, K. etal. Five colourvariants of bright Luminescent protein for real-time multicolour bioimaging. Nat. Common. 7, 13718 (2016). Weihs, F. & Dacres, H. Red-shifted bioluminescence Resonance Energy Transfer: Improved tools and materials for analytical in vivo approaches. Trends Analyt. Chem. 116, 61-73 (2019). Brennan, C. K. etal. Multiplexed bioluminescence imagingwith a substrate unmixing platform. Cell Chem. Biol. 29, 1649-166O.e4 (2022). Bioluminescence: Fundamentals and Applications in Biotechnology- Volume 2. (Springer, Berlin, Germany, 2016). Calabretta, M. M. & Michelini, E. Current advances in the use of bioluminescence assays for drug discovery: an update of the last ten years. Expert Opin. Drug Discov. 1- 11 (2023). Branchini, B. R. etal. A firefly Luciferase Dual color bioluminescence reporter assay using two substrates to simultaneously monitor two gene expression events. Sci. Rep. 8, 5990 (2018). Yeh, H.-W., Wu, T., Chen, M. & Ai, H.-W. Identification of Factors Complicating Bioluminescence Imaging. Biochemistry 58, 1689-1697 (2019). Kleinovink, J. W. etal. A dual-color bioluminescence reporter mouse for simultaneous in vivo imaging of T cell localization and function. Front. Immunol. 9, 3097 (2018). Liu, S. etal. Molecular imaging reveals a high degree of cross-seeding of spontaneous metastases in a novel mouse model of synchronous bilateral breast cancer. Mol. Imaging Biol. 24, 104-114 (2022). Zambito, G. etal. Red-shifted click beetle luciferase mutant expands the multicolor bioluminescent palette for deep tissue imaging. iScience 24, 101986 (2021 ). Su, Y. C. et al. Novel NanoLuc substrates enable bright two-population bioluminescence imaging in animals. Nat. Methods 17, 852-860 (2020).Atty. Docket: UCSC-412WO Takai, A. etal. Expanded palette of Nano-lanterns for real-time multicolor luminescence imaging. Proc. Natl. Acad. Sci. U. S. A. 112, 4352-4356 (2015). Marusyk, A., Janiszewska, M. & Polyak, K. Intratumor heterogeneity: The Rosetta stone of therapy resistance. Cancer Cell 37, 471-484 (2020). Wang, L. etal. Application of bioluminescence resonance energy transfer-based cell tracking approach in bone tissue engineering. J. Tissue Eng. 12, 2041731421995465 (2021 ). Hikita, T., Miyata, M., Watanabe, R. & Oneyama, C. In vivo imaging of long-term accumulation of cancer-derived exosomes using a BRET-based reporter. Sci. Rep. 10, 16616 (2020). Boitet, M. et al. Biolum’ RGB: A low-cost, versatile, and sensitive bioluminescence imaging instrument for a broad range of users. ACS Sens. 7, 2556-2566 (2022). Watson, J. L. etal. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089-1100 (2023). Adhikari, S. Generalized biomolecular modeling and design with RoseTTAFold allatom. (2024) doi:10.1242 / prelights.36373. Dauparas, J. etal. Atomic context-conditioned protein sequence design using LigandMPNN. b / oRx / v(2023) doi:10.1101 / 2023.12.22.573103.
Claims
Atty. Docket: UCSC-412WOWhat is claimed is:CLAIMS1 . A polypeptide having luciferase activity and comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 85% identity to any one of the following amino acid sequences:1c1 (neoLuxI )MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTG AKRKIESVEIKDGEAWKVTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPL (SEQ ID NO:1); 2b9MSAAEVRDFVDRFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGP GSARREITAFEVSDGVAHVDWLRAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL (SEQ ID NO:2);2c7MSPEEKRVFVERFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGP GSARREITAFEVSDGVAHVDWLIAKLAGGAEVTVRLHHTFFFRPDENRLVRVEVEVEPL (SEQ ID NO:3):1 c3MSPEEIRDFVKRFYEALDAGDAETAAQLLWDAGCRRIELWDGTVFEGPDVRDQFVAWFRALQASV TGAKREILKVEVKDGTVAWEVRLTATYKATGKTFWRLTHVFTFDPETGELVEVKVTLTPL (SEQ ID NO:4);1 b12MSEEIREFVDRFYAALDAGDADTAADLLFSSGCKKIHLWDGTVFDGDKEAFKAWFEDLFAKSEGAT RRVTSFAVDLDGLPRADVEVELTTTIDGKEVRVRLRHTFYFDAEGRLVEWVERLPL (SEQ ID NO:5); 1f3MSEEEKREFVERFYAALDKGGEEGAEEAADLLFSSGCKEIHLWDGRVFTSKEEFKAWFVELWASLG EKGARREVTAFEVNEDGTAVVDVVLTAEWKDGTVRWRLRHVFHFEDGKLVRVEVERLPL (SEQ ID NO:6); and2a 1MSEEEMREFVERFYAALDAGDAETASSLLFDSGCKKIHLWDGRVFTSKEEFKDWFRHLHEDVLEG AVRKVTSFEVDPEKGVAWDWLTARVKATGEEVQVRLRHTFYFEEGKLVEVWERLPL (SEQ ID NO:7).Atty. Docket: UCSC-412WO2. The polypeptide of claim 1, wherein the amino acid sequence has at least 85% identity to SEQ ID NO:1.
3. The polypeptide of claim 2, wherein the amino acid sequence comprises one or more of the following amino acids: D at position 18, V at position 83, and L at position 100, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
4. The polypeptide of claim 2, wherein the amino acid sequence comprises one or more of the following amino acids: E at position 18, L at position 83, and I at position 100, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
5. The polypeptide of claim 2, wherein the amino acid sequence comprises one or more of the following amino acids: L at position 17, D at position 18, and V at position 83, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
6. The polypeptide of claim 2, wherein the amino acid sequence comprises one or more of the following amino acids: I at position 17, E at position 18, and L at position 83, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
7. The polypeptide of claim 2, wherein the amino acid sequence comprises E at position 18, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
8. The polypeptide of claim 2, wherein the amino acid sequence comprises L at position 17 and V at position 83, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
9. The polypeptide of claim 2, wherein the amino acid sequence comprises I at position 17 and L at position 83, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
10. The polypeptide of claim 2, wherein the amino acid sequence comprises one or more of the following amino acids: D at position 18, V at position 83, and V at position 98, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.Atty. Docket: UCSC-412WO11 . The polypeptide of claim 2, wherein the amino acid sequence comprises one or more of the following amino acids: E at position 18, L at position 83, and L at position 98, wherein the numbering of the amino acid position is based on the numbering of amino acid positions in SEQ ID NO:1.
12. The polypeptide of claim 2, wherein the amino acid sequence comprises D at position 18 and I at position 37, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
13. The polypeptide of claim 2, wherein the amino acid sequence comprises E at position 18 and V at position 37, and I at position 85, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
14. The polypeptide of claim 2, wherein the amino acid sequence comprises D at position 18 and V at position 98, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
15. The polypeptide of claim 2, wherein the amino acid sequence comprises E at position 18 and L at position 98, wherein the numbering of the amino acid position is based on the numbering of the amino acid position in SEQ ID NO:1 .
16. The polypeptide of claim 2, wherein the amino acid sequence comprises V at position 83.
17. The polypeptide of claim 2, wherein the amino acid sequence comprises L at position 83.
18. The polypeptide of any one of claims 2-17, wherein the amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:1 or wherein the amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequences set forth in Figs. 15A-15C.
19. The polypeptide of claim 1, wherein the amino acid sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:2 or wherein the amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequences set forth in Figs. 16A-16C.
20. The polypeptide of claim 1 , wherein the amino acid sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at leastAtty. Docket: UCSC-412WO98%, at least 99% identical to SEQ ID NO:3 or wherein the amino acid sequence is at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequences set forth in Figs. 17A-17C.21 . The polypeptide of claim 1 , wherein the amino acid sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:4.
22. The polypeptide of claim 1 , wherein the amino acid sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:5.
23. The polypeptide of claim 1 , wherein the amino acid sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:6.
24. The polypeptide of any one of claims 1-23, comprisingthe secondary structure arrangement H1 -L1 -H2-L2-B1 -L3-B2-L4-H3-L5-B3-L6-B4-L7-B5-L8-B6-L9, wherein “H” is a helical domain, “L” is a loop domain, and “B” is a beta strand domain and comprising catalytic dyads of (i) D residue at position 20 and R residue at position 69; and (ii) Y residue at position 16 and H residue at position 104, wherein the numbering of the amino acid position is based on the numberingof the amino acid position in SEQ ID NO:1 .
25. The polypeptide of claim 24, wherein the H1 domain is at least or up to 20 amino acids in length; the L1 domain is at least or up to 4-5 amino acids in length; the H2 domain is at least or up to 9 amino acids in length; the L2 domain is at least or up to 5-6 amino acids in length; the B1 domain is at least or up to 2 amino acids in length; the L3 domain is at least or up to 3 amino acids in length; the B2 domain is at least or up to 4 amino acids in length; the L4 domain is at least or up to 4-5 amino acids in length; the H3 domain is at least or up to 12-16 amino acids in length; the L5 domain is at least or up to 5-8 amino acids in length; the B3 domain is at least or up to 8-10 amino acids in length; the L6 domain is at least or up to 3-7 amino acids in length;Atty. Docket: UCSC-412WO the B4 domain is at least or up to 10-12 amino acids in length; the L7 domain is at least or up to 4-6 amino acids in length; the B5 domain is at least or up to 12-14 amino acids in length; the L8 domain is at least or up to 4-7 amino acids in length; the B6 domain is at least or up to 6-9 amino acids in length; and the L9 domain is at least or up to 3 amino acids in length.
26. The polypeptide of claim 24 or 25, wherein one or more of the following is true: the H1 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: MSGTDEEIAEFVKAFYEAID (SEQ ID NO:29); MSAAEVRDFVDRFYAALD (SEQ ID NO:30); or MSPEEKRVFVERFYAALD (SEQ ID NO:31 ); the H2 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: ETAADLLFG (SEQ ID NO:32), AAASGLL (SEQ ID NO:33), or AAASGL (SEQ ID NO:34); the B1 domain comprises an amino acid sequence having at least 50% or 100% identity to the amino acid sequence: IH or IE; the B2 domain comprises an amino acid sequence having at least 50% or 100% identity to the amino acid sequence: GTV or GT; the H3 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: EGFETWFKKLQS (SEQ ID NO:35), VLEQFETWFNGLLST (SEQ ID NO:36), or DDVLEQFETWFNGLLS (SEQ ID NO:37); the B3 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: RKIESVEI (SEQ ID NO:38), RREITAFEV (SEQ ID NO:39), or RREITAFEV (SEQ ID NO:39); the B4 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: AWKLTLTAT (SEQ ID NO:40), VAHVDWLRAK (SEQ ID NO:41 ), or AHVDWLIAK (SEQ ID NO:42);Atty. Docket: UCSC-412WO the B5 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: KFWELEHTFTF (SEQ ID NO:43), AEVTVRLHHTFFFR (SEQ ID NO:44), or VTVRLHHTFFF (SEQ ID NO:45); and the B6 domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence: ELVEVKVTA (SEQ ID NO:46), RLVRVEVEV (SEQ ID NO:47), or RLVRVEVEV (SEQ ID NO:47).
27. The polypeptide of any one of claims 24-26, wherein the L1 , L2, L3, L4, L5, L6, L7, L8, and L9 domains are at least 1 , 2, 3, 4, 5, 6, 7, or 8 amino acids in length and comprise any amino acid and optionally are up to 5, 6, 7, or 8 amino acids in length.
28. A fusion protein comprising the polypeptide of any one of claims 1 -27 and another polypeptide.
29. The fusion protein of claim 28, wherein the other polypeptide is a fluorescent polypeptide.
30. The fusion protein of claim 29, wherein the fluorescent polypeptide is mNeonGreen, Gold, Kok, Kate, CyOF1 , or CyOFP.31 . The fusion protein of claim 29, wherein the fluorescent polypeptide is mNeonGreen and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLNSLPAT HELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGM SPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTA ADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHS KTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:55);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLFDNM ASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLP YPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTAtty. Docket: UCSC-412WONSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKT ELKHSKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:56);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPRDDN MASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQY LPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFR KTELKHSKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:57);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPTGMAS LPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYP DGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTN SLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTE LKHSKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:58);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPISSLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDG MSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLT AADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKH SKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:59);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLNSLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPD GMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSL TAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELK HSKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:60); orMGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPPYSLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLT AADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKH SKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:61),Atty. Docket: UCSC-412WO or wherein the fusion protein comprises a fusion of any one of the polypeptides of any one of claims 1 -27 and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to: SLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPY PDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMT NSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKT ELKHSKTELNFKEWQKAFTDVMGMDELYK (SEQ ID NO:62), optionally wherein (i) fluorescent protein lacks the amino acids correspondingto the first 1 -5, 1 -8, or 1 -1 1 amino acids at the N-terminus of the mNeonGreen fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, 1 -6, 1-8, 1 -12, 1 -14, 1 -16, or 1-18 amino acids.
32. The fusion protein of claim 29, wherein the fluorescent polypeptide is Gold and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYFTGWP ILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDHMK QHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYN SHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPN EKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:63);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATGSELFT GVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPD HMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEY NYNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSK DPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:64);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATHGELFTAtty. Docket: UCSC-412WOGWPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPD HMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEY NYNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSK DPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:65);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPWLLFT GWPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPD HMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEY NYNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:66);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPQLLFT GWPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPD HMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEY NYNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:67);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPWDTGV VPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDH MKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYN YNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:68); orMGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYFTGV VPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARYPDH MKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYN YNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:69), or wherein the fusion protein comprises a fusion of any one of the polypeptides of any one of claims 1 -27 and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, atAtty. Docket: UCSC-412WO least 97%, at least 98%, at least 99%, or 100% to::FTGWPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTSLGYGLQCFARY PDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHK LEYNYNSHNVYITADKQKNGIKANFKIRHNIEDGGVQLADHYQQNTPIGDGPVLLPDNHYLSYQSK LSKDPNEKRDHMVLLEFVTAAGITLGMDELYK (SEQ ID NO:70), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -5, 1 -8, or 1 -9 amino acids at the N-terminus of the Gold fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -3, 1 -4, 1 -6, 1-8, 1 -13, or 1 -15 amino acids.
33. The fusion protein of claim 29, wherein the fluorescent polypeptide is mKok and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATIQAEMKM RYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPE EIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVD WEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNIT EQVEDAVAHS (SEQ ID NO:71 );MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATITPEMK MRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKY PEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSV DWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGN ITEQVEDAVAHS (SEQ ID NO:72);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATIAKEMK MRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKY PEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSV DWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGN ITEQVEDAVAHS (SEQ ID NO:73);Atty. Docket: UCSC-412WOMGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATIQAEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGN ITEQVEDAVAHS (SEQ ID NO:74);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATLTHEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQVEDAVAHS (SEQ ID NO:75),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPNPIKPEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGN ITEQVEDAVAHS (SEQ ID NO:76),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPSPIKPEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGN ITEQVEDAVAHS (SEQ ID NO:77),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPPPKPEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGN ITEQVEDAVAHS (SEQ ID NO:78),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPAIKPEMKMRYAtty. Docket: UCSC-412WOYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEI PDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWE PSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQ VEDAVAHS (SEQ ID NO:79), MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVT GAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPILPEMKMRYY MDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIP DYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEP STEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQV EDAVAHS (SEQ ID NO:80), or MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVT GAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIQPEMKMRYY MDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHRVFTKYPEEIP DYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQNQSVDWEP STEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRKTEGNITEQV EDAVAHS (SEQ ID NO:81 ), or wherein the fusion protein comprises a fusion of any one of the polypeptides of any one of claims 1 -27 and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to: PEMKMRYYMDGSVNGHEFTIEGEGTGRPYEGHQEMTLRVTMAEGGPMPFAFDLVSHVFCYGHR VFTKYPEEIPDYFKQAFPEGLSWERSLEFEDGGSASVSAHISLRGNTFYHKSKFTGVNFPADGPIMQ NQSVDWEPSTEKITASDGVLKGDVTMYLKLEGGGNHKCQFKTTYKAAKEILEMPGDHYIGHRLVRK TEGNITEQVEDAVAHS (SEQ ID NO:82), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -4, 1 -5, or 1 -6 amino acids at the N-terminus of the mKok fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 - 2, 1 -3, 1-4, 1 -5, 1-6, 1 -8, or 1-11 amino acids.
34. The fusion protein of claim 29, wherein the fluorescent polypeptide is cyOFP and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%,Atty. Docket: UCSC-412WO at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYIKENMR SKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLP DYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWE PSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETY VEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:83);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATDSIKEN MRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPA DLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLG WEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADN ETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:84);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYIKEN MRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPA DLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLG WEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADN ETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:85);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPFVKEN MRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPA DLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLG WEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADN ETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:86);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPSKKEN MRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPA DLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLG WEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADN ETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:87);Atty. Docket: UCSC-412WOMGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPILENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:88);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIRENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:89),MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTETPIKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGW EPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:90);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPHIKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:91 ); orMGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATSPIKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKVVEGGPLPFAFDILATHFMYGSKVFIKYPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKKTLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKEADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:92); or wherein the fusion protein comprises a fusion of any one of the polypeptides of any one of claims 1 -27 and a fluorescent protein comprising an amino acid sequenceAtty. Docket: UCSC-412WO having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:IKENMRSKLYLEGSVNGHQFKCTHEGEGKPYEGKQTNRIKWEGGPLPFAFDILATHFMYGSKVFIK YPADLPDYFKQSFPEGFTWERVMVFEDGGVLTATQDTSLQDGELIYNVKVRGVNFPANGPVMQKK TLGWEPSTETMYPADGGLEGRCDKALKLVGGGHLHVNFKTTYKSKKPVKMPGVHYVDRRLERIKE ADNETYVEQYEHAVARYSNLGGGMDELYK (SEQ ID NO:93), optionally wherein (i) fluorescent protein lacks the amino acids correspondingto the first 1 -5, 1 -8, 1 -9, 1 -10, or 1 -1 1 amino acids at the N-terminus of the cyOFP fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, 1 -6, 1 -8, 1-13, or 1 -15 amino acids.
35. The fusion protein of claim 29, wherein the fluorescent polypeptide is mKate and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQ SAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIRENMH MKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIP DFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEA STETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKET YVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:94);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPCLIKEN MHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQ GIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGW EASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADK ETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:95);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKL QSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPPLIKEN MHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQ GIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWAtty. Docket: UCSC-412WOEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:96);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPILENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:97);MGSGTDEEIAEFVKAFYEAIDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPIRENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:98),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYPLIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLXDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:99),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPNPLIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NQ:100),MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYYIKENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:101),Atty. Docket: UCSC-412WOMGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVT GAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPMFIKENMHMK LYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDF FKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTE TLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVE QHEVAVARYCDLPSKLGHR (SEQ ID NQ:102), MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVT GAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPYIKENMHMKL YMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFF KQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTET LYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVE QHEVAVARYCDLPSKLGHR (SEQ ID NQ:103), or MGSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVT GAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLIKENMHMKL YMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINHTQGIPDFF KQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTLGWEASTET LYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKEADKETYVE QHEVAVARYCDLPSKLGHR (SEQ ID NO:104); or wherein the fusion protein comprises a fusion of any one of the polypeptides of any one of claims 1 -27 and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to: ENMHMKLYMEGTVNNHHFKCTSEGEGKPYEGTQTMRIKAVEGGPLPFAFDILATSFMYGSKTFINH TQGIPDFFKQSFPEGFTWERVTTYEDGGVLTATQDTSLQDGCLIYNVKIRGVNFPSNGPVMQKKTL GWEASTETLYPADGGLEGRADMALKLVGGGHLICNLKTTYRSKKPAKNLKMPGVYYVDRRLERIKE ADKETYVEQHEVAVARYCDLPSKLGHR (SEQ ID NO:105), optionally wherein (i) fluorescent protein lacks the amino acids correspondingto the first 1 -4, 1 -5, 1 -6, or 1 -7 amino acids at the N-terminus of the mKate fluorescent protein and / or (ii) the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, or 1-6 amino acids.Atty. Docket: UCSC-412WO36. The fusion protein of claim 29, wherein the fluorescent polypeptide is Neon-PEST and the fusion protein comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:MTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSA VTGAKRKIESVEIKDGEAWKLTLTATYKATGKKFWELEHTFTFDRERNELVEVKVTATPLNSLPATHE LHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSP FQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTAAD WCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKT ELNFKEWQKAFTDVMGMDELYKNSHGFPPEVEEQAAGTLPMSCAQESGMDRHPAACASARINV (SEQ ID NO:106) or wherein the fusion protein comprises a fusion of any one of the polypeptides of any one of claims 1-27 and a fluorescent protein comprising an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to:SLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGF HQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEGSHIKGEAQVKGTGFPAD GPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPM YVFRKTELKHSKTELNFKEWQKAFTDVMGMDELYKNSHGFPPEVEEQAAGTLPMSCAQESGMDR HPAACASARINV (SEQ ID NG:107), optionally wherein (i) fluorescent protein lacks the amino acids corresponding to the first 1 -6, 1 -8, 1 -10, or 1 -11 amino acids at the N-terminus of the Neon-PEST fluorescent protein and / or (ii) wherein the C-terminus of any one of the polypeptides is fused to the N-terminus of the fluorescent protein via a linker, wherein the linker has a length of 1 -2, 1 -4, or 1-6 amino acids.
37. The fusion protein of any one of claims 28-36, wherein the C-terminus of the polypeptide of any one of claims 1 -27 is fused to the N-terminus of a fluorescent protein via a linker, optionally wherein the fluorescent protein includes a deletion of at least or up to the first 6- 11 amino acids at the N-terminus relative to the corresponding full-length fluorescent protein.Atty. Docket: UCSC-412WO38. A self-complementing multipartite protein having luciferase activity, comprising at least a first polypeptide component and a second polypeptide component, wherein the at least first polypeptide component and the second polypeptide component are not covalently linked or are covalently linked via a cleavable linker, wherein in total the first polypeptide component and the second polypeptide component comprise the secondary structure arrangement H1 -L1 -H2-L2-B1 -L3-B2-L4-H3-L5-B3-L6-B4-L7-B5-L8-B6-L9, wherein each domain is as defined in any one of claims 24-37; wherein (a) each H and B domain is fully present within one polypeptide component of eitherthe first polypeptide component orthe second polypeptide component, (b) the first polypeptide component and the second polypeptide component do not include all of the H and B domains, (c) the relative order of the H, L, and B domains when present in either of the first polypeptide component orthe second polypeptide component is unchanged with reference to the protein as defined in any one of claims 24-37, and (d) the first component and the second component when not present in the self-complementing multipartite protein do not possess detectable luciferase activity or have luciferase activity lower than the luciferase activity of the self-complementing multipartite protein.
39. The self-complementing multipartite protein of claim 38, wherein the first polypeptide component and the second polypeptide component comprise the secondary structure arrangement as set forth in Table 2:Table 2:Atty. Docket: UCSC-412WOwherein the L domain in parenthesis is (i) present in one but not both of the first and second polypeptide components, (ii) is split between the first and second polypeptide components, or (iii) absent.
40. The self-complementing multipartite protein of claim 38 or 39, wherein(i) one or both of the first polypeptide component and the second polypeptide component comprises an additional domain, ii) one or both of the first polypeptide component and the second polypeptide component comprises an additional domain covalently linked to one or both of the first polypeptide component and the second polypeptide component,(iii) the first polypeptide component is a fusion protein comprising a first domain and the second component is a fusion protein comprising a second domain, wherein the first and second domains can associate with each other; or(iv) the self-complementing multipartite protein comprises from N-terminus to C-terminus,(a) the first polypeptide component, a linker, and the second component or(b) the second component, a linker, and the first component, wherein the selfcomplementing multipartite protein has luciferase activity and wherein upon cleavage ofAtty. Docket: UCSC-412WO the linker, the self-complementing multipartite protein has substantially reduced cleavage activity or substantially undetectable cleavage activity.41 . The self-complementing multipartite protein of any one of claims 38-40, wherein the first polypeptide component comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to MSGTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEIHLWDGTVFTSKEGFETWFKKLQSAVTGAKRKIESVEIKDGEAWKVTLTATYKATGKKFWELEHTFTFDRE (SEQ ID NO:109);MSEEEQREFVDRFYAALDAGDAETASALFPDGTKIYLWDGKVFTTREEFRAWFEKLYSTSENAKRH WSFKVDGNKADVEWLHANINGEKKTVRLRHVFYFEG (SEQ ID NO:110);MSAEQQREFVKRFYEALDAGDADTASALFPDGTEIHLWDGTTFRTRAEFRAWFEELYSTSENASRE VTSFSVDGDVADVEWLRANLGGEDRTVSLRHVFHFAG (SEQ ID NO:111 );MSSDAQRAFVDRFYRALDAGDAETASALFPDGTRIHLWDGTTFTTREEFRAWFVDLRSRSENAARE WSFDVDGDVAHVEWLKAVIEGEEVWRLRHVFEWEGD (SEQ ID NO:1 12);MSAEAQRRFVDRFYAALDAGDADTASALFPDGTEIHLWDGRTFRTRAEFRAWFRELRARSDNARR EWAFEVDGDTAHVEWLRASIDGEERWRLRHTFYFEG (SEQ ID NO:1 13);MSEEEIREFVRRFYEALDAGDAATASALFPDGTEIHLWDGTTFRTQAQFRAWFERLRAQSANARREI VDLKVEGDRAKVEVILRASFDGEEKWNLTHEFLFEGD (SEQ ID NO:114);MSAAEVRDFVDRFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGP GSARREITAFEVSDGVAHVDWLRAKLAGGAEVTVRLHHTFFFRPD (SEQ ID NO:115);MTDEEIAEFVKAFYEALDAGDAETAADLLFGAGCKEVHLWDGTVFTSKEGFETWFKKLQSAVTGAK RKIESVEIKDGEAWKLTITATYKATGKKFWEIENTFTFDRE (SEQ ID NO:1 16);MSAEAHRRFVDRFYAALDAGDADTASALFPDGTEIHLWDGRTFRTRAEFRTWFRELRARSDNARR EWAFEVDGDTAHVEWLRASIDGEERWRLRHTFYFEG (SEQ ID NO:1 17);MSEEIREFVDRFYAALDAGDADTAADLLFSSGCKKIHLWDGTVFDGDKEAFKAWFEDLFAKSEGAT RRVTSFAVDLDGLPRADVEVELTTTIDGKEVRVRLRHTFYFDAEG (SEQ ID NO:118);MSPEEIRDFVKRFYEALDAGDAETAAQLLWDAGCRRIELWDGTVFEGPDVRDQFVAWFRALQASV TGAKREILKVEVKDGTVAWEVRLTATYKATGKTFWRLTHVFTFDPETG (SEQ ID NO:1 19);MSPEEKKTFVDRFYAALDAGDAKTAADLLFGDDGKCRIRLWDGREFVDDKEAFERWFEGLLSLTEP GTAKREWAFEVDENGRAHVDWLTARVKGSADEFFVRLHHTFYFEDG (SEQ ID NG:120);Atty. Docket: UCSC-412WOMSEEEKREFVERFYAALDKGGEEGAEEAADLLFSSGCKEIHLWDGRVFTSKEEFKAWFVELWASLG EKGARREVTAFEVNEDGTAWDWLTAEWKDGTVRWRLRHVFHFEDG (SEQ ID NO:121 );MSEEEMREFVERFYAALDAGDAETASSLLFDSGCKKIHLWDGRVFTSKEEFKDWFRHLHEDVLEG AVRKVTSFEVDPEKGVAWDWLTARVKATGEEVQVRLRHTFYFEEG (SEQ ID NO:122);MSEEEQREFVARFYAALDAGDAETASALFPDGTEIHLWDGKTFTTRAEFRAWFEKLHSLSDNASRH VTSFKVDGNVAEVEWLHADFKGKKLTVKLRHRYQFEG (SEQ ID NO:123);MSKEEQEEFVKQFYEALDAGDAETASALFPDGTVIHLWDGKTFHTQAEFRAWFEELKSTSENAKRE VTKFEVDGDVADVEWLKANINGEEKWNLKHKFKFEG (SEQ ID NO:124);MSEEEIKEFVKRFYEALDAGDAETASALFPDGTRIYLWDGRVFRTRAEFRAWFVELHSTSEDAKREVI ELKVEGNVAKVKWLHANINGEKKTVLLEHYFEFEG (SEQ ID NO:125);MSEEEIREFVRRFYEALDAGDAETASALFKDGTKIYLWDGTVFETREEFRAWFVELYSKSENARRRV VSFKVDGNVAEVEWLHASFQGEDKWRLKHRFKFEG (SEQ ID NO:126);MSEESQREFWKRFYAALDAGDAETASALFPDGTEIHLWDGTVFRTRAEFRAWFVDLHSKSDNASR EITSFKVEGNKALVEWLHASFKGEERTVKLTHVFEFEG (SEQ ID NO:127); orMSPEEKRVFVERFYAALDAGDAAAASGLLEGIEEIHLWDGTVFSGDDVLEQFETWFNGLLSTLTGP GSARREITAFEVSDGVAHVDWLIAKLAGGAEVTVRLHHTFFFRPDEN (SEQ ID NO:128).
42. The self-complementing multipartite protein of any one of claims 38-41 , wherein the second polypeptide component comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% toRNELVEVKVTATPL (SEQ ID NO:129);DKLVEVKVEIKPL (SEQ ID NQ:130);DRLVRVEVSIRPL (SEQ ID NO:131 );RLVEVYVEIDPL (SEQ ID NO:132);DRLVRVEVEIEPL (SEQ ID NO:133);RLVRVSVTITPL (SEQ ID NO:134);ENRLVRVEVEVEPL (SEQ ID NO:135);RNELVEMKATATPL (SEQ ID NO:136);DRLVRVEVEIEPG (SEQ ID NO:137);RLVEVWERLPL (SEQ ID NO:138);Atty. Docket: UCSC-412WOELVEVKVTLTPL (SEQ ID NO:139);KLVEVDVEAEPL (SEQ ID NO:140);KLVRVEVERLPL (SEQ ID NO:141);KLVEVWERLPL (SEQ ID NO:142);DKWEVWVEIEPL (SEQ ID NO:143);DRLVRVDVEIFPL (SEQ ID NO:144);DRLVEVRVEIKPL (SEQ ID NO:145);DEWEVEVDIEPL (SEQ ID NO:146); orDRLVRVEVEIKPL (SEQ ID NO:147).
43. The self-complementing multipartite protein of any one of claims 38-42, wherein the first polypeptide component complements with a second component, wherein the first and second polypeptide components pairs are:Atty. Docket: UCSC-412WOAtty. Docket: UCSC-412WOor wherein the first polypeptide complement and the second polypeptide complement of a pair comprises an amino acid sequence having a sequence identity of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to the first polypeptide component and the second polypeptide component, respectively, as set forth above.
44. A nucleic acid comprising a nucleotide sequence encoding the polypeptide of any one of claims 1 -27, the fusion protein of any one of claims 28-37, or the first polypeptide component and the second polypeptide component of any one of claims 38-43.
45. An expression vector comprising the nucleic acid of claim 44 operatively linked to an expression control element.
46. A recombinant host cell comprising the polypeptide of any one of claims 1 -27, the fusion protein of any one of claims 28-37, the first polypeptide component and the second polypeptide component of any one of claims 38-43, the nucleic acid of claim 44, or the expression vector of claim 45.Atty. Docket: UCSC-412WO47. A first nucleic acid comprising a nucleotide sequence encoding the first polypeptide component of any one of claims 38-43 and a second nucleic acid comprising a nucleotide sequence encoding the second polypeptide component of any one of claims 38-43.
48. A first vector comprising the first nucleic acid of claim 47 and a second vector comprising the second nucleic acid of claim 47.
49. A kit comprising the nucleic acid of claim 44, the vector of claim 45, the recombinant host cell of claim 46, the first nucleic acid and the second nucleic acid of claim 48, or first vector and the second vector of claim 45, optionally wherein (i) the first nucleic acid and the second nucleic acid are in separate containers or (ii) first vector and the second vector are in separate containers.
50. The kit of claim 49, wherein the vector or the first vector or the second vector comprises a multiple cloning site.51 . The kit of claim 49 or 50, comprising a substrate for the polypeptide of any one of claims 1- 27, the fusion protein of any one of claims 28-37, or the self-complementing multipartite protein of any one of claims 38-43, optionally, wherein the substrate is Diphenylterazine (DTZ).
Citation Information
Patent Citations
Red-shifted luciferase-luciferin pairs for enhanced bioluminescence
US20180057801A1
De novo designed luciferase
WO2023137417A2