High-yield method for developing genetically encoded fluorescent indicators based on orphan bacterial proteins

A high-throughput method using bioinformatics and multiwell plate readers to identify and annotate orphan bacterial proteins, facilitating the development of GEFIs for metabolic monitoring and detection of biomedically relevant molecules.

WO2026050877A1PCT designated stage Publication Date: 2026-03-12SAN SEBASTIAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Current methods for identifying effector/ligand interactions with orphan bacterial proteins are low-throughput, costly, and require specialized equipment, limiting the development of Genetically Encoded Fluorescent Indicators (GEFIs) for metabolite monitoring.

Method used

A high-throughput method using bioinformatics tools to identify orphan bacterial proteins, followed by fusion with fluorescent proteins and screening with standard multiwell plate readers to detect conformational changes induced by ligand binding, enabling functional annotation and development of GEFIs.

Benefits of technology

Enables the identification and functional annotation of orphan bacterial proteins, allowing for the development of novel GEFIs that can monitor cellular metabolism with high spatial and temporal resolution and detect biomedically important molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000016_0001
    Figure IMGF000016_0001
  • Figure 00000019_0000
    Figure 00000019_0000
  • Figure 00000019_0001
    Figure 00000019_0001
Patent Text Reader

Abstract

The present invention relates to a method for developing genetically encoded fluorescent indicators based on orphan bacterial proteins. This approach uses a high-yield, unbiased method to identify orphan protein effectors or ligands. The technique is based on detecting variations in the spectroscopic signal that are due to the interaction of a test molecule with a protein, which is fused with one or two fluorescent proteins. The method can be implemented using purified proteins from systems such as bacteria, yeasts, insects or mammal cells, or using fusion proteins targeting the periplasm of Escherichia coli. The assessment is performed by optical system readers in 96- or 384-well formats, which allows for an efficient and scalable analysis. Once the specific ligand has been identified, the corresponding bacterial protein serves as a scaffold to design genetically encoded fluorescent indicators. The method incorporates a bioinformatic algorithm that facilitates the identification of orphan protein families based on genome or metagenome resources. These proteins are subsequently fused with fluorescent proteins and are assessed by comparing them with molecule libraries in order to determine their functionality. This approach allows the functional annotation of orphan proteins as transcription factors or periplasmic proteins, and their use in designing, developing and optimising fluorescent indicators. As a result, the ability to monitor emerging metabolites and metabolic networks with high precision and specificity is improved, in addition to facilitating the development of protocols to detect biomedically relevant molecules in biological fluids.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] HIGH-YIELD METHOD FOR DEVELOPING GENETICALLY ENCODED FLUORESCENT INDICATORS BASED ON ORPHAN BACTERIAL PROTEINS

[0002] FIELD OF INVENTION

[0003] The present invention comprises a method for identifying orphan bacterial protein effectors / ligands using fluorescence, and employing these proteins as scaffolds to develop Genetically Encoded Fluorescent Indicators (GEFIs). This method exploits variations in the fluorescence signal induced by the binding of a test molecule to the orphan protein fused to one or two fluorescent proteins. Specifically, ligand binding modulates the electrical charge of the chromophore solvent and / or the distance or orientation of the fluorescent proteins, affecting the fluorescence readout. This perturbation can be used to identify the orphan protein effector / ligand from a library of small molecules. The method can be performed with purified protein from bacteria, yeast, insects, or mammalian cells, or with periplasmic fusion proteins in Escherichia coli.Screening can be performed in multiwell plate format using standard fluorescence readers or by using a photodetection system, where such a system is any instrument capable of measuring one or more properties of photons, including, but not limited to, their intensity, wavelength, lifetime, polarization, or their spatial and temporal distribution. Once the ligand is identified, the corresponding bacterial protein can be used as a scaffold to insert circularly permuted fluorescent proteins into its backbone and develop highly sensitive single-fluorophore indicators for a given ligand.

[0004] BACKGROUND OF THE INVENTION

[0005] Genetically encoded fluorescent markers (GEFIs) are illuminating cellular energy metabolism. These fusion proteins comprise a bacterial ligand-binding protein fused to one or two fluorescent proteins. The binding of the molecule of interest to the bacterial domain induces a conformational change that alters the physical properties of the fluorescent protein(s). GEFIs can be expressed non-invasively in cells, targeted to different cell types or specific compartments to achieve high spatiotemporal resolution, or used as purified protein to detect molecules of medical interest in biological fluids. Selecting a suitable ligand-binding domain is critical for developing metabolite-targeted GEFIs. Classically, bacterial periplasmic proteins and transcription factors have been successfully used to develop metabolite-targeted GEFIs.Currently, only a very small fraction of the metabolic network can be explored with GEFIs, hindering the dynamic assessment of nodal points in cellular metabolism and, ultimately, the monitoring of the entire metabolic network and its temporal and spatial plasticity. The main reason for the scarcity of GEFIs for metabolites is the lack of recognition domains with known ligands, despite the extraordinary accumulation of biodiversity resulting from whole-genome sequencing and metagenomics. Thousands of DNA sequences encoding, for example, periplasmic proteins and transcription factors remain orphaned. Bacterial protein biodiversity is a fertile source of potential new scaffolds for developing GEFIs targeting emerging metabolites.

[0006] The development and sophistication of whole-genome sequencing systems and metagenomic platforms have led to the accumulation of an extraordinary number of sequences encoding unknown proteins, and therefore, proteins with no defined function. In the post-genomic era, one of the greatest challenges is understanding the meaning of the genetic information collected from genomic and metagenomic resources. Approximately 32% of the sequences submitted to UniProtKB, one of the largest protein databases, are labeled as unknown protein, and between 30% and 40% of functionally identified protein modules are reported as incorrectly annotated. The identification and functional annotation of bacterial proteins allows for the characterization of new enzymes and input / output modules useful for creating recombinant proteins with applications in biotechnology and basic science.Identifying the ligand of a given protein is a crucial step for its functional characterization.

[0007] Bioinformatics tools, in vitro effector / ligand interaction methods, and genetic disruption followed by functional screening have been used to identify effector / ligand molecules, but they offer low throughput and sensitivity. However, no GEFIs based on orphan bacterial proteins exist in the state of the art. Likewise, no functional protein annotation methods using steady-state fluorescence as the output readout, detected through standard fluorescence readers for a multiwell plate format, have been described. Nevertheless, related documents exist in the art, which will be described below according to the following references: REFERENCES

[0008] 1. Koskinen, P.; Toronen, P.; Nokso-Koivisto, J.; Holm, L., PANNZER: high-throughput functional annotation of uncharacterized proteins in an error-prone environment. Bioinformatics 2015, 31 (10), 1544-52.

[0009] 2. Schuller, A.; Slater, A. W.; Norambuena, T.; Cifuentes, J. J.; Almonacid, L. I.; Melo, F., Computer-based annotation of putative AraC / XylS-family transcription factors of known structure but unknown function. J Biomed Biotechnol 2012, 2012, 103132.

[0010] 3. de Lorimier, R. M.; Smith, J. J.; Dwyer, M. A.; Looger, L. L.; Sali, K. M.; Paavola,

[0011] C. D.; Rizk, S. S.; Sadigov, S.; Conrad, D. W.; Loew, L.; Hellinga, H. W., Construction of a fluorescent biosensor family. Protein Sci 2002, 11 (11), 2655-75.

[0012] 4. Dwyer, M. A.; Hellinga, H. W., Periplasmic binding proteins: a versatile superfamily for protein engineering. Curr Opin Struct Biol 2004, 14 (4), 495-504.

[0013] 5. Peroza, E. A.; Boumezbeur, A. H.; Zamboni, N., Rapid, randomized development of genetically encoded FRET sensors for small molecules. Analyst 2015, 140 (13), 4540-8.

[0014] 6. Vetting, M. W.; Al-Obaidi, N.; Zhao, S.; San Francisco, B.; Kim, J.; Wichelecki,

[0015] D. J.; Bouvier, J. T.; Solbiati, J. O.; Vu, H.; Zhang, X.; Rodionov, D. A.; Love, J. D.; Hillerich, B. S.; Seidel, R. D.; Quinn, R. J.; Osterman, A. L.; Cronan, J. E.; Jacobson, M. P.; Gerlt, J. A.; Almo, S. C., Experimental strategies for functional annotation and metabolism discovery: targeted screening of solute binding proteins and unbiased panning of metabolomes. Biochemistry 2015, 54 (3), 909-31.

[0016] 7. Chai, Y.; Kolter, R.; Losick, R., A widely conserved gene cluster required for lactate utilization in Bacillus subtilis and its involvement in biofilm formation. J Bacteriol 2009, 191 (8), 2423-30.

[0017] 8. Nadler, D.C.; Morgan, S.A.; Flamholz, A.; Kortright, K.E.; Savage, D.F., Rapid construction of metabolite biosensors using domain-insertion profiling. Nat Commun 2016, 7, 12266.

[0018] 9. Bailey, T.L.; Boden, M.; Buske, F.A.; Frith, M.; Grant, C.E.; Clementi, L.; Ren, J.; Li, W.W.; Noble, WS, MEME SUITE: tools for motif discovery and searching. Nucleic Acids Res 2009, 37 (Web Server issue), W202-8. 10. Bailey, T.L.; Elkan, C., Fitting a mixture model by expectation maximization to discover motifs in biopolymers. Proc Int Conf Intell Syst Mol Biol 1994, 2, 28-36.

[0019] 11. Grant, C.E.; Bailey, T.L.; Noble, WS, FIMO: scanning for occurrences of a given motif. Bioinformatics 2011, 27 (7), 1017-8.

[0020] The current state of the art to explore the protein-effector / ligand interaction includes a broad range of methodologies within bioinformatic analysis 1,2 , fluorescence with organic dyes3-5 , TSA (Thermal Shift Assay), MST (Microscale Thermophoresis), ITC (Isothermal Titration Calorimetry), SPR (Surface Plasmon Resonance), DPI (Dual Polarization Interferometry), DSF (Deferential Scanning Fluorometry) 6 and functional screening after genetic interruption 7 However, most are not suitable for high throughput, require expensive and / or dedicated equipment, and are costly to implement. Furthermore, fluorescence-based methods have been described for the rapid construction of high-throughput metabolite indicators, but these require ligand-binding domains with known ligands. 5,8To our knowledge, no technology is currently available that allows for the identification of an effector ligand for an orphan protein using steady-state fluorescence in a high-throughput format with standard multiwell plate readers. Furthermore, each new functional annotation of an orphan protein, e.g., bacterial periplasmic protein, histidine kinase, or transcription factor, opens the door to the development of a novel GEFI (Global Effector Inhibitor) with a high degree of novelty.

[0021] We present this methodology as a novel technique for developing GEFIs from orphan bacterial proteins. Its main strengths lie in the unbiased approach provided by a high-throughput, detectable readout using non-dedicated equipment, such as multiwell plate readers. Furthermore, the described invention is low-cost, as the protein extract can be produced in Escherichia coli expression systems and could even be carried out using intact bacterial cells or other biological systems.

[0022] BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The invention is illustrated in the accompanying drawings, where:

[0024] Figure 1 shows the basic scheme of the methodology for identifying effector / ligand of bacterial orphan proteins.

[0025] Figure 2 shows the bioinformatics algorithm used to identify a family of orphan bacterial proteins from genomic or metagenomic databases (e.g., the GntR transcription factor family).

[0026] Figure 3 shows a screening of amino acid-binding proteins; the heat map represents Z-score values ​​(light gray tones indicate low values ​​and dark gray tones high values). Red dots are hits with Z-score > 3 and CV% < 15.

[0027] Figure 4 shows a heat map where warm colors correspond to hits with Z-score values ​​> 3.

[0028] Figure 5 shows heat maps where warm colors correspond to hits with CV% < 15.

[0029] Figure 6, proof of concept showing the alignment of amino acid sequences of orphan GntR transcription factors fused to the FRET pair mTFP / Vcniis. The ligands identified using this methodology are: Asn_GAJ93016 (asparagine), Tyr_GAJ93204 (tyrosine), Asp_GAJ95440 (aspartate), and Trp_GAJ93613 (tryptophan).

[0030] Figure 7 shows structural data, expression patterns, and functional data of the single-fluorophore GEFI indicators based on the newly functionally annotated bacterial proteins for tyrosine, tryptophan, and asparagine, identified by the method.

[0031] DISCLOSURE OF THE INVENTION

[0032] The objective of the present invention is to provide a method for identifying effector molecules / ligands for orphan proteins, such as histidine kinases, bacterial transcription factors, and periplasmic proteins, using fluorescence, thereby enabling their functional annotation and the development of GEFIs (Global Effector Factors). The methodology involves identifying a bacterial protein from genomic or metagenomic resources using bioinformatics tools, flanking or inserting one or more fluorescent, bioluminescent, anisotropic, or phosphorescent proteins into the backbone of the identified protein, and exposing the recombinant chimera to a library of molecules. Binding of the test molecule to the recombinant protein will induce a measurable change in the fluorescent signal, allowing for the identification of the effector and its functional annotation.Since this methodology can be performed with purified protein or recombinant proteins targeting the periplasm in Escherichia coli, the assay can be carried out using standard 96- or 384-well plate readers, drastically reducing the assay cost. Subsequently, a unique fluorophore indicator is developed using the functionally annotated bacterial transcription factor by inserting circularly permuted fluorescent, bioluminescent, anisotropic, or phosphorescent proteins into hot spots defined in silico within the bacterial cytoskeleton.

[0033] BRIEF DESCRIPTION OF THE INVENTION

[0034] The present invention comprises an algorithm for identifying orphan protein families from genomic or metagenomic resources and a method for identifying their effector molecule using steady-state fluorescence as a readout of the protein-effector / ligand interaction. Therefore, conformational change, a universal mechanism of bacterial effector-protein docking, is employed to modulate the fluorescent signal and reveal the nature of the effector / ligand and its protein partner.

[0035] The bioinformatics algorithm for in silico searching comprises: 1) identification of open reading frames (ORFs) from genomic or metagenomic assemblies; 2) generation of amino acid sequence motifs or signatures from known members of a protein family using the MEME tool 9 11 and 3) search for reasons of exclusive occurrence using the FIMO tool 11and removal of repeated sequences by manual or semi-automatic curing.

[0036] Once family members are identified from genomic or metagenomic data using in silico DNA sequencing, all genes undergo gene synthesis or PCR amplification from genomic DNA. With the gene cloned into a plasmid vector and using standard recombinant DNA technology, all orphan proteins are flanked with a pair of fluorescent, bioluminescent, anisotropic, or phosphorescent proteins, or a single fluorescent, bioluminescent, anisotropic, or phosphorescent protein is inserted into the orphan protein backbone. The recombinant proteins are then expressed and purified from Escherichia coli or targeted to the periplasm for experiments in intact bacteria. Both the purified protein and the recombinant protein targeted to the periplasm are fused to fluorescent, bioluminescent, anisotropic, or phosphorescent proteins. They are then exposed to a library of molecules of interest.The protein-ligand interaction can be detected by the change in the fluorescent signal using a standard multiwell plate reader in high-throughput format.

[0037] According to the configuration of the method for the identification of effectors / ligands:

[0038] • First realization: use of the method for functional annotation of orphan bacterial proteins by identifying their effector / ligand, using steady-state fluorescence.

[0039] • Second realization: use of the method to develop fluorescent indicators for a specific effector molecule / ligand, based on a bacterial protein previously without a known function.

[0040] Both the preceding summary and the following detailed description provide examples for illustrative purposes only and should not be considered restrictive. Furthermore, additional features or variations may be offered. For example, certain embodiments may be directed to various combinations and sub-combinations of the features described in the detailed description.

[0041] DETAILED DESCRIPTION OF THE INVENTION

[0042] The following detailed description refers to the accompanying drawings. Although specific embodiments of the invention are described, further modifications, adaptations, and implementations are possible. For example, substitutions, additions, or modifications may be made to the illustrated elements, and the described methods may be modified by rearranging, substituting, or adding steps, without departing from the scope of the invention. Although the method is described in terms of "comprising" various elements or steps, the methodology may also "consist essentially of" or "include" such elements or steps, unless otherwise stated. Furthermore, the terms "a," "an," and "the" are intended to include plural alternatives, e.g., at least one, unless otherwise stated. The method comprises the following general steps:

[0043] 1. In silico identification of a putative family of bacterial proteins from genomic or metagenomic resources.

[0044] 2. Amplification or synthetic synthesis of the identified bacterial proteins.

[0045] 3. Fusion of fluorescent protein(s) to each identified bacterial protein.

[0046] 4. Expression and / or purification of recombinant proteins.

[0047] 5. Unbiased screening assay by exposing each recombinant protein to a library of small molecules.

[0048] 6. Detection of the reading using a standard multi-well plate reader.

[0049] 7. Use of positive hits to develop single fluorophore indicators by inserting circularly permuted fluorescent, bioluminescent, anisotropic or phosphorescent proteins into the corresponding bacterial protein backbone.

[0050] The bioinformatics algorithm for in silico search consists of: 1) identification of ORFs from genomic or metagenomic assembly data (processed with Python or another tool capable of handling a large number of sequences); 2) generation of motifs from known members of a protein family using MEME (which returns an amino acid motif or sequence signature based on the selected sequences); and 3) search for uniquely occurring motifs with FIMO (which returns sequences containing the motif with a p-value < 10). 9 ) and removal of repeated sequences by manual or semi-automatic curing.

[0051] Once the family members are identified, they are synthesized or amplified by PCR, cloned into a plasmid vector, and, using standard recombinant DNA techniques, flanked with a pair of fluorescent, bioluminescent, anisotropic, or phosphorescent proteins, or a single fluorescent protein is inserted. They are then expressed and purified in Escherichia coli or oriented to the periplasm for experiments in intact bacterial cells. All fused orphan proteins are exposed to an unbiased library of small molecules. The protein-ligand interaction is detected by the change in the fluorescent signal using a standard multiwell plate reader. To establish statistical criteria for successful identification, the Z-score can be calculated. A protein extract or intact cell exposed to a given effector molecule / ligand that shows a shift of 3 standard deviations from the mean and a CV < 15% is considered a successful identification.As a proof of concept, representatives of the GntR transcription factor family were identified in the Rhizobium rhizogenes genome using amino acid sequence signatures. Subsequently, orphan bacterial proteins were fused to two fluorescent reporters expressed in Escherichia coli, purified, and exposed to a small library of amino acids and related molecules. From a total of 85 identified orphan transcription factors belonging to the GntR family, after manual curation to remove repetitive sequences, a total of 65 orphan transcription factors were exposed to 28 molecules. This methodology allowed the identification of the effector / ligand for four R. rhizogenes orphan proteins: asparagine (GAJ93016), tyrosine (GAJ93204), aspartic acid (GAJ95440), and tryptophan (GAJ93613). Therefore, the method allowed the assignment of function to 4 orphan transcriptional factors.In each case, the chimeric proteins—a given bacterial protein fused to a fluorescent module—can be considered fluorescent reporters for their ligand. The current state of the art does not include genetically encoded fluorescent nanosensors for detecting and quantifying asparagine, tyrosine, aspartic acid, and tryptophan based on unknown proteins, making these unique and highly inventive.

[0052] This methodology can be used to explore and identify effector molecules / ligands from various bacterial species by exposing these proteins to libraries of biomedically relevant molecules. Since recombinant proteins rely on fluorescent readout and can be expressed in diverse biological systems, these methodologies are useful for developing a novel set of GEFIs that explore cellular energy metabolism with high spatial and temporal resolution and for producing purified GEFI preparations to detect biomedically important molecules in biological fluids.

[0053] EXAMPLES

[0054] Materials and methods

[0055] Standard reagents, such as amino acids and structurally related molecules, were acquired from USBiological, and buffers from ThermoFisher.

[0056] Workflow

[0057] A total of 53 ORFs corresponding to Rhizobium rhizogenes were obtained from the Bioproject database (Agrobacterium rhizogenes NBRC 13257), which contains 13,077 ORFs (open reading frames). These 53 selected ORFs are predicted to belong to the GntR family and contain the FCD motif. To complement the previous data and ensure the identification of all GntR family members from the same dataset, we used an algorithm based on the MEME suite (http: / / meme-suite.org / ). This consisted of the following steps:

[0058] 1. Identification of ORFs from genomic assembly data. Starting from the original 13,077 ORFs in the same database, the ORFs were filtered by size using standard Biopython tools; only amino acid sequences between 220 and 300 residues were used.

[0059] 2. Generation of motifs or sequence signatures from known members of the GntR family using the MEME suite. A motif signature for the GntR family was generated by analyzing 36 previously annotated representative genes. The motif was generated using the Motif Discovery tool of the MEME suite.

[0060] 3. Search for the exclusive occurrence of the motif using MEME. Using the FIMO tool of the MEME suite, all ORFs in the database containing the motif GntR were selected. The statistical threshold was set at 10. 9 for FIMO. Repeated sequences were removed.

[0061] 4. Search for FCD or UTRA motifs: From the preselected ORFs, sequences showing a putative FCD or UTRA motif were selected. The presence of motifs was determined using the NCBI conserved domain search tool (https: / / www.ncbi.nlm.nih.gov / Structure / cdd / wrpsb.cgi). A total of 40 ORFs were selected using this approach. Comparing the sequences of the initial 53 ORFs obtained from Bioproject annotations with the 40 ORFs predicted by the MEME method, we found that 26 ORFs were common to both groups. After removing duplicate sequences, 67 selected genes were obtained. Of these 67 genes, we were able to clone 65 using the USER method.

[0062] All data handling, such as filtering by size, removing duplicates, and general manipulation, was done with Microsoft Excel and Biopython 1.77.

[0063] Cloning of identified open read frames

[0064] This stage involved the construction of DNA sequences encoding putative reporters based on the FRET pair mTFP / Vcnus and 65 orphan transcription factors belonging to the GntR family of environmental bacteria such as Rhizobium rhizogenes NBRC 13257. The bacterial species were obtained from ATCC (American Type Culture Collection, USA) and NBRC (NITE Biological Resource Center, Japan).

[0065] To construct the DNA sequence of each reporter, it is necessary to fuse each identified orphan transcription factor with the FRET pair mTFP / Vcnus. The complete workflow consists of: i) primer design; ii) PCR amplification of GntR members from bacterial genomic DNA; iii) PCR cloning of the amplicon into a compatible expression vector; and iv) sequencing verification.

[0066] The USER method is a ligase-independent cloning technique based on uracil excision. This system requires: primers modified with int-DeoxyUridine, a USER-compatible expression vector, a cleaving enzyme, DNA polymerase capable of overcoming uracil stalling, and the USER™ enzyme mix. With this recombination system, expression clones can be obtained directly from the recombination between PCR amplicons and the USER-compatible expression vector, enabling the generation of putative reporter DNA in a single step, saving time and resources.

[0067] To perform USER cloning, a USER-compatible expression vector and an amplicon amplified with internally uracil-modified primers are required. To construct the expression vector, we selected the pRSETB vector as a scaffold. This vector contains ampicillin resistance and a T7 promoter, making it ideal for protein production and purification in bacteria using sepharose-nickel technology. Additionally, we fused a DNA sequence encoding mTFP and Venus downstream of the T7 promoter and in-frame with the HisTag peptide. To obtain a USER-compatible vector, it is necessary to clone a USER recombination cassette between the mTFP and Venus coding sequences. This allows us to directly clone the amplicon produced by PCR amplification of TF GntR genes. We used the Xhol restriction site to clone the USER cassette. Our final USER-compatible plasmid vector was named pUSERBact.

[0068] A total of 65 orphan TF genes were amplified from bacterial genomic DNA. The primers used to amplify the genes must meet USER recombination compatibility requirements. All primers are 27 nt long. For further details, see the following example:

[0069] Primer design example: ggcgtt / ideoxyU / atggccgcgacagcaaccct

[0070] To avoid errors during the primer design process, we programmed an Excel spreadsheet to semi-automatically generate the primer sequence for each TF gene. This reduced both the time required and the potential for errors. A total of 65 primer sets were designed to amplify members of the TF GntR family, plus 6 primers to amplify a control on each plate. We selected E. coli PdhR as our positive control on each plate.

[0071] Once the primers were designed, they were synthesized at IDT (Integrated DNA Technologies, USA) in 96-well plates. Two plates were obtained in total: three with forward primers and three with reverse primers, covering the 65 transcription factors and the positive controls. To establish the PCR conditions, we performed a polymerase chain reaction with different DNA polymerases. We obtained a sharp amplicon band with the Pfu Turbo Cx polymerase, which is consistent with its ability to overcome uracil stalling. After testing the polymerase chain reaction and the primers, we performed USER cloning using pUSERBact as the expression vector and the LldR gene amplified by PCR from a plasmid vector and genomic DNA. In our preparative experiments, all selected clones were positive for LldR gene cloning. Therefore, USER cloning exhibits high efficiency.

[0072] Next, we performed PCR amplifications using some of the genes from the 96-well plate to test amplification efficiency under less stringent conditions. We obtained specific and high-concentration PCR products for some genes, but in other cases, we obtained weak or nonexistent products. To improve PCR amplification, we decided to perform the PCR reactions with an enhancer. We used 5% DMSO in the individual reactions, and the differences in PCR performance were quantitatively greater. The next step was to amplify the genes using the primers in the 96-well plate format. On our first attempt, we obtained PCR amplification signals for 90% of the target genes, demonstrating the efficiency of our molecular biology workflow.

[0073] Purification of recombinant proteins

[0074] Once all the fusion proteins were constructed, each was expressed and purified from Escherichia coli. In summary, all recombinant chimeras were transformed into Escherichia coli BL21 (DE3). A single colony was inoculated into 100 ml of LB medium containing 100 mg / ml of ampicillin (without IPTG) and shaken in the dark for 2–3 days. Cells were collected by centrifugation at 5000 rpm (4 °C) for 10 min and disrupted by sonication (Hielscher Ultrasound Technology) in 5 ml of Tris-HCl buffer, pH 8.0. A cell-free extract was obtained by centrifugation at 10,000 rpm (4 °C) for 1 hour and filtration of the supernatant. Proteins were purified using a nickel resin (Novagen His-Bind) according to the manufacturer's recommendations. The purified proteins were quantified using the Biuret method and stored at -20 °C in 20% glycerol.

[0075] Fluorescence measurements

[0076] Nickel-purified proteins were resuspended at 100 nM in an intracellular buffer containing (in mM): 10 NaCl, 130 KCl, 1.25 MgSCh, and 10 HEPES (pH 7.0). Fluorescence was measured using a TECAN microplate reader: samples were excited at 430 nm, and the emission intensities of mTFP and Venus were recorded at 485 nm and 528 nm, respectively.

[0077] High-performance screening

[0078] The constructs were exposed to a library of different amino acids and structurally related molecules at a concentration of 2 mM. The FRET ratio change was measured in a 96-well plate reader.

[0079] To establish statistical criteria for identifying “hits,” the Z-score was calculated. The Z-score is a dimensionless parameter obtained by subtracting an individual raw value from the population mean and then dividing the difference by the population standard deviation. This parameter allows us to evaluate the number of standard deviations by which the specific data point differs from the population mean, in this case obtained from the small amino acid library. The Z-score was calculated using the formula: where p is the mean and i the standard deviation of the population. Results within 3 standard deviations above or below the mean were considered “hits” to be validated. In addition, the coefficient of variation (CV) was calculated to obtain information on the variability between replicated wells. We conducted a screening campaign evaluating one protein per plate. The selection criteria for hits were a Z-score > 3 and CV% < 15.

[0080] A fluorescence-based method for identifying the effector / ligand of orphan proteins. Bacterial proteins are fused to fluorescent protein(s) and exposed to a library of small molecules. Changes in the fluorescent signal induced by a specific ligand provide a rapid readout of the binding and allow for the functional identification of the unknown protein.

[0081] Identification of effector molecules / ligands for an orphan bacterial protein by using changes in the fluorescent signal as readout. Each hit can be used as a scaffold protein for the development of a GELI.

[0082] While certain embodiments of the invention have been described, others may exist. Furthermore, any step of the disclosed method may be modified by rearranging, inserting, or deleting it, without departing from the invention. Although the specification includes a detailed description and associated drawings, the scope of the invention is indicated in the following claims.

[0083] After applying the methods described above, a family of orphan bacterial proteins was identified from genomic or metagenomic databases (e.g., the GntR transcription factor family). A summary of the results obtained at each step is shown in Figure 2.

[0084] Using the screening technique of the invention, amino acid-binding proteins were obtained. The classification is shown in Figure 3, where the heat map represents Z-score values ​​(light gray tones indicate low values ​​and dark gray tones high values). Red dots represent matches with a Z-score > 3 and a CV% < 15. The obtained proteins were classified according to the level of matches, considering Z-score values ​​greater than 3 or less than 15, as shown in Figures 4 and 5. Once the sequences of the amino acid-binding proteins were obtained, they were aligned, yielding ligands identified by means of the invention. Specifically, the following were identified: Asn_GAJ93016 (asparagine) from SEQ ID NO: 1 and its amino acid sequence SEQ ID NO: 2, Tyr_GAJ93204 (tyrosine) from SEQ ID NO: 3 and amino acid sequence SEQ ID NO: 4, Trp_GAJ93613 (tryptophan) from SEQ ID NO: 5 and amino acid sequence SEQ ID NO: 6.

[0085] Finally, expression patterns and functional data were obtained for the single fluorophore GEI indicators, as shown in Figure 7.

Claims

CLAIMS 1) A high-throughput, unbiased method for identifying effectors / ligands of an orphan bacterial protein using steady-state fluorescence, CHARACTERIZED in that it comprises the steps: a) Identifying a putative family of bacterial proteins from genomic or metagenomic resources; b) Amplifying or synthesizing the identified bacterial proteins; c) Fusing one or more fluorescent reporter proteins to each identified bacterial protein; d) Expressing and / or purifying recombinant proteins in biological systems; e) Assaying an unbiased screening by exposing each recombinant protein to a library of small molecules; f) Detecting the steady-state fluorescence readout: bioluminescence, anisotropy, FLIM, or phosphorescence.; g) Use positive hits to develop single fluorophore indicators by inserting circularly permuted fluorescent, bioluminescent, anisotropic or phosphorescent proteins into the corresponding bacterial protein backbone. 2) Genetically encoded single fluorophore indicator (GEFI), CHARACTERIZED in that it corresponds to a fusion protein between an orphan protein, identified by the method of claim 1, and a fluorescent reporter protein. 3) Genetically encoded single fluorophore indicator (GEFI) according to claim 2, CHARACTERIZED in that for asparagine based on the orphan bacterial transcription factor GAJ93016 and encoded by a nucleic acid sequence with at least 50%, 60%, 70%, 80%, 85%, 90%, 95% or 99% identity with SEQ ID NO 1. 4) Genetically encoded single fluorophore indicator (GEFI) according to claim 2, CHARACTERIZED in that for asparagine based factor orphan bacterial transcription factor GAJ93016 with at least 50%, 60%, 70%, 80% 85%, 90%, 95% or 99% amino acid sequence identity with SEQ ID NO 2. 5) Genetically encoded single fluorophore indicator (GEFI) according to claim 2, CHARACTERIZED in that for tyrosine based on the orphan bacterial transcription factor GAJ93204 and encoded by a nucleic acid sequence with at least 50%, 60%, 70%, 80%, 85%, 90%, 95% or 99% identity with SEQ ID NO 3. 6) Genetically encoded single fluorophore indicator (GEFI) according to claim 2, CHARACTERIZED in that for tyrosine based orphan bacterial transcription factor GAJ93204 with at least 50%, 60%, 70%, 80%, 85%, 90%, 95% or 99% amino acid sequence identity with SEQ ID NO 4. 7) Genetically encoded single fluorophore indicator (GEFI) according to claim 2, CHARACTERIZED in that for tryptophan based on the orphan bacterial transcription factor GAJ93613 and encoded by a nucleic acid sequence with at least 50%, 60%, 70%, 80%, 85%, 90%, 95% or 99% identity with SEQ ID NO 5. 8) Genetically encoded single fluorophore indicator (GEFI) according to claim 2, CHARACTERIZED in that for tryptophan based orphan bacterial transcription factor GAJ93613 with at least 50%, 60%, 70%, 80%, 85%, 90%, 95% or 99% amino acid sequence identity with SEQ ID NO 6. 9) Use of a fusion protein between an orphan protein and a fluorescent reporter protein, CHARACTERIZED in that it serves as a novel scaffold for developing genetically encoded fluorescent indicators (GEFIs) intended to detect molecules of biomedical relevance at the single-cell level and / or in biological fluids.