Modified olfactory receptors
Modified olfactory receptors with a specific C-terminal domain improve functional expression and sensitivity in assays, addressing inefficiencies in current screening methods by enabling comprehensive ligand identification and uniform receptor expression.
Patent Information
- Application Number
- JP2025534935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-25
- Filing Date
- 2023-12-15
- Publication Date
- 2025-12-24
AI Technical Summary
Current methods for expressing olfactory receptors in host cells are inefficient, leading to low sensitivity in ligand screening assays, particularly for poorly soluble and cytotoxic ligands, and result in incomplete coverage of the receptor-ligand space, with many receptors having unidentified ligands.
Modified olfactory receptor proteins with a specific C-terminal domain, combined with nucleic acid molecules and expression vectors, enable improved functional expression and sensitivity in assays, allowing for the identification of novel ligands, enhancers, and antagonists, even at physiologically relevant concentrations.
The modified receptors enhance assay sensitivity, enabling the identification of a broader range of ligands, including poorly soluble and cytotoxic ones, and provide uniform expression levels across the receptor library, improving the accuracy and throughput of screening processes.
Smart Images

Figure 2025542013000057 
Figure 2025542013000058 
Figure 2025542013000059
Abstract
Description
[Technical Field]
[0001] Field Aspects and embodiments described herein relate to the fields of biotechnology and flavors and fragrances, and in particular to modified olfactory receptors, and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses for expressing olfactory receptors and for identifying novel olfactory receptors and novel olfactory receptor ligands, enhancers, and antagonists. [Background technology]
[0002] background Olfactory receptors, or odorant receptors (ORs), are expressed in olfactory neurons in the olfactory epithelium and are responsible for detecting odorants. Olfactory receptors belong to the G protein-coupled receptor superfamily (GPCR). Activation of ORs by odorants (ligands) activates olfactory-specific G proteins, which then promote the production of cyclic AMP (cAMP) via type III adenylate cyclase. Increased intracellular cAMP levels lead to the opening of cyclic nucleotide-gated ion channels, which allow calcium ions to enter the cells, depolarize olfactory neurons, and trigger action potentials that carry information to the glomeruli in the olfactory bulb.
[0003] The human genome encodes approximately 400 different functional olfactory receptors. A particular olfactory receptor may be activated by more than one ligand molecule, and a particular ligand molecule may activate multiple olfactory receptors. This creates a highly complex interaction network between ORs and the repertoire of ligands. Elucidating these interactions may enable the discovery of compounds such as novel flavor and fragrance ingredients, or odor enhancers that are more sustainable and / or easier to produce than currently used compounds. For many of the approximately 400 different olfactory receptor genes, different variants, called alleles or haplotypes, exist in the human population. The protein products of different alleles or haplotypes of the same gene may have different ligand selectivity or sensitivity.
[0004] In addition to their expression in olfactory neurons of the olfactory epithelium, OR expression has also been found in many other cells and tissues (Feldmesser et al. Widespread ectopic expression of olfactory receptor genes. BMC Genomics 20067: p. 121; Massberg, D. and H. Hatt. Human olfactory receptors: novel cellular functions outside of the nose. Physiol Rev 2018;98(3):1739-1763). These ORs have been found to be associated with many important diseases. Therefore, in addition to their primary interest as targets for odorant ligands, ORs are also of great interest as targets for discovering agonists and antagonists to treat various diseases (Lee et al. Therapeutic potential of ectopic olfactory and taste receptors. Nature Reviews Drug Discovery 2019;18(2):116-138). In particular, OR signaling has been associated with decreased or increased cell proliferation in bladder, colon, prostate, lung, and liver cancer cells. Similarly, expression of ORs in inflammatory cells has been associated with the regulation of inflammation. Thus, screening for ORs has multiple applications beyond the scope of olfaction.
[0005] Efficient screening of olfactory receptors requires their expression in cultured cell lines, which generally involves introducing olfactory receptor genes into cells and then stably or transiently overexpressing them. In general, functional expression of olfactory receptors in host cells has proven to be very difficult, because it has proven difficult to obtain correct folding of the receptor and / or correct insertion of the receptor into the cell membrane. Therefore, several approaches have been attempted to improve the functional heterologous expression of olfactory receptors.
[0006] Expression of functional heterologous olfactory receptors using currently available expression systems generally requires coexpression of accessory proteins from the receptor-transporting protein (RTP) family, such as RTP1S and RTP2, which are typically expressed in olfactory neurons and facilitate OR trafficking to the cell surface membrane (Yu et al. Receptor-transporting protein (RTP) family members play divergent roles in the functional expression of odorant receptors. PLoS One 2017;12(6):e0179067). Furthermore, it has been found that fusing olfactory receptor genes with a sequence encoding the N-terminal sequence of rhodopsin (initial methionine and the following 19 amino acids) (the rho tag) enhances olfactory receptor expression (Krautwurst et al. Identification of ligands for olfactory receptors by functional expression of a receptor library. Cell. 1998;95(7):917-26).
[0007] However, even with RTP proteins and N-terminal rho tags, more than half of the known olfactory receptors cannot be functionally expressed using currently available nucleic acid constructs, cell lines, and methods. As a result, the coverage of the currently available receptor-ligand space is limited, and there are multiple receptors with unidentified ligands (orphan receptors), limiting the industrial application of the method. Even for receptors expressed using the method, expression levels can be low, resulting in low sensitivity of screening assays. In fact, many receptors functionally expressed in current expression systems are only strongly activated at relatively high ligand concentrations, e.g., 10-300 μM, which is a concentration at which many ligands already evoke a sensory experience at much lower concentrations in vivo. Thus, receptors expressed in current systems are often not activated at physiologically relevant concentrations. This indicates that current systems are often not sensitive enough to mimic the in vivo situation. This also poses practical problems. This is because ligands that are cytotoxic or poorly soluble in cell culture medium (weak) are unable to cause receptor activation in current screening cell lines, either because they are not sufficiently dissolved (most odorants are non-polar molecules with limited solubility in water) or because their cytotoxicity at the high test concentrations applied directly leads to inactivation of the cell line.
[0008] Classical OR screening assays rely on an approach in which clonal populations of cells are transfected with DNA expression constructs encoding specific receptors and / or accessory molecules, typically one at a time, and then tested for functional activation by various ligands. The assays also typically involve coexpression of a luciferase gene operably linked to a cAMP-inducible promoter, used as a reporter gene (Saito et al. RTP family members induce functional expression of mammalian odorant receptors. Cell 2004;119(5):679-691). Activation of olfactory receptors and the subsequent increase in intracellular cAMP result in the expression of luciferase. Luciferase-catalyzed oxidation of luciferin results in luminescence, which can then be detected and quantified. Classical OR screening assays are limited in their sensitivity, can result in widely varying expression levels for different receptors, and are often not compatible with high-throughput screening and selection methods, such as screening libraries of volatile flavor and fragrance compounds, which contain ligands with moderate activity, cytotoxic properties, and limited solubility in cell culture media. Summary of the Invention
[0009] overview In view of all of the above, there is a need for improved olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses for expressing olfactory receptors and identifying novel olfactory receptor ligands, enhancers, and antagonists. More specifically, such improved olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses should enhance assay sensitivity, allowing testing at low concentrations, thereby enabling the full ligand spectrum of ORs to be described, even for poorly soluble and cytotoxic ligands. Such improved sensitivity should also enable the identification of new ligand-OR pairs to deorphanize receptors and better describe the full receptor space (not only additional ligands, but also antagonists and enhancers) of already deorphanized receptors.
[0010] In addition, for large-scale deorphanization of olfactory receptors (i.e., for the identification of ligands for all receptors with previously unknown ligands), and to find all active ORs, and especially the ORs most sensitive to a given ligand of interest, it is necessary to express the entire library of many / all human olfactory receptors. To find the truly most important receptor for a given ligand, all receptors should be expressed at similar levels, preferably at least as fully expressed as possible. Otherwise, false positives will be observed, whereby a highly expressed receptor appears to be the most sensitive receptor for the ligand of interest, and the truly most sensitive receptor will be missed due to its lower functional expression. Therefore, for such screening operations involving multiple receptors, it is desirable to normalize the functional expression of different receptors and minimize expression differences between receptors. Therefore, a library of human ORs optimized to ensure that the functional expression of all receptors to be used in OR expression assays is comparable is needed.
[0011] The present invention provides olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses that are particularly useful in the context of ORs that are difficult to express using conventional approaches, or that result in assays with limited sensitivity using conventional approaches, and in identifying novel cognate receptor-ligand pairs. The present invention also provides olfactory receptor variants with improved functional expression, which enable more sensitive assays for detecting OR-ligand interactions. The olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses described herein exhibit at least one of the following advantages over the prior art: - Improved functional heterologous expression - Enables deorphanization of receptors with no known ligands - Increase assay sensitivity to determine the complete ligand spectrum of a receptor - Increase assay sensitivity to measure ligand binding at more physiologically relevant concentrations - Increasing the sensitivity of the assay to identify antagonists and enhancers - Enables functional expression of olfactory receptors that was not possible using conventional approaches - Allows identification of novel cognate receptor-ligand pairs - enhancing the sensitivity of assays using olfactory receptors (this can be measured by a lower ligand concentration to achieve similar activity, or a lower EC50 value; i.e., increased potency, a significantly lower detection threshold, or increased efficacy). - Ability to test more cytotoxic molecules - Ability to test poorly soluble molecules - Ability to detect ligands in complex test mixtures and unpurified synthetic samples - Ability to screen for ligands with specific odor classes in complex test mixtures and crude synthetic samples - the ability to identify the more olfactory active isomer or enantiomer in racemic mixtures and biodegradable compounds and compositions; - Allows for the creation of OR libraries with better and more uniform functional expression compared to using wild-type OR genes - Increased throughput capacity for olfactory receptor and ligand screening
[0012] As demonstrated in the experimental section herein, the application of the olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses described herein are associated with some of the above advantages, and thus provide highly significant improvements compared to conventional approaches.Therefore, the aspects and embodiments of the present invention described herein solve at least some of the problems and needs discussed herein.
[0013] One aspect of the present invention relates to an olfactory receptor protein, wherein the protein has a modified C-terminal domain comprising the following amino acid sequence motif: RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1). In some embodiments, the olfactory receptor protein of the present invention is such that the modified C-terminal domain is fused to the seventh transmembrane helix (TM7) of the protein. In some embodiments, the olfactory receptor protein of the present invention is such that the protein is a class I or class II olfactory receptor with a modified C-terminal domain, preferably a human, canine, or feline class I or class II olfactory receptor with a modified C-terminal domain, more preferably a human class I or class II olfactory receptor with a modified C-terminal domain. In some embodiments, the olfactory receptor protein of the present invention is a class II receptor, selected from the group consisting of OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A 25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2, preferably wherein the class II receptor is selected from the group consisting of OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2, and OR5A2. In some embodiments, the olfactory receptor protein of the present invention is such that the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1. In some embodiments, the olfactory receptor protein of the present invention is such that the sequence motif is: RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 5).
[0014] In some embodiments, the olfactory receptor proteins of the present invention are such that x is not proline. In some embodiments, the olfactory receptor proteins of the present invention are such that x is not proline or tryptophan. In some embodiments, the olfactory receptor proteins of the present invention are such that x is selected from D, K, R, E, N, V, A, Q, or G, preferably where x is selected from D, K, R, E, N, V, A, or Q.
[0015] In some embodiments, the olfactory receptor protein of the present invention is such that the amino acid sequence motif consists of 1 to 6 additional C-terminal amino acid residues, optionally wherein: - the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, more preferably the first additional amino acid residue is C, R or K, most preferably C; - the second additional amino acid residue is selected from C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably C or R; - the third additional amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K; the fourth, fifth and sixth amino acid residues are selected from K and R;
[0016] In some embodiments, the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164), and CRRRKK (SEQ ID NO: 165).
[0017] In some embodiments, the olfactory receptor proteins of the present invention are such that the sequence motif is selected from the group consisting of SEQ ID NOs: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, and 820-890.
[0018] In some embodiments, the olfactory receptor protein of the present invention further comprises an N-terminal tag peptide, and preferably, the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag, and an M3 tag.
[0019] Another aspect of the present invention relates to a nucleic acid molecule comprising a nucleotide sequence encoding the olfactory receptor protein of the present invention. In some embodiments, the nucleic acid molecule of the present invention further comprises a promoter sequence, preferably a constitutive promoter sequence. In some embodiments, the nucleic acid molecule of the present invention further comprises a terminator sequence. In some embodiments, the nucleic acid molecule of the present invention further comprises a nucleotide sequence encoding an N-terminal signal peptide, preferably a leucine-rich signal peptide such as MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77).
[0020] Another aspect of the present invention relates to an expression vector comprising the nucleic acid molecule of the present invention. In some embodiments, the expression vector of the present invention is a plasmid.
[0021] Another aspect of the present invention relates to a recombinant host cell comprising a nucleic acid molecule of the present invention or an expression vector of the present invention, preferably wherein the cell expresses an olfactory receptor protein of the present invention. In some embodiments, the recombinant host cell of the present invention is such that the cell further expresses one or more olfactory receptor accessory proteins. In some embodiments, the one or more olfactory receptor accessory proteins are RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf , Giα, and functional variants thereof, and preferably selected from the group consisting of RTP1S, RTP2, and functional variants thereof. In some embodiments, the recombinant host cell of the present invention is such that the cell is a HEK293 cell or a HEK293T cell.
[0022] Another aspect of the present invention relates to a library comprising a diverse repertoire of olfactory receptor proteins of the present invention, nucleic acid molecules of the present invention, expression vectors of the present invention, or recombinant host cells of the present invention. In some embodiments, the library of the present invention is such that the diverse repertoire of olfactory receptor proteins, the diverse repertoire of olfactory receptor proteins encoded by nucleic acid molecules or expression vectors, or the diverse repertoire of olfactory receptor proteins expressed by recombinant host cells share the same modified C-terminal domain.
[0023] Another aspect of the present invention relates to the use of the olfactory receptor protein of the present invention, the nucleic acid molecule of the present invention, the expression vector of the present invention, the recombinant host cell of the present invention, or the library of the present invention to identify a ligand, enhancer, or antagonist of an olfactory receptor.
[0024] Another aspect of the present invention relates to the use of the library of the present invention to identify olfactory receptors capable of binding to a target ligand.
[0025] Another aspect of the present invention relates to a method for identifying an olfactory receptor ligand, said method comprising: a) providing an olfactory receptor protein of the present invention or a recombinant host cell expressing an olfactory receptor protein of the present invention; b) contacting the receptor or recombinant host cell with a test compound or composition; and c) Detecting activation of olfactory receptors.
[0026] Another aspect of the present invention relates to a method for identifying an enhancer or antagonist of an olfactory receptor, said method comprising: a) providing an olfactory receptor protein of the present invention or a cell expressing an olfactory receptor protein of the present invention; b) contacting the receptor or recombinant host cell with a cognate ligand and a test compound or composition; and c) Detecting increased or decreased activation of olfactory receptors compared to a ligand-only control.
[0027] In some embodiments of the method for identifying an olfactory receptor ligand and the method for identifying an enhancer or antagonist of an olfactory receptor, the olfactory receptor is selected from the group consisting of OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)), OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25 (OR2A25(S75N, A209P)), OR11G2 (or OR11G2(I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3(S73G)), OR10G9, OR2L5, OR8H1, OR OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2, and OR2AG2 (preferably OR2AG2(Y28C)).
[0028] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1, and OR4S2, preferably OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and OR5V1, more preferably OR2M2 or OR2V1.
[0029] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, wherein the olfactory receptor is OR2M2 or OR2V1, and step b) further comprises contacting the receptor or recombinant host cell with a copper salt. In some embodiments of the method for identifying an OR2M2 or OR2V1 antagonist, the cognate ligand is selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, and 4-methoxy-2-methylpentane-2-thiol.
[0030] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR51B2 (preferably OR51B2 (C120R, L134F, C209S)). In some embodiments of the method for identifying an OR51B2 antagonist, the cognate ligand is 3-methyl-2-hexenoic acid.
[0031] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR5V1. In some embodiments of the method for identifying an OR5V1 antagonist, the cognate ligand is 2,4,6-trichloroanisole.
[0032] Another aspect of the present invention relates to a method for identifying an olfactory receptor capable of binding to a target ligand, said method comprising: a) providing a library of the present invention; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins with a target ligand; and d) Identifying the olfactory receptors activated by target ligands.
[0033] Another aspect of the present invention relates to a method for generating an objective representation of the olfactory properties of a test compound or composition, said method comprising: a) providing a library of the present invention; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a test compound or composition; and d) Detecting the activation of each of the olfactory receptor proteins.
[0034] Another aspect of the invention relates to a method for assessing the differences or similarities between two or more test compounds or compositions, said method comprising: a) providing a library of the present invention; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of two or more test compounds or compositions; d) detecting activation of each of the olfactory receptor proteins for each of two or more test compounds or compositions; and e) Comparing the activated olfactory receptor protein between each of two or more test compounds or compositions.
[0035] explanation Various features of the aspects and embodiments of the present disclosure are further described below, with the understanding that the headings used throughout this specification are merely for navigational aids and should not be construed as definitive, and that features described in different sections may relate to all aspects and embodiments described herein and may therefore be combined as appropriate.
[0036] Introduction Next to modifying the N-terminus of olfactory receptors, for example by adding a rho tag, other changes have also been made to the sequences of olfactory receptors in an attempt to improve their expression.
[0037] In one attempt, consensus sequences were derived from each of the different classes of olfactory receptors. The resulting consensus sequences were found to produce highly expressed receptors in six of the nine cases, generally better expressed than individual receptors in that class (Ikegami et al. Structural instability and divergence from conserved residues underlying intracellular retention of mammalian odorant receptors. PNAS 2020;117(6):2957-2967). This consensus approach corresponded to the full-length receptor sequence. This means that the resulting consensus receptor does not correspond to the ligand affinity of a single native receptor, but rather is a completely synthetic receptor with a novel ligand spectrum, making it unsuitable for screening commercially interesting odorants perceived by the human nose.
[0038] There have also been reports of modifying the C-terminal sequence of OR genes. As shown below, in most cases, changing or truncating the C-terminal sequence reduced activity, and only in a few cases did it result in higher signal amplitude. To date, altering the C-terminus of ORs has not resulted in assays with significantly higher sensitivity, i.e., assays that allow testing at significantly lower test concentrations.
[0039] Kotthoff et al. found that truncating human OR8D1 by three amino acids reduced signal amplitude, and truncating it by seven, 11, or 15 amino acids eliminated the signal; however, they did not find any enhancement of the function of the truncated C-terminus (Kotthoff et al. The FASEB Journal 2021;35:e21274). Altering the C-terminus had a negative effect on OR8D1. However, altering the parent C-terminal sequence toward the consensus sequence had a positive effect in two cases: thus, altering the C-terminus of human OR2M3 toward the consensus sequence by three amino acids enhanced signal amplitude threefold, but did not significantly enhance sensitivity / potency. That is, it did not shift the dose-response curve to lower concentrations. Altering mouse olfr16 toward the consensus doubled the signal amplitude, but only modestly enhanced sensitivity (EC50 from 33.99 to 20.42 micromolar, Table S19 in the reference). All other changes introduced at the C-terminus (>70 investigated variants) invariably resulted in decreased activity and / or reduced surface expression, indicating that altering the C-terminus is not sufficient to significantly improve receptor expression and assay sensitivity.
[0040] In a detailed analysis of two closely related mouse receptors (mouse Olfr539 and Olfr541) (Ikegami et al. 2020), Olfr541 was poorly expressed, whereas Olfr539 was highly expressed. Strong expression was achieved by exchanging a portion of the central domain of Olfr541, or even a single amino acid, with the Olfr539 sequence. However, replacing the C-terminus of the poorly expressed Olfr541 with the C-terminus of the highly expressed Olfr539 did not enhance Olfr541 expression. This suggests that other portions of the sequence, other than the C-terminus, are important for conferring functional expression.
[0041] In another example, the sequence of human OR5A2 was modified at both the C-terminus and N-terminus by exchanging the OR5A2 sequence with the OR2A5 sequence to generate a chimeric receptor (Example 5 in WO2019110630A1). The resulting receptor responded to musk compounds as well as the wild-type receptor, and the modified C-terminus did not alter the dose-response curve or enhance sensitivity.
[0042] Furthermore, in a frequently cited study (Krautwurst et al. Identification of ligands for olfactory receptors by functional expression of a receptor library. Cell. 1998;95(7):917-26), Krautwurst et al. generated chimeric receptors using the N- and C-terminal sequences of the well-expressed mouse receptor M4 and inserting sequences from TM2 to TM7 of other tested receptors into this framework. Chimeric receptors with this M4 framework responded comparably to wild-type receptors, but the chimeric receptors did not demonstrate enhanced signal or better functionality. In one example, a response was reported at even lower concentrations (1 μM instead of 10 μM) for the wild-type mouse I7 receptor compared to the chimeric receptor, indicating that activity can be largely maintained, but not increased, by exchanging the C-terminus with that of a well-expressed receptor.
[0043] Katada et al. found that truncating mouse mOR-EG by 3 amino acids reduced the signal amplitude, truncating it by 6 amino acids reduced the signal amplitude and sensitivity, and truncating it by 12 amino acids eliminated the signal; however, they did not find any enhanced function of the truncated C-terminus, indicating that the full-length C-terminus is required in the cases investigated (Kadata et al. Structural determinants for membrane trafficking and G protein selectivity of a mouse olfactory receptor. Journal of Neurochemistry 2004;90(6):1453-1463).
[0044] Kato et al. found that mutating K296, K299, K303, K304, and K309 at the C-terminus of mouse mOR-EG to proline reduced or eliminated activity, whereas mutation of these residues to Ala or Arg did not significantly affect activity (except for K296A, which reduced sensitivity). However, this study did not find any enhancement of function (enhanced amplitude or sensitivity) of the mutated C-terminal sequence, indicating that most of these basic residues are not essential for functional expression (Kato et al. Amino acids involved in conformational dynamics and G protein coupling of an odorant receptor: targeting gain-of-function mutation. Journal of Neurochemistry 2008;107(5):1261-1270).
[0045] Finally, Sato et al. investigated the C-terminal sequences of olfactory receptors and which residues are important. They found that residues 299, 300, 303, and 304 at the C-terminus of mOR-S6 are important for maintaining function, but no enhanced function was found due to C-terminal modifications (Sato et al. Functional Role of the C-Terminal Amphipathic Helix 8 of Olfactory Receptors and Other G Protein-Coupled Receptors. Int. J. Mol. Sci. 2016;17(11):1930).
[0046] In summary, although it has previously been known that altering the N-terminus by adding a rho tag improves olfactory receptor expression, attempts to alter the C-terminus did not result in significant improvements in the assay, primarily demonstrating that the C-terminus is sensitive to changes in structure, which in most cases results in loss of function, not gain of function. Based on these teachings, it was not expected that altering the C-terminal sequence would result in significant advances in the functional expression of olfactory receptors. Rather, this detailed analysis showed that truncating the C-terminus of the receptor or altering the C-terminus toward a consensus sequence could not significantly enhance the sensitivity of the assay, and while signal amplitude could be enhanced in a few selected cases, the majority of alterations to the C-terminus had negative effects.
[0047] olfactory receptor proteins The present disclosure generally relates to olfactory receptor proteins (ORs) having modified C-terminal domains. Accordingly, provided herein are olfactory receptor proteins, wherein the proteins have modified C-terminal domains. Preferably, the olfactory receptor proteins are mammalian, more preferably human olfactory receptor proteins. Preferably, the olfactory receptor proteins correspond to class I or class II ORs, as described herein below.
[0048] The term "olfactory receptor" or "odorant receptor" (OR), as used herein, has its conventional meaning as commonly understood by those skilled in the art in light of the present disclosure. It refers to a receptor belonging to the seven-transmembrane-domain G protein-coupled receptor superfamily (GPCR), which is typically expressed in the plasma membrane of olfactory receptor neurons. Seven predicted transmembrane (TM) domains, TM I to TM VII, are connected by three predicted internal (IC) loop domains (IC I to IC III) and three predicted external (EC) loop domains (EC I to EC III). ORs typically contain olfactory receptor-specific amino acid motifs. Examples of such motifs include the MAYDRYVAIC (SEQ ID NO:2) motif that overlaps TM III and IC II, the FSTCSSH (SEQ ID NO:3) motif that overlaps IC III and TM VI, the PMLNPFIY (SEQ ID NO:4) motif in TM VII, and the presence of three conserved C residues in EC II and a highly conserved GN residue in TM I, as discussed in Zhang and Firestein (2002) Nature Neurosci 5(2): 124-33 and Malnic et al. (2004) PNAS 101(8):2584-9, both of which are incorporated herein by reference.
[0049] The C-terminal domain of an olfactory receptor begins immediately after the end of the seventh TM helix (TM7). Those skilled in the art can accurately determine this location based on their common general knowledge. For example, the seventh transmembrane region (TM7) can be easily recognized in any OR based on sequence alignment or by using well-known publicly available databases. For example, the "GPCR Prediction Ensemble Database (GPCR-PEnDB)" available at https: / / gpcr.utep.edu / database shows the amino acid position and native C-terminus length of TM7 for many ORs.
[0050] The same information can also be derived from the "HORDE" (The Human Olfactory Data Explorer) database maintained by the Weizmann Institute of Science and available at https: / / genome.weizmann.ac.il / horde / , which is described in Olender et al. (2013) Methods Mol Biol 1003:23-38, incorporated herein by reference. In HORDE, all residues TM1-TM7 are annotated.
[0051] The same information can also be found in common sequence databases such as UniProt (The UniProt Consortium, UniProt: the universal protein knowledgebase in 2021, Nucleic Acids Research, Vol. 49, No. D1, January 8, 2021, pp. D480-D489; available at www.uniprot.org).
[0052] TM7 in class II olfactory receptors typically ends with the consensus sequence NPLIYSL (SEQ ID NO: 225), and the last of these seven amino acids is usually located at residues 292-298 of the full-length receptor, so the end of TM7 can be easily identified for a given OR. For improved functional expression, the modified C-terminus is fused to the olfactory receptor immediately after this sequence (or at the position where the native C-terminus begins, as shown at https: / / gpcr.utep.edu / database).
[0053] TM7 in class I olfactory receptors typically ends with the consensus sequence NPIIYSL (SEQ ID NO: 226) or NPIIYSGL (SEQ ID NO: 227), with the last of these seven amino acids usually located at residue 297 (290-300). For improved functional expression, the modified C-terminus is fused to the olfactory receptor immediately after this sequence (or at the position where the native C-terminus begins, as shown at https: / / gpcr.utep.edu / database).
[0054] Therefore, in some embodiments, the C-terminal domain as described herein is configured to start immediately after the last residue of the seventh transmembrane helix (TM7).In some embodiments, the C-terminal domain is fused to the seventh transmembrane helix of olfactory receptor protein.In some embodiments, the C-terminal domain as described herein can also be referred to as cytoplasmic domain or intracellular domain.
[0055] Thus, in some embodiments, the olfactory receptors described herein comprise the amino acid sequence of an olfactory receptor up to and including the last residue of the seventh transmembrane helix (TM7) of the olfactory receptor, followed by the amino acid sequence of a modified C-terminal domain comprising an amino acid sequence motif as described herein.
[0056] The modified C-terminal domain comprises or consists of 10 to 40 amino acids, preferably 12 to 30 amino acids, more preferably 14 to 26 amino acids, and even more preferably 16 to 22 amino acids. In some embodiments, the minimum length of the modified C-terminal domain of an olfactory receptor of the present disclosure is 10, 11, 12, 13, 14, 15, or 16 amino acids; and the maximum length of the modified C-terminal domain of an olfactory receptor of the present disclosure is 26, 25, 24, 23, or 22 amino acids.
[0057] The present inventors have found that the modified C-terminal domain rich in basic amino acids exhibits the particularly advantageous effects described herein.As used herein, basic amino acids are understood to refer to amino acids with side chains that are protonated at neutral pH.Basic amino acids include lysine (Lys, K), arginine (Arg, R) and histidine (His, H).In the context of the present disclosure, preferred basic amino acids are lysine and arginine.
[0058] Therefore, in some embodiments, the modified C-terminal domain described herein is a modified C-terminal domain that comprises at least 5 basic residues, preferably at least 6 basic residues, more preferably at least 7 basic residues.In some embodiments, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, or at least 43% of amino acids are basic amino acids (i.e., lysine, arginine, or histidine).In a preferred embodiment, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 32% of amino acids are basic amino acids (i.e., lysine, arginine, or histidine). In a preferred embodiment, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 35% of the amino acids are basic amino acids (i.e., lysine, arginine, or histidine). In a preferred embodiment, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 38% of the amino acids are basic amino acids (i.e., lysine, arginine, or histidine).
[0059] In one aspect, an olfactory receptor protein is provided, wherein the protein has the amino acid sequence motif: RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 1) It has a modified C-terminal domain comprising:
[0060] A "sequence motif" may also be referred to as a "sequence pattern" or simply as a "sequence." A "sequence motif" or "sequence pattern" has its conventional and ordinary meaning as understood by those skilled in the art in light of the present disclosure. It refers to an amino acid (or nucleotide) sequence that occurs repeatedly, with some variation, in multiple sites in a molecule or in several different molecules, and that has (or is suspected to have, or is suspected to be associated with) a biological significance or exhibits a biological activity as described herein. The biological significance or biological activity of the sequence motif described herein is preferably its ability to improve the functional heterologous expression of (a nucleotide sequence encoding) an olfactory receptor protein that contains the motif as its C-terminal domain.
[0061] As used herein, the term "expression" or "heterologous expression" of a DNA molecule by a cell encompasses any step involved in the production of a polypeptide by a cell, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, transport to the cell membrane, and secretion. Expression may be assessed by any method known to those of skill in the art. For example, expression may be assessed by measuring the level of gene expression in transduced cells at the mRNA or protein level by standard assays known to those of skill in the art, such as qPCR, RNA sequencing, Northern blot analysis, Western blot analysis, mass spectrometry of protein-derived peptides, fluorescence-activated cell sorting (FACS), immunostaining, or ELISA.
[0062] As used herein, the term "functional expression" or "functional heterologous expression" refers to the production of a polypeptide by a cell, wherein the polypeptide exhibits biological activity. For example, an olfactory receptor is functionally expressed in a cell if, following its production, it is transported to and integrated into the cell membrane, and can trigger its corresponding signal transduction cascade following its activation by a ligand. A conventional method for evaluating the functional expression of an olfactory receptor involves the expression of the OR with the coexpression of a luciferase gene operably linked to a cAMP-inducible promoter used as a reporter gene (Saito et al. (2004) Cell 119(5): 679-691). When an olfactory receptor is functionally expressed, its activation and the resulting increase in intracellular cAMP lead to an increase in luciferase expression. In a standard assay, the oxidation of luciferin by luciferase results in luminescence, which can then be detected and quantified.
[0063] The sensory potency of an odorant, and therefore the potency of odorant-OR interaction in vivo, is most commonly described as the odor detection threshold (OTH) of a certain odorant, that is, the minimum concentration in the gas phase that can be detected by the human nose.The terms "odor detection threshold" and "odor threshold", also abbreviated as "OTH" herein, are synonymous and are well-established terms in the field of fragrance.See, for example, "The Measuring of Odors" by Neuner-Jehle, N., Etzweiler, F., in Perfumes, Springer, Dordrecht, edited by Mailer, PM, Lamparsky, D. (1994), the entire contents of which are incorporated herein by reference.Odor detection threshold can be measured by methods and means commonly available in the art, for example, by using an olfactometer in conjunction with a human subject. Another possibility to measure the odor detection threshold is to inject a dilution series of a defined amount (measured in ng) of a ligand into a gas chromatograph (GC), whereby a human panelist sniffs the molecules released from the GC column in the sniff port and indicates whether they are detectable by the nose, thereby obtaining the GC-threshold (GCCO) in ng.
[0064] In the OR screening assay, which applies OR in an in vitro system, the ligand is dissolved in a liquid medium. Nevertheless, similar to the in vivo odor threshold experiment, the minimum dose at which an odorant begins to activate receptors can be determined in an in vitro experiment, whereby the sensitivity in an in vitro system is reported as the lowest concentration dissolved in the medium that begins to activate receptors, while the sensitivity in vivo is expressed as the lowest detectable concentration in the gas phase. In practical terms, the detection threshold in vitro can be defined as the concentration that produces a two-fold response relative to background signals, for example, in a luciferase assay as described herein.
[0065] The functional expression of the olfactory receptor polypeptide as described herein is improved (increased) compared to baseline functional expression, resulting in improved (increased) biological activity, for example, compared to the corresponding unmodified receptor gene. The functional expression may be improved (increased) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 100%, at least 200%, at least 300%, or at least 500% compared to the corresponding unmodified receptor gene. Such improved functional expression may also be manifested in an increased sensitivity of recombinant host cells expressing the olfactory receptor polypeptide as described herein, such that the receptor detection threshold in an in vitro assay is reduced by at least 2-fold, at least 3.1-fold, at least 10-fold, at least 20-fold, at least 50-fold, preferably at least 3.1-fold (i.e., the detection sensitivity of the tested ligand is increased). Such improved functional expression may also manifest itself in an increased sensitivity of recombinant host cells expressing olfactory receptor polypeptides as described herein (as described elsewhere herein), such that the EC50, i.e., the concentration required to reach 50% of the maximum signal amplitude, is reduced by at least 2-fold, at least 3.1-fold, at least 10-fold, at least 20-fold, or at least 50-fold, preferably at least 3.1-fold (i.e., the efficacy of the tested ligand is enhanced). Such improved functional expression may also manifest itself in an increased sensitivity of recombinant host cells expressing olfactory receptor polypeptides as described herein, such that the maximum signal amplitude is enhanced by at least 30%, at least 50%, at least 80%, at least 100%, at least 200%, or at least 500% (i.e., the efficacy of the tested ligand is increased). The improvement in functional expression may also be of such magnitude that functional expression of an olfactory receptor that was not possible using conventional approaches is achieved using the olfactory receptors, nucleic acid molecules, expression vectors, recombinant host cells, and methods and uses of the present disclosure (an all-or-nothing effect that cannot be quantified in numbers).
[0066] Several conventional notations for describing sequence motifs are in use and known to those skilled in the art, most of which are variations of the standard notation of regular expressions and use at least the following conventions: - There is a single-letter alphabet, i.e., the standard IUPAC one-letter codes for amino acids, each of which represents a specific amino acid or set of amino acids; - The strings created from the alphabet represent the corresponding amino acid sequences; - the string between the square brackets matches any one of the listed amino acids, i.e. it represents a sequence ambiguity or alternative, e.g. the hypothetical notation [XYZ] means X or Y or Z; - A string of characters between braces (curly brackets) ("{}") stands for any amino acid except the listed amino acids, e.g., the hypothetical notation {X} stands for any amino acid except X.
[0067] Thus, the 16 amino acids in the sequence motif above: RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 1) is defined as follows: - the first residue is R; - the second residue is N; - the third residue is K or R; - the fourth residue is E or D or Q; - the fifth residue is V or M or I or L; - the sixth residue is K or R; - the seventh residue can be any amino acid; - the eighth residue is A; - the 9th residue is L or I or V; - the 10th residue is K or R or H; - the 11th residue is K or R; - residue 12 is L or I; - residue 13 is L or I or F; - residue 14 is K or R or G; - the 15th residue is K or R; and - Residue 16 is K or R.
[0068] Thus, the sequence motif mentioned above: RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 1) Alternatively, RNX1X2X3X4X5AX6X7X8X9X 10 X 11 X 12 X 13 may be written as: - X1 is K or R; - X2 is E or D or Q; - X3 is V or M or I or L; - X4 is K or R; - X5 is any amino acid; - X6 is L or I or V; - X7 is K or R or H; - X8 is K or R; - X9 is L or I; - X 10 is L or I or F; - X 11 is K or R or G; - X 12 is K or R; and - X 13 is K or R.
[0069] Considering that the standard IUPAC one-letter code for amino acids includes the symbol "J" to represent leucine (L) or isoleucine (I), the sequence motif described above can alternatively be RNX1X2X3X4X5AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 1) may be written as: - X1 is K or R; - X2 is E or D or Q; - X3 is V or M or I or L; - X4 is K or R; - X5 is any amino acid; - X6 is L or I or V; - X7 is K or R or H; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 12 is K or R; and - X 13 is K or R.
[0070] All of the above notations are fully equivalent and represent the same sequence (SEQ ID NO: 1); therefore, those skilled in the art will understand that they can be used interchangeably. The same is true for the other sequence motifs described herein.
[0071] Olfactory receptors can be classified into class I or "fish-like" receptors and class II or "tetrapod" receptors. Class I or "fish-like" receptors are an evolutionarily older receptor class known to be particularly responsive to more soluble compounds such as carboxylic acids, making them particularly interesting for screening antagonists against malodorous substances. The majority of olfactory receptors belong to class II or "tetrapod" receptors, which include important receptors for most aromatic molecules. Although there are significant differences in the C-terminal sequences between class I and class II receptors, the present disclosure targets both class I and class II olfactory receptors.
[0072] Thus, in some embodiments, the olfactory receptor protein as described herein is a class I or class II olfactory receptor having a modified C-terminal domain.
[0073] The olfactory receptor proteins of the present disclosure have a modified C-terminal domain. Thus, the C-terminal domain of the olfactory receptor proteins of the present disclosure is different from the natural or homologous C-terminal domain with which the olfactory receptor is normally associated. Therefore, the olfactory receptor proteins of the present disclosure are not naturally occurring proteins. The olfactory receptor proteins described herein can be characterized as "modified" olfactory receptor proteins, "engineered" olfactory receptor proteins, "hybrid" olfactory receptor proteins, "chimeric" olfactory receptor proteins, "non-natural" olfactory receptor proteins, or similar expressions and combinations thereof. Throughout this disclosure, those skilled in the art will understand that the term "olfactory receptor protein" may be replaced with the term "olfactory receptor protein with a modified C-terminal domain."
[0074] The present disclosure encompasses olfactory receptor proteins from any mammal.Therefore, in some embodiments, the olfactory receptor protein as described herein is a mammalian olfactory receptor with modified C-terminal domain, preferably a mammalian class I or class II olfactory receptor with modified C-terminal domain.Preferred mammals in the context of the present disclosure are pet or companion animals and humans, more preferably humans.Among pet or companion animals, cats and dogs are particularly preferred.Therefore, in some embodiments, the olfactory receptor protein as described herein is a human, canine or feline olfactory receptor with modified C-terminal domain, preferably a human olfactory receptor with modified C-terminal domain.In some embodiments, the olfactory receptor protein as described herein is a human, canine or feline class I or class II olfactory receptor with modified C-terminal domain, preferably a human class I or class II olfactory receptor with modified C-terminal domain.
[0075] Mammalian and human olfactory receptors are discussed in publications such as Mainland et al. (2015) Sci Data 2:150002 (incorporated herein by reference) and in publicly available databases such as the HORDE (The Human Olfactory Data Explorer) database maintained by the Weizmann Institute of Science, available at https: / / genome.weizmann.ac.il / horde / , and described in Olender et al. (2013) Methods Mol Biol 1003:23-38 (incorporated herein by reference).Other relevant publicly available databases exist, for example, as described in Marenco et al. Database (Oxford) 2016:baw132 and Han et al. Science China. Life sciences, 10.1007 / s11427-021-2081-6.
[0076] A complete list of human olfactory receptors and their alleles or haplotypes can be found in the HORDE database mentioned above. A major allele or haplotype is defined by the HORDE database as one with a frequency of 20% or more in the human population. A list of the complete amino acid sequences of the majority of the major alleles or haplotypes of human olfactory receptors is available to those skilled in the art and can be found as supplementary information in Ikegami et al. 2020 (supra), while the nucleotide sequences of 625 genes covering the majority of human OR genes and their major alleles or haplotypes are included as supplementary information in Mainland et al. 2015 (supra). All known human olfactory receptors can also be found in general sequence databases, such as NCBI Genbank, available at https: / / www.ncbi.nlm.nih.gov / genbank / .
[0077] In some embodiments, the olfactory receptor protein as described herein is selected from the group consisting of: OR10A2, OR10A3, OR10A4, OR10A5, OR10A6, OR10A7, OR10AD1, OR10AG1, OR10C1, OR10D3, OR10G2, OR10G3, OR10G4, OR10G6, OR10G7, OR10G8, OR10G9, OR10H1, OR10H2, OR10H3, OR10H4, OR10H5, OR10J1, OR10J3, OR10J5, OR10K1, OR10K2, OR1 0P1, OR10P2, OR10Q1, OR10R2, OR10S1, OR10T2, OR10V1, OR10W1, OR10X1, OR10Z1, OR11A1, OR11G2, OR11H1, OR11H2, OR11H4, OR11H6, OR11L1, OR12D2, OR 12D3, OR13C2, OR13C3, OR13C4, OR13C5, OR13C8, OR13C9, OR13D1, OR13F1, OR13G1, OR13H1, OR13J1, OR14A16, OR14A2, OR14C36, OR14I1, OR14L1P, OR1A1 , OR1A2, OR1B1, OR1C1, OR1D2, OR1D4, OR1D5, OR1E1, OR1E2, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L6, OR1 L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1S1, OR1S2, OR2A1, OR2A12, OR2A14, OR2A2, OR2A25, OR2A4, OR2A42, OR2A5, OR2A7, OR2A9P, ORAE1, OR2AG1, OR2AG2, O R2AJ1, OR2AK2, OR2AP1, OR2AT4, OR2B11, OR2B2, OR2B3, OR2B6, OR2B8P, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2F2, OR2G2, OR2G3, OR2G6, OR2H1, OR2H2, OR2J1P, OR2J2, OR2J3, OR2K2, OR2L13, OR2L2, OR2L3, OR2L5, OR2L8, OR2M2, OR2M3, OR2M4, OR2M5, OR2M7, OR2S2, OR2T1, OR2T10, OR2T11, OR2T12, OR2T2,OR2T27、OR2T29、OR2T3、OR2T33、OR2T34、OR2T35、OR2T4OR2T5、OR2T6、OR2T7、OR2T8、OR2V1、OR2V2、OR2W1、OR2W3、OR2W5、OR2Y1、OR2Z1、OR3A1、OR3A2、OR3A3、OR3A4、OR4A15、OR4A16、OR4A4、R4A47、OR4A5、OR4B1、OR4C11、OR4C12、OR4C13、OR4C15、OR4C16、OR4C、OR4C45、OR4C46、OR4C5、OR4C6、OR4D1、OR4D10、OR4D11、OR4D2、OR4D5、OR4D6、OR4D9、OR4E2、OR4F15、OR4F16、OR4F17、OR4F21、OR4F29、OR4F3、OR4F4、OR4F5、OR4F6、OR4K1、OR4K13、OR4K14、OR4K15、OR4K17、OR4K2、OR4K3P、OR4K5、OR4L1、OR4M1、OR4M2、OR4N2、OR4N4、OR4N5、OR4P4、OR4Q3、OR4S1、OR4S2、OR4X1、OR4X2、OR51A2、OR51A4、OR51A7、OR51B2、OR51B4、OR51B5、OR51B6、OR51D1、OR51E1、OR51E2、OR51F1、OR51F2、OR51G1、OR51G2、OR51H1P、OR51I1、OR51I2、OR51J1、OR51L1、OR51M1、OR51Q1、OR51S1、OR51T1、OR51V1、OR52A1、OR52A4、OR52A5、OR52B2、OR52B4、OR52B6、OR52D1、OR52E2、OR52E4、OR52E5、OR52E6、OR52E8、OR52H1、OR52I1、OR52I2、OR52J3、OR52K1、OR52K2、OR52L1、OR52M1、OR52N1、OR52N2、OR52N4、OR52N5、OR52P1P、OR52R1、OR52W1、OR56A1、OR56A3、OR56A4、OR56A5、OR56B1、OR56B4、OR5A1、OR5A2、OR5AC2、OR5AK2、OR5AL1P、OR5AN1、OR5AP2、OR5AR1、OR5AS1、OR5AU1、OR5B12、OR5B17、OR5B2、OR5B21、OR5B3、OR5C1、OR5D13、OR5D14、OR5D16, OR5D18, OR5F1, OR5H1, OR5H14, OR5H15, OR5H2, OR5H6, OR5I1, OR5J2, OR5K1, OR5K2, OR5 K3, OR5K4, OR5L1, OR5L2, OR5M1, OR5M10, OR5M11, OR5M3, OR5M8, OR5M9, OR5P2, OR5P3, OR5R1, OR5 T1, OR5T2, OR5T3, OR5V1, OR5W2, OR6A2, OR6B1, OR6B2, OR6B3, OR6C1, OR6C2, OR6C3, OR6C4, OR6C 6, OR6C65, OR6C68, OR6C70, OR6C74, OR6C75, OR6C76, OR6F1, OR6J1, OR6K2, OR6K3, OR6K6, OR6M1, OR6N1, OR6N2, OR6P1, OR6Q1, OR6S1, OR6T1, OR6X1, OR6Y1, OR7A10, OR7A17, OR7A5, OR7C1, OR7C2 , OR7D2, OR7D4, OR7E24, OR7G1, OR7G2, OR7G3, OR8A1, OR8B12, OR8B2, OR8B3, OR8B4, OR8B8, OR8D1 , OR8D2, OR8D4, OR8G1, OR8G5, OR8H1, OR8H2, OR8H3, OR8I2, OR8J1, OR8J3, OR8K1, OR8K3, OR8K5, OR8S1, OR8U1, OR8U8, OR8U9, OR9A2, OR9A4, OR9G1, OR9G4, OR9G9, OR9I1, OR9K2, OR9Q1, and OR9Q2. Variants and haplotypes of these receptors are also encompassed.
[0078] In some embodiments, the olfactory receptor protein as described herein is selected from the group consisting of: a class II olfactory receptor listed under "receptor" in Table 15, or - Class I olfactory receptors listed as "receptors" in Table 22 is.
[0079] The specific sequences mentioned in Table 15 (for DNA encoding the modified receptor comprising an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 221)) and Table 22 (for DNA encoding the modified receptor comprising an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 741)) are specific specific embodiments, and are not limiting in this context. The presence of the N-terminal tag is optional, and a different N-terminal tag can also be used, as described elsewhere herein. Similarly, any modified C-terminal domain disclosed herein can be used in place of the modified C-terminus of SEQ ID NO: 221 or SEQ ID NO: 741.
[0080] The modified C-terminal domain as described herein may comprise the sequence RNKEVKDALKRLLKRK (SEQ ID NO: 10). In some embodiments, the sequence RNKEVKDALKRLLKRK (SEQ ID NO: 10) may contain amino acid substitutions at up to 1, 2, 3, 4, 5, 6, 7, 8, or 9 positions. In some embodiments, the sequence may contain amino acid substitutions at up to 1, 2, 3, 4, or 5 positions, but the first (R), second (N), and eighth (A) residues are fixed. In some embodiments, the sequence may contain amino acid substitutions at up to 1, 2, 3, 4, 5, 6, 7, 8, or 9 positions, but the first (R), second (N), fourth (E), and eighth (A) residues are fixed. The amino acid substitutions are preferably conservative amino acid substitutions, as described elsewhere herein. Examples of particularly suitable amino acid substitutions in this context include replacing K with R and R with K.
[0081] In some embodiments, the olfactory receptor protein as described herein is the class II olfactory receptor with modified C-terminal domain.In some embodiments, the class II olfactory receptor is selected from the group consisting of OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2 and OR4S2, and preferably selected from the group consisting of OR7A17, OR2M2 and OR5A2. In some embodiments, the class II olfactory receptor is selected from the group consisting of OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5A1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2.In some embodiments, the class II olfactory receptor is OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5 , OR2A25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2, and preferably selected from the group consisting of OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2, and OR5A2.
[0082] OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, O R1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3(S73G), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2 are human class II olfactory receptors.
[0083] OR2A25 may also be OR2A25(S75N, A209P). OR1N2 may also be OR1N2(W23R, V230G, T287M). OR7E24 may also be OR7E24(P242S). OR2AG2 may also be OR2AG2(Y28C). OR8K3 may also be OR8K3(L122R). OR2L2 may also be OR2L2(V259L). OR11G2 may also be OR11G2(I65N, V82I). OR10G7 may also be OR10G7(T5S). OR2AK2 may also be OR2AK2(S84N). OR10A6 may also be OR10A6(A117V, V140G, L287P). OR10J1 may also be OR10J1(M51I, I92M). OR10G3 may also be OR10G3(S73G).
[0084] In some embodiments, (wild-type) OR7C1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26664 (NCBI Reference Sequence NP_001357414.2). In some embodiments, (wild-type) OR9Q2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219957 (NCBI Reference Sequence: NP_001005283.1). In some embodiments, (wild-type) OR8K3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219473 (NCBI Reference Sequence: NP_001005202.1). In some embodiments, (wild-type) OR10J5 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 127385 (NCBI Reference Sequence: NP_001004469.1). In some embodiments, (wild-type) OR1C1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26188 (NCBI Reference Sequence: NP_036485.2). In some embodiments, (wild-type) OR7D4 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 125958 (NCBI Reference Sequence: NP_001005191.1). In some embodiments, (wild-type) OR2T4 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 127074 (NCBI Reference Sequence: NP_001004696.2). In some embodiments, (wild-type) OR5B12 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390191 (NCBI Reference Sequence: NP_001004733.1). In some embodiments, (wild-type) OR7A17 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26333 (NCBI Reference Sequence: NP_112163.1). In some embodiments, (wild-type) OR10H5 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 284433 (NCBI Reference Sequence: NP_001004466.1).In some embodiments, (wild-type) OR5A2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219981 (NCBI Reference Sequence: NP_001001954.1). In some embodiments, (wild-type) OR5A1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219982 (NCBI Reference Sequence: NP_001004728.1). In some embodiments, (wild-type) OR1N2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 138882 (NCBI Reference Sequence: NP_001004457.2). In some embodiments, (wild-type) OR2C1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 4993 (NCBI Reference Sequence: NP_036500.2). In some embodiments, (wild-type) OR2T11 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 127077 (NCBI Reference Sequence: NP_001001964.1). In some embodiments, (wild-type) OR2M2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 391194 (NCBI Reference Sequence: NP_001004688.1). In some embodiments, (wild-type) OR4S2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219431 (NCBI Reference Sequence: NP_001004059.2). In some embodiments, (wild-type) OR2L3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 391192 (NCBI Reference Sequence: NP_001004687.1). In some embodiments, (wild-type) OR2AG2 has the amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 338755 (NCBI Reference Sequence: NP_001004490.1). In some embodiments, (wild-type) OR7A5 has the amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26659 (NCBI Reference Sequence: NP_001357409.1).In some embodiments, (wild-type) OR7E24 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26648 (NCBI Reference Sequence: NP_001073404.1). In some embodiments, (wild-type) OR7A10 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390892 (NCBI Reference Sequence: NP_001005190.1). In some embodiments, (wild-type) OR10H2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26538 (NCBI Reference Sequence: NP_039227.1). In some embodiments, (wild-type) OR10H1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26539 (NCBI Reference Sequence: NP_039228.1). In some embodiments, (wild-type) OR10D3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26497 (NCBI Reference Sequence: NP_001342142.1). In some embodiments, (wild-type) OR1D2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 4991 (NCBI Reference Sequence: NP_001373017.1). In some embodiments, (wild-type) OR2A5 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 393046 (NCBI Reference Sequence: NP_036497.1). In some embodiments, (wild-type) OR2A25 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 392138 (NCBI Reference Sequence: NP_001004488.1). In some embodiments, (wild-type) OR11G2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390439 (NCBI Reference Sequence: NP_001005503.2). In some embodiments, (wild-type) OR14J1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 442191 (NCBI Reference Sequence: NP_112208.1).In some embodiments, (wild-type) OR5M3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219482 (NCBI Reference Sequence: NP_001004742.2). In some embodiments, (wild-type) OR8D1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 283159 (NCBI Reference Sequence: NP_001002917.1). In some embodiments, (wild-type) OR10G3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26533 (NCBI Reference Sequence: NP_001005465.1). In some embodiments, (wild-type) OR10G9 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219870 (NCBI Reference Sequence: NP_001001953.1). In some embodiments, (wild-type) OR2L5 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 81466 (NCBI Reference Sequence: NP_001245213.1). In some embodiments, (wild-type) OR8H1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 219469 (NCBI Reference Sequence: NP_001005199.1). In some embodiments, (wild-type) OR10K1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 391109 (NCBI Reference Sequence: NP_001004473.1). In some embodiments, (wild-type) OR11A1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26531 (NCBI Reference Sequence: NP_001381757.1). In some embodiments, (wild-type) OR2V1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26693 (NCBI Reference Sequence: NP_001245212.1). In some embodiments, (wild-type) OR5P3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 120066 (NCBI Reference Sequence: NP_703146.1).In some embodiments, (wild-type) OR6P1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 128366 (NCBI Reference Sequence: NP_001153797.1). In some embodiments, (wild-type) OR2L2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26246 (NCBI Reference Sequence: NP_001004686.1). In some embodiments, (wild-type) OR10G7 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390265 (NCBI Reference Sequence: NP_001004463.1). In some embodiments, (wild-type) OR5AN1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390195 (NCBI Reference Sequence: NP_001004729.1). In some embodiments, (wild-type) OR5V1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 81696 (NCBI Reference Sequence: NP_110503.3). In some embodiments, (wild-type) OR2AK2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 391191 (NCBI Reference Sequence: NP_001004491.2). In some embodiments, (wild-type) OR10A3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26496 (NCBI Reference Sequence: NP_001003745.1). In some embodiments, (wild-type) OR10A6 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390093 (NCBI Reference Sequence: NP_001004461.1). In some embodiments, (wild-type) OR10J1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26476 (NCBI Reference Sequence: NP_001350486.1). In some embodiments, (wild-type) OR2J2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26707 (NCBI Reference Sequence: NP_112167.2).These OR major and functional alleles or haplotypes were used herein.
[0085] In some embodiments, the olfactory receptor protein as described herein is a class I olfactory receptor with a modified C-terminal domain.In some embodiments, the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1.OR51B2 may preferably be OR51B2(C120R, L134F, C209S).OR52K1 may also be OR52K1(Q52R).OR51B5 may also be OR51B5(G5S).OR56A3 may also be OR56A3(M51T).
[0086] OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1 are human class I olfactory receptors. In some embodiments, (wild-type) OR52A5 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390054 (NCBI Reference Sequence: NP_001005160.1). In some embodiments, (wild-type) OR52E8 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390079 (NCBI Reference Sequence: NP_001005168.2). In some embodiments, (wild-type) OR56A4 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 120793 (NCBI Reference Sequence: NP_001005179.3). In some embodiments, (wild-type) OR51B2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 79345 (NCBI Reference Sequence: NP_149420.4). In some embodiments, (wild-type) OR52K1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390036 (NCBI Reference Sequence: NP_001005171.2). In some embodiments, (wild-type) OR56A1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 120796 (NCBI Reference Sequence: NP_001001917.3). In some embodiments, (wild-type) OR51B5 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 282763 (NCBI Reference Sequence: NP_001005567.2). In some embodiments, (wild-type) OR56A3 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 390083 (NCBI Reference Sequence: NP_001003443.2). In some embodiments, (wild-type) OR51L1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 119682 (NCBI Reference Sequence: NP_001004755.1).
[0087] As detailed in the experimental section, the inventors have found that a particular sequence corresponding to the sequence motif of SEQ ID NO: 1 is particularly advantageous.Then, a more preferred sequence motif has been identified.Therefore, in a preferred embodiment, the olfactory receptor protein as described herein has the sequence motif: RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (Sequence number 5) This is what happens.
[0088] Sequence motif: RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (Sequence number 5) Alternatively, RNX1EX 3' X4X5AX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 5) may be written as: - X1 is K or R; - X3' is V or M or I; - X4 is K or R; - X5 is any amino acid; - X6 is L or I or V; - X 7' is K or R; - X8 is K or R; - X 10 is L or I or F; - X 11' is K or R; - X 12 is K or R; and - X 13 is K or R.
[0089] In some embodiments, the olfactory receptor protein as described herein comprises the sequence motif: RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 824) This is what happens. Sequence motif: RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 824) Alternatively, RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 824) may be written as: - X1 is K or R; - X5 is any amino acid; - X6 is L or I or V; - X7 is K or R or H; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 12 is K or R; and - X 13 is K or R.
[0090] A modified C-terminal domain containing the motif of SEQ ID NO: 824 may be particularly advantageous in the case of class I olfactory receptors.
[0091] The sequence motifs of SEQ ID NO: 1, SEQ ID NO: 5, and SEQ ID NO: 824 contain one residue that can be any amino acid, which is herein designated as "x" or "X5" above. The inventors have found that certain residues at this position result in particularly advantageous sequences. Therefore, in a further preferred embodiment, the olfactory receptor proteins described herein and having a modified C-terminal domain containing any of the above amino acid sequence motifs are such that "x" or "X5" is selected from any amino acid except proline (Pro, P).
[0092] In light of the above, in a preferred embodiment, the olfactory receptor protein as described herein comprises the sequence motif: RN[KR][EDQ][VMIL][KR]{P}A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 311) This is what happens.
[0093] Sequence motif: RN[KR][EDQ][VMIL][KR]{P}A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 311) Alternatively, RNX1X2X3X4X 5'''' AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 311) may be written as: - X1 is K or R; - X 2は、 E or D or Q; - X3 is V or M or I or L; - X4 is K or R; - X 5'''' is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y (any amino acid except P); - X6 is L or I or V; - X7 is K or R or H; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 12 is K or R; and - X 13 is K or R.
[0094] In other further preferred embodiments, the olfactory receptor as described herein has the sequence motif: RNX1EX 3' X4X 5'''' AX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 312) wherein: - X1 is K or R; - X 3' is V or M or I; - X4 is K or R; - X 5'''' is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y (any amino acid except P); - X6 is L or I or V; - X 7' is K or R; - X8 is K or R; - X 10 is L or I or F; - X 11' is K or R; - X 12 is K or R; and - X 13 is K or R.
[0095] In some embodiments, the olfactory receptor proteins described herein and having a modified C-terminal domain containing any of the above amino acid sequence motifs are such that "x" or "X5" is selected from any amino acid except proline (Pro, P) and tryptophan (Trp, W).
[0096] In some embodiments, the olfactory receptor protein having a modified C-terminal domain comprising any of the above amino acid sequence motifs as described herein, "x" or "X5" is selected from D, K, R, E, N, V, A, Q, or G, preferably "x" or "X5" is selected from D, K, R, E, N, V, A, or Q, more preferably "x" or "X5" is selected from D, K, or R. In other words, in a further preferred embodiment, the olfactory receptor protein as described herein has a sequence motif: RN[KR][EDQ][VMIL][KR][DKRENVAQG]A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 6), Preferably: RN[KR][EDQ][VMIL][KR][DKRENVAQ]A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 7), more preferably: RN[KR][EDQ][VMIL][KR][DKR]A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 166) Or will it be; Or the sequence motif is RN[KR]E[VMI][KR][DKRENVAQG]A[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 8), preferably: RN[KR]E[VMI][KR][DKRENVAQ]A[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 9), more preferably: RN[KR]E[VMI][KR][DKR]A[LIV][KR][KR]L[LIF][KR][KR][KR] (Sequence number 167) This is what happens.
[0097] In some embodiments, the olfactory receptor protein as described herein comprises the sequence motif: RN[KR][EDQ][VMIL][KR]KA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (Sequence number 820) This is what happens. The sequence motif RN[KR][EDQ][VMIL][KR]KA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 820) can alternatively be RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 820) may be written as: - X1 is K or R; - X2 is E or D or Q; - X3 is V or M or I or L; - X4 is K or R; - X6 is L or I or V; - X7 is K or R or H; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 12 is K or R; and - X 13 is K or R.
[0098] In some embodiments, the olfactory receptor protein as described herein comprises the sequence motif: RN[KR]E[VMI][KR]KA[LIV][KR][KR]L[LIF][KR][KR][KR] (Sequence number 821) This is what happens.
[0099] Sequence motif: RN[KR]E[VMI][KR]KA[LIV][KR][KR]L[LIF][KR][KR][KR] (Sequence number 821) Alternatively, RNX1EX 3' X4KAX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 821) may be written as: - X1 is K or R; - X 3' is V or M or I; - X4 is K or R; - X5 is any amino acid; - X6 is L or I or V; - X 7' is K or R; - X8 is K or R; - X 10 is L or I or F; - X 11' is K or R; - X 12 is K or R; and - X 13 is K or R.
[0100] In some embodiments, the olfactory receptor protein as described herein comprises the sequence motif: RNKEVKX 5'''' ALKRLLKRK (SEQ ID NO: 319) where: X 5''''is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y.
[0101] Some specific and particularly advantageous sequences corresponding to the sequence motifs described herein have been identified. Thus, in some embodiments, the olfactory receptor protein as described herein has a sequence motif selected from the group consisting of: RNKEVKDALKRLLKRK (SEQ ID NO: 10), RNREVKDALKRLLKRK (SEQ ID NO: 11), RNKEIKDALKRLLKRK (SEQ ID NO: 12), RNKEMKDALKRLLKRK (SEQ ID NO: 13), RNKEVRDALKRLLKRK (SEQ ID NO: 14), RNKEVKKALKRLLKRK (SEQ ID NO: 15), RNKEVKEALKRLLKRK (SEQ ID NO: 16), RNKEVKNALKRLLKRK (SEQ ID NO: 17), RNKEVKRALKRLLKRK (SEQ ID NO: 18), RNKEVKVALKRLLKRK (SEQ ID NO: 19), RNKEVKAALKRLLKRK (SEQ ID NO: 20), RNKEVKQALKRLLKRK (SEQ ID NO: 21), RNKEVKDAVKRLLKRK (SEQ ID NO: 22), RNKEVKDAIKRLLKRK (SEQ ID NO: 23), RNKEVKDALRRLLKRK (SEQ ID NO: 24), RNKEVKDALKKLLKRK (SEQ ID NO: 25), RNKEVKDALKRLIKRK (SEQ ID NO: 26), RNKEVKDALKRLFKRK (SEQ ID NO: 27), RNKEVKDALKRLLRRK (SEQ ID NO: 28), RNKEVKDALKRLLKKK (SEQ ID NO: 29), RNKEVKDALKRLLKRR (SEQ ID NO: 30), RNKDVKDALKRLLKRK (SEQ ID NO: 31), RNKQVKDALKRLLKRK (SEQ ID NO: 32), RNKELKDALKRLLKRK (SEQ ID NO: 33), RNKEVKGALKRLLKRK (SEQ ID NO: 34), RNKEVKDALHRLLKRK (SEQ ID NO: 35), RNKEVKDALKRILKRK (SEQ ID NO: 36), RNKEVKDALKRLLGRK (SEQ ID NO: 37), RNKEVKRAIKRLLKRK (SEQ ID NO: 38), RNKEVKKAIKRLLKRK (SEQ ID NO: 39), RNKEVKRAIKRLFKRK (SEQ ID NO: 40), RNKEVKKAIKRLFKRK (SEQ ID NO: 41), RNKEVKRAIRKLLKRK (SEQ ID NO: 42), RNKEVKDALRKLLKRK (SEQ ID NO: 43), RNKEVKDALKRLLRRR (SEQ ID NO: 44), RNREMRKALHRLLGKK (SEQ ID NO: 827), RNREVKKAIHKLIGRK (SEQ ID NO: 828), RNREVRKAVHRLFKRK (SEQ ID NO: 829), RNKEMKKAIHKLFGKK (SEQ ID NO: 830), RNRDVKKAVHKLFRRK (SEQ ID NO: 831), RNRDMKKAVHKLFGKR (SEQ ID NO: 832), RNKELRKALHKLLGRK (SEQ ID NO: 833), RNRDVRKALRRILRRR (SEQ ID NO: 834), RNKDVRKAVRKLIRRR (SEQ ID NO: 835), RNRDVRKAVRRLFRKR (SEQ ID NO: 836), RNKDIKKAVKKLIKKK (SEQ ID NO: 837), RNRELRKAVRRLFKRR (SEQ ID NO: 838), RNKELRKAVRKIIKKK (SEQ ID NO: 839), RNRDVKKAVRRLFRRK (SEQ ID NO: 840), RNREVRKALRRIIRKR (SEQ ID NO: 841), RNKDIRKAVKKIFRRK (SEQ ID NO: 842), RNKDVRKAVRRLIKRK (SEQ ID NO: 843), RNRDLRKAVRKLFKKK (SEQ ID NO: 844), RNRDLRKALRRIFKRR (SEQ ID NO: 845), RNRDVRKAIKKLIRKR (SEQ ID NO: 846), RNKELKKAIKRILKKK (SEQ ID NO: 847), RNRDVRKAIRKLLKRK (SEQ ID NO: 848), RNRDLRKAVRRIFKKR (SEQ ID NO: 849), RNRDVRKAVRKLFKRR (SEQ ID NO: 850), RNRDVRKALRRLFKKR (SEQ ID NO: 851), RNKELKKALRKLIGKK (SEQ ID NO: 852), RNREMRKAIKKIIKKK (SEQ ID NO: 853), RNKEIKKAIKKIIKKR (SEQ ID NO: 854), RNRDVKKAIRRLFRRR (SEQ ID NO: 855), RNREVKKAVKKLIGKR (SEQ ID NO: 856), RNREMRKALRRLFRKR (SEQ ID NO: 857), RNKELKKALRRLIGRR (SEQ ID NO: 858), RNRDVKKALRKLIGKR (SEQ ID NO: 859), RNREVKKAVKKLIRRK (SEQ ID NO: 860), RNKEVRKALKKLFGKK (SEQ ID NO: 861), RNKEIRKALRRLFGKK (SEQ ID NO: 862), RNKDVKKALRRLFGKK (SEQ ID NO: 863), RNKELKKAIKRLIRRK (SEQ ID NO: 864), RNKDVRKAVKRLLKKR (SEQ ID NO: 865), RNKELRKAIRRLLRRR (SEQ ID NO: 866), RNRDIRKALRKLFKKK (SEQ ID NO: 867), RNRELKKALRRLLRRR (SEQ ID NO: 868), RNREVKKALRRLFGKK (SEQ ID NO: 869), RNRDVRKALKRLLKRK (SEQ ID NO: 870), RNRDMRKAIRKLFGRK (SEQ ID NO: 871), RNRELKKAIRKLLKRK (SEQ ID NO: 872), RNRDIRKAVKKLFGKK (SEQ ID NO: 873), RNKEVKKAIRKLFGRR (SEQ ID NO: 874), RNREVRKAVRKLFRRK (SEQ ID NO: 875), RNRDMKKALKKLFRRR (SEQ ID NO: 876), RNRDVRKALKRLLGRR (SEQ ID NO: 877), RNKDLKKAVKKLFGRK (SEQ ID NO: 878), RNKDVRKAVRRLFGRR (SEQ ID NO: 879), RNKEVKCALKRLLKRK (SEQ ID NO: 880), RNKEVKFALKRLLKRK (SEQ ID NO: 881), RNKEVKHALKRLLKRK (SEQ ID NO: 882), RNKEVKIALKRLLKRK (SEQ ID NO: 883), RNKEVKLALKRLLKRK (SEQ ID NO: 884), RNKEVKMALKRLLKRK (SEQ ID NO: 885), RNKEVKSALKRLLKRK (SEQ ID NO: 886), RNKEVKTALKRLLKRK (SEQ ID NO: 887), RNKEVKWALKRLLKRK (SEQ ID NO: 888), and RNKEVKYALKRLLKRK (SEQ ID NO: 889).
[0102] As detailed in the experimental section of the present disclosure, it has been shown that the olfactory receptor with the modified C-terminal domain, which comprises any of these specific sequences listed above, exhibits advantageous and surprising technical effects.Based on these findings, those skilled in the art can design other advantageous sequences that are similarly compatible with the scope of sequence motifs described herein.Some illustrative and non-limiting examples of such sequences are as follows:
[0103] RNKEVKRALKRLLRRR (SEQ ID NO: 45) RNKEVKKALKRLLRRR (SEQ ID NO: 46) RNREVKRAIKRLLKRK (SEQ ID NO: 47) RNREVKKAIKRLLKRK (SEQ ID NO: 48) RNREVKRAIKRLFKRK (SEQ ID NO: 49) RNREVKKAIKRLFKRK (SEQ ID NO: 50) RNREVKRAIRKLLKRK (SEQ ID NO: 51) RNREVKDALRKLLKRK (SEQ ID NO: 52) RNREVKDALKRLLRRR (SEQ ID NO: 53) RNKEVKKAIKRLLRRK (SEQ ID NO: 54) RNKEVKKAIKRLLKKK (SEQ ID NO: 55) RNKEVKKAIKRLLKRR (SEQ ID NO: 56) RNKEVKRAIKRLLRRK (SEQ ID NO: 57) RNKEVKRAIKRLLKKK (SEQ ID NO: 58) RNKEVKRAIKRLLKRR (SEQ ID NO: 59) RNKEVKKAIKRLFRRK (SEQ ID NO: 60) RNKEVKKAIKRLFKKK (SEQ ID NO: 61) RNKEVKKAIKRLFKRR (SEQ ID NO: 62) RNKEVKRAIKRLFRRK (SEQ ID NO: 63) RNKEVKRAIKRLFKKK (SEQ ID NO: 64) RNKEVKRAIKRLFKRR (SEQ ID NO: 65) RNREVKRAIKRLLRKK (SEQ ID NO: 66) RNREVKKAIKRLLRKK (SEQ ID NO: 67) RNREVKRAIKRLFRRR (SEQ ID NO: 68) RNREVKKAIKRLFRRR (SEQ ID NO: 69) RNREVKKAIKRLFRRK (SEQ ID NO: 70) RNREVKKAIKRLFKKK (SEQ ID NO: 71) RNREVKKAIKRLFKRR (SEQ ID NO: 72) RNREVKRAIKRLFRRK (SEQ ID NO: 73) RNREVKRAIKRLFKKK (SEQ ID NO: 74) RNREVKRAIKRLFKRR (SEQ ID NO: 75)
[0104] In some embodiments, the olfactory receptor protein as described herein has the sequence motif RNRDVRKALRRLFRKK (SEQ ID NO: 307) or RNRDVRRALRRLFRKK (SEQ ID NO: 308) Such a C-terminal motif can be particularly advantageous because it contains the statistically optimal amino acid at each individual position.
[0105] In some embodiments, the olfactory receptor protein as described herein has the sequence motif RNKQIRDALKRLLKRK (SEQ ID NO: 890). Such a C-terminal motif may be particularly advantageous for class I olfactory receptors.
[0106] The 16-amino acid sequence motifs described herein may optionally contain additional C-terminal residues. Indeed, the inventors have found that, although not required, such additional C-terminal residues can provide additional beneficial effects, as described in detail in the experimental section. If present, the number of additional C-terminal residues is not particularly limited, but is preferably less than 10. More preferably, if present, the number of additional C-terminal residues is 1 to 6. Thus, depending on the number of additional C-terminal residues added to the 16-amino acid sequence motifs described herein, the total length of the sequence motifs described herein, when any additional C-terminal residues are present, may be 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 amino acids, preferably 17, 18, 19, 20, 21, or 22 amino acids.
[0107] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the amino acid sequence motif comprises 1 to 10, preferably 1 to 6, additional C-terminal amino acid residues. In some embodiments, the additional C-terminal amino acid residues are as follows: - the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from C, R, K, E, G, H, F, P and Y, more preferably the first additional amino acid residue is C, R or K, most preferably the first additional amino acid residue is C; - the second amino acid residue is selected from C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second amino acid residue is selected from C, R, K, N, G, I, L, F, P, T and Y, more preferably the second amino acid residue is C, R or K, most preferably the second amino acid residue is C or R; - the third amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S and G, more preferably the third amino acid residue is R or K; the 4th to 10th, or the 4th, 5th and 6th, preferably the 4th, 5th and 6th, additional amino acid residues are selected from K and R.
[0108] In some embodiments, the amino acid sequence motif comprises one additional C-terminal residue, which preferably corresponds to the first additional amino acid residue defined above, and more preferably is selected from C, R, P, L, K, G, Y, F, M, or W.
[0109] In some embodiments, the amino acid sequence motif comprises two additional C-terminal residues.The first and second additional amino acid residues are preferably as defined above.Specific advantageous examples of the combination of the first and second additional amino acid residues include CC, SI, YP, PQ, FR, CR, RR, EK, PR, CG, FK, RG, RC, RT, RF, GG, YR, GC, TG, PC, HP, PG, KY, CP, YY, FF, CF, NP, YL, IC, HC, CL, YC, ER, RP, PA, FC and RY.Particularly preferred in this context is the CC sequence.
[0110] Thus, in some embodiments, the olfactory receptor protein as described herein has a sequence motif of: RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 CC (SEQ ID NO: 822), RNX1EX 3’ X4KAX6X 7’ X8LX 10 X 11’ X 12 X 13 CC (SEQ ID NO: 823), and; RNKEVKX 5'''' ALKRLLKRKCC (SEQ ID NO: 320) such that the compound is selected from the group consisting of: - X1 is K or R; - X2 is E or D or Q; - X3 is V or M or I or L; - X 3' is V or M or I; - X4 is K or R; - X 5'''' is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y; - X6 is L or I or V; - X7 is K or R or H; - X 7' is K or R; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 11' is K or R; - X 12 is K or R; and - X 13 is K or R.
[0111] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of: RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), RNREVKDALKRLLKRKCC (SEQ ID NO: 87), RNKEIKDALKRLLKRKCC (SEQ ID NO: 88), RNKEMKDALKRLLKRKCC (SEQ ID NO: 89), RNKEVRDALKRLLKRKCC (SEQ ID NO: 90), RNKEVKKALKRLLKRKCC (SEQ ID NO: 91), RNKEVKEALKRLLKRKCC (SEQ ID NO: 92), RNKEVKNALKRLLKRKCC (SEQ ID NO: 93), RNKEVKRALKRLLKRKCC (SEQ ID NO: 94), RNKEVKVALKRLLKRKCC (SEQ ID NO: 95), RNKEVKAALKRLLKRKCC (SEQ ID NO: 96), RNKEVKQALKRLLKRKCC (SEQ ID NO: 97), RNKEVKDAVKRLLKRKCC (SEQ ID NO: 98), RNKEVKDAIKRLLKRKCC (SEQ ID NO: 99), RNKEVKDALRRLLKRKCC (SEQ ID NO: 100), RNKEVKDALKKLLKRKCC (SEQ ID NO: 101), RNKEVKDALKRLIKRKCC (SEQ ID NO: 102), RNKEVKDALKRLFKRKCC (SEQ ID NO: 103), RNKEVKDALKRLLRRKCC (SEQ ID NO: 104), RNKEVKDALKRLLKKKCC (SEQ ID NO: 105), RNKEVKDALKRLLKRRCC (SEQ ID NO: 106), RNKDVKDALKRLLKRKCC (SEQ ID NO: 107), RNKQVKDALKRLLKRKCC (SEQ ID NO: 108), RNKELKDALKRLLKRKCC (SEQ ID NO: 109), RNKEVKGALKRLLKRKCC (SEQ ID NO: 110), RNKEVKDALHRLLKRKCC (SEQ ID NO: 111), RNKEVKDALKRILKRKCC (SEQ ID NO: 112), RNKEVKDALKRLLGRKCC (SEQ ID NO: 113), RNKEVKRAIKRLLKRKCC (SEQ ID NO: 114), RNKEVKKAIKRLLKRKCC (SEQ ID NO: 115), RNKEVKRAIKRLFKRKCC (SEQ ID NO: 116), RNKEVKKAIKRLFKRKCC (SEQ ID NO: 117), RNKEVKRAIRKLLKRKCC (SEQ ID NO: 118), RNKEVKDALRKLLKRKCC (SEQ ID NO: 119), RNKEVKDALKRLLRRRCC (SEQ ID NO: 120), RNREMRKALHRLLGKKCC (SEQ ID NO: 254), RNREVKKAIHKLIGRKCC (SEQ ID NO: 255), RNREVRKAVHRLFKRKCC (SEQ ID NO: 256), RNKEMKKAIHKLFGKKCC (SEQ ID NO: 257), RNRDVKKAVHKLFRRKCC (SEQ ID NO: 258), RNRDMKKAVHKLFGKRCC (SEQ ID NO: 259), RNKELRKALHKLLGRKCC (SEQ ID NO: 260), RNRDVRKALRRILRRRCC (SEQ ID NO: 261), RNKDVRKAVRKLIRRRCC (SEQ ID NO: 262), RNRDVRKAVRRLFRKRCC (SEQ ID NO: 263), RNKDIKKAVKKLIKKKCC (SEQ ID NO: 264), RNRELRKAVRRLFKRRCC (SEQ ID NO: 265), RNKELRKAVRKIIKKKCC (SEQ ID NO: 266), RNRDVKKAVRRLFRRKCC (SEQ ID NO: 267), RNREVRKALRRIIRKRCC (SEQ ID NO: 268), RNKDIRKAVKKIFRRKCC (SEQ ID NO: 269), RNKDVRKAVRRLIKRKCC (SEQ ID NO: 270), RNRDLRKAVRKLFKKKCC (SEQ ID NO: 271), RNRDLRKALRRIFKRRCC (SEQ ID NO: 272), RNRDVRKAIKKLIRKRCC (SEQ ID NO: 273), RNKELKKAIKRILKKKCC (SEQ ID NO: 274), RNRDVRKAIRKLLKRKCC (SEQ ID NO: 275), RNRDLRKAVRRIFKKRCC (SEQ ID NO: 276), RNRDVRKAVRKLFKRRCC (SEQ ID NO: 277), RNRDVRKALRRLFKKRCC (SEQ ID NO: 278), RNKELKKALRKLIGKKCC (SEQ ID NO: 279), RNREMRKAIKKIIKKKCC (SEQ ID NO: 280), RNKEIKKAIKKIIKKRCC (SEQ ID NO: 281), RNRDVKKAIRRLFRRRCC (SEQ ID NO: 282), RNREVKKAVKKLIGKRCC (SEQ ID NO: 283), RNREMRKALRRLFRKRCC (SEQ ID NO: 284), RNKELKKALRRLIGRRCC (SEQ ID NO: 285), RNRDVKKALRKLIGKRCC (SEQ ID NO: 286), RNREVKKAVKKLIRRKCC (SEQ ID NO: 287), RNKEVRKALKKLFGKKCC (SEQ ID NO: 288), RNKEIRKALRRLFGKKCC (SEQ ID NO: 289), RNKDVKKALRRLFGKKCC (SEQ ID NO: 290), RNKELKKAIKRLIRRKCC (SEQ ID NO: 291), RNKDVRKAVKRLLKKRCC (SEQ ID NO: 292), RNKELRKAIRRLLRRRCC (SEQ ID NO: 293), RNRDIRKALRKLFKKKCC (SEQ ID NO: 294), RNRELKKALRRLLRRRCC (SEQ ID NO: 295), RNREVKKALRRLFGKKCC (SEQ ID NO: 296), RNRDVRKALKRLLKRKCC (SEQ ID NO: 297), RNRDMRKAIRKLFGRKCC (SEQ ID NO: 298), RNRELKKAIRKLLKRKCC (SEQ ID NO: 299), RNRDIRKAVKKLFGKKCC (SEQ ID NO: 300), RNKEVKKAIRKLFGRRCC (SEQ ID NO: 301), RNREVRKAVRKLFRRKCC (SEQ ID NO: 302), RNRDMKKALKKLFRRRCC (SEQ ID NO: 303), RNRDVRKALKRLLGRRCC (SEQ ID NO: 304), RNKDLKKAVKKLFGRKCC (SEQ ID NO: 305), RNKDVRKAVRRLFGRRCC (SEQ ID NO: 306), RNRDVRKALRRLFRKKCC (SEQ ID NO: 309), RNRDVRRALRRLFRKKCC (SEQ ID NO: 310), RNKEVKCALKRLLKRKCC (SEQ ID NO: 321), RNKEVKFALKRLLKRKCC (SEQ ID NO: 322), RNKEVKHALKRLLKRKCC (SEQ ID NO: 323), RNKEVKIALKRLLKRKCC (SEQ ID NO: 324), RNKEVKLALKRLLKRKCC (SEQ ID NO: 325), RNKEVKMALKRLLKRKCC (SEQ ID NO: 326), RNKEVKSALKRLLKRKCC (SEQ ID NO: 328), RNKEVKTALKRLLKRKCC (SEQ ID NO: 329), RNKEVKWALKRLLKRKCC (SEQ ID NO: 330), RNKEVKYALKRLLKRKCC (SEQ ID NO: 331), and RNKQIRDALKRLLKRKCC (SEQ ID NO: 740).
[0112] A C-terminal motif having the sequence RNRDVRKALRRLFRKKCC (SEQ ID NO: 309) or RNRDVRRALRRLFRKKCC (SEQ ID NO: 310) may be particularly advantageous because it is composed of the statistically optimal amino acid at each individual position. A C-terminal motif having the sequence RNKQIRDALKRLLKRKCC (SEQ ID NO: 740) may be particularly advantageous in the case of class I olfactory receptors.
[0113] In addition to the CC sequence, other preferred combinations of the first and second additional amino acid residues are CR, RR, YP, RF and FK. Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of: RNKEVKRAIKRLLKRKCR (SEQ ID NO: 121), RNKEVKKAIKRLLKRKCR (SEQ ID NO: 122), RNKEVKRALKRLLKRKRR (SEQ ID NO: 123), RNKEVKRALKRLLKRKYP (SEQ ID NO: 124), RNKEVKRALKRLLKRKRF (SEQ ID NO: 125), RNKEVKRALKRLLKRKFK (SEQ ID NO: 126), RNKEVKKALKRLLKRKRR (SEQ ID NO: 127), RNKEVKKALKRLLKRKYP (SEQ ID NO: 128), RNKEVKKALKRLLKRKRF (SEQ ID NO: 129), and RNKEVKKALKRLLKRKFK (SEQ ID NO: 130).
[0114] In some embodiments, the amino acid sequence motif comprises three additional C-terminal residues. The first, second, and third amino acid residues are preferably as defined above. In some embodiments, the first and second additional amino acid residues are CC, and the third additional amino acid residue is as defined above. Specific advantageous examples of combinations of the first, second, and third additional amino acid residues include CRR, CCC, CCF, CCL, CCM, CCS, CCP, CCA, CCY, CCH, CCN, CCD, CCK, CCR, and CCG.
[0115] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of: RNKEVKDALKRLLKRKCRR (SEQ ID NO: 133), RNKEVKDALKRLLKRKCCC (SEQ ID NO: 134), RNKEVKDALKRLLKRKCCF (SEQ ID NO: 135), RNKEVKDALKRLLKRKCCL (SEQ ID NO: 136), RNKEVKDALKRLLKRKCCM (SEQ ID NO: 137), RNKEVKDALKRLLKRKCCS (SEQ ID NO: 138), RNKEVKDALKRLLKRKCCP (SEQ ID NO: 139), RNKEVKDALKRLLKRKCCA (SEQ ID NO: 140), RNKEVKDALKRLLKRKCCY (SEQ ID NO: 141), RNKEVKDALKRLLKRKCCH (SEQ ID NO: 142), RNKEVKDALKRLLKRKCCN (SEQ ID NO: 143), RNKEVKDALKRLLKRKCCD (SEQ ID NO: 144), RNKEVKDALKRLLKRKCCK (SEQ ID NO: 145), RNKEVKDALKRLLKRKCCR (SEQ ID NO: 146), and RNKEVKDALKRLLKRKCCG (SEQ ID NO: 147).
[0116] In some embodiments, the amino acid sequence motif comprises four additional C-terminal residues. The first, second, third and fourth additional amino acid residues are preferably as defined above. Specific advantageous examples of combinations of the first, second, third and fourth additional amino acid residues include CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CCRR (SEQ ID NO: 161), CCRK (SEQ ID NO: 228) and CCKR (SEQ ID NO: 229), among which CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160) and CCRR (SEQ ID NO: 161) are preferred. Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of: RNKEVKDALKRLLKRKCRRR (SEQ ID NO: 149), RNKEVKDALKRLLKRKCRKK (SEQ ID NO: 150), and RNKEVKDALKRLLKRKCCRR (SEQ ID NO: 151).
[0117] In some embodiments, the amino acid sequence motif comprises five additional C-terminal residues. The first, second, third, fourth, and fifth additional amino acid residues are preferably as defined above. Specific advantageous examples of combinations of the first, second, third, fourth, and fifth additional amino acid residues include CRRRR (SEQ ID NO: 162), CCRRR (SEQ ID NO: 163), CCKRR (SEQ ID NO: 230), CCRKR (SEQ ID NO: 231), CCRRK (SEQ ID NO: 232), CCRKK (SEQ ID NO: 233), CCKRK (SEQ ID NO: 234), CCKKR (SEQ ID NO: 235), and CCKKK (SEQ ID NO: 236), among which CRRRR (SEQ ID NO: 162) and CCRRR (SEQ ID NO: 163) are preferred.
[0118] In some embodiments, the olfactory receptor protein as described herein has a sequence motif of the following: RNX1X2X3X4X5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 825), and RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 826) is selected from the group consisting of: - X1 is K or R; - X2 is E or D or Q; - X3 is V or M or I or L; - X4 is K or R; - X 5は、 Any amino acid; - X6 is L or I or V; - X7 is K or R or H; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 12 is K or R; and - X 13 is K or R. Preferably, X5 is any amino acid except proline (Pro, P).
[0119] In some embodiments, the olfactory receptor protein as described herein has a sequence motif of the following: RNKEVKDALKRLLKRKCRRRR (SEQ ID NO: 154), RNKEVKDALKRLLKRKCCRRR (SEQ ID NO: 156), RNKEVKKAIKRLLKRKCCRRR (SEQ ID NO: 220), RNKEVKRAIKRLLKRKCCRRR (SEQ ID NO: 237), RNKEVKKAIKRLFKRKCCRRR (SEQ ID NO: 221), RNKEVKRAIKRLFKRKCCRRR (SEQ ID NO: 238), and RNKQIRDALKRLLKRKCCRRR (SEQ ID NO: 741), and preferably such that the sequence motif is RNKEVKKAIKRLFKRKCCRRR (SEQ ID NO: 221). A C-terminal motif having the sequence RNKQIRDALKRLLKRKCCRRR (SEQ ID NO: 741) may be particularly advantageous in the case of class I olfactory receptors.
[0120] In some embodiments, the amino acid sequence motif comprises six additional C-terminal residues. The first, second, third, fourth, fifth, and sixth additional amino acid residues are preferably as defined above. Specific advantageous examples of combinations of the first, second, third, fourth, fifth, and sixth additional amino acid residues include CRRRRR (SEQ ID NO: 164), CRRRKK (SEQ ID NO: 165), and CCRRRR (SEQ ID NO: 224).
[0121] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of: RNKEVKDALKRLLKRKCRRRRR (SEQ ID NO: 157), RNKEVKDALKRLLKRKCRRRKK (SEQ ID NO: 158), and RNKEVKDALKRLLKRKCCRRRR (SEQ ID NO: 219).
[0122] As described in detail in the experimental section of this disclosure, olfactory receptors having modified C-terminal domains containing any of these specific sequences listed above have been shown to exhibit advantageous and surprising technical effects.
[0123] In some embodiments, the olfactory receptor protein as described herein may be such that the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164), and CRRRKK (SEQ ID NO: 165).
[0124] In some embodiments, the olfactory receptor described herein may further comprise an N-terminal signal peptide. Typically, the N-terminal signal peptide is a cleavable peptide. This means that it cannot be cleaved from the mature protein and alter OR-ligand binding and signal transduction. Therefore, it is understood by those skilled in the art that the N-terminal signal peptide described herein is not usually part of the mature olfactory receptor. In some embodiments, the N-terminal signal peptide is a leucine-rich peptide, preferably MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77). MRPQILLLLALLTLGLA (SEQ ID NO: 76) and MSHQILLLLALLTLGLA (SEQ ID NO: 77) are known as Lucy tags. MRPQILLLLALLTLGLA (SEQ ID NO: 76) is a human Lucy tag, while MSHQILLLLALLTLGLA (SEQ ID NO: 77) is a mouse Lucy tag. Also encompassed are the sequences of SEQ ID NOs: 76 and 77 in which 1, 2, 3, 4 or up to 5 amino acids have been substituted, deleted, added or inserted. Substitutions, and especially conservative substitutions, are preferred.
[0125] In some embodiments, the olfactory receptor described herein further comprises an N-terminal tag peptide. Typically, the N-terminal tag peptide is a non-cleavable peptide. The N-terminal tag peptide may be an epitope tag used to purify or capture the protein. An example of such an epitope tag is a FLAG tag, which is further described below. The N-terminal tag peptide may also be a peptide that enhances expression. Examples of such tag peptides that enhance expression are rhodopsin (rho) tag, SST3 tag, and M3 tag.
[0126] In some embodiments, the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag (the 45-N-terminal amino acid of the somatostatin 3 receptor; an example of which is SEQ ID NO: 223), an M3 tag (the 61-N-terminal amino acid of the muscarinic acetylcholine receptor M3; an example of which is SEQ ID NO: 222), a c-myc tag, and an HA tag, preferably selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag, and an M3 tag. Preferably, the olfactory receptor protein as described herein further comprises an N-terminal tag peptide, more preferably selected from the group consisting of FLAG tag, rhodopsin (rho) tag, IL-6 tag, SST3 tag, M3 tag, c-myc tag, and HA tag, more preferably selected from the group consisting of FLAG tag, rhodopsin (rho) tag, SST3 tag, and M3 tag, even more preferably selected from the group consisting of FLAG tag and rhodopsin (rho) tag, most preferably rhodopsin (rho) tag. Thus, in some embodiments, the olfactory receptor protein as described herein further comprises an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of FLAG tag, rhodopsin (rho) tag, IL-6 tag, SST3 tag, M3 tag, c-myc tag, and HA tag, more preferably selected from the group consisting of FLAG tag, rhodopsin (rho) tag, SST3 tag, and M3 tag, even more preferably selected from the group consisting of FLAG tag and rhodopsin (rho) tag, most preferably rhodopsin (rho) tag.
[0127] In a preferred embodiment, the N-terminal tag peptide comprises at least a tag peptide selected from the group consisting of a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, preferably a rho tag. Optionally, a FLAG peptide may further be present.
[0128] Combinations of the above tag peptides can also be used.Typically, this combination comprises at least one of rho tag, SST3 tag and M3 tag, preferably at least rho tag.For example, in some embodiments, N-terminal tag peptide is the combination of Rho tag and FLAG tag.In other words, in some embodiments, the olfactory receptor protein as described herein further comprises N-terminal tag peptide, where N-terminal tag peptide comprises Rho tag and FLAG tag.Preferably, in this situation, FLAG tag is located at the N-terminal side from Rho tag, for example, as shown in SEQ ID NO:81.
[0129] The FLAG tag and rho tag are described in Shepard et al. (2013) PloS One 8(7): e68758, Zhuang and Matsunami (2007) J Biol Chem 282(20): 15284-15293, and WO2014 / 037800, each of which is incorporated herein by reference. The IL-6 tag is described in Noe et al. A bi-functional IL-6-HaloTag® as a tool to measure the cell-surface expression of recombinant odorant receptors and to facilitate their activity quantification. J Biol Methods. 2017, 4(4): e82, which is incorporated herein by reference. The SST3 tag and M3 tag are described in Tan et al. (2022) Scientific reports 12:17658, which is incorporated herein by reference.
[0130] In some embodiments, the FLAG tag as described herein has the sequence of SEQ ID NO: 78. In some embodiments, the rho tag as described herein has the sequence of SEQ ID NO: 79. In some embodiments, the IL-6 tag as described herein is IL-6-HaloTag® as described in Noe et al. (supra) A bi-functional IL-6-HaloTag® as a tool to measure the cell-surface expression of recombinant odorant receptors and to facilitate their activity quantification. J Biol Methods. 2017, 4(4):e82, which is incorporated herein by reference. In some embodiments, the SST3 tag as described herein has the sequence of SEQ ID NO: 223. In some embodiments, the M3 tag as described herein has the sequence of SEQ ID NO: 222. Also encompassed are the sequences of SEQ ID NOs: 78, 79, 222, and 223, in which up to 1, 2, 3, 4, or 5 amino acids have been substituted, deleted, added, or inserted. Substitutions, and especially conservative substitutions, are preferred.
[0131] As described herein, the N-terminal signal peptide and as described herein, the N-terminal tag peptide can be advantageously used in combination with each other.Usually, the N-terminal signal peptide will be located upstream (N-terminal side) from the N-terminal tag.For example, as described herein, the olfactory receptor can further comprise an N-terminal signal peptide and one or more N-terminal tag peptides.In some embodiments, as described herein, the olfactory receptor can further comprise: - human or mouse Lucy signal peptide (such as SEQ ID NO: 76 or 77); - a FLAG tag (such as SEQ ID NO: 78) or an IL-6 tag, preferably a FLAG tag (such as SEQ ID NO: 78); and a rho tag (such as SEQ ID NO: 79), an SST3 tag or an M3 tag, preferably a rho tag (such as SEQ ID NO: 79). SEQ ID NO: 80 is an example of a nucleotide sequence encoding a combination of mouse Lucy signal peptide, FLAG tag peptide, and rho tag peptide (SEQ ID NO: 81).
[0132] In some embodiments, the olfactory receptor as described herein is modified to include one or more additional N-terminal glycosylation sites.Such glycosylation sites exist, for example, in the M3 and SST3 tags (Tan et al., Scientific Reports (2022) 12:17658) and N-terminal rho tags (Kaushal et al., 1998, Proc Natl Acad Sci USA 91(9):4024-4028), as described elsewhere herein.
[0133] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of SEQ ID NOs: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, and 820-890. Also encompassed in this context are the sequences of SEQ ID NOs: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, and 820-890, in which 1, 2, 3, 4, 5, 6, 7, 8, or up to 9 amino acids are substituted, deleted, added, or inserted. Preferably, sequences including substitutions, deletions, additions, and / or insertions still correspond to the general sequence motif of SEQ ID NO: 1. Substitutions, and especially conservative substitutions, are preferred. Examples of particularly suitable amino acid substitutions in this context include substituting K for R and R for K.
[0134] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of SEQ ID NOs: 1, 10-75, 86-130, 133-147, 149-151, 154, 156-158, 198, 219-221, 254-310, 321-326, 328-331, 740-741, and 827-890. Also included in this context are the sequences of SEQ ID NOs: 1, 10-75, 86-130, 133-147, 149-151, 154, 156-158, 198, 219-221, 254-310, 321-326, 328-331, 740-741, and 827-890, in which 1, 2, 3, 4, 5, 6, 7, 8, or up to 9 amino acids are substituted, deleted, added, or inserted. Substitutions, and especially conservative substitutions, are preferred. Examples of particularly suitable amino acid substitutions in this context include replacing K with R and R with K.
[0135] It is understood that in the context of any of the olfactory receptors described throughout this disclosure, the term "comprising" may be replaced with the term "essentially consisting of" or "consisting of." In other words, in some embodiments, the olfactory receptors described herein have a modified C-terminal domain that consists essentially of the amino acid sequence motifs disclosed herein or that consists of the amino acid sequence motifs disclosed herein.
[0136] nucleic acid molecule In another aspect, the present disclosure relates to a nucleic acid molecule comprising a nucleotide sequence encoding any of the olfactory receptor proteins described herein. The nucleotide sequence encoding an olfactory receptor may also be referred to as the "gene" encoding the olfactory receptor or the "coding sequence" for the olfactory receptor. The nucleotide sequence encoding an olfactory receptor is part of common general knowledge and can be obtained from general and specific sequence databases known by those skilled in the art, as described elsewhere herein.
[0137] The nucleic acid molecules of the present disclosure do not encode wild-type olfactory receptors, instead, they encode olfactory receptors with modified C-terminal domains, as detailed in the previous section.Therefore, the nucleic acid molecules of the present disclosure do not exist in nature.The nucleic acid molecules described herein can also be characterized as " modified " nucleic acid molecules, " engineered " nucleic acid molecules, " hybrid " nucleic acid molecules, " chimeric " nucleic acid molecules, " non-natural " nucleic acid molecules, or similar expressions and combinations thereof.
[0138] Exemplary nucleic acid molecules of the present disclosure are provided as SEQ ID NOs: 331-739 and 742-819. Thus, in some embodiments, the olfactory receptors described herein are at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119 ... Encoded by a nucleic acid molecule comprising a nucleotide sequence that comprises at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity.
[0139] SEQ ID NOs: 331-739 represent DNA encoding modified receptors including an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 221). SEQ ID NOs: 742-819 represent DNA encoding modified receptors including an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 741). As explained in this disclosure, the presence of the N-terminal tag is optional, and a different N-terminal tag can be used as described elsewhere herein. Similarly, any modified C-terminal domain disclosed herein may be used in place of the modified C-terminus of SEQ ID NO: 221 or SEQ ID NO: 741. SEQ ID NOs: 331-739 and SEQ ID NOs: 742-819 also include a 5' BamHI restriction site (GGATCC) and Kozak sequence (GCCACC), and a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes. The presence of these sequences is entirely optional.
[0140] The nucleic acid molecules of the present disclosure may contain additional sequence elements. Typically, the additional sequence elements may be sequence elements commonly used to support the expression of nucleotide sequences, such as promoters, nuclear localization signals, Kozak sequences, polyA tails, transcription terminators, etc. When one or more of such additional sequence elements are present, the nucleic acid molecules described herein may also be referred to as "nucleic acid constructs" or "gene constructs." It is understood that different sequence elements may be "operably linked" to each other to achieve a functional nucleic acid molecule. The term "operably linked" is provided elsewhere herein in the section entitled "General Information."
[0141] As used herein, "nucleic acid construct" refers to a DNA molecule comprising a region (coding region or ORF) that is transcribed into an RNA molecule (e.g., an mRNA molecule) in a cell, operably linked to suitable regulatory regions, such as, but not limited to, a promoter and / or enhancer sequence. A nucleic acid construct generally comprises multiple operably linked fragments, such as a promoter, an enhancer, a 5' leader sequence, a coding region, and / or a 3' untranslated region (3' end) (e.g., including a polyadenylation site and / or a transcription termination site). A nucleic acid construct may be recombinant, i.e., one that is not normally found in nature, for example, a nucleic acid construct in which a promoter is not naturally associated with part or all of the coding region. Molecular toolbox techniques for the preparation of nucleic acid constructs are well known in the art and are discussed in standard handbooks such as Ausubel et al., Current Protocols in Molecular Biology, 3rd Edition (2003), John Wiley & Sons Inc, and Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012), Cold Spring Harbor Laboratory Press, both of which are incorporated herein by reference in their entireties. Non-limiting examples of such techniques, some of which are presented herein in the Experimental Section, include fusion PCR, restriction enzyme digestion, Golden-gate cloning, etc.
[0142] In some embodiments, the nucleic acid molecule as described herein further comprises a promoter sequence. In other words, the present disclosure encompasses a nucleic acid molecule as described herein, wherein the nucleotide sequence is operably linked to a promoter sequence. In some embodiments, the promoter sequence as described herein is a constitutive promoter sequence.
[0143] As used herein, the term "promoter" or "transcriptional regulatory sequence" refers to a nucleic acid sequence that functions to control the transcription (i.e., expression) of one or more coding sequences, that is located upstream, in the direction of transcription, of a transcription start site of the coding sequence, and that is structurally identified by the presence of a binding site for DNA-dependent RNA polymerase, a transcription start site, and any other DNA sequences (including, but not limited to, transcription factor binding sites, repressor and activator protein binding sites, and other sequences of nucleotides known to those of skill in the art to act directly or indirectly to regulate the amount of transcription from the promoter).
[0144] In some embodiments, the promoter sequence as described herein that is a constitutive promoter sequence is a CMV promoter. In some embodiments, the CMV promoter may have the nucleotide sequence of SEQ ID NO: 82, or a sequence having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 82.
[0145] In some embodiments, a nucleic acid molecule as described herein further comprises an enhancer sequence. As used herein, the term "enhancer" refers to a nucleic acid sequence that can stimulate the transcription of a sequence to which it is operably linked. An operably linked enhancer does not necessarily have to be contiguous with the coding sequence whose transcription it controls. Enhancers may be used as a single sequence or may be included in a fusion nucleotide sequence with other enhancers and / or promoters, as described herein.
[0146] In some embodiments, the nucleic acid molecule as described herein further comprises a terminator sequence. In other words, the present disclosure encompasses the nucleic acid molecule as described herein, wherein the nucleotide sequence is operably linked to a terminator sequence. The "terminator sequence" may alternatively be referred to herein as a "transcription terminator," a "transcription terminator sequence," or simply a "terminator." In some embodiments, the terminator sequence is a bovine growth hormone (bgh) terminator sequence. In some embodiments, the bgh terminator sequence may have the nucleotide sequence of SEQ ID NO:83 or a sequence having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO:83.
[0147] In some embodiments, the nucleic acid molecule as described herein further comprises a nucleotide sequence encoding an N-terminal signal peptide. Suitable N-terminal signal peptides have been previously discussed herein. In some embodiments, the N-terminal signal peptide is a leucine-rich signal peptide, preferably a human Lucy tag or a mouse Lucy tag, more preferably a human Lucy tag or a mouse Lucy tag represented by the amino acid sequence MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77). Also encompassed are the sequences of SEQ ID NOs: 76 and 77, in which up to 1, 2, 3, 4, or 5 amino acids have been substituted, deleted, added, or inserted. Substitutions, and especially conservative substitutions, are preferred.
[0148] In some embodiments, the nucleic acid molecule as described herein further comprises a nucleotide sequence encoding an N-terminal tag peptide. Suitable N-terminal tag peptides have been previously discussed herein.
[0149] In some embodiments, the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag (the 45-N-terminal amino acid of the somatostatin 3 receptor; an example of which is SEQ ID NO: 223), an M3 tag (the 61-N-terminal amino acid of the muscarinic acetylcholine receptor M3; an example of which is SEQ ID NO: 222), a c-myc tag, and an HA tag, preferably selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag, and an M3 tag. More preferably, the tag is selected from the group consisting of a FLAG tag and a rhodopsin (rho) tag, and most preferably, a rhodopsin (rho) tag.
[0150] Thus, in some embodiments, the nucleic acid sequence as described herein further comprises a nucleic acid sequence encoding an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag, an M3 tag, a c-myc tag, and an HA tag, more preferably from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, even more preferably from the group consisting of a FLAG tag and a rhodopsin (rho) tag, and most preferably is a rhodopsin (rho) tag.
[0151] In a preferred embodiment, the N-terminal tag peptide comprises at least a tag peptide selected from the group consisting of a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, preferably a rho tag. Optionally, a FLAG peptide may further be present.
[0152] The combination of the above tag peptides can also be used.Typically, this combination comprises at least one of rho tag, SST3 tag and M3 tag, preferably at least rho tag.For example, in some embodiments, N-terminal tag peptide is the combination of rho tag and FLAG tag.In other words, in some embodiments, the nucleic acid molecule as described herein further comprises the nucleotide sequence encoding N-terminal tag peptide, wherein N-terminal tag peptide comprises rho tag and FLAG tag.Preferably, in this situation, FLAG tag is located at the N-terminal side from rho tag, for example as shown in SEQ ID NO:81.
[0153] In some embodiments, the FLAG tag encoded by the nucleic acid molecule as described herein has the sequence of SEQ ID NO: 78. In some embodiments, the rho tag encoded by the nucleic acid molecule as described herein has the sequence of SEQ ID NO: 79. In some embodiments, the IL-6 tag encoded by the nucleic acid molecule as described herein is IL-6-HaloTag® as described in Noe et al. In some embodiments, the SST3 tag encoded by the nucleic acid molecule as described herein has the sequence of SEQ ID NO: 223. In some embodiments, the M3 tag encoded by the nucleic acid molecule as described herein has the sequence of SEQ ID NO: 222. Also encompassed are the sequences of SEQ ID NOs: 78, 79, 222, and 223, in which up to 1, 2, 3, 4, or 5 amino acids have been substituted, deleted, added, or inserted. Substitutions, and especially conservative substitutions, are preferred.
[0154] The nucleotide sequence encoding the N-terminal signal peptide as described herein and the N-terminal tag peptide as described herein can advantageously be used in combination with each other.Typically, the nucleotide sequence encoding the N-terminal signal peptide is located upstream (N-terminal side) from the nucleotide sequence encoding the N-terminal tag.For example, the nucleic acid molecule as described herein can further comprise the nucleotide sequence encoding the N-terminal signal peptide and one or more N-terminal tag peptides.In some embodiments, the nucleic acid molecule as described herein can further comprise: - a nucleotide sequence encoding a human or mouse Lucy signal peptide (such as SEQ ID NO: 76 or 77); - a FLAG tag (such as SEQ ID NO: 78) or an IL-6 tag, preferably a FLAG tag (such as SEQ ID NO: 78); and a rho tag (such as SEQ ID NO: 79), an SST3 tag or an M3 tag, preferably a rho tag (such as SEQ ID NO: 79). SEQ ID NO: 80 is an example of a nucleotide sequence encoding a combination of mouse Lucy signal peptide, FLAG tag peptide, and rho tag peptide (SEQ ID NO: 81).
[0155] In some embodiments, a nucleic acid molecule as described herein is modified to include a nucleotide sequence encoding one or more additional N-terminal glycosylation sites.
[0156] In some embodiments, the nucleic acid molecule as described herein further comprises a nucleotide sequence encoding one or more olfactory receptor accessory proteins. An olfactory receptor "accessory protein" or "chaperone" is a protein or peptide that can assist in the expression, trafficking, and / or signal transduction of an olfactory receptor to the surface of a cell that expresses the olfactory receptor.
[0157] Non-limiting examples of accessory proteins encompassed by the present disclosure include RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf , Giα, or functional variants thereof, and are further described in WO2006 / 002161 and WO2014 / 037800, which are incorporated by reference in their entireties. Preferred accessory proteins are RTP1S and / or RTP2, preferably human RTP1S and / or human RTP2.
[0158] The accessory proteins described herein also encompass functional variants of their wild-type counterparts, i.e., accessory molecules that are modified compared to the corresponding naturally occurring or wild-type sequences.In this context, RTP1S, as used herein, encompasses the RTP1S V227I variant, and RTP2, as used herein, encompasses the RTP2 L220R variant.Preferred accessory proteins are the human RTP1S V227I variant (SEQ ID NO: 84) and the human RTP2 L220R variant (SEQ ID NO: 85).
[0159] Thus, in some embodiments, one or more olfactory receptor "accessory proteins" as described herein include RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf, Giα, and functional variants thereof, and preferably selected from the group consisting of RTP1S, RTP2, and functional variants thereof. In some embodiments, the one or more olfactory receptor "accessory proteins" as described herein are the RTP1S V227I variant and the RTP2 L220R variant. In some embodiments, the nucleotide sequence encoding one or more olfactory receptor accessory proteins comprises a nucleotide sequence encoding a polypeptide represented by SEQ ID NO: 84 and / or 85, or a nucleotide sequence encoding a polypeptide having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity or similarity to SEQ ID NO: 84 and / or 85.
[0160] The nucleotide sequences described herein may be codon-optimized for expression in host cells, preferably eukaryotic cells, more preferably human cells. Suitable host cells are also described elsewhere herein. See, for example, the "Cells" section. "Codon optimization," as used herein, refers to a process employed to modify an existing coding sequence or design a coding sequence, for example, to improve translation in an expression host cell or organism of an RNA molecule that is a transcript transcribed from the coding sequence, or to improve transcription of the coding sequence. Codon optimization includes, but is not limited to, the process of selecting codons for a coding sequence to match the codon preferences of the expression host cell or organism. Codon optimization also removes elements that potentially negatively affect RNA stability and / or translation (e.g., termination sequences, TATA boxes, splice sites, ribosome entry sites, repeat and / or GC-rich sequences, and RNA secondary structure or instability motifs). In some embodiments, the codon-optimized sequence exhibits at least a 3%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or more increase in gene expression, transcription, RNA stability, and / or translation compared to the original non-codon-optimized sequence.
[0161] Expression vector The nucleic acid molecules as described herein can be placed in an expression vector. Thus, in another aspect, there is provided an expression vector comprising any of the nucleic acid molecules as described herein.
[0162] An "expression vector," alternatively referred to herein as a "vector" or a "delivery vector," refers to a molecular biology tool used to obtain expression of a coding region (such as a gene) in a host cell, e.g., by introducing a nucleotide sequence capable of affecting the expression of the gene or coding sequence (in a host cell compatible with that sequence). The expression vector may be stabilized in the host cell and remain episomal. Alternatively, the vector may be integrated into the genome of the host cell, e.g., through homologous recombination, non-homologous end joining, etc. A description of suitable "host cells" in the context of the present disclosure is provided elsewhere herein.
[0163] Suitable expression vectors may be selected from any genetic element known in the art that can facilitate the transfer of nucleic acids between cells, including, but not limited to, plasmids, phages, transposons, cosmids, chromosomes, artificial chromosomes, viruses (e.g., but not limited to, retroviruses, lentiviruses, etc.), virions, etc. Expression vectors may also be chemical vectors, such as lipid complexes or naked DNA. "Naked DNA" or "naked nucleic acid" refers to a nucleic acid molecule that is not contained in an encapsulation vehicle and that facilitates delivery of the nucleic acid into the cytoplasm of a target host cell. Naked DNA may be circular or linear (linearized DNA sequence). Optionally, naked nucleic acids may be associated with standard means used in the art to facilitate delivery of the nucleic acid to the target host cell, for example, to facilitate transport of the nucleic acid across a cell membrane.
[0164] The preferred expression vector is a plasmid. Suitable plasmids are known in the art and are described in standard handbooks such as Ausubel et al. and Sambrook and Green (supra). Suitable plasmids may also be selected from commercially available vectors, such as the pcDNA3.1(+) series (Invitrogen, MA, USA) or pGL4.29 series vectors (Promega, WI, USA).
[0165] cell The nucleic acid molecules and expression vectors described herein are particularly useful for introduction into host cells.Therefore, in another aspect, a recombinant host cell is provided, comprising the nucleic acid molecule or expression vector as described herein before.Preferably, the recombinant host cell described herein expresses or is capable of expressing the olfactory receptor protein as described herein.
[0166] In some embodiments, the host cell of the present disclosure contains a plurality, i.e., two or more, of the nucleic acid molecules and / or expression vectors described herein. Thus, in such cases, the recombinant host cell expresses or is capable of expressing a plurality, i.e., two or more, of the olfactory receptor proteins described herein. In such host cells, two or more olfactory receptors are preferably activated by odorants with similar odors. Odorants with similar odors are typically described by the same odor descriptor. An "odor descriptor" is a general term used by trained perfumers and evaluators to describe a specific odor sensation common to a group of ligands or a mixture of ligands. For example, odor descriptors are "green" for an odor reminiscent of fresh crushed leaves and "floral" for a floral scent. They can also be more specific, such as "floral-rosy" for notes reminiscent of the scent of roses or "white floral" for notes reminiscent of ylang-ylang or jasmine.
[0167] In some embodiments, the combination of two or more olfactory receptor proteins in this context may be selected from the group consisting of: - OR7A17 and OR7C1. These ORs are activated by molecules with the odor descriptor "woody-ambery" and can be co-expressed in cells to detect woody-ambery notes. - OR5A2, OR5AN1, and OR1N2. These ORs are activated by molecules with "musk-like" odor descriptors, and two or more olfactory receptor proteins from this group can be co-expressed in cells to detect musk-like odors. - OR2L2, OR2L3, OR2L5, OR2AK2, and OR11G2. These ORs are activated by molecules with the odor descriptor "fruity-ester," and two or more olfactory receptor proteins from this group can be co-expressed in cells to detect fruity-ester notes. - OR10A3, OR10A6, and OR10J1. These ORs are activated by molecules with the odor descriptor "fruity-lactonic," and two or more olfactory receptor proteins from this group can be co-expressed in cells to detect fruity-lactonic notes. - OR10H1, OR10H2, OR10H5, and OR10K1. These ORs are activated by molecules with the odor descriptor "marine," and two or more olfactory receptor proteins from this group can be co-expressed in cells to detect marine notes. - OR10G7 and OR10D3. These ORs are activated by molecules with the odor descriptor "spicy" and can be co-expressed in cells to detect spicy notes.
[0168] "Host cells," alternatively referred to herein as "recombinant host cells," "engineered cells," or simply "cells," refer to cells that have been manipulated by the introduction of nucleic acid molecules and / or expression vectors as defined herein. Host cells may also refer to isolated cells or cells in culture. Host cells may also be "transformed cells," in which the cells have been infected with, for example, a modified virus. As a non-limiting example, lentivirus may be used, although other suitable viruses, such as retroviruses and others, are also contemplated. Introduction of a nucleic acid construct may also be achieved by non-viral methods, such as transfection. "Transfection" refers to non-viral methods of introducing DNA (or RNA) into cells so that the introduced nucleic acid sequence is expressed. Transfection methods and protocols are well known in the art; non-limiting examples include calcium phosphate transfection, PEG transfection, and liposome or lipoplex transfection, and are discussed in standard handbooks such as Ausubel et al. and Sambrook and Green (supra). Further examples of transfection methods are described herein in the exemplary section. Transfection can be transient or stable, the latter referring to when cells have the nucleic acid construct integrated into their genome. Thus, host cells containing the nucleic acid construct as described herein can also be "stably transfected cells" or "transiently transfected cells."
[0169] The host cell may be further genetically modified, for example, by introducing one or more genetic modifications, including but not limited to, mutation, substitution, insertion and / or deletion of nucleotides in its genome, and / or by introducing additional nucleic acid constructs.The modifications may be contained in the nucleotide sequence encoding the olfactory receptor, accessory molecule, and / or another genomic region, and may result in functional expression or improved functional expression of the olfactory receptor and / or accessory molecule.The definition of functional expression is provided elsewhere in this specification.
[0170] Modifications of nucleic acid sequences may be made using any recombinant DNA technique known in the art, for example, as described in standard handbooks such as Ausubel et al., and Sambrook and Green (supra). See also Kunkel (1985) Proc. Natl. Acad. Sci. 82:488 (describing site-directed mutagenesis) and Roberts et al. (1987) Nature 328:731 734 or Wells, JA, et al. (1985) Gene 34: 315 (describing cassette mutagenesis).
[0171] Host cells may contain epigenetic modifications in nucleic acid molecules encoding olfactory receptors, accessory proteins, and / or other genomic regions that can result in the functional expression of the olfactory receptors and / or accessory proteins, or can improve their functional expression.As used herein, the term "epigenetic modification" has its conventional meaning as generally understood by those skilled in the art in light of the present disclosure.It refers to chemical modifications of DNA or histone proteins that do not change their nucleotide sequence itself.Non-limiting examples of epigenetic modifications include methylation, acetylation, phosphorylation, serotonylation, citrullination, ubiquitination, sumoylation, and ribosylation of nucleic acids.
[0172] Advantageously, in some embodiments, the recombinant host cell as described herein further expresses one or more olfactory receptor accessory proteins as described herein. To achieve this, additional nucleic acid molecules or expression vectors encoding one or more olfactory receptor accessory proteins may be included in the recombinant host cell. These additional nucleic acid molecules or expression vectors may be stably integrated into chromosomes, or they may be introduced for transient expression, for example, by transfection. Suitable olfactory receptor accessory proteins have been described elsewhere herein. Thus, in some embodiments, one or more olfactory receptor "accessory proteins" as described herein include RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf , Giα, and functional variants thereof, and preferably selected from the group consisting of RTP1S, RTP2, and functional variants thereof. In some embodiments, the one or more olfactory receptor "accessory proteins" as described herein are the RTP1S V227I variant and the RTP2 L220R variant. In some embodiments, the nucleotide sequence encoding one or more olfactory receptor accessory proteins comprises a nucleotide sequence encoding a polypeptide represented by SEQ ID NO: 84 and / or 85, or a nucleotide sequence encoding a polypeptide having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity or similarity to SEQ ID NO: 84 and / or 85.
[0173] In some embodiments, the recombinant host cells as described herein further express one or more reporter genes, such as a luciferase gene. Reporter genes are described in more detail elsewhere herein.
[0174] The recombinant host cells as described herein may be prokaryotic or eukaryotic cells, preferably they are eukaryotic cells.Suitable prokaryotic cells may be selected from bacteria and archaea.Suitable eukaryotic cells may be selected from insect, plant, yeast, fungus, algae, mammalian and human cells, of which human cells are preferred.
[0175] Suitable host cells include HEK293, HEK293T, HeLa, CHO, OP6, HeLa-S3, HEKn, HEKa, PC-3, Calul, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial cells, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-L1, 132-d5 human embryonic fibroblasts, 10.1 mouse fibroblasts, 293-T, 3T3, BHK, BHK-21, BR 293, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, and CHO Dhfr − / −. These include, but are not limited to, COS-7, HL-60, LNCap, MCF-7, MCF-IOA, MDCK II, SkBr3, Vero cells, primary olfactory cells, immortalized olfactory cells, immortalized taste cells, and transgenic variants thereof, among which HEK293 and HEK293T are preferred. Thus, in some embodiments, the recombinant host cells described herein are HEK293 or HEK293T cells. Between HEK293 and HEK293T, HEK293T is more preferred. Cell lines are available from various conveniently available culture collections, such as the American Type Culture Collection (VA, USA).
[0176] Library In another aspect, the present disclosure relates to a library comprising a diverse repertoire of olfactory receptor proteins as described herein, nucleic acid molecules as described herein, expression vectors as described herein, or recombinant host cells as described herein.
[0177] In some embodiments of the library as described herein, the various repertoires of olfactory receptor proteins, the various repertoires of olfactory receptor proteins encoded by nucleic acid molecules or expression vectors, or the various repertoires of olfactory receptor proteins expressed by recombinant host cells share the same modified C-terminal domain.In a specific, but non-limiting example, they share the modified C-terminal domain represented by SEQ ID NO: 221 (particularly in the context of class II olfactory receptors) or SEQ ID NO: 741 (particularly in the context of class I olfactory receptors).
[0178] The library of olfactory receptors with the same C-terminal domain advantageously makes the functional activity of all receptors more uniform.The library of most receptors with wild-type C-terminal has significantly different functional expression levels, while in the library with the same C-terminal sequence, the functional expression between receptors is more similar.Therefore, by testing ligands against the normalized library with the same C-terminal, it is possible to find the most sensitive receptor that is activated by a given ligand, which is very important for the later screening of new ligands within a given odor type.On the other hand, if target receptors are identified from libraries with significantly different functional expression, receptor selection may be skewed by any functional expression in in vitro systems rather than by true ligand affinity.
[0179] The library as described herein is not particularly limited in terms of the number of distinct olfactory receptor proteins, the nucleic acid molecules or expression vectors that encode distinct olfactory receptor proteins, or the recombinant host cells that express distinct olfactory receptor proteins.However, in some embodiments, the library as described herein comprises a diverse repertoire of at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, or at least 400 distinct olfactory receptor proteins, the nucleic acid molecules or expression vectors that encode distinct olfactory receptor proteins, or the recombinant host cells that express distinct olfactory receptor proteins.
[0180] In some embodiments, the library described herein comprises a diverse repertoire of about 250 to about 800 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding distinct olfactory receptor proteins, or recombinant host cells expressing distinct olfactory receptor proteins. Such a library size allows coverage of most of canine olfactory receptors.
[0181] In some embodiments, the library described herein comprises a diverse repertoire of about 250 to about 670 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding distinct olfactory receptor proteins, or recombinant host cells expressing distinct olfactory receptor proteins. Such a library size allows coverage of most feline olfactory receptors.
[0182] In some embodiments, the library described herein comprises a diverse repertoire of about 250 to about 500 or about 250 to about 400 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding distinct olfactory receptor proteins, or recombinant host cells expressing distinct olfactory receptor proteins. Such a library size allows coverage of most human olfactory receptors.
[0183] In some embodiments, the library described herein comprises a diverse repertoire of about 400 to about 800 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding distinct olfactory receptor proteins, or recombinant host cells expressing distinct olfactory receptor proteins. Such a library size allows coverage of the majority of human olfactory receptors, including major alternative alleles or haplotypes.
[0184] A library may contain both class I and class II ORs, or the library may be specialized for class I ORs or class II ORs. Thus, in some embodiments, the library described herein comprises distinct class II olfactory receptor proteins, preferably human class II olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins. A library of class II ORs described herein comprises a diverse repertoire of at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, or at least 400 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the distinct olfactory receptor proteins, or recombinant host cells expressing the distinct olfactory receptor proteins. A preferred number for a library of human class II ORs is 400-450. Preferably, such a library contains at least one, preferably all, of the class II olfactory receptors listed under "receptor" in Table 15. In such a library, the class II olfactory receptor may differ from a specific SEQ ID NO in Table 15 in its N-terminal tag or its modified C-terminal domain. For example, instead of the modified C-terminal domain contained by the olfactory receptor encoded by the nucleic acid molecule listed in Table 15, the receptor may contain a different modified C-terminal domain as described herein. SEQ ID NOs: 331-739 also contain a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC), as well as a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes. The presence of these sequences is entirely optional.
[0185] In some embodiments, the library described herein comprises distinct class I olfactory receptor proteins, preferably human class I olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins. The library of class II ORs described herein comprises a diverse repertoire of at least 10, at least 25, at least 50, or at least 75 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the distinct olfactory receptor proteins, or recombinant host cells expressing the distinct olfactory receptor proteins. A preferred number for a library of human class II ORs is 70 to 100. Preferably, such a library comprises at least one, preferably all, of the class II olfactory receptors listed under "Receptor" in Table 22. In such a library, the class II olfactory receptor may differ from a specific SEQ ID NO in Table 22 in its N-terminal tag or its modified C-terminal domain. For example, instead of the modified C-terminal domain contained by the olfactory receptor encoded by the nucleic acid molecule listed in Table 22, the receptor may contain a different modified C-terminal domain as described herein. SEQ ID NOs: 742-819 also contain a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC), and a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes. The presence of these sequences is entirely optional.
[0186] In some embodiments, the library described herein comprises distinct Class I and Class II olfactory receptor proteins, preferably human Class I and Class II olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins. In this context, the library preferably comprises a diverse repertoire of 470 to 550 distinct olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the distinct olfactory receptor proteins, or recombinant host cells expressing the distinct olfactory receptor proteins.
[0187] Methods and Uses The olfactory receptors, nucleic acid molecules, recombinant host cells, and libraries described herein allow for the functional expression of olfactory receptors that have been impossible to express using conventional approaches or that have limited sensitivity when used in conventional approaches, allowing for the identification of novel cognate receptor-ligand pairs.Therefore, they are particularly useful for expressing olfactory receptors and for identifying novel olfactory receptors, as well as novel olfactory receptor ligands, enhancers, and antagonists.
[0188] In one aspect, the use of an olfactory receptor protein as described herein, a nucleic acid molecule as described herein, an expression vector as described herein, a cell as described herein, or a library as described herein to identify an olfactory receptor ligand, enhancer, or antagonist is provided.
[0189] In one aspect, there is provided the use of a library as described herein to identify olfactory receptors capable of binding to a target ligand.
[0190] In one aspect, a method for identifying an olfactory receptor ligand is provided, the method comprising: a) providing an olfactory receptor protein as described herein, or a cell expressing an olfactory receptor protein as described herein; b) contacting the receptor or cell with a test compound or composition; and c) Detecting activation of olfactory receptors.
[0191] In one aspect, a method for identifying an enhancer or antagonist of an olfactory receptor is provided, said method comprising: a) providing an olfactory receptor protein as described herein, or a cell expressing an olfactory receptor protein as described herein; b) contacting the receptor or cell with a cognate ligand and a test compound or composition; and c) Detecting increased or decreased activation of olfactory receptors compared to a ligand-only control.
[0192] As used herein, an "antagonist" of an olfactory receptor refers to a compound that reduces the activation of a given olfactory receptor by an OR ligand. As used herein, an "enhancer" of an olfactory receptor refers to a compound that increases the activation of a given olfactory receptor by an OR ligand.
[0193] In some embodiments of the methods for identifying olfactory receptor ligands and methods for identifying enhancers or antagonists of olfactory receptors, the olfactory receptor is selected from the group consisting of OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2 V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)), OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25 (OR2A25 (S75N, A209P)), OR11G2 (or OR11G2 (I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3 (S73G)), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2 (S84N)), OR10A3, OR10A6 (preferably OR10A6 (A117V, V140G, L287P)), OR10J1 (preferably OR10J1 (M51I, I92M)), OR2J2, and OR2AG2 (preferably OR2AG2 (Y28C)). For these receptors, it is particularly interesting to identify olfactory receptor ligands and olfactory receptor enhancers. In some embodiments of the methods for identifying olfactory receptor ligands and methods for identifying olfactory receptor enhancers or antagonists, the olfactory receptor is OR5A2, OR5A1, OR7A17, OR7C1, OR8K3, OR1N2, OR10J5, OR5B12 or OR10H5.
[0194] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1, and OR4S2, preferably OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and OR5V1. OR2M2 and OR2V1 are more preferred. Of these two, OR2M2 is more preferred.
[0195] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR2M2 or OR2V1, preferably OR2M2. In this situation, optionally, step b) further comprises contacting the receptor or recombinant host cell with a copper salt. Thus, in some embodiments, the present disclosure provides a method for identifying an olfactory receptor antagonist, the method comprising: a) providing an OR2M2 or OR2V1 olfactory receptor protein as described herein, or a cell expressing an OR2M2 or OR2V1 olfactory receptor protein as described herein; b) contacting the receptor or cell with a cognate ligand, a test compound or composition, and a copper salt; and c) Detecting increased or decreased activation of olfactory receptors compared to a ligand-only control.
[0196] In some embodiments, copper salts may be used at a concentration of 1-100 μM, preferably 10-100 μM, for example, 30 μM. Suitable copper salts include copper(II) salts such as CuCl, CuSO, Cu(OH), and copper acetate.
[0197] For such methods for identifying antagonists of OR2M2 or OR2V1, the cognate ligand is preferably selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, and 4-methoxy-2-methylpentane-2-thiol.
[0198] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR51B2, preferably OR51B2 (C120R, L134F, C209S). For such a method for identifying an OR51B2 antagonist, the cognate ligand is preferably 3-methyl-2-hexenoic acid.
[0199] In some embodiments of the method for identifying an enhancer or antagonist of an olfactory receptor, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR5V1. For such a method for identifying an OR5V1 antagonist, the cognate ligand is preferably 2,4,6-trichloroanisole.
[0200] As previously described herein, the present disclosure encompasses libraries containing a diverse repertoire of olfactory receptor proteins, nucleic acid molecules or expression vectors expressing olfactory receptor proteins, and recombinant host cells expressing olfactory receptor proteins, preferably wherein each of the olfactory receptor proteins shares the same modified C-terminal domain. Advantageously, such a library of olfactory receptors with identical C-terminal domains provides homogeneous and uniform functional expression of all receptors when detecting the activation of the olfactory receptors.
[0201] Thus, in one aspect, there is provided a method for generating an objective representation of the olfactory properties of a test compound or composition, said method comprising: a) providing a library as described herein; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a test compound or composition; and d) Detecting the activation of each of the olfactory receptor proteins. Also encompassed is the use of libraries as described herein to generate an objective representation of the olfactory properties of a test compound or composition.
[0202] This objective expression of olfactory characteristics can be called the "fingerprint" of OR activation. The activation level of each olfactory receptor can be expressed as an n-dimensional vector, where n is the number of distinct olfactory receptors contained in the library. More information on how to measure and express the activation level of olfactory receptors is provided elsewhere in this specification.
[0203] This type of objective expression of olfactory properties also allows for the comparison of olfactory properties between two or more test compounds or compositions in an objective manner.
[0204] Thus, in one aspect, there is provided a method for assessing the differences or similarities between two or more test compounds or compositions, said method comprising: a) generating an objective representation of the olfactory properties of two or more test compounds or compositions as described herein; and b) Comparing objective expressions of olfactory properties between two or more test compounds or compositions.
[0205] In some embodiments, a method is provided for assessing the differences or similarities between two or more test compounds or compositions, said method comprising: a) providing a library as described herein; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of two or more test compounds or compositions; d) detecting activation of each of the olfactory receptor proteins for each of two or more test compounds or compositions; and e) Comparing the activated olfactory receptor protein between each of two or more test compounds or compositions. Also encompassed is the use of libraries as described herein to assess the differences or similarities between two or more test compounds or compositions.
[0206] In some embodiments, the two or more test compounds or compositions may comprise a first test composition and a second test composition. In some embodiments, the second composition lacks one or more compounds present in the first composition, but otherwise comprises the same compounds as the first composition; or the second composition contains one or more alternative compounds for one or more compounds present in the first composition, but otherwise comprises the same compounds as the first composition.
[0207] In some embodiments, the final step of comparing the objective expression of the olfactory properties of activated olfactory receptor proteins between two or more test compounds or compositions involves calculating a distance measurement.Suitable distance measurement methods are known to those skilled in the art.As an example, the activation level of each olfactory receptor may be represented as an n-dimensional vector, and the distance measure may be a measure of the distance between two vectors.In some embodiments, the distance between two vectors may be based on the so-called 1-norm or L1-norm.The L1-norm is a standard measure in mathematics and is calculated as the sum of the absolute values of vectors.The distance between two vectors can then be calculated based on the norm of their difference.Thus, the distance between a first test compound or composition and a second test compound or composition may be calculated according to the following formula:
number
[0208] This distance is also known as the Euclidean distance. More information on how to measure and express the level of olfactory receptor activation is provided elsewhere in this specification.
[0209] In some embodiments, the test compound or composition involved in the method of the present disclosure is a mixture of odorants.In fact, the increased sensitivity of the method disclosed herein allows the detection of ligands, enhancers and antagonists from complex samples, even in the presence of strong odorant matrices.In some embodiments, the test compound or composition is a fragrance composition.In some embodiments, the test compound or composition is a composition such as a perfume composition, which comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45 or at least 50 odorants.
[0210] In some embodiments, the test compounds or compositions involved in the methods of the present disclosure are unpurified synthetic compounds. Automated parallel synthesis of single compounds has dramatically increased the number of new organic compounds available for biological assays. However, a disadvantage of automated parallel synthesis is that a difficult and expensive subsequent purification step is usually required before they can be used in biological assays. This is particularly true in sensory testing, where the olfactory evaluation of new fragrance or flavor ingredients is difficult due to the presence of odorous impurities that cause off-flavors or mask the odor of the desired molecule. Advantageously, the increased sensitivity of the methods disclosed herein allows the use of unpurified synthetic compounds. Thus, in some embodiments, the test compounds or compositions involved in the methods of the present disclosure include unpurified synthetic molecules. In this case, the synthetic molecules may have a purity level of less than 90%, less than 85%, less than 80%, less than 75%, or less than 70% (the % purity level is the ratio of the desired product to the combined impurities, typically measured by liquid chromatography, gas chromatography, or quantitative NMR).
[0211] In accordance with the above, in some embodiments, the test compound or composition involved in the methods of the present disclosure is a mixture of odorants (as described above) or an unpurified synthetic compound (as described above).
[0212] In some embodiments, the test compound or composition involved in the method of the present disclosure is a racemic mixture of odorants. In some embodiments, the test compound or composition involved in the method of the present disclosure is an isolated or synthetic isomer or enantiomer of such a racemic mixture.
[0213] As described elsewhere, olfactory receptors have also been found to be associated with various diseases.Therefore, in some embodiments, the test compound or composition may comprise a therapeutic candidate, such as an anti-cancer drug candidate.In some embodiments, the test compound or composition may comprise a pharmacologically active agent.
[0214] The ligands, enhancers, and antagonists described herein are preferably biodegradable. Indeed, there is growing interest in this field in moving toward biodegradable components. Thus, in some embodiments, the test compounds or compositions involved in the methods of the present disclosure are biodegradable compounds or compositions. As used herein, a compound or composition is considered biodegradable if it meets the acceptance criteria according to the OECD manometric respirometry method, and in particular the OECD 301F method, which is well known in the art. In this method, the acceptance level for a compound to be considered as having "the ability to be readily biodegraded" or "easily biodegradable" is to reach 60% of the theoretical oxygen demand and / or chemical oxygen demand. This acceptance value must be reached within a 10-day window of the 28-day test period. The 10-day window begins when the degree of biodegradation reaches 10% of the theoretical oxygen demand and / or chemical oxygen demand, and must end before the 28th day of the test. If a test for the ability to be readily biodegraded yields a positive result, it can be assumed that the compound will undergo rapid and eventual biodegradation in the environment (OECD 25: Guidelines for the Testing of Chemicals, Chapter 3, Part 1, Preface: Principles and Strategies Related to the Testing of Degradation of Organic Chemicals, adopted July 2003). The assessment of "essentially biodegradable" can also be performed using OECD Method 301F, although the pass criteria are different. More specifically, the pass criteria are 60% of the theoretical oxygen demand and / or chemical oxygen demand. This pass value can also be achieved after a 28-day test period, which is usually extended to 60 days. The 10-day window does not apply.
[0215] In some embodiments, the test compound or composition involved in the disclosed methods is a compound or composition that is readily biodegradable. In some embodiments, the test compound or composition involved in the disclosed methods is an inherently biodegradable compound or composition.
[0216] In one aspect, a method for identifying an olfactory receptor capable of binding to a target ligand is provided, said method comprising: a) providing a library as described herein; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a target ligand; and d) Identifying the olfactory receptors activated by target ligands.
[0217] All methods described herein involve detecting the activation of olfactory receptors. Detecting the activation of olfactory receptors may refer to measuring the level of activation of olfactory receptors. Thus, throughout this disclosure, "detecting activation" may be replaced with "measuring the level of activation" or similar expressions. Means and methods for detecting the activation of olfactory receptors are generally available in the art.
[0218] For example, as described in detail in the experimental section herein, detecting the activation of olfactory receptors can involve the co-introduction or use of a luciferase gene operably linked to a cAMP-responsive promoter / element (Saito et al. (2004) Cell 119(5): 679-691, incorporated herein by reference in its entirety), which is used as a reporter gene. Activation of olfactory receptors and the subsequent increase in intracellular cAMP result in the expression of luciferase. Cleavage of luciferin by luciferase results in luminescence, which can then be detected and quantified.
[0219] To measure and report the (level of) olfactory receptor activation, for example, the luciferase fold induction can be calculated. Typically, the induction ratio is obtained relative to a solvent-only control. Usually, a background control containing all other reagents but no cells is also obtained, and the value of this control will be subtracted from all measurements to account for background luminescence. For the solvent-only control, cells expressing OR and luciferase genes are treated with solvent only (excluding test compounds or compositions). All values from experiments using test compounds and compositions are then divided by the average value of these solvent-only control measurements to calculate the luciferase induction ratio. Thus, solvent controls and test compounds and compositions that do not cause OR activation will obtain a value of 1, indicating no luciferase induction. A value significantly greater than 1 indicates luciferase gene activation and thereby enhanced cAMP production due to OR activation.
[0220] Once a potent ligand for a given OR and the concentration of that ligand for maximal OR activation are known, this ligand (tested at its maximal induction concentration) can be introduced into the experiment as a positive control, and the luciferase ratio can be reported as % of the activation of the positive control according to the following formula:
number
[0221] Based on this calculation, the potency of the ligands can then be compared by applying a sigmoidal curve fit according to the Hill equation and calculating, for example, the EC20% or EC50% values, i.e., the concentrations that result in 20% or 50% activation of the OR compared to the positive control.
[0222] In experiments without a positive control, such as experiments using a library of receptors, other types of normalization can be used. Thus, since the dynamic range (maximum efficacy) of different receptors can vary considerably, it may be appropriate to use the logarithm of the induction ratio to report the data. Other options are normalization to the highest induction ratio for each receptor when multiple samples are tested, or normalization to the historical value of maximum efficacy for a given OR.
[0223] Other receptor genes that can be coupled to cAMP-responsive promoters / elements include green fluorescent protein. There are also several methods for directly measuring changes in intracellular cAMP concentration based on the binding of antibodies to cAMP. Another method for detecting OR activation involves coupling a hybrid G protein to an OR, which activates calcium release from intracellular stores. Changes in calcium concentration can then be measured with a chemiluminescent probe sensitive to changes in calcium concentration, or with recombinant fluorescent or luminescent proteins that can sense differences in calcium concentration.
[0224] Further means and methods for detecting the activation of olfactory receptors are known to those skilled in the art, and include, among others, GTPase / GTP binding assays, aequorin-based assays, fluorescence-based assays, membrane depolarization assays, melanophore assays, PKC activation assays, PKA activation assays, and kinase assays, as described in WO2019 / 110630, the entire contents of which are incorporated herein by reference. An overview of some methods for detecting the activation of olfactory receptors in heterologous cells by odorants is provided in the review by Veithen et al., 2017, Springer Handbook of Odor Chapter 22.2, Springer International Publishing (CH), the entire contents of which are incorporated herein by reference.
[0225] The ligand and / or test compound or composition as described herein can be added to existing culture, or the medium of existing culture can be replaced with fresh medium containing the ligand.Suitable ligands can be selected from any chemical compounds known in the art that can activate olfactory receptors (alternatively referred to as "fragrance compounds" or "odorants"), which are discussed in standard handbooks such as Buettner (2017), Springer Handbook of Odor, Springer International Publishing (CH), the entire contents of which are incorporated herein by reference.Suitable compounds can also be found in publicly available databases, such as "OlfactionBase" available at https: / / olfab.iiita.ac.in / olfactionbase / and discussed in Sharma et al. OlfactionBase: a repository to explore odors, odorants, olfactory receptors and odorant-receptor interactions. Nucleic Acids Res. 2022 Jan 7;50(D1):D678-D686.
[0226] Non-limiting examples of suitable ligands include esters (e.g., geranyl acetate, methyl formate, methyl acetate, methyl propionate, methyl butyrate, ethyl acetate, ethyl butyrate, isoamyl acetate, pentyl butyrate, pentyl valerate, octyl acetate, benzyl acetate, methyl anthranilate, hexyl acetate), linear terpenes (e.g., myrcene, geraniol, nerol, citral, citronellal, citronellol, linalool, nerolidol, ocimene), cyclic terpenes (e.g., limonene, camphor, methol, carboplatin, methyl methyl methacrylate ...methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methacrylate, methyl methyl methyl methyl methacrylate, methyl methyl methyl methyl methacrylate, methyl methyl methyl methyl methacrylate, methyl methyl methyl methyl methyl methyl methyl methyl methyl bon, terpineol, alpha-ronone, thujone, eucalyptol, jasmine), aromatic compounds (e.g., benzaldehyde, eugenol, isoeugenol, cinnamaldehyde, ethyl maltol, ethyl vanillin, anisole, anethole, estragole, thymol), amines (e.g., trimethylamine, putrescine, cadaverine, pyridine, indole, skatole), alcohols (e.g., furaneol, 1-hexanol, ethanol), aldehydes (e.g., acetonitrile, acetone, methylparaben ... Aldehydes (e.g., hexanal, furfural, hexylcinnamaldehyde, isovaleraldehyde, anisaldehyde, cuminaldehyde), ketones (e.g., dihydrojasmone, 2-acetyl-1-pyrroline, 6-acetyl-2,3,4,5-tetrahydropyridine), lactones (e.g., gamma-decalactone, gamma-nonalactone, delta-octalactone, jasmine lactone, massoialactone, wine lactone, sotolon), thiols (e.g., thioacetone, allylthiol , ethanethiol, 2-methyl-2-propanethiol, butane-1-thiol, mercaptan, methanethiol, furan-2-ylmethanethiol, benzyl mercaptan), musks (e.g., nitromusk, polycyclic musks, macrocyclic musks, linear / alicyclic musks, musk ketone, musk ambrette, musk mosquene, musk tibetene, musk xylene), cresols (e.g., vanilla cresol (ultra vanilla)), propenyl guaethol (vanitrope), carboxylic acids, and the like.
[0227] Preferred ligands include those having musky, woody, lily-of-the-valley, floral, green, balsamic, spicy, or fruity notes. "Musky," "woody," "lily-of-the-valley," "floral," "green," "balsamic," "spicy," and "fruity" are terms recognized in the art in the context of odorant molecules. Ligands with musky, woody, lily-of-the-valley, floral, green, balsamic, spicy, and / or fruity notes can be identified by those skilled in the art in publicly available databases, such as "OlfactionBase" (available at https: / / olfab.iiita.ac.in / olfactionbase / and discussed in Sharma et al., "OlfactionBase: a repository to explore odors, odorants, olfactory receptors and odorant-receptor interactions. Nucleic Acids Res. 2022 Jan 7;50(D1):D678-D686," incorporated herein by reference in its entirety). Related musk compounds are also described in WO2019 / 11630, incorporated herein by reference in its entirety.
[0228] Further examples of preferred ligands are Ambermax®, para-cresol, menthol, (S)-menthol, menthone, Mahonial®, Nympheal®, linalool, androstenone, androstenol, cyclopentanethiol, Hedione®, Hedione® HC, Ambrofix®, calone, 4-ethyloctanoic acid, Galaxolide®, Galaxolide® S, ethyl vanillin, beta-ionone, ambrettolide, 3-methyl-3-hydroxy-hexanoic acid, 3-methyl-2-hexenoic acid, nonanoic acid, decanoic acid, undecanoic acid, ethyl 3-mercaptopropionate, diallyl disulfide, benzothiazole, 2-methyl-3-tetrahydrofuranthiol, muscone, dipropyl disulfide, musk ketone, Arborone, Georgywood®, iso E Super, Cedrol, Mercapto-8-p-menthan-3-one, 2-Naphthalenethiol, 3-(Methylthio)propionaldehyde, 1,3-Propanedithiol, Benzothiazole, Allyl Sulfide, Allyl Mercaptan, Hydroxyethylmethylthiazole, Dimethyl Trisulfide, Thioglycolic Acid, 3-Mercapto-2-pentanone, 2-((Methyldisulfanyl)methylfuran), Bis(methylthio)methane, Ethyl 2-Mercaptopropionate, Methyl Thiobutyrate, 3-Methoxypropanol Mercapto-3-methylbutyl formate, methyl 3-mercaptopropionate, butyl 3-mercaptopropionate, 3-mercaptopropionic acid, dimethyl disulfide, 2-mercaptopropionic acid, 2-methyl-3-furanthiol, benzyl mercaptan, 2-mercapto-2-methyl-1-pentanol, allyl isothiocyanate, 2-mercaptobutanone, 2-heptanethiol, 2-methyl-3-tetrahydrofuranthiol, cis-2-isobutyl-4,5-dimethyl-2,5-Dihydrothiazole, 3-methyl-3-sulfanylhexan-1-ol, (rac)-3-mercapto-2-methyl-1-pentanol, 1-hexanethiol, cyclopentanethiol, 2-methyl-2-propanethiol, 2-methyl-3-(methyldithio)furan, 2-methyl-3-buten-1-ol, sodium methanethiolate, sodium hydrosulfide hydrate, 4-methoxy-2-methylpentane-2-thiol, blackcurrant body body), furfuryl mercaptan, Anjeruk®, 3-mercaptohexyl acetate, methoxymethyl butanethiol, 3-mercaptohexanol, dimethyl sulfide, acetylthiazole, mercapto-8-methene-1-para, 2-methyl-3-sulfanyl-pentanol, (E,S)-3,7-dimethylnon-6-en-1-ol, 3-mercapto-3-methylhexan-1-ol, 7-(3-methylbutyl)-benzo[b][1,4]dioxepin-3-one, benzyl salicylate, delta-damascone, cyclohexanecarbo Ethyl phosphate, geosmin, 3-(4-isobutyl-2-methylphenyl)propanal, patchoulol, Peonile®, rotundone, heliotropin, methyl salicylate, Galbanone®, Javanol, Timberol®, trans-2,cis-6-nonadienal, Esterly, Cascalone®, Azurone®, Rosabloom™, Rosyfolia®, β-ionone, sotolone, 3-methyl-3-hydroxyhexanoic acid, 2-methylundecanoic acid.
[0229] Those skilled in the art will appreciate that the amount of ligand required for activation of an olfactory receptor may vary depending on the olfactory receptor and the ability of the ligand to physically associate with said olfactory receptor. A ligand may be selected such that it has an EC 50 EC values, typically between 1 nM and 1 mM 50An olfactory receptor is considered to be "specific" for a given olfactory receptor (i.e., specific for that receptor) if it is capable of physically associating (i.e., binding) with said receptor at a value of 0.01. 50 means the concentration of a ligand at which a given activation of an olfactory receptor is 50% of the maximum for that olfactory receptor, which can be measured using methods as described elsewhere herein.
[0230] Aspects and Aspects Aspects and embodiments of the present invention are described in the following numbered paragraphs, which form an integral part of this specification. Each of the features, aspects and embodiments described below is further described in the description above.
[0231] 1. An olfactory receptor protein having a modified C-terminal domain in which at least 32% of the amino acids are lysine, arginine, or histidine.
[0232] 2. An olfactory receptor protein having a modified C-terminal domain in which at least 35% of the amino acids are lysine, arginine, or histidine.
[0233] 3. The olfactory receptor protein described in paragraph 1 or 2, having a modified C-terminal domain in which at least 38% of the amino acids are lysine, arginine, or histidine.
[0234] 4. An olfactory receptor protein described in any of the preceding paragraphs, having a modified C-terminal domain containing the amino acid sequence motif RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1).
[0235] 5. The olfactory receptor protein of any of the preceding paragraphs, wherein the modified C-terminal domain is fused to the seventh transmembrane helix of the protein.
[0236] 6. The olfactory receptor protein described in any of the preceding paragraphs, which is a class I or class II olfactory receptor having a modified C-terminal domain, preferably a human, canine or feline class I or class II olfactory receptor having a modified C-terminal domain, more preferably a human class I or class II olfactory receptor having a modified C-terminal domain.
[0237] 7. Proteins include OR10A2, OR10A3, OR10A4, OR10A5, OR10A6, OR10A7, OR10AD1, OR10AG1, OR10C1, OR10D3, OR10G2, OR10G3, OR10G4, OR10G6, OR10G7, OR10G8, OR10G9, OR10H1, OR10H2, OR10H3, OR10H4, OR10H5, OR10J1, OR10J3, OR10J5, OR10K1, OR10K2, OR10P1, OR10P2, OR10Q1, OR10R2, OR10S1, OR10T2, and OR10V1. , OR10W1, OR10X1, OR10Z1, OR11A1, OR11G2, OR11H1, OR11H2, OR11H4, OR11H6, OR11L1, OR12D2, OR12D3, OR13C2, OR13C3, OR13C4, OR13C5, OR13C8, OR13 C9, OR13D1, OR13F1, OR13G1, OR13H1, OR13J1, OR14A16, OR14A2, OR14C36, OR14I1, OR14L1P, OR1A1, OR1A2, OR1B1, OR1C1, OR1D2, OR1D4, OR1D5, OR1E1, O R1E2, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L6, OR1L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1S1, OR1S2, OR2A 1, OR2A12, OR2A14, OR2A2, OR2A25, OR2A4, OR2A42, OR2A5, OR2A7, OR2A9P, ORAE1, OR2AG1, OR2AG2, OR2AJ1, OR2AK2, OR2AP1, OR2AT4, OR2B11, OR2B2, OR 2B3, OR2B6, OR2B8P, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2F2, OR2G2, OR2G3, OR2G6, OR2H1, OR2H2, OR2J1P, OR2J2, OR2J3, OR2K2, OR2L13, OR2L2, OR2 L3, OR2L5, OR2L8, OR2M2, OR2M3, OR2M4, OR2M5, OR2M7, OR2S2, OR2T1, OR2T10, OR2T11, OR2T12, OR2T2, OR2T27, OR2T29, OR2T3, OR2T33, OR2T34, OR2T35,OR2T4OR2T5、OR2T6、OR2T7、OR2T8、OR2V1、OR2V2、OR2W1、OR2W3、OR2W5、OR2Y1、OR2Z1、OR3A1、OR3A2、OR3A3、OR3A4、OR4A15、OR4A16、OR4A4、R4A47、OR4A5、OR4B1、OR4C11、OR4C12、OR4C13、OR4C15、OR4C16、OR4C、OR4C45、OR4C46、OR4C5、OR4C6、OR4D1、OR4D10、OR4D11、OR4D2、OR4D5、OR4D6、OR4D9、OR4E2、OR4F15、OR4F16、OR4F17、OR4F21、OR4F29、OR4F3、OR4F4、OR4F5、OR4F6、OR4K1、OR4K13、OR4K14、OR4K15、OR4K17、OR4K2、OR4K3P、OR4K5、OR4L1、OR4M1、OR4M2、OR4N2、OR4N4、OR4N5、OR4P4、OR4Q3、OR4S1、OR4S2、OR4X1、OR4X2、OR51A2、OR51A4、OR51A7、OR51B2、OR51B4、OR51B5、OR51B6、OR51D1、OR51E1、OR51E2、OR51F1、OR51F2、OR51G1、OR51G2、OR51H1P、OR51I1、OR51I2、OR51J1、OR51L1、OR51M1、OR51Q1、OR51S1、OR51T1、OR51V1、OR52A1、OR52A4、OR52A5、OR52B2、OR52B4、OR52B6、OR52D1、OR52E2、OR52E4、OR52E5、OR52E6、OR52E8、OR52H1、OR52I1、OR52I2、OR52J3、OR52K1、OR52K2、OR52L1、OR52M1、OR52N1、OR52N2、OR52N4、OR52N5、OR52P1P、OR52R1、OR52W1、OR56A1、OR56A3、OR56A4、OR56A5、OR56B1、OR56B4、OR5A1、OR5A2、OR5AC2、OR5AK2、OR5AL1P、OR5AN1、OR5AP2、OR5AR1、OR5AS1、OR5AU1、OR5B12、OR5B17、OR5B2、OR5B21、OR5B3、OR5C1、OR5D13、OR5D14、OR5D16、OR5D18、OR5F1、OR5H1、OR5H14、OR5H15、OR5H2, OR5H6, OR5I1, OR5J2, OR5K1, OR5K2, OR5K3, OR5K4, OR5L1, OR5L2, OR5M1, OR5M10, OR5M11, OR5 M3, OR5M8, OR5M9, OR5P2, OR5P3, OR5R1, OR5T1, OR5T2, OR5T3, OR5V1, OR5W2, OR6A2, OR6B1, OR6B2, OR6 B3, OR6C1, OR6C2, OR6C3, OR6C4, OR6C6, OR6C65, OR6C68, OR6C70, OR6C74, OR6C75, OR6C76, OR6F1, OR 6J1, OR6K2, OR6K3, OR6K6, OR6M1, OR6N1, OR6N2, OR6P1, OR6Q1, OR6S1, OR6T1, OR6X1, OR6Y1, OR7A10, O R7A17, OR7A5, OR7C1, OR7C2, OR7D2, OR7D4, OR7E24, OR7G1, OR7G2, OR7G3, OR8A1, OR8B12, OR8B2, OR8 B3, OR8B4, OR8B8, OR8D1, OR8D2, OR8D4, OR8G1, OR8G5, OR8H1, OR8H2, OR8H3, OR8I2, OR8J1, OR8J3, OR8 The olfactory receptor protein of any of the preceding paragraphs is a human class II or class I olfactory receptor selected from the group consisting of OR8K1, OR8K3, OR8K5, OR8S1, OR8U1, OR8U8, OR8U9, OR9A2, OR9A4, OR9G1, OR9G4, OR9G9, OR9I1, OR9K2, OR9Q1, OR9Q2, or variants thereof.
[0238] 8. Class II receptors: OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5M3 , OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2, preferably wherein the class II receptor is selected from the group consisting of OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2, and OR5A2.
[0239] 9. The olfactory receptor protein of any one of paragraphs 1 to 7, wherein the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1.
[0240] 10. An olfactory receptor protein described in any one of paragraphs 4 to 9, wherein the sequence motif is RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (sequence number 5).
[0241] 11. The sequence motif is: RNX1X2X3X4X 5'''' AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 311), RNX1EX 3' X4X 5'''' AX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 312), RNKEVKX 5'''' ALKRLLKRK (SEQ ID NO: 319), RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 820), RNX1EX 3' X4KAX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 821), and; RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 824), wherein: - X1 is K or R; - X2 is E or D or Q; - X3 is V or M or I or L; - X 3' is V or M or I; - X4 is K or R; - X5 is any amino acid; - X 5'''' is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y; - X6 is L or I or V; - X7 is K or R or H; - X 7' is K or R; - X8 is K or R; - X 10 is L or I or F; - X 11 is K or R or G; - X 11' is K or R; - X 12 is K or R; and - X 13 is K or R, 11. The olfactory receptor protein according to any one of paragraphs 4 to 10.
[0242] 12. An olfactory receptor protein described in any one of paragraphs 4 to 11, wherein x or X5 is not proline.
[0243] 13. An olfactory receptor protein described in any one of paragraphs 4 to 11, wherein x is not proline or tryptophan.
[0244] 14.x, X5, or X 5'''' is selected from D, K, R, E, N, V, A, Q or G, and preferably x, X5 or X 5'''' is selected from D, K, R, E, N, V, A, or Q.
[0245] 15. The olfactory receptor protein of any one of paragraphs 4 to 14, wherein the amino acid sequence motif comprises 1 to 6 additional C-terminal amino acid residues, and optionally wherein: - the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, more preferably the first additional amino acid residue is C, R or K, most preferably C; - the second additional amino acid residue is selected from C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably C or R; - the third additional amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K; - the fourth, fifth and sixth additional amino acid residues are selected from K and R; The olfactory receptor protein.
[0246] 16. The amino acid sequence motif is: - CC, SI, YP, PQ, FR, CR, RR, EK, PR, CG, FK, RG, RC, RT, RF, GG, YR, GC, TG, PC, HP, PG, KY, CP, YY, FF, CF, NP, YL, IC, HC, CL, YC, ER, RP, PA, FC, RY; - CRR, CCC, CCF, CCL, CCM, CCS, CCP, CCA, CCY, CCH, CCN, CCD, CCK, CCR, CCG; - CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CCRR (SEQ ID NO: 161), CCRK (SEQ ID NO: 228), CCKR (SEQ ID NO: 229), - CRRRR (SEQ ID NO: 162), CCRRR (SEQ ID NO: 163), CCKRR (SEQ ID NO: 230), CCRKR (SEQ ID NO: 231), CCRRK (SEQ ID NO: 232), CCRKK (SEQ ID NO: 233), CCKRK (SEQ ID NO: 234), CCKKR (SEQ ID NO: 235), CCKKK (SEQ ID NO: 236), CRRRRR (SEQ ID NO: 164), CRRRKK (SEQ ID NO: 165), and CCRRRR (SEQ ID NO: 224) and an additional C-terminal amino acid residue selected from the group consisting of: Preferably, wherein the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164), and CRRRKK (SEQ ID NO: 165), 16. The olfactory receptor protein of paragraph 15.
[0247] 17. The sequence motif is: RNKEVKX 5'''' ALKRLLKRKCC (SEQ ID NO: 320), RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 CC (SEQ ID NO: 822), RNX1EX 3’ X4KAX6X 7’ X8LX 10 X 11’ X 12 X 13 CC (SEQ ID NO: 823), RNX1X2X3X4X5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 825), and RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 826) 17. The olfactory receptor protein of paragraph 15 or 16, selected from the group consisting of:
[0248] 18. The sequence motif is: RNKEVKDALKRLLKRK (SEQ ID NO: 10), RNREVKDALKRLLKRK (SEQ ID NO: 11), RNKEIKDALKRLLKRK (SEQ ID NO: 12), RNKEMKDALKRLLKRK (SEQ ID NO: 13), RNKEVRDALKRLLKRK (SEQ ID NO: 14), RNKEVKKALKRLLKRK (SEQ ID NO: 15), RNKEVKEALKRLLKRK (SEQ ID NO: 16), RNKEVKNALKRLLKRK (SEQ ID NO: 17), RNKEVKRALKRLLKRK (SEQ ID NO: 18), RNKEVKVALKRLLKRK (SEQ ID NO: 19), RNKEVKAALKRLLKRK (SEQ ID NO: 20), RNKEVKQALKRLLKRK (SEQ ID NO: 21), RNKEVKDAVKRLLKRK (SEQ ID NO: 22), RNKEVKDAIKRLLKRK (SEQ ID NO: 23), RNKEVKDALRRLLKRK (SEQ ID NO: 24), RNKEVKDALKKLLKRK (SEQ ID NO: 25), RNKEVKDALKRLIKRK (SEQ ID NO: 26), RNKEVKDALKRLFKRK (SEQ ID NO: 27), RNKEVKDALKRLLRRK (SEQ ID NO: 28), RNKEVKDALKRLLKKK (SEQ ID NO: 29), RNKEVKDALKRLLKRR (SEQ ID NO: 30), RNKDVKDALKRLLKRK (SEQ ID NO: 31), RNKQVKDALKRLLKRK (SEQ ID NO: 32), RNKELKDALKRLLKRK (SEQ ID NO: 33), RNKEVKGALKRLLKRK (SEQ ID NO: 34), RNKEVKDALHRLLKRK (SEQ ID NO: 35), RNKEVKDALKRILKRK (SEQ ID NO: 36), RNKEVKDALKRLLGRK (SEQ ID NO: 37), RNKEVKRAIKRLLKRK (SEQ ID NO: 38), RNKEVKKAIKRLLKRK (SEQ ID NO: 39), RNKEVKRAIKRLFKRK (SEQ ID NO: 40), RNKEVKKAIKRLFKRK (SEQ ID NO: 41), RNKEVKRAIRKLLKRK (SEQ ID NO: 42), RNKEVKDALRKLLKRK (SEQ ID NO: 43), RNKEVKDALKRLLRRR (SEQ ID NO: 44), RNKEVKRALKRLLRRR (SEQ ID NO: 45), RNKEVKKALKRLLRRR (SEQ ID NO: 46), RNREVKRAIKRLLKRK (SEQ ID NO: 47), RNREVKKAIKRLLKRK (SEQ ID NO: 48), RNREVKRAIKRLFKRK (SEQ ID NO: 49), RNREVKKAIKRLFKRK (SEQ ID NO: 50), RNREVKRAIRKLLKRK (SEQ ID NO: 51), RNREVKDALRKLLKRK (SEQ ID NO: 52), RNREVKDALKRLLRRR (SEQ ID NO: 53), RNKEVKKAIKRLLRRK (SEQ ID NO: 54), RNKEVKKAIKRLLKKK (SEQ ID NO: 55), RNKEVKKAIKRLLKRR (SEQ ID NO: 56), RNKEVKRAIKRLLRRK (SEQ ID NO: 57), RNKEVKRAIKRLLKKK (SEQ ID NO: 58), RNKEVKRAIKRLLKRR (SEQ ID NO: 59), RNKEVKKAIKRLFRRK (SEQ ID NO: 60), RNKEVKKAIKRLFKKK (SEQ ID NO: 61), RNKEVKKAIKRLFKRR (SEQ ID NO: 62), RNKEVKRAIKRLFRRK (SEQ ID NO: 63), RNKEVKRAIKRLFKKK (SEQ ID NO: 64), RNKEVKRAIKRLFKRR (SEQ ID NO: 65), RNREVKRAIKRLLRKK (SEQ ID NO: 66), RNREVKKAIKRLLRKK (SEQ ID NO: 67), RNREVKRAIKRLFRRR (SEQ ID NO: 68), RNREVKKAIKRLFRRR (SEQ ID NO: 69), RNREVKKAIKRLFRRK (SEQ ID NO: 70), RNREVKKAIKRLFKKK (SEQ ID NO: 71), RNREVKKAIKRLFKRR (SEQ ID NO: 72), RNREVKRAIKRLFRRK (SEQ ID NO: 73), RNREVKRAIKRLFKKK (SEQ ID NO: 74), RNREVKRAIKRLFKRR (SEQ ID NO: 75), RNREMRKALHRLLGKK (SEQ ID NO: 827), RNREVKKAIHKLIGRK (SEQ ID NO: 828), RNREVRKAVHRLFKRK (SEQ ID NO: 829), RNKEMKKAIHKLFGKK (SEQ ID NO: 830), RNRDVKKAVHKLFRRK (SEQ ID NO: 831), RNRDMKKAVHKLFGKR (SEQ ID NO: 832), RNKELRKALHKLLGRK (SEQ ID NO: 833), RNRDVRKALRRILRRR (SEQ ID NO: 834), RNKDVRKAVRKLIRRR (SEQ ID NO: 835), RNRDVRKAVRRLFRKR (SEQ ID NO: 836), RNKDIKKAVKKLIKKK (SEQ ID NO: 837), RNRELRKAVRRLFKRR (SEQ ID NO: 838), RNKELRKAVRKIIKKK (SEQ ID NO: 839), RNRDVKKAVRRLFRRK (SEQ ID NO: 840), RNREVRKALRRIIRKR (SEQ ID NO: 841), RNKDIRKAVKKIFRRK (SEQ ID NO: 842), RNKDVRKAVRRLIKRK (SEQ ID NO: 843), RNRDLRKAVRKLFKKK (SEQ ID NO: 844), RNRDLRKALRRIFKRR (SEQ ID NO: 845), RNRDVRKAIKKLIRKR (SEQ ID NO: 846), RNKELKKAIKRILKKK (SEQ ID NO: 847), RNRDVRKAIRKLLKRK (SEQ ID NO: 848), RNRDLRKAVRRIFKKR (SEQ ID NO: 849), RNRDVRKAVRKLFKRR (SEQ ID NO: 850), RNRDVRKALRRLFKKR (SEQ ID NO: 851), RNKELKKALRKLIGKK (SEQ ID NO: 852), RNREMRKAIKKIIKKK (SEQ ID NO: 853), RNKEIKKAIKKIIKKR (SEQ ID NO: 854), RNRDVKKAIRRLFRRR (SEQ ID NO: 855), RNREVKKAVKKLIGKR (SEQ ID NO: 856), RNREMRKALRRLFRKR (SEQ ID NO: 857), RNKELKKALRRLIGRR (SEQ ID NO: 858), RNRDVKKALRKLIGKR (SEQ ID NO: 859), RNREVKKAVKKLIRRK (SEQ ID NO: 860), RNKEVRKALKKLFGKK (SEQ ID NO: 861), RNKEIRKALRRLFGKK (SEQ ID NO: 862), RNKDVKKALRRLFGKK (SEQ ID NO: 863), RNKELKKAIKRLIRRK (SEQ ID NO: 864), RNKDVRKAVKRLLKKR (SEQ ID NO: 865), RNKELRKAIRRLLRRR (SEQ ID NO: 866), RNRDIRKALRKLFKKK (SEQ ID NO: 867), RNRELKKALRRLLRRR (SEQ ID NO: 868), RNREVKKALRRLFGKK (SEQ ID NO: 869), RNRDVRKALKRLLKRK (SEQ ID NO: 870), RNRDMRKAIRKLFGRK (SEQ ID NO: 871), RNRELKKAIRKLLKRK (SEQ ID NO: 872), RNRDIRKAVKKLFGKK (SEQ ID NO: 873), RNKEVKKAIRKLFGRR (SEQ ID NO: 874), RNREVRKAVRKLFRRK (SEQ ID NO: 875), RNRDMKKALKKLFRRR (SEQ ID NO: 876), RNRDVRKALKRLLGRR (SEQ ID NO: 877), RNKDLKKAVKKLFGRK (SEQ ID NO: 878), RNKDVRKAVRRLFGRR (SEQ ID NO: 879), RNKEVKCALKRLLKRK (SEQ ID NO: 880), RNKEVKFALKRLLKRK (SEQ ID NO: 881), RNKEVKHALKRLLKRK (SEQ ID NO: 882), RNKEVKIALKRLLKRK (SEQ ID NO: 883), RNKEVKLALKRLLKRK (SEQ ID NO: 884), RNKEVKMALKRLLKRK (SEQ ID NO: 885), RNKEVKSALKRLLKRK (SEQ ID NO: 886), RNKEVKTALKRLLKRK (SEQ ID NO: 887), RNKEVKWALKRLLKRK (SEQ ID NO: 888), RNKEVKYALKRLLKRK (SEQ ID NO: 889), RNKQIRDALKRLLKRK (SEQ ID NO: 890). RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), RNREVKDALKRLLKRKCC (SEQ ID NO: 87), RNKEIKDALKRLLKRKCC (SEQ ID NO: 88), RNKEMKDALKRLLKRKCC (SEQ ID NO: 89), RNKEVRDALKRLLKRKCC (SEQ ID NO: 90), RNKEVKKALKRLLKRKCC (SEQ ID NO: 91), RNKEVKEALKRLLKRKCC (SEQ ID NO: 92), RNKEVKNALKRLLKRKCC (SEQ ID NO: 93), RNKEVKRALKRLLKRKCC (SEQ ID NO: 94), RNKEVKVALKRLLKRKCC (SEQ ID NO: 95), RNKEVKAALKRLLKRKCC (SEQ ID NO: 96), RNKEVKQALKRLLKRKCC (SEQ ID NO: 97), RNKEVKDAVKRLLKRKCC (SEQ ID NO: 98), RNKEVKDAIKRLLKRKCC (SEQ ID NO: 99), RNKEVKDALRRLLKRKCC (SEQ ID NO: 100), RNKEVKDALKKLLKRKCC (SEQ ID NO: 101), RNKEVKDALKRLIKRKCC (SEQ ID NO: 102), RNKEVKDALKRLFKRKCC (SEQ ID NO: 103), RNKEVKDALKRLLRRKCC (SEQ ID NO: 104), RNKEVKDALKRLLKKKCC (SEQ ID NO: 105), RNKEVKDALKRLLKRRCC (SEQ ID NO: 106), RNKDVKDALKRLLKRKCC (SEQ ID NO: 107), RNKQVKDALKRLLKRKCC (SEQ ID NO: 108), RNKELKDALKRLLKRKCC (SEQ ID NO: 109), RNKEVKGALKRLLKRKCC (SEQ ID NO: 110), RNKEVKDALHRLLKRKCC (SEQ ID NO: 111), RNKEVKDALKRILKRKCC (SEQ ID NO: 112), RNKEVKDALKRLLGRKCC (SEQ ID NO: 113), RNKEVKRAIKRLLKRKCC (SEQ ID NO: 114), RNKEVKKAIKRLLKRKCC (SEQ ID NO: 115), RNKEVKRAIKRLFKRKCC (SEQ ID NO: 116), RNKEVKKAIKRLFKRKCC (SEQ ID NO: 117), RNKEVKRAIRKLLKRKCC (SEQ ID NO: 118), RNKEVKDALRKLLKRKCC (SEQ ID NO: 119), RNKEVKDALKRLLRRRCC (SEQ ID NO: 120), RNREMRKALHRLLGKKCC (SEQ ID NO: 254), RNREVKKAIHKLIGRKCC (SEQ ID NO: 255), RNREVRKAVHRLFKRKCC (SEQ ID NO: 256), RNKEMKKAIHKLFGKKCC (SEQ ID NO: 257), RNRDVKKAVHKLFRRKCC (SEQ ID NO: 258), RNRDMKKAVHKLFGKRCC (SEQ ID NO: 259), RNKELRKALHKLLGRKCC (SEQ ID NO: 260), RNRDVRKALRRILRRRCC (SEQ ID NO: 261), RNKDVRKAVRKLIRRRCC (SEQ ID NO: 262), RNRDVRKAVRRLFRKRCC (SEQ ID NO: 263), RNKDIKKAVKKLIKKKCC (SEQ ID NO: 264), RNRELRKAVRRLFKRRCC (SEQ ID NO: 265), RNKELRKAVRKIIKKKCC (SEQ ID NO: 266), RNRDVKKAVRRLFRRKCC (SEQ ID NO: 267), RNREVRKALRRIIRKRCC (SEQ ID NO: 268), RNKDIRKAVKKIFRRKCC (SEQ ID NO: 269), RNKDVRKAVRRLIKRKCC (SEQ ID NO: 270), RNRDLRKAVRKLFKKKCC (SEQ ID NO: 271), RNRDLRKALRRIFKRRCC (SEQ ID NO: 272), RNRDVRKAIKKLIRKRCC (SEQ ID NO: 273), RNKELKKAIKRILKKKCC (SEQ ID NO: 274), RNRDVRKAIRKLLKRKCC (SEQ ID NO: 275), RNRDLRKAVRRIFKKRCC (SEQ ID NO: 276), RNRDVRKAVRKLFKRRCC (SEQ ID NO: 277), RNRDVRKALRRLFKKRCC (SEQ ID NO: 278), RNKELKKALRKLIGKKCC (SEQ ID NO: 279), RNREMRKAIKKIIKKKCC (SEQ ID NO: 280), RNKEIKKAIKKIIKKRCC (SEQ ID NO: 281), RNRDVKKAIRRLFRRRCC (SEQ ID NO: 282), RNREVKKAVKKLIGKRCC (SEQ ID NO: 283), RNREMRKALRRLFRKRCC (SEQ ID NO: 284), RNKELKKALRRLIGRRCC (SEQ ID NO: 285), RNRDVKKALRKLIGKRCC (SEQ ID NO: 286), RNREVKKAVKKLIRRKCC (SEQ ID NO: 287), RNKEVRKALKKLFGKKCC (SEQ ID NO: 288), RNKEIRKALRRLFGKKCC (SEQ ID NO: 289), RNKDVKKALRRLFGKKCC (SEQ ID NO: 290), RNKELKKAIKRLIRRKCC (SEQ ID NO: 291), RNKDVRKAVKRLLKKRCC (SEQ ID NO: 292), RNKELRKAIRRLLRRRCC (SEQ ID NO: 293), RNRDIRKALRKLFKKKCC (SEQ ID NO: 294), RNRELKKALRRLLRRRCC (SEQ ID NO: 295), RNREVKKALRRLFGKKCC (SEQ ID NO: 296), RNRDVRKALKRLLKRKCC (SEQ ID NO: 297), RNRDMRKAIRKLFGRKCC (SEQ ID NO: 298), RNRELKKAIRKLLKRKCC (SEQ ID NO: 299), RNRDIRKAVKKLFGKKCC (SEQ ID NO: 300), RNKEVKKAIRKLFGRRCC (SEQ ID NO: 301), RNREVRKAVRKLFRRKCC (SEQ ID NO: 302), RNRDMKKALKKLFRRRCC (SEQ ID NO: 303), RNRDVRKALKRLLGRRCC (SEQ ID NO: 304), RNKDLKKAVKKLFGRKCC (SEQ ID NO: 305), RNKDVRKAVRRLFGRRCC (SEQ ID NO: 306), RNRDVRKALRRLFRKKCC (SEQ ID NO: 309), RNRDVRRALRRLFRKKCC (SEQ ID NO: 310), RNKEVKCALKRLLKRKCC (SEQ ID NO: 321), RNKEVKFALKRLLKRKCC (SEQ ID NO: 322), RNKEVKHALKRLLKRKCC (SEQ ID NO: 323), RNKEVKIALKRLLKRKCC (SEQ ID NO: 324), RNKEVKLALKRLLKRKCC (SEQ ID NO: 325), RNKEVKMALKRLLKRKCC (SEQ ID NO: 326), RNKEVKSALKRLLKRKCC (SEQ ID NO: 328), RNKEVKTALKRLLKRKCC (SEQ ID NO: 329), RNKEVKWALKRLLKRKCC (SEQ ID NO: 330), RNKEVKYALKRLLKRKCC (SEQ ID NO: 331), RNKQIRDALKRLLKRKCC (SEQ ID NO: 740), RNKEVKRAIKRLLKRKCR (SEQ ID NO: 121), RNKEVKKAIKRLLKRKCR (SEQ ID NO: 122), RNKEVKRALKRLLKRKRR (SEQ ID NO: 123), RNKEVKRALKRLLKRKYP (SEQ ID NO: 124), RNKEVKRALKRLLKRKRF (SEQ ID NO: 125), RNKEVKRALKRLLKRKFK (SEQ ID NO: 126), RNKEVKKALKRLLKRKRR (SEQ ID NO: 127), RNKEVKKALKRLLKRKYP (SEQ ID NO: 128), RNKEVKKALKRLLKRKRF (SEQ ID NO: 129), RNKEVKKALKRLLKRKFK (SEQ ID NO: 130), RNKEVKDALKRLLKRKCRR (SEQ ID NO: 133), RNKEVKDALKRLLKRKCCC (SEQ ID NO: 134), RNKEVKDALKRLLKRKCCF (SEQ ID NO: 135), RNKEVKDALKRLLKRKCCL (SEQ ID NO: 136), RNKEVKDALKRLLKRKCCM (SEQ ID NO: 137), RNKEVKDALKRLLKRKCCS (SEQ ID NO: 138), RNKEVKDALKRLLKRKCCP (SEQ ID NO: 139), RNKEVKDALKRLLKRKCCA (SEQ ID NO: 140), RNKEVKDALKRLLKRKCCY (SEQ ID NO: 141), RNKEVKDALKRLLKRKCCH (SEQ ID NO: 142), RNKEVKDALKRLLKRKCCN (SEQ ID NO: 143), RNKEVKDALKRLLKRKCCD (SEQ ID NO: 144), RNKEVKDALKRLLKRKCCK (SEQ ID NO: 145), RNKEVKDALKRLLKRKCCR (SEQ ID NO: 146), RNKEVKDALKRLLKRKCCG (SEQ ID NO: 147), RNKEVKDALKRLLKRKCRRR (SEQ ID NO: 149), RNKEVKDALKRLLKRKCRKK (SEQ ID NO: 150), RNKEVKDALKRLLKRKCCRR (SEQ ID NO: 151), RNKEVKDALKRLLKRKCRRRR (SEQ ID NO: 154), RNKEVKDALKRLLKRKCCRRR (SEQ ID NO: 156), RNKEVKKAIKRLLKRKCCRRR (SEQ ID NO: 220), RNKEVKRAIKRLLKRKCCRRR (SEQ ID NO: 237), RNKEVKKAIKRLFKRKCCRRR (SEQ ID NO: 221), RNKEVKRAIKRLFKRKCCRRR (SEQ ID NO: 238), RNKQIRDALKRLLKRKCCRRR (SEQ ID NO: 741), RNKEVKDALKRLLKRKCRRRRR (SEQ ID NO: 157), RNKEVKDALKRLLKRKCRRRKK (SEQ ID NO: 158), and RNKEVKDALKRLLKRKCCRRRR (SEQ ID NO: 219) 18. The olfactory receptor protein according to any one of paragraphs 4 to 17, selected from the group consisting of:
[0249] 19. The olfactory receptor protein of any one of paragraphs 4 to 17, wherein the sequence motif is RNRDVRKALRRLFRKK (sequence number 307) or RNRDVRRALRRLFRKK (sequence number 308).
[0250] 20. The olfactory receptor protein of any one of paragraphs 4 to 18, wherein the sequence motif is RNKEVKKAIKRLFKRKCCRRR (sequence number 221) or RNKQIRDALKRLLKRKCCRRR (sequence number 741).
[0251] 21. The olfactory receptor protein of any one of paragraphs 4 to 20, wherein the sequence motif is selected from the group consisting of SEQ ID NOs: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, and 820-890.
[0252] 22. The olfactory receptor protein of any of the preceding paragraphs, wherein the olfactory receptor further comprises an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (Rho) tag, an SST3 tag, and an M3 tag.
[0253] 23. A nucleic acid molecule comprising a nucleotide sequence encoding an olfactory receptor protein as described in any of the preceding paragraphs.
[0254] 24. The nucleic acid molecule according to paragraph 23, further comprising a promoter sequence, preferably a constitutive promoter sequence.
[0255] 25. The nucleic acid molecule of paragraph 23 or 24, further comprising a terminator sequence.
[0256] 26. The nucleic acid molecule according to any one of paragraphs 23 to 25, further comprising a nucleotide sequence encoding an N-terminal signal peptide, preferably a leucine-rich signal peptide such as MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77).
[0257] 27. The nucleotide sequence is selected from the group consisting of SEQ ID NOs: 331-739 and SEQ ID NOs: 742-819, and at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 27. The nucleic acid molecule of any one of paragraphs 23 to 26, comprising at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity.
[0258] 28. An expression vector comprising a nucleic acid molecule according to any one of paragraphs 23 to 27.
[0259] 29. The expression vector of paragraph 28, which is a plasmid.
[0260] 30. A recombinant host cell comprising a nucleic acid molecule as described in any one of paragraphs 23 to 27 or an expression vector as described in paragraph 28 or 29, preferably wherein the cell expresses an olfactory receptor protein as described in any one of paragraphs 1 to 22.
[0261] 31. The recombinant host cell of paragraph 30, wherein the cell further expresses one or more olfactory receptor accessory proteins.
[0262] 32. One or more olfactory receptor accessory proteins are RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, or Gα olf32. The recombinant host cell of paragraph 31, wherein the RTP1S is selected from the group consisting of RTP1S, RTP2, and functional variants thereof, and preferably selected from the group consisting of RTP1S, RTP2, and functional variants thereof.
[0263] 33. The recombinant host cell of any one of paragraphs 30 to 32, wherein the cell is a HEK293 cell or a HEK293T cell.
[0264] 34. A library comprising a diverse repertoire of olfactory receptor proteins as described in any one of paragraphs 1 to 22, nucleic acid molecules as described in any one of paragraphs 23 to 27, expression vectors as described in paragraph 28 or 29, or recombinant host cells as described in any one of paragraphs 30 to 33.
[0265] 35. The library described in paragraph 34, wherein a diverse repertoire of olfactory receptor proteins, olfactory receptor proteins encoded by nucleic acid molecules or expression vectors, or olfactory receptor proteins expressed by recombinant host cells share the same modified C-terminal domain.
[0266] 36. The library of paragraph 34 or 35, wherein the library comprises at least 25 (preferably 400 to 450) distinct class II olfactory receptor proteins, preferably human class II olfactory receptor proteins, nucleic acid molecules or expression vectors encoding said olfactory receptor proteins, or recombinant host cells expressing said olfactory receptor proteins, more preferably wherein the library comprises at least one, and most preferably all, of the class II olfactory receptors listed under "Receptor" in Table 15.
[0267] 37. The library of any one of paragraphs 34 to 36, wherein the library comprises at least 25 (preferably 70 to 100) distinct class I olfactory receptor proteins, preferably human class I olfactory receptor proteins, nucleic acid molecules or expression vectors encoding said olfactory receptor proteins, or recombinant host cells expressing said olfactory receptor proteins, more preferably wherein the library comprises at least one, and preferably all, of the class I olfactory receptors listed under "Receptor" in Table 22.
[0268] 38. The library of paragraph 34 or 35, wherein the library comprises a diverse repertoire of at least 250 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins.
[0269] 39. Use of an olfactory receptor protein as described in any one of paragraphs 1 to 22, a nucleic acid molecule as described in any one of paragraphs 23 to 27, an expression vector as described in paragraph 28 or 29, a recombinant host cell as described in any one of paragraphs 30 to 33, or a library as described in any one of paragraphs 34 to 38 to identify a ligand, enhancer or antagonist of an olfactory receptor.
[0270] 40. Use of a library as described in any one of paragraphs 34 to 38 for identifying olfactory receptors capable of binding to a target ligand.
[0271] 41. A method for identifying an olfactory receptor ligand, comprising: a) providing a recombinant host cell expressing an olfactory receptor protein as described in any one of paragraphs 1 to 22 or an olfactory receptor protein as described in any one of paragraphs 30 to 33; b) contacting the receptor or recombinant host cell with a test compound or composition; and c) detecting activation of olfactory receptors The method comprising:
[0272] 42. A method for identifying an enhancer or antagonist of an olfactory receptor, comprising: a) providing a cell expressing an olfactory receptor protein as described in any one of paragraphs 1 to 22 or an olfactory receptor protein as described in any one of paragraphs 30 to 33; b) contacting the receptor or recombinant host cell with a cognate ligand and a test compound or composition; and c) detecting increased or decreased activation of olfactory receptors compared to a ligand-only control; The method comprising:
[0273] 43. The method of paragraph 41 or 42, preferably wherein the method is for identifying an olfactory receptor ligand or enhancer, wherein the olfactory receptor is selected from the group consisting of OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)), OR R7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25 (OR2A25 (S75N, A209P)), OR11G2 (or OR11G2 (I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3 (S73G)), OR10G9, OR2L5, OR8H1, OR10K 1. The method of claim 1, wherein the target gene is selected from the group consisting of OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2, and OR2AG2 (preferably OR2AG2(Y28C)).
[0274] 44. The method according to paragraph 42, wherein the method is for identifying an olfactory receptor antagonist, and wherein the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1, and OR4S2, preferably OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and OR5V1, more preferably OR2M2 or OR2V1.
[0275] 45. The method of paragraph 42, wherein the method is for identifying an olfactory receptor antagonist, wherein the olfactory receptor is OR2M2 or OR2V1, and optionally, step b) further comprises contacting the receptor or recombinant host cell with a copper salt.
[0276] 46. The method of paragraph 45, wherein the cognate ligand is selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, and 4-methoxy-2-methylpentane-2-thiol.
[0277] 47. The method according to paragraph 42, wherein the method is for identifying an olfactory receptor antagonist, and wherein the olfactory receptor is OR51B2 (preferably OR51B2 (C120R, L134F, C209S)).
[0278] 48. The method of paragraph 47, wherein the cognate ligand is 3-methyl-2-hexenoic acid.
[0279] 49. The method according to paragraph 42, wherein the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR5V1.
[0280] 50. The method of paragraph 49, wherein the cognate ligand is 2,4,6-trichloroanisole.
[0281] 51. A method for identifying an olfactory receptor capable of binding to a target ligand, comprising: a) provide a library as set out in any one of paragraphs 34 to 38; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins with a target ligand; and d) Identifying the olfactory receptors activated by target ligands The method comprising:
[0282] 52. A method for generating an objective representation of the olfactory properties of a test compound or composition, comprising: a) provide a library as set out in any one of paragraphs 34 to 38; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a test compound or composition; and d) detecting the activation of each of the olfactory receptor proteins The method comprising:
[0283] 53. A method for assessing the differences or similarities between two or more test compounds or compositions, comprising: a) provide a library as set out in any one of paragraphs 34 to 38; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of two or more test compounds or compositions; d) detecting activation of each of the olfactory receptor proteins for each of two or more test compounds or compositions; and e) comparing the activated olfactory receptor protein between each of two or more test compounds or compositions; The method comprising:
[0284] General information Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly and commonly understood by one of ordinary skill in the art to which this invention belongs and as read in light of the present disclosure.
[0285] Sequence identity It should be understood that each nucleic acid molecule or protein fragment or polypeptide or peptide or derived peptide or construct identified herein by a given sequence identity number (SEQ ID NO:) is not limited to this particular sequence as disclosed. Each coding sequence as identified herein encodes a given protein fragment or polypeptide or peptide or derived peptide or construct, or is itself a protein fragment or polypeptide or construct or peptide or derived peptide.
[0286] Throughout this application, whenever a reference is made to the SEQ ID NO of a particular nucleotide sequence encoding a given protein fragment or polypeptide or peptide or derived peptide (for example, SEQ ID NO:X), it may be replaced by: i. a nucleotide sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 99% sequence identity to SEQ ID NO: X; ii. a nucleotide sequence whose sequence differs from that of the nucleic acid molecule in (i) due to the degeneracy of the genetic code; or iii. A nucleotide sequence encoding an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% amino acid identity or similarity to the amino acid sequence encoded by nucleotide sequence SEQ ID NO:X.
[0287] A preferred level of sequence identity is 70%. Another preferred level of sequence identity or similarity is 75%. Another preferred level of sequence identity or similarity is 80%. Another preferred level of sequence identity or similarity is 85%. Another preferred level of sequence identity or similarity is 90%. Another preferred level of sequence identity or similarity is 95%. Another preferred level of sequence identity or similarity is 99%.
[0288] Throughout this application, whenever an amino acid sequence of a particular SEQ ID NO (for example, SEQ ID NO: Y) is mentioned, it may be replaced with a polypeptide represented by an amino acid sequence comprising a sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% sequence identity or similarity to the amino acid sequence of SEQ ID NO: Y. A preferred level of sequence identity is 70%. Another preferred level of sequence identity or similarity is 75%. Another preferred level of sequence identity or similarity is 80%. Another preferred level of sequence identity or similarity is 85%. Another preferred level of sequence identity or similarity is 90%. Another preferred level of sequence identity or similarity is 95%. Another preferred level of sequence identity or similarity is 99%.
[0289] Each nucleotide or amino acid sequence described herein by a percentage of identity or similarity to a given nucleotide or amino acid sequence, respectively, is in further preferred embodiments at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133 4%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity or similarity.
[0290] Each non-coding nucleotide sequence (i.e., that of a promoter or that of another regulatory region) can be replaced by a nucleotide sequence containing a nucleotide sequence having at least 60% sequence identity or similarity with the sequence number of a specific nucleotide sequence (e.g., SEQ ID NO: A). Preferred nucleotide sequences have at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity to SEQ ID NO:A. In preferred embodiments, such non-coding nucleotide sequences, such as promoters, exhibit or exert at least an activity of such non-coding nucleotide sequences (such as the activity of a promoter as known to those skilled in the art).
[0291] The terms "homology," "sequence identity," and the like are used interchangeably herein. Sequence identity is described herein as the relationship between two or more amino acid (polypeptide or protein) sequences or two or more nucleic acid (polynucleotide) sequences, which is determined by comparing the sequences. In a preferred embodiment, sequence identity is calculated based on the full length of the two given sequences or a portion thereof, more preferably based on the full length of the two given sequences. A portion thereof preferably means at least 50%, 60%, 70%, 80%, 90%, or 100% of both sequences. In the art, "identity" also refers to the degree of sequence relatedness between amino acid sequences or nucleic acid sequences, as the case may be, as determined by the match between the strings of such sequences. The "similarity" between two amino acid sequences is determined by comparing the amino acid sequence of one polypeptide and its conservative amino acid substitutes with the sequence of a second polypeptide. "Identity" and "similarity" can be readily calculated by known methods, including but not limited to those described in Bioinformatics and the Cell: Modern Computational Approaches in Genomics, Proteomics and Transcriptomics, Xia X., Springer International Publishing, New York, 2018; and Bioinformatics: Sequence and Genome Analysis, Mount D., Cold Spring Harbor Laboratory Press, New York, 2004, each of which is incorporated herein by reference.
[0292] "Sequence identity" and "sequence similarity" can be determined by aligning two peptide or two nucleotide sequences using a global or local alignment algorithm, depending on the length of the two sequences. Sequences of similar length are preferably aligned using a global alignment algorithm (e.g., Needleman-Wunsch) that optimally aligns the sequences over their entire length, while sequences of substantially different lengths are preferably aligned using a local alignment algorithm (e.g., Smith-Waterman). Sequences are then said to be "substantially identical" or "essentially similar" if they share at least a certain minimum percentage of sequence identity (as described below) when optimally aligned (e.g., using the programs EMBOSS needle or EMBOSS water with default parameters).
[0293] Global alignment is preferably used to determine sequence identity when two sequences have similar lengths. When sequences have substantially different overall lengths, local alignment, such as that using the Smith-Waterman algorithm, is preferred. EMBOSS needle uses the Needleman-Wunsch global alignment algorithm to align two sequences over their entire length (full length), maximizing the number of matches and minimizing the number of gaps. EMBOSS water uses the Smith-Waterman local alignment algorithm. Generally, the default parameters of EMBOSS needle and EMBOSS water are used: gap open penalty = 10 (nucleotide sequence) / 10 (protein), gap extension penalty = 0.5 (nucleotide sequence) / 0.5 (protein). For nucleotide sequences, the default scoring matrix used is DNAfull, and for proteins, the default scoring matrix is Blosum62 (Henikoff & Henikoff, 1992, PNAS 89, 915-919, incorporated herein by reference).
[0294] Alternatively, the percentage of similarity or identity may be determined by searching against public databases using algorithms such as FASTA or BLAST. Thus, the nucleic acid and protein sequences of some embodiments of the present disclosure can further be used as "query sequences" to search against public databases, for example, to identify other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) of Altschul et al. (1990) J. Mol. Biol. 215:403-10, which are incorporated herein by reference. To obtain nucleotide sequences homologous to the nucleic acid molecules of the present disclosure, BLAST nucleotide searches can be performed with the NBLAST program, score = 100, word length = 12. To obtain amino acid sequences homologous to the protein molecules of the present disclosure, BLAST protein searches can be performed with the BLASTx program, score = 50, word length = 3. To obtain gapped alignments for comparison purposes, Gapped BLAST can be used as described in Altschulet et al. (1997) Nucleic Acids Res. 25(17):3389-3402, which is incorporated herein by reference. When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., BLASTx and BLASTn) can be used. See the homepage of the National Center for Biotechnology Information, accessible on the World Wide Web (located at www.ncbi.nlm.nih.gov / ).
[0295] Optionally, when determining the degree of amino acid similarity, those skilled in the art may also take into account so-called conservative amino acid substitutions. As used herein, "conservative" amino acid substitution refers to the interchangeability of residues with similar side chains. Examples of classes of amino acid residues for conservative substitution are provided in the table below. [Table A]
[0296] Alternative conservative amino acid residue substitution classes: [Table B]
[0297] Physical and functional classification of alternative amino acid residues: [Table C]
[0298] For example, the group of amino acids with aliphatic side chains is glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic-hydroxyl side chains is serine and threonine; the group of amino acids with amide-containing side chains is asparagine and glutamine; the group of amino acids with aromatic side chains is phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains is lysine, arginine, and histidine; the group of amino acids with sulfur-containing side chains is cysteine and methionine.Preferred conservative amino acid substitution groups are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.Substitution variants of the amino acid sequences disclosed herein are those in which at least one residue in the disclosed sequence is removed and a different residue is inserted in its place.Preferably, the amino acid change is conservative. Preferred conservative substitutions for each of the naturally occurring amino acids are as follows: Ala to Ser; Arg to Lys; Asn to Gln or His; Asp to Glu; Cys to Ser or Ala; Gln to Asn; Glu to Asp; Gly to Pro; His to Asn or Gln; Ile to Leu or Val; Leu to Ile or Val; Lys to Arg; Gln or Glu; Met to Leu or Ile; Phe to Met, Leu, or Tyr; Ser to Thr; Thr to Ser; Trp to Tyr; Tyr to Trp or Phe; and Val to Ile or Leu. Particularly preferred amino acid substitutions in the present disclosure include substitutions of Lys for Arg and Arg for Lys.
[0299] Gene or coding nucleotide sequence The term "gene" refers to a DNA fragment comprising a region (transcribed region) that is transcribed into an RNA molecule (e.g., mRNA) within a cell and operably linked to a suitable regulatory region (e.g., a promoter). The coding nucleotide sequence may comprise a sequence native to the cell, a sequence that does not naturally occur in the cell, or it may comprise a combination of both.
[0300] operably linked As used herein, the term "operably linked" refers to the linkage of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, a transcriptional regulatory sequence is operably linked to a coding sequence if it affects the transcription of the coding sequence. Operably linked means that the DNA sequences being linked are typically contiguous, and, where necessary to join two protein-coding regions, contiguous and in reading frame. Linking can be accomplished by ligation at convenient restriction sites, or with adapters or linkers inserted instead, or by gene synthesis.
[0301] Proteins and Amino Acids The terms "protein" or "peptide" or "polypeptide" or "amino acid sequence" are used interchangeably and refer to a molecule consisting of a chain of amino acids, regardless of its particular mechanism of action, size, three-dimensional structure, or origin. In an amino acid sequence as described herein, the amino acids or "residues" are designated by their three-letter symbols. These three letter symbols, as well as the corresponding one letter symbols, are well known to those skilled in the art and have the following meanings: A (Ala) is alanine, C (Cys) is cysteine, D (Asp) is aspartic acid, E (Glu) is glutamic acid, F (Phe) is phenylalanine, G (Gly) is glycine, H (His) is histidine, I (Ile) is isoleucine, K (Lys) is lysine, L (Leu) is leucine, M (Met) is methionine, N (Asn) is asparagine, P (Pro) is proline, Q (Gln) is glutamine, R (Arg) is arginine, S (Ser) is serine, T (Thr) is threonine, V (Val) is valine, W (Trp) is tryptophan, and Y (Tyr) is tyrosine. The residue may be any proteinogenic amino acid, as well as any non-proteinogenic amino acid, such as D-amino acids and modified amino acids formed by post-translational modifications, and also any unnatural amino acid. In a preferred embodiment, an amino acid in this disclosure may refer to any one of the 20 standard proteinogenic amino acids (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y).
[0302] In this specification and the claims, the verb "comprise" and its conjugations are used in their open-ended sense to mean including or containing items that follow the word, but not excluding items not specifically mentioned. Thus, the terms "comprising," "including," "comprising," and the like, when used herein, are synonymous with "including," "comprises," or "containing," and are inclusive or open-ended and do not exclude additional, unrecited members, elements, or method steps.
[0303] Additionally, the verb "consisting of" may be replaced by "consisting essentially of," meaning that the subject matter as described herein may include additional component(s) not specifically identified, and that said additional component(s) do not alter the inherent characteristics of the disclosure. Additionally, the verb "consisting of" may be replaced by "consisting essentially of," meaning that the method as described herein may include additional step(s) not specifically identified, and that said additional step(s) do not alter the inherent characteristics of the disclosure.
[0304] Throughout this disclosure, the term "comprising" may be replaced by the terms "consisting essentially of" or "consisting of." As used herein, the singular forms "a," "an," and "the" include both the singular and plural referents unless the context clearly dictates otherwise. Thus, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein. As used herein, the term "at least" refers to a particular value or more than that particular value. For example, "at least 2" is understood to be the same as "2 or more," i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, ...
[0305] Furthermore, terms such as first, second, third, etc. in the description and in the claims are used to distinguish between similar elements and not necessarily to describe a sequential or chronological order, with the understanding that the terms so used are interchangeable under appropriate circumstances and that the aspects described herein are capable of operating in sequences other than those described or illustrated herein.
[0306] The term "about" or "approximately," when used in connection with a numerical value (e.g., about 10), preferably means that the value may be 1% more or less than the given value (10). As used herein, the term "and / or" indicates that one or more of the stated cases may occur alone or in combination with at least one of the stated cases, up to and including all of the stated cases.
[0307] Various aspects are described herein. Each aspect as identified herein may be combined together unless otherwise indicated. Titles, subtitles, headings, etc. are used herein merely for ease of reading and are not intended to limit or restrict the present disclosure in any way.
[0308] All patent applications, patents, and printed publications cited herein are incorporated herein by reference in their entirety, except for any definitions, disclaimers or disclaimers of subject matter, and except to the extent that the incorporated material is inconsistent with the explicit disclosure herein, in which case the language of the present disclosure will control.
[0309] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of the present invention. Indeed, the present invention is in no way limited to the methods and materials described.
[0310] The present invention is further illustrated by the following examples, which should not be construed as limiting the scope of the present invention. Generalizations of the specific aspects and features disclosed in the following examples to the foregoing description are part of this disclosure. [Brief explanation of the drawings]
[0311] [Figure 1]Comparison of wild-type OR5A2 with the wild-type C-terminal sequence (SEQ ID NO: 239) and chimeric OR5A2 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). Dose-dependent luciferase induction by various ligands is shown.
[0312] [Figure 2] Example of dose-response analysis of wild-type and chimeric OR7C1 and OR9Q2 with optimized C-terminal sequences (SEQ ID NO: 86) used to derive the detection threshold. EC50 values are summarized in Table 1.
[0313] [Figure 3] Dose-response analysis comparing wild-type OR5B12 with the wild-type C-terminal sequence (SEQ ID NO: 241) and chimeric OR5B12 with an optimized C-terminal sequence (SEQ ID NO: 86) with a variant in which the wild-type C-terminal sequence was changed to the consensus sequence (SEQ ID NO: 242) by changing the amino acid sequence to leucine residues at two positions, making all conserved residues identical to the consensus sequence.
[0314] [Figure 4] Dose-response analysis comparing wild-type OR5A2 with the wild-type C-terminal sequence (SEQ ID NO: 239) and chimeric OR5A2 with an optimized C-terminal sequence (SEQ ID NO: 86) with a variant in which the wild-type C-terminal sequence was changed to the consensus sequence (SEQ ID NO: 240) by changing the amino acid sequence at three positions to make all conserved residues identical to the consensus sequence.
[0315] [Figure 5] Dose-response analysis comparing wild-type OR7A17 with the wild-type C-terminal sequence (SEQ ID NO: 243) and chimeric OR7A17 with an optimized C-terminal sequence (SEQ ID NO: 86) with a variant in which the wild-type C-terminal sequence was changed to the consensus sequence (SEQ ID NO: 244) by changing the amino acid sequence at five positions to make all conserved residues identical to the consensus sequence.
[0316] [Figure 6] Dose-response analysis of wild-type and chimeric O8K3 with various ligands. The weaker agonist, menthone, can only be identified by the more sensitive variant with an optimized C-terminus (SEQ ID NO: 86).
[0317] [Figure 7] Dose-response analysis of wild-type and chimeric O7D4 with various ligands. The weaker agonist, androstenol, can only be identified by the more sensitive chimeric variant with an optimized C-terminus (SEQ ID NO: 86).
[0318] [Figure 8] Screening of OR2T4 wild type and OR2T4 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) against 52 sulfur compounds.
[0319] [Figure 9] Dose-response analysis of wild-type and chimeric OR2T4 with various ligands. No ligands are known for OR2T4, and screening of the wild-type receptor with candidate ligands did not result in deorphanization. However, testing of a chimeric variant with an optimized C-terminus (SEQ ID NO: 86) revealed that the receptor is specifically activated by certain sulfur-containing compounds, such as cyclopentanethiol.
[0320] [Figure 10] Screening of OR2T11 wild type and OR2T11 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) against 52 sulfur compounds.
[0321] [Figure 11]Example of dose-response analysis of chimeric OR7C1 with wild-type and C-terminal sequence (SEQ ID NO: 86) and optimized C-terminal sequence variants RNKEVKRAIRKLLKRKCC (SEQ ID NO: 118) and RNKEVKRAIKRLLKRKCR (SEQ ID NO: 121) that combine multiple sequence variations described in Examples 6 and 7.
[0322] [Figure 12] Activation of OR52E8 by odorant acids present in human sweat. The wild-type was not activated by acid. Replacing the wild-type with the functional C-terminal sequence of OR51E1 did not improve expression. However, using the optimized sequence from Examples 1-5 (SEQ ID NO: 86) resulted in a strong signal upon addition of 3-methyl-3-hydroxyhexanoic acid.
[0323] [Figure 13] Activation of OR56A4 by acids of different chain lengths. The wild-type was activated by decanoic acid and undecanoic acid at high concentrations, but not by nonanoic acid. Replacing the wild-type with the functional C-terminal sequence of OR51E1 reduced activity. However, using the optimized sequence from Examples 1-5 (SEQ ID NO: 86) resulted in a strong signal and a much lower detection threshold upon addition of all three acids.
[0324] [Figure 14] Various chimeric variants of OR8K3: The wild-type C-terminal sequence was replaced with either the optimized sequence (SEQ ID NO: 86) from Examples 1-5 or the C-terminal sequences of two functional receptors, OR1N2 and OR5AN1. Chimeric receptors with C-terminal sequences from other functional receptors did not provide functional expression. Dose-dependent luciferase induction by various ligands is shown.
[0325] [Figure 15]Chimeric variants of OR5AN1: The wild-type C-terminal sequence was replaced with the C-terminal sequence of the functional receptor OR1N2. Chimeric receptors with the C-terminal sequence from OR1N2 did not provide functional expression. Dose-dependent luciferase induction by various ligands is shown.
[0326] [Figure 16] Chimeric variant of OR5A2: The wild-type C-terminal sequence was replaced with the C-terminal sequence of the functional receptor OR1N2. The chimeric receptor with the C-terminal sequence from OR1N2 provided functional expression, which was approximately 100-fold weaker than OR5A2 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). Dose-dependent luciferase induction by various ligands was demonstrated.
[0327] [Figure 17] Arborone (left) compared to Iso E Super (OTNE, right) for activation of OR7A17 with the optimized C-terminus RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). Dose-dependent luciferase induction by the indicated ligands is shown.
[0328] [Figure 18] Chemicals tested for odor threshold in vivo (OTH_mean, ng / L) vs. EC50% (μM) determined for activation of OR7A17 with the optimized C-terminus RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) in vitro.
[0329] [Figure 19] Screening of OR2M2 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) against 52 sulfur compounds in the presence of 30 μM copper.
[0330] [Figure 20]Dose-response analysis of OR2M2 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) to 3-methyl-3-sulfanyl-hexanol, a major human malodorous sulfur compound, and related chemicals, in the presence and absence of copper.
[0331] [Figure 21] Screening of the entire library of class II OR variants (n=408) with 50 μM patchoulol. The Y-axis shows fold luciferase induction. The OR numbers on the X-axis correspond to the numbers in Table 15.
[0332] [Figure 22] Dose-response analysis of the library compared the improved OR sequence with the wild-type, patchoulol-OR14J1, and found that the wild-type was inactive and deorphanization of the OR was only possible with the modified sequence.
[0333] [Figure 23] Screening of a library (n=77) of class I OR variants with 500 μM 3-methyl-2-hexenoic acid. The Y-axis shows fold luciferase induction. The OR numbers on the X-axis correspond to the numbers in Table 22.
[0334] [Figure 24] Screening of the entire library of class II OR variants (n=408) with three different fragrances tested at 10 ppm. The Y-axis indicates fold luciferase induction. On the X-axis are represented ORs that resulted in at least a 2-fold induction of luciferase with at least one fragrance. Only activated ORs are shown.
[0335] [Figure 25]Geranium oil supplemented with various levels of Ambrofix® (0.1% to 10%) was tested in a dose-response analysis for OR7A17, which has the C-terminal domain of SEQ ID NO: 221 (the full-length DNA sequence encoding the modified receptor is SEQ ID NO: 611). The Y-axis shows fold luciferase induction. On the X-axis, the concentration of geranium oil (ppm) is shown.
[0336] [Figure 26] Screening of Iso E super and Ambermax® in cells expressing either OR7A17 or OR7C1 or both, which have the C-terminal domain of SEQ ID NO: 221. The Y-axis shows fold luciferase induction. Concentration (μM) is shown on the X-axis.
[0337] [Figure 27] Activation of OR10J5, which has the C-terminal domain of SEQ ID NO: 86, by two structural isomers, (S,E)-10-hydroxy-4,8-dimethyldec-4-enal and (R,E)-10-hydroxy-4,8-dimethyldec-4-enal. The Y-axis shows the luciferase induction rate of the positive control (Mahonial®). The X-axis shows the concentration of the test compound in micromolar. DETAILED DESCRIPTION OF THE INVENTION
[0338] example A general approach to OR expression To generate expression plasmids, OR coding sequences fused to the desired C-terminal sequence at the end of TM7 were synthesized by a DNA synthesis service provider (BioCat GmbH, Germany) and inserted into pcDNA3.1(+) (Invitrogen, MA, USA) downstream of the CMV promoter sequence (SEQ ID NO: 82) using the BamHI and NotI restriction sites. All synthetic OR nucleotide sequences described below in Examples 1-15 further contain a nucleotide sequence encoding a signal peptide (mmLucy-FLAG-rho, SEQ ID NOs: 80 and 81) at their N-terminus and a bgh terminator sequence (SEQ ID NO: 83) at their C-terminus. All OR expression plasmids also contain a Kozak sequence (GCCACC) between the BamHI restriction site and the start codon of the signal peptide. Thus, these plasmids contain constitutively expressed OR genes.
[0339] Expression of OR genes was typically performed in HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). These cells were seeded at a density of 10,000 cells / well in polyethyleneimine-coated 96-well plates (100 μl / well) and grown for 24 hours at 37° C. in the presence of 5% CO2.
[0340] 0.625 μg of OR expression plasmid, 1 μg of empty pcDNA3.1(+) vector, and 1 μg of pGL4.29 (Promega) carrying CRE-inducible luciferase were diluted in 0.25 ml of OptiMEM medium (Gibco™, ThermoFisher Scientific, MA, USA). In parallel, 15 μl of Lipofectamine 2000 (Invitrogen) was diluted in 0.25 ml of OptiMEM medium, and after 5 minutes of preincubation, the two mixtures were combined to prepare the transfection mixture, which was then incubated for another 25 minutes.
[0341] 50 μl of growth medium was replaced with fresh DMEM containing 9% fetal bovine serum (FBS). The pre-incubated transfection mixture was diluted to 5 ml in OptiMEM medium, and 50 μl of the diluted mixture was added per well (total final volume 150 μl). Cells were further cultured for 24 hours at 37°C in the presence of 5% CO2 to allow DNA uptake and OR expression.
[0342] Functional expression and response to ligand were tested by removing 100 μl of growth medium and adding 50 μl of DMEM containing 9% FBS, ligand, and a maximum of 1% DMSO. Cells were stimulated for 4.5 hours, and then the luciferase signal induced by OR-dependent cAMP production was measured.
[0343] Example 1: Improved functional expression of modified OR5A2 using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) Wild-type OR5A2 and OR5A2 genes modified with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were transfected into HEK293T cells stably integrated with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85).
[0344] No activation of the wild-type receptor was detected by either the polycyclic musk Galaxolide® or the macrocyclic musk muscone (Figure 1, left panel). However, the modified receptor was activated by both musk compounds. As shown in Figure 1 (right panel), activation occurred at low concentrations (<0.01 μM for Galaxolide® and 1 μM for muscone) and was specific to the musk compounds; no response was recorded with ethyl vanillin.
[0345] Therefore, only by using modified variants of OR5A2 with modified C-terminal sequences, rather than the wild-type C-terminal sequence, could we efficiently screen compounds with potentially musk-like odors against OR5A2 expressed in cell lines with RTP1S and RTP2 and a Lucy-FLAG-rho-tag.
[0346] Example 2: Improving the sensitivity of extensively modified ORs using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) As shown for OR5A2 in Example 1, modified versions of various OR genes were generated by replacing the C-terminal sequence after TM7 with the optimized sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). Cells were transfected with either the wild-type OR gene or the respective modified versions. The transfected cells were stimulated with the cognate ligands of these receptors in a dose-response analysis, and the lowest concentration at which the ligand induced a luciferase signal twice the background level was determined from the dose-response curve as a measure of the lower detection threshold. In parallel, the EC50, i.e., the concentration at which 50% of maximum activation (potency) was reached, was calculated. Finally, the maximum luciferase induction ratio (efficacy) compared to the solvent control was determined and compared between the wild-type and modified versions.
[0347] As can be seen in Table 1 and Figure 2, as well as in the examples, the detection threshold (concentration (μM) for 2-fold induction) was clearly lowered for all receptors tested using the optimized sequence, SEQ ID NO: 86; for many receptors, this increase in sensitivity was 10- to 100-fold or more. This increase in sensitivity was also observed by comparing EC50 values, which were also affected by the observed maximum induction (efficacy), which for some receptors was lower for the wild-type than for the modified versions. Thus, for OR52E8, OR2T4, and OR5A2, no induction was observed for the wild-type, while functional expression was achieved for the modified versions (an all-or-nothing effect). Other receptors (e.g., OR8K3, OR5B12, OR7C1, OR7D4) also had much lower efficacies for the wild-type compared to the modified versions (in addition to a much higher detection threshold for the wild-type).
[0348] Table 1. Improved functional expression of modified OR genes containing the truncated and optimized C-terminus RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]
[0349] Example 3: Improved functional expression of modified ORs using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) was compared to OR genes optimized by altering the C-terminus towards consensus. One previously investigated option for optimizing expression (Kotthoff et al., 2021, supra) was based on modifying the parent C-terminal sequence of a receptor toward a consensus sequence by restoring consensus at a single non-conserved amino acid, while leaving any residues in the C-terminal sequence for which no clear consensus residue existed from multiple sequence alignment unchanged. This resulted in a specific C-terminal sequence for each receptor derived from that receptor's specific parent sequence. This approach was compared to an approach in which each receptor was modified with the same optimized C-terminus as described in Examples 1 and 2. Thus, the C-terminal sequence of OR5B12 was modified by introducing two leucine residues at the positions where this amino acid was most common in the consensus sequence. At all other positions, OR5B12 already contains residues typical of the consensus sequence (Kotthoff et al., 2021, supra). As shown in Figure 3, altering the sequence of OR5B12 toward the consensus had no significant effect, while the optimized sequence resulted in stronger functional expression compared to the wild-type. Thus, restoring consensus with conserved amino acids in the parent sequence, as done by Kotthoff et al. (2021, supra), is not a generally applicable method for improving functional expression, whereas generating modified receptors with optimized C-terminal sequences as described herein is an applicable method.
[0350] Similarly, for OR5A2, changing the native C-terminus toward consensus by exchanging three amino acids had no effect, and similar to the wild type, functional expression was not achieved by this approach, whereas adding an optimized C-terminus did result in functional expression (Figure 4).
[0351] In the case of OR7A17, which already had a relatively low detection threshold, two-fold lower than the wild-type at 1.9 μM for Ambrofix®, changing the sequence toward consensus by exchanging five amino acids indeed had a positive effect, lowering the detection threshold three-fold to 0.6 μM; however, by using an optimized C-terminal sequence, a much better effect was achieved, as the detection threshold could be lowered 18-fold (Table 1 and FIG. 5). Thus, the optimal improvement in activity is not achieved by simply using the parent C-terminus and mutating it toward consensus, but rather by exchanging the C-terminal sequence with the optimized sequence described herein.
[0352] Example 4: A more comprehensive ligand spectrum and discovery of novel ligand-OR pairs by using the optimized C-terminus RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) The ligand spectrum of the modified OR with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) was further tested with multiple ligands and compared to the OR with the wild-type sequence.
[0353] As shown in Figure 6, the wild-type receptor OR8K3 showed only a weak response to menthol (both natural menthol and (S)-menthol ("Menthol Laevo")), while the related molecule menthone was inactive up to 316 μM. Using the modified receptor variant, stronger activation by menthol was observed, while in parallel, menthone also activated the receptor, albeit with lower potency. Thus, the improved sensitivity provided by the modified variant allows the assay to detect activation of OR8K3 by a broader set of mint-smelling molecules.
[0354] Similarly, Figure 7 shows the dose response of OR7D4 when tested with both androstenone and the closely related molecule androstenol. Significant activation by androstenone is demonstrated for both variants, whereas androstenol can activate only the modified version of the receptor: androstenol is a significantly weaker agonist, but would appear to lack activation of OR7D4 if tested against only the weakly expressed wild-type variant. Thus, it is possible to better understand the receptor space of OR7D4 based on the improved sensitivity of the assay provided by the modified variant.
[0355] Example 5: Identification of OR-ligand pairs using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) Little is known about the binding of major sulfur odorants to human ORs. Therefore, a library of sulfur odorants was screened against a selection of candidate sulfur compound receptors. As shown in Figure 8, no clear OR-ligand pairs were identified for OR2T4 with the wild-type sequence, whereas a more sensitive modified version with an optimized C-terminus identified diallyl disulfide, dipropyl disulfide, 2-methyl-3-tetrahydrofuranthiol, and cyclopentanethiol as good ligands for this receptor. The difference in the effectiveness of this screening was also verified by dose-response analysis using the wild-type and receptor variants (Figure 9). Similarly, for OR2T11, none of the 52 sulfur compounds were active against the wild-type, whereas a more sensitive modified variant with an optimized C-terminus identified eight ligands with more than four-fold induction for this receptor (Figure 10). Thus, the optimized C-terminal sequence facilitates deorphanization of the receptor.
[0356] Example 6: Flexibility of optimized C-terminal sequences: base substitutions in C-terminal sequence motifs As shown in Examples 1-5, the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) allows for improved heterologous OR activation experiments against a variety of ORs, allowing for deorphanization of ORs, discovery of novel ligands for the deorphaned ORs, testing at lower ligand concentrations with fewer solubility issues and less cytotoxicity, and obtaining higher efficacy due to better signal-to-noise ratios.
[0357] To test for potential variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that would result in similar functional expression improvements, various point mutations were introduced into the first 16 amino acids of this sequence (RNKEVKDALKRLLKRK, corresponding to SEQ ID NO: 10), and the variants were fused to OR7C1 after TM7. As shown in Figure 2, OR7C1 with the C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) results in approximately 10-fold luciferase induction at 1 μM Ambermax® and 20-fold induction at 3.1 μM Ambermax®. The activity of the various modified variants was then compared to the variant with the optimized sequence described above at these two concentrations. Therefore, all variants were transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). Cells were then stimulated with 1 μM and 3.1 μM Ambermax®, and the luciferase induction rate was compared to that of the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which was set to 100%, and to that in an equivalent experiment performed with the wild-type. Test concentrations were selected to fall within the range of partial activity of the wild-type sequence.
[0358] Sequence modifications of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that conferred greater than 66% activation at 1 μM Ambermax® were considered active variants that could be preferentially employed for optimized expression. Sequence modifications that conferred 33-66% activation of the reference sequence at 1 μM Ambermax® were considered partially active variants and could optionally be employed for optimized expression. As shown in Table 2, the optimized sequence employed in Examples 1-5 is one possible option, but other variants with single amino acid substitutions are similarly active and, in some cases, may even provide improved activity.
[0359] Table 2. Ambermax® activation of OR7C1 variants with optimized C-termini containing single base substitutions in the first 16 amino acids of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). [Table 2-1] [Table 2-2]
[0360] Example 7: C-terminal sequence flexibility - base substitutions in amino acids at positions 17 and 18 of the C-terminal sequence motif To further test for potential variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that would provide similar benefits, 120 different point mutations were introduced into the last two amino acids of this sequence (positions 17 and 18), and all variants were fused after TM7 of OR7C1. The activity of the various modified variants was then compared to variants with the optimized sequence described above in Examples 1-5.
[0361] Therefore, all variants were transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). Cells were then stimulated with 1 μM and 3.1 μM Ambermax®, and the induction rate of luciferase was compared to that of the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which was set to 100%, and to that in an equivalent experiment performed with the wild type.
[0362] As shown in Table 3, there is flexibility in the last two amino acids and many variants are active, but there can also be complete inactivation due to incorrect amino acids at these positions.
[0363] Table 3. Ambermax® activation of optimized C-terminal variants fused to OR7C1: Effect of base substitutions at amino acid positions 17 and 18 of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) [Table 3-1] [Table 3-2]
[0364] Example 8: C-terminal sequence flexibility - truncation of amino acids at positions 17 and 18 of the C-terminal sequence motif To further test potential variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that confer similar advantages in functional expression, additional truncations of one or two amino acids were introduced and the sequence was fused to either OR7C1, OR2T4, or OR5B12 after TM7. The activity of the various modified variants was then compared to variants with the optimized sequence described above. For optimal activity, 30 μM copper (added as CuCl) was added to the assay with OR2T4.
[0365] As shown in Table 4, a sequence motif having the first 16 amino acids of RNKEVKDALKRLLKRKCC (RNKEVKDALKRLLKRK, SEQ ID NO: 10) is sufficient for high functional expression of OR7C1, indicating that the additional amino acids at positions 17 and 18 are optional and do not need to be introduced for functional expression of all ORs. However, omitting these two additional amino acids resulted in reduced activity in the case of OR2T4 and a complete loss of activity in OR5B12, indicating that adding additional amino acids at positions 17 and 18 is beneficial when these ORs are employed, or when libraries of many or all ORs are generated, because the additional amino acids allow for more extensive improvement of expression of various ORs.
[0366] Table 4. Activation of modified OR7C1, OR2T4, and OR5B12 variants with optimized C-termini: Effect of truncation of any amino acid at positions 17 and 18 of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) [Table 4]
[0367] Example 9: C-terminal sequence flexibility - Addition of additional amino acids at positions 19-22 As shown in Examples 1-5 and in Example 6, two cysteine residues at positions 17 and 18 of the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) provide improved functional expression of various receptors. As shown in Example 8, these amino acids are optional and are not required for improved expression of all ORs. As shown in Example 7, flexibility exists in these two final amino acids, and substituting one of the Cys residues with a basic amino acid can maintain or even improve activity, particularly for OR7C1, where adding an arginine as the second residue improved activity. This effect is further enhanced by variants extended with additional basic amino acids. For ease of reference, positions corresponding to extended variants are displayed using the positioning of SEQ ID NO: 86 (having 18 amino acids) as a reference; thus, positions 19-22 of SEQ ID NO: 86, as used herein, correspond to the positions present in the extended variant. In other words, reference to a variant, for example, position 19 of SEQ ID NO: 86 means that this variant is extended by one amino acid, and so on.
[0368] Thus, as shown in Table 5, having one to six basic amino acids at positions 17-22 improved the activity of O7C1. Similarly, having two to three basic amino acid residues at positions 18-20 improved the activity of OR5B12 compared to only two Cs at positions 17 and 18 (Table 6). For OR2T4, having two to three basic amino acid residues at positions 18-20 did not further enhance activity, but it was not detrimental to activity either (Table 7). This indicates that in libraries where all ORs have the same C-terminal addition, these additions are preferred because they only improve or maintain activity, not impair activity for the receptors tested. Furthermore, various amino acids added at position 19 improved the activity of OR5B12 (Table 6), which depended particularly on the additional amino acids at positions 16 and beyond (see Example 8). The methods used and the calculation of results were the same as in Examples 6-8.
[0369] Table 5. Activation by Ambermax® of OR7C1 variants with optimized C-termini containing additional basic residues at amino acid positions 18-22. [Table 5]
[0370] Table 6. Activation by Hedione® HC of OR5B12 variants with optimized C-termini containing additional basic residues at amino acid positions 17-20. [Table 6]
[0371] Table 7. Activation by cyclopentanethiol of OR2T4 variants with optimized C-termini containing additional basic residues at amino acid positions 17–20. [Table 7]
[0372] Example 10: C-terminal sequence flexibility: Combining functional base substitutions As shown in Examples 6-9, following the optimized sequences of Examples 1-5, functional variants were identified at single amino acid positions within the first 16 amino acids of RNKEVKDALKRLLKRKCC (SEQ ID NO:86) (corresponding to RNKEVKDALKRLLKRK (SEQ ID NO:10) and within the optional additional C-terminal cysteine residues at positions 17-22). These functional variants can be combined with each other to provide many different functional variants of the C-terminus, and in some combinations can even provide variants with further improved activity compared to SEQ ID NO:86.
[0373] Table 8 lists the C-terminal sequence variants described in Examples 6, 7, and 9 combined into the same C-terminal sequence fused to OR7C1. The methods used and the methods for calculating the results are the same as in Examples 6-8.
[0374] As can be seen from the results, functional variants combined at the same C-terminal sequence yielded all functional combinations, and combining variants with enhanced functionality yielded variants with even enhanced C-terminal sequences. In addition to the improved response at the two specific screening concentrations as shown in Table 8, this improvement is also evident from the dose-response analysis, as shown in the representative example in Figure 11.
[0375] Table 8. Ambermax® activation of OR7C1 variants with optimized C-termini containing multiple base substitutions selected from the active variants identified in Examples 6, 7, and 9. Residues of the variants compared to SEQ ID NO: 86 are in bold and underlined. [Table 8]
[0376] Example 11: C-terminal sequence flexibility: testing functional variants with different parent receptors As shown in Example 2, the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) functionally improves the response of multiple receptors. As further shown in Examples 6-10, functional variants of the optimized C-terminal sequences of Examples 1-5 can be identified that are still active or even have improved activity when tested with OR7C1. Furthermore, functional variants, and particularly combinations of variants from Examples 6-10, were tested for broad applicability to different ORs.
[0377] A variant of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) was fused to OR5B12 and transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). Cells were then stimulated with 31 μM and 62.5 μM Hedione® HC, and the induction rate of luciferase was compared to that for a receptor with the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which was set to 100%, and to that in an equivalent experiment performed with the wild type.
[0378] As shown in Table 9, the variant combination identified to maintain or improve expression of OR7C1 (Example 10) also significantly further improved expression of OR5B12, demonstrating that the identified variants can be generalized to other ORs.
[0379] Table 9. Activation by Hedione® HC of OR5B12 variants with optimized C-termini containing multiple base substitutions selected from the variants identified in Examples 6-9. [Table 9]
[0380] Similarly, variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were fused to OR2T4 and transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). Cells were then stimulated with 4 μM and 8 μM 2-methyl-3-tetrahydrofuranthiol, and the induction rate of luciferase was compared to that of the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which was set to 100%, and to that in equivalent experiments performed with the wild type.
[0381] As shown in Table 10, the variant combinations identified to maintain or improve expression of OR7C1 (Example 9) were all active when tested with OR2T4: for this receptor, they did not further improve functional activity; however, they did not impair activity either. This indicates that in a library in which all ORs have the same C-terminal appendage, these variants may still be preferable because they only improve or maintain activity, not impair receptor activity.
[0382] Table 10. Activation by 2-methyl-3-tetrahydrofuranthiol of OR2T4 variants with optimized C-termini containing multiple base substitutions selected from the variants identified in Examples 6-9. [Table 10]
[0383] Variants of the sequence motif RNKEVKDALKRLLKRKCC were further tested in OR5A2, as shown in Table 11. Similar to OR2T4, the identified variant was also fully functional when applied to OR5A2, but it did not further enhance the activity (which was already very high upon induction by muscone starting at 0.03 μM).
[0384] Table 11. Dose-dependent induction of OR5A2 by two different variants of the C-terminal sequence motif (induction ratio of luciferase is shown) [Table 11]
[0385] Example 12: Replacing the C-terminal sequence of a class I olfactory receptor with an optimal C-terminal sequence originally derived from a class II OR The optimized sequences in Examples 1-5 were originally derived by mutating the C-terminal sequence of class II OR5AN1. The C-terminal sequences of class I ORs diverge widely from class II sequences, resulting in different C-terminal consensus sequences for class I receptors (Kotthoff et al. 2021, supra). Therefore, we further tested whether the optimized C-terminal sequences developed starting from class II receptors could also be applied to class I receptors. Surprisingly, this sequence also resulted in a clear enhancement of the expression of several class I receptors, as already summarized in Example 2 for OR52A5, OR52E8, and OR56A4. A more detailed comparison is shown in Figures 12 and 13.
[0386] Activation of OR52E8 by odorant acids present in human sweat (Natsch et al. A specific bacterial aminoacylase cleaves odorant precursors secreted in the human axilla. The Journal of Biological Chemistry 2003; 278(8):5718-5727) was examined using various variants of the parent receptor sequence. The wild-type variant was not activated by acid. Replacing the wild-type C-terminus with the functional OR51E1 C-terminal sequence did not provide functional OR52E8. However, using the optimized sequences from Examples 1-5 resulted in a strong signal upon addition of 3-methyl-3-hydroxyhexanoic acid.
[0387] Activation of OR56A4 by acids of different chain lengths was further tested. The wild-type was activated by decanoic acid and undecanoic acid at high concentrations, but not by nonanoic acid. Replacing the wild-type with the functional OR51E1 C-terminal sequence reduced activity. However, using the optimized sequences from Examples 1-5 resulted in a strong signal and a much lower detection threshold upon addition of all three acids.
[0388] Example 13: Replacing the C-terminal sequence of a less functional receptor with the C-terminal sequence of a functional receptor is compared to adding the optimal C-terminal sequence of the present disclosure. Given the importance of the C-terminal sequence for correct expression, as observed in the previous example, an obvious way to improve expression may be to simply provide a hybrid receptor by replacing the C-terminal domain of a non-functional or poorly functional receptor with that of a functional receptor. This approach has been tested previously (see Ikegami et al., 2020, supra), whereby the C-terminus of the poorly expressed mouse Olf541 was replaced with that of the well-expressed Olfr539, but this did not result in enhanced Olfr541 expression. However, this approach was further tested in comparative examples using different human ORs. Thus, the C-terminal sequence of OR8K3 was replaced with that of the ambrettride receptor OR1N2 or the musk receptor OR5AN1. The wild-type versions of both of these receptors were functional, indicating that they could function correctly with their native C-terminal sequences and, therefore, that this C-terminal sequence could provide largely functional expression. However, modified OR8K3 variants with either of these C-terminal sequences showed no functional response to the cognate ligand, menthol (Figure 14), whereas modified receptors with C-terminal sequences according to Examples 1-5 resulted in robust functional expression. Thus, the improvements achieved by the modified C-terminal sequences of the present disclosure cannot be achieved by simply generating chimeric receptors that combine the functional C-terminal domain of a functional receptor with a poorly expressed receptor.
[0389] Similarly, the C-terminal sequence of OR5AN1 was replaced with that of the ambretride receptor OR1N2. The wild-type versions of both of these receptors were functional, indicating that they could function correctly with their native C-terminal sequences. However, chimeric OR5AN1 variants with the C-terminal sequence of OR1N2 showed no functional response to the cognate ligands, musk ketone or muscone (Figure 15).
[0390] Similarly, the musk receptor OR5A2, which has the C-terminal sequence of OR1N2, has lower sensitivity compared to the variant receptor of the present disclosure, which has the C-terminal sequence RNKEVKDALKRLLKRKCC (Figure 16). Thus, the improvement achieved by the modified C-terminal sequence of the present disclosure cannot be achieved by simply generating a chimeric receptor that combines the functional C-terminal domain of a functional receptor with a poorly expressed receptor.
[0391] Example 14: Screening for odorants with desired Arborone odor qualities using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) A modified variant of OR7A17 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) as described in Example 2 was tested against Arborone, the main characteristic odor component of the composite Iso E Super (Hong & Corey. Enantioselective syntheses of georgyone, arborone, and structural relatives. Relevance to the molecular-level understanding of olfaction. Journal of the American Chemical Society 2006; 128(4): 1346-1352). As shown in Figure 17, Arborone activated OR7A17 with the optimized C-terminus already at concentrations as low as 1 nanomolar, i.e., approximately 50-fold lower than Iso E Super, indicating that the hybrid OR7A17 with the optimized C-terminus is a powerful tool for screening for the desired woody characteristics of Arborone. Thus, a number (n=22) of woody fragrance molecules were tested against OR7A17, which has the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), for both in vivo odor detection thresholds (OTH) in humans and in vitro EC50. As shown in Figure 18, the in vitro data well predicted the low odor detection thresholds of Ambrofix®, Georgywood®, and Iso E super (highlighted, with chemical structures at the bottom left), distinguishing them from odorants of intermediate potency and the much less active (both in vitro and in vivo) cedrol (highlighted at the top right).
[0392] Example 15. Screening for receptor activation by major human sweat odorants The most pungent odorant in human axillary sweat, with an ODT of 1 picogram per liter of air, is 3-methyl-3-sulfanylhexanol (also known as 3-methyl-3-mercaptohexanol) (Natsch et al. Identification of Odoriferous Sulfanylalkanols in Human Axilla Secretions and Their Formation through Cleavage of Cysteine Precursors by a CS Lyase Isolated from Axilla Bacteria. Chemistry and Biochemistry 2004; 1:1058-1072). To date, no OR has been reported that can selectively detect this major odorant responsible for undesirable axillary odor in human subjects. Screening a library of receptor variants with the C-terminal sequence of SEQ ID NO:86 against a library of 52 sulfur compounds (see Example 5) in the presence of copper identified the previously orphan receptor OR2M2, which is activated by 3-methyl-3-mercapto-hexanol and closely related chemicals, but not by other sulfur odorants (Figure 19). Dose-response analysis (Figure 20) showed that these identified ligands were able to activate OR2M2 at concentrations as low as 1 μM only when tested in the presence of copper, an effect that has been observed for other receptors that respond to sulfur compounds. Thus, using the improved sequence of the present disclosure, it is now possible for the first time to deorphanize OR2M2, and the modified OR2M2 with an optimized C-terminus can then be used as a screening target for the most potent human axillary malodors in combination with one of its cognate ligands, 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, or 4-methoxy-2-methylpentane-2-thiol, in the presence of copper.
[0393] The complete amino acid and DNA sequences of the OR2M2 modified sequence having the modified C-terminal domain of SEQ ID NO:86 are provided as SEQ ID NOs:245 and 246.
[0394] Example 16: Random permutations of variable positions within the modified C-terminus of the present disclosure (SEQ ID NO: 1) General linked to OR7D4 [ka] Random permutations of OR7D4 were generated based on the degenerate oligonucleotide for the C-terminal domain. For this experiment, the MfeI / XhoI fragment from pcDNA3.1(+)-mmLucy-FLAG-rho-OR7D4, containing the CMV promoter and OR coding sequence, was used to replace the EcoRI / SalI fragment of pRDVCCB-CMV-dCas9-VPH-2A-Blast (Cellecta, Inc., Mountain View, USA). The coding region of the C-terminal domain of OR7D4 (between Bsu36I and NotI) was then replaced with a complementary degenerate oligonucleotide. Random clones were selected, sequenced, and tested for the correct open reading frame. The correct open reading frame of OR7D4 and OR7D4 were confirmed. As the C-terminal domain [ka] A total of 55 clones with permutations of the wild-type sequence were obtained. All the different clones were compared to the wild-type sequence in a dose-response study using the ligand androstenone in a functional assay in HEK cells. The wild-type sequence was not significantly induced by 1 μM of the ligand, whereas the induction of the different variants by androstenone was 10.1- to 73.6-fold at this concentration. The EC2 for a 2-fold induction of luciferase for the wild-type was 1.71 μM, but decreased to 0.003- to 0.13 μM for the different variants, corresponding to a 13- to 545-fold improvement in assay sensitivity with the modified C-terminal sequence, with a median sensitivity improvement of 125-fold. These results demonstrate that all of the tested random permutations of the common sequence resulted in functional assays with significantly improved sensitivity (more than 10-fold).
[0395] [Table 12-1] [Table 12-2] [Table 12-3]
[0396] The results in Table 12 were evaluated to determine the best amino acid residue (i.e., the residue resulting in the lowest median EC2 for sequences containing it at a given position) at each variable position in the generic sequence. Based on this analysis, the statistically optimal sequence among the generic sequences was RNRDVRKALRRLFRKK (SEQ ID NO: 307). Because R at position 7 is at least as active as K (see Table 2 in Example 6), the optimal sequence was also RNRDVRRALRRLFRKK (SEQ ID NO: 308). The data from Table 12 was then further analyzed for the number of residues that differ from the statistically optimal sequence for each sequence permutation. As shown in Table 13, sequences closest to SEQ ID NO: 307 have good overall activity. Thus, while all tested random permutations within the generic sequence are active, the most active sequences are clustered among those closest to the optimal sequence. Thus, by way of example, sequences with 2-4 residue differences from SEQ ID NO: 307 are particularly active, resulting in a greater than 100-fold improvement in sensitivity over wild-type, with only one exception.
[0397] Table 13. Relationship between distance and activity between random sequence permutations in Table A and optimal SEQ ID NO: 307. [Table 13]
[0398] Example 17: Improved functional expression with a different amino acid at position 7 of the generic sequence [ka] To test for possible variants that result in similar improved functional expression of [ka] A different amino acid was introduced at position 7 in OR7C1 (represented by x), and the variant was fused after TM7 with OR7C1. As shown in Figure 2, OR7C1 with the C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) shows approximately 10-fold luciferase induction at 1 μM Ambermax® and approximately 20-fold induction at 3.1 μM Ambermax®, while the wild-type is inactive. The activity of the different modified variants was then compared as described in Example 6. The results show that position 7 of SEQ ID NO: 86 has high flexibility, and all 20 amino acids except proline show superior activity compared to the wild-type.
[0399] Table 14. Ambermax® activation of OR7C1 variants with optimized C-termini containing a single base substitution at position 7 of the C-terminal motif RNKEVKDALKRLLKRKCCKCC (SEQ ID NO: 86). [Table 14]
[0400] Example 18: Screening a library of class II ORs all having the same improved C-terminal domain with a single odorant All human class II ORs and all major genetic variants of these ORs were identified in the same [ka] A library of class II ORs was generated in the vector pcDNA3.1(+) using the sequence encoding the following:
[0401] Table 15. Class II OR library members with modified C-terminal domains [Table 15-1] [Table 15-2] [Table 15-3] [Table 15-4] [Table 15-5] [Table 15-6] [Table 15-7]
[0402] SEQ ID NOs: 331-739 also contain a 5' BamHI restriction site (GGATCC) and Kozak sequence (GCCACC), and a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes.
[0403] Parallel transfection experiments were performed using all these plasmids (n = 408), in which cells expressing various ORs were stimulated with a given ligand one at a time, at the highest concentration tested without apparent cytotoxicity. A total of 26 ligands were tested against the complete library of class II receptors. For 25 of these ligands, at least one receptor was identified that showed luciferase induction greater than fourfold over background, indicating that these improved OR libraries are capable of identifying cognate receptors for the vast majority of compounds. In general, screening with this library yields a very good signal-to-noise ratio, clearly separating positive screening hits from inactive ligand-OR associations. As an example, screening of the complete library with the ligand patchoulol is shown in Figure 21. Patchoulol resulted in strong induction of OR14J1 and weak induction of OR11A1 and OR7A17. The complete set of ORs identified from this deorphanization campaign is shown in Table 16.
[0404] The selected ligand-OR pairs were then tested in a confirmatory study using dose-response analysis. For all tested ligand-OR pairs from the newly deorphanized receptors (n=31), a clear dose response was obtained when these screening hits were validated in a confirmatory study (see, for example, the dose response of patchoulol for ORs listed in Example 19 and OR14J1 in Figure 22). This indicates that with this improved OR library with a high signal-to-noise ratio, this screening yields highly reproducible ligand-OR associations. This can be compared to the state-of-the-art technique described in Mainland et al. (2015) Sci Data 2:150002, which uses OR expression using RTP1S and an N-terminal rho tag but employs wild-type OR sequences. In this screening, a comprehensive OR library (class I and class II; 511 clones covering major variants) was tested against 73 odorants. Although the initial screen yielded 1572 odorant / receptor pair hits (covering 394 ORs), only 63 clones representing only 27 ORs could be validated for association in a dose-response analysis, indicating significant noise in the primary screen due to poor or no expression of wild-type sequences.
[0405] Table 16. Ligands tested against the complete library of ORs and the cognate receptors identified for these ligands [Table 16]
[0406] Example 19: Activation of ORs Identified in a Screen Using a Library of All Human ORs with Improved C-Terminal Domains - Comparison of Wild-Type and Improved ORs [ka] The improved functional expression of receptors from the OR library described in Example 18, which has the following structure, was further analyzed. For selected receptors, wild-type sequences were synthesized and cloned into pcDNA3.1(+). Dose-response analysis with the cognate ligand was then performed on both the wild-type and modified sequences. The data were analyzed for induction threshold (concentration for 2-fold luciferase induction), EC50 (potency), and maximal induction (efficacy), as also performed in Table 1. As shown in Table 17, for the 13 ORs identified in the screen in Example C, the wild-type was completely inactive, whereas ORs with optimized C-terminal domains were activated by their cognate ligands in the low micromolar range (an all-or-nothing effect). This demonstrates that these ORs can be deorphanized only by this library of ORs with improved expression due to C-terminal modifications. For the nine additional tested ORs, the improved sequences result in clearly lowered detection thresholds (5.2-163 fold) and increased efficacy (1.9-16.8 fold).
[0407] [Table 17-1] [Table 17-2]
[0408] Example 20: Optimization of expression with SEQ ID NO: 221 compared to SEQ ID NO: 86 All tested receptors in Table 1 resulted in improved functional expression when the wild-type C-terminal domain was replaced with SEQ ID NO: 86. As shown for OR7C1 in Table 8 and OR5B12 in Table 9, this functional expression was further improved when SEQ ID NO: 86 was replaced with SEQ ID NO: 221. This additional improvement was further tested with OR10H5 (Table 18) and OR7A17 (Table 19). In both cases, SEQ ID NO: 221 further improved functional expression compared to the already strongly improved functional expression over wild-type when SEQ ID NO: 86 was used. The complete DNA sequences encoding the modified receptors are SEQ ID NO: 684 for OR10H5 and SEQ ID NO: 611 for OR7A17.
[0409] Table 18. Improved functional expression of OR10H5 [Table 18]
[0410] Table 19. Improved functional expression of OR7A17 [Table 19]
[0411] Example 21: Improved C-terminal domains of class I ORs A further variant of the improved C-terminal domain sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which corresponds to the general sequence RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1) and has additional amino acids at the C-terminus (CC), was tested against the class I receptor OR52A5, which is activated by 4-ethyloctanoic acid. As shown in Table 20, SEQ ID NO: 86 results in clearly improved expression compared to the wild-type. This activation is further improved by replacing amino acid positions 4-6 with the sequence QIR or by adding the terminal sequence RRR.
[0412] These two improvements were further combined and the combination was tested against OR52A5 and two additional class I receptors. As shown in Table 21, all three class I receptors: [ka] In the case of OR52A4, 4-ethyloctanoic acid, which is detected at low concentrations by the human nose, was detected by the receptor with the C-terminal modification with a high luciferase response already at 0.19 μM, demonstrating highly sensitive detection of this carboxylic acid by the modified class I receptor.
[0413] Table 20. Optimized C-terminal domain for class I OR52A5 [Table 20]
[0414] Table 21. Optimized C-terminal domains for class I OR52E8, OR56A4, and OR52A5 [Table 21]
[0415] Thus, the general sequences of further improved C-terminal domains are RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 822), RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR]RRR (SEQ ID NO: 823), and RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR]RRR (SEQ ID NO: 824).
[0416] Example 22: Screening a library of class I ORs all with the same improved C-terminal domain The same in the vector pcDNA3.1(+) [ka] Using the sequences encoding all human class I ORs and all major genetic variants of these ORs, a library of all class I ORs with improved C-terminal domains was synthesized. A complete list of class II ORs included in this library is shown in Table 22:
[0417] Table 22. Library members of class I ORs [Table 22-1] [Table 22-2] [Table 22-3]
[0418] SEQ ID NOs:742-819 also contain a 5' BamHI restriction site (GGATCC) and Kozak sequence (GCCACC), and a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes.
[0419] Parallel transfection experiments were performed using all of these plasmids (n=77, SEQ ID NO: 760; OR51F1 was excluded due to poor plasmid yield), and cells expressing various ORs were stimulated with one given ligand at a time. A total of 10 different carboxylic acid ligands were tested against the complete library of class I receptors. For eight of these ligands, at least one receptor was identified with luciferase induction greater than three-fold over background (Table 23), indicating that these improved OR libraries can identify cognate receptors, particularly for carboxylic acids.
[0420] As an example, Figure 23 shows the screening of the complete library with 3-methyl-2-hexenoic acid, an important component of human sweat (Natsch et al. A specific bacterial aminoacylase cleaves odorant precursors secreted in the human axilla. Journal of biological chemistry 2003; 278(8):5718-5727). 3-Methyl-2-hexenoic acid strongly induced the OR51B2 (C120R, L134F, C209S) variant but not the wild-type OR51B2 variant. This contradicts Li et al. (PLoS Genet. 2022 Feb 3;18(2):e1009564), who reported that the wild-type, but not the (C120R, L134F, C209S) mutant variant of OR51B2, was activated by 3-methyl-2-hexenoic acid. Therefore, contrary to the expression data in that publication, the (C120R, L134F, C209S) variant of OR51B2 is a preferred target for screening for antagonists against human sweat odor. In addition, 3-methyl-2-hexenoic acid also activates both tested variants of OR52K1, thus making them new targets for screening for malodor antagonists. Next to 3-methyl-2-hexenoic acid, 4-methyl-3-hexenoic acid also activates the (C120R, L134F, C209S) variant of OR51B2. This acid also activates both tested variants of OR51B5. 4-methyl-3-hexenoic acid is an important cause of typical malodors in laundry (Kubota et al., Appl Environ Microbiol. 2012;78(9):3317-24). Therefore, additional screening for antagonists against OR51B5 can be used to target this specific malodor.Screening the complete library of class I ORs for 3-methyl-3-hydroxyhexanoic acid, the most dominant carboxylic acid in human sweat odor, yielded only one hit, OR52E8, confirming the above results that the OR52E8 variant with improved expression due to an optimized C-terminal domain is an important target for screening for antagonists against this acid. Next to ORs for these short-chain odorant acids, OR52K1, OR52K1(Q52R), OR56A1, OR56A3, OR56A3(M51T), and OR56A4 are activated by the C8-C11 chain acids tested in this screen.
[0421] Table 23. Ligands tested against the complete library of class I ORs and the cognate receptors identified for these ligands. [Table 23]
[0422] Example 23: Screening and matching of complex fragrances (i.e., complex odorant mixtures) using a library of class II ORs that all have the same improved C-terminal domain The human class II OR library of Example 18 was further tested with three fragrance oils, i.e., complex mixtures of over 24 single fragrance oil components mixed in specific ratios to produce a desired overall odor. "Fragrance A" and "Fragrance B" are independent fragrance formulations with different overall odor impressions, while "Fragrance A mod" shares 66% of its components with "Fragrance A" and retains the overall odor of "Fragrance A" but is modified by replacing 33% of the single components. A total of 18 ORs and OR variants were activated by at least one of these oils, with at least two-fold luciferase induction. The induction of activated subsets of ORs by these three oils is shown in Figure 24. As is evident from the black and gray bars, "Perfume A" and "Perfume A mod" induce similar patterns of OR activation, whereas "Perfume B" is clearly different, especially with regard to the activation of OR10G7, OR11G2, OR1N2, OR2AJ1, OR2J2, OR3A3, OR5AN1, OR5B12, and OR5P3. Thus, Perfume B has a different "fingerprint" of OR activation, reflecting its different olfactory properties. The difference / similarity of different oils can also be quantified using distance measurements. For example, the distance between two compositions i and j can be calculated, for example, according to the following formula:
number
[0423] For the 18 ORs in Figure 24, the distance between "Fragrance A" and "Fragrance A mod" is 2.2, while the distance between "Fragrance A" and "Fragrance B" is 6.4 and the distance between "Fragrance A mod" and "Fragrance B" is 5.5, indicating the similarity between "Fragrance A" and "Fragrance A mod" and the difference between both of them from "Fragrance B".
[0424] These results demonstrate that the approach of screening a sensitive library of ORs with modified C-terminal domains in complex mixtures can be used to measure the similarity of complex odorant mixtures, such as fragrances or flavors, and to provide an objective representation of the olfactory profile of such mixtures.
[0425] In some countries, due to regulatory pressure, enigmatic fragrance ingredients are sometimes banned. Therefore, it is desirable to design fragrances that smell very similar to the consumer but replace the banned ingredient. Classically, after the removal of the banned ingredient, a single ingredient was used to replace the olfactory deficiency. However, a single ingredient may activate multiple ORs, and in many cases, multiple ingredients must be removed. Therefore, it is important to replace the OR activation pattern of the removed ingredient and reconstruct the overall OR activation pattern of the original fragrance, which can be done with one or more ingredients.
[0426] This can be achieved using a method in which single components are screened for OR activation against a complete OR library with modified C-terminal domains, as shown in Examples 18 (Class II) and 22 (Class I), and their OR activation patterns are recorded to create a database of single-component activation patterns. The original perfume oil and the original perfume oil with the scrutinized component removed are then both screened against the complete OR library, as shown above. After adding the selected replacement component, the resulting perfume can then be validated by testing it again against the complete OR library, as shown above, and its similarity to the original perfume can be calculated in an objective manner.
[0427] Example 24: Detection of homologous odorants in complex mixtures or reaction mixtures of low purity Classically, perfume ingredients and experimental fragrance ingredients (research materials) must be highly pure to be evaluated by perfumers. This is because the human nose has approximately 400 simultaneously functional ORs, and any odorant impurities will therefore affect the overall olfactory impression of the sample, making it difficult for humans to judge samples of limited purity. Assays based on a single or small number of expressed receptors, on the other hand, focus on one specific odor type and can, in principle, detect odorants with that specific odor type against a complex background. However, in the case of poorly expressed receptors, the matrix quickly reaches cytotoxic concentrations, hindering the assay when the active ingredient directed toward the target odor to be explored is present at low concentrations.
[0428] To investigate the sensitivity of the modified olfactory receptor of the present disclosure to detect odorants in a complex odorant matrix, the ligand Ambrofix® was added to Egyptian geranium oil, which contains more than 1% of 13 different materials and serves as an example of a complex essential oil, i.e., a complex background matrix of potent odorant materials. The Ambrofix® addition levels were 0%, 0.1%, 0.316%, 1%, 3.16%, and 10%. These mixtures were tested with OR7A17 (the DNA sequence encoding the modified receptor is included as SEQ ID NO: 611), which has a C-terminal domain of SEQ ID NO: 221, as described in Example 20. As shown in Figure 25, the added oil significantly induced luciferase expression over background. Thus, for a potent ligand such as Ambrofix®, a highly sensitive assay can detect concentrations of at least 0.1% of the target component in a complex matrix.
[0429] Example 25: Screening a library of ORs with a mixture of odorants with a specific odor class. As shown by the case of Ambrofix® in Example 18, which activates specific receptors OR7E24 and OR7A17, a single odorant may cause activation of multiple ORs. On the other hand, multiple odorants may be perceived and described as a common odor descriptor due to their common activation of a given set of ORs. Therefore, to induce a specific odor sensation, a specific set of ORs may need to be activated. This specific set of ORs can be identified by mixing several odorants with a given odor type. This mixture is then screened against a complete library of ORs. The identified set of ORs can then be further used to screen for that specific odor type. For example, to identify a set of representative ORs for the "fruity ester" odor category, a mixture of the following compounds was created: pentanoic acid, 2-methylethyl ester; butanoic acid, 3-methyl-, ethyl ester; acetic acid, phenoxy-, 2-propenyl ester; hexanoic acid, ethyl ester; acetic acid, (3-methylbutoxy)-, 2-propenyl ester; acetic acid, (cyclohexyloxy)-, 2-propenyl ester; butanoic acid, 2-methyl-, (3Z / E)-3-hexenyl ester; cyclohexanecarboxylic acid, ethyl ester; and oxiranecarboxylic acid, 3-phenyl-, ethyl ester. This mixture was screened against the full library of ORs as described in Example 18. This approach identified a specific set of ORs activated by ligands for this specific odor class: OR11G2, OR11G2(I65N, V82I), OR1D2, OR2AK2(S84N), and OR2L5 were identified as the set of ORs most strongly activated by the fruity ester mixture.
[0430] Table 24. ORs identified by screening of class II ORs with a mixture of fruity esters [Table 24]
[0431] Similarly, a mixture of odorants with the fruity lactonic descriptor containing equal amounts of dodecalactone delta, heptalactone gamma, octalactone gamma, nonalactone gamma, undecalactone gamma, decalactone gamma, dodecalactone gamma, decalactone delta, methyl tuberate, and nectaryl was prepared and screened against a full library of class II ORs. This mixture was found to activate a set of ORs: OR10A3, OR10A6 (A117V, V140G, L287P), OR10J1 (M51I, I92M), OR1D2, and OR2J2.
[0432] Table 25. ORs identified by screening of class II ORs with a mixture of fruity lactones [Table 25]
[0433] In subsequent deconvolution experiments, single components were tested against all identified ORs in dose-response studies. Table 26 lists the concentrations for 2-fold luciferase induction (10-fold in the case of OR10A3). These data indicate that all identified ORs are activated by single components of the mixture, albeit with some differences in specificity. Thus, OR2AP1 and OR2J2 are particularly strongly activated by the long-chain dodecalactone gamma, while OR10A3 is highly sensitive to several longer-chain lactones. Overall, undecalactone gamma was the most potent ligand for four of the five identified ORs, consistent with the fact that, among lactones, undecalactone gamma has the lowest olfactory detection threshold in vivo. Among the five ORs, OR10A3 was the most sensitive. Thus, for the most potent ligand, undecalactone gamma, a 10.2-fold activation is observed already at 0.31 μM.
[0434] Table 26. ORs identified by screening of class II ORs with a mixture of fruity lactones were tested with individual lactones in a dose-response analysis. [Table 26] OR10A3 is a highly sensitive and highly potent receptor, and therefore, the EC10, the concentration for 10-fold OR activation, is shown here.
[0435] Example 26: Screening of odorants using cells expressing multiple receptors To screen for a particular odor class, some or all members of a given set of ORs identified as specific for a given odor class can be combined in a screen for new odorants or new odorant mixtures, as shown in Example 25. Screening for different ORs can be performed with individual ORs of the OR set in serial or parallel assays. In addition, some or all of the ORs of the OR set specific for that odor class can also be co-expressed in a single cell line. Thus, to screen for woody ambery molecules, cells were transfected separately or simultaneously with two plasmids encoding optimized OR7A17 and OR7C1, both of which have the C-terminal domain of SEQ ID NO: 221. The cells were then stimulated with Iso E super, a specific ligand for OR7A17, and Ambermax®, a specific ligand for OR7C1. As shown in Figure 26, cells expressing only OR7A17 respond only to Iso E super, while cells expressing OR7C1 respond only to Ambermax®. Cells expressing both receptors are functional sensors for the Amberly Woody odor, detecting both ligands down to low concentrations.
[0436] Example 27. Determination of the more potent ligand for OR10J5, (S,E)-10-hydroxy-4,8-dimethyldec-4-enal and (R,E)-10-hydroxy-4,8-dimethyldec-4-enal Activation of OR10J5, which has the C-terminal domain of SEQ ID NO: 86, by (S,E)-10-hydroxy-4,8-dimethyldec-4-enal or (R,E)-10-hydroxy-4,8-dimethyldec-4-enal was measured in a dose-response analysis.
[0437] The potencies of the two tested OR10J5 ligands were expressed as EC20% values, which are the concentrations resulting in a 20% increase in luciferase activity compared to the positive control (100 μM, Mahonial®; (4E)-9-hydroxy-5,9-dimethyl-4-decenal). (S,E)-10-hydroxy-4,8-dimethyldec-4-enal had an EC20% value of 10.1 μM, while (R,E)-10-hydroxy-4,8-dimethyldec-4-enal had an EC20% value of 24.4 μM. Due to the difference in EC20%, the concentration of the (S,E)-isomer required to activate the receptor was much lower than that of the (R,E)-isomer. Figure 28 graphically shows the responses of both test compounds. This example further demonstrates the utility of the improved assay for determining the most odor-active isomer.
Claims
1. An olfactory receptor protein having a modified C-terminal domain containing the amino acid sequence motif RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1).
2. The olfactory receptor protein of claim 1 , wherein the modified C-terminal domain is fused to the seventh transmembrane helix of the protein.
3. The olfactory receptor protein according to claim 1 or 2, which is a class I or class II olfactory receptor having a modified C-terminal domain, preferably a human, canine or feline class I or class II olfactory receptor having a modified C-terminal domain, more preferably a human class I or class II olfactory receptor having a modified C-terminal domain.
4. Class II receptors include OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR 5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2, preferably wherein the class II receptor is selected from the group consisting of OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2, and OR5A2.
5. The olfactory receptor protein of claim 4, wherein the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1.
6. The olfactory receptor protein according to any one of claims 1 to 5, wherein the sequence motif is RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 5).
7. The olfactory receptor protein according to any one of claims 1 to 6, wherein x is not proline.
8. The olfactory receptor protein according to any one of claims 1 to 7, wherein x is neither proline nor tryptophan.
9. The olfactory receptor protein according to any one of claims 1 to 8, wherein x is selected from D, K, R, E, N, V, A, Q or G, and preferably x is selected from D, K, R, E, N, V, A or Q.
10. 10. The olfactory receptor protein of any one of claims 1 to 9, wherein the amino acid sequence motif comprises 1 to 6 additional C-terminal amino acid residues, and optionally wherein: - the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from C, R, K, E, G, H, F, P, Y, more preferably the first additional amino acid residue is C, R or K, most preferably C; - the second additional amino acid residue is selected from C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably C or R; - the third additional amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K; the fourth, fifth and sixth amino acid residues are selected from K and R; Preferably, wherein the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRR (SEQ ID NO: 164), and CRRRKK (SEQ ID NO: 165), The olfactory receptor protein.
11. The olfactory receptor protein according to any one of claims 1 to 10, wherein the sequence motif is selected from the group consisting of SEQ ID NOs: 1, 5 to 75, 86 to 130, 133 to 147, 149 to 151, 154, 156 to 158, 166, 167, 198, 219 to 221, 254 to 312, 319 to 326, 328 to 331, 740 to 741, and 820 to 890.
12. A nucleic acid molecule comprising a nucleotide sequence encoding an olfactory receptor protein as described in any one of claims 1 to 11, wherein the nucleic acid molecule preferably further comprises a promoter sequence, more preferably a constitutive promoter sequence.
13. 13. An expression vector comprising a nucleic acid molecule as defined in claim 12, preferably wherein the expression vector is a plasmid.
14. A recombinant host cell, preferably a HEK293 cell or a HEK293T cell, comprising a nucleic acid molecule as described in claim 12 or an expression vector as described in claim 13, wherein the cell preferably expresses an olfactory receptor protein as described in any one of claims 1 to 11.
15. The recombinant host cell of claim 14 , wherein the cell further expresses one or more olfactory receptor accessory proteins.
16. A library comprising a diverse repertoire of olfactory receptor proteins as described in any one of claims 1 to 11, nucleic acid molecules as described in claim 12, expression vectors as described in claim 13, or recombinant host cells as described in claim 14 or 15.
17. The library of claim 16, wherein a diverse repertoire of olfactory receptor proteins, olfactory receptor proteins encoded by nucleic acid molecules or expression vectors, or olfactory receptor proteins expressed by recombinant host cells share the same modified C-terminal domain.
18. 1. A method for identifying an olfactory receptor ligand, comprising: a) providing a recombinant host cell expressing an olfactory receptor protein as defined in any one of claims 1 to 11 or an olfactory receptor protein as defined in claim 14 or 15; b) contacting the receptor or recombinant host cell with a test compound or composition; and c) Detecting the activation of olfactory receptors Including, Preferably, the olfactory receptor is selected from the group consisting of OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)), OR7A10, OR10H2, OR10H1, OR10D3, OR 1D2, OR2A5, OR2A25 (OR2A25(S75N, A209P)), OR11G2 (or OR11G2(I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3(S73G)), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2, and OR2AG2 (preferably OR2AG2(Y28C)), The method.
19. 1. A method for identifying an enhancer or antagonist of an olfactory receptor, comprising: a) providing a cell expressing an olfactory receptor protein as defined in any one of claims 1 to 11 or an olfactory receptor protein as defined in claim 14 or 15; b) contacting the receptor or recombinant host cell with a cognate ligand and a test compound or composition; and c) detecting increased or decreased activation of olfactory receptors compared to a ligand-only control; Including, Preferably, the olfactory receptor is selected from the group consisting of OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)), OR7A10, OR10H2, OR10H1, OR10D3, OR 1D2, OR2A5, OR2A25 (OR2A25(S75N, A209P)), OR11G2 (or OR11G2(I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3(S73G)), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2, and OR2AG2 (preferably OR2AG2(Y28C)), The method.
20. 20. The method of claim 19, wherein the method is for identifying an olfactory receptor antagonist, and wherein the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1, and OR4S2, preferably OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and OR5V1, more preferably OR2M2 or OR2V1.
21. 20. The method of claim 19, for identifying an olfactory receptor antagonist, wherein the olfactory receptor is OR2M2 or OR2V1, and wherein step b) further comprises contacting the receptor or recombinant host cell with a copper salt, preferably wherein the cognate ligand is selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, and 4-methoxy-2-methylpentane-2-thiol.
22. 20. The method of claim 19, wherein the method is for identifying an olfactory receptor antagonist, wherein the olfactory receptor is OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and preferably wherein the cognate ligand is 3-methyl-2-hexenoic acid.
23. 20. The method of claim 19, wherein the method is for identifying an olfactory receptor antagonist, wherein the olfactory receptor is OR5V1, and preferably wherein the cognate ligand is 2,4,6-trichloroanisole.
24. 1. A method for identifying an olfactory receptor capable of binding to a target ligand, comprising: a) providing a library according to claim 16 or 17; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins with a target ligand; and d) Identifying the olfactory receptors activated by the target ligand The method comprising:
25. 1. A method for generating an objective representation of the olfactory properties of a test compound or composition, comprising: a) providing a library according to claim 16 or 17; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a test compound or composition; and d) Detecting the activation of each of the olfactory receptor proteins The method comprising:
26. 1. A method for assessing differences or similarities between two or more test compounds or compositions, comprising: a) providing a library according to claim 16 or 17; b) optionally obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from said library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of two or more test compounds or compositions; d) detecting activation of each of the olfactory receptor proteins for each of two or more test compounds or compositions; and e) Comparing the activated olfactory receptor protein between each of two or more test compounds or compositions. The method comprising: