Modified olfactory receptors
By modifying the C-terminal domain of the olfactory receptor protein, combining it with appropriate nucleic acid molecules and recombinant host cells, an olfactory receptor library is constructed, which solves the problem of low sensitivity in olfactory receptor expression and screening in the existing technology and achieves more sensitive olfactory receptor expression and screening at physiologically relevant concentrations.
Patent Information
- Application Number
- CN202380092741.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-25
- Filing Date
- 2023-12-15
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies make it difficult to effectively express and screen olfactory receptors at low concentrations, resulting in low sensitivity, inability to fully cover the receptor-ligand space, and difficulty in identifying new olfactory receptor ligands and antagonists.
By modifying the C-terminal domain of the olfactory receptor protein and combining it with appropriate nucleic acid molecules, expression vectors and recombinant host cells, an olfactory receptor library is constructed, the functional expression and screening methods of the olfactory receptors are optimized, and the sensitivity of the assay is enhanced.
This enables more sensitive olfactory receptor expression and screening at physiologically relevant concentrations, enabling the identification of more olfactory receptor-ligand pairs, including poorly soluble and cytotoxic ligands, and improving the accuracy and efficiency of screening.
Smart Images

Figure BDA0005522380490000671 
Figure BDA0005522380490000971 
Figure BDA0005522380490000981
Abstract
Description
Technical Field
[0001] Various aspects and embodiments described herein relate to the fields of biotechnology and flavors and fragrances, and in particular to modified olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses for expressing olfactory receptors and for identifying novel olfactory receptors and novel olfactory receptor ligands, enhancers, and antagonists. Background Art
[0002] Olfactory or odorant receptors (ORs) are expressed in the olfactory sensory neurons of the olfactory epithelium and are responsible for the detection of odors. Olfactory receptors belong to the G protein-coupled receptor superfamily (GPCR). Activation of ORs by odorants (ligands) activates olfactory-specific G proteins, which then promote the production of cyclic AMP (cAMP) via type III adenylate cyclase. The level of intracellular cAMP increases, leading to the opening of cyclic nucleotide-gated ion channels, which allow calcium ions to enter the cell, depolarize olfactory sensory neurons, and trigger action potentials that carry information to the olfactory corpuscles of the olfactory bulb.
[0003] The human genome encodes approximately 400 different functional olfactory receptors. A specific olfactory receptor can be activated by more than one ligand molecule, and a specific ligand molecule can activate multiple olfactory receptors, resulting in a highly complex network of interactions between the OR and the ligand repertoire. Elucidating these interactions could allow the discovery of new flavor and fragrance ingredients, or compounds that are more sustainable and / or easier to produce than currently used compounds, such as odor enhancers. For many of the approximately 400 different olfactory receptor genes, different variants, called alleles or haplotypes, exist in the human population. Protein products from different alleles or haplotypes of the same gene can have different ligand selectivity or sensitivity.
[0004] In addition to being expressed in olfactory sensory neurons of the olfactory epithelium, OR expression is also found in many other cells and tissues (Feldmesser et al., Widespread ectopic expression of olfactory receptor genes. BMC Genomics 2006; 7: p. 121; Massberg, D. and H. Hatt. Human Olfactory Receptors: Novel Cellular Functions Outside of the Nose. Physiol Rev 2018; 98(3): 1739-1763). These ORs have been found to be associated with many important diseases, so in addition to their primary interest as odor ligand targets, ORs are also of great interest as targets for the search for agonists and antagonists to treat various diseases (Lee et al., Therapeutic potential of ectopic olfactory and taste receptors. Nature Reviews Drug Discovery 2019; 18(2): 116-138). In particular, OR signaling has been associated with decreased or increased cell proliferation in bladder, colon, prostate, lung, and liver cancer cells. Similarly, OR expression in inflammatory cells has been associated with the regulation of inflammation. Thus, OR screening has multiple applications beyond olfaction.
[0005] Effective screening of olfactory receptors requires their expression in cultured cell lines, which typically involves introducing the olfactory receptor gene into cells, followed by stable or transient overexpression. Functional expression of olfactory receptors in host cells has generally proven difficult, as obtaining correct folding of the receptors and / or proper insertion of the receptors into the cell membrane has proven challenging. Consequently, several approaches have been attempted to improve functional heterologous expression of olfactory receptors.
[0006] Functional heterologous olfactory receptor expression using expression systems currently available in the art generally requires the co-expression of accessory proteins of the receptor transporter (RTP) family, such as RTP1S and RTP2 (Yu et al., Receptor-transporting protein (RTP) family members play divergent roles in the functional expression of odorant receptors. PLoS One 2017; 12(6): e0179067), which are typically expressed in olfactory sensory neurons and facilitate the transport of ORs to the cell surface membrane. In addition, fusion of an olfactory receptor gene with a sequence encoding the N-terminal sequence of rhodopsin (initial methionine and 19 subsequent amino acids) (= rho-tag) was found to promote the expression of olfactory receptors (Krautwurst et al., Identification of ligands for olfactory receptors by functional expression of a receptor library. Cell. 1998; 95(7): 917-26).
[0007] However, even when using RTP proteins and N-terminal rho-tags, more than half of the known olfactory receptors cannot be functionally expressed using currently available nucleic acid constructs, cell lines and methods, resulting in limited coverage of the currently available receptor-ligand space, multiple receptors without identified ligands (orphan receptors), and limited industrial application of the methods. Even for receptors expressed using the methods described, expression may be very low, resulting in low sensitivity of screening assays. In fact, many of those receptors that are functionally expressed in current expression systems are strongly activated only at relatively high ligand concentrations (e.g., at 10-300 μM), while many ligands have already triggered sensory experiences in vivo at much lower concentrations. Therefore, receptors as expressed in current systems are typically not activated at physiologically relevant concentrations. This suggests that current systems are generally not sensitive enough to mimic the situation in vivo. This also leads to practical problems, as (weak) ligands that are cytotoxic or poorly soluble in cell culture medium cannot trigger receptor activation in current screening cell lines because they are not sufficiently soluble (most odorants are polar molecules with limited solubility in water) or directly lead to cell line inactivation through cytotoxicity at the high test concentrations applied.
[0008] Classical OR screening assays rely on such methods, in which the clonal colony of cells typically receives a DNA expression construct encoding a specific receptor and / or auxiliary molecule at a time, and then tests its functional activation with various ligands. The assay further typically involves the coexpression of a luciferase gene operably connected to a cAMP-inducible promoter (Saito et al., RTP family members induce functional expression of mammalian odorant receptors. Cell 2004; 119 (5): 679-691), which is used as a reporter gene. The activation of olfactory receptors and the subsequent increase in intracellular cAMP lead to the expression of luciferase. The oxidation of luciferin catalyzed by luciferase leads to the emission of light, which can then be detected and quantified. Classical OR screening assays are limited in their sensitivity, may result in different heights of different receptors expressed, and are generally incompatible with high-throughput screening and selection methods, such as screening libraries of volatile spices and flavor compounds including moderately active ligands, ligands with cytotoxic properties, and ligands with limited solubility in cell culture media. Summary of the Invention
[0009] In view of all of the above, there is a need for improved olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses for expressing olfactory receptors and for identifying novel olfactory receptor ligands, enhancers, and antagonists. More specifically, such improved olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses should enhance the sensitivity of the assay to allow testing at lower concentrations, and thereby also describe the complete ligand spectrum of the OR with respect to poorly soluble and cytotoxic ligands. This improvement in sensitivity should also allow the identification of novel ligand-OR pairs, both deorphaning the receptor and better describing the complete receptive space of the already deorphaned receptor (additional ligands but also antagonists and enhancers).
[0010] In addition, for the large-scale de-orphanization of olfactory receptors (i.e., identifying ligands for all receptors that have no known ligands so far) and for finding all active ORs, especially the most sensitive ORs, for a given ligand of interest, it is necessary to express the entire library of many / all human olfactory receptors. In order to find the truly most important receptor for a given ligand, all receptors should be expressed at similar levels, preferably at least to the greatest extent possible. Otherwise, false positive reactions are observed, whereby the strongly expressed receptors appear as the most sensitive receptors to the ligand of interest, and the truly most sensitive receptors are missed due to lower functional expression. Therefore, for this screening activity of multiple receptors, it is desirable to standardize the functional expression of different receptors and minimize the expression differences between receptors. Therefore, it is necessary to optimize the human OR library in a way to give similar functional expression of all receptors to be used in OR expression determination.
[0011] The present invention provides olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses that are particularly useful in the case of ORs that are difficult to express using conventional methods or in the case of ORs that have limited sensitivity when measured using conventional methods, as well as in identifying novel cognate receptor-ligand pairs. The present invention further provides olfactory receptor variants with improved functional expression that allow more sensitive assays to detect OR-ligand interactions. Compared to the prior art, the olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses described herein exhibit, for example, at least one of the following benefits:
[0012] - Improved functional heterologous expression
[0013] - Allows deorphanization of receptors with no known ligands
[0014] - Enhanced assay sensitivity to determine the complete ligand spectrum of a receptor
[0015] - Enhance the sensitivity of the assay to measure ligand binding at more physiologically relevant concentrations
[0016] - Enhanced assay sensitivity to identify antagonists and enhancers
[0017] - Achieve functional expression of olfactory receptors, which is otherwise impossible using conventional methods
[0018] - Enables identification of novel cognate receptor-ligand pairs
[0019] - Enhanced sensitivity of olfactory receptor assays (as measured by lowering ligand concentration to achieve similar activity or lowering the EC50 value; i.e., enhanced potency, significantly lowered detection threshold, or increased efficacy)
[0020] - Ability to test more cytotoxic molecules
[0021] - Ability to test poorly soluble molecules
[0022] - Ability to detect ligands in complex test mixtures and unpurified synthetic samples
[0023] - Ability to screen for ligands with specific odor profiles in complex test mixtures and unpurified synthetic samples
[0024] - Ability to identify the more olfactory active isomer or enantiomer in racemic mixtures; and biodegradable compounds and compositions
[0025] - Allows the generation of OR libraries with better and more consistent functional expression compared to using wild-type OR genes
[0026] - Increased production capacity for olfactory receptor and ligand screening
[0027] As demonstrated in the experimental section herein, the use of the olfactory receptors and related nucleic acid molecules, expression vectors, recombinant host cells, libraries, and methods and uses described herein are associated with several of the above benefits and thus provide highly significant improvements over conventional methods. Thus, the various aspects and embodiments of the invention as described herein address at least some of the problems and needs discussed herein.
[0028] One aspect of the present invention relates to an olfactory receptor protein, wherein the protein has an amino acid sequence motif comprising RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR]
[0029] (SEQ ID NO: 1). In some embodiments, the olfactory receptor protein of the present invention is such that the modified C-terminal domain is fused to the seventh transmembrane helix (TM7) of the protein. In some embodiments, the olfactory receptor protein of the present invention is such that the protein is a class I or class II olfactory receptor with a modified C-terminal domain, preferably a human, dog or cat class I or class II olfactory receptor with a modified C-terminal domain, more preferably a human class I or class II olfactory receptor with a modified C-terminal domain. In some embodiments, the olfactory receptor protein of the present invention is such that the class II receptor is selected from the group consisting of: OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D 3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2, preferably wherein the class II receptor is selected from the group consisting of: OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2, and OR5A2. In some embodiments, the olfactory receptor protein of the present invention is such that the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3 and OR51L1. In some embodiments, the olfactory receptor protein of the present invention is such that the sequence motif is RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 5).
[0030] In some embodiments, the olfactory receptor protein of the present invention is such that x is not proline. In some embodiments, the olfactory receptor protein of the present invention is such that x is not proline or tryptophan. In some embodiments, the olfactory receptor protein of the present invention is such that x is selected from D, K, R, E, N, V, A, Q or G, preferably wherein x is selected from D, K, R, E, N, V, A or Q.
[0031] In some embodiments, the olfactory receptor protein of the invention is such that the amino acid sequence motif comprises 1 to 6 additional C-terminal amino acid residues, optionally wherein:
[0032] - the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, more preferably the first additional amino acid residue is C, R or K, most preferably C;
[0033] - the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably C or R;
[0034] - the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K;
[0035] - the fourth, fifth and sixth additional amino acid residues are selected from K and R.
[0036] In some embodiments, the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164), and CRRRKK (SEQ ID NO: 165).
[0037] In some embodiments, the olfactory receptor protein of the present invention is such that the sequence motif is selected from the group consisting of SEQ ID NO: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, 820-890.
[0038] In some embodiments, the olfactory receptor protein of the present invention is such that it further comprises an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rho tag, an SST3 tag and an M3 tag.
[0039] Another aspect of the present invention relates to nucleic acid molecules comprising a nucleotide sequence encoding an olfactory receptor protein of the present invention. In some embodiments, the nucleic acid molecules of the present invention are such that they further comprise a promoter sequence, preferably a constitutive promoter sequence. In some embodiments, the nucleic acid molecules of the present invention are such that they further comprise a terminator sequence. In some embodiments, the nucleic acid molecules of the present invention are such that they further comprise a nucleotide sequence encoding an N-terminal signal peptide, preferably a leucine-rich signal peptide, such as MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77).
[0040] Another aspect of the present invention relates to an expression vector comprising a nucleic acid molecule of the present invention. In some embodiments, the expression vector of the present invention is a plasmid.
[0041] Another aspect of the present invention relates to a recombinant host cell comprising a nucleic acid molecule of the present invention or an expression vector of the present invention, preferably wherein the cell expresses an olfactory receptor protein of the present invention. In some embodiments, the recombinant host cell of the present invention is such that the cell further expresses one or more olfactory receptor accessory proteins. In some embodiments, the one or more olfactory receptor accessory proteins are selected from the group consisting of RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf , Giα and functional variants thereof, preferably selected from the group consisting of RTP1S, RTP2 and functional variants thereof. In some embodiments, the recombinant host cell of the present invention is such that the cell is a HEK293 cell or a HEK293T cell.
[0042] Another aspect of the present invention relates to a diverse repertoire comprising an olfactory receptor protein of the present invention, a nucleic acid molecule of the present invention, an expression vector of the present invention, or a recombinant host cell of the present invention. In some embodiments, the library of the present invention is such that the diverse repertoire of olfactory receptor proteins, olfactory receptor proteins encoded by nucleic acid molecules or expression vectors, or olfactory receptor proteins expressed by recombinant host cells share the same modified C-terminal domain.
[0043] Another aspect of the present invention relates to the use of the olfactory receptor protein of the present invention, the nucleic acid molecule of the present invention, the expression vector of the present invention, the recombinant host cell of the present invention or the library of the present invention for identifying olfactory receptor ligands, enhancers or antagonists.
[0044] Another aspect of the present invention relates to the use of the library of the present invention for identifying an olfactory receptor capable of binding a target ligand.
[0045] Another aspect of the present invention relates to a method for identifying an olfactory receptor ligand, the method comprising:
[0046] a) providing the olfactory receptor protein of the present invention or a recombinant host cell expressing the olfactory receptor protein of the present invention;
[0047] b) contacting the receptor or recombinant host cell with a test compound or composition; and
[0048] c) Detecting activation of olfactory receptors.
[0049] Another aspect of the present invention relates to a method for identifying an olfactory receptor enhancer or antagonist, the method comprising:
[0050] a) providing the olfactory receptor protein of the present invention or a cell expressing the olfactory receptor protein of the present invention;
[0051] b) contacting the receptor or recombinant host cell with a cognate ligand and a test compound or composition; and
[0052] c) Detecting increased or decreased activation of the olfactory receptor compared to a ligand-only control.
[0053] In some embodiments of the methods for identifying an olfactory receptor ligand and the methods for identifying an olfactory receptor enhancer or antagonist, the olfactory receptor is selected from the group consisting of: OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)), OR 7A10,OR10H2,OR10H1,OR10D3,OR1D2,OR2A5,OR2A25(OR2A25(S75N,A209P)),OR11G2(or OR11G2(I65N,V82I)),OR14J1,OR5M3,OR8D1,OR10G3(preferably OR10G3(S73G)),OR10G9,OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2, and OR2AG2 (preferably OR2AG2(Y28C)).
[0054] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1, and OR4S2, preferably selected from the group consisting of OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and OR5V1, more preferably OR2M2 or OR2V1.
[0055] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is for identifying an olfactory receptor antagonist, the olfactory receptor is OR2M2 or OR2V1, and step b) further comprises contacting the receptor or the recombinant host cell with a copper salt. In some embodiments of the method for identifying an OR2M2 or OR2V1 antagonist, the cognate ligand is selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, and 4-methoxy-2-methylpentane-2-thiol.
[0056] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR51B2 (preferably OR51B2 (C120R, L134F, C209S)). In some embodiments of the method for identifying an OR51B2 antagonist, the cognate ligand is 3-methyl-2-hexenoic acid.
[0057] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR5V1. In some embodiments of the method for identifying an OR5V1 antagonist, the cognate ligand is 2,4,6-trichloroanisol.
[0058] Another aspect of the present invention relates to a method for identifying an olfactory receptor capable of binding a target ligand, the method comprising:
[0059] a) providing a library of the present invention;
[0060] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins from the library;
[0061] c) contacting a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins with the target ligand; and
[0062] d) identifying an olfactory receptor activated by said target ligand.
[0063] Another aspect of the invention relates to a method for generating an objective representation of the olfactory properties of a test compound or composition, the method comprising:
[0064] a) providing a library of the present invention;
[0065] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0066] c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with the test compound or composition; and
[0067] d) detecting activation of each of said olfactory receptor proteins.
[0068] Another aspect of the present invention relates to a method for evaluating differences or similarities between two or more test compounds or compositions, the method comprising:
[0069] a) providing a library of the present invention;
[0070] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0071] c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of the two or more test compounds or compositions;
[0072] d) detecting activation of each olfactory receptor protein by each of the two or more test compounds or compositions; and
[0073] e) comparing the activated olfactory receptor protein between each of the two or more test compounds or compositions.
[0074] describe
[0075] The various features of various aspects and embodiments of the present disclosure are further described below. It should be noted that the headings used throughout this specification are only used to assist navigation and should not be interpreted as definitive, and that the features described in different sections may be relevant to all aspects and embodiments described herein and therefore may be appropriately combined.
[0076] introduce
[0077] In addition to modifying the N-terminus of the olfactory receptor by adding, for example, a rho-tag, other changes have been made to the sequence of the olfactory receptor in an attempt to improve its expression.
[0078] In one attempt, consensus sequences were derived from each of the different classes of olfactory receptors, and it was found that for the resulting nine consensus receptors, highly expressed receptors were produced in six of the nine cases, which were generally better expressed than the individual receptors in the class (Ikegami et al., Structural instability and divergence from conserved residues underlie intracellular retention of mammalian odorant receptors. PNAS 2020; 117(6): 2957-2967). This consensus approach resolved the full-length sequence of the receptor, which means that the resulting consensus receptors do not correspond to the ligand affinity of a single natural receptor, but are fully synthetic receptors with a novel ligand spectrum, and are therefore not interesting for screening commercially interesting odorants perceived by the human nose.
[0079] Modification of the C-terminal sequence of the OR gene has also been reported. As shown below, in most cases, changing or truncating the C-terminal sequence resulted in reduced activity and only in a few cases resulted in higher signal amplitudes. To date, altering the C-terminal end of the OR has never resulted in a significantly more sensitive assay, i.e., an assay that allows testing at significantly lower test concentrations.
[0080] Kotthoff et al. found that truncating human OR8D1 by three amino acids reduced the signal amplitude, and truncating it by 7 amino acids, or 11 or 15 amino acids eliminated the signal, but they did not find any functional enhancement of the truncated C-terminus (Kotthoff et al., The FASEB Journal 2021; 35: e21274.). Changing the C-terminus had a negative effect in OR8D1. However, changing the parent C-terminal sequence toward the consensus sequence had a positive effect in two cases: Thus, changing the C-terminus of human OR2M3 toward the consensus sequence by changing three amino acids did increase the signal amplitude threefold, but it did not significantly improve sensitivity / efficacy, i.e., it did not shift the dose-response curve to lower concentrations. Changing mouse olfr16 toward the consensus sequence doubled the signal amplitude, but it only slightly increased sensitivity (EC50 from 33.99 micromolar to 20.42 micromolar, Table S19 in the reference). All other changes introduced into the C-terminus (>70 variants studied) invariably resulted in reduced activity and / or reduced surface expression, overall indicating that receptor expression and assay sensitivity could not be substantially improved by altering the C-terminus.
[0081] In a very detailed analysis in the aforementioned reference (Ikegami et al., 2020), two closely related mouse receptors (mouse Olfr539 and Olfr541) were compared, with Olfr541 being poorly expressed and Olfr539 being well expressed. Strong expression could be achieved by exchanging a portion of the sequence, or even a single amino acid, of the central domain of Olfr541 with a sequence from Olfr539, whereas replacing the C-terminus of the poorly expressed Olf541 with the C-terminus of the well-expressed Olfr539 did not enhance Olfr541 expression, suggesting that other parts of the sequence, rather than the C-terminus, are critical for conferring functional expression.
[0082] In another example, the human OR5A2 sequence was modified by exchanging the OR5A2 sequence with the OR2A5 sequence at the C-terminus and N-terminus to generate a chimeric receptor (Example 5 in WO2019110630A1). The resulting receptor responded to musk compounds as well as the wild-type receptor, and the dose-response curve was not changed, and the sensitivity was not improved by the change in the C-terminus.
[0083] Furthermore, in the often cited work (Krautwurst et al., Identification of ligands forolfactory receptors by functional expression of a receptor library. Cell. 1998; 95(7): 917-26), Krautwurst et al. used the N-terminal and C-terminal sequences of the well-expressed mouse receptor M4 and entered the sequences of TM2-TM7 from other test receptors into this backbone to generate chimeric receptors. The chimeric receptors with this M4 backbone responded the same as the wild-type receptor, but the chimeric receptors did not show enhanced signal or better functionality, and in one instance, the wild-type mouse I7 receptor showed a response at even lower concentrations (1 μM instead of 10 μM) than the chimeric receptor, indicating that exchanging the C-terminus with the C-terminus of a well-expressed receptor can in most cases maintain rather than increase activity.
[0084] Katada et al. found that truncating mouse mOR-EG by 3 amino acids reduced signal amplitude, truncating by 6 amino acids reduced signal amplitude and sensitivity, and truncating by 12 amino acids abolished the signal, but they did not find enhanced function of the truncated C-terminus, indicating that the full-length C-terminus is required under the conditions studied (Kadata et al., Structural determinants for membrane trafficking and G protein selectivity of a mouse olfactory receptor. Journal of Neurochemistry 2004;90(6):1453–1463).
[0085] Kato et al. found that mutating K296, K299, K303, K304, and K309 in the C-terminus of mouse mOR-EG to proline reduced / eliminated activity, and mutation of these residues to Ala or Arg did not significantly affect activity (except that K296A reduced sensitivity), but the study did not find enhanced function (enhanced amplitude or enhanced sensitivity) of the mutated C-terminal sequences and suggested that most of these basic residues appear to be non-essential for functional expression (Kato et al., Amino acids involved in conformational dynamics and G protein coupling of an odorant receptor: targeting gain-of-function mutation. Journal of Neurochemistry 2008; 107(5): 1261-1270).
[0086] Finally, Sato et al. studied the C-terminal olfactory receptor sequence and which residues are critical, and they found that residues 299, 300, 303, and 304 in the C-terminus of mOR-S6 are important for maintaining function, but they did not find enhanced function of C-terminal modifications (Sato et al., Functional Role of the C-Terminal Amphipathic Helix 8 of Olfactory Receptors and Other G Protein-Coupled Receptors. Int. J. Mol. Sci. 2016; 17(11): 1930).
[0087] In summary, while it is well known that altering the N-terminus by adding a rho-tag improves olfactory receptor expression, attempts to alter the C-terminus have not resulted in significantly improved assays, primarily showing that the C-terminus is sensitive to structural changes that, in most cases, result in loss-of-function rather than gain-of-function modifications. Based on these teachings, it would not be expected that significant advances in functional expression of olfactory receptors would result from altering the C-terminal sequence. Instead, this detailed analysis did not demonstrate that truncating the C-terminus of the receptor or shifting the C-terminus toward a consensus sequence could significantly enhance assay sensitivity, but rather showed that signal amplitude could be enhanced in a few selected cases, while most alterations to the C-terminus had a negative impact.
[0088] Olfactory receptor protein
[0089] The present disclosure generally relates to olfactory receptor proteins (ORs) having a modified C-terminal domain. Thus, an olfactory receptor protein is provided herein, wherein the protein has a modified C-terminal domain. Preferably, the olfactory receptor protein is a mammalian, more preferably a human olfactory receptor protein. Preferably, the olfactory receptor protein corresponds to a class I or class II OR, as described later herein.
[0090] As used herein, the term "olfactory receptor" or "odorant receptor" (OR) has its customary meaning as generally understood by those skilled in the art in view of the present disclosure. It refers to a receptor belonging to the seven-transmembrane domain G protein-coupled receptor superfamily (GPCR), which is typically expressed in the cell membrane of olfactory receptor neurons. The predicted seven transmembrane (TM) domains TM I to TM VII are connected by three predicted internal (IC) loop domains (IC I to IC III) and three predicted external (EC) loop domains (EC I to EC III). ORs typically contain olfactory receptor-specific amino acid motifs. Examples of such motifs are the MAYDRYVAIC (SEQ ID NO: 2) motif overlapping with TMIII and ICII, the FSTCSSH (SEQ ID NO: 3) motif overlapping with IC III and TM VI, the PMLNPFIY (SEQ ID NO: 4) motif in TM VII and the three conserved C residues in ECII, and the presence of the highly conserved GN residue in TMI, described in Zhang and Firestein (2002) Nature Neurosci 5(2): 124-33 and Malnic et al., (2004) PNAS 101(8): 2584-9, both incorporated herein by reference.
[0091] The C-terminal domain of the olfactory receptor begins just after the end of the seventh TM helix (TM7). A skilled person can determine this position without any doubt based on their common general knowledge. For example, the seventh transmembrane region (TM7) can be easily identified within any OR based on sequence alignment or by using well-known and publicly available databases. For example, the "GPCR Prediction Ensemble Database (GPCR-PEnDB)" available at https: / / gpcr.utep.edu / database indicates the amino acid position of TM7 and the length of the native-C terminus of many ORs.
[0092] The same information can be derived from the "HORDE" (The Human Olfactory Data Explorer) database maintained by the Weizmann Institute of Science and available at https: / / genome.weizmann.ac.il / horde / , which is described in Olender et al., (2013) Methods Mol Biol 1003:23-38, which is incorporated herein by reference. In HORDE, all residues TM1-TM7 are annotated.
[0093] The same information can also be found in universal sequence databases, such as Uniprot (The UniProt Consortium, UniProt: the universal protein knowledgebase in 2021, Nucleic Acids Research, Volume 49, Issue D1, January 8, 2021, pages D480–D489, available at www.uniprot.org).
[0094] TM7 in class II olfactory receptors typically ends with the consensus sequence NPLIYSL (SEQ ID NO: 225). The last of these seven amino acids is typically located at residues 292-298 of the full-length receptor, making it easy to identify the TM7 end of any given OR. Directly following this sequence (or at the beginning of the native C-terminus, as shown at https: / / gpcr.utep.edu / database), the modified C-terminus described above is fused to the olfactory receptor to improve functional expression.
[0095] TM7 in class I olfactory receptors typically ends with the consensus sequence NPIIYSL (SEQ ID NO: 226) or NPIIYSGL (SEQ ID NO: 227), with the last of these seven amino acids typically located at residue 297 (290-300). Directly following this sequence (or at the beginning of the native C-terminus, as shown at https: / / gpcr.utep.edu / database), the modified C-terminus is fused to the olfactory receptor to improve functional expression.
[0096] Thus, in some embodiments, the C-terminal domain described herein is such that it begins after the last residue of the seventh transmembrane helix (TM7). In some embodiments, the C-terminal domain is fused to the seventh transmembrane helix of the olfactory receptor protein. In some embodiments, the C-terminal domain described herein may also be referred to as a cytoplasmic domain or an intracellular domain.
[0097] Thus, in some embodiments, the olfactory receptor described herein comprises the amino acid sequence of the olfactory receptor up to and including the last residue of the seventh transmembrane helix (TM7) of the olfactory receptor, followed by the amino acid sequence of a modified C-terminal domain comprising the amino acid sequence motif described herein.
[0098] The modified C-terminal domain may comprise or consist of 10 to 40 amino acids, preferably 12 to 30 amino acids, more preferably 14 to 26 amino acids, and even more preferably 16 to 22 amino acids. In some embodiments, the minimum length of the modified C-terminal domain of the olfactory receptor of the present disclosure is 10, 11, 12, 13, 14, 15, or 16 amino acids; the maximum length of the modified C-terminal domain of the olfactory receptor of the present disclosure is 26, 25, 24, 23, or 22 amino acids.
[0099] The present inventors have found that the particularly advantageous effects described herein are observed for modified C-terminal domains rich in basic amino acids. As used herein, basic amino acids are understood to refer to amino acids whose side chains are protonated at neutral pH. Basic amino acids include lysine (Lys, K), arginine (Arg, R) and histidine (His, H). Preferred basic amino acids in the context of the present disclosure are lysine and arginine.
[0100] Therefore, in some embodiments, the modified C-terminal domain described herein is a modified C-terminal domain comprising at least 5 basic residues, preferably at least 6 basic residues, more preferably at least 7 basic residues. In some embodiments, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42% or at least 43% of the amino acids are basic amino acids (i.e., lysine, arginine or histidine). In preferred embodiments, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 32% of the amino acids are basic amino acids (i.e., lysine, arginine or histidine). In preferred embodiments, the modified C-terminal domain described herein is a modified C-terminal domain in which at least 35% of the amino acids are basic amino acids (i.e., lysine, arginine or histidine). In a preferred embodiment, the modified C-terminal domain described herein is a modified C-terminal domain wherein at least 38% of the amino acids are basic amino acids (ie, lysine, arginine, or histidine).
[0101] In one aspect, an olfactory receptor protein is provided, wherein the protein has a modified C-terminal domain comprising the following amino acid sequence motif:
[0102] RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ IDNO: 1).
[0103] "Sequence motif" may also be expressed as "sequence pattern" or simply "sequence". "Sequence motif" or "sequence pattern" has its customary and ordinary meaning as understood by those skilled in the art in view of the present disclosure. It refers to an amino acid (or nucleotide) sequence that is repeated at multiple sites in a molecule or multiple different molecules with a certain degree of variation, and the sequence has (or is speculated to have or is assumed to be associated with) the biological significance described herein or exhibits the biological activity described herein. The biological significance or biological activity of the sequence motif described herein is preferably its ability to improve the functional heterologous expression of an olfactory receptor protein (nucleotide sequence encoding the protein) that contains the motif as a C-terminal domain.
[0104] As used herein, the term "expression" or "heterologous expression" of a DNA molecule by a cell includes any step involved in the production of a polypeptide by the cell, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, transport to the cell membrane, and secretion. Expression can be assessed by any method known to those skilled in the art. For example, expression can be assessed by measuring the level of gene expression in transduced cells at the mRNA or protein level by standard assays known to those skilled in the art, such as qPCR, RNA sequencing, Northern blot analysis, Western blot analysis, mass spectrometry analysis of protein-derived peptides, fluorescence activated cell sorting (FACS), immunostaining, or ELISA.
[0105] As used herein, the term "functional expression" or "functional heterologous expression" refers to a cell producing a polypeptide, wherein the polypeptide exhibits biological activity. For example, when an olfactory receptor is produced, transported and integrated into the cell membrane, and is able to trigger its corresponding signal cascade after being activated by a ligand, the receptor is functionally expressed by the cell. Conventional methods for evaluating the functional expression of olfactory receptors include expressing the receptor and co-expressing a luciferase gene operably linked to a cAMP-inducible promoter (Saito et al., (2004) Cell 119 (5): 679-691), which is used as a reporter gene. If the olfactory receptor is functionally expressed, its activation and the subsequent increase in intracellular cAMP will lead to increased luciferase expression. In standard assays, luciferase oxidizes luciferin to cause luminescence, and the light can then be detected and quantified.
[0106] The perceived potency of an odor, and the potency of the odorant-OR interaction in vivo, is most commonly described as the odorant's odor detection threshold (OTH), i.e., the lowest concentration detected by the human nose in the gas phase. The terms "odor detection threshold" and "odor threshold" (also abbreviated herein as "OTH") are synonymous and are commonly used terms in the field of fragrances, see, for example, Neuner-Jehle, N., Etzweiler, F. "The Measuring of Odors" (1994). See: Mailer, PM, Lamparsky, D. (eds) Perfumes. Springer, Dordrecht, the entire contents of which are incorporated herein by reference. The odor detection threshold can be measured by methods and means commonly used in the art, such as using an olfactometer in conjunction with human subjects. Another method for measuring the odor detection threshold is to inject a series of specified amounts (measured in ng) of ligand dilutions into a gas chromatograph (GC), followed by a human panelist sniffing the molecules released from the GC column at the sniffing port and indicating whether they can be detected by the nose. This yields the GC threshold value (GCO) in ng.
[0107] In OR screening assays using ORs in in vitro systems, the ligand is dissolved in a liquid medium. However, similar to in vivo olfactory threshold experiments, it is possible to determine the lowest dose of the odorant at which receptor activation begins in the in vitro assay. Thus, the sensitivity of the in vitro system is reported as the lowest concentration dissolved in the medium that begins to activate the receptor, while in vivo sensitivity is expressed as the lowest detectable concentration in the gas phase. As a practical approach, the in vitro detection threshold can be defined as the concentration that results in a response twice the background signal in, for example, the luciferase assay described herein.
[0108] The functional expression of the olfactory receptor polypeptides described herein is improved (increased) relative to baseline functional expression, thereby resulting in improved (increased) biological activity, for example, relative to an unmodified corresponding receptor gene. The functional expression can be improved (increased) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 100%, at least 200%, at least 300%, or at least 500% relative to an unmodified corresponding receptor gene. Such improved functional expression can also be manifested in an increased sensitivity of recombinant host cells expressing the olfactory receptor polypeptides described herein, such that the detection threshold of the receptor in an in vitro assay is reduced by at least 2-fold, at least 3.1-fold, at least 10-fold, at least 20-fold, or at least 50-fold, preferably at least 3.1-fold (i.e., the sensitivity of detection of the ligand being tested is increased). This improvement in functional expression can also be reflected in an increase in the sensitivity of recombinant host cells expressing the olfactory receptor polypeptides described herein (described elsewhere herein), such that the EC50 (i.e., the concentration required to achieve 50% of the maximum signal amplitude) is reduced by at least 2-fold, at least 3.1-fold, at least 10-fold, at least 20-fold, or at least 50-fold, preferably at least 3.1-fold (i.e., the potency of the tested ligand is enhanced). This improvement in functional expression can also be reflected in an increase in the sensitivity of recombinant host cells expressing the olfactory receptor polypeptides described herein, such that the maximum signal amplitude is increased by at least 30%, at least 50%, at least 80%, at least 100%, at least 200%, or at least 500% (i.e., the potency of the tested ligand is enhanced). The improvement in functional expression can also be achieved to the extent that the olfactory receptors, nucleic acid molecules, expression vectors, recombinant host cells, and methods and uses of the present disclosure achieve functional expression of the olfactory receptor that is not achievable using conventional methods (an all-or-nothing effect that cannot be quantified numerically).
[0109] Several customary notations for describing sequence motifs are in use and are well known to the skilled artisan, most of which are variations of the standard notation for regular expressions and use at least the following conventions:
[0110] - has a single-letter alphabet, the standard IUPAC single-letter code for amino acids, where each letter represents a specific amino acid or group of amino acids;
[0111] - a string of characters drawn from the alphabet representing the corresponding amino acid sequence;
[0112] - the string of characters between the square brackets matches any of the listed amino acids, i.e., it represents a sequence ambiguity or alternative, e.g., the notation [XYZ] is assumed to mean X or Y or Z;
[0113] - A string of characters between braces / curly brackets ("{}") represents any amino acid other than the listed amino acid, for example, assume that the symbol {X} represents any amino acid other than X.
[0114] Thus, the 16 amino acids in the sequence motif described above:
[0115] RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR](SEQ IDNO:1) is defined as follows:
[0116] - the first residue is R;
[0117] - the second residue is N;
[0118] - The third residue is K or R;
[0119] - the fourth residue is E or D or Q;
[0120] - the fifth residue is V or M or I or L;
[0121] - the sixth residue is K or R;
[0122] - The 7th residue can be any amino acid;
[0123] - the 8th residue is A;
[0124] - the 9th residue is L, I, or V;
[0125] - the 10th residue is K, R, or H;
[0126] - the 11th residue is K or R;
[0127] - the 12th residue is L or I;
[0128] - the 13th residue is L, I, or F;
[0129] - the 14th residue is K, R, or G;
[0130] - the 15th residue is K or R; and
[0131] - The 16th residue is K or R.
[0132] Thus, the sequence motif described above:
[0133] RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1) may also be alternatively described as:
[0134] RNX1X2X3X4X5AX6X7X8X9X 10 X 11 X 12 X 13 ,in
[0135] -X1 is K or R;
[0136] -X2 is E or D or Q;
[0137] -X3 is V or M or I or L;
[0138] -X4 is K or R;
[0139] -X5 is any amino acid;
[0140] -X6 is L, I or V;
[0141] -X7 is K or R or H;
[0142] -X8 is K or R;
[0143] -X9 is L or I;
[0144] -X 10 It is L or I or F;
[0145] -X 11 It is K or R or G;
[0146] -X 12 is K or R; and
[0147] -X 13 It is K or R.
[0148] Taking into account that the standard IUPAC single-letter code for amino acids includes the symbol "J" for leucine (L) or isoleucine (I), the above sequence motif can alternatively be described as:
[0149] RNX1X2X3X4X5AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 1), wherein:
[0150] -X1 is K or R;
[0151] -X2 is E or D or Q;
[0152] -X3 is V or M or I or L;
[0153] -X4 is K or R;
[0154] -X5 is any amino acid;
[0155] -X6 is L, I or V;
[0156] -X7 is K or R or H;
[0157] -X8 is K or R;
[0158] -X 10 It is L or I or F;
[0159] -X 11 It is K or R or G;
[0160] -X 12 is K or R; and
[0161] -X 13 It is K or R.
[0162] All of the above symbols are fully equivalent and represent the same sequence (SEQ ID NO: 1); therefore, the skilled person understands that they can be used interchangeably. The same applies to any other sequence motifs described herein.
[0163] Olfactory receptors can be divided into Class I or "fish-like" receptors, and Class II or "quadrangular pyramidal" receptors. Class I or "fish-like" receptors are an evolutionarily older class of receptors known to respond particularly to more soluble compounds (such as carboxylic acids), which are of particular interest for screening antagonists of off-flavor odorants. Most olfactory receptors belong to Class II or "quadrangular pyramidal" receptors, which contain the key receptors for most aromatic molecules. There are considerable differences in the C-terminal sequences between Class I and Class II receptors, however, the present disclosure targets both Class I and Class II olfactory receptors.
[0164] Thus, in some embodiments, the olfactory receptor protein as described herein is a Class I or Class II olfactory receptor with a modified C-terminal domain.
[0165] The olfactory receptor proteins of the present disclosure have a modified C-terminal domain. Therefore, the C-terminal domain of the olfactory receptor proteins of the present disclosure is different from the natural or homologous C-terminal domain typically associated with the olfactory receptor. Therefore, the olfactory receptor proteins of the present disclosure are non-naturally occurring proteins. It can be seen that the olfactory receptor proteins described herein can be characterized as "modified" olfactory receptor proteins, "engineered" olfactory receptor proteins, "hybrid" olfactory receptor proteins, "chimeric" olfactory receptor proteins, "non-natural" olfactory receptor proteins, or similar expressions and combinations thereof. Throughout this disclosure, the skilled person understands that the term "olfactory receptor protein" can be replaced with the term "olfactory receptor protein with a modified C-terminal domain".
[0166] The present disclosure includes olfactory receptor proteins from any mammal. Thus, in some embodiments, the olfactory receptor protein as described herein is a mammalian olfactory receptor with a modified C-terminal domain, preferably a mammalian Class I or Class II olfactory receptor with a modified C-terminal domain. In the context of the present disclosure, preferred mammals are pets or companion animals and humans, more preferably humans. Among pets or companion animals, cats and dogs are particularly preferred. Thus, in some embodiments, the olfactory receptor protein as described herein is a human, dog or cat olfactory receptor with a modified C-terminal domain, preferably a human olfactory receptor with a modified C-terminal domain. In some embodiments, the olfactory receptor protein as described herein is a human, dog or cat Class I or Class II olfactory receptor with a modified C-terminal domain, preferably a human Class I or Class II olfactory receptor with a modified C-terminal domain.
[0167] Mammalian and human olfactory receptors are discussed in publications such as Mainland et al., (2015) Sci Data 2: 150002 (which are incorporated herein by reference) and in publicly available databases such as the HORDE (Human Olfactory Data Explorer) database maintained by the Weizmann Institute of Science, available at https: / / genome.weizmann.ac.il / horde / , described in Olender et al., (2013) Methods Mol Biol 1003: 23-38, which are incorporated herein by reference. Other relevant publicly available databases exist, for example, as described in Marenco et al., Database (Oxford) 2016: baw132 and Han et al., Science China. Life sciences, 10.1007 / s11427-021-2081-6.
[0168] A complete list of human olfactory receptors and their alleles or haplotypes can be found in the HORDE database mentioned above. Major alleles or haplotypes are defined as those with a frequency ≥ 20% in the human population according to the HORDE database. A list of all amino acid sequences of the major alleles or haplotypes of the vast majority of human olfactory receptors is available to the skilled person and is included as supporting information for Ikegami et al. 2020 (see above), while the nucleotide sequences of 625 genes covering the vast majority of human OR genes and their major alleles or haplotypes are included as supporting information for Mainland et al., 2015 (see above). All known human olfactory receptors can also be obtained from universal sequence databases, such as the NCBI Genbank available at https: / / www.ncbi.nlm.nih.gov / genbank / .
[0169] In some embodiments, the olfactory receptor protein as described herein is selected from the group consisting of: OR10A2, OR10A3, OR10A4, OR10A5, OR10A6, OR10A7, OR10AD1, OR10AG1, OR10C1, OR10D3, OR10G2, OR10G3, OR10G4, OR10G6, OR10G7, OR10G8, OR10G9, OR10H1, OR10H2, OR10H3, OR10H4, OR10H5, OR10J1, OR10J3, OR10J5, OR10K1, OR10K2, OR10P1, OR10P2, OR10Q1, OR10Q2. R2,OR10S1,OR10T2,OR10V1,OR10W1,OR10X1,OR10Z1,OR11A1,OR11G2,OR11H1,OR11H2,OR11H4,OR11H6,OR11L1,OR12D2,OR12D3,OR13C2,OR13C3,OR 13C4,OR13C5,OR13C8,OR13C9,OR13D1,OR13F1,OR13G1,OR13H1,OR13J1,OR14A16,OR14A2,OR14C36,OR14I1,OR14L1P,OR1A1,OR1A2,OR1B1,OR1C1,OR 1D2,OR1D4,OR1D5,OR1E1,OR1E2,OR1F1,OR1F12,OR1G1,OR1I1,OR1J1,OR1J2,OR1J4,OR1K1,OR1L1,OR1L3,OR1L4,OR1L6,OR1L8,OR1M1,OR1N1,OR1N2 ,OR1Q1,OR1S1,OR1S2,OR2A1,OR2A12,OR2A14,OR2A2,OR2A25,OR2A4,OR2A42,OR2A5,OR2A7,OR2A9P,ORAE1,OR2AG1,OR2AG2,OR2AJ1,OR2AK2,OR2AP1, OR2AT4,OR2B11,OR2B2,OR2B3,OR2B6,OR2B8P,OR2C1,OR2C3,OR2D2,OR2D3,OR2F1,OR2F2,OR2G2,OR2G3,OR2G6,OR2H1,OR2H2,OR2J1P,OR2J2,OR2J3,O R2K2,OR2L13,OR2L2,OR2L3,OR2L5,OR2L8,OR2M2,OR2M3,OR2M4,OR2M5,OR2M7,OR2S2,OR2T1,OR2T10,OR2T11,OR2T12,OR2T2,OR2T27,OR2T29,OR2T3,OR2T33,OR2T34,OR2T35,OR2T4OR2T5,OR2T6,OR2T7,OR2T8,OR2V1,OR2V2,OR2W1,OR2W3,OR2W5,OR2Y1,OR2Z1,OR3A1,OR3A2,OR3A3,OR3A4,OR4A15,OR4A16,OR4A4,R4A47,OR4A5,OR4B1,OR4C11,OR4C12,OR4C13,OR4C15,OR4C16,OR4C,OR4C45,OR4C46,OR4C5,OR4C6,OR4D1,OR4D10,OR4D11,OR4D2,OR4D5,OR4D6,OR4D9,OR4E2,OR4F15,OR4F16,OR4F17,OR4F21,OR4F29,OR4F3,OR4F4,OR4F5,OR4F6,OR4K1,OR4K13,OR4K14,OR4K15,OR4K17,OR4K2,OR4K3P,OR4K5,OR4L1,OR4M1,OR4M2,OR4N2,OR4N4,OR4N5,OR4P4,OR4Q3,OR4S1,OR4S2,OR4X1,OR4X2,OR51A2,OR51A4,OR51A7,OR51B2,OR51B4,OR51B5,OR51B6,OR51D1,OR51E1,OR51E2,OR51F1,OR51F2,OR51G1,OR51G2,OR51H1P,OR51I1,OR51I2,OR51J1,OR51L1,OR51M1,OR51Q1,OR51S1,OR51T1,OR51V1,OR52A1,OR52A4,OR52A5,OR52B2,OR52B4,OR52B6,OR52D1,OR52E2,OR52E4,OR52E5,OR52E6,OR52E8,OR52H1,OR52I1,OR52I2,OR52J3,OR52K1,OR52K2,OR52L1,OR52M1,OR52N1,OR52N2,OR52N4,OR52N5,OR52P1P,OR52R1,OR52W1,OR56A1,OR56A3,OR56A4,OR56A5,OR56B1,OR56B4,OR5A1,OR5A2,OR5AC2,OR5AK2,OR5AL1P,OR5AN1,OR5AP2,OR5AR1,OR5AS1,OR5AU1,OR5B12,OR5B17,OR5B2,OR5B21,OR5B3,OR5C1,OR5D13,OR5D14,OR5D16,OR5D18,OR5F1,OR5H1,OR5H14,OR5H15,OR5H2,OR5H6,OR5I1,OR5J2,OR5K1,OR5K2,OR5K3,OR5K4,OR5L1,OR 5L2,OR5M1,OR5M10,OR5M11,OR5M3,OR5M8,OR5M9,OR5P2,OR5P3,OR5R1,OR5T1,OR5T2,OR5T3 ,OR5V1,OR5W2,OR6A2,OR6B1,OR6B2,OR6B3,OR6C1,OR6C2,OR6C3,OR6C4,OR6C6,OR6C65,OR6 C68,OR6C70,OR6C74,OR6C75,OR6C76,OR6F1,OR6J1,OR6K2,OR6K3,OR6K6,OR6M1,OR6N1,OR6 N2,OR6P1,OR6Q1,OR6S1,OR6T1,OR6X1,OR6Y1,OR7A10,OR7A17,OR7A5,OR7C1,OR7C2,OR7D2 ,OR7D4,OR7E24,OR7G1,OR7G2,OR7G3,OR8A1,OR8B12,OR8B2,OR8B3,OR8B4,OR8B8,OR8D1,OR 8D2, OR8D4, OR8G1, OR8G5, OR8H1, OR8H2, OR8H3, OR8I2, OR8J1, OR8J3, OR8K1, OR8K3, OR8K5, OR8S1, OR8U1, OR8U8, OR8U9, OR9A2, OR9A4, OR9G1, OR9G4, OR9G9, OR9I1, OR9K2, OR9Q1, and OR9Q2. Variants and haplotypes of these receptors are also contemplated.
[0170] In some embodiments, the olfactory receptor protein as described herein is:
[0171] - a class II olfactory receptor listed as a "receptor" in Table 15, or
[0172] - A class I olfactory receptor listed as "Receptor" in Table 22.
[0173] The specific sequences mentioned in Table 15 (for DNA encoding a modified receptor, including an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 221)) and Table 22 (for DNA encoding a modified receptor, including an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 741)) are specific embodiments that are not limiting in this context. The presence of an N-terminal tag is optional, and different N-terminal tags can be used, as described elsewhere herein. Similarly, any modified C-terminal domain disclosed herein can be used in place of the modified C-terminus of SEQ ID NO: 221 or SEQ ID NO: 741.
[0174] Modified C- terminal domain as described herein can include sequence RNKEVKDALKRLLKRK (SEQ ID NO: 10). In some embodiments, sequence RNKEVKDALKRLLKRK (SEQ ID NO: 10) can include amino acid replacement at 1, 2, 3, 4, 5, 6, 7, 8 or up to 9 positions. In some embodiments, the sequence can include amino acid replacement at 1, 2, 3, 4 or up to 5 positions, but the residue at the first (R), second (R) and eighth (A) position is fixed. In some embodiments, the sequence can include amino acid replacement at 1, 2, 3, 4, 5, 6, 7, 8 or up to 9 positions, but the residue at the first (R), second (N), fourth (E) and the first (A) position is fixed. Amino acid replacement can preferably be conservative amino acid replacement, as described in other parts of this paper. In this context, the example of particularly suitable amino acid replacement includes K replacing R, and R replacing K.
[0175] In some embodiments, the olfactory receptor protein as described herein is a class II olfactory receptor with a modified C-terminal domain. In some embodiments, the class II olfactory receptor is selected from the group consisting of OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2 and OR4S2, preferably selected from the group consisting of OR7A17, OR2M2 and OR5A2. In some embodiments, the class II olfactory receptor is selected from the group consisting of OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5A1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1, and OR2J2. In some embodiments, a class II olfactory receptor is selected from the group consisting of OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR1D 2, OR2A5, OR2A25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1 and OR2J2, preferably selected from the group consisting of OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2 and OR5A2.
[0176] OR7C1,OR9Q2,OR8K3,OR10J5,OR1C1,OR7D4,OR2T4,OR5B12,OR7A17,OR10H5,OR5A1,OR5A 2,OR1N2,OR2C1,OR2T11,OR2M2,OR4S2,OR2L3,OR2AG2,OR7A5,OR7E24,OR7A10,OR10H2,O R10H1, OR10D3, OR1D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3(S73G), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1 and OR2J2 are human class II olfactory receptors.
[0177] OR2A25 may also be OR2A25(S75N, A209P). OR1N2 may also be OR1N2(W23R, V230G, T287M). OR7E24 may also be OR7E24(P242S). OR2AG2 may also be OR2AG2(Y28C). OR8K3 may also be OR8K3(L122R). OR2L2 may also be OR2L2(V259L). OR11G2 may also be OR11G2(I65N, V82I). OR10G7 may also be OR10G7(T5S). OR2AK2 may also be OR2AK2(S84N). OR10A6 may also be OR10A6(A117V, V140G, L287P). OR10J1 may also be OR10J1(M51I, I92M). OR10G3 can also be OR10G3(S73G).
[0178] In some embodiments, (wild type) OR7C1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26664 (NCBI Reference Sequence NP_001357414.2). In some embodiments, (wild type) OR9Q2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219957 (NCBI Reference Sequence: NP_001005283.1). In some embodiments, (wild type) OR8K3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219473 (NCBI Reference Sequence: NP_001005202.1). In some embodiments, (wild type) OR10J5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 127385 (NCBI Reference Sequence: NP_001004469.1). In some embodiments, (wild type) OR1C1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26188 (NCBI Reference Sequence: NP_036485.2). In some embodiments, (wild type) OR7D4 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 125958 (NCBI Reference Sequence: NP_001005191.1). In some embodiments, (wild type) OR2T4 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 127074 (NCBI Reference Sequence: NP_001004696.2). In some embodiments, (wild type) OR5B12 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390191 (NCBI Reference Sequence: NP_001004733.1). In some embodiments, (wild type) OR7A17 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26333 (NCBI Reference Sequence: NP_112163.1). In some embodiments, (wild type) OR10H5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 284433 (NCBI Reference Sequence: NP_001004466.1). In some embodiments, (wild type) OR5A2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219981 (NCBI Reference Sequence: NP_001001954.1). In some embodiments, (wild type) OR5A1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219982 (NCBI Reference Sequence: NP_001004728.1). In some embodiments, the (wild-type) OR1N2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 138882 (NCBI Reference Sequence: NP_001004457.2).In some embodiments, (wild type) OR2C1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 4993 (NCBI Reference Sequence: NP_036500.2). In some embodiments, (wild type) OR2T11 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 127077 (NCBI Reference Sequence: NP_001001964.1). In some embodiments, (wild type) OR2M2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 391194 (NCBI Reference Sequence: NP_001004688.1). In some embodiments, (wild type) OR4S2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219431 (NCBI Reference Sequence: NP_001004059.2). In some embodiments, (wild type) OR2L3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 391192 (NCBI Reference Sequence: NP_001004687.1). In some embodiments, (wild type) OR2AG2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 338755 (NCBI Reference Sequence: NP_001004490.1). In some embodiments, (wild type) OR7A5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26659 (NCBI Reference Sequence: NP_001357409.1). In some embodiments, (wild type) OR7E24 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26648 (NCBI Reference Sequence: NP_001073404.1). In some embodiments, (wild type) OR7A10 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390892 (NCBI Reference Sequence: NP_001005190.1). In some embodiments, (wild type) OR10H2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26538 (NCBI Reference Sequence: NP_039227.1). In some embodiments, (wild type) OR10H1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26539 (NCBI Reference Sequence: NP_039228.1). In some embodiments, (wild type) OR10D3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26497 (NCBI Reference Sequence: NP_001342142.1). In some embodiments, the (wild-type) OR1D2 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 4991 (NCBI Reference Sequence: NP_001373017.1).In some embodiments, (wild type) OR2A5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 393046 (NCBI Reference Sequence: NP_036497.1). In some embodiments, (wild type) OR2A25 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 392138 (NCBI Reference Sequence: NP_001004488.1). In some embodiments, (wild type) OR11G2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390439 (NCBI Reference Sequence: NP_001005503.2). In some embodiments, (wild type) OR14J1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 442191 (NCBI Reference Sequence: NP_112208.1). In some embodiments, (wild type) OR5M3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219482 (NCBI Reference Sequence: NP_001004742.2). In some embodiments, (wild type) OR8D1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 283159 (NCBI Reference Sequence: NP_001002917.1). In some embodiments, (wild type) OR10G3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26533 (NCBI Reference Sequence: NP_001005465.1). In some embodiments, (wild type) OR10G9 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219870 (NCBI Reference Sequence: NP_001001953.1). In some embodiments, (wild type) OR2L5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 81466 (NCBI Reference Sequence: NP_001245213.1). In some embodiments, (wild type) OR8H1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 219469 (NCBI Reference Sequence: NP_001005199.1). In some embodiments, (wild type) OR10K1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 391109 (NCBI Reference Sequence: NP_001004473.1). In some embodiments, (wild type) OR11A1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26531 (NCBI Reference Sequence: NP_001381757.1). In some embodiments, the (wild-type) OR2V1 has an amino acid sequence encoded by the nucleotide sequence of NCBI Gene ID 26693 (NCBI Reference Sequence: NP_001245212.1).In some embodiments, (wild type) OR5P3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 120066 (NCBI Reference Sequence: NP_703146.1). In some embodiments, (wild type) OR6P1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 128366 (NCBI Reference Sequence: NP_001153797.1). In some embodiments, (wild type) OR2L2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26246 (NCBI Reference Sequence: NP_001004686.1). In some embodiments, (wild type) OR10G7 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390265 (NCBI Reference Sequence: NP_001004463.1). In some embodiments, (wild type) OR5AN1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390195 (NCBI Reference Sequence: NP_001004729.1). In some embodiments, (wild type) OR5V1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 81696 (NCBI Reference Sequence: NP_110503.3). In some embodiments, (wild type) OR2AK2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 391191 (NCBI Reference Sequence: NP_001004491.2). In some embodiments, (wild type) OR10A3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26496 (NCBI Reference Sequence: NP_001003745.1). In some embodiments, (wild type) OR10A6 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390093 (NCBI Reference Sequence: NP_001004461.1). In some embodiments, (wild type) OR10J1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26476 (NCBI Reference Sequence: NP_001350486.1). In some embodiments, (wild type) OR2J2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 26707 (NCBI Reference Sequence: NP_112167.2). The major and functional alleles or haplotypes of these ORs are used herein.
[0179] In some embodiments, the olfactory receptor protein described herein is a Class I olfactory receptor with a modified C-terminal domain. In some embodiments, the Class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1. OR51B2 may preferably be OR51B2(C120R, L134F, C209S). OR52K1 may also be OR52K1(Q52R). OR51B5 may also be OR51B5(G5S). OR56A3 may also be OR56A3(M51T).
[0180] OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3, and OR51L1 are human class I olfactory receptors. In some embodiments, (wild-type) OR52A5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390054 (NCBI Reference Sequence: NP_001005160.1). In some embodiments, (wild-type) OR52E8 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390079 (NCBI Reference Sequence: NP_001005168.2). In some embodiments, (wild-type) OR56A4 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 120793 (NCBI Reference Sequence: NP_001005179.3). In some embodiments, (wild type) OR51B2 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 79345 (NCBI Reference Sequence: NP_149420.4). In some embodiments, (wild type) OR52K1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390036 (NCBI Reference Sequence: NP_001005171.2). In some embodiments, (wild type) OR56A1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 120796 (NCBI Reference Sequence: NP_001001917.3). In some embodiments, (wild type) OR51B5 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 282763 (NCBI Reference Sequence: NP_001005567.2). In some embodiments, (wild-type) OR56A3 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 390083 (NCBI Reference Sequence: NP_001003443.2). In some embodiments, (wild-type) OR51L1 has an amino acid sequence encoded by the nucleotide sequence of NCBI gene ID 119682 (NCBI Reference Sequence: NP_001004755.1).
[0181] As explained in the experimental section, the inventors have found that certain sequences corresponding to the sequence motif of SEQ ID NO: 1 are particularly advantageous. On this basis, more preferred sequence motifs have been identified. Therefore, in a preferred embodiment, the olfactory receptor protein as described herein is such that the sequence motif is:
[0182] RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 5).
[0183] The sequence motif:
[0184] RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO:5) may alternatively be described as:
[0185] RNX1EX 3' X4X5AX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 5), wherein:
[0186] -X1 is K or R;
[0187] -X 3' It is V or M or I;
[0188] -X4 is K or R;
[0189] -X5 is any amino acid;
[0190] -X6 is L, I or V;
[0191] -X 7' It is K or R;
[0192] -X8 is K or R;
[0193] -X 10 It is L or I or F;
[0194] -X 11' It is K or R;
[0195] -X 12 is K or R; and
[0196] -X 13 It is K or R.
[0197] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is:
[0198] RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 824).
[0199] RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO:824) can alternatively be described as:
[0200] RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 824), wherein:
[0201] -X1 is K or R;
[0202] -X5 is any amino acid;
[0203] -X6 is L, I or V;
[0204] -X7 is K or R or H;
[0205] -X8 is K or R;
[0206] -X 10 It is L or I or F;
[0207] -X 11 It is K or R or G;
[0208] -X 12 is K or R; and
[0209] -X 13 It is K or R.
[0210] A modified C-terminal domain comprising the motif of SEQ ID NO: 824 may be particularly advantageous in the case of class I olfactory receptors.
[0211] The sequence motifs of SEQ ID NO: 1, SEQ ID NO: 5, and SEQ ID NO: 824 contain a residue, which can be any amino acid, denoted as "x" or "X5" above. The inventors have discovered that certain residues at this position produce particularly favorable sequences. Therefore, in a further preferred embodiment, the olfactory receptor protein as described herein and having a modified C-terminal domain comprising any of the above amino acid sequence motifs is such that "x" or "X5" is selected from any amino acid except proline (Pro, P).
[0212] In accordance with the above, in a preferred embodiment, the olfactory receptor protein as described herein is such that the sequence motif is
[0213] RN[KR][EDQ][VMIL][KR]{P}A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 311). The sequence motif
[0214] RN[KR][EDQ][VMIL][KR]{P}A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 311) can alternatively be described as:
[0215] RNX1X2X3X4X 5”” AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 311), wherein:
[0216] X1 is K or R;
[0217] X2 is E or D or Q
[0218] X3 is V, M, I, or L;
[0219] X4 is K or R;
[0220] ·X 5”” is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y (any amino acid except P);
[0221] X6 is L, I, or V;
[0222] X7 is K, R, or H;
[0223] X8 is K or R;
[0224] ·X 10 It is L or I or F;
[0225] ·X 11 It is K or R or G;
[0226] ·X 12 is K or R; and
[0227] ·X 13 It is K or R.
[0228] In another further preferred embodiment, the olfactory receptor as described herein is such that the sequence motif is
[0229] RNX1EX 3' X4X 5”” AX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO:312), wherein:
[0230] X1 is K or R;
[0231] ·X 3' It is V or M or I;
[0232] X4 is K or R;
[0233] ·X 5”” is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y (any amino acid except P);
[0234] X6 is L, I, or V;
[0235] ·X 7' It is K or R;
[0236] X8 is K or R;
[0237] ·X 10 It is L or I or F;
[0238] ·X 11' It is K or R;
[0239] ·X 12 is K or R; and
[0240] ·X 13 It is K or R.
[0241] In some embodiments, the olfactory receptor protein as described herein and having a modified C-terminal domain comprising any of the above amino acid sequence motifs is such that "x" or "X5" is selected from any amino acid except proline (Pro, P) and tryptophan (Trp, W).
[0242] In some embodiments, the olfactory receptor protein as described herein and having a modified C-terminal domain comprising any of the above amino acid sequence motifs is such that "x" or "X5" is selected from D, K, R, E, N, V, A, Q or G, preferably "x" or "X5" is selected from D, K, R, E, N, V, A or Q, more preferably "x" or "X5" is selected from D, K or R. In other words, in a further preferred embodiment, the olfactory receptor protein as described herein is such that the sequence motif is:
[0243] RN[KR][EDQ][VMIL][KR][DKRENVAQG]A[LIV][KRH][KR][LI][LIF][KRG][KR][KR](SEQ ID NO:6),
[0244] Preferably:
[0245] RN[KR][EDQ][VMIL][KR][DKRENVAQ]A[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 7), more preferably:
[0246] RN[KR][EDQ][VMIL][KR][DKR]A[LIV][KRH][KR][LI][LIF][KRG][KR][KR](SEQID NO:166);
[0247] Or the sequence motif is:
[0248] RN[KR]E[VMI][KR][DKRENVAQG]A[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ IDNO: 8), preferably:
[0249] RN[KR]E[VMI][KR][DKRENVAQ]A[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 9), more preferably:
[0250] RN[KR]E[VMI][KR][DKR]A[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 167).
[0251] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is:
[0252] RN[KR][EDQ][VMIL][KR]KA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 820).
[0253] The sequence motif RN[KR][EDQ][VMIL][KR]KA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 820)
[0254] Can alternatively be described as RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 820), wherein:
[0255] -X1 is K or R;
[0256] -X2 is E or D or Q;
[0257] -X3 is V or M or I or L;
[0258] -X4 is K or R;
[0259] -X6 is L, I or V;
[0260] -X7 is K or R or H;
[0261] -X8 is K or R;
[0262] -X 10 It is L or I or F;
[0263] -X 11 It is K or R or G;
[0264] -X 12 is K or R; and
[0265] -X 13 It is K or R.
[0266] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is:
[0267] RN[KR]E[VMI][KR]KA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 821).
[0268] The sequence motif:
[0269] RN[KR]E[VMI][KR]KA[LIV][KR][KR]L[LIF][KR][KR][KR](SEQ ID NO:821)
[0270] It can alternatively be described as:
[0271] RNX1EX 3' X4KAX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 821), wherein:
[0272] -X1 is K or R;
[0273] -X 3' It is V or M or I;
[0274] -X4 is K or R;
[0275] -X5 is any amino acid;
[0276] -X6 is L, I or V;
[0277] -X 7' It is K or R;
[0278] -X8 is K or R;
[0279] -X 10 It is L or I or F;
[0280] -X 11' It is K or R;
[0281] -X 12 is K or R; and
[0282] -X 13 It is K or R.
[0283] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is:
[0284] RNKEVKX 5”” ALKRLLKRK (SEQ ID NO: 319), wherein:
[0285] -X 5”” It is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y.
[0286] Several specific and particularly advantageous sequences corresponding to the sequence motifs described herein have been identified. Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0287] RNKEVKDALKRLLKRK(SEQ ID NO:10),
[0288] RNREVKDALKRLLKRK(SEQ ID NO:11),
[0289] RNKEIKDALKRLLKRK(SEQ ID NO:12),
[0290] RNKEMKDALKRLLKRK(SEQ ID NO:13),
[0291] RNKEVRDALKRLLKRK(SEQ ID NO:14),
[0292] RNKEVKKALKRLLKRK(SEQ ID NO:15),
[0293] RNKEVKEALKRLLKRK(SEQ ID NO:16),
[0294] RNKEVKNALKRLLKRK(SEQ ID NO:17),
[0295] RNKEVKRALKRLLKRK(SEQ ID NO:18),
[0296] RNKEVKVALKRLLKRK(SEQ ID NO:19),
[0297] RNKEVKAALKRLLKRK(SEQ ID NO:20),
[0298] RNKEVKQALKRLLKRK(SEQ ID NO:21),
[0299] RNKEVKDAVKRLLKRK(SEQ ID NO:22),
[0300] RNKEVKDAIKRLLKRK(SEQ ID NO:23),
[0301] RNKEVKDALRRLLKRK(SEQ ID NO:24),
[0302] RNKEVKDALKKLLKRK(SEQ ID NO:25),
[0303] RNKEVKDALKRLIKRK(SEQ ID NO:26),
[0304] RNKEVKDALKRLFKRK(SEQ ID NO:27),
[0305] <h2 style=";text-align:left;direction:ltr">SEQ ID NO:33),RNKEVKGALKRLLKRK(SEQ ID NO:34),RNKEVKDALHRLLKRK(SEQ ID NO:35),RNKEVKDALKRILKRK(SEQ ID NO:36),RNKEVKDALKRLLGRK(SEQ ID NO:37),RNKEVKRAIKRLLKRK(SEQ ID NO:38),RNKEVKKAIKRLLKRK(SEQ ID NO:39),RNKEVKRAIKRLFKRK(SEQ ID NO:40),RNKEVKKAIKRLFKRK(SEQ ID NO:41),RNKEVKRAIRKLLKRK(SEQ ID NO:42),RNKEVKDALRKLLKRK(SEQ ID NO:43),RNKEVKDALKRLLRRR(SEQ ID NO:44),RNREMRKALHRLLGKK(SEQ ID NO:827),RNREVKKAIHKLIGRK(SEQ ID NO:828),RNREVRKAVHRLFKRK(SEQ ID NO:829),RNKEMKKAIHKLFGKK(SEQ ID NO:830),RNRDVKKAVHKLFRRK(SEQ ID NO:831),RNRDMKKAVHKLFGKR(SEQ ID NO:832),RNKELRKALHKLLGRK(SEQ ID NO:833),RNRDVRKALRRILRRR(SEQ ID NO:834),RNKDVRKAVRKLIRRR(SEQ ID NO:835),RNRDVRKAVRRLFRKR(SEQ ID NO:836),RNKDIKKAVKKLIKKK(SEQ ID NO:837),RNRELRKAVRRLFKRR(SEQ ID NO:838),RNKELRKAVRKIIKKK(SEQ ID NO:839),RNRDVKKAVRRLFRRK(SEQ ID NO:840),<h2 style=";text-align:left;direction:ltr">RNREVRKALRRIIRKR(SEQ ID NO:841),RNKDIRKAVKKIFRRK(SEQ ID NO:842),RNKDVRKAVRRLIKRK(SEQ ID NO:843),RNRDLRKAVRKLFKKK(SEQ ID NO:844),RNRDLRKALRRIFKRR(SEQ ID NO:845),RNRDVRKAIKKLIRKR(SEQ ID NO:846),RNKELKKAIKRILKKK(SEQ ID NO:847),RNRDVRKAIRKLLKRK(SEQ ID NO:848),RNRDLRKAVRRIFKKR(SEQ ID NO:849),RNRDVRKAVRKLFKRR(SEQ ID NO:850),RNRDVRKALRRLFKKR(SEQ ID NO:851),RNKELKKALRKLIGKK(SEQ ID NO:852),RNREMRKAIKKIIKKK(SEQ ID NO:853),RNKEIKKAIKKIIKKR(SEQ ID NO:854),RNRDVKKAIRRLFRRR(SEQ ID NO:855),RNREVKKAVKKLIGKR(SEQ ID NO:856),RNREMRKALRRLFRKR(SEQ ID NO:857),RNKELKKALRRLIGRR(SEQ ID NO:858),RNRDVKKALRKLIGKR(SEQ ID NO:859),RNREVKKAVKKLIRRK(SEQ ID NO:860),RNKEVRKALKKLFGKK(SEQ ID NO:861),RNKEIRKALRRLFGKK(SEQ ID NO:862),RNKDVKKALRRLFGKK(SEQ ID NO:863),RNKELKKAIKRLIRRK(SEQ ID NO:864),RNKDVRKAVKRLLKKR(SEQ ID NO:865),RNKELRKAIRRLLRRR(SEQ ID NO:866),RNRDIRKALRKLFKKK(SEQ ID NO:867),RNRELKKALRRLLRRR(SEQ ID NO:868),RNREVKKALRRLFGKK(SEQ ID NO:869),RNRDVRKALKRLLKRK(SEQ ID NO:870),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0306] RNRDMRKAIRKLFGRK(SEQ ID NO:871),
[0307] RNRELKKAIRKLLKRK(SEQ ID NO:872),
[0308] RNRDIRKAVKKLFGKK(SEQ ID NO:873),
[0309] RNKEVKKAIRKLFGRR(SEQ ID NO:874),
[0310] RNREVRKAVRKLFRRK(SEQ ID NO:875),
[0311] RNRDMKKALKKLFRRR(SEQ ID NO:876),
[0312] RNRDVRKALKRLLGRR(SEQ ID NO:877),
[0313] RNKDLKKAVKKLFGRK(SEQ ID NO:878),
[0314] RNKDVRKAVRRLFGRR(SEQ ID NO:879),
[0315] RNKEVKCALKRLLKRK(SEQ ID NO:880),
[0316] RNKEVKFALKRLLKRK(SEQ ID NO:881),
[0317] RNKEVKHALKRLLKRK(SEQ ID NO:882),
[0318] RNKEVKIALKRLLKRK(SEQ ID NO:883),
[0319] RNKEVKLALKRLLKRK(SEQ ID NO:884),
[0320] RNKEVKMALKRLLKRK(SEQ ID NO:885),
[0321] RNKEVKSALKRLLKRK(SEQ ID NO:886),
[0322] RNKEVKTALKRLLKRK(SEQ ID NO:887),
[0323] RNKEVKWALKRLLKRK (SEQ ID NO: 888), and
[0324] RNKEVKYALKRLLKRK (SEQ ID NO:889).
[0325] Olfactory receptors having modified C-terminal domains containing any of the specific sequences listed above have been shown to exhibit advantageous and surprising technical effects, as described in detail in the experimental section of this disclosure. Based on these insights, a skilled person can design further advantageous sequences that are correspondingly suitable for the sequence motifs described herein. Some illustrative and non-limiting examples of such sequences are as follows:
[0326] RNKEVKRALKRLLRRR(SEQ ID NO:45)
[0327] RNKEVKKALKRLLRRR(SEQ ID NO:46)
[0328] RNREVKRAIKRLLKRK(SEQ ID NO:47)
[0329] RNREVKKAIKRLLKRK(SEQ ID NO:48)
[0330] RNREVKRAIKRLFKRK(SEQ ID NO:49)
[0331] RNREVKKAIKRLFKRK(SEQ ID NO:50)
[0332] RNREVKRAIRKLLKRK(SEQ ID NO:51)
[0333] RNREVKDALRKLLKRK(SEQ ID NO:52)
[0334] RNREVKDALKRLLRRR(SEQ ID NO:53)
[0335] RNKEVKKAIKRLLRRK(SEQ ID NO:54)
[0336] RNKEVKKAIKRLLKKK(SEQ ID NO:55)
[0337] RNKEVKKAIKRLLKRR(SEQ ID NO:56)
[0338] RNKEVKRAIKRLLRRK(SEQ ID NO:57) <h2 style=";text-align:left;direction:ltr">
[0339] <h2 style=";text-align:left;direction:ltr"> RNKEVKRAIKRLLKKK(SEQ ID NO:58)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0340] <h2 style=";text-align:left;direction:ltr"> RNKEVKRAIKRLLKRR(SEQ ID NO:59)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0341] <h2 style=";text-align:left;direction:ltr"> RNKEVKKAIKRLFRRK(SEQ ID NO:60)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0342] <h2 style=";text-align:left;direction:ltr"> RNKEVKKAIKRLFKKK(SEQ ID NO:61)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0343] <h2 style=";text-align:left;direction:ltr"> RNKEVKKAIKRLFKRR(SEQ ID NO:62)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0344] <h2 style=";text-align:left;direction:ltr"> RNKEVKRAIKRLFRRK(SEQ ID NO:63)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0345] <h2 style=";text-align:left;direction:ltr"> RNKEVKRAIKRLFKKK(SEQ ID NO:64)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0346] <h2 style=";text-align:left;direction:ltr"> RNKEVKRAIKRLFKRR(SEQ ID NO:65)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0347] <h2 style=";text-align:left;direction:ltr"> RNREVKRAIKRLLRKK(SEQ ID NO:66)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0348] <h2 style=";text-align:left;direction:ltr"> RNREVKKAIKRLLRKK(SEQ ID NO:67)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0349] <h2 style=";text-align:left;direction:ltr"> RNREVKRAIKRLFRRR(SEQ ID NO:68)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0350] <h2 style=";text-align:left;direction:ltr"> RNREVKKAIKRLFRRR(SEQ ID NO:69)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0351] <h2 style=";text-align:left;direction:ltr"> RNREVKKAIKRLFRRK(SEQ ID NO:70)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0352] <h2 style=";text-align:left;direction:ltr"> RNREVKKAIKRLFKKK(SEQ ID NO:71)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0353] <h2 style=";text-align:left;direction:ltr"> RNREVKKAIKRLFKRR(SEQ ID NO:72)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0354] <h2 style=";text-align:left;direction:ltr"> RNREVKRAIKRLFRRK(SEQ ID NO:73)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0355] <h2 style=";text-align:left;direction:ltr"> RNREVKRAIKRLFKKK(SEQ ID NO:74)<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0356] <h2 style=";text-align:left;direction:ltr"> RNREVKRAIKRLFKRR(SEQ ID NO:75)
[0357] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is
[0358] RNRDVRKALRRLFRKK (SEQ ID NO: 307) or RNRDVRRALRRLFRKK (SEQ ID NO: 308).Such C-terminal motifs may be particularly advantageous because they contain statistically optimal amino acids at each individual position.
[0359] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is RNKQIRDALKRLLKRK (SEQ ID NO: 890).Such a C-terminal motif may be particularly advantageous in the case of class I olfactory receptors.
[0360] The 16 amino acid long sequence motifs described herein may optionally include additional C-terminal residues. In fact, the inventors have found that, although not required, such additional C-terminal residues may lead to further beneficial effects, as described in detail in the experimental section. The number of additional C-terminal residues (if present) is not particularly limited, but is preferably less than 10. More preferably, the number of additional C-terminal residues (if present) is 1 to 6. Therefore, depending on the number of additional C-terminal residues added to the 16 amino acid long sequence motifs described herein, when the optional additional C-terminal residues are present, the total length of the sequence motifs described herein can be 17, 18, 19, 20, 21, 22, 23, 24, 25 or 26 amino acids, preferably 17, 18, 19, 20, 21 or 22 amino acids.
[0361] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the amino acid sequence motif comprises 1 to 10, preferably 1 to 6, additional C-terminal amino acid residues. In some embodiments, the additional C-terminal amino acid residues are as follows:
[0362] - the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P and Y, more preferably the first additional amino acid residue is C, R or K, most preferably the first additional amino acid residue is C;
[0363] - the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably the second additional amino acid residue is C or R;
[0364] - the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K;
[0365] - the fourth to tenth, or the fourth, fifth and sixth, preferably the fourth, fifth and sixth additional amino acid residues are selected from K and R.
[0366] In some embodiments, the amino acid sequence motif comprises an additional C-terminal residue. The additional amino acid residue preferably corresponds to the first additional amino acid residue defined above, more preferably selected from C, R, P, L, K, G, Y, F, M or W.
[0367] In some embodiments, the amino acid sequence motif comprises two additional C-terminal residues. The first and second additional amino acid residues are preferably as defined above. Specific favorable examples of the combination of the first and second additional amino acid residues include CC, SI, YP, PQ, FR, CR, RR, EK, PR, CG, FK, RG, RC, RT, RF, GG, YR, GC, TG, PC, HP, PG, KY, CP, YY, FF, CF, NP, YL, IC, HC, CL, YC, ER, RP, PA, FC, and RY. In this case, the CC sequence is particularly preferred.
[0368] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0369] RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 CC (SEQ ID NO: 822),
[0370] RNX1EX 3’ X4KAX6X 7’ X8LX 10 X 11’ X 12 X 13 CC (SEQ ID NO: 823), and;
[0371] RNKEVKX 5”” ALKRLLKRKCC (SEQ ID NO: 320), wherein:
[0372] -X1 is K or R;
[0373] -X2 is E or D or Q;
[0374] -X3 is V or M or I or L;
[0375] -X 3' It is V or M or I;
[0376] -X4 is K or R;
[0377] -X 5”” is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y;
[0378] -X6 is L, I or V;
[0379] -X7 is K or R or H;
[0380] -X 7' It is K or R;
[0381] -X8 is K or R;
[0382] -X 10 It is L or I or F;
[0383] -X 11 It is K or R or G;
[0384] -X 11' It is K or R;
[0385] -X 12 is K or R; and
[0386] -X 13 It is K or R.
[0387] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0388] RNKEVKDALKRLLKRKCC(SEQ ID NO:86),
[0389] RNREVKDALKRLLKRKCC(SEQ ID NO:87),
[0390] RNKEIKDALKRLLKRKCC(SEQ ID NO:88),
[0391] RNKEMKDALKRLLKRKCC(SEQ ID NO:89),
[0392] RNKEVRDALKRLLKRKCC(SEQ ID NO:90),
[0393] RNKEVKKALKRLLKRKCC(SEQ ID NO:91),
[0394] RNKEVKEALKRLLKRKCC(SEQ ID NO:92),
[0395] RNKEVKNALKRLLKRKCC(SEQ ID NO:93),
[0396] RNKEVKRALKRLLKRKCC(SEQ ID NO:94),
[0397] RNKEVKVALKRLLKRKCC(SEQ ID NO:95),
[0398] RNKEVKAALKRLLKRKCC(SEQ ID NO:96),
[0399] RNKEVKQALKRLLKRKCC(SEQ ID NO:97),
[0400] RNKEVKDAVKRLLKRKCC(SEQ ID NO:98),
[0401] RNKEVKDAIKRLLKRKCC(SEQ ID NO:99),
[0402] RNKEVKDALRRLLKRKCC(SEQ ID NO:100),
[0403] RNKEVKDALKKLLKRKCC(SEQ ID NO:101),
[0404] RNKEVKDALKRLIKRKCC(SEQ ID NO:102),
[0405] RNKEVKDALKRLFKRKCC(SEQ ID NO:103),
[0406] RNKEVKDALKRLLRRKCC(SEQ ID NO:104),
[0407] RNKEVKDALKRLLKKKCC(SEQ ID NO:105),
[0408] RNKEVKDALKRLLKRRCC(SEQ ID NO:106),
[0409] RNKDVKDALKRLLKRKCC(SEQ ID NO:107),
[0410] RNKQVKDALKRLLKRKCC(SEQ ID NO:108),
[0411] RNKELKDALKRLLKRKCC(SEQ ID NO:109),
[0412] <h2 style=";text-align:left;direction:ltr">RNKEVKGALKRLLKRKCC(SEQ ID NO:110),RNKEVKDALHRLLKRKCC(SEQ ID NO:111),RNKEVKDALKRILKRKCC(SEQ ID NO:112),RNKEVKDALKRLLGRKCC(SEQ ID NO:113),RNKEVKRAIKRLLKRKCC(SEQ ID NO:114),RNKEVKKAIKRLLKRKCC(SEQ ID NO:115),RNKEVKRAIKRLFKRKCC(SEQ ID NO:116),RNKEVKKAIKRLFKRKCC(SEQ ID NO:117),RNKEVKRAIRKLLKRKCC(SEQ ID NO:118),RNKEVKDALRKLLKRKCC(SEQ ID NO:119),RNKEVKDALKRLLRRRCC(SEQ ID NO:120),RNREMRKALHRLLGKKCC(SEQ ID NO:254),RNREVKKAIHKLIGRKCC(SEQ ID NO:255),RNREVRKAVHRLFKRKCC(SEQ ID NO:256),RNKEMKKAIHKLFGKKCC(SEQ ID NO:257),RNRDVKKAVHKLFRRKCC(SEQ ID NO:258),RNRDMKKAVHKLFGKRCC(SEQ ID NO:259),RNKELRKALHKLLGRKCC(SEQ ID NO:260),RNRDVRKALRRILRRRCC(SEQ ID NO:261),RNKDVRKAVRKLIRRRCC(SEQ ID NO:262),RNRDVRKAVRRLFRKRCC(SEQ ID NO:263),RNKDIKKAVKKLIKKKCC(SEQ ID NO:264),RNRELRKAVRRLFKRRCC(SEQ ID NO:265),RNKELRKAVRKIIKKKCC(SEQ ID NO:266),RNRDVKKAVRRLFRRKCC(SEQ ID NO:267),RNREVRKALRRIIRKRCC(SEQ ID NO:268),RNKDIRKAVKKIFRRKCC(SEQ ID NO:269),RNKDVRKAVRRLIKRKCC(SEQ ID NO:270),RNRDLRKAVRKLFKKKCC(SEQ ID NO:271),<h2 style=";text-align:left;direction:ltr">RNRDLRKALRRIFKRRCC(SEQ ID NO:272),RNRDVRKAIKKLIRKRCC(SEQ ID NO:273),RNKELKKAIKRILKKKCC(SEQ ID NO:274),RNRDVRKAIRKLLKRKCC(SEQ ID NO:275),RNRDLRKAVRRIFKKRCC(SEQ ID NO:276),RNRDVRKAVRKLFKRRCC(SEQ ID NO:277),RNRDVRKALRRLFKKRCC(SEQ ID NO:278),RNKELKKALRKLIGKKCC(SEQ ID NO:279),RNREMRKAIKKIIKKKCC(SEQ ID NO:280),RNKEIKKAIKKIIKKRCC(SEQ ID NO:281),RNRDVKKAIRRLFRRRCC(SEQ ID NO:282),RNREVKKAVKKLIGKRCC(SEQ ID NO:283),RNREMRKALRRLFRKRCC(SEQ ID NO:284),RNKELKKALRRLIGRRCC(SEQ ID NO:285),RNRDVKKALRKLIGKRCC(SEQ ID NO:286),RNREVKKAVKKLIRRKCC(SEQ ID NO:287),RNKEVRKALKKLFGKKCC(SEQ ID NO:288),RNKEIRKALRRLFGKKCC(SEQ ID NO:289),RNKDVKKALRRLFGKKCC(SEQ ID NO:290),RNKELKKAIKRLIRRKCC(SEQ ID NO:291),RNKDVRKAVKRLLKKRCC(SEQ ID NO:292),RNKELRKAIRRLLRRRCC(SEQ ID NO:293),RNRDIRKALRKLFKKKCC(SEQ ID NO:294),RNRELKKALRRLLRRRCC(SEQ ID NO:295),RNREVKKALRRLFGKKCC(SEQ ID NO:296),RNRDVRKALKRLLKRKCC(SEQ ID NO:297),RNRDMRKAIRKLFGRKCC(SEQ ID NO:298),RNRELKAIRKLLKRKCC(SEQ ID NO:299),RNRDIRKAVKKLFGKKCC(SEQ ID NO:300),<h2 style=";text-align:left;direction:ltr">RNKEVKKAIRKLFGRRCC(SEQ ID NO:301),RNREVRKAVRKLFRRKCC(SEQ ID NO:302),RNRDMKKALKKLFRRRCC(SEQ ID NO:303),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0413] <h2 style=";text-align:left;direction:ltr"> RNRDVRKALKRLLGRRCC(SEQ ID NO:304),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0414] <h2 style=";text-align:left;direction:ltr"> RNKDLKKAVKKLFGRKCC(SEQ ID NO:305),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0415] <h2 style=";text-align:left;direction:ltr"> RNKDVRKAVRRLFGRCC(SEQ ID NO:306),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0416] <h2 style=";text-align:left;direction:ltr"> RNRDVRKALRRLFRKKCC(SEQ ID NO:309),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0417] <h2 style=";text-align:left;direction:ltr"> RRNRDVRRALRRLFRKCC(SEQ ID NO:310),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0418] <h2 style=";text-align:left;direction:ltr"> RNKEVKCALKRLLKRKCC(SEQ ID NO:321),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0419] <h2 style=";text-align:left;direction:ltr"> RNKEVKFALKRLLKRKCC(SEQ ID NO:322),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0420] <h2 style=";text-align:left;direction:ltr"> RNKEVKHALKRLLKRKCC(SEQ ID NO:323),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0421] <h2 style=";text-align:left;direction:ltr"> RNKEVKIALKRLLKRKCC(SEQ ID NO:324),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0422] <h2 style=";text-align:left;direction:ltr"> RNKEVKLALKRLLKRKCC(SEQ ID NO:325),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0423] <h2 style=";text-align:left;direction:ltr"> RNKEVKMALKRLLKRKCC(SEQ ID NO:326),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0424] <h2 style=";text-align:left;direction:ltr"> RNKEVKSALKRLLKRKCC(SEQ ID NO:328),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0425] <h2 style=";text-align:left;direction:ltr"> RNKEVKTALKRLLKRKCC(SEQ ID NO:329),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0426] <h2 style=";text-align:left;direction:ltr"> RNKEVKWALKRLLKRKCC(SEQ ID NO:330),<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0427] RNKEVKYALKRLLKRKCC (SEQ ID NO: 331), and
[0428] RNKQIRDALKRLLKRKCC (SEQ ID NO:740).
[0429] C-terminal motifs having the sequence RNRDVRKALRRLFRKKCC (SEQ ID NO: 309) or RNRDVRRALRRLFRKKCC (SEQ ID NO: 310) may be particularly advantageous because they contain statistically optimal amino acids at each individual position. A C-terminal motif having the sequence RNKQIRDALKRLLKRKCC (SEQ ID NO: 740) may be particularly advantageous in the context of class I olfactory receptors.
[0430] In addition to the CC sequence, other preferred combinations of the first and second additional amino acid residues are CR, RR, YP, RF and FK. Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0431] RNKEVKRAIKRLLKRKCR(SEQ ID NO:121),
[0432] RNKEVKKAIKRLLKRKCR(SEQ ID NO:122),
[0433] RNKEVKRALKRLLKRKRR(SEQ ID NO:123),
[0434] RNKEVKRALKRLLKRKYP(SEQ ID NO:124),
[0435] RNKEVKRALKRLLKRKRF(SEQ ID NO:125),
[0436] RNKEVKRALKRLLKRKFK(SEQ ID NO:126),
[0437] RNKEVKKALKRLLKRKRR(SEQ ID NO:127),
[0438] RNKEVKKALKRLLKRKYP(SEQ ID NO:128),
[0439] RNKEVKKALKRLLKRKRF(SEQ ID NO:129), and
[0440] RNKEVKKALKRLLKRKFK (SEQ ID NO: 130).
[0441] In some embodiments, the amino acid sequence motif comprises three additional C-terminal residues. The first, second, and third additional amino acid residues are preferably as defined above. In some embodiments, the first and second additional amino acid residues are CC, and the third additional amino acid residue is as defined above. Specific advantageous examples of combinations of the first, second, and third additional amino acid residues include CRR, CCC, CCF, CCL, CCM, CCS, CCP, CCA, CCY, CCH, CCN, CCD, CCK, CCR, and CCG.
[0442] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0443] RNKEVKDALKRLLKRKCRR(SEQ ID NO:133),
[0444] RNKEVKDALKRLLKRKCCC(SEQ ID NO:134),
[0445] RNKEVKDALKRLLKRKCCF(SEQ ID NO:135),
[0446] RNKEVKDALKRLLKRKCCL(SEQ ID NO:136),
[0447] RNKEVKDALKRLLKRKCCM(SEQ ID NO:137),
[0448] RNKEVKDALKRLLKRKCCS(SEQ ID NO:138),
[0449] RNKEVKDALKRLLKRKCCP(SEQ ID NO:139),
[0450] RNKEVKDALKRLLKRKCCA(SEQ ID NO:140),
[0451] RNKEVKDALKRLLKRKCCY(SEQ ID NO:141),
[0452] RNKEVKDALKRLLKRKCCH(SEQ ID NO:142),
[0453] RNKEVKDALKRLLKRKCCN(SEQ ID NO:143),
[0454] RNKEVKDALKRLLKRKCCD(SEQ ID NO:144),
[0455] RNKEVKDALKRLLKRKCCK(SEQ ID NO:145),
[0456] RNKEVKDALKRLLKRKCCR (SEQ ID NO: 146), and
[0457] RNKEVKDALKRLLKRKCCG (SEQ ID NO: 147).
[0458] In some embodiments, the amino acid sequence motif comprises four additional C-terminal residues. The first, second, third and fourth additional amino acid residues are preferably as defined above. Specific advantageous examples of the combination of the first, second, third and fourth additional amino acid residues include CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CCRR (SEQ ID NO: 161), CCRK (SEQ ID NO: 228) and CCKR (SEQ ID NO: 229), of which CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160) and CCRR (SEQ ID NO: 161) are preferred. Preferably, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0459] RNKEVKDALKRLLKRKCRRR(SEQ ID NO:149),
[0460] RNKEVKDALKRLLKRKCRKK(SEQ ID NO:150), and
[0461] RNKEVKDALKRLLKRKCCRR (SEQ ID NO: 151).
[0462] In some embodiments, the amino acid sequence motif comprises 5 additional C-terminal residues. The first, second, third, fourth and fifth additional amino acid residues are preferably as defined above. Specific favorable examples of the first, second, third, fourth and fifth additional amino acid residue combinations include CRRRR (SEQ ID NO: 162), CCRRR (SEQ ID NO: 163), CCKRR (SEQ ID NO: 230), CCRKR (SEQ ID NO: 231), CCRRK (SEQ ID NO: 232), CCRKK (SEQ ID NO: 233), CCKRK (SEQ ID NO: 234), CCKKR (SEQ ID NO: 235), and CCKKK (SEQ ID NO: 236), wherein CRRRR (SEQ ID NO: 162) and CCRRR (SEQ ID NO: 163) are preferred.
[0463] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0464] RNX1X2X3X4X5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 825), and
[0465] RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 826), wherein:
[0466] -X1 is K or R;
[0467] -X2 is E or D or Q;
[0468] -X3 is V or M or I or L;
[0469] -X4 is K or R;
[0470] -X5 is any amino acid;
[0471] -X6 is L, I or V;
[0472] -X7 is K or R or H;
[0473] -X8 is K or R;
[0474] -X 10 It is L or I or F;
[0475] -X 11 It is K or R or G;
[0476] -X 12 is K or R; and
[0477] -X 13 It is K or R.
[0478] Preferably, X5 is any amino acid except proline (Pro, P).
[0479] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0480] RNKEVKDALKRLLKRKCRRRR(SEQ ID NO:154),
[0481] RNKEVKDALKRLLKRKCCRRR(SEQ ID NO:156),
[0482] RNKEVKKAIKRLLKRKCCRRR(SEQ ID NO:220),
[0483] RNKEVKRAIKRLLKRKCCRRR(SEQ ID NO:237),
[0484] RNKEVKKAIKRLFKRKCCRRR(SEQ ID NO:221),
[0485] RNKEVKRAIKRLFKRKCCRRR (SEQ ID NO:238), and
[0486] RNKQIRDALKRLLKRKCCRRR(SEQ ID NO:741),
[0487] Preferably, the sequence motif is RNKEVKKAIKRLFKRKCCRRR (SEQ ID NO: 221).
[0488] A C-terminal motif having the sequence RNKQIRDALKRLLKRKCCRRR (SEQ ID NO: 741) may be particularly advantageous in the case of class I olfactory receptors.
[0489] In some embodiments, the amino acid sequence motif comprises six additional C-terminal residues. The first, second, third, fourth, fifth, and sixth additional amino acid residues are preferably as defined above. Specific advantageous examples of combinations of the first, second, third, fourth, fifth, and sixth additional amino acid residues include CRRRRR (SEQ ID NO: 164), CRRRKK (SEQ ID NO: 165), and CCRRRR (SEQ ID NO: 224).
[0490] Thus, in some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of:
[0491] RNKEVKDALKRLLKRKCRRRRR(SEQ ID NO:157),
[0492] RNKEVKDALKRLLKRKCRRRKK (SEQ ID NO:158), and
[0493] RNKEVKDALKRLLKRKCCRRRR (SEQ ID NO: 219).
[0494] Olfactory receptors having modified C-terminal domains comprising any of these specific sequences listed above have been shown to exhibit advantageous and surprising technical effects, as described in detail in the experimental section of this disclosure.
[0495] In some embodiments, the olfactory receptor protein as described herein can be such that the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164), CRRRKK (SEQ ID NO: 165).
[0496] In some embodiments, the olfactory receptors described herein further include an N-terminal signal peptide. Typically, the N-terminal signal peptide is a cleavable peptide. This means that it is cleaved from the mature protein and cannot alter OR-ligand binding and signaling. Therefore, the skilled artisan will appreciate that the N-terminal signal peptides described herein are generally not part of the mature olfactory receptor. In some embodiments, the N-terminal signal peptide is a leucine-rich signal peptide, preferably MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77). MRPQILLLLALLTLGLA (SEQ ID NO: 76) and MSHQILLLLALLTLGLA (SEQ ID NO: 77) are referred to as so-called Lucy-tags. MRPQILLLLALLTLGLA (SEQ ID NO: 76) is a human Lucy-tag, while MSHQILLLLALLTLGLA (SEQ ID NO: 77) is a mouse Lucy-tag. Also encompassed are sequences of SEQ ID NOs: 76 and 77 in which 1, 2, 3, 4 or up to 5 amino acids are substituted, deleted, added or inserted. Substitutions, especially conservative substitutions, are preferred.
[0497] In some embodiments, the olfactory receptors described herein further comprise an N-terminal tag peptide. Typically, the N-terminal tag peptide is a non-cleavable peptide.
[0498] The N-terminal tag peptide can be an epitope tag for purification or capture of proteins. An example of such an epitope tag is the FLAG tag, which will be further described below.
[0499] N-terminal tag peptides may also be peptides that promote expression. Examples of such tag peptides that promote expression are rhodopsin (rho) tags, SST3 tags, and M3 tags, as described below.
[0500] In some embodiments, the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag (the 45-N-terminal amino acid of the somatostatin 3 receptor; an example thereof is SEQ ID NO: 223), an M3-Tag (the 61-N-terminal amino acid of the muscarinic acetylcholine receptor M3; an example thereof is SEQ ID NO: 222), a c-myc tag, and an HA tag, preferably selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, even more preferably selected from the group consisting of a FLAG tag and a rhodopsin (rho) tag, and most preferably a rhodopsin (rho) tag. Therefore, in some embodiments, the olfactory receptor protein as described herein further comprises an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag, an M3-Tag, a c-myc tag and an HA tag, more preferably selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag and an M3 tag, even more preferably selected from the group consisting of a FLAG tag and a rhodopsin (rho) tag, and most preferably a rhodopsin (rho) tag.
[0501] In a preferred embodiment, the N-terminal tag peptide comprises at least one tag peptide selected from the group consisting of a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, preferably a rho tag. Optionally, a FLAG peptide may further be present.
[0502] Combinations of the above-mentioned tag peptides can also be used. Typically, such a combination will include at least one of a rho tag, an SST3 tag, and an M3-Tag, preferably at least a rho tag. For example, in some embodiments, the N-terminal tag peptide is a combination of a Rho tag and a FLAG tag. In other words, in some embodiments, the olfactory receptor protein as described herein further includes an N-terminal tag peptide, wherein the N-terminal tag peptide includes a Rho tag and a FLAG tag. Preferably, in this case, the FLAG tag is located at the N-terminus of the Rho tag, for example as shown in SEQ ID NO:81.
[0503] The FLAG tag and rho tag are described in Shepard et al. (2013) PloS One 8(7):e68758, Zhuang and Matsunami (2007) J Biol Chem 282(20):15284-15293, and WO2014 / 037800, each of which is incorporated herein by reference. The IL-6 tag is described in Noe et al. A bi-functional IL-6- As a tool to measure the cell-surface expression of recombinant odorant receptors and to facilitate their activity quantification. J Biol Methods. 2017, 4 (4): e82 is described, which is incorporated herein by reference. SST3 labeling and M3 labeling are described in Tan et al., (2022) Scientific reports 12: 17658, which is incorporated herein by reference.
[0504] In some embodiments, a FLAG tag as described herein has the sequence of SEQ ID NO: 78. In some embodiments, a rho tag as described herein has the sequence of SEQ ID NO: 79. In some embodiments, an IL-6 tag as described herein is a bi-functional IL-6- as a tool to measure the cell-surface expression of recombinant odorant receptors and tofacilitate their activity quantification.J Biol Methods.2017,4(4):e82 IL-6- This document is incorporated herein by reference. In some embodiments, the SST3 tag as described herein has the sequence of SEQ ID NO: 223. In some embodiments, the M3 tag as described herein has the sequence of SEQ ID NO: 222. Also contemplated are sequences of SEQ ID NOs: 78, 79, 222, and 223, wherein 1, 2, 3, 4, or up to 5 amino acids are substituted, deleted, added, or inserted. Substitutions are preferred, particularly conservative substitutions.
[0505] The N-terminal signal peptide as described herein and the N-terminal tag peptide as described herein can be advantageously used in combination with each other. Typically, the N-terminal signal peptide will be located upstream (N-terminal) of the N-terminal tag. For example, the olfactory receptor as described herein may further comprise an N-terminal signal peptide and one or more N-terminal tag peptides. In some embodiments, the olfactory receptor as described herein may further comprise:
[0506] - human or mouse Lucy signal peptide (such as SEQ ID NO: 76 or 77);
[0507] - a FLAG tag (such as SEQ ID NO: 78) or an IL-6 tag, preferably a FLAG tag (such as SEQ ID NO: 78); and
[0508] - a rho tag (such as SEQ ID NO: 79), an SST3 tag, or an M3 tag, preferably a rho tag (such as SEQ ID NO: 79).
[0509] SEQ ID NO: 80 is an example of a nucleotide sequence encoding a combination of mouse Lucy signal peptide, FLAG tag peptide, and rho tag peptide (SEQ ID NO: 81).
[0510] In some embodiments, the olfactory receptors described herein are modified to include one or more additional N-terminal glycosylation sites. For example, such glycosylation sites are present in M3 and SST3 tags (Tan et al., Scientific Reports (2022) 12: 17658) and N-terminal rho tags (Kaushal et al., 1998, Proc Natl Acad Sci USA 91 (9): 4024-4028), described elsewhere herein.
[0511] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of SEQ ID NO: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, 820-890. In this context, sequences of SEQ ID NOs: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741 and 820-890 are also encompassed, in which 1, 2, 3, 4, 5, 6, 7, 8 or up to 9 amino acids are substituted, deleted, added or inserted. Preferably, the sequence comprising substitutions, deletions, additions and / or insertions still corresponds to the general sequence motif of SEQ ID NO: 1. Substitutions, in particular conservative substitutions, are preferred. In this context, examples of particularly suitable amino acid substitutions include substitutions of K for R and substitutions of R for K.
[0512] In some embodiments, the olfactory receptor protein as described herein is such that the sequence motif is selected from the group consisting of SEQ ID NO: 1, 10-75, 86-130, 133-147, 149-151, 154, 156-158, 198, 219-221, 254-310, 321-326, 328-331, 740-741 and 827-890. In this context, the sequences of SEQ ID NOs: 1, 10-75, 86-130, 133-147, 149-151, 154, 156-158, 198, 219-221, 254-310, 321-326, 328-331, 740-741 and 827-890 are also encompassed, in which 1, 2, 3, 4, 5, 6, 7, 8 or up to 9 amino acids are substituted, deleted, added or inserted. Substitutions, especially conservative substitutions, are preferred. In this context, examples of particularly suitable amino acid substitutions include substitutions of K for R and substitutions of R for K.
[0513] It is to be understood that in the context of any olfactory receptor described throughout this disclosure, the term "comprising" can be replaced with the term "consisting essentially of" or "consisting of. In other words, in some embodiments, the olfactory receptors described herein have a modified C-terminal domain that consists essentially of, or consists of, the amino acid sequence motifs disclosed herein.
[0514] Nucleic acid molecules
[0515] In another aspect, the present disclosure relates to nucleic acid molecules comprising a nucleotide sequence encoding any of the olfactory receptor proteins described herein. A nucleotide sequence encoding an olfactory receptor may also be referred to as a "gene" encoding an olfactory receptor or a "coding sequence" of an olfactory receptor. Nucleotide sequences encoding olfactory receptors are part of the common general knowledge and can be obtained by a person skilled in the art from well-known general and specific sequence databases, as described elsewhere herein.
[0516] The nucleic acid molecules of the present disclosure do not encode wild-type olfactory receptors, but rather encode olfactory receptors having modified C-terminal domains, as described in detail in the previous sections. Therefore, the nucleic acid molecules of the present disclosure are non-naturally occurring. Therefore, the nucleic acid molecules described herein can be similarly characterized as "modified" nucleic acid molecules, "engineered" nucleic acid molecules, "hybrid" nucleic acid molecules, "chimeric" nucleic acid molecules, "non-natural" nucleic acid molecules, or similar expressions and combinations thereof.
[0517] Exemplary nucleic acid molecules of the present disclosure are provided as SEQ ID NOs: 331-739 and SEQ ID NOs: 742-819. Thus, in some embodiments, the olfactory receptors described herein are encoded by a nucleic acid molecule comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 331-739 and SEQ ID NOs: 742-819. The sequences of the group consisting of NOs:742-819 have nucleotide sequences having at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity.
[0518] SEQ ID NOs 331-739 represent DNA encoding modified receptors comprising an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 221). SEQ ID NOs 742-819 represent DNA encoding modified receptors comprising an N-terminal mmLucy-FLAG-rho tag (SEQ ID NO: 81) and a modified C-terminus (SEQ ID NO: 741). As explained in the present disclosure, the presence of an N-terminal tag is optional, and different N-terminal tags can be used, as described elsewhere herein. Similarly, any modified C-terminal domain disclosed herein can be used in place of the modified C-terminus of SEQ ID NO: 221 or SEQ ID NO: 741. SEQ ID NOs 331-739 and SEQ ID NOs 742-819 also contain a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC) for cloning and expression purposes, and a 3' NotI restriction site (GCGGCCGC). The presence of these sequences is entirely optional.
[0519] The nucleic acid molecules of the present disclosure may contain other sequence elements. Typically, the other sequence elements may be sequence elements commonly used to assist in the expression of nucleotide sequences, such as promoters, nuclear localization signals, kozak sequences, polyA tails, transcription terminators, and the like. When one or more such other sequence elements are present, the nucleic acid molecules described herein may also be referred to as "nucleic acid constructs" or "gene constructs." It should be understood that different sequence elements may be "operably linked" to each other to form functional nucleic acid molecules. A description of "operably linked" is provided elsewhere in the section entitled "General Information" herein.
[0520] As used herein, "nucleic acid construct" refers to a DNA molecule comprising a region (coding region or open reading frame), which is transcribed into an RNA molecule (e.g., an mRNA molecule) in a cell and operably linked to a suitable regulatory region (such as, but not limited to, a promoter and / or enhancer sequence). Nucleic acid constructs typically comprise multiple operably linked fragments, such as, a promoter, an enhancer, a 5' leader sequence, a coding region, and / or a 3' untranslated region (3'-end), for example, comprising polyadenylation and / or transcription termination sites. Nucleic acid constructs can be recombinant, i.e., generally not present in nature, such as, a nucleic acid construct in which a promoter is not associated with part or all of the coding region in nature. Molecular toolbox techniques for preparing nucleic acid constructs are well known in the art and are discussed in standard manuals such as Ausubel et al., Current Protocols in Molecular Biology, 3rd Edition (2003), John Wiley & Sons Inc, and Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012), Cold Spring Harbor Laboratory Press; both of which are incorporated herein by reference in their entirety. Non-limiting examples of such techniques include fusion PCR, restriction digestion, Golden-gate cloning, and the like, some of which are demonstrated in the experimental section herein.
[0521] In some embodiments, the nucleic acid molecules described herein further comprise a promoter sequence. In other words, the present disclosure encompasses nucleic acid molecules described herein, wherein the nucleotide sequence is operably linked to a promoter sequence. In some embodiments, the promoter sequence described herein is a constitutive promoter sequence.
[0522] As used herein, the term "promoter" or "transcriptional regulatory sequence" refers to a nucleic acid sequence used to control the transcription (i.e., expression) of one or more coding sequences, which is located upstream in the transcription direction of the coding sequence transcription start site and is structurally recognized by the presence of a DNA-dependent RNA polymerase binding site, a transcription start site, and any other DNA sequence (including but not limited to transcription factor binding sites, repressor and activator protein binding sites, and any other nucleotide sequence known to those skilled in the art that directly or indirectly regulates the amount of promoter transcription).
[0523] In some embodiments, the promoter sequence described herein, i.e., the constitutive promoter sequence described herein, is a CMV promoter. In some embodiments, the CMV promoter can have the nucleotide sequence of SEQ ID NO: 82, or a sequence having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 82.
[0524] In some embodiments, the nucleic acid molecules described herein further comprise an enhancer sequence.
[0525] As used herein, the term "enhancer" refers to a nucleic acid sequence that can stimulate transcription of a sequence to which it is operably linked. An operably linked enhancer need not necessarily be adjacent to the coding sequence whose transcription it controls. An enhancer can be used as a single sequence or can be included in a fusion nucleotide sequence with other enhancers and / or promoters described herein.
[0526] In some embodiments, the nucleic acid molecules described herein further comprise a terminator sequence. In other words, the present disclosure encompasses nucleic acid molecules described herein, wherein the nucleotide sequence is operably linked to a terminator sequence. "Terminator sequence" may alternatively be referred to herein as "transcription terminator," "transcription terminator sequence," or simply "terminator." In some embodiments, the terminator sequence is a bovine growth hormone (BGH) terminator sequence. In some embodiments, the bgh terminator sequence can have the nucleotide sequence of SEQ ID NO:83, or a sequence at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:83.
[0527] In some embodiments, the nucleic acid molecules described herein further comprise a nucleotide sequence encoding an N-terminal signal peptide. Suitable N-terminal signal peptides have been discussed previously herein. In some embodiments, the N-terminal signal peptide is a leucine-rich signal peptide, preferably a human Lucy tag or a mouse Lucy tag, more preferably a human Lucy tag or a mouse Lucy tag represented by the amino acid sequence MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77). Sequences of SEQ ID NOs: 76 and 77 are also contemplated, in which 1, 2, 3, 4 or a maximum of 5 amino acids are replaced, deleted, added or inserted. Substitutions, particularly conservative substitutions, are preferred.
[0528] In some embodiments, the nucleic acid molecule described herein further comprises a nucleotide sequence encoding an N-terminal tag peptide. Suitable N-terminal tag peptides are discussed previously herein.
[0529] In some embodiments, the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag (the 45 N-terminal amino acids of the somatostatin 3 receptor, an example is SEQ ID NO: 223), an M3 tag (the 61 N-terminal amino acids of the muscarinic acetylcholine receptor M3; an example is SEQ ID NO: 222), a c-myc tag, and an HA tag, preferably selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, even more preferably selected from the group consisting of a FLAG tag and a rhodopsin (rho) tag, and most preferably a rhodopsin (rho) tag.
[0530] Therefore, in some embodiments, the nucleic acid molecules described herein further comprise a nucleotide sequence encoding an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an IL-6 tag, an SST3 tag, an M3 tag, a c-myc tag and an HA tag, more preferably selected from the group consisting of a FLAG tag, a rhodopsin (rho) tag, an SST3 tag and an M3 tag, even more preferably selected from the group consisting of a FLAG tag and a rhodopsin (rho) tag, and most preferably a rhodopsin (rho) tag.
[0531] In a preferred embodiment, the N-terminal tag peptide comprises at least one tag peptide selected from the group consisting of a rhodopsin (rho) tag, an SST3 tag, and an M3 tag, preferably a rho tag. Optionally, a FLAG peptide may also be present.
[0532] Combinations of the above-mentioned tag peptides can also be used. Typically, such combinations will include at least one of a rho tag, an SST3 tag, and an M3 tag, preferably at least a rho tag. For example, in some embodiments, the N-terminal tag peptide is a combination of a rho tag and a FLAG tag. In other words, in some embodiments, the nucleic acid molecules described herein further comprise a nucleotide sequence encoding an N-terminal tag peptide, wherein the N-terminal tag peptide comprises a rho tag and a FLAG tag. Preferably, in this case, the FLAG tag is located at the N-terminus of the rho tag, for example as shown in SEQ ID NO:81.
[0533] In some embodiments, the FLAG tag encoded by the nucleic acid molecules described herein has the sequence of SEQ ID NO: 78. In some embodiments, the Rho tag encoded by the nucleic acid molecules described herein has the sequence of SEQ ID NO: 79. In some embodiments, the IL-6 tag encoded by the nucleic acid molecules described herein is an IL-6-tag as described in Noe et al. In some embodiments, the nucleic acid molecules described herein encode an SST3 tag having the sequence of SEQ ID NO: 223. In some embodiments, the nucleic acid molecules described herein encode an M3 tag having the sequence of SEQ ID NO: 222. Also contemplated are sequences of SEQ ID NOs: 78, 79, 222, and 223, wherein 1, 2, 3, 4, or up to 5 amino acids are substituted, deleted, added, or inserted. Substitutions, particularly conservative substitutions, are preferred.
[0534] The nucleotide sequences encoding the N-terminal signal peptide described herein and the N-terminal tag peptide described herein can be advantageously used in combination with each other. Typically, the nucleotide sequence encoding the N-terminal signal peptide will be located upstream (N-terminus) of the nucleotide sequence encoding the N-terminal tag. For example, the nucleic acid molecules described herein may further comprise a nucleotide sequence encoding an N-terminal signal peptide and one or more N-terminal tag peptides. In some cases, the nucleic acid molecules described herein may further comprise:
[0535] - a nucleotide sequence encoding a human or mouse Lucy signal peptide (such as SEQ ID NO: 76 or 77);
[0536] - a FLAG tag (such as SEQ ID NO: 78) or an IL-6 tag, preferably a FLAG tag (such as SEQ ID NO: 78); and
[0537] - a rho tag (such as SEQ ID NO: 79), an SST3 tag or an M3 tag, preferably a rho tag (such as SEQ ID NO: 79).
[0538] SEQ ID NO: 80 is an example of a nucleotide sequence encoding a combination of mouse Lucy signal peptide, FLAG tag peptide, and rho tag peptide (SEQ ID NO: 81).
[0539] In some embodiments, the nucleic acid molecules described herein are modified to contain a nucleotide sequence encoding one or more additional N-terminal glycosylation sites.
[0540] In some embodiments, the nucleic acid molecules described herein further comprise a nucleotide sequence encoding one or more olfactory receptor accessory proteins. Olfactory receptor "accessory proteins" or "molecular chaperones" are proteins or peptides that assist in the expression, trafficking, and / or signaling of olfactory receptors to the surface of cells expressing the olfactory receptors.
[0541] Non-limiting examples of accessory proteins encompassed by the present disclosure include RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf, Giα or its functional variants, etc., and are further described in WO2006 / 002161 and WO2014 / 037800, the entire contents of which are incorporated herein by reference. Preferred accessory proteins are RTP1S and / or RTP2, preferably human RTP1S and / or human RTP2.
[0542] The accessory proteins described herein also encompass functional variants of their wild-type counterparts, i.e., accessory molecules that have been modified compared to the corresponding naturally occurring or wild-type sequence. In this context, RTP1S as used herein includes the RTP1SV227I variant, and RTP2 as used herein includes the RTP2L220R variant. Preferred accessory proteins are human RTP1S V227I variant (SEQ ID NO: 84) and human RTP2 L220R variant (SEQ ID NO: 85).
[0543] Thus, in some embodiments, one or more olfactory receptor "accessory proteins" described herein are selected from the group consisting of RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf , Giα and functional variants thereof, preferably selected from the group consisting of RTP1S, RTP2 and functional variants thereof. In some embodiments, the one or more olfactory receptor "accessory proteins" described herein are RTP1S V227I variant and RTP2 L220R variant. In some embodiments, the nucleotide sequence encoding one or more olfactory receptor accessory proteins comprises a nucleotide sequence encoding a polypeptide as set forth in SEQ ID NO:84 and / or 85, or a nucleotide sequence encoding a polypeptide having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity or similarity to SEQ ID NO:84 and / or 85.
[0544] Nucleotide sequence as herein described can be codon optimized, so that in host cell (preferably, eukaryotic cell, more preferably, human cell), express.Suitable host cell is also described in other parts of this paper, for example, referring to " cell " chapters and sections.As used herein, " codon optimized " refers to be used for modifying existing coding sequence or design coding sequence, for example to improve the translation of the transcript RNA molecule transcribed by this coding sequence in expression host cell or organism or improve the process that coding sequence is transcribed.Codon optimized includes but is not limited to, comprises the codon of selection coding sequence to adapt to the process of codon preference of expression host cell or organism.Codon optimized also eliminates the element (for example, terminator sequence, TATA box, splice site, ribosome entry site, tumor-necrosis factor glycoproteins and / or sequence and RNA secondary structure or instability motif) that RNA stability and / or translation have a negative impact potentially. In some embodiments, the codon-optimized sequence exhibits at least 3%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or more increase in gene expression, transcription, RNA stability and / or translation compared to the original, non-codon-optimized sequence.
[0545] expression vector
[0546] The nucleic acid molecules described herein can be placed in an expression vector. Thus, in another aspect, an expression vector comprising any of the nucleic acid molecules described herein is provided.
[0547] "Expression vector", alternatively referred to herein as "vector" or "delivery vector", refers to a molecular biology tool for obtaining expression of a coding region (such as a gene) in a host cell, for example by introducing a nucleotide sequence that can achieve expression of a gene or coding sequence in a host cell compatible with the sequence. The expression vector can be stable and remain episomal in the host cell. Alternatively, the vector can be integrated into the genome of the host cell, for example, by homologous recombination, non-homologous end joining or other means. A description of a suitable "host cell" in the context of the present disclosure is provided elsewhere herein.
[0548] Suitable expression vectors can be selected from any genetic element known in the art that can promote nucleic acid transfer between cells, such as, but not limited to, plasmids, phages, transposons, cosmids, chromosomes, artificial chromosomes, viruses (such as, but not limited to retroviruses, slow viruses, etc.), virions, etc. The expression vector can also be a chemical carrier, such as, a lipoplex or naked DNA. "Naked DNA" or "naked nucleic acid" refers to a nucleic acid molecule that is not contained in a packaging device that promotes nucleic acid delivery to the target host cell cytoplasm. Naked DNA can be circular or linear (linearized DNA sequence). Optionally, naked nucleic acid can be associated with a standard device used in the art that promotes its delivery of nucleic acid to the target host cell, for example, to promote the transport of nucleic acid across the cell membrane.
[0549] Preferred expression vectors are plasmids. Suitable plasmids are known in the art and are described in standard manuals, such as the works of Ausubel et al. and Sambrook and Green (see above). Suitable plasmids can also be selected from commercially available vectors, such as the pcDNA3.1(+) series (Invitrogen, MA, USA) or the pGL4.29 series vectors (Promega, WI, USA).
[0550] cell
[0551] The nucleic acid molecules and expression vectors described herein are particularly suitable for introduction into host cells. Therefore, in another aspect, a recombinant host cell comprising the nucleic acid molecules or expression vectors described herein is provided. Preferably, the recombinant host cell described herein expresses or is capable of expressing the olfactory receptor protein described herein.
[0552] In some embodiments, the host cell of the present disclosure comprises a plurality (i.e., two or more) of nucleic acid molecules and / or expression vectors as described herein. Therefore, in this case, the recombinant host cell expresses or is capable of expressing a plurality (i.e., two or more) of olfactory receptor proteins as described herein. In the case of such host cells, the two or more olfactory receptors are preferably activated by odorants with similar odors. Odorants with similar odors are usually described with the same odor descriptor. "Odor descriptor" is a common term used by trained perfumers and evaluators to describe the specific odor sensation shared by a group of ligands or ligand mixtures. Odor descriptors such as "green fragrance" represent the smell reminiscent of freshly crushed leaves, or "floral fragrance" represents the fragrance of flowers. They can also be more specific, such as "floral rose" represents the fragrance reminiscent of the smell of roses, or "white flower fragrance" represents the smell reminiscent of ylang-ylang or jasmine, and so on.
[0553] In some embodiments, the combination of two or more olfactory receptor proteins in this context may be selected from the group consisting of:
[0554] -OR7A17 and OR7C1. These ORs are activated by molecules with the odor descriptor "woody-amber" and can be co-expressed in cells to detect woody-amber notes.
[0555] These ORs are activated by molecules with the odor descriptor "musk," and two or more olfactory receptor proteins in this group can be co-expressed in cells to detect musk notes.
[0556] These ORs are activated by molecules with the odor descriptor "fruity ester," and two or more olfactory receptor proteins in this group can be co-expressed in cells to detect fruity ester notes.
[0557] These ORs are activated by molecules with the odor descriptor "fruity-lactone," and two or more olfactory receptor proteins in this group can be co-expressed in cells to detect fruity-lactone notes.
[0558] These ORs are activated by molecules with the odor descriptor "marine," and two or more olfactory receptor proteins in this group can be co-expressed in cells to detect marine notes.
[0559] - OR10G7 and OR10D3. These ORs are activated by molecules with the odor descriptor "spicy" and can be co-expressed in cells to detect spicy notes.
[0560] "Host cell", alternatively referred to herein as "recombinant host cell", "engineered cell" or simply "cell", refers to a cell that has been engineered by the introduction of a nucleic acid molecule and / or expression vector as defined herein. Host cells may refer to isolated or cultured cells. Host cells may be "transduced cells", wherein these cells have been infected, for example, with a modified virus. As a non-limiting example, lentiviruses may be used, but other suitable viruses, such as retroviruses or other viruses, are also contemplated. The introduction of nucleic acid constructs may also be performed by non-viral methods, such as transfection. "Transfection" refers to a non-viral method of transferring DNA (or RNA) into a cell, thereby causing the expression of the transferred nucleic acid sequence. Transfection methods and protocols are well known in the art, non-limiting examples being calcium phosphate transfection, PEG transfection, and liposome or lipoplex transfection, and are discussed in standard manuals such as the works of Ausubel et al. and Sambrook and Green (see above). The illustrative section herein provides another example of a transfection method. Transfection may be transient or stable, the latter referring to a situation in which the nucleic acid construct is integrated into the genome of the cell. Thus, a host cell comprising a nucleic acid construct described herein may also be a "stably transfected cell" or a "transiently transfected cell."
[0561] The host cell can be further genetically modified, for example by introducing one or more genetic modifications, including but not limited to nucleotide mutations, substitutions, insertions and / or deletions in its genome, and / or introducing additional nucleic acid constructs. The modifications can be included in the nucleotide sequences and / or other genomic regions encoding the olfactory receptors, auxiliary molecules, and can result in functional expression or enhanced functional expression of the olfactory receptors and / or auxiliary molecules. The definition of functional expression is provided elsewhere herein.
[0562] The modification of the nucleic acid sequence can be carried out using any recombinant DNA technology known in the art, such as, for example, the techniques described in standard manuals (such as the works by Ausubel et al. and Sambrook and Green) (see above). See also Kunkel (1985) Proc. Natl. Acad. Sci. 82:488 (describing site-directed mutagenesis) and Roberts et al., (1987) Nature 328:731 734 or Wells, JA et al., (1985) Gene 34:315 (describing cassette mutagenesis).
[0563] Host cells may include epigenetic modifications in nucleic acid molecules encoding olfactory receptors, auxiliary proteins, and / or other genomic regions that may result in functional expression or enhanced functional expression of the olfactory receptors and / or auxiliary proteins. As used herein, the term "epigenetic modification" has the customary meaning commonly understood by those skilled in the art based on this disclosure. It refers to chemical modifications of DNA or histones that do not change the nucleotide sequence itself. Non-limiting examples of epigenetic modifications include nucleic acid methylation, acetylation, phosphorylation, serotonylation, citrullination, ubiquitination, sumoylation, and ribosylation.
[0564] Advantageously, in some embodiments, the recombinant host cells described herein further express one or more olfactory receptor accessory proteins described herein. To this end, additional nucleic acid molecules or expression vectors encoding one or more olfactory receptor accessory proteins can be included in the recombinant host cells. These additional nucleic acid molecules or expression vectors can be stably integrated into the chromosome, or they can be introduced for transient expression, for example, by transfection. Suitable olfactory receptor accessory proteins have been described elsewhere herein. Thus, in some embodiments, one or more olfactory receptor "accessory proteins" described herein are selected from RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf , Giα and functional variants thereof, preferably selected from the group consisting of RTP1S, RTP2 and functional variants thereof. In some embodiments, the one or more olfactory receptor "accessory proteins" described herein are RTP1S V227I variant and RTP2 L220R variant. In some embodiments, the nucleotide sequence encoding one or more olfactory receptor accessory proteins comprises a nucleotide sequence encoding a polypeptide as set forth in SEQ ID NO:84 and / or 85, or a nucleotide sequence encoding a polypeptide having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity or similarity to SEQ ID NO:84 and / or 85.
[0565] In some embodiments, the recombinant host cells described herein also express one or more reporter genes, such as a luciferase gene. Reporter genes are described in more detail elsewhere herein.
[0566] The recombinant host cells described herein can be prokaryotic cells or eukaryotic cells, preferably, they are eukaryotic cells. Suitable prokaryotic cells can be selected from bacteria and archaea. Suitable eukaryotic cells can be selected from insect, plant, yeast, fungi, algae, mammals and human cells, wherein human cells are preferred.
[0567] Suitable host cells include, but are not limited to, HEK293, HEK293T, HeLa, CHO, OP6, HeLa-S3, HEKn, HEKa, PC-3, Calul, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial cells, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-L1, 132-d5 human fetal fibroblasts, 10.1 mouse fibroblasts, 293-T, 3T3, BHK, BHK-21, BR 293, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr- / -COS-7, HL-60, LNCap, MCF-7, MCF-10A, MDCKII, SkBr3, Vero cells, primary olfactory cells, immortalized olfactory cells, immortalized taste cells, and transgenic varieties thereof, with HEK293 and HEK293T being preferred. Therefore, in some embodiments, the recombinant host cells described herein are HEK293 or HEK293T cells. Of the two, HEK293T is more preferred. Cell lines are available from various public culture collections, such as the American Type Culture Collection (VA, USA).
[0568] library
[0569] In another aspect, the present disclosure relates to libraries comprising a diverse repertoire of the olfactory receptor proteins described herein, the nucleic acid molecules described herein, the expression vectors described herein, or the recombinant host cells described herein.
[0570] In some embodiments of the libraries described herein, the diverse repertoire of olfactory receptor proteins, olfactory receptor proteins encoded by nucleic acid molecules or expression vectors, or olfactory receptor proteins expressed by recombinant host cells share the same modified C-terminal domain. In a specific but non-limiting example, they share the modified C-terminal domain represented by SEQ ID NO: 221 (particularly in the context of class II olfactory receptors) or SEQ ID NO: 741 (particularly in the context of class I olfactory receptors).
[0571] This library of olfactory receptors with identical C-terminal domains has the advantage of providing more uniform functional activity across all receptors. While most libraries of receptors with wild-type C-termini exhibit significant variation in functional expression levels, functional expression is more comparable across receptors in libraries with identical C-terminal sequences. Therefore, testing ligands using this standardized library with identical C-termini allows for finding the most sensitive receptor activated by a given ligand, which is crucial for subsequent screening of novel ligands within a given odor profile. On the other hand, if target receptors are identified from libraries with widely varying functional expression, receptor selection may be biased by arbitrary functional expression in the in vitro system, rather than true ligand affinity.
[0572] The libraries described herein are not particularly limited in the number of different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. However, in some embodiments, the libraries described herein comprise a diverse repertoire of at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, or at least 400 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins.
[0573] In some embodiments, the libraries described herein comprise a diverse repertoire of about 250 to about 800 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. Such library size allows coverage of most canine olfactory receptors.
[0574] In some embodiments, the libraries described herein comprise a diverse repertoire of about 250 to about 670 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. Such library size allows coverage of most cat olfactory receptors.
[0575] In some embodiments, the libraries described herein comprise a diverse repertoire of about 250 to about 500 or 250 to about 400 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. Such library size allows coverage of most human olfactory receptors.
[0576] In some embodiments, the libraries described herein comprise a diverse repertoire of about 400 to about 800 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. Such library size allows coverage of most human olfactory receptors, including major alternative alleles or haplotypes.
[0577] The library may comprise both Class I and Class II ORs, or the library may focus on Class I ORs or Class II ORs. Thus, in some embodiments, the libraries described herein comprise different Class II olfactory receptor proteins, preferably human Class II olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins. The Class II OR libraries described herein may comprise a diverse repertoire of at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350 or at least 400 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. In the case of a human Class II OR library, the preferred number is 400 to 450. Preferably, such libraries comprise at least one, preferably all, Class II olfactory receptors listed as "receptors" in Table 15. It should be understood that in such libraries, the N-terminal tag of the Class II olfactory receptor or its modified C-terminal domain may be different from the specific SEQ ID NOs in Table 15, for example, the receptor may comprise a different modified C-terminal domain as described herein, rather than the modified C-terminal domain comprised by the olfactory receptor encoded by the nucleic acid molecules listed in Table 15. SEQ ID NOs 331-739 also contain a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC), as well as a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes. The presence of these sequences is entirely optional.
[0578] In some embodiments, the libraries described herein comprise different Class I olfactory receptor proteins, preferably human Class I olfactory receptor proteins, nucleic acid molecules or expression vectors encoding said olfactory receptor proteins, or recombinant host cells expressing said olfactory receptor proteins. The Class I OR libraries described herein may comprise a diverse repertoire of at least 10, at least 25, at least 50, or at least 75 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins. In the case of human Class II OR libraries, the preferred number is 70 to 100. Preferably, such libraries comprise at least one, preferably all, Class I olfactory receptors listed as "receptors" in Table 22. It should be understood that in such libraries, the N-terminal tag or modified C-terminal domain of the Class I olfactory receptor may differ from the SEQ ID NOs specified in Table 22, for example, the receptor may comprise a different modified C-terminal domain described herein, rather than the modified C-terminal domain contained in the olfactory receptor encoded by the nucleic acid molecules listed in Table 22. SEQ ID NOs 742-819 also contain a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC), as well as a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes. The presence of these sequences is entirely optional.
[0579] In some embodiments, the libraries described herein comprise different Class I and Class II olfactory receptor proteins, preferably human Class I and Class II olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins. In this case, the library may preferably comprise a diverse repertoire of 470 to 550 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins.
[0580] Methods and uses
[0581] The olfactory receptors, nucleic acid molecules, recombinant host cells, and libraries described herein enable functional expression of olfactory receptors that cannot be expressed using conventional methods or that result in limited assay sensitivity using conventional methods, and can be used to identify new cognate receptor-ligand pairs. Therefore, they are particularly useful for methods and uses of expressing olfactory receptors, as well as for identifying new olfactory receptors and new olfactory receptor ligands, enhancers, and antagonists.
[0582] In one aspect, provided is a use of the olfactory receptor protein described herein, the nucleic acid molecule described herein, the expression vector described herein, the cell described herein, or the library described herein for identifying an olfactory receptor ligand, enhancer, or antagonist.
[0583] In one aspect, there is provided the use of a library as described herein for identifying an olfactory receptor capable of binding a target ligand.
[0584] In one aspect, a method of identifying an olfactory receptor ligand is provided, the method comprising:
[0585] a) providing an olfactory receptor protein as described herein or a cell expressing an olfactory receptor protein as described herein;
[0586] b) contacting the receptor or cell with a test compound or composition; and
[0587] c) Detecting activation of olfactory receptors.
[0588] In one aspect, a method of identifying an olfactory receptor enhancer or antagonist is provided, the method comprising:
[0589] a) providing an olfactory receptor protein as described herein or a cell expressing an olfactory receptor protein as described herein;
[0590] b) contacting the receptor or cell with a cognate ligand and a test compound or composition; and
[0591] c) Detecting increased or decreased activation of the olfactory receptor compared to a ligand-only control.
[0592] As used herein, an olfactory receptor "antagonist" is a compound that reduces activation of a given olfactory receptor by an OR ligand. As used herein, an olfactory receptor "enhancer" is a compound that enhances activation of a given olfactory receptor by an OR ligand.
[0593] In some embodiments of the method for identifying an olfactory receptor ligand and the method for identifying an olfactory receptor enhancer or antagonist, the olfactory receptor is selected from the group consisting of: OR7C1, OR8K3 (preferably OR8K3 (L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2 (W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2 (V259L)), OR10G7 (preferably OR10G7 (T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2 (Y28C)), OR7A5, OR7E24 (or OR7E24 (P242S)), OR7A10, OR10H2, OR10H1, OR10D3, OR1D2, OR2A5, OR2A25 (OR2A25(S75N, A209P)), OR11G2 (or OR11G2(I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3(S73G)), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2 and OR2AG2 (preferably OR2AG2(Y28C)). For these receptors, identifying olfactory receptor ligands and identifying olfactory receptor enhancers are particularly important. In some embodiments of the method for identifying an olfactory receptor ligand and the method for identifying an olfactory receptor enhancer or antagonist, the olfactory receptor is OR5A2, OR5A1, OR7A17, OR7C1, OR8K3, OR1N2, OR10J5, OR5B12 or OR10H5.
[0594] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is used to identify an olfactory receptor antagonist, and the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1, and OR4S2, preferably selected from the group consisting of OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), and OR5V1. More preferably, OR2M2 and OR2V1 are selected. Of the two, OR2M2 is more preferred.
[0595] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is used to identify an olfactory receptor antagonist, and the olfactory receptor is OR2M2 or OR2V1, preferably OR2M2. In this case, optionally, step b) further comprises contacting the receptor or recombinant host cell with a copper salt. Therefore, in some embodiments, the present disclosure provides a method for identifying an olfactory receptor antagonist, the method comprising:
[0596] a) providing the OR2M2 or OR2V1 olfactory receptor protein described herein or a cell expressing the OR2M2 or OR2V1 olfactory receptor protein described herein;
[0597] b) contacting the receptor or cell with a cognate ligand, a test compound or composition, and a copper salt; and
[0598] c) Detecting increased or decreased activation of the olfactory receptor compared to a ligand-only control.
[0599] In some embodiments, copper salts may be used at a concentration between 1 and 100 μM, preferably between 10 and 100 μM, for example 30 μM. Suitable copper salts include copper (II) salts such as CuCl 2 , CuSO 4 , Cu(OH) 2 and copper acetate.
[0600] In the case of such methods for identifying OR2M2 or OR2V1 antagonists, the cognate ligand is preferably selected from the group consisting of 3-methyl-3-thio-hexanol, 2-mercapto-2-methyl-pentanol and 4-methoxy-2-methylpentane-2-thiol.
[0601] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR51B2, preferably OR51B2 (C120R, L134F, C209S). In the case of such a method for identifying an OR51B2 antagonist, the cognate ligand is preferably 3-methyl-2-hexenoic acid.
[0602] In some embodiments of the method for identifying an olfactory receptor enhancer or antagonist, the method is for identifying an olfactory receptor antagonist, and the olfactory receptor is OR5V1. In the context of such a method for identifying an OR5V1 antagonist, the cognate ligand is preferably 2,4,6-trichloroanisole.
[0603] As explained above, the present disclosure encompasses libraries comprising a diverse repertoire of olfactory receptor proteins, nucleic acid molecules or expression vectors expressing olfactory receptor proteins, and recombinant host cells expressing olfactory receptor proteins, preferably wherein each olfactory receptor protein shares the same modified C-terminal domain. Advantageously, such libraries of olfactory receptors having the same C-terminal domain provide uniform and consistent functional expression of all receptors when assaying for activation of the olfactory receptors.
[0604] Thus, in one aspect, there is provided a method for generating an objective representation of the olfactory properties of a test compound or composition, the method comprising:
[0605] a) providing a library as described herein;
[0606] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0607] c) contacting the diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a test compound or composition; and
[0608] d) Detection of activation of each olfactory receptor protein.
[0609] Also contemplated is the use of the libraries described herein to generate an objective representation of the olfactory properties of a test compound or composition.
[0610] This objective representation of olfactory properties can be called an OR activation "fingerprint." It is possible to represent the activation level of each olfactory receptor as an n-dimensional vector, where n is the number of different olfactory receptors included in the library. More information on how to measure and express olfactory receptor activation levels is provided elsewhere in this article.
[0611] This objective representation of olfactory properties also allows for the comparison of olfactory properties between two or more test compounds or compositions in an objective manner.
[0612] Thus, in one aspect, there is provided a method of assessing differences or similarities between two or more test compounds or compositions, the method comprising:
[0613] a) generating an objective representation of the olfactory properties of two or more test compounds or compositions, as described herein; and
[0614] b) Comparison of an objective representation of the olfactory properties of two or more test compounds or compositions.
[0615] In some embodiments, a method for assessing differences or similarities between two or more test compounds or compositions is provided, the method comprising:
[0616] a) providing a library as described herein;
[0617] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0618] c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of two or more test compounds or compositions;
[0619] d) detecting activation of the olfactory receptor protein by each of the two or more test compounds or compositions; and
[0620] e) comparing the activated olfactory receptor protein between each of the two or more test compounds or compositions.
[0621] Also contemplated are the use of the libraries described herein to assess differences or similarities between two or more test compounds or compositions.
[0622] In some embodiments, the two or more test compounds or compositions may include a first test composition and a second test composition. In some embodiments, the second composition lacks one or more compounds present in the first composition, but otherwise comprises the same compounds as the first composition; or the second composition comprises one or more compounds that can replace one or more compounds present in the first composition, but otherwise comprises the same compounds as the first composition.
[0623] In some embodiments, the final step of comparing the objective representation of the olfactory properties of the activated olfactory receptor proteins between each of the two or more test compounds or compositions includes calculating a distance metric. Suitable distance metrics are known to those skilled in the art. As an example, the activation level of each olfactory receptor can be represented as an n-dimensional vector, and the distance metric can be a measure of the distance between the two vectors. In some embodiments, the distance between the two vectors can be based on the so-called 1-norm or L1-norm. The L1-norm is a standard metric in mathematics that is calculated as the sum of the absolute values of the vectors. The distance between the two vectors can then be calculated based on the norm of the difference between them. Therefore, the distance between the first test compound or composition and the second test compound or composition can be calculated according to the following formula:
[0624]
[0625] This distance is also known as the Euclidean distance. Further information on how to measure and express olfactory receptor activation levels is provided elsewhere in this article.
[0626] In some embodiments, the test compound or composition involved in the methods of the present disclosure is a mixture of odorants. Indeed, the increased sensitivity of the methods disclosed herein enables detection of ligands, enhancers, and antagonists from complex samples against a matrix of strong odorants. In some embodiments, the test compound or composition is a fragrance composition. In some embodiments, the test compound or composition is a composition comprising at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 odorants, such as a fragrance composition.
[0627] In some embodiments, the test compound or composition involved in the method of the present disclosure is an unpurified synthetic compound. The automated parallel synthesis of a single compound greatly increases the number of new organic compounds that can be used for biological assays. However, the disadvantage of automated parallel synthesis is that it is usually necessary to perform difficult and expensive subsequent purification steps before it can be used for biological assays. This is especially true for sensory assays because the olfactory evaluation of new spices or fragrance ingredients is very difficult in the presence of olfactory impurities (these impurities can produce odors or mask the smell of the desired molecule). Advantageously, the increased sensitivity of the methods disclosed herein enables the use of unpurified synthetic compounds. Therefore, in some embodiments, the test compound or composition involved in the method of the present disclosure comprises an unpurified synthetic molecule. In this case, the purity level of the synthetic molecule can be less than 90%, less than 85%, less than 80%, less than 75% or less than 70% (the purity percentage level is the ratio of the desired product to the comprehensive impurities, usually measured by liquid chromatography, gas chromatography or quantitative NMR).
[0628] In accordance with the above, in some embodiments, the test compound or composition involved in the methods of the present disclosure is a mixture of odorants (as described above) or an unpurified synthetic compound (as described above).
[0629] In some embodiments, the test compound or composition involved in the methods of the present disclosure is a racemic mixture of an odorant. In some embodiments, the test compound or composition involved in the methods of the present disclosure is an isolated or synthesized isomer or enantiomer of such a racemic mixture.
[0630] As explained elsewhere, olfactory receptors have also been found to be associated with various diseases. Thus, in some embodiments, the test compound or composition may comprise a candidate therapeutic agent, such as a candidate anticancer agent. In some embodiments, the test compound or composition may comprise a pharmacologically active agent.
[0631] The ligands, enhancers, and antagonists described herein are preferably biodegradable. In fact, there is growing interest in biodegradable ingredients in the field. Therefore, in some embodiments, the test compounds or compositions involved in the methods of the present disclosure are biodegradable compounds or compositions. As used herein, a compound or composition is considered biodegradable if it meets the qualification criteria of the OECD manometric respirometry method (particularly the OECD 301F method, which is well known in the art). In this method, a compound is considered to have "readily biodegradable" or the qualification level of "readily biodegradable" is reaching 60% of the theoretical oxygen demand and / or chemical oxygen demand. This qualification value must be achieved within a 10-day window within a 28-day test period. The 10-day window begins when the degree of biodegradation reaches 10% of the theoretical oxygen demand and / or chemical oxygen demand and must end before the 28th day of the test. If a positive result is obtained in the ready biodegradability test, it can be assumed that the compound will undergo rapid and ultimate biodegradation in the environment (OECD 25 Guidance for the Testing of Chemicals, Section 3, Part 1: Principles and Strategies Related to Degradation Testing of Organic Chemicals; Accepted: July 2003). An "inherently biodegradable" assessment can also be performed using OECD Method 301F, but the acceptance criteria are different. More specifically, the acceptance criteria are 60% of the theoretical oxygen demand and / or chemical oxygen demand. This acceptance value can be achieved after a 28-day test period and is often extended to 60 days. The 10-day window does not apply.
[0632] In some embodiments, the test compound or composition involved in the methods of the present disclosure is a readily biodegradable compound or composition. In some embodiments, the test compound or composition involved in the methods of the present disclosure is an inherently biodegradable compound or composition.
[0633] In one aspect, a method for identifying an olfactory receptor capable of binding a target ligand is provided, the method comprising:
[0634] a) providing a library as described herein;
[0635] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0636] c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a target ligand; and
[0637] d) Identification of olfactory receptors activated by target ligands.
[0638] All methods described herein relate to detecting olfactory receptor activation. It should be understood that detecting olfactory receptor activation can refer to measuring the activation level of an olfactory receptor. Therefore, in this disclosure, "detecting activation" can be replaced with "measuring the activation level" or similar expressions. Means and methods for detecting olfactory receptor activation are common in the art.
[0639] For example, as described in detail in the experimental section herein, detecting the activation of olfactory receptors can involve co-transfection or use of a luciferase gene operably linked to a cAMP-responsive promoter / element (Saito et al., (2004) Cell 119(5):679-691, the entire contents of which are incorporated herein by reference), which is used as a reporter gene. Activation of the olfactory receptors and subsequent increase in intracellular cAMP lead to expression of luciferase. Luciferase cleavage of luciferin leads to luminescence, which can then be detected and quantified.
[0640] For example, in order to measure and report the activation (level) of olfactory receptors, the induction fold of luciferase can be calculated. Typically, the induction fold is taken relative to a control value containing only solvent. Typically, a background control without cells but containing all other reagents is also set together, and the numerical value of the control is subtracted from all measured values to explain background luminescence. For controls containing only solvent, cells expressing OR and luciferase genes are treated with solvent only (excluding test compounds or compositions). All experimental values using test compounds and compositions are then divided by the mean value of these control measurements containing only solvent to calculate the luciferase induction fold. Therefore, solvent controls and test compounds and compositions that do not cause OR activation will obtain a value of 1, indicating that there is no luciferase induction. Numerical values significantly higher than 1 represent activation of the luciferase gene, and cAMP generation is enhanced due to OR activation.
[0641] Once a strong ligand for a given OR is known, as well as the concentration of that ligand that results in maximal OR activation, that ligand (tested at its maximally inducing concentration) can be introduced into the experiment as a positive control, and the luciferase fold can be reported as a percentage of activation of the positive control according to the following formula:
[0642] Activation % = (luciferase activation fold of test compound - 1) / (luciferase activation fold of positive control - 1) × 100
[0643] Based on this calculation, the potency of the ligands can be compared by applying a sigmoidal curve fit to the Hill equation and calculating, for example, EC20% or EC50% values (ie, the concentration that results in 20% or 50% activation of the OR compared to the positive control).
[0644] In experiments without positive controls, such as those using receptor libraries, other types of normalization can be used. Therefore, because the dynamic range (maximum efficacy) of different receptors can vary greatly, it may be appropriate to report data using the logarithm of fold induction. Other options are:
[0645] When multiple samples were tested, normalization was performed to the highest fold induction for each receptor or to the historical value of maximum efficacy for a given OR.
[0646] Other reporter genes that can be coupled to cAMP-responsive promoters / elements include green fluorescent protein. There are also various methods that directly measure changes in intracellular cAMP concentrations based on antibody binding to cAMP. Other methods for detecting OR activation include coupling hybrid G proteins to ORs, whereby the hybrid G-protein activates calcium release from intracellular reservoirs. Changes in calcium concentration are then measured using chemical fluorescent probes sensitive to changes in calcium concentration or recombinant fluorescent or luminescent proteins that can sense differences in calcium concentration.
[0647] Further means and methods for detecting olfactory receptor activation are known to those skilled in the art, including: GTPase / GTP binding assays, aequorin-based assays, fluorescence-based assays, membrane depolarization assays, melanocyte assays, PKC activation assays, PKA activation assays, kinase assays, and the like, as described, for example, in WO2019 / 110630, the entire contents of which are incorporated herein by reference. Veithen et al., 2017, Springer Handbook of Odor Chapter 22.2, Springer International publishing (CH), provides an overview of some methods for detecting olfactory receptor activation by odorants in heterologous cells, the entire contents of which are incorporated herein by reference.
[0648] The ligand and / or test compound or composition as described herein can be added to the existing culture, or alternatively the culture medium of the existing culture can be replaced with a fresh culture medium comprising the ligand. Suitable ligands can be selected from any compound known in the art that can activate olfactory receptors (alternatively referred to as "aromatic compounds" or "odorants"), which are discussed in standard manuals, such as Buettner (2017), Springer Handbook of Odor, Springer International publishing (CH), the entire contents of which are incorporated herein by reference. Suitable compounds can also be found in public databases, such as "OlfactionBase", available at https: / / olfab.iiita.ac.in / olfactionbase / , and discussed in Sharma et al., OlfactionBase:arepository to explore odors, odorants, olfactory receptors and odorant-receptor interactions.Nucleic Acids Res.2022Jan 7; 50(D1):D678-D686.
[0649] Non-limiting examples of suitable ligands include esters (e.g., geranyl acetate, methyl formate, methyl acetate, methyl propionate, methyl butyrate, ethyl acetate, ethyl butyrate, isoamyl acetate, amyl butyrate, amyl valerate, octyl acetate, benzyl acetate, methyl anthranilate, hexyl acetate), linear terpenes (e.g., myrcene, geraniol, nerol, citral, citronellal, citronellol, linalool, nerolidol, ocimene), cyclic Terpenes (limonene, camphor, methanol, carvone, terpineol, α-ionone, thujone, eucalyptol, jasmine flavor), aromatic compounds (e.g., benzaldehyde, eugenol, isoeugenol, cinnamaldehyde, ethyl maltol, ethyl vanillaldehyde, anisole, estragole, thymol), amines (e.g., trimethylamine, putrescine, cadaverine, pyridine, indole, skatole), alcohols (e.g., furanone, 1-hexanol, ethanol), aldehydes (e.g., , acetaldehyde, hexanal, furfural, hexylcinnamaldehyde, isovaleraldehyde, anisaldehyde, p-isopropylbenzaldehyde), ketones (e.g., dihydrojasmone, 2-acetyl-1-pyrrolidine, 6-acetyl-2,3,4,5-tetrahydropyridine), lactones (e.g., γ-decanolide, γ-nonalactone, δ-octanolactone, jasmine lactone, massoia lactone, wine lactone, fenugreek lactone), thiols (e.g., propylthiophene, allylthiol, ethanethiol, 2-methyl-2-propanethiol, 1-butanethiol, mercaptan, methyl mercaptan, furan-2-ylmethyl mercaptan, benzyl mercaptan), musk (e.g., nitromusk, polycyclic musk, macrocyclic musk, linear / alicyclic musk, musk ketone, ambrette musk, umbelliferous musk, Tibetan musk, xylene musk), cresol (e.g., ultravanil), propenyl ethyl guaiacol (vanillin), carboxylic acid, etc.
[0650] Preferred ligands include those with musk, woody, lily-of-the-valley, floral, green, balsamic, spicy, or fruity notes. "Musk," "woody," "lily-of-the-valley," "floral," "green," "balsamic," "spicy," and "fruity" are generally accepted terms in the field of odor molecules. One skilled in the art can identify ligands with musk, woody, lily-of-the-valley, floral, green, balsamic, spicy, and / or fruity notes in public databases such as "OlfactionBase," available at https: / / olfab.iiita.ac.in / olfactionbase / and discussed in Sharma et al., OlfactionBase: a repository to explore odors, odorants, olfactory receptors and odorant-receptor interactions. Nucleic Acids Res. 2022 Jan 7; 50(D1): D678-D686, the entire contents of which are incorporated herein by reference. Related musk compounds are also described in WO2019 / 11630, the entire contents of which are incorporated herein by reference.
[0651] Further examples of preferred ligands are Ambermax, p-cresol, menthol, (S)-menthol, menthone, Mahonial, Nympheal, linalool, androstenone, androstenol, cyclopentanethiol, methyl dihydrojasmonate, methyl dihydrojasmonate HC, Ambrofix, watermelon ketone, 4-ethyloctanoic acid, galaxol, galaxol S, ethyl vanillal, β-ionone, ambrette lactone, 3-methyl-3-hydroxy-hexanoic acid, 3-methyl-2-hexenoic acid, nonanoic acid, decanoic acid, undecanoic acid, ethyl 3-mercaptopropionate, diallyl disulfide, benzothiazole, 2-methyl-3-tetrahydrofuranthiol, musk ketone, dipropyl disulfide, musk ketone, Arborone, georgywood, iso E super, cedrol, p-menthane-8-thiol-3-one, 2-naphthalenethiol, 3-(methylthio)propanal, 1,3-propanedithiol, benzothiazole, allyl sulfide, allyl mercaptan, hydroxyethylmethylthiazole, dimethyl trisulfide, thioglycolic acid, 3-mercapto-2-pentanone, 2-(methyldisulfide)methylfuran, di(methylthio)methane, ethyl 2-mercaptopropionate, methyl thiobutyrate, 3-mercapto-3 -Methylbutyl formate, methyl 3-mercaptopropionate, butyl 3-mercaptopropionate, 3-mercaptopropionic acid, dimethyl disulfide, 2-mercaptopropionic acid, 2-methyl-3-furanthiol, benzyl mercaptan, 2-mercapto-2-methyl-1-pentanol, allyl isothiocyanate, 2-mercaptobutanone, 2-heptyl mercaptan, 2-methyl-3-tetrahydrofuranthiol, cis-2-isobutyl-4,5-dimethyl-2,5-dihydrothiazole, 3-Methyl-3-sulfanylhexan-1-ol, (racemic)-3-mercapto-2-methyl-1-pentanol, 1-hexanethiol, cyclopentanethiol, 2-methyl-2-propanethiol, 2-methyl-3-(methyldithio)furan, 2-methyl-3-buten-1-ol, sodium methylthiolate, sodium hydrosulfide hydrate, 4-methoxy-2-methylpentane-2-thiol, blackcurrant dy), furfurylthiol, anjeruk, 3-mercaptohexyl acetate, methoxymethylbutylthiol, 3-mercaptohexanol, dimethyl sulfide, acetylthiazole, 1-p-menene-8-thiol, 2-methyl-3-thiopentanol, (E,S)-3,7-dimethylnon-6-en-1-ol, 3-mercapto-3-methylhexan-1-ol, 7-(3-methylbutyl)-phenylpropyl[b][1,4]dioxepin -3-Ketone, benzyl salicylate, δ-damascone, ethyl cyclohexanecarboxylate, geosmin, 3-(4-isobutyl-2-methylphenyl)propionaldehyde, patchouli alcohol, peony nitrile, cyperone, heliotropin, methyl salicylate, geranone, Java sandalwood, wood alcohol, trans-2, cis-6-nonadienal, esterly, cascalone, azurone, rosabloom, rosyfolia, β-ionone, fenugreek lactone, 3-methyl-3-hydroxyhexanoic acid, 2-methylundecanoic acid.
[0652] Those skilled in the art will appreciate that the amount of ligand required to activate an olfactory receptor may vary depending on the olfactory receptor and the ability of the ligand to physically associate with the olfactory receptor. 50 Value (usually EC 50 If a ligand (e.g., a ligand with an EC value between 1 nM and 1 mM) physically associates with (i.e., binds to) a given olfactory receptor, then the ligand is considered to be a "ligand for" (i.e., specific for) that receptor. In the context of a ligand for an olfactory receptor, the EC 50 It refers to the concentration of ligand that results in 50% of the maximum activation of a given olfactory receptor for that receptor, which can be measured using the methods described elsewhere herein.
[0653] Various aspects and implementation plans
[0654] Various aspects and embodiments of the present invention are set out in the following numbered paragraphs, which form an integral part of this specification.Each feature, aspect and embodiment described below has been further described in the above description.
[0655] 1. An olfactory receptor protein, wherein the protein has a modified C-terminal domain, wherein at least 32% of the amino acids in the C-terminal domain are lysine, arginine or histidine.
[0656] 2. An olfactory receptor protein, wherein the protein has a modified C-terminal domain, wherein at least 35% of the amino acids in the C-terminal domain are lysine, arginine or histidine.
[0657] 3. The olfactory receptor protein of paragraph 1 or 2, wherein the protein has a modified C-terminal domain in which at least 38% of the amino acids are lysine, arginine, or histidine.
[0658] 4. The olfactory receptor protein of any preceding paragraph, wherein the protein has a modified C-terminal domain comprising the amino acid sequence motif RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1).
[0659] 5. The olfactory receptor protein according to any of the preceding paragraphs, wherein the modified C-terminal domain is fused to the seventh transmembrane helix of the protein.
[0660] 6. The olfactory receptor protein according to any of the preceding paragraphs, wherein the protein is a class I or class II olfactory receptor with a modified C-terminal domain, preferably a class I or class II olfactory receptor of humans, dogs or cats with a modified C-terminal domain, more preferably a class I or class II olfactory receptor of humans with a modified C-terminal domain.
[0661] 7. The olfactory receptor protein of any of the preceding paragraphs, wherein the protein is a human class II or class I olfactory receptor selected from the group consisting of: OR10A2, OR10A3, OR10A4, OR10A5, OR10A6, OR10A7, OR10AD1, OR10AG1, OR10C1, OR10D3, OR10G2, OR10G3, OR10G4, OR10G6, OR10G7, OR10G8, OR10G9, OR10H1, OR10H2, OR10H3, OR10H4, OR10H5, OR10J1, OR10J3, OR10J5, OR10K1, OR10K2, OR10P 1,OR10P2,OR10Q1,OR10R2,OR10S1,OR10T2,OR10V1,OR10W1,OR10X1,OR10Z1,OR11A1,OR11G2,OR11H1,OR11H2,OR11H4,OR11H6,OR11L1,OR12D2,OR1 2D3,OR13C2,OR13C3,OR13C4,OR13C5,OR13C8,OR13C9,OR13D1,OR13F1,OR13G1,OR13H1,OR13J1,OR14A16,OR14A2,OR14C36,OR14I1,OR14L1P,OR1A1, OR1A2,OR1B1,OR1C1,OR1D2,OR1D4,OR1D5,OR1E1,OR1E2,OR1F1,OR1F12,OR1G1,OR1I1,OR1J1,OR1J2,OR1J4,OR1K1,OR1L1,OR1L3,OR1L4,OR1L6,OR1 L8,OR1M1,OR1N1,OR1N2,OR1Q1,OR1S1,OR1S2,OR2A1,OR2A12,OR2A14,OR2A2,OR2A25,OR2A4,OR2A42,OR2A5,OR2A7,OR2A9P,ORAE1,OR2AG1,OR2AG2,O R2AJ1,OR2AK2,OR2AP1,OR2AT4,OR2B11,OR2B2,OR2B3,OR2B6,OR2B8P,OR2C1,OR2C3,OR2D2,OR2D3,OR2F1,OR2F2,OR2G2,OR2G3,OR2G6,OR2H1,OR2H2, OR2J1P,OR2J2,OR2J3,OR2K2,OR2L13,OR2L2,OR2L3,OR2L5,OR2L8,OR2M2,OR2M3,OR2M4,OR2M5,OR2M7,OR2S2,OR2T1,OR2T10,OR2T11,OR2T12,OR2T2,OR2T27,OR2T29,OR2T3,OR2T33,OR2T34,OR2T35,OR2T4OR2T5,OR2T6,OR2T7,OR2T8,OR2V1,OR2V2,OR2W1,OR2W3,OR2W5,OR2Y1,OR2Z1,OR3A1,OR3A2,OR3A3,OR3A4,OR4A15,OR4A16,OR4A4,R4A47,OR4A5,OR4B1,OR4C11,OR4C12,OR4C13,OR4C15,OR4C16,OR4C,OR4C45,OR4C46,OR4C5,OR4C6,OR4D1,OR4D10,OR4D11,OR4D2,OR4D5,OR4D6,OR4D9,OR4E2,OR4F15,OR4F16,OR4F17,OR4F21,OR4F29,OR4F3,OR4F4,OR4F5,OR4F6,OR4K1,OR4K13,OR4K14,OR4K15,OR4K17,OR4K2,OR4K3P,OR4K5,OR4L1,OR4M1,OR4M2,OR4N2,OR4N4,OR4N5,OR4P4,OR4Q3,OR4S1,OR4S2,OR4X1,OR4X2,OR51A2,OR51A4,OR51A7,OR51B2,OR51B4,OR51B5,OR51B6,OR51D1,OR51E1,OR51E2,OR51F1,OR51F2,OR51G1,OR51G2,OR51H1P,OR51I1,OR51I2,OR51J1,OR51L1,OR51M1,OR51Q1,OR51S1,OR51T1,OR51V1,OR52A1,OR52A4,OR52A5,OR52B2,OR52B4,OR52B6,OR52D1,OR52E2,OR52E4,OR52E5,OR52E6,OR52E8,OR52H1,OR52I1,OR52I2,OR52J3,OR52K1,OR52K2,OR52L1,OR52M1,OR52N1,OR52N2,OR52N4,OR52N5,OR52P1P,OR52R1,OR52W1,OR56A1,OR56A3,OR56A4,OR56A5,OR56B1,OR56B4,OR5A1,OR5A2,OR5AC2,OR5AK2,OR5AL1P,OR5AN1,OR5AP2,OR5AR1,OR5AS1,OR5AU1,OR5B12,OR5B17,OR5B2,OR5B21,OR5B3,OR5C1,OR5D13,OR5D14,OR5D16,OR5D18,OR5F1,OR5H1,OR5H14,OR5H15,OR5H2,OR5H6,OR5I1,OR5J2,OR5K1,OR5K2,OR5 K3,OR5K4,OR5L1,OR5L2,OR5M1,OR5M10,OR5M11,OR5M3,OR5M8,OR5M9,OR5P2,OR5P3,OR5R1,OR5 T1,OR5T2,OR5T3,OR5V1,OR5W2,OR6A2,OR6B1,OR6B2,OR6B3,OR6C1,OR6C2,OR6C3,OR6C4,OR6C6 ,OR6C65,OR6C68,OR6C70,OR6C74,OR6C75,OR6C76,OR6F1,OR6J1,OR6K2,OR6K3,OR6K6,OR6M1,O R6N1,OR6N2,OR6P1,OR6Q1,OR6S1,OR6T1,OR6X1,OR6Y1,OR7A10,OR7A17,OR7A5,OR7C1,OR7C2,O R7D2,OR7D4,OR7E24,OR7G1,OR7G2,OR7G3,OR8A1,OR8B12,OR8B2,OR8B3,OR8B4,OR8B8,OR8D1,O R8D2, OR8D4, OR8G1, OR8G5, OR8H1, OR8H2, OR8H3, OR8I2, OR8J1, OR8J3, OR8K1, OR8K3, OR8K5, OR8S1, OR8U1, OR8U8, OR8U9, OR9A2, OR9A4, OR9G1, OR9G4, OR9G9, OR9I1, OR9K2, OR9Q1, OR9Q2, or variations thereof.
[0662] 8. The olfactory receptor protein according to any of the preceding paragraphs, wherein the class II receptor is selected from the group consisting of OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10D3, OR10H1, OR10D3. D2, OR2A5, OR2A25, OR11G2, OR14J1, OR5M3, OR8D1, OR10G3, OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2, OR10A3, OR10A6, OR10J1 and OR2J2, preferably wherein said class II receptor is selected from the group consisting of OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2 and OR5A2.
[0663] 9. The olfactory receptor protein according to any one of paragraphs 1-7, wherein the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR52K1, OR56A1, OR51B5, OR56A3 and OR51L1.
[0664] 10. The olfactory receptor protein of any of paragraphs 4-9, wherein the sequence motif is RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 5).
[0665] 11. The olfactory receptor protein of any one of paragraphs 4-10, wherein the sequence motif is selected from the group consisting of:
[0666] RNX1X2X3X4X 5”” AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO:311),
[0667] RNX1EX 3' X4X5”” AX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO:312),
[0668] RNKEVKX 5”” ALKRLLKRK (SEQ ID NO: 319),
[0669] RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO:820),
[0670] RNX1EX 3' X4KAX6X 7' X8LX 10 X 11' X 12 X 13 (SEQ ID NO: 821), and;
[0671] RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 (SEQ ID NO: 824), wherein:
[0672] -X1 is K or R;
[0673] -X2 is E or D or Q;
[0674] -X3 is V or M or I or L;
[0675] -X 3' It is V or M or I;
[0676] -X4 is K or R;
[0677] -X5 is any amino acid;
[0678] -X 5”” is D or K or E or N or R or V or A or Q or G or C or F or H or I or L or M or S or T or W or Y;
[0679] -X6 is L, I or V;
[0680] -X7 is K or R or H;
[0681] -X 7' It is K or R;
[0682] -X8 is K or R;
[0683] -X 10 It is L or I or F;
[0684] -X 11 It is K or R or G;
[0685] -X 11' It is K or R;
[0686] -X 12 is K or R; and
[0687] -X 13 It is K or R.
[0688] 12. The olfactory receptor protein of any of paragraphs 4-11, wherein x or X5 is not proline.
[0689] 13. The olfactory receptor protein of any of paragraphs 4-11, wherein x is not proline or tryptophan.
[0690] 14. The olfactory receptor protein according to any one of 4-13, wherein x, X5 or X 5”” is selected from D, K, R, E, N, V, A, Q or G, preferably wherein x, X5 or X 5”” Selected from D, K, R, E, N, V, A or Q.
[0691] 15. The olfactory receptor protein of any one of paragraphs 4-14, wherein the amino acid sequence motif comprises 1 to 6 additional C-terminal amino acid residues, optionally wherein:
[0692] - the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, more preferably the first additional amino acid residue is C, R or K, most preferably C;
[0693] - the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably C or R;
[0694] - the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K;
[0695] - the fourth, fifth and sixth additional amino acid residues are selected from K and R.
[0696] 16. The olfactory receptor protein of paragraph 15, wherein the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of:
[0697] -CC,SI,YP,PQ,FR,CR,RR,EK,PR,CG,FK,RG,RC,RT,RF,GG,YR,GC,TG,PC,HP,PG,KY,CP,YY,FF,CF,NP,YL,IC,HC,CL,YC,ER,RP,PA,FC,RY;
[0698] -CRR,CCC,CCF,CCL,CCM,CCS,CCP,CCA,CCY,CCH,CCN,CCD,CCK,CCR,CCG;
[0699] -CRRR(SEQ ID NO:159),CRKK(SEQ ID NO:160),CCRR(SEQ ID NO:161),CCRK(SEQ ID NO:228),CCKR(SEQ ID NO:229),
[0700] -CRRRR(SEQ ID NO:162),CCRRR(SEQ ID NO:163).CCKRR(SEQ ID NO:230),CCRKR(SEQ ID NO:231),CCRRK(SEQ ID NO:232),CCRKK(SEQ ID NO:233),CCKRK(SEQ ID NO:234),CCKKR(SEQ ID NO:235),CCKKK(SEQ ID NO:236),
[0701] -CRRRRR(SEQ ID NO:164), CRRRKK(SEQ ID NO:165) and CCRRRR(SEQ ID NO:224),
[0702] Preferably, wherein the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164) and CRRRKK (SEQ ID NO: 165).
[0703] 17. The olfactory receptor protein of paragraph 15 or 16, wherein the sequence motif is selected from the group consisting of:
[0704] RNKEVKX 5”” ALKRLLKRKCC (SEQ ID NO: 320),
[0705] RNX1X2X3X4KAX6X7X8JX 10 X 11 X 12 X 13 CC (SEQ ID NO: 822),
[0706] RNX1EX 3’ X4KAX6X 7’ X8LX 10 X 11’ X 12 X 13 CC (SEQ ID NO: 823),
[0707] RNX1X2X3X4X5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 825), and
[0708] RNX1QIRX5AX6X7X8JX 10 X 11 X 12 X 13 CCRRR (SEQ ID NO: 826).
[0709] 18. The olfactory receptor protein of any of paragraphs 4-17, wherein the sequence motif is selected from the group consisting of:
[0710] RNKEVKDALKRLLKRK(SEQ ID NO:10),
[0711] RNREVKDALKRLLKRK(SEQ ID NO:11),
[0712] RNKEIKDALKRLLKRK(SEQ ID NO:12),
[0713] RNKEMKDALKRLLKRK(SEQ ID NO:13),
[0714] RNKEVRDALKRLLKRK(SEQ ID NO:14),
[0715] RNKEVKKALKRLLKRK(SEQ ID NO:15),
[0716] RNKEVKEALKRLLKRK(SEQ ID NO:16),
[0717] RNKEVKNALKRLLKRK(SEQ ID NO:17),
[0718] RNKEVKRALKRLLKRK(SEQ ID NO:18),
[0719] RNKEVKVALKRLLKRK(SEQ ID NO:19),
[0720] RNKEVKAALKRLLKRK(SEQ ID NO:20),
[0721] RNKEVKQALKRLLKRK(SEQ ID NO:21),
[0722] RNKEVKDAVKRLLKRK(SEQ ID NO:22),
[0723] RNKEVKDAIKRLLKRK(SEQ ID NO:23),
[0724] RNKEVKDALRRLLKRK(SEQ ID NO:24),
[0725] RNKEVKDALKKLLKRK(SEQ ID NO:25),
[0726] RNKEVKDALKRLIKRK(SEQ ID NO:26),
[0727] RNKEVKDALKRLFKRK(SEQ ID NO:27),
[0728] RNKEVKDALKRLLRRK(SEQ ID NO:28),
[0729] RNKEVKDALKRLLKKK(SEQ ID NO:29),
[0730] RNKEVKDALKRLLKRR(SEQ ID NO:30),
[0731] RNKDVKDALKRLLKRK(SEQ ID NO:31),
[0732] RNKQVKDALKRLLKRK(SEQ ID NO:32),
[0733] RNKELKDALKRLLKRK(SEQ ID NO:33),
[0734] RNKEVKGALKRLLKRK(SEQ ID NO:34),
[0735] RNKEVKDALHRLLKRK(SEQ ID NO:35),
[0736] RNKEVKDALKRILKRK(SEQ ID NO:36),
[0737] RNKEVKDALKRLLGRK(SEQ ID NO:37),RNKEVKRAIKRLLKRK(SEQ ID NO:38),RNKEVKKAIKRLLKRK(SEQ ID NO:39),RNKEVKRAIKRLFKRK(SEQ ID NO:40),RNKEVKKAIKRLFKRK(SEQ ID NO:41),RNKEVKRAIKRLLKRK(SEQ ID NO:41),RNKEVKRAIKRLLKRK(SEQ ID NO:41),RNKEVKKAIKRLLKRK(SEQ ID NO:41),RNKEVKRAIKRLLKRK(SEQ ID NO:41), NO:42),RNKEVKDALRKLLKRK(SEQ ID NO:43),RNKEVKDALKRLLRRR(SEQ ID NO:44),RNKEVKRALKRLLRRR(SEQ ID NO:45),RNKEVKKALKRLLRRR(SEQ ID NO:46),RNREVKRAIKRLLKRK(SEQ ID NO:46),RNREVKRAIKRLLKRK(SEQ ID NO:47),RNREVKRAIKRLLKRK(SEQ ID NO:46),RNKEVKRAIKRLLKLL NO:48),RNREVKRAIKRLFKRK(SEQ ID NO:49),RNREVKKAIKRLFKRK(SEQ ID NO:50),RNREVKRAIRKLLKRK(SEQ ID NO:51),RNREVKDALRKLLKRK(SEQ ID NO:52),RNREVKDALRKLLKRK(SEQ ID NO:52),RNREVKDALRKLLKRK(SEQ ID NO:53),RNREVKRAIKRLFKRK(SEQ ID NO:3), NO:54),RNKEVKKAIKRLLKKK(SEQ ID NO:55),RNKEVKKAIKRLLKRR(SEQ ID NO:56),RNKEVKRAIKRLLRRK(SEQ ID NO:57),RNKEVKRAIKRLLKKK(SEQ ID NO:58),RNKEVKRAIKRLLKRR(SEQ ID NO:58),RNKEVKRAIKRLLKRR(SEQ ID NO:59),RNKEVKRAIKRLLKKK(SEQ ID NO:59),RNKEVKRAIKRLLKKK(SEQ ID NO:59),RNKEVKRAIKRLLKKK(SEQ ID NO:59),RNKEVKRAIKRLLKRR(SEQ ID NO:59), NO:60),RNKEVKKAIKRLFKKK(SEQ ID NO:61),RNKEVKKAIKRLFKRR(SEQ ID NO:62),RNKEVKRAIKRLFRRK(SEQ ID NO:63),RNKEVKRAIKRLFKKK(SEQ ID NO:64),RNKEVKRAIKRLFKRR(SEQ ID NO:65),RNKEVKRAIKRLFKRR(SEQ ID NO:65),RNKRAIKRLFKLR(SEQ ID NO:65), NO:66),RNREVKKAIKRLLRKK(SEQ ID NO:67),RNREVKRAIKRLFRRR(SEQ ID NO:68),<h2 style=";text-align:left;direction:ltr">RNREVKKAIKRLFRRR(SEQ ID NO:69),RNREVKKAIKRLFRRK(SEQ ID NO:70),RNREVKKAIKRLFKKK(SEQ ID NO:71),RNREVKKAIKRLFKRR(SEQ ID NO:72),RNREVKRAIKRLFRRK(SEQ ID NO:73),RNREVKRAIKRLFKKK(SEQ ID NO:74),RNREVKRAIKRLFKRR(SEQ ID NO:75),RNREMRKALHRLLGKK(SEQ ID NO:827),RNREVKKAIHKLIGRK(SEQ ID NO:828),RNREVRKAVHRLFKRK(SEQ ID NO:829),RNKEMKKAIHKLFGKK(SEQ ID NO:830),RNRDVKKAVHKLFRRK(SEQ ID NO:831),RNRDMKKAVHKLFGKR(SEQ ID NO:832),RNKELRKALHKLLGRK(SEQ ID NO:833),RNRDVRKALRRILRRR(SEQ ID NO:834),RNKDVRKAVRKLIRRR(SEQ ID NO:835),RNRDVRKAVRRLFRKR(SEQ ID NO:836),RNKDIKKAVKKLIKKK(SEQ ID NO:837),RNRELRKAVRRLFKRR(SEQ ID NO:838),RNKELRKAVRKIIKK(SEQ ID NO:839),RNRDVKKAVRRLFRRK(SEQ ID NO:840),RNREVRKALRRIIRKR(SEQ ID NO:841),RNKDIRKAVKKIFRRK(SEQ ID NO:842),RNKDVRKAVRRLIKRK(SEQ ID NO:843),RNRDLRKAVRKLFKKK(SEQ ID NO:844),RNRDLRKALRRIFKRR(SEQ ID NO:845),RNRDVRKAIKKLIRKR(SEQ ID NO:846),RNKELKKAIKRILKKK(SEQ ID NO:847),RNRDVRKAIRKLLKRK(SEQ ID NO:848),RNRDLRKAVRRIFKKR(SEQ ID NO:849),RNRDVRKAVRKLFKRR(SEQ ID NO:850),<h2 style=";text-align:left;direction:ltr">RNRDVRKALRRLFKKR(SEQ ID NO:851),RNKELKKALRKLIGKK(SEQ ID NO:852),RNREMRKAIKKIIKKK(SEQ ID NO:853),RNKEIKKAIKKIIKKR(SEQ ID NO:854),RNRDVKKAIRRLFRRR(SEQ ID NO:855),RNREVKKAVKKLIGKR(SEQ ID NO:856),RNREMRKALRRLFRKR(SEQ ID NO:857),RNKELKKALRRLIGRR(SEQ ID NO:858),RNRDVKKALRKLIGKR(SEQ ID NO:859),RNREVKKAVKKLIRRK(SEQ ID NO:860),RNKEVRKALKKLFGKK(SEQ ID NO:861),RNKEIRKALRRLFGKK(SEQ ID NO:862),RNKDVKKALRRLFGKK(SEQ ID NO:863),RNKELKKAIKRLIRRK(SEQ ID NO:864),RNKDVRKAVKRLLKKR(SEQ ID NO:865),RNKELRKAIRRLLRRR(SEQ ID NO:866),RNRDIRKALRKLFKKK(SEQ ID NO:867),RNRELKKALRRLLRRR(SEQ ID NO:868),RNREVKKALRRLFGKK(SEQ ID NO:869),RNRDVRKALKRLLKRK(SEQ ID NO:870),RNRDMRKAIRKLFGRK(SEQ ID NO:871),RNRELKKAIRKLLKRK(SEQ ID NO:872),RNRDIRKAVKKLFGKK(SEQ ID NO:873),RNKEVKKAIRKLFGRR(SEQ ID NO:874), RNREVRKAVRKLFRRK(SEQ ID NO:875), RNRDMKKALKKLFRRR(SEQ ID NO:876), RNRDVRKALKRLLGRR(SEQ ID NO:877), RNKDLKKAVKKLFGRK(SEQ ID NO:878), RNKDVRKAVRRLFGRR(SEQ ID NO:879), RNKEVKCALKRLLKRK(SEQ ID NO:880), RNKEVKFALKRLLKRK(SEQ ID NO:881),RNKEVKHALKRLLKRK(SEQ ID NO:882),RNKEVKIALKRLLKRK(SEQ ID NO:883),RNKEVKLALKRLLKRK(SEQ ID NO:884),RNKEVKMALKRLLKRK(SEQ ID NO:885),RNKEVKSALKRLLKRK(SEQ ID NO:886),RNKEVKTALKRLLKRK(SEQ ID NO:887),RNKEVKWALKRLLKRK(SEQ ID NO:888),RNKEVKYALKRLLKRK(SEQ ID NO:889),RNKQIRDALKRLLKRK(SEQ ID NO:890).RNKEVKDALKRLLKRKCC(SEQ ID NO:86),RNREVKDALKRLLKRKCC(SEQ ID NO:87),RNKEIKDALKRLLKRKCC(SEQ ID NO:88),RNKEMKDALKRLLKRKCC(SEQ ID NO:89),RNKEVRDALKRLLKRKCC(SEQ ID NO:90),RNKEVKKALKRLLKRKCC(SEQ ID NO:91),RNKEVKEALKRLLKRKCC(SEQ ID NO:92),RNKEVKNALKRLLKRKCC(SEQ ID NO:93),RNKEVKRALKRLLKRKCC(SEQ ID NO:94),RNKEVKVALKRLLKRKCC(SEQ ID NO:95),RNKEVKAALKRLLKRKCC(SEQ ID NO:96),RNKEVKQALKRLLKRKCC(SEQ ID NO:97),RNKEVKDAVKRLLKRKCC(SEQ ID NO:98),RNKEVKDAIKRLLKRKCC(SEQ ID NO:99),RNKEVKDALRRLLKRKCC(SEQ ID NO:100),RNKEVKDALKKLLKRKCC(SEQ ID NO:101),RNKEVKDALKRLIKRKCC(SEQ ID NO:102),RNKEVKDALKRLFKRKCC(SEQ ID NO:103),RNKEVKDALKRLLRRKCC(SEQ ID NO:104),RNKEVKDALKRLLKKKCC(SEQ ID NO:105),RNKEVKDALKRLLKRRCC(SEQ ID NO:106),<h2 style=";text-align:left;direction:ltr">RNKDVKDALKRLLKRKCC(SEQ ID NO:107),RNKQVKDALKRLLKRKCC(SEQ ID NO:108),RNKELKDALKRLLKRKCC(SEQ ID NO:109),RNKEVKGALKRLLKRKCC(SEQ ID NO:110),RNKEVKDALHRLLKRKCC(SEQ ID NO:111),RNKEVKDALKRILKRKCC(SEQ ID NO:112),RNKEVKDALKRLLGRKCC(SEQ ID NO:113),RNKEVKRAIKRLLKRKCC(SEQ ID NO:114),RNKEVKKAIKRLLKRKCC(SEQ ID NO:115),RNKEVKRAIKRLFKRKCC(SEQ ID NO:116),RNKEVKKAIKRLFKRKCC(SEQ ID NO:117),RNKEVKRAIRKLLKRKCC(SEQ ID NO:118),RNKEVKDALRKLLKRKCC(SEQ ID NO:119),RNKEVKDALKRLLRRRCC(SEQ ID NO:120),RNREMRKALHRLLGKKCC(SEQ ID NO:254),RNREVKKAIHKLIGRKCC(SEQ ID NO:255),RNREVRKAVHRLFKRKCC(SEQ ID NO:256),RNKEMKKAIHKLFGKKCC(SEQ ID NO:257),RNRDVKKAVHKLFRRKCC(SEQ ID NO:258),RNRDMKKAVHKLFGKRCC(SEQ ID NO:259),RNKELRKALHKLLGRKCC(SEQ ID NO:260),RNRDVRKALRRILRRRCC(SEQ ID NO:261),RNKDVRKAVRKLIRRRCC(SEQ ID NO:262),RNRDVRKAVRRLFRKRCC(SEQ ID NO:263),RNKDIKKAVKKLIKKKCC(SEQ ID NO:264),RNRELRKAVRRLFKRRCC(SEQ ID NO:265),RNKELRKAVRKIIKKKCC(SEQ ID NO:266),RNRDVKKAVRRLFRRKCC(SEQ ID NO:267),RNREVRKALRRIIRKRCC(SEQ ID NO:268),<h2 style=";text-align:left;direction:ltr">RNKDIRKAVKKIFRRKCC(SEQ ID NO:269),RNKDVRKAVRRLIKRKCC(SEQ ID NO:270),RNRDLRKAVRKLFKKKCC(SEQ ID NO:271),RNRDLRKALRRIFKRRCC(SEQ ID NO:272),RNRDVRKAIKKLIRKRCC(SEQ ID NO:273),RNKELKKAIKRILKKKCC(SEQ ID NO:274),RNRDVRKAIRKLLKRKCC(SEQ ID NO:275),RNRDLRKAVRRIFKKRCC(SEQ ID NO:276),RNRDVRKAVRKLFKRRCC(SEQ ID NO:277),RNRDVRKALRRLFKKRCC(SEQ ID NO:278),RNKELKKALRKLIGKKCC(SEQ ID NO:279),RNREMRKAIKKIIKKKCC(SEQ ID NO:280),RNKEIKKAIKKIIKKRCC(SEQ ID NO:281),RNRDVKKAIRRLFRRRCC(SEQ ID NO:282),RNREVKKAVKKLIGKRCC(SEQ ID NO:283),RNREMRKALRRLFRKRCC(SEQ ID NO:284),RNKELKKALRRLIGRRCC(SEQ ID NO:285),RNRDVKKALRKLIGKRCC(SEQ ID NO:286),RNREVKKAVKKLIRRKCC(SEQ ID NO:287),RNKEVRKALKKLFGKKCC(SEQ ID NO:288),RNKEIRKALRRLFGKKCC(SEQ ID NO:289),RNKDVKKALRRLFGKKCC(SEQ ID NO:290),RNKELKKAIKRLIRRKCC(SEQ ID NO:291),RNKDVRKAVKRLLKKRCC(SEQ ID NO:292),RNKELRKAIRRLLRRRCC(SEQ ID NO:293),RNRDIRKALRKLFKKKCC(SEQ ID NO:294),RNRELKKALRRLLRRRCC(SEQ ID NO:295),RNREVKKALRRLFGKKCC(SEQ ID NO:296),RNRDVRKALKRLLKRKCC(SEQ ID NO:297),<h2 style=";text-align:left;direction:ltr">RNRDMRKAIRKLFGRKCC(SEQ ID NO:298),RNRELKKAIRKLLKRKCC(SEQ ID NO:299),RNRDIRKAVKKLFGKKCC(SEQ ID NO:300),RNKEVKKAIRKLFGRRCC(SEQ ID NO:301),RNREVRKAVRKLFRRRKCC(SEQ ID NO:302),RNRDMKKALKKLFRRRRCC(SEQ ID NO:303),RNRDVRKALKRLLGRRCC(SEQ ID NO:304),RNKDLKKAVKKLFGRKCC(SEQ ID NO:305),RNKDVRKAVRRLFGRRCC(SEQ ID NO:306),RNRDVRKALRRLFRKKCC(SEQ ID NO:309),RNRDVRRALRRLFRKKCC(SEQ ID NO:310),RNKEVKCALKRLLKRKCC(SEQ ID NO:321),RNKEVKFALKRLLKRKCC(SEQ ID NO:322),RNKEVKHALKRLLKRKCC(SEQ ID NO:323),RNKEVKIALKRLLKRKCC(SEQ ID NO:324),RNKEVKLALKRLLKRKCC(SEQ ID NO:325),RNKEVKMALKRLLKRKCC(SEQ ID NO:326),RNKEVKSALKRLLKRKCC(SEQ ID NO:328),RNKEVKTALKRLLKRKCC(SEQ ID NO:329),RNKEVKWALKRLLKRKCC(SEQ ID NO:330),RNKEVKYALKRLLKRKCC(SEQ ID NO:331),RNKQIRDALKRLLKRKCC(SEQ ID NO:740),RNKEVKRAIKRLLKRKCR(SEQ ID NO:121),RNKEVKKAIKRLLKRKCR(SEQ ID NO:122),RNKEVKRALKRLLKRKRR(SEQ ID NO:123),RNKEVKRALKRLLKRKYP(SEQ ID NO:124),RNKEVKRALKRLLKRKRF(SEQ ID NO:125),RNKEVKRALKRLLKRKFK(SEQ ID NO:126),RNKEVKKALKRLLKRKRR(SEQ ID NO:127),RNKEVKKALKRLLKRKYP(SEQ ID NO:128),RNKEVKKALKRLLKRKRF(SEQ ID NO:129),RNKEVKKALKRLLKRKFK(SEQ ID NO:130),RNKEVKDALKRLLKRKCRR(SEQ ID NO:133),RNKEVKDALKRLLKRKCCC(SEQ ID NO:134),RNKEVKDALKRLLKRKCCF(SEQ ID NO:135),RNKEVKDALKRLLKRKCCL(SEQ ID NO:136),RNKEVKDALKRLLKRKCCM(SEQ ID NO:137),RNKEVKDALKRLLKRKCCS(SEQ ID NO:138),RNKEVKDALKRLLKRKCCP(SEQ ID NO:139),RNKEVKDALKRLLKRKCCA(SEQ ID NO:140),RNKEVKDALKRLLKRKCCY(SEQ ID NO:141),RNKEVKDALKRLLKRKCCH(SEQ ID NO:142),RNKEVKDALKRLLKRKCCN(SEQ ID NO:143),RNKEVKDALKRLLKRKCCD(SEQ ID NO:144),RNKEVKDALKRLLKRKCCK(SEQ ID NO:145),RNKEVKDALKRLLKRKCCR(SEQ ID NO:146),RNKEVKDALKRLLKRKCCG(SEQ ID NO:147),RNKEVKDALKRLLKRKCRRR(SEQ ID NO:149),RNKEVKDALKRLLKRKCRKK(SEQ ID NO:150),RNKEVKDALKRLLKRKCCRR(SEQ ID NO:151),RNKEVKDALKRLLKRKCRRRR(SEQ ID NO:154),RNKEVKDALKRLLKRKCCRRR(SEQ ID NO:156),
[0738] RNKEVKKAIKRLLKRKCCRRR(SEQ ID NO:220),
[0739] RNKEVKRAIKRLLKRKCCRRR(SEQ ID NO:237),
[0740] RNKEVKKAIKRLFKRKCCRRR(SEQ ID NO:221),
[0741] RNKEVKRAIKRLFKRKCCRRR(SEQ ID NO:238),
[0742] RNKQIRDALKRLLKRKCCRRR(SEQ ID NO:741),
[0743] RNKEVKDALKRLLKRKCRRRRR(SEQ ID NO:157),
[0744] RNKEVKDALKRLLKRKCRRRKK (SEQ ID NO:158), and
[0745] RNKEVKDALKRLLKRKCCRRRR (SEQ ID NO: 219).
[0746] 19. The olfactory receptor protein of any of paragraphs 4-17, wherein the sequence motif is RNRDVRKALRRLFRKK (SEQ ID NO: 307) or RNRDVRRALRRLFRKK (SEQ ID NO: 308).
[0747] 20. The olfactory receptor protein of any of paragraphs 4-18, wherein the sequence motif is RNKEVKKAIKRLFKRKCCRRR (SEQ ID NO: 221) or RNKQIRDALKRLLKRKCCRRR (SEQ ID NO: 741).
[0748] 21. The olfactory receptor protein of any of paragraphs 4-20, wherein the sequence motif is selected from the group consisting of SEQ ID NOs: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, and 820-890.
[0749] 22. The olfactory receptor protein according to any of the preceding paragraphs, wherein the olfactory receptor further comprises an N-terminal tag peptide, preferably wherein the N-terminal tag peptide is selected from the group consisting of a FLAG tag, a rhodopsin (Rho) tag, an SST3 tag and an M3 tag.
[0750] 23. A nucleic acid molecule comprising a nucleotide sequence encoding the olfactory receptor protein of any preceding paragraph.
[0751] 24. The nucleic acid molecule according to paragraph 23, further comprising a promoter sequence, preferably a constitutive promoter sequence.
[0752] 25. The nucleic acid molecule of paragraph 23 or 24, further comprising a terminator sequence.
[0753] 26. The nucleic acid molecule according to any of paragraphs 23-25, further comprising a nucleotide sequence encoding an N-terminal signal peptide, preferably a leucine-rich signal peptide, such as MRPQILLLLALLTLGLA (SEQ ID NO: 76) or MSHQILLLLALLTLGLA (SEQ ID NO: 77).
[0754] 27. The nucleic acid molecule according to any one of paragraphs 23-26, wherein the nucleotide sequence comprises a sequence selected from the group consisting of SEQ ID NOs: 331-739 and SEQ ID %, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence of the group consisting of NOs:742-819.
[0755] 28. An expression vector comprising the nucleic acid molecule of any of paragraphs 23-27.
[0756] 29. The expression vector of paragraph 28, wherein the expression vector is a plasmid.
[0757] 30. A recombinant host cell comprising the nucleic acid molecule of any of paragraphs 23-27 or the expression vector of paragraph 28 or 29, preferably wherein the cell expresses the olfactory receptor protein of any of paragraphs 1-22.
[0758] 31. The recombinant host cell of paragraph 30, wherein the cell further expresses one or more olfactory receptor accessory proteins.
[0759] 32. The recombinant host cell of paragraph 31, wherein the one or more olfactory receptor accessory proteins are selected from the group consisting of RTP1, RTP1S, RTP2, REEP, β-adrenergic receptor, heat shock protein 70, Ric8b, Gα olf, Giα and functional variants thereof, preferably selected from the group consisting of RTP1S, RTP2 and functional variants thereof.
[0760] 33. The recombinant host cell of any one of paragraphs 30-32, wherein the cell is a HEK293 or HEK293T cell.
[0761] 34. A library comprising a diverse repertoire of the olfactory receptor protein of any of paragraphs 1-22, the nucleic acid molecule of any of paragraphs 23-27, the expression vector of paragraph 28 or 29, or the recombinant host cell of any of paragraphs 30-33.
[0762] 35. The library of paragraph 34, wherein the olfactory receptor proteins, the olfactory receptor proteins encoded by the nucleic acid molecules or expression vectors, or the diverse repertoire of olfactory receptor proteins expressed by the recombinant host cells share the same modified C-terminal domain.
[0763] 36. The library according to paragraph 34 or 35, wherein the library comprises at least 25 (preferably 400-450) different class II olfactory receptor proteins, preferably human class II olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins, more preferably wherein the library comprises at least one, and most preferably all, class II olfactory receptors listed as "receptors" in Table 15.
[0764] 37. The library according to any of paragraphs 34-36, wherein the library comprises at least 25 (preferably 70-100) different class I olfactory receptor proteins, preferably human class I olfactory receptor proteins, nucleic acid molecules or expression vectors encoding the olfactory receptor proteins, or recombinant host cells expressing the olfactory receptor proteins, more preferably wherein the library comprises at least one, preferably all, class I olfactory receptors listed as "receptors" in Table 22.
[0765] 38. The library of paragraph 34 or 35, wherein the library comprises a diverse repertoire of at least 250 different olfactory receptor proteins, nucleic acid molecules or expression vectors encoding different olfactory receptor proteins, or recombinant host cells expressing different olfactory receptor proteins.
[0766] 39. Use of the olfactory receptor protein of any of paragraphs 1-22, the nucleic acid molecule of any of paragraphs 23-27, the expression vector of paragraph 28 or 29, the recombinant host cell of any of paragraphs 30-33, or the library of any of paragraphs 34-38 for identifying an olfactory receptor ligand, enhancer, or antagonist.
[0767] 40. Use of the library of any of paragraphs 34-38 to identify an olfactory receptor capable of binding a target ligand.
[0768] 41. A method for identifying an olfactory receptor ligand, the method comprising:
[0769] a) providing the olfactory receptor protein of any one of paragraphs 1-22 or the recombinant host cell expressing the olfactory receptor protein of any one of paragraphs 30-33;
[0770] b) contacting the receptor or recombinant host cell with a test compound or composition; and
[0771] c) Detecting activation of olfactory receptors.
[0772] 42. A method for identifying an olfactory receptor enhancer or antagonist, the method comprising:
[0773] a) providing the olfactory receptor protein as described in any one of paragraphs 1 to 22 or the cell expressing the olfactory receptor protein as described in any one of paragraphs 30 to 33;
[0774] b) contacting the receptor or recombinant host cell with a cognate ligand and a test compound or composition; and
[0775] c) Detecting increased or decreased activation of the olfactory receptor compared to a ligand-only control.
[0776] 43. The method according to paragraph 41 or 42, preferably wherein the method is used to identify an olfactory receptor ligand or enhancer, wherein the olfactory receptor is selected from the group consisting of: OR7C1, OR8K3 (preferably OR8K3(L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2(W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2(V259L)), OR10G7 (preferably OR10G7(T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2(Y28C)), OR7A5, OR7E24 (or OR7E24(P242S)) ,OR7A10,OR10H2,OR10H1,OR10D3,OR1D2,OR2A5,OR2A25(OR2A25(S75N,A209P)),OR11G2(or OR11G2(I65N,V82I)),OR14J1,OR5M3,OR8D1,OR10G3(preferably OR10G3(S73G)),OR10G9,OR 2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2 and OR2AG2 (preferably OR2AG2(Y28C)).
[0777] 44. The method of paragraph 42, wherein the method is used to identify an olfactory receptor antagonist, and wherein the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1 and OR4S2, preferably selected from the group consisting of OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)) and OR5V1, more preferably OR2M2 or OR2V1.
[0778] 45. The method of paragraph 42, wherein the method is used to identify an olfactory receptor antagonist, wherein the olfactory receptor is OR2M2 or OR2V1, and optionally wherein step b) further comprises contacting the receptor or recombinant host cell with a copper salt.
[0779] 46. The method of paragraph 45, wherein the cognate ligand is selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol, and 4-methoxy-2-methylpentane-2-thiol.
[0780] 47. The method of paragraph 42, wherein the method is used to identify an olfactory receptor antagonist, wherein the olfactory receptor is OR51B2 (preferably OR51B2 (C120R, L134F, C209S)).
[0781] 48. The method of paragraph 47, wherein the cognate ligand is 3-methyl-2-hexenoic acid.
[0782] 49. The method of paragraph 42, wherein the method is used to identify an olfactory receptor antagonist, wherein the olfactory receptor is OR5V1.
[0783] 50. The method of paragraph 49, wherein the cognate ligand is 2,4,6-trichloroanisole.
[0784] 51. A method for identifying an olfactory receptor capable of binding a target ligand, the method comprising:
[0785] a) providing a library as described in any of paragraphs 34-38;
[0786] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins from the library;
[0787] c) contacting a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins with a target ligand; and
[0788] d) identifying an olfactory receptor activated by said target ligand.
[0789] 52. A method for generating an objective representation of the olfactory properties of a test compound or composition, the method comprising:
[0790] a) providing a library as described in any of paragraphs 34-38;
[0791] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0792] c) contacting the diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with a test compound or composition; and
[0793] d) Detection of activation of each olfactory receptor protein.
[0794] 53. A method for evaluating differences or similarities between two or more test compounds or compositions, the method comprising:
[0795] a) providing a library as described in any of paragraphs 34-38;
[0796] b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library;
[0797] c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of two or more test compounds or compositions;
[0798] d) detecting activation of each olfactory receptor protein by each of the two or more test compounds or compositions; and
[0799] e) comparing the activated olfactory receptor protein between each of the two or more test compounds or compositions.
[0800] General information
[0801] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as those customary and commonly understood by one of ordinary skill in the art to which this invention belongs upon reading the contents of this disclosure.
[0802] Sequence identity
[0803] It should be understood that each nucleic acid molecule or protein fragment or polypeptide or peptide or derivative peptide or construct identified herein by a given sequence identification number (SEQ ID NO) is not limited to the specific sequence disclosed. Each coding sequence identified herein encodes a given protein fragment or polypeptide or peptide or derivative peptide or construct, or is itself a protein fragment or polypeptide or construct or peptide or derivative peptide.
[0804] Throughout this application, whenever a specific nucleotide sequence SEQ ID NO (taking SEQ ID NO: X as an example) encoding a given protein fragment, polypeptide, peptide, or derivative peptide is mentioned, it can be replaced by:
[0805] i. a nucleotide sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 99% sequence identity to SEQ ID NO: X;
[0806] ii. a nucleotide sequence whose sequence differs from the sequence of a nucleic acid molecule due to (i) the degeneracy of the genetic code; or
[0807] iii. A nucleotide sequence encoding an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 99% amino acid identity or similarity to the amino acid sequence encoded by the nucleotide sequence of SEQ ID NO: X.
[0808] A preferred level of sequence identity or similarity is 70%. Another preferred level of sequence identity or similarity is 75%. Another preferred level of sequence identity or similarity is 80%. Another preferred level of sequence identity or similarity is 85%. Another preferred level of sequence identity or similarity is 90%. Another preferred level of sequence identity or similarity is 95%. Another preferred level of sequence identity or similarity is 99%.
[0809] Throughout this application, whenever a specific amino acid sequence SEQ ID NO is mentioned (using SEQ ID NO: Y as an example), it can be replaced by: a polypeptide represented by an amino acid sequence comprising a sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 99% sequence identity or similarity to the amino acid sequence SEQ ID NO: Y. A preferred level of sequence identity or similarity is 70%. Another preferred level of sequence identity or similarity is 75%. Another preferred level of sequence identity or similarity is 80%. Another preferred level of sequence identity or similarity is 85%. Another preferred level of sequence identity or similarity is 90%. Another preferred level of sequence identity or similarity is 95%. Another preferred level of sequence identity or similarity is 99%.
[0810] In further preferred embodiments, each nucleotide sequence or amino acid sequence described herein has at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, %, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity or similarity.
[0811] Each non-coding nucleotide sequence (i.e., a non-coding nucleotide sequence of a promoter or another regulatory region) can be replaced by a nucleotide sequence comprising a nucleotide sequence having at least 60% sequence identity or similarity with a specific nucleotide sequence SEQ ID NO (taking SEQ ID NO: A as an example). Preferred nucleotide sequences are at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: A. In a preferred embodiment, such a non-coding nucleotide sequence, such as a promoter, at least exhibits or exerts an activity of such a non-coding nucleotide sequence, such as a promoter activity known to those skilled in the art.
[0812] The terms "homology", "sequence identity" and the like are used interchangeably herein. Sequence identity is described herein as the relationship between two or more amino acid (polypeptide or protein) sequences or two or more nucleic acid (polynucleotide) sequences, determined by comparing the sequences. In a preferred embodiment, sequence identity is calculated based on the full length of the two given sequences or a portion thereof, more preferably based on the full length of the two given sequences. A portion thereof preferably refers to at least 50%, 60%, 70%, 80%, 90% or 100% of the two sequences. In the art, "identity" also refers to the degree of sequence relatedness between amino acid or nucleic acid sequences, as the case may be, as determined by the degree of matching between such sequence strings. The "similarity" between two amino acid sequences is determined by comparing the amino acid sequence of one polypeptide and its conserved amino acid substituents with the sequence of a second polypeptide. "Identity" and "similarity" can be easily calculated by known methods, including but not limited to those described in Bioinformatics and the Cell: Modern Computational Approaches in Genomics, Proteomics and transcriptomics, Xia X., Springer International Publishing, New York, 2018; and Bioinformatics: Sequence and Genome Analysis, Mount D., Cold Spring Harbor Laboratory Press, New York, 2004, each of which is incorporated herein by reference.
[0813] "Sequence identity" and "sequence similarity" can be determined by comparing two peptide sequences or two nucleotide sequences using a global or local alignment algorithm, depending on the length of the two sequences. Sequences of similar length are preferably compared using a global alignment algorithm (e.g., Needleman-Wunsch), which optimally aligns the sequences over their entire lengths; whereas sequences with larger length differences are preferably compared using a local alignment algorithm (e.g., Smith-Waterman). When multiple sequences (e.g., as optimally aligned using the default parameters of the EMBOSS Needle or EMBOSS Water programs) share at least a certain minimum sequence identity percentage (as described below), they can be said to be "substantially identical" or "substantially similar."
[0814] When two sequence lengths are similar, global comparison is applicable to determining sequence identity.When the total sequence length difference is larger, preferably local comparison (such as the comparison using the Smith-Waterman algorithm).EMBOSS Needle uses the Needleman-Wunsch global comparison algorithm to compare two sequences over their entire length (full length), thus maximizing the number of matches and minimizing the number of gaps.EMBOSS Water uses the Smith-Waterman local comparison algorithm.Usually, EMBOSS Needle and EMBOSS Water use default parameters, gap open penalty=10 (nucleotide sequence) / 10 (protein), gap extension penalty=0.5 (nucleotide sequence) / 0.5 (protein).For nucleotide sequence, the default scoring matrix used is DNAfull, and for protein, the default scoring matrix used is Blosum62 (Henikoff&Henikoff, 1992, PNAS 89,915-919, incorporated herein by reference).
[0815] Alternatively, algorithms such as FASTA, BLAST, etc. can be used to determine similarity or identity percentage by searching public databases. Therefore, the nucleic acid and protein sequences of some embodiments of the present disclosure can be further used as "query sequences", to search public databases, for example, to identify other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) in Altschul et al., (1990) J.Mol.Biol.215:403-10, which are incorporated herein by reference. BLAST nucleotide searches can be performed using the NBLAST program (score = 100, word length = 12) to obtain nucleotide sequences homologous to the nucleic acid molecules of the present disclosure. BLAST protein searches can be performed using the BLASTx program (score = 50, word length = 3) to obtain amino acid sequences homologous to the protein molecules of the present disclosure. To obtain gapped alignments for comparison, Gapped BLAST can be used, as described in Altschul et al. (1997) Nucleic Acids Res. 25(17):3389-3402, which is incorporated herein by reference. When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., BLASTx and BLASTn) can be used. See the homepage of the National Center for Biotechnology Information at www.ncbi.nlm.nih.gov / .
[0816] Optionally, when determining amino acid similarity, those skilled in the art may also consider so-called conservative amino acid substitutions. As used herein, "conservative" amino acid substitutions refer to the interchangeability between residues with similar side chains. The following table lists examples of amino acid residue classes for conservative substitutions.
[0817]
[0818] Alternative conservative amino acid residue substitution categories:
[0819] 1 A S T 2 D E 3 N Q 4 R K 5 I L M 6 F Y W
[0820] Alternative physical and functional classifications of amino acid residues:
[0821]
[0822] For example, the group of amino acids with aliphatic side chains includes glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic hydroxyl side chains includes serine and threonine; the group of amino acids with amide-containing side chains includes asparagine and glutamine; the group of amino acids with aromatic side chains includes phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains includes lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains includes cysteine and methionine. Preferred conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine. Substitution variants of the amino acid sequences disclosed herein refer to variants in which at least one residue in the disclosed sequence is removed and a different residue is inserted in its place. Preferably, the amino acid changes are conservative. Preferred conservative substitutions for each naturally occurring amino acid are as follows: Ala for Ser; Arg for Lys; Asn for Gln or His; Asp for Glu; Cys for Ser or Ala; Gln for Asn; Glu for Asp; Gly for Pro; His for Asn or Gln; Ile for Leu or Val; Leu for Ile or Val; Lys for Arg; Gln or Glu; Met for Leu or Ile; Phe for Met, Leu or Tyr; Ser for Thr; Thr for Ser; Trp for Tyr; Tyr for Trp or Phe; and Val for Ile or Leu. Particularly suitable amino acid substitutions in the present disclosure include substitutions of Arg with Lys and Lys with Arg.
[0823] Gene or coding nucleotide sequence
[0824] The term "gene" refers to a DNA segment comprising a region (transcribed region) that is transcribed into an RNA molecule (e.g., mRNA) in a cell and operably linked to an appropriate regulatory region (e.g., a promoter). The coding nucleotide sequence may comprise a sequence naturally occurring in the cell, a sequence not naturally occurring in the cell, or a combination of the two.
[0825] operably connected
[0826] As used herein, the term "operably linked" refers to polynucleotide elements being linked in a functional relationship. A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, if a transcriptional regulatory sequence affects the transcription of a coding sequence, the transcriptional regulatory sequence is operably linked to the coding sequence. Operably linked means that the DNA sequences being linked are generally continuous, and when it is necessary to connect two protein coding regions, they are also continuous and in reading frame. Connection can be achieved by ligation at convenient restriction sites, or alternatively by ligation at inserted adapters or joints, or by gene synthesis.
[0827] Protein and amino acids
[0828] The terms "protein," "peptide," "polypeptide," or "amino acid sequence" are used interchangeably to refer to a molecule composed of a chain of amino acids, without regard to its specific mode of action, size, three-dimensional structure, or origin. In the amino acid sequences described herein, amino acids or "residues" are represented by three-letter symbols. These three-letter symbols and the corresponding one-letter symbols are well known to those skilled in the art and have the following meanings: A (Ala) is alanine, C (Cys) is cysteine, D (Asp) is aspartic acid, E (Glu) is glutamic acid, F (Phe) is phenylalanine, G (Gly) is glycine, H (His) is histidine, I (Ile) is isoleucine, K (Lys) is lysine, L (Leu) is leucine, M (Met) is methionine, N (Asn) is asparagine, P (Pro) is proline, Q (Gln) is glutamine, R (Arg) is arginine, S (Ser) is serine, T (Thr) is threonine, V (Val) is valine, W (Trp) is tryptophan, and Y (Tyr) is tyrosine. The residue can be any proteinogenic amino acid or any non-proteinogenic amino acid, such as D-amino acids and modified amino acids formed by post-translational modifications, as well as any unnatural amino acid. In a preferred embodiment, the amino acid in the present disclosure may refer to any one of the 20 standard protein amino acids (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y).
[0829] In this document and its claims, the verb "to comprise" and its conjugations are used in a non-limiting sense to mean that the items following the word are included or comprised, but items not expressly mentioned are not excluded. Thus, the terms "comprise," "comprising," "consisting of," and the like as used herein are synonymous with "including," "includes," or "containing," "contains," and are inclusive or open-ended and do not exclude additional, unmentioned members, elements, or method steps.
[0830] Furthermore, the verb "consisting of can be replaced with "consisting essentially of," which means that the subject matter described herein can include additional components beyond those explicitly specified, without altering the unique characteristics of the present disclosure. Furthermore, the verb "consisting of can be replaced with "consisting essentially of," which means that the methods described herein can include additional steps beyond those explicitly specified, without altering the unique characteristics of the present disclosure.
[0831] Throughout this disclosure, the term "comprising" may be replaced by the term "consisting essentially of" or "consisting of.
[0832] As used herein, the singular forms "a," "an," and "the" refer to both the singular and the plural, unless the context clearly dictates otherwise. Thus, the terms "a," "an," "one or more," and "at least one" are used interchangeably herein.
[0833] As used herein, "at least" a particular value refers to that particular value or a greater value. For example, "at least 2" should be understood to be equivalent to "2 or greater," i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc.
[0834] Furthermore, the terms "first," "second," "third," etc., in the description and claims are used to distinguish similar elements and are not necessarily used to describe a sequential or chronological order. It is understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments described herein are capable of operation in sequences other than those described or illustrated herein.
[0835] The word "about" or "approximately" when used in conjunction with a numerical value (eg, about 10) preferably means that the value may be 1% more or less than the given value (10).
[0836] As used herein, the term "and / or" means that one or more of the stated conditions may occur alone or in combination with at least one of the stated conditions, until all of the stated conditions occur.
[0837] Various embodiments are described herein. Unless otherwise indicated, each embodiment described herein may be combined together. Titles, subtitles, headings, etc. are used herein only for ease of reading and are not intended to limit or restrict the present disclosure in any way.
[0838] All patent applications, patents, and printed publications cited herein are incorporated by reference in their entirety, except to the extent of any definitions, subject matter disclaimers, or disclaimers, and unless the incorporated material is inconsistent with the explicit disclosure herein, in which case the language of the disclosure controls.
[0839] Those skilled in the art will recognize many methods and materials similar or equivalent to those described herein that can be used to practice the present invention. In fact, the present invention is in no way limited to the methods and materials described.
[0840] The present invention will be further described by the following examples, but these examples should not be considered as limiting the scope of the present invention. Certain aspects and features disclosed in the following examples summarize the above description and form part of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0841] Figure 1 Comparison of wild-type OR5A2 with wild-type C-terminal sequence (SEQ ID NO: 239) and chimeric OR5A2 with optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). Shown is the dose-dependent luciferase induction of different ligands.
[0842] Figure 2 Example of dose response analysis of wild type and chimeric OR7C1 and OR9Q2 with optimized C-terminal sequence (SEQ ID NO: 86) used to derive detection thresholds and EC50 values as summarized in Table 1.
[0843] Figure 3 Dose response analysis of wild-type OR5B12 with a wild-type C-terminal sequence (SEQ ID NO: 241) and a chimeric OR5B12 with an optimized C-terminal sequence (SEQ ID NO: 86) compared to a variant that altered the wild-type C-terminal sequence toward the consensus sequence (SEQ ID NO: 242) by changing the amino acid sequence at two positions to leucine residues, thereby leaving all conserved residues identical to the consensus sequence.
[0844] Figure 4Dose response analysis of wild-type OR5A2 with a wild-type C-terminal sequence (SEQ ID NO: 239) and a chimeric OR5A2 with an optimized C-terminal sequence (SEQ ID NO: 86) compared to a variant that altered the wild-type C-terminal sequence toward the consensus sequence (SEQ ID NO: 240) by changing the amino acid sequence at three positions such that all conserved residues are identical to the consensus sequence.
[0845] Figure 5 Dose response analysis of wild-type OR7A17 with a wild-type C-terminal sequence (SEQ ID NO: 243) and a chimeric OR7A17 with an optimized C-terminal sequence (SEQ ID NO: 86) compared to a variant that altered the wild-type C-terminal sequence toward the consensus sequence (SEQ ID NO: 244) by changing the amino acid sequence at five positions such that all conserved residues are identical to the consensus sequence.
[0846] Figure 6 Dose response analysis of wild-type and chimeric O8K3 with different ligands. The weaker agonist, menthone, could only be identified using the more sensitive variant with an optimized C-terminus (SEQ ID NO: 86).
[0847] Figure 7 Dose response analysis of wild-type and chimeric O7D4 with different ligands. The weaker agonist androstenol could only be identified using the more sensitive chimeric variant with an optimized C-terminus (SEQ ID NO: 86).
[0848] Figure 8 OR2T4 wild type and OR2T4 with an optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were screened against 52 sulfur compounds.
[0849] Figure 9 Dose response analysis of wild-type and chimeric OR2T4 with various ligands. OR2T4 has no known ligands, and screening of the wild-type receptor with potential ligands did not result in deorphanization. However, testing of a chimeric variant with an optimized C-terminus (SEQ ID NO: 86) revealed that the receptor is specifically activated by specific sulfur-containing compounds, such as cyclopentanethiol.
[0850] Figure 10 OR2T11 wild type and OR2T11 with an optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were screened against 52 sulfur compounds.
[0851] Figure 11.Example of dose response analysis of wild type and chimeric OR7C1 with C-terminal sequence (SEQ ID NO:86) and variants with optimized C-terminal sequences RNKEVKRAIRKLLKRKCC (SEQ ID NO:118) and RNKEVKRAIKRLLKRKCR (SEQ ID NO:121), which variants combine multiple sequence variations described in Examples 6 and 7.
[0852] Figure 12 Activation of OR52E8 by odorous acids present in human sweat. The wild type was not activated by acid. Replacing the wild type with the C-terminal sequence of a functional OR51E1 did not improve expression. However, using the optimized sequence from Examples 1-5 (SEQ ID NO: 86) resulted in a strong signal after addition of 3-methyl-3-hydroxyhexanoic acid.
[0853] Figure 13 Activation of OR56A4 by acids of different chain lengths. The wild type was activated by decanoic and undecanoic acids at high concentrations, but not by nonanoic acid. Replacing the wild type with the C-terminal sequence of functional OR51E1 did reduce activity. However, using the optimized sequence from Examples 1-5 (SEQ ID NO: 86) resulted in strong signals and much lower detection thresholds after addition of all three acids.
[0854] Figure 14 Different chimeric variants of OR8K3: The wild-type C-terminal sequence was replaced with the optimized sequence from Examples 1-5 (SEQ ID NO: 86) or the C-terminal sequences of two functional receptors (i.e., OR1N2 and OR5AN1). Chimeric receptors with C-terminal sequences from other functional receptors did not provide functional expression. Dose-dependent induction of luciferase by different ligands was demonstrated.
[0855] Figure 15 Chimeric variants of OR5AN1: The wild-type C-terminal sequence was replaced with the C-terminal sequence of the functional receptor OR1N2. The chimeric receptor with the C-terminal sequence from OR1N2 did not provide functional expression. Dose-dependent induction of luciferase by different ligands was demonstrated.
[0856] Figure 16 Chimeric variants of OR5A2: The wild-type C-terminal sequence was replaced with the C-terminal sequence of the functional receptor OR1N2. While the chimeric receptor with the C-terminal sequence from OR1N2 did provide functional expression, its functional expression was approximately 100-fold weaker than that of OR5A2 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO:86). Dose-dependent induction of luciferase by the different ligands is shown.
[0857] Figure 17 Activation of OR7A17 with an optimized C-terminal RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) by Arborone (left) compared to Iso E Super (OTNE, right). Shown is the dose-dependent induction of luciferase by the mentioned ligands.
[0858] Figure 18 Comparison of in vivo odor thresholds (OTH_avg, ng / L) of tested chemicals with EC50% (in μM) of in vitro activation assay against OR7A17 with optimized C-terminal RNKEVKDALKRLLKRKCC (SEQ ID NO: 86).
[0859] Figure 19 OR2M2 with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) was screened against 52 sulfur compounds in the presence of 30 μM copper.
[0860] Figure 20 Dose response analysis of OR2M2 with an optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) against the key human odorant sulfur compound 3-methyl-3-thio-hexanol and related chemicals in the presence and absence of copper.
[0861] Figure 21 The entire library of class II OR variants (n=408) was screened with 50 μM patchouli alcohol. The Y axis shows the fold induction of luciferase. The OR numbers on the X axis correspond to those in Table 15.
[0862] Figure 22 Dose response analysis of the library hit patchouliol-OR14J1, comparing the modified OR sequence to the wild type. The wild type is inactive, and OR deorphanization is only possible with the modified sequence.
[0863] Figure 23 A library of class I OR variants (n=77) was screened with 500 μM 3-methyl-2-hexenoic acid. The Y-axis shows fold induction of luciferase. The OR numbers on the X-axis correspond to those in Table 22.
[0864] Figure 24 The entire library of class II OR variants (n=408) was screened using three different fragrance oils tested at 10 ppm. The Y-axis shows fold induction of luciferase. On the X-axis are depicted ORs that resulted in at least 2-fold luciferase induction by at least one fragrance. Only activated ORs are shown.
[0865] Figure 25 Geranium oil spiked with varying levels of ambroxan (0.1%-10%) was tested in a dose-response assay on OR7A17, which has the C-terminal domain of SEQ ID 221 (the full DNA sequence encoding the modified receptor is SEQ ID NO: 611). The Y-axis shows fold induction of luciferase. The X-axis indicates the concentration of geranium oil in ppm.
[0866] Figure 26 Iso E super and Ambermax were screened on cells expressing either or both OR7A17 or OR7C1 with the C-terminal domain of SEQ ID 221. The Y axis shows fold induction of luciferase. Concentration in μM is indicated on the X axis.
[0867] Figure 27 OR10J5, which has a C-terminal domain of SEQ ID 86, is activated by two structural isomers, (S,E)-10-hydroxy-4,8-dimethyldec-4-enal and (R,E)-10-hydroxy-4,8-dimethyldec-4-enal. The Y-axis shows the % luciferase induction of the positive control (Mahonial). The X-axis indicates the concentration of the test compound in micromolar. Example
[0868] General method of OR expression
[0869] To create expression plasmids, the OR coding sequence fused to the desired C-terminal sequence at the TM7 terminus was synthesized by a DNA synthesis service provider (BioCat GmbH, Germany) and inserted into pcDNA3.1 (+) downstream of the CMV promoter sequence (SEQ ID NO: 82) using BamHI and NotI restriction sites (Invitrogen, MA, USA). All synthetic OR nucleotide sequences described in Examples 1-15 below also contain a nucleotide sequence encoding a signal peptide (mmLucy-FLAG-rho, SEQ ID NO: 80, 81) at its N-terminus and a bgh terminator sequence (SEQ ID NO: 83) at the C-terminus. All OR expression plasmids also contain a Kozak sequence (GCCACC) between the BamHI restriction site and the start codon of the signal peptide. Thus, these plasmids contain constitutively expressed OR genes.
[0870] OR gene expression was typically performed in HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). These cells were seeded at a density of 10,000 cells / well into polyethyleneimine-coated 96-well plates (100 μl / well) and grown at 37° C. in the presence of 5% CO 2 for 24 hours.
[0871] 0.625 μg of OR expression plasmid, 1 μg of empty pcDNA3.1(+) vector and 1 μg of pGL4.29 (Promega) carrying CRE inducible luciferase were mixed in 0.25 ml of OptiMEM medium (Gibco TM At the same time, 15 μl of Lipofectamine 2000 (Invitrogen) was diluted in 0.25 ml of OptiMEM medium, and after a 5-minute preincubation, the two mixtures were combined to prepare a transfection mixture, which was incubated for another 25 minutes.
[0872] 50 μl of growth medium was replaced with fresh DMEM containing 9% fetal bovine serum (FBS). The pre-incubated transfection mixture was diluted to 5 ml in OptiMEM medium, and 50 μl of the diluted mixture was added to each well (total final volume 150 μl). The cells were further incubated at 37°C for 24 h in the presence of 5% CO2 to allow DNA uptake and expression of ORs.
[0873] Functional expression and responsiveness to ligand were tested by removing 100 μl of growth medium and adding 50 μl of DMEM containing 9% FBS, ligand and a maximum of 1% DMSO. Cells were stimulated for 4.5 h and then luciferase signaling based on induction of OR-dependent cAMP production was measured.
[0874] Example 1: Using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) Improved functional expression of OR5A2
[0875] Wild-type OR5A2 and OR5A2 modified with an optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were transfected into HEK293T cells that had stably integrated DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85).
[0876] No activation of the wild-type receptor by the polycyclic musk galaxolide or the macrocyclic musk muscone was detected ( Figure 1 , left). However, the modified receptor was activated by two musk compounds. Figure 1 As shown in (right panel), activation occurred at low concentrations (<0.01 μM for galaxolide and 1 μM for muscone) and was specific for musk compounds, with no response recorded for ethyl vanillaldehyde.
[0877] Therefore, OR5A2 expressed together with RTP1S and RTP2 and Lucy-FLAG-rho-Tag in cell lines can be effectively screened for potential musk odor compounds only by using modified variants of OR5A2 with modified C-terminal sequences instead of the wild-type C-terminal sequence.
[0878] Example 2: Modification of a large number of genes using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) The sensitivity of the modified OR is improved
[0879] Modified forms of different OR genes were generated by exchanging the C-terminal sequence after TM7 with the optimized sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), as shown for OR5A2 in Example 1. Cells were transfected with wild-type OR genes or the corresponding modified forms. In a dose-response analysis, the transfected cells were stimulated with the cognate ligands of these receptors, and the lowest concentration of ligand that triggered a 2-fold induction of the luciferase signal over the background was determined from the dose-response curve as a measure of the lower limit of detection. In parallel, the EC50, i.e., the concentration that achieved 50% of the maximum activation (potency), was calculated. Finally, the maximum fold induction of luciferase over the solvent control (efficacy) was determined and compared between the wild-type and modified forms.
[0880] As shown in Table 1 and Figure 2 As can be observed in the illustrated examples, for all receptors tested, the detection threshold (2-fold induction concentration in μM) was significantly reduced using the optimized sequence SEQ ID NO: 86, and for many receptors, this increase in sensitivity was 10-fold to more than 100-fold. This increased sensitivity was also observed by comparing the EC50 values, however, the EC50 values were also affected by the maximum induction (efficacy) observed, with the maximum induction of the wild type being lower than that of the modified form for several receptors. Thus, for OR52E8, OR2T4, and OR5A2, no induction of the wild type was observed, while functional expression (all or nothing effect) was achieved with the modified form. Other receptors (e.g., OR8K3, OR5B12, OR7C1, OR7D4) also had much lower efficacy for the wild type (much higher detection threshold close to that of the wild type) than the modified form.
[0881] Table 1. Increased functional expression of modified OR genes containing a truncated, optimized C-terminal RNKEVKDALKRLLKRKCC (SEQ ID NO: 86)
[0882]
[0883]
[0884] There was no luciferase induction above background at the maximum concentration tested.
[0885] na Not applicable, insufficient induction to calculate EC50
[0886] *Higher EC50 due to much higher efficacy of the optimized C-terminus
[0887] Example 3: Compared with OR genes optimized by changing the C-terminus towards the consensus sequence, the use of optimized C-terminus Improved functional expression of ORs modified with the terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86)
[0888] One option for optimizing expression that has been studied previously (Kotthoff et al. 2021, see above) is based on the parental C-terminal sequence of the receptor, which is varied towards the consensus sequence by restoring the consensus sequence at a single non-conserved amino acid, but all residues in the C-terminal sequence remain unchanged for which multiple sequence alignment has no clear consensus residues. Therefore, for each receptor, a specific C-terminal sequence is derived from the specific parental sequence of that receptor. This approach was compared with a method in which the same optimized, modified C-terminus as described in Examples 1 and 2 was added to each receptor. Therefore, the C-terminal sequence of OR5B12 was altered by introducing two leucine residues at the positions of the most common amino acids in the consensus sequence. At all other positions, OR5B12 already contains residues typical of the consensus sequence (Kotthoff et al. 2021, see above). As Figure 3 As shown, changing the sequence of OR5B12 towards the consensus sequence had no significant effect, whereas the optimized sequence resulted in strong functional expression compared to wild type. Therefore, restoring the consensus sequence at the conserved amino acids in the parental sequence as done in (Kotthoff et al. 2021, supra) is not a generally applicable approach for improved functional expression, but rather generating modified receptors with optimized C-terminal sequences as described herein is a generally applicable approach.
[0889] Similarly, for OR5A2, altering the native C-terminus towards the consensus sequence by exchanging three amino acids had no effect, and similar to the wild type, no functional expression was achieved by this approach, whereas addition of the optimized C-terminus resulted in functional expression ( Figure 4 ).
[0890] In the case of OR7A17, which as wild type had a relatively low detection threshold of 1.9 μM for a 2-fold induction by ambroxan, changing the sequence towards the consensus sequence by exchanging 5 amino acids did have a positive effect, lowering the detection threshold 3-fold to 0.6 μM, whereas using an optimized C-terminal sequence achieved a better effect, as the detection threshold could be lowered 18-fold (Tables 1 and Figure 5 ). Therefore, the best improvement in activity is not achieved by simply using the parent C-terminus and mutating it towards the consensus sequence, but rather by exchanging the C-terminal sequence with the optimized sequence described herein.
[0891] Example 4: Use of the optimized C-terminal RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) to obtain a more comprehensive Ligand profiling and discovery of new ligand-OR pairs
[0892] The ligand profile of the modified OR with the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) was further tested using various ligands and compared to the OR with the wild-type sequence.
[0893] like Figure 6 As shown, the wild-type receptor OR8K3 exhibited only a weak response to menthols (natural menthol and (S)-menthol ("L-menthol")), while the related molecule menthone was inactive up to 316 μM. Using the modified receptor variants, stronger activation by menthol was observed, while menthone also activated the receptor, albeit with less potency. Thus, the improved assay sensitivity provided by the modified variants allowed the detection of activation of OR8K3 by a broader set of minty flavor molecules.
[0894] Similarly, in Figure 7 In Figure 2, the dose response of OR7D4 is shown when tested with androstenone and the closely related molecule androstenol. Significant activation of androstenone is shown for both variants, while androstenol can only activate the modified form of the receptor: androstenol is a significantly weaker agonist, but appears to lack activation of OR7D4 if tested only on the poorly expressed wild-type variant. Thus, a better understanding of the receptor space for OR7D4 can be achieved based on the improved assay sensitivity provided by the modified variants.
[0895] Example 5: Using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) identification OR-matching Body pair
[0896] Little is known about the binding of key sulfur odorants to human ORs. Therefore, a library of sulfur odorants was screened for potential sulfur compound receptors. Figure 8As shown, for OR2T4 with the wild-type sequence, no clear OR-ligand pair was identified, whereas for the more sensitive modified variants with optimized C-termini, diallyl disulfide, dipropyl disulfide, 2-methyl-3-tetrahydrofuranthiol, and cyclopentanethiol were identified as good ligands for this receptor. This difference in screening efficacy was also confirmed by dose-response analysis using wild-type and receptor variants ( Figure 9 Similarly, for OR2T11, none of the 52 sulfur compounds tested were active when tested on wild type, whereas a more sensitive modified variant with an optimized C-terminus identified 8 ligands with >4-fold induction for this receptor ( Figure 10 ). Thus, the optimized C-terminal sequence promotes receptor deorphanisation.
[0897] Example 6: Optimized C-terminal sequence flexibility: base substitutions in the C-terminal sequence motif
[0898] As shown in Examples 1-5, the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) allows for improved heterologous OR activation experiments with a large number of ORs, allows for the deorphanization of ORs, the discovery of new ligands for the deorphaned ORs, testing at lower ligand concentrations with fewer solubility issues and lower cytotoxicity and achieving higher efficacy to obtain a better signal-to-noise ratio.
[0899] To test possible variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that resulted in similarly improved functional expression, different point mutations were introduced into the first 16 amino acids of this sequence (corresponding to RNKEVKDALKRLLKRK, SEQ ID NO: 10) and the variants were fused to OR7C1 after TM7. Figure 2As shown, OR7C1 with C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) produces about 10 times of luciferase induction under 1 μM Ambermax, and produces 20 times of induction under 3.1 μM Ambermax. Then the activation of the variants of different modifications is compared with the variants with the above-mentioned optimized sequence at these two concentrations. Therefore, all variants are transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). The cells are then stimulated with 1 μM and 3.1 μM Ambermax, and the induction fold of luciferase is compared with the induction fold in the same experiment performed on the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which is set to 100% and compared with the wild type. The test concentration is selected to fall within the partial activity range of the wild type sequence.
[0900] Sequence modifications that give >66% activation of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) at 1 μM Ambermax are considered active variants and can be preferentially used for optimized expression. Sequence modifications that give 33-66% activation of the reference sequence at 1 μM Ambermax are considered partially active variants and can optionally be used for optimized expression. As shown in Table 2, although the optimized sequence used in Examples 1-5 is one possible choice, other variants with single amino acid substitutions have similar activity and, in some cases, even provide improved activity.
[0901] Table 2. Ambermax activation of OR7C1 variants with optimized C-termini containing single base substitutions in the first 16 amino acids of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86).
[0902]
[0903]
[0904]
[0905] Example 7: Flexibility of the C-terminal sequence - Amino acids at positions 17 and 18 of the C-terminal sequence motif base substitution
[0906] To further test possible variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that resulted in similar benefits, 120 different point mutations were introduced into the optional last two amino acids (positions 17 and 18) of the sequence, and all variants were fused after TM7 of OR7C1. The activation of the different modified variants was then compared with the variants having the optimized sequence described in Examples 1-5 above.
[0907] Therefore, all variants were transfected into HEK293T cells that had been stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). The cells were then stimulated with 1 μM and 3.1 μM Ambermax, and the fold induction of luciferase was compared with the fold induction in the same experiment performed on the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which was set to 100%, and compared with the wild type.
[0908] As shown in Table 3, there is flexibility in the last two amino acids, and many variants are active, but having inappropriate amino acids at these positions can also be completely inactive.
[0909] Table 3. Ambermax activation of variants fused to the optimized C-terminus of OR7C1: Effect of base substitutions at amino acid positions 17 and 18 of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86)
[0910]
[0911]
[0912] Example 8: C-terminal sequence flexibility - amino acid truncation at positions 17 and 18 of the C-terminal sequence motif
[0913] To further test possible variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) that produced similar benefits in functional expression, further truncations of one or two amino acids were introduced and the sequence was fused to OR7C1, OR2T4, or OR5B12 after TM7. The activation of the different modified variants was then compared with that of the variants with the optimized sequence described above. For optimal activity, 30 μM copper was added to the assay with OR2T4 (added as CuCl2).
[0914] As shown in Table 4, a sequence motif having the first 16 amino acids of RNKEVKDALKRLLKRKCC (RNKEVKDALKRLLKRK, SEQ ID NO: 10) was sufficient for high functional expression of OR7C1, indicating that the additional amino acids at positions 17 and 18 are optional and do not need to be introduced for functional expression of all ORs. However, in the case of OR2T4, omitting these two additional amino acids resulted in reduced activity, and in OR5B12 there was a complete loss of activity, indicating that when using these ORs or if a library of many or all ORs is generated, it is beneficial to add additional amino acids at positions 17 and 18, as the optional amino acids allow for a broader improvement in the expression of different ORs.
[0915] Table 4. Activation of modified OR7C1, OR2T4 and OR5B12 variants with optimized C-terminus: Effect of truncation of optional amino acids at positions 17 and 18 of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86)
[0916]
[0917]
[0918] Example 9: Flexibility of the C-terminal sequence - addition of additional amino acids at positions 19-22
[0919] As shown in Examples 1-5 and 6, the two cysteine residues at positions 17 and 18 of the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO:86) provide improved functional expression of multiple receptors. As shown in Example 8, these amino acids are optional and are not essential for the improved expression of all ORs. As shown in Example 7, these two last amino acids have flexibility, and particularly maintain or even improve activity by replacing one of the Cys-residues with basic amino acids, thereby particularly adding arginine as the second residue to improve the activity of OR7C1. Variants extended with more basic amino acids further utilize this effect. For ease of reference, the positioning of SEQ ID NO:86 (with 18 amino acids) is used as a reference to represent the position corresponding to the extended variant, so the position 19-22 of SEQ ID NO:86 used herein corresponds to the position present in the extended variant. In other words, mentioning the SEQ ID NO:86 position 19 in the variant means that the variant has been extended by 1 amino acid, etc.
[0920] Thus, having 1-6 basic amino acids at amino acid positions 17-22 improves the activity of O7C1, as shown in Table 5. Similarly, having 2-3 basic amino acid residues at positions 18-20 improves the activity of OR5B12 (Table 6) compared to having only two Cs at positions 17 and 18. For OR2T4, having 2-3 basic amino acid residues at positions 18-20 does not further increase activity, but it also does not harm activity (Table 7), showing that in the library to which the same C-terminal end is added to each OR, these additions are preferred because they only improve or maintain activity, but do not impair the activity of the test receptor. In addition, the different amino acids added at position 19 improve the activity of OR5B12 (Table 6), which particularly depends on the additional amino acids outside position 16 (see Example 8). The method used and the mode of calculation of the results are the same as in Example 6-8.
[0921] Table 5. Activation of OR7C1 variants with optimized C-termini containing additional basic residues at amino acid positions 18-22 by Ambermax.
[0922]
[0923] Table 6. Activation of OR5B12 variants with optimized C-termini containing additional residues at amino acid positions 17-20 by dihydrojasmonate HC.
[0924]
[0925]
[0926] Table 7. Activation of OR2T4 variants with optimized C-termini containing additional basic residues at amino acid positions 17-20 by cyclopentanethiol
[0927]
[0928]
[0929] *For optimal activity, 30 μM copper was added to the assay with OR2T4 (added as CuCl2).
[0930] Example 10: C-terminal sequence flexibility: combinations of functional base substitutions
[0931] As shown in Examples 6-9, following the optimized sequences of Examples 1-5, functional variants were identified at single amino acid positions within the first 16 amino acids of RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), corresponding to RNKEVKDALKRLLKRK (SEQ ID NO: 10), and within the optional additional C-terminal cysteine residues at positions 17 to 22. It is further demonstrated herein that these functional variants can be combined with each other to provide many different functional variants of the C-terminus, and in some combinations even variants with further improved activity compared to SEQ ID NO: 86.
[0932] Table 8 lists variants of the C-terminal sequences described in Examples 6, 7, and 9, which were combined into the same C-terminal sequence and fused to OR7C1. The methods used and the way the results were calculated were the same as in Examples 6-8.
[0933] It is evident from the results that functional variants combined within the same C-terminal sequence gave all functional combinations and that by combining variants with enhanced functionality even further enhanced variants of the C-terminal sequence were obtained. In addition to the improved responses at two specific screening concentrations as shown in Table 8, this improvement was also evident from the Figure 11 This is also evident in the dose response analysis shown in the representative examples.
[0934] Table 8. Activation by Ambermax of OR7C1 variants with optimized C-termini containing multiple base substitutions selected from active variants identified in Examples 6, 7 and 9. Variant residues compared to SEQ ID NO: 86 are shown in bold and underlined.
[0935]
[0936]
[0937] Example 11: C-terminal sequence flexibility: testing functional variants with different parent receptors As shown in Example 2, the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) functionally improves the responses of multiple receptors. As further shown in Examples 6-10, functional variants of the optimized C-terminal sequences of Examples 1-5 can be identified that are still active or even have improved activity when tested with OR7C1. The functional variants of Examples 6-10, and in particular combinations of variants, were further tested to determine whether they are also broadly applicable to different ORs.
[0938] Variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were fused to OR5B12 and transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). The cells were then stimulated with 31 μM and 62.5 μM dihydrojasmonate HC, and the fold induction of luciferase was compared with the fold induction in the same experiment with the receptor having the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), the fold induction of the latter being set to 100%, and compared with the wild type.
[0939] As shown in Table 9, the combination of variants identified as maintaining or improving OR7C1 expression (Example 10) also significantly further improved OR5B12 expression, demonstrating that the identified variants can be generalized and applied to other ORs.
[0940] Table 9. Activation of OR5B12 variants with optimized C-termini containing multiple base substitutions selected from the variants identified in Examples 6-9 by dihydrojasmonate HC.
[0941]
[0942]
[0943] Similarly, variants of the sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) were fused to OR2T4 and transfected into HEK293T cells stably transfected with DNA sequences encoding functional variants of human RTP1S (V227I, SEQ ID NO: 84) and RTP2 (L220R, SEQ ID NO: 85). The cells were then stimulated with 4 μM and 8 μM 2-methyl-3-tetrahydrofuranthiol, and the fold induction of luciferase was compared with the fold induction in the same experiment performed on the standard sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86), which was set to 100%, and compared with the wild type.
[0944] As shown in Table 10, the variant combinations identified as maintaining or improving OR7C1 expression (Example 9) were all active when tested with OR2T4: they did not further improve functional activity for this receptor; however, they were not detrimental to activity either, suggesting that in a library where the same C-terminus is added to each OR, these variants may still be preferred because they only improve or maintain activity but do not compromise the activity of the receptor.
[0945] Table 10. Activation of OR2T4 variants with optimized C-termini containing multiple base substitutions selected from the variants identified in Examples 6-9 by 2-methyl-3-tetrahydrofuranthiol.
[0946]
[0947]
[0948] Variants of the sequence motif RNKEVKDALKRLLKRKCC were also tested with OR5A2 as shown in Table 11. Similar to OR2T4, the identified variant was also fully functional when applied to OR5A2, although it did not further enhance the activity (which was already very high, induced by musk ketone starting from 0.03 μM).
[0949] Table 11. Dose-dependent induction of OR5A2 with two different variants of the C-terminal sequence motif (fold induction of luciferase is shown)
[0950]
[0951] Example 12: Replacement of the C-terminus in a Class I Olfactory Receptor with an Optimal C-terminal Sequence Originally Derived from a Class II OR sequence
[0952] The optimized sequences in Examples 1-5 were initially obtained by mutating the C-terminal sequence of class II OR5AN1. The C-terminal sequence of class I ORs differs greatly from that of class II, also leading to a different C-terminal consensus sequence for class I receptors (Kotthoff et al. 2021, supra). It was therefore further tested whether the optimized C-terminal sequence developed starting from class II receptors could also be applied to class I receptors. Surprisingly, this sequence in fact also led to a clear increase in the expression of several class I receptors, as already summarized in Example 2 for OR52A5, OR52E8 and OR56A4. A more detailed comparison is given in Figure 12 and Figure 13 Given in.
[0953] Activation of OR52E8 by odorant acids present in human sweat (Natsch et al., A specific bacterial aminoacylase cleaves odorant precursors secreted in the human axilla. The Journal of biological chemistry 2003; 278(8): 5718–5727) was investigated using different variants of the parent receptor sequence. The wild-type variant was not activated by acid. Replacing the wild-type C-terminus with a functional OR51E1 C-terminal sequence did not provide a functional OR52E8. However, using the optimized sequence from Examples 1-5 resulted in a strong signal after addition of 3-methyl-3-hydroxyhexanoic acid.
[0954] We further tested the activation of OR56A4 by acids of varying chain lengths. The wild-type was activated by decanoic and undecanoic acids at high concentrations, but not by nonanoic acid. Replacing the wild-type with the C-terminal sequence of a functional OR51E1 did reduce activity. However, using the optimized sequence from Examples 1-5 resulted in strong signals and a much lower detection threshold after addition of all three acids.
[0955] Example 13: Compared with the addition of the optimal C-terminal sequence of the present disclosure, the C-terminal sequence of the functional receptor Replace the C-terminal sequence of poorly functioning receptors
[0956] As observed in the preceding examples, knowing the importance of the C-terminal sequence for correct expression, an obvious approach to improving expression might simply be to provide hybrid receptors whereby the C-terminal domain in a non-functional or poorly functional receptor is replaced with the C-terminal domain of a functional receptor. Although this approach has been tested in the past (Ikegami et al. 2020, supra), whereby the C-terminus of the poorly expressing mouse Olf541 was replaced with the C-terminus of the well-expressing Olfr539 and did not result in enhanced expression of Olfr541, this approach was further tested in comparative examples with different human ORs. Thus, the C-terminal sequence of OR8K3 was replaced with the C-terminal sequence of either the tansylide-receptor OR1N2 or the musk receptor OR5AN1. For both receptors, the wild type was functional, indicating that they can function correctly with their native C-terminal sequence and that this C-terminal sequence could, in principle, provide functional expression. However, modified OR8K3 variants with either of these C-terminal sequences did not show any functional response to the cognate ligand menthol ( Figure 14), whereas the modified receptors with the C-terminal sequences according to Examples 1-5 resulted in strong functional expression. Thus, the improvements achieved with the modified C-terminal sequences of the present disclosure cannot be achieved by simply creating a chimeric receptor that combines the functional C-terminal domain of a functional receptor with a poorly expressed receptor.
[0957] Similarly, the C-terminal sequence of OR5AN1 was replaced with the C-terminal sequence of the ambrette lactone receptor OR1N2. For both receptors, the wild type was functional, indicating that they can function correctly with their native C-terminal sequences. However, the chimeric OR5AN1 variant with the C-terminal sequence of OR1N2 did not show any functional response to the cognate ligands musk ketone or muscone ( Figure 15 ).
[0958] Similarly, the musk receptor OR5A2 having the C-terminal sequence of OR1N2 has poorer sensitivity compared to the variant receptor of the present disclosure having the C-terminal sequence RNKEVKDALKRLLKRKCC ( Figure 16 ). Thus, the improvements achieved with the modified C-terminal sequences of the present disclosure cannot be achieved by simply creating a chimeric receptor combining a functional C-terminal domain of a functional receptor with a poorly expressed receptor.
[0959] Example 14: Using the optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) Screening Arborone's odorous substances of desired odor qualities
[0960] A modified variant of OR7A17 with an optimized C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) as described in Example 2 was tested on Arborone (Hong & Corey. Enantioselective syntheses of georgyone, arborone, and structural relatives. Relevance to the molecular-level understanding of olfaction. Journal of the American Chemical Society 2006; 128(4): 1346–1352), which is the key characteristic odor component of the complex Iso E Super. Figure 17As shown, arborone activates OR7A17 with an optimized C-terminus at concentrations as low as 1 nanomolar, approximately 50-fold lower than IsoE super, indicating that hybrid OR7A17 with an optimized C-terminus is a powerful tool for screening arborone for desired woody notes. Therefore, a number of (n=22) woody fragrance molecules were tested for their odor detection threshold (OTH) in humans and for the in vitro EC50 of OR7A17 with an optimized C-terminal sequence, RNKEVKDALKRLLKRKCC (SEQ ID NO: 86). Figure 18 As shown, in vitro data well predicted the low odor detection thresholds for Ambroxan, Georgywood, and Iso E super (highlighted in the lower left with chemical structures), distinguishing them from odorants of moderate potency and the much less active (in vitro and in vivo) cedrol (highlighted in the upper right).
[0961] Example 15. Screening for receptor activation by key human sweat odorants
[0962] The most pungent odorant in human underarm sweat with an ODT of 1 pg / L air is 3-methyl-3-mercapto-hexanol (also known as 3-methyl-3-mercapto-hexanol) (Natsch et al. Identification of Odoriferous Sulfanylalkanols in Human Axilla Secretions and Their Formation through Cleavage of Cysteine Precursors by a CS Lyase Isolated from Axillabacteria. Chemistry and Biochemistry 2004; 1: 1058-1072). To date, no OR has been described that can selectively detect this key odorant that causes undesirable underarm odor in human subjects. By screening a library of receptor variants with the C-terminal sequence SEQ ID NO: 86 against a library of 52 sulfur compounds in the presence of copper (see Example 5), the hitherto orphan receptor OR2M2 was identified, which is activated by 3-methyl-3-mercapto-hexanol and closely related chemicals, but not by other sulfur odorants ( Figure 19 ). Dose-response analysis ( Figure 20) showed that these identified ligands could activate OR2M2 at concentrations as low as 1 μM when tested in the presence of copper alone, an effect observed for other receptors that respond to sulfur compounds. Thus, using the improved sequences of the present disclosure, it was possible for the first time to deorphanize OR2M2, and the modified OR2M2 with an optimized C-terminus could therefore be used in combination with one of its cognate ligands, 3-methyl-3-thio-hexanol, 2-mercapto-2-methyl-pentanol, or 4-methoxy-2-methylpentane-2-thiol, in the presence of copper as a screening target for the most effective human underarm odor substances.
[0963] The complete amino acid and DNA sequences of the OR2M2 modified sequence having the modified C-terminal domain of SEQ ID NO: 86 are provided as SEQ ID NOs: 245 and 246.
[0964] Example 16: Random Permutation of Variable Positions within the Modified C-Terminus of the Disclosure (SEQ ID NO: 1)
[0965] Generation of a universal C-terminal sequence linked to OR7D4 based on a degenerate oligonucleotide in the C-terminal domain For this experiment, the EcoRI / SalI fragment of pRDVCCB-CMV-dCas9-VPH-2A-Blast (Cellecta, Inc., Mountain View, USA) was replaced with the MfeI / XhoI fragment from pcDNA3.1(+)-mmLucy-FLAG-rho-OR7D4 (containing the CMV promoter and OR coding sequence). The coding region of the C-terminal domain of OR7D4 (between Bsu36I and NotI) was then replaced with a complementary degenerate oligonucleotide. Random clones were selected, sequenced, and tested for the correct open reading frame. The correct open reading frame and universal sequence of OR7D4 were used. As the C-terminal domain, a total of 55 clones were obtained. By functional assays tested in a dose-response test with the ligand androstenone in HEK cells, all different clones were compared with the wild-type sequence. Although the wild-type sequence was not significantly induced by 1 μM of ligand, the induction of androstenone to different variants was 10.1-73.6 times at this concentration. The EC2 for two-fold induction of wild-type luciferase was 1.71 μM, while the EC2 for different variants was reduced to 0.003 μM to 0.13 μM, which corresponds to an increase in the sensitivity of the assay of the modified C-terminal sequence by 13-545 times, with a median of 125 times the increased sensitivity. These results show that random permutations of the universal sequences tested produce significantly increased (>10 times) the sensitivity of the functional assay.
[0966] Table 12. Functional assay using androstenone to activate OR7D4 variants utilizing the common sequence Different random arrangements of as the C-terminal domain.
[0967]
[0968]
[0969]
[0970]
[0971] The results of Table 12 were evaluated to derive the optimal amino acid residue at each variable position in the universal sequence (i.e., the residue that gave the lowest median EC2 for sequences containing that residue at a given position). Based on this analysis, the statistically optimal sequence within the universal sequence was And - since R at position 7 is at least as active as K (see Table 2 in Example 6) - the optimal sequence is also The data from Table 12 were then further analyzed for each sequence permutation with respect to how many residues differed from the statistically optimal sequence. As shown in Table 13, those sequences closest to SEQ ID NO: 307 had overall better activity. Thus, while all random permutations tested within the common sequence were active, the most active sequences clustered among those closest to the optimal sequence. Thus, for example, sequences that differed by 2-4 residues from SEQ ID NO: 307 were particularly active, with sensitivity increased >100-fold over wild type, with only one exception.
[0972] Table 13. Relationship between the distances between the random sequence permutations in Table A and the best SEQ ID NO: 307 and activity.
[0973]
[0974]
[0975] Example 17: Improved functional expression by different amino acids at position 7 of the common sequence
[0976] To test sequences that result in similarly improved functional expression Possible variants in the general sequence A different amino acid was introduced at position 7 (indicated by x) in , and the variant was fused to OR7C1 after TM7. Figure 2As shown, OR7C1 with the C-terminal sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) gave approximately 10-fold luciferase induction at 1 μM Ambermax and 20-fold induction at 3.1 μM Ambermax, while the wild type was inactive. The activation of the differently modified variants was then compared as described in Example 6. The results showed that position 7 of SEQ ID NO: 86 has high flexibility, with all 20 amino acids except proline showing superior activity compared to the wild type.
[0977] Table 14. Activation by Ambermax of OR7C1 variants with an optimized C-terminus containing a single base substitution at position 7 of the C-terminal motif RNKEVKDALKRLLKRKCC (SEQ ID NO: 86).
[0978]
[0979]
[0980] Example 18: Screening of all class II ORs with the same modified C-terminal domain using a single odorant Library
[0981] The same modified C-terminal domain was encoded in the vector pcDNA3.1(+). The sequences of all human class II ORs and all key gene variants of these ORs were synthesized to generate a library of all class II ORs with improved C-terminal domains. The complete list of class II ORs included in this library is given in Table 15:
[0982] Table 15. Class II OR library members with modified C-terminal domains
[0983]
[0984]
[0985]
[0986]
[0987]
[0988]
[0989]
[0990]
[0991] SEQ ID NOs 331-739 also include a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC), and a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes.
[0992] Parallel transfection experiments were performed with all of these plasmids (n=408), and cells expressing different ORs were stimulated with one given ligand at a time, at the maximal, clearly non-cytotoxic concentration tested. A total of 26 different ligands were tested on the full library of class II receptors. For 25 of these ligands, at least one receptor was identified with >4-fold luciferase induction relative to background, indicating that cognate receptors can be identified for most odorants with these improved OR libraries. In general, screening with this library resulted in a very good signal-to-noise ratio, clearly separating positive screening hits from inactive ligand-OR associations. As an example, screening of the full library with the ligand patchouliol is shown in Figure 21 Patchouli alcohol strongly induced OR14J1 and weakly induced OR11A1 and OR7A17. The full set of ORs identified from this deorphanization effort is shown in Table 16.
[0993] The selected ligand-OR pairs were then tested in confirmatory assays with dose response analysis. For all tested ligand-OR pairs from newly deorphaned receptors (n=31), a clear dose response was obtained when these screening hits were validated in confirmatory assays (e.g., see OR and SEQ ID NO: 19 listed in Example 19). Figure 22 (dose response of OR14J1 to patchouliol in
[15] ), indicating that this screen yields very reproducible ligand-OR associations using this improved OR library with a high signal-to-noise ratio. This can be compared to the prior art described in Mainland et al. (2015) Sci Data 2:150002, which employed OR expression with RTP1S and an N-terminal rho-tag, but employed wild-type OR sequences. In this screen, a comprehensive OR library (class I and class II; 511 clones, including key variants) was tested against 73 odorants. The initial screen yielded 1572 odorant / receptor pair hits (covering 394 ORs), but these associations could be validated in dose-response analyses for only 63 clones representing only 27 ORs, indicating significant noise in the primary screen due to poor or absent expression of the wild-type sequence.
[0994] Table 16. Ligands tested on the full library of ORs and cognate receptors identified for these ligands
[0995]
[0996] Example 19: Activation of ORs identified in a screen using a library of all human ORs with improved C-terminal domains Comparison of live-wild-type OR and improved OR
[0997] The C-terminal domain described in Example 18 was further analyzed. Improved functional expression of receptors of OR library of selected receptors. For the selected receptors, wild-type sequences were synthesized and cloned into pcDNA3.1 (+). Then, dose-response analysis was performed on the wild-type and modified sequences with cognate ligands. As done in Table 1, the induction threshold (2 times the concentration of luciferase induction) and EC50 (potency) and maximum induction (efficacy) of the analysis data were analyzed. As shown in Table 17, for the 13 ORs identified in the screening in Example C, the wild-type was completely inactive, while the OR with optimized C-terminal domain was activated by cognate ligands in the low micromolar range (all or no effect). This proves that these ORs can only be orphaned by this OR library with improved expression by C-terminal modification. For another 9 ORs tested, the improved sequence resulted in a significant reduction in the detection threshold (5.2–163 times) and an increase in efficacy (1.9–16.8 times).
[0998] Table 17. Contains optimized C-terminal compared to the wild-type sequence Improved functional expression of modified OR genes
[0999]
[1000]
[1001] na is not applicable because the wild type is inactive and cannot be calculated
[1002] Example 20: Optimized expression using SEQ ID NO: 221 compared to SEQ ID NO: 86
[1003] When the wild-type C-terminal domain was replaced with SEQ ID NO:86, all tested receptors in Table 1 gave improved functional expression. This functional expression was further improved when SEQ ID NO:221 was used to replace SEQ ID NO:86, as shown in Table 8 for OR7C1 and Table 9 for OR5B12. This additive improvement was further tested with OR10H5 (Table 18) and OR7A17 (Table 19). In both cases, SEQ ID NO:221 further improved functional expression compared to the already strongly improved functional expression relative to wild-type when SEQ ID NO:86 was used. The complete DNA sequences encoding the modified receptors are SEQ ID NO:684 for OR10H5 and SEQ ID NO:611 for OR7A17.
[1004] Table 18. Improved functional expression of OR10H5
[1005]
[1006] Table 19. Improved functional expression of OR7A17
[1007]
[1008] Example 21: Improved C-terminal domain for class I ORs
[1009] The modified C-terminal domain sequence RNKEVKDALKRLLKRKCC (SEQ ID NO: 86) and other variants falling within the general sequence RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1) and having additional amino acids at the C-terminus (CC) were tested on the class I receptor OR52A5 activated by 4-ethyloctanoic acid. As shown in Table 20, SEQ ID NO: 86 gave significantly improved expression compared to wild type. This activation was further improved by replacing amino acid positions 4-6 with the sequence QIR or by adding the terminal sequence RRR.
[1010] These two improvements were further combined and tested on OR52A5 and two other class I receptors. As shown in Table 21, all three class I receptors showed significant differences in the C-terminal sequence of the modified In the case of OR52A4, 4-ethyloctanoic acid, which is detected at low concentrations by the human nose, showed a high luciferase response at 0.19 μM with the receptor having the C-terminal modification, indicating that the modified class I receptor is very sensitive to the detection of this carboxylic acid.
[1011] Table 20. Optimized C-terminal domain of class I OR52A5
[1012]
[1013] Table 21. Optimized C-terminal domains of class I OR52E8, OR56A4, and OR52A5
[1014]
[1015]
[1016] Thus, additional improved C-terminal domain common sequences are RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO:822), RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR]RRR (SEQ ID NO:823), and RN[KR]QIRxA[LIV][KRH][KR][LI][LIF][KRG][KR][KR]RRR (SEQ ID NO:824).
[1017] Example 22: Screening of a library of all class I ORs with the same modified C-terminal domain
[1018] The same C-terminal domain was encoded in the vector pcDNA3.1(+). The sequences of all human class I ORs and all key gene variants of these ORs were synthesized to generate a library of all class I ORs with improved C-terminal domains. The complete list of class II ORs included in this library is given in Table 22:
[1019] Table 22. Class I OR library members
[1020]
[1021]
[1022]
[1023] SEQ ID NOs 742-819 also include a 5' BamHI restriction site (GGATCC) and a Kozak sequence (GCCACC) and a 3' NotI restriction site (GCGGCCGC) for cloning and expression purposes.
[1024] Parallel transfection experiments were performed with all of these plasmids (n=77, SEQ ID NO: 760, OR51F1 was excluded due to poor plasmid yield), and cells expressing different ORs were stimulated with a given ligand one at a time. A total of 10 different carboxylic acid ligands were tested on the full library of class I receptors. For 8 of these ligands, at least one receptor was identified with >3-fold luciferase induction over background (Table 23), demonstrating that these improved OR libraries can be used to identify cognate receptors, particularly for carboxylic acids.
[1025] As an example, the results of screening the full library with 3-methyl-2-hexenoic acid are shown in Figure 233-Methyl-2-hexenoic acid is an important component of human sweat (Natsch et al., A specific bacterial aminoacylase cleaves odorant precursors secreted in the human axilla. The Journal of biological chemistry 2003; 278(8): 5718–5727). 3-Methyl-2-hexenoic acid strongly induces the OR51B2 (C120R, L134F, C209S) variant but not the wild-type variant of OR51B2. This contradicts the report of Li et al. (PLoS Genet. 2022 Feb 3; 18(2): e1009564), who reported that the wild-type OR51B2 but not the (C120R, L134F, C209S) variant was activated by 3-methyl-2-hexenoic acid. Therefore, contrary to the expression data in this publication, the (C120R, L134F, C209S) variant of OR51B2 is a preferred target for screening human sweat odor antagonists. In addition, 3-methyl-2-hexenoic acid also activated two tested variants of OR52K1, and is therefore a further new target for screening odor antagonists. In addition to 3-methyl-2-hexenoic acid, 4-methyl-3-hexenoic acid also activated the (C120R, L134F, C209S) variant of OR51B2. This acid also activated two tested variants of OR51B5. 4-methyl-3-hexenoic acid is an important contributor to typical odor substances in laundry (Kubota et al., Appl Environ Microbiol. 2012; 78(9): 3317-24). Therefore, screening for additional antagonists of OR51B5 can be used to target this specific odor substance. Screening the complete library of class I ORs for 3-methyl-3-hydroxyhexanoic acid, the most dominant carboxylic acid in human sweat odor, yielded only one hit, OR52E8, confirming the above results that OR52E8 variants with improved expression through optimized C-terminal domains are key targets for screening antagonists of this acid. In addition to the ORs for these short-chain odorant acids, OR52K1, OR52K1(Q52R), OR56A1, OR56A3, OR56A3(M51T), and OR56A4 were activated by the C8-C11 chain acids tested in this screen.
[1026] Table 23. Ligands tested on the full library of class I ORs and cognate receptors identified for these ligands
[1027]
[1028]
[1029] Example 23: Screening and matching complexes using a library of all class II ORs with the same modified C-terminal domain Fragrance (i.e. mixture of complex odorous substances)
[1030] The human class II OR library of Example 18 was further tested with three fragrance oils, i.e., complex mixtures of ≥24 individual fragrance ingredients, which were mixed in specific ratios to produce the desired overall odor. "Fragrance A" and "Fragrance B" are independent fragrance preparations with different overall odor impressions, while "Modified Fragrance A" shares 66% of its ingredients with "Fragrance A" and has been modified to maintain the overall odor of Fragrance A, but replacing 33% of the individual ingredients. Overall, 18 ORs and OR variants were activated by at least one of these oils with at least 2-fold luciferase induction. The induction of these activated OR subsets by these three oils is shown in Figure 24 . As is apparent from the black and grey bars, "Fragrance A" and "Modified Fragrance A" elicit similar OR activation patterns, whereas "Fragrance B" differs markedly, particularly with respect to the activation of OR10G7, OR11G2, OR1N2, OR2AJ1, OR2J2, OR3A3, OR5AN1, OR5B12, and OR5P3. Thus, Fragrance B has a different OR activation "fingerprint" reflecting its different olfactory profile. The differences / similarity of different oils can also be quantified using a distance metric. For example, the distance between two compositions i and j can be calculated, for example, according to the following formula:
[1031]
[1032] for Figure 24 Among the 18 ORs, the distance between "spice A" and "modified spice A" is 2.2, while the distance between "spice A" and "spice B" is 6.4, and the distance between "modified spice A" and "spice B" is 5.5, indicating the similarity between "spice A" and "modified spice A" and the difference between them and "spice B".
[1033] These results indicate that the method of screening sensitive libraries of ORs with modified C-terminal domains using complex mixtures can be used to measure the similarity of complex odorant mixtures (such as fragrances or flavorings) and give an objective representation of the olfactory profile of such mixtures.
[1034] Due to regulatory pressure, secret fragrance ingredients are banned in some countries. Therefore, it's desirable to design a fragrance that closely resembles the scent of consumers but replaces the banned ingredient. Traditionally, a single ingredient has been used to attempt to replace the anosmia after removing the banned ingredient. However, because a single ingredient may activate multiple ORs, and multiple ingredients often have to be removed, it's crucial to replace the OR activation pattern of the removed ingredient and reconstruct the overall OR activation pattern of the original fragrance. This can be accomplished with one or more ingredients.
[1035] This can be achieved using the method of screening single ingredients for OR activation on a full OR library with a modified C-terminal domain, as shown in Examples 18 (Class II) and 22 (Class I), and recording their OR activation patterns to generate a database of activation patterns for single ingredients. The original fragrance oil and the original fragrance oil with the ingredient under investigation removed are then screened on the full OR library as shown above. After adding the selected alternative ingredient, the resulting fragrance can then be validated by testing it again on the full OR library as shown above, and the similarity to the original fragrance can be calculated in an objective manner.
[1036] Example 24: Detection of homologous odorants in complex mixtures or reaction mixtures of low purity
[1037] Traditionally, fragrance ingredients and experimental fragrance ingredients (research samples) have required a high degree of purification in order to be evaluated by perfumers – this is because all approximately 400 ORs of the human nose are functional simultaneously, and any odorant impurities will therefore affect the overall olfactory impression of the sample, and it is therefore difficult for humans to judge samples of limited purity. On the other hand, assays with a single or a few expressed receptors focus on a specific odor profile and can therefore, in principle, detect odorants with this specific odor profile against a complex background. However, in the case of poorly expressed receptors, the matrix will interfere with the assay, as matrix components will quickly reach cytotoxic concentrations when the active ingredient is present at low concentrations in the direction of the target odor being searched for.
[1038] To investigate the sensitivity of the modified olfactory receptors of the present disclosure to detect odorants in a complex odorant matrix, the ligand ambroxan was spiked into a composite essential oil (i.e., geranium oil from Egypt) containing 13 different components at greater than 1% as an example of a complex background matrix of strong odorant materials. The levels of ambroxan spiked in were 0%, 0.1%, 0.316%, 1%, 3.16%, and 10%. These mixtures were tested as described in Example 20 using OR7A17 (containing the DNA sequence encoding the modified receptor as SEQ ID NO: 611) having a C-terminal domain of SEQ ID 221. Figure 25 As shown, the spiked oil significantly induced luciferase expression above background. Thus, for a potent ligand like ambroxan, a sensitive assay can detect concentrations as low as at least 0.1% of the target component in a complex matrix.
[1039] Example 25: Screening of an OR library using a mixture of odorants with a specific odor description
[1040] A single odorant can trigger the activation of multiple ORs, as demonstrated in Example 18 with ambroxan, which activates specific receptors OR7E24 and OR7A17. On the other hand, multiple odorants can be perceived and described by a common odor descriptor due to their collective activation of a given set of ORs. To trigger a specific odor sensation, a specific set of ORs may need to be activated. This specific set of ORs can be identified by mixing several odorants with a given odor description. This mixture can then be screened against a comprehensive library of ORs. The identified set of ORs can then be further used to screen for that specific odor description. For example, to identify a representative OR group for the odor description "fruity esters," a mixture of the following compounds was prepared: 2-methylethyl valeric acid; 3-methylethyl butyric acid; 2-propenyl acetic acid; ethyl hexanoic acid; 2-propenyl (3-methylbutyloxy) acetate; 2-propenyl (cyclohexyloxy) acetate; 2-methyl (3Z / E)-3-hexenyl butyric acid; ethyl cyclohexanecarboxylate; and 3-phenyl oxadiazine carboxylate. This mixture was screened against a full library of ORs as described in Example 18. Using this approach, a specific OR group activated by the ligands for this particular odor description was identified: thus, OR11G2, OR11G2(I65N, V82I), OR1D2, OR2AK2(S84N), and OR2L5 were identified as the OR group most strongly activated by the mixture of fruity esters.
[1041] Table 24. ORs identified by screening Class II ORs with a mixture of fruity esters.
[1042]
[1043] Similarly, a mixture of odorants with a fruity-lactone descriptor was prepared, containing equal amounts of δ-dodecalactone, γ-heptanolactone, γ-octanolactone, γ-nonanolactone, γ-undecalactone, γ-decalactone, γ-dodecalactone, δ-decalactone, methyl tuberose, and schonolacetone. This mixture was used to screen the entire library of class II ORs, revealing that a specific set of ORs, namely OR10A3, OR10A6 (A117V, V140G, L287P), OR10J1 (M51I, I92M), OR1D2, and OR2J2, were activated.
[1044] Table 25. ORs identified by screening class II ORs with a mixture of fruity lactones
[1045]
[1046] In subsequent deconvolution experiments, single components were tested in dose-response assays on all identified ORs. Table 26 lists the concentrations that resulted in 2-fold luciferase induction (10-fold in the case of OR10A3). These data show that all identified ORs were activated by the single components of the mixture, with some differences in specificity. Thus, OR2AP1 and OR2J2 were particularly strongly activated by long-chain γ-dodecalactone, while OR10A3 was very sensitive to several longer-chain lactones. Overall, γ-undecalactone was the most potent ligand for four of the five identified ORs, consistent with the fact that among lactones, γ-undecalactone has the lowest olfactory detection threshold in vivo. Of the five ORs, OR10A3 was the most sensitive. Thus, for the most potent ligand, γ-undecalactone, a 10.2-fold activation was observed at 0.31 μM.
[1047] Table 26. ORs identified by screening of class II ORs using a mixture of fruity lactones tested with individual lactones in a dose-response analysis.
[1048]
[1049]
[1050] 1) Assume that the arbitrary molecular weight of the lactone mixture is 200
[1051] OR10A3 is a very sensitive receptor with high efficacy, therefore EC10, the concentration for 10-fold OR activation, is indicated here.
[1052] Example 26: Screening of odorants using cells expressing multiple receptors
[1053] In order to screen for a specific odor description, some or all members of a given set of ORs identified as specific for a given odor description as shown in Example 25 can be combined in the screening for new odorants or new odorant mixtures. The screening of different ORs can be performed sequentially with single ORs of the OR group or in parallel assays. In addition, some or all ORs of the OR group specific for the odor description can also be co-expressed in a single cell line. Therefore, in order to screen for costus amber molecules, cells are transfected separately or simultaneously with two plasmids encoding optimized OR7A17 and OR7C1 (both having the C-terminal domain of SEQ ID NO: 221). The cells are then stimulated with the ligands Iso E super (which is a specific ligand for OR7A17) and Ambermax (which is a specific ligand for OR7C1). As Figure 26As shown, cells expressing only OR7A17 responded only to IsoE super, while cells expressing OR7C1 responded only to Ambermax. Cells expressing both receptors became functional sensors for the amber-woody odor profile and detected both ligands down to low concentrations.
[1054] Example 27. In (S,E)-10-hydroxy-4,8-dimethyldec-4-enal or (R,E)-10-hydroxy-4,8-dimethyl Identification of a more potent ligand for OR10J5 in 4-dec-enal
[1055] Activation of OR10J5 having the C-terminal domain of SEQ ID NO: 86 by (S,E)-10-hydroxy-4,8-dimethyldec-4-enal or (R,E)-10-hydroxy-4,8-dimethyldec-4-enal was measured in a dose response analysis.
[1056] The potency of the two OR10J5 ligands tested is expressed as EC20% values, which are the concentrations that result in a 20% increase in luciferase activity relative to the positive control (100 μM, Mahonial; (4E)-9-hydroxy-5,9-dimethyl-4-decenal). The EC20% value for (S,E)-10-hydroxy-4,8-dimethyldec-4-enal is 10.1 μM, while the EC20% value for (R,E)-10-hydroxy-4,8-dimethyldec-4-enal is 24.4 μM. Due to the difference in EC20%, lower concentrations of the (S,E)-isomer are required to activate the receptor compared to the concentration required for the (R,E)-isomer. Figure 28 illustrates the responses of the two test compounds. This example further illustrates the usefulness of the improved assay...
Claims
1. An olfactory receptor protein, wherein the protein has a modified C-terminal domain comprising the amino acid sequence motif RN[KR][EDQ][VMIL][KR]xA[LIV][KRH][KR][LI][LIF][KRG][KR][KR] (SEQ ID NO: 1). 2 . The olfactory receptor protein according to claim 1 , wherein the modified C-terminal domain is fused to the seventh transmembrane helix of the protein.
3. The olfactory receptor protein according to claim 1 or 2, wherein the protein is a class I or class II olfactory receptor with a modified C-terminal domain, preferably a human, dog or cat class I or class II olfactory receptor with a modified C-terminal domain, more preferably a human class I or class II olfactory receptor with a modified C-terminal domain.
4. The olfactory receptor protein of claim 3, wherein the class II receptor is selected from the group consisting of OR7C1, OR9Q2, OR8K3, OR10J5, OR1C1, OR7D4, OR2T4, OR5B12, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2, OR2C1, OR2T11, OR2M2, OR4S2, OR2V1, OR5P3, OR6P1, OR2L2, OR10G7, OR5AN1, OR5V1, OR2L3, OR2AG2, OR7A5, OR7E24, OR7A10, OR10H2, OR10H1, OR10D3, OR10D4, OR10H5 OR7A17, OR7C1, OR2A25, OR7E24, OR10H1, OR10K1, OR2AG2, OR10H2, OR10H5, OR10D3, OR14J1, OR7A10, OR2L5, OR2M2, and OR5A2.
5. The olfactory receptor protein according to claim 4, wherein the class I receptor is selected from the group consisting of OR52A5, OR52E8, OR56A4, OR51B2, OR52K1, OR56A1, OR51B5, OR56A3 and OR51L1.
6. The olfactory receptor protein according to any one of claims 1 to 5, wherein the sequence motif is RN[KR]E[VMI][KR]xA[LIV][KR][KR]L[LIF][KR][KR][KR] (SEQ ID NO: 5).
7. The olfactory receptor protein according to any one of claims 1 to 6, wherein x is not proline.
8. The olfactory receptor protein according to any one of claims 1 to 7, wherein x is not proline or tryptophan.
9. The olfactory receptor protein according to any one of claims 1 to 8, wherein x is selected from D, K, R, E, N, V, A, Q or G, preferably wherein x is selected from D, K, R, E, N, V, A or Q.
10. The olfactory receptor protein according to any one of claims 1 to 9, wherein the amino acid sequence motif comprises 1 to 6 additional C-terminal amino acid residues, optionally wherein: - the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, W, M and N, preferably the first additional amino acid residue is selected from the group consisting of C, R, K, E, G, H, F, P, Y, more preferably the first additional amino acid residue is C, R or K, most preferably C; - the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T, Y and Q, preferably the second additional amino acid residue is selected from the group consisting of C, R, K, N, G, I, L, F, P, T and Y, more preferably the second additional amino acid residue is C, R or K, most preferably C or R; - the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S, G, H and N, preferably the third additional amino acid residue is selected from the group consisting of R, K, C, L, F, M, Y, A, P, S and G, more preferably the third additional amino acid residue is R or K; - the fourth, fifth and sixth additional amino acid residues are selected from K and R, Preferably, wherein the amino acid sequence motif comprises an additional C-terminal amino acid residue selected from the group consisting of: CC, CCR, CCRR (SEQ ID NO: 161), CCRRR (SEQ ID NO: 163), CCRRRR (SEQ ID NO: 224), CR, CRR, CRRR (SEQ ID NO: 159), CRKK (SEQ ID NO: 160), CRRRR (SEQ ID NO: 162), CRRRRR (SEQ ID NO: 164) and CRRRKK (SEQ ID NO: 165).
11. The olfactory receptor protein according to any one of claims 1 to 10, wherein the sequence motif is selected from the group consisting of SEQ ID NO: 1, 5-75, 86-130, 133-147, 149-151, 154, 156-158, 166, 167, 198, 219-221, 254-312, 319-326, 328-331, 740-741, 820-890.
12. A nucleic acid molecule comprising a nucleotide sequence encoding the olfactory receptor protein according to any one of claims 1 to 11, preferably further comprising a promoter sequence, more preferably a constitutive promoter sequence.
13. An expression vector comprising the nucleic acid molecule according to claim 12, preferably wherein the expression vector is a plasmid.
14. A recombinant host cell, preferably a HEK293 or HEK293T cell, comprising the nucleic acid molecule according to claim 12 or the expression vector according to claim 13, preferably wherein the cell expresses the olfactory receptor protein according to any one of claims 1 to 11.
15. The recombinant host cell of claim 14, wherein the cell further expresses one or more olfactory receptor accessory proteins.
16. A library comprising a diverse repertoire of the olfactory receptor protein according to any one of claims 1 to 11, the nucleic acid molecule according to claim 12, the expression vector according to claim 13, or the recombinant host cell according to claim 14 or 15.
17. The library of claim 16, wherein the olfactory receptor proteins, the olfactory receptor proteins encoded by the nucleic acid molecules or expression vectors, or the diverse repertoire of olfactory receptor proteins expressed by the recombinant host cells share the same modified C-terminal domain.
18. A method for identifying an olfactory receptor ligand, the method comprising: a) providing the olfactory receptor protein according to any one of claims 1 to 11 or a recombinant host cell expressing the olfactory receptor protein according to claim 14 or 15; b) contacting the receptor or recombinant host cell with a test compound or composition; and c) detecting activation of the olfactory receptor, Preferably, wherein the olfactory receptor is selected from the group consisting of: OR7C1, OR8K3 (preferably OR8K3 (L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2 (W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2 (V259L)), OR10G7 (preferably OR10G7 (T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2 (Y28C)), OR7A5, OR7E24 (or OR7E24 (P242S)), OR7A10, OR10H2, OR10 H1, OR10D3, OR1D2, OR2A5, OR2A25 (OR2A25(S75N, A209P)), OR11G2 (or OR11G2(I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3(S73G)), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2 and OR2AG2 (preferably OR2AG2(Y28C)).
19. A method for identifying an olfactory receptor enhancer or antagonist, the method comprising: a) providing the olfactory receptor protein according to any one of claims 1 to 11 or the cell expressing the olfactory receptor protein according to claim 14 or 15; b) contacting the receptor or recombinant host cell with a cognate ligand and a test compound or composition; and c) detecting an increase or decrease in activation of the olfactory receptor compared to a ligand-only control, Preferably, wherein the olfactory receptor is selected from the group consisting of: OR7C1, OR8K3 (preferably OR8K3 (L122R)), OR10J5, OR7A17, OR10H5, OR5A1, OR5A2, OR1N2 (preferably OR1N2 (W23R, V230G, T287M)), OR2M2, OR2V1, OR5P3, OR6P1, OR2L2 (or OR2L2 (V259L)), OR10G7 (preferably OR10G7 (T5S)), OR5AN1, OR5V1, OR2L3, OR2AG2 (preferably OR2AG2 (Y28C)), OR7A5, OR7E24 (or OR7E24 (P242S)), OR7A10, OR10H2, OR10 H1, OR10D3, OR1D2, OR2A5, OR2A25 (OR2A25(S75N, A209P)), OR11G2 (or OR11G2(I65N, V82I)), OR14J1, OR5M3, OR8D1, OR10G3 (preferably OR10G3(S73G)), OR10G9, OR2L5, OR8H1, OR10K1, OR11A1, OR2AK2 (preferably OR2AK2(S84N)), OR10A3, OR10A6 (preferably OR10A6(A117V, V140G, L287P)), OR10J1 (preferably OR10J1(M51I, I92M)), OR2J2 and OR2AG2 (preferably OR2AG2(Y28C)).
20. The method according to claim 19, wherein the method is used to identify an olfactory receptor antagonist, and wherein the olfactory receptor is selected from the group consisting of OR52A5, OR52E8, OR56A1, OR56A3, OR56A4, OR52K1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), OR51B5, OR9Q2, OR7D4, OR2T4, OR2C1, OR2T11, OR2M2, OR2V1, OR5V1 and OR4S2, preferably selected from the group consisting of OR2M2, OR2V1, OR51B2 (preferably OR51B2 (C120R, L134F, C209S)) and OR5V1, more preferably OR2M2 or OR2V1.
21. The method according to claim 19, wherein the method is used to identify olfactory receptor antagonists, wherein the olfactory receptor is OR2M2 or OR2V1, and wherein step b) further comprises contacting the receptor or the recombinant host cell with a copper salt, preferably wherein the cognate ligand is selected from the group consisting of 3-methyl-3-sulfanyl-hexanol, 2-mercapto-2-methyl-pentanol and 4-methoxy-2-methylpentane-2-thiol.
22. The method according to claim 19, wherein the method is used to identify an olfactory receptor antagonist, wherein the olfactory receptor is OR51B2 (preferably OR51B2 (C120R, L134F, C209S)), preferably wherein the cognate ligand is 3-methyl-2-hexenoic acid.
23. The method of claim 19, wherein the method is used to identify an olfactory receptor antagonist, wherein the olfactory receptor is OR5V1, preferably wherein the cognate ligand is 2,4,6-trichloroanisole.
24. A method for identifying an olfactory receptor capable of binding a target ligand, the method comprising: a) providing a library according to claim 16 or 17; b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins from the library; c) contacting a diverse repertoire of olfactory receptor proteins or recombinant host cells expressing olfactory receptor proteins with the target ligand; and d) identifying an olfactory receptor activated by said target ligand.
25. A method for generating an objective representation of the olfactory properties of a test compound or composition, the method comprising: a) providing a library according to claim 16 or 17; b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with the test compound or composition; and d) detecting activation of each of said olfactory receptor proteins.
26. A method for evaluating differences or similarities between two or more test compounds or compositions, the method comprising: a) providing a library according to claim 16 or 17; b) optionally, obtaining a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins from the library; c) contacting a diverse repertoire of olfactory receptor proteins or cells expressing olfactory receptor proteins with each of the two or more test compounds or compositions; d) detecting activation of each olfactory receptor protein by each of the two or more test compounds or compositions; as well as e) comparing the activated olfactory receptor protein between each of the two or more test compounds or compositions.
Citation Information
Patent Citations
Modulators of odorant receptors
WO2006002161A2
Compositions and methods for increasing the expression and signalling of proteins on cell surfaces
WO2014037800A2
Seed treatment with natamycin
WO2019011630A1
Olfactory receptor involved in the perception of MUSK fragrance and the use thereof
WO2019110630A1