Method for predicting the presence or absence of odor characteristics or olfactory receptor activation characteristics in a substance

By employing stereochemical similarity and machine learning models based on olfactory receptor activation data, the method enhances the accuracy of predicting aroma characteristics and olfactory receptor activation, addressing the limitations of existing techniques.

JP7735996B2Active Publication Date: 2025-09-09AJINOMOTO CO INC

Patent Information

Application Number
JP2022512182
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2021-03-29
Publication Date
2025-09-09
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Existing methods for predicting aroma characteristics and olfactory receptor activation in substances face challenges such as the need for human expertise, low throughput, and failure to account for multiple conformations in molecular structures, leading to inaccurate predictions.

Method used

Predicting aroma characteristics and olfactory receptor activation by determining the maximum stereochemical similarity between a test substance and a control substance, and using machine learning to generate models based on olfactory receptor activation data.

Benefits of technology

Improves prediction accuracy by considering stereochemical structures and using machine learning, enabling effective screening and design of substances with desired aroma properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007735996000005
    Figure 0007735996000005
  • Figure 0007735996000006
    Figure 0007735996000006
  • Figure 0007735996000007
    Figure 0007735996000007
Patent Text Reader

Abstract

Provided is a technique for predicting the presence or absence of aroma properties or olfactory receptor activation properties in a substance. In this technique, the presence or absence of the aforesaid target properties in a test substance is predicted on the basis of the maximum similarity in three-dimensional chemical structure between the test substance and a control substance.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] In one aspect, the present invention relates to a technique for predicting the presence or absence of an aroma characteristic or olfactory receptor activation characteristic in a substance. In another aspect, the present invention relates to a technique for predicting the presence or absence of components such as aroma characteristics or molecular structure in a substance. In another aspect, the present invention relates to a technique for predicting the compatibility of a substance with an aroma characteristic. [Background technology]

[0002] Aroma is an important factor that influences the palatability of foods, cosmetics, etc. Therefore, techniques for screening aroma components necessary to reproduce a desired aroma and techniques for reproducing a desired aroma by combining aroma components are industrially important techniques for developing foods, cosmetics, etc.

[0003] Conventionally, screening of aroma compounds has been carried out by humans evaluating the aroma of test substances through sensory testing, but sensory testing has problems such as the need to train experts who can evaluate aromas and low throughput.

[0004] In mammals such as humans, odors are perceived when molecules of odor components bind to olfactory receptors on olfactory nerve cells present in the olfactory epithelium in the upper part of the nasal cavity, and the receptor's response to the molecules is transmitted to the central nervous system. In recent years, methods have been reported for screening substances that exhibit a desired odor using the response of olfactory receptors as an indicator (Patent Document 1, etc.).

[0005] In recent years, with the advancement of machine learning technology, research has been conducted into predicting aroma characteristics directly from the structure of a compound (Non-Patent Documents 1-3). The technical key points for improving the accuracy of prediction models can be broadly divided into three areas: quantification of molecular structure, prediction algorithms, and the quantity and quality of data used for learning. Among these, existing methods for quantifying molecular structure include methods that calculate physicochemical features from molecular structure (e.g., Dragon and EPI Suite), methods that create molecular fingerprints (e.g., MACCS Keys and Morgan fingerprints) that represent the presence or absence of partial molecular structures as bits (1 / 0) and then calculate the structural similarity between molecules, and methods that treat molecular structures as graphs (networks) or images and forcibly quantify them using neural network technology. In all of these methods, one compound is assumed to have one structure, and information about multiple conformations is ignored. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent Publication No. 2019-037197 [Non-patent literature]

[0007] [Non-Patent Document 1] Kobi Snitz et. al., Predicting Odor Perceptual Similarity from Odor Structure. PLoS Comput Biol 9(9): e1003184, September 2013. [Non-patent document 2] Andreas Keller et. al., Predicting human olfactory perception from chemical features of odor molecules. Science, 355(6327):820-826, February 2017. [Non-patent document 3] Benjamin Sanchez-Lengeling et. al., Machine Learning for Scent: Learning Generalizable Perceptual Representations of Small Molecules. arXiv:1910.10685v1, October 2019. Summary of the Invention [Problem to be solved by the invention]

[0008] In one aspect, the present invention aims to provide a technique for predicting the presence or absence of an aroma characteristic or olfactory receptor activation characteristic in a substance. In another aspect, the present invention aims to provide a technique for predicting the presence or absence of components such as aroma characteristics or molecular structure in a substance. In another aspect, the present invention aims to provide a technique for predicting the compatibility of a substance with an aroma characteristic. [Means for solving the problem]

[0009] The present inventors have found that it is possible to predict whether a test substance has the above-mentioned target properties based on the highest similarity in stereochemical structure between the test substance and a control substance, and have thus completed one aspect of the present invention.

[0010] That is, in one aspect, the present invention can be exemplified as follows. [1] 1. A method for predicting the presence or absence of a property of interest for a test substance, comprising: predicting the presence or absence of the desired property for the test substance based on the maximum similarity of the stereochemical structure between the test substance and the reference substance; Including, The method, wherein the property is an odor property or an olfactory receptor activation property. [2] The method, wherein the control substance comprises a positive control for the property of interest. [3] The method, wherein the control substance is one type of substance. [4] The method, wherein the control substance is a combination of two or more substances. [5] the control material comprises a positive control for the property of interest; The method predicts that the test substance has the desired property when the maximum similarity of the stereochemical structure between the test substance and the positive control is high. [6] The method, wherein the prediction comprises clustering the test substance and the control substance based on the greatest similarity of stereochemical structure between the test substance and the control substance. [7] the control material comprises a positive control for the property of interest; The method predicts that the test substance has the desired property if the test substance is clustered into a cluster that includes the positive control. [8] The method further comprising, before the prediction, calculating the maximum similarity. [9] 1. A method for screening for a substance having a desired property, comprising: predicting the presence or absence of the property of interest for a test substance using the method; and a step of selecting the test substance predicted to have the desired property as a substance having the desired property; Including, The method, wherein the property is an odor property or an olfactory receptor activation property.

[10] The method further comprises a step of confirming the presence or absence of the desired property for a test substance predicted to have the desired property.

[11] The method, wherein the maximum similarity is used in the prediction in combination with a structural similarity between the test substance and the control substance other than the maximum similarity.

[12] 1. A method for designing a material with desired properties, comprising: a step of designing a substance to be designed based on the maximum similarity of the stereochemical structure between the substance to be designed and a reference substance; Including, The method, wherein the property is an odor property or an olfactory receptor activation property.

[13] the control material comprises a positive control for the property of interest; The design is performed such that the substance to be designed is clustered into a cluster that includes the positive control; The method, wherein the clustering comprises a step of clustering the substance to be designed and the control substance based on the maximum similarity of the stereochemical structures between the substance to be designed and the control substance.

[0011] Furthermore, the present inventors have discovered that a model for predicting the presence or absence of an aroma characteristic or molecular structure in a substance can be generated by machine learning, thereby completing another aspect of the present invention.

[0012] That is, in another aspect, the present invention can be exemplified as follows. [1] 1. A method for producing a model that predicts the presence or absence of a constituent of interest in a test substance, comprising: The model includes a decision tree that outputs a classification result regarding the presence or absence of the component of interest in the test substance based on test olfactory receptor activation data of the test substance; the method includes generating the decision tree by machine learning; the constituent element is an aroma characteristic or a molecular structure, The method, wherein the test olfactory receptor activation data is data regarding activation of a test olfactory receptor by the test substance. [2] The machine learning is performed using a dataset including control substance component data and control olfactory receptor activation data; the component data is data regarding the component of interest in the control material; the control olfactory receptor activation data is data regarding activation of a control olfactory receptor by the control substance; the control substance is a combination of two or more substances, including a positive control and a negative control; The above method, wherein the control olfactory receptor is a combination of two or more olfactory receptors including the test olfactory receptor. [3] the component data is data indicating the presence or absence of the target component in the control substance; The method, wherein the control olfactory receptor activation data is data indicating the degree of activation of the control olfactory receptor by the control substance. [4] The method, wherein the machine learning is performed using the component data as a dependent variable and the control olfactory receptor activation data as an explanatory variable. [5] The method, wherein the machine learning is performed by CART. [6] The method, wherein the machine learning is performed by ensemble learning. [7] The method, wherein the control substance is a combination of 500 or more substances. [8] The method, wherein 50% or more of the total number of control substances are selected from compounds listed in The Good Scents Company. [9] The above method, wherein the test olfactory receptor is one type of olfactory receptor or a combination of two or more types of olfactory receptors.

[10] The above method, wherein the control olfactory receptor is a combination of 300 or more types of olfactory receptors.

[11] 50% or more of the total number of the control olfactory receptors is OR1A1, OR1A2, OR1B1, OR1C1, OR1D2, OR1D5, OR1E1, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1R1P, OR1S1, OR2A1, OR2A2, OR2A4, OR2A5, OR2A12, OR2A14, OR2A25, OR2AE1, OR2AG1, OR2AG2, OR2AJ1P, OR2 AK2, OR2AP1, OR2AT4, OR2B2, OR2B3, OR2B6, OR2B11, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2G2, OR2G3, OR2G6, OR2H1, OR2H2, OR2J2, OR2J3, OR2K2, OR2 L2, OR2L8, OR2L13, OR2M2, OR2M4, OR2M7, OR2S2, OR2T1, OR2T2, OR2T5, OR2T6, OR2T8, OR2T10, OR2T11, OR2T27, OR2T34, OR2V2, OR2W1, OR2W3, OR2Y1, OR2 Z1, OR3A1, OR3A2, OR3A3, OR3A4, OR4A5, OR4A15, OR4A16, OR4A47, OR4B1, OR4C3, OR4C5, OR4C6, OR4C11, OR4C12, OR4C13, OR4C15, OR4C16, OR4C46, OR4D 1, OR4D2, OR4D5, OR4D6, OR4D9, OR4D10, OR4D11, OR4E2, OR4F3, OR4F5, OR4F6, OR4F14P, OR4F15, OR4G11P, OR4H12P, OR4K1, OR4K2, OR4K5, OR4K13, OR4K1 4, OR4K15, OR4K17, OR4L1, OR4M1, OR4N2, OR4N4, OR4N5, OR4P4, OR4Q3, OR4S1, OR4S2, OR4X1, OR4X2, OR5A1, OR5A2, OR5AC2, OR5AK2, OR5AK3P, OR5AN1, O R5AP2, OR5AR1, OR5AS1, OR5AU1, OR5B2, OR5B3, OR5B12, OR5B17, OR5B21, OR5C1, OR5D13, OR5D14, OR5D16, OR5D18, OR5F1, OR5H1, OR5H2, OR5H6, OR5H14,OR5I1、OR5J2、OR5K1、OR5K3、OR5K4、OR5L2、OR5M3、OR5M8、OR5M9、OR5M10、OR5M11、OR5P3、OR5R1、OR5T1、OR5T2、OR5T3、OR5V1、OR5W2、OR6A2、OR6B1、OR6B2、OR6C1、OR6C2、OR6C3、OR6C4、OR6C6、OR6C65、OR6C66P、OR6C68、OR6C70、OR6C74、OR6C75、OR6C76、OR6F1、OR6J1、OR6K2、OR6K3、OR6K6、OR6M1、OR6N1、OR6N2、OR6P1、OR6Q1、OR6S1、OR6T1、OR6V1、OR6X1、OR6Y1、OR7A3P、OR7A5、OR7A10、OR7A17、OR7C1、OR7C2、OR7D2、OR7D4、OR7E24、OR7G1、OR7G2、OR7G3、OR8A1、OR8B3、OR8B4、OR8B8、OR8B12、OR8D1、OR8D2、OR8D4、OR8G2、OR8G5、OR8H3、OR8I2、OR8J1、OR8J3、OR8K1、OR8K3、OR8K5、OR8S1、OR8U1、OR9A4、OR9G1、OR9G4、OR9I1、OR9K2、OR9Q1、OR9Q2、OR10A3、OR10A4、OR10A5、OR10A6、OR10A7、OR10AD1、OR10AG1、OR10C1、OR10D3、OR10D4P、OR10G2、OR10G3、OR10G4、OR10G6、OR10G7、OR10G9、OR10H2、OR10H4、OR10J1、OR10J3、OR10J5、OR10K1、OR10K2、OR10P1、OR10Q1、OR10R2、OR10S1、OR10T2、OR10V1、OR10W1、OR10X1、OR10Z1、OR11A1、OR11G2、OR11H4、OR11H6、OR11H12、OR11L1、OR12D2、OR12D3、OR13A1、OR13C2、OR13C3、OR13C4、OR13C8、OR13D1、OR13F1、OR13G1、OR13H1、OR13J1、OR14A2、OR14A16、OR14C36、OR14I1、OR14J1、OR14K1、OR14L1P、OR51A1P、OR51A4、OR51A7、OR51B2、OR51B4、OR51B5、OR51B6、OR51D1、OR51E1、OR51E2, OR51F1, OR51F2, OR51F5P, OR51G1, OR51G2, OR51H1, OR51I1, OR51I2, OR51L1, OR51M1, OR51Q1, O R51S1, OR51T1, OR51V1, OR52A1, OR52A4, OR52A5, OR52B2, OR52B4, OR52B6, OR52D1, OR52E2, OR52E4, OR52 E5, OR52E8, OR52H1, OR52I2, OR52J3, OR52K2, OR52L2P, OR52M1, OR52N1, OR52N2, OR52N4, OR52N5, OR52P2P, OR52R1, OR52W1, OR52Z1P, OR56A1, OR56A3, OR56A4, OR56A5, OR56B1, OR56B2P, OR56B4.

[12] The method, wherein the control olfactory receptor is a human olfactory receptor.

[13] A model produced by the method.

[14] 1. A method for predicting the presence or absence of a component of interest in a test substance, comprising: predicting the presence or absence of the component of interest for the test substance based on the test olfactory receptor activation data of the test substance and the model. Including, The method, wherein the component is an aroma property or a molecular structure.

[15] 1. A method for screening a substance having a component of interest, comprising: predicting the presence or absence of the component of interest for a test substance based on the test olfactory receptor activation data of the test substance and the model; and a step of selecting the test substance predicted to have the component of interest as a substance having the component of interest; Including, The method, wherein the component is an aroma property or a molecular structure.

[16] The method predicts that the test substance has the component of interest if the test substance is classified into a leaf node where the ratio of positive controls is 50% or more.

[17] The method further comprises the step of confirming the presence or absence of the component of interest in a test substance predicted to contain the component of interest.

[0013] Furthermore, the present inventors have discovered that a model that predicts the degree of compatibility with the aroma characteristics of a substance can be generated by machine learning, thereby completing another aspect of the present invention.

[0014] That is, in another aspect, the present invention can be exemplified as follows. [1] 1. A method for preparing a model that predicts the degree of fit of a test substance to a desired odor profile, comprising: the model includes a regression equation that outputs a predicted value of the goodness of fit of the test substance based on test olfactory receptor activation data of the test substance; the method includes generating the regression equation by machine learning; The method, wherein the test olfactory receptor activation data is data regarding activation of a test olfactory receptor by the test substance. [2] The method, wherein the regression equation is a linear regression equation. [3] The machine learning is performed using a dataset including odor characteristic data of a control substance and control olfactory receptor activation data; the aroma characteristic data is data indicating the degree of conformity of the control substance to the target aroma characteristic, the control olfactory receptor activation data is data regarding activation of a control olfactory receptor by the control substance; the control substance is a combination of two or more substances; The above method, wherein the control olfactory receptor is a combination of two or more olfactory receptors including the test olfactory receptor. [4] The method, wherein the control olfactory receptor activation data is data indicating the degree of activation of the control olfactory receptor by the control substance. [5] The method, wherein the machine learning is performed using the aroma characteristic data as a response variable and the control olfactory receptor activation data as an explanatory variable. [6] The method, wherein the control substance is a combination of 100 or more substances. [7] The method, wherein 50% or more of the total number of control substances are selected from compounds listed in the Atlas of Odor Character Profiles. [8] The method, wherein the odor characteristic data is a percentage of applicability value calculated according to the standards described in the Atlas of Odor Character Profiles. [9] The above method, wherein the test olfactory receptor is a combination of 10 or more types of olfactory receptors.

[10] The above method, wherein the control olfactory receptor is a combination of 300 or more types of olfactory receptors.

[11] 50% or more of the total number of the control olfactory receptors is OR1A1, OR1A2, OR1B1, OR1C1, OR1D2, OR1D5, OR1E1, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1R1P, OR1S1, OR2A1, OR2A2, OR2A4, OR2A5, OR2A12, OR2A14, OR2A25, OR2AE1, OR2AG1, OR2AG2, OR2AJ1P, OR2 AK2, OR2AP1, OR2AT4, OR2B2, OR2B3, OR2B6, OR2B11, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2G2, OR2G3, OR2G6, OR2H1, OR2H2, OR2J2, OR2J3, OR2K2, OR2 L2, OR2L8, OR2L13, OR2M2, OR2M4, OR2M7, OR2S2, OR2T1, OR2T2, OR2T5, OR2T6, OR2T8, OR2T10, OR2T11, OR2T27, OR2T34, OR2V2, OR2W1, OR2W3, OR2Y1, OR2 Z1, OR3A1, OR3A2, OR3A3, OR3A4, OR4A5, OR4A15, OR4A16, OR4A47, OR4B1, OR4C3, OR4C5, OR4C6, OR4C11, OR4C12, OR4C13, OR4C15, OR4C16, OR4C46, OR4D 1, OR4D2, OR4D5, OR4D6, OR4D9, OR4D10, OR4D11, OR4E2, OR4F3, OR4F5, OR4F6, OR4F14P, OR4F15, OR4G11P, OR4H12P, OR4K1, OR4K2, OR4K5, OR4K13, OR4K1 4, OR4K15, OR4K17, OR4L1, OR4M1, OR4N2, OR4N4, OR4N5, OR4P4, OR4Q3, OR4S1, OR4S2, OR4X1, OR4X2, OR5A1, OR5A2, OR5AC2, OR5AK2, OR5AK3P, OR5AN1, O R5AP2, OR5AR1, OR5AS1, OR5AU1, OR5B2, OR5B3, OR5B12, OR5B17, OR5B21, OR5C1, OR5D13, OR5D14, OR5D16, OR5D18, OR5F1, OR5H1, OR5H2, OR5H6, OR5H14,OR5I1、OR5J2、OR5K1、OR5K3、OR5K4、OR5L2、OR5M3、OR5M8、OR5M9、OR5M10、OR5M11、OR5P3、OR5R1、OR5T1、OR5T2、OR5T3、OR5V1、OR5W2、OR6A2、OR6B1、OR6B2、OR6C1、OR6C2、OR6C3、OR6C4、OR6C6、OR6C65、OR6C66P、OR6C68、OR6C70、OR6C74、OR6C75、OR6C76、OR6F1、OR6J1、OR6K2、OR6K3、OR6K6、OR6M1、OR6N1、OR6N2、OR6P1、OR6Q1、OR6S1、OR6T1、OR6V1、OR6X1、OR6Y1、OR7A3P、OR7A5、OR7A10、OR7A17、OR7C1、OR7C2、OR7D2、OR7D4、OR7E24、OR7G1、OR7G2、OR7G3、OR8A1、OR8B3、OR8B4、OR8B8、OR8B12、OR8D1、OR8D2、OR8D4、OR8G2、OR8G5、OR8H3、OR8I2、OR8J1、OR8J3、OR8K1、OR8K3、OR8K5、OR8S1、OR8U1、OR9A4、OR9G1、OR9G4、OR9I1、OR9K2、OR9Q1、OR9Q2、OR10A3、OR10A4、OR10A5、OR10A6、OR10A7、OR10AD1、OR10AG1、OR10C1、OR10D3、OR10D4P、OR10G2、OR10G3、OR10G4、OR10G6、OR10G7、OR10G9、OR10H2、OR10H4、OR10J1、OR10J3、OR10J5、OR10K1、OR10K2、OR10P1、OR10Q1、OR10R2、OR10S1、OR10T2、OR10V1、OR10W1、OR10X1、OR10Z1、OR11A1、OR11G2、OR11H4、OR11H6、OR11H12、OR11L1、OR12D2、OR12D3、OR13A1、OR13C2、OR13C3、OR13C4、OR13C8、OR13D1、OR13F1、OR13G1、OR13H1、OR13J1、OR14A2、OR14A16、OR14C36、OR14I1、OR14J1、OR14K1、OR14L1P、OR51A1P、OR51A4、OR51A7、OR51B2、OR51B4、OR51B5、OR51B6、OR51D1、OR51E1、OR51E2, OR51F1, OR51F2, OR51F5P, OR51G1, OR51G2, OR51H1, OR51I1, OR51I2, OR51L1, OR51M1, OR51Q1, O R51S1, OR51T1, OR51V1, OR52A1, OR52A4, OR52A5, OR52B2, OR52B4, OR52B6, OR52D1, OR52E2, OR52E4, OR52 E5, OR52E8, OR52H1, OR52I2, OR52J3, OR52K2, OR52L2P, OR52M1, OR52N1, OR52N2, OR52N4, OR52N5, OR52P2P, OR52R1, OR52W1, OR52Z1P, OR56A1, OR56A3, OR56A4, OR56A5, OR56B1, OR56B2P, OR56B4.

[12] The method, wherein the control olfactory receptor is a human olfactory receptor.

[13] The method, wherein the control olfactory receptor activation data for the control olfactory receptors for which the absolute value of the correlation coefficient between the odor characteristic data and the control olfactory receptor activation data exceeds 0.2 is used as an explanatory variable in the machine learning.

[14] The method, wherein the step includes a step of calculating a correlation coefficient between the odor characteristic data and the control olfactory receptor activation data prior to the machine learning.

[15] A model produced by the method.

[16] 1. A method for predicting the degree of fit of a test substance to a desired odor profile, comprising: predicting the degree of conformance of the test substance to the target odor characteristics based on the test olfactory receptor activation data of the test substance and the model; A method comprising:

[17] A method for screening a substance that has a high degree of compatibility with a target aroma characteristic, comprising the steps of: predicting the degree of conformance of the test substance to the desired odor profile based on the test olfactory receptor activation data of the test substance and the model; and a step of selecting test substances predicted to have a high degree of compatibility with the target odor characteristics as substances with a high degree of compatibility with the target odor characteristics; A method comprising:

[18] The method further comprises a step of confirming the degree of suitability of test substances predicted to have a high degree of suitability for the target aroma characteristics. [Brief explanation of the drawings]

[0015] [Figure 1] A heat map of the stereochemical similarity matrix (grayscale image). [Figure 2] Figure showing the results of visualizing the stereochemical similarity space using t-SNE (grayscale image). [Figure 3] Diagram showing the distribution of "OR4S2 activity" in stereochemical structure similarity space (grayscale image). [Figure 4] Diagram showing the distribution of "OR5K1 activity" in stereochemical structure similarity space (grayscale image). [Figure 5] Diagram showing the distribution of "OR10G4 activity" in stereochemical structure similarity space (grayscale image). [Figure 6] A diagram (gray-tone image) showing the distribution of the aroma attribute "onion" in the stereochemical structure similarity space. [Figure 7] A diagram (halftone image) showing the distribution of the aroma attribute "nutty" in stereochemical structure similarity space. [Figure 8] A diagram (gray-tone image) showing the distribution of the aroma characteristic "phenolic" in the stereochemical structure similarity space. [Figure 9] A figure (grayscale image) showing the results of evaluating the correlation between odor similarity and stereochemical structure similarity and molecular fingerprint similarity when mixed at various mixing ratios. [Figure 10] A figure (grayscale image) showing the results of evaluating the correlation between odor similarity and stereochemical structure similarity and molecular fingerprint similarity when mixed at various mixing ratios. [Figure 11] A diagram showing the tree model of the aroma attribute "burnt." [Figure 12] A diagram showing the tree model for the aroma attribute "sweet." [Figure 13] A diagram showing the tree model for the aroma characteristic "nuts." [Figure 14] A diagram showing a tree model of the pyrazine skeleton. [Figure 15] A diagram showing a tree model of the aldehyde group. [Figure 16] A diagram showing a tree model of an ester bond. [Figure 17] FIG. 1 is a graph showing the relationship between the measured PA value and the predicted PA value for the aroma characteristic "STARWBERRY." [Figure 18] Graph showing the relationship between measured PA values ​​and predicted PA values ​​for the aroma characteristic "ANISE (LICORICE)." [Figure 19] FIG. 1 is a graph showing the relationship between the measured PA value and the predicted PA value for the aroma characteristic "NEW RUBBER." DETAILED DESCRIPTION OF THE INVENTION

[0016] (A) First Aspect of the Invention The first aspect of the present invention, specifically the prediction method of the present invention and the design method of the present invention relating to the first aspect of the present invention, will be described below.

[0017] <1> Prediction method of the present invention according to the first aspect of the present invention The prediction method of the present invention is a method for predicting the presence or absence of a target property for a test substance. "Predicting the presence or absence of a target property for a test substance" means predicting whether the test substance has the target property. Predicting the presence or absence of a target property for a test substance will hereinafter also be referred to simply as "prediction." Prediction can be performed based on the maximum similarity of the stereochemical structures between the test substance and a control substance. That is, the prediction method of the present invention may include a step of predicting the presence or absence of a target property for a test substance based on the maximum similarity of the stereochemical structures between the test substance and a control substance. This step is also referred to as a "prediction step." By performing prediction based on the maximum similarity of the stereochemical structures between the test substance and a control substance, the accuracy of prediction can be improved compared to prediction based on structural similarity between substances that does not take multiple conformations into account, such as molecular fingerprint similarity.

[0018] Furthermore, by predicting the presence or absence of a desired property for a test substance, it is possible to screen for a substance having the desired property. That is, a test substance predicted to have the desired property can be selected as a substance having the desired property, thereby screening for a substance having the desired property. That is, the prediction method of the present invention may be a method for screening for a substance having the desired property. That is, the prediction method of the present invention may further include a step of selecting a test substance predicted to have the desired property as a substance having the desired property. That is, the screening method may be a method for screening for a substance having the desired property, including a step of predicting the presence or absence of the desired property for a test substance based on the maximum similarity of the stereochemical structure between the test substance and a control substance, and a step of selecting the test substance predicted to have the desired property as a substance having the desired property. In other words, the screening method may be a method for screening for a substance having the desired property, including a step of predicting the presence or absence of the desired property for a test substance using the prediction method of the present invention, and a step of selecting the test substance predicted to have the desired property as a substance having the desired property.

[0019] The prediction method of the present invention may further comprise, prior to the prediction step, a step of calculating the maximum similarity of the stereochemical structures between the test substance and the control substance. This step is also referred to as a "calculation step."

[0020] <1-1> Target characteristics The term "target property" refers to the property to be predicted, such as an odor property or an olfactory receptor activation property.

[0021] The term "aroma characteristics" refers to the properties of exhibiting an aroma. The type of aroma is not particularly limited. Aromas include absinthe, acacia, acai, acerola, acetic, acetone, acidic, acorn, acrylate, agarwood, alcoholic, aldehydic, alfalfa, algae, alliaceous, allspice, almond, almond bitter almond, almond roasted almond, almond toasted. almond, amber, ambergris, ambrette, ammoniacal, angelica, animal, anise, anisic, apple, apple cooked apple, apple dried apple, apple green apple, apple red apple, apple skin, apricot, aromatic, arrack, artichoke, asafetida, asparagus, astringent, autumn, avocado, bacon, baked, balsamic, banana, banana peel, banana ripe banana, banana unripe banana, barley roasted barley, basil, bay, bean green bean, beany, beef juice, beefy, beefy roasted beefy, beer, beeswax, benzoin, bergamot, berry, berry ripe berry, bitter, blackberry, bloody, blueberry, bois de rose, boronia, bouillon, boysenberry, brandy, bread baked, bread crust, bread rye bread, bready, broccoli, brothy, brown, bubble gum, buchu, burnt, butterrancid、buttermilk、butterscotch、buttery、cabbage、calamus、camphoreous、cananga、candy、cantaloupe、capers、caramellic、caraway、cardamom、carnation、carrot、carrot seed、carvone、cascarilla、cashew、cassia、castoreum、catty、cauliflower、cedar、cedarwood、celery、cereal、chamomile、charred、cheesy、cheesy bleu cheese、cheesy cheddar cheese、cheesy feta cheese、cheesy gorgonzola cheese、cheesy gouda cheese、cheesy limburger cheese、cheesy parmesan cheese、cheesy roquefort cheese、chemical、cherry、cherry maraschino cherry、chervil、chestnut、chicken、chicken coup、chicken fat、chicken roasted chicken、chicory、chive、chocolate、chocolate dark chocolate、chocolate white chocolate、chrysanthemum、cider、cilantro、ciltrano、cinnamon、cinnamyl、cistus、citronella、citrus、citrus peel、citrus rind、civet、clam、clean、cloth laundered cloth、clove、clover、cocoa、coconut、coffee、coffee roasted coffee、cognac、cologne、cooked、cookie、cooling、copaiba、coriander、corn、corn chip、cornmeal、cornmint、cortex、costus、cottoncandy、coumarinic、cranberry、creamy、cubeb、cucumber、cucumber skin、cumin、currant black currant、currant bud black currant bud、currant red currant、curry、custard、cyclamen、cypress、dairy、date、davana、deertongue、dewy、dill、dirty、dragon fruit、dry、durian、dusty、earthy、egg nog、egg yolk、eggy、elderberry、elderflower、elemi、estery、ethereal、eucalyptus、fatty、fecal、fennel、fenugreek、fermented、fig、filbert、fir needle、fishy、fleshy、floral、foliage、forest、fougere、frankincense、freesia、fresh、fresh outdoors、fried、fruit dried fruit、fruit overripe fruit、fruit ripe fruit、fruit tropical fruit、fruity、fudge、fungal、fusel、galanga、galbanum、gardenia、garlic、gasoline、gassy、genet、geranium、ginger、ginseng、goaty、goji berry、gooseberry、gourmand、graham cracker、grain、grain toasted grain、grape、grape skin、grapefruit、grapefruit peel、grassy、gravy、greasy、green、grilled、guaiacol、guaiacwood、guava、hairy、ham、harsh、hawthorn、hay、hay new mown hay、hazelnut、hazelnut roastedhazelnut、heather、heliotrope、herbal、hibiscus、honey、honeydew、honeysuckle、hops、horehound、horseradish、huckleberry、humus、hyacinth、hyssop、immortelle、incense、jackfruit、jammy、jasmin、jonquil、juicy、juicy fruit、juniper、ketonic、kimchi、kiwi、kokumi、kumquat、labdanum、lachrymatory、lactonic、lamb、lard、lavandin、lavender、lavender spike lavender、leafy、leathery、leek、lemon、lemon peel、lemongrass、lettuce、licorice、licorice black licorice、lilac、lily、lily of the valley、lime、linden flower、lingonberry、liver、lobster、loganberry、lovage、lychee、macadamia、mace、magnolia、mahogany、malty、mandarin、mango、maple、marigold、marine、marjoram、marshmallow、marzipan、mastic、meaty、meaty roasted meaty、medicinal、melon、melon rind、melon unripe melon、mentholic、metallic、milky、mimosa、minty、molasses、moldy、mossy、muguet、mulberry、mushroom、musk、mustard、musty、mutton、myrrh、naphthyl、narcissus、nasturtium、natural、neroli、noni fruit、nut flesh、nut skin、nutmeg、nutty、oakmoss、oatmeal、oats、ocean、oily、onion、onion cooked onion、onion greenonion、opoponax、orange、orange bitter orange、orange peel、orange rind、orangeflower、orchid、oriental、origanum、orris、osmanthus、oyster、ozone、painty、palmarosa、papaya、paper、parsley、passion fruit、patchouli、pea green pea、peach、peanut、peanut butter、peanut roasted peanut、pear、pear skin、pecan、peely、pennyroyal、peony、pepper bell pepper、pepper black pepper、peppermint、peppery、peru balsam、petal、petitgrain、petroleum、phenolic、pimenta、pine、pineapple、pistachio、plastic、plum、plum skin、pomegranate、popcorn、pork、potato、potato baked potato、potato chip、potato raw potato、powdery、praline、privet、privetblossom、prune、pulpy、pumpkin、pungent、quince、radish、rain、raisin、rancid、raspberry、raw、reseda、resinous、rhubarb、rindy、ripe、roasted、root beer、rooty、rose、rose dried rose、rose red rose、rose tea rose、rose white rose、rosemary、rubbery、rue、rummy、saffron、sage、sage clary sage、salmon、salty、sandalwood、sandy、sappy、sarsaparilla、sassafrass、sauerkraut、sausage、sausage smokedsausage、savory、sawdust、scallion、seafood、seashore、seaweed、seedy、sesame、sharp、shellfish、shrimp、skunk、smoky、soapy、soft、solvent、soup、sour、spearmint、spicy、spinach、spruce、starchy、starfruit、storax、strawberry、stringent、styrene、sugar、sugar brown sugar、sugar burnt sugar、sulfurous、sweaty、sweet、sweet pea、taco、tagette、tallow、tamarind、tangerine、tansy、tarragon、tart、tea、tea black tea、tea green tea、tea rooibos tea、tea white tea、tequila、terpenic、thujonic、thyme、toasted、tobacco、toffee、tolu balsam、tomato、tomato leaf、tonka、tropical、truffle、tuberose、tuna、turkey、turmeric、turnup、tutti frutti、umami、urine、valerian root、vanilla、vegetable、verbena、vetiver、vinegar、violet、violet leaf、walnut、warm、wasabi、watercress、watermelon、watermelon rind、watery、waxy、weedy、wet、whiskey、winey、wintergreen、woody、woody burnt wood、woody oak wood、woody old wood、wormwood、yeasty、ylang、yogurt、yuzu、zedoary、zesty、bark、birch bark、blood、raw meat、burnt candle、burnt milk、burnt pepper、burnt rubber、cadaverous (dead animal)、cardboard、catExamples of aromas include urine, chalky, cleaning fluid, cooked vegetables, cork, creosote, crushed grass, crushed weeds, dirty linen, disinfectant, carbohydrate, fermented (rotten) fruit, fragrant, fresh green vegetables, fresh tobacco smoke, fried chicken, heavy, household gas, kerosene, kippery (smoked fish), laurel leaves, light, mothballs, mouse, nail polish remover, new rubber, peanut butter, perfumery, putrid, four, decayde, rope, seasoning (for meat), seminal, sperm-like, sewer, sickening, sooty, sour milk, stale, stale tobacco smoke, tab, tea leaves, turpentine (pine oil), varnish, wet paper, wet wool, and wet dog. The aroma may be a single aroma or a combination of two or more aromas. In other words, the "presence or absence of aroma characteristics" may refer to the presence or absence of the property of presenting any one type of aroma, or the presence or absence of the property of presenting each of two or more types of aromas (i.e., the pattern of which aromas are present and which are not present for two or more types of aromas).

[0022] The term "olfactory receptor activating properties" refers to the property of activating an olfactory receptor. The type of olfactory receptor is not particularly limited.

[0023] The olfactory receptors are OR1A1, OR1A2, OR1B1, OR1C1, OR1D2, OR1D4, OR1D5, OR1E1, OR1E2, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L6, OR1L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1R1P, OR1S1, OR1S2, OR2A1, OR2A2, OR2A4, OR2A5, OR2A7, OR2A12, OR2A14, OR2A25, OR2AE1, and OR2AG. 1, OR2AG2, OR2AJ1P, OR2AK2, OR2AP1, OR2AT4, OR2B2, OR2B3, OR2B6, OR2B11, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2F2, OR2G2, OR2G3, OR2G6, OR2H1, O R2H2, OR2J1P, OR2J2, OR2J3, OR2K2, OR2L2, OR2L3, OR2L5, OR2L8, OR2L13, OR2M2, OR2M3, OR2M4, OR2M5, OR2M7, OR2S2, OR2T1, OR2T2, OR2T3, OR2T4, OR2T 5, OR2T6, OR2T7, OR2T8, OR2T10, OR2T11, OR2T12, OR2T27, OR2T29, OR2T33, OR2T34, OR2T35, OR2V1, OR2V2, OR2W1, OR2W3, OR2Y1, OR2Z1, OR3A1, OR3A2, OR3A3, OR3A4, OR4A4P, OR4A5, OR4A15, OR4A16, OR4A47, OR4B1, OR4C3, OR4C5, OR4C6, OR4C11, OR4C12, OR4C13, OR4C15, OR4C16, OR4C45, OR4C46, OR4D1, OR4D2, OR4D5, OR4D6, OR4D9, OR4D10, OR4D11, OR4E2, OR4F3, OR4F4, OR4F5, OR4F6, OR4F14P, OR4F15, OR4F17, OR4F21, OR4G11P, OR4H12P, OR4K1, OR4K2, OR4K5, OR4K13, OR4K14, OR4K15, OR4K17, OR4L1, OR4M1, OR4M2, OR4N2, OR4N4, OR4N5, OR4P4, OR4Q3, OR4S1, OR4S2, OR4X1, OR4X2, OR5A1, OR5A2, OR5AC2,OR5AK2、OR5AK3P、OR5AN1、OR5AP2、OR5AR1、OR5AS1、OR5AU1、OR5B2、OR5B3、OR5B12、OR5B17、OR5B21、OR5C1、OR5D13、OR5D14、OR5D16、OR5D18、OR5F1、OR5H1、OR5H2、OR5H6、OR5H14、OR5H15、OR5I1、OR5J2、OR5K1、OR5K2、OR5K3、OR5K4、OR5L1、OR5L2、OR5M1、OR5M3、OR5M8、OR5M9、OR5M10、OR5M11、OR5P2、OR5P3、OR5R1、OR5T1、OR5T2、OR5T3、OR5V1、OR5W2、OR6A2、OR6B1、OR6B2、OR6B3、OR6C1、OR6C2、OR6C3、OR6C4、OR6C6、OR6C65、OR6C66P、OR6C68、OR6C70、OR6C74、OR6C75、OR6C76、OR6F1、OR6J1、OR6K2、OR6K3、OR6K6、OR6M1、OR6N1、OR6N2、OR6P1、OR6Q1、OR6S1、OR6T1、OR6V1、OR6X1、OR6Y1、OR7A3P、OR7A5、OR7A10、OR7A17、OR7C1、OR7C2、OR7D2、OR7D4、OR7E24、OR7G1、OR7G2、OR7G3、OR8A1、OR8B2、OR8B3、OR8B4、OR8B8、OR8B12、OR8D1、OR8D2、OR8D4、OR8G1、OR8G2、OR8G5、OR8H1、OR8H2、OR8H3、OR8I2、OR8J1、OR8J3、OR8K1、OR8K3、OR8K5、OR8S1、OR8U1、OR8U8、OR9A2、OR9A4、OR9G1、OR9G4、OR9I1、OR9K2、OR9Q1、OR9Q2、OR10A2、OR10A3、OR10A4、OR10A5、OR10A6、OR10A7、OR10AD1、OR10AG1、OR10C1、OR10D3、OR10D4P、OR10G2、OR10G3、OR10G4、OR10G6、OR10G7、OR10G8、OR10G9、OR10H1、OR10H2、OR10H3、OR10H4、OR10H5、OR10J1、OR10J3、OR10J5、OR10K1、OR10K2、OR10P1、OR10Q1、OR10R2、OR10S1、OR10T2、OR10V1、OR10W1、OR10X1, OR10Z1, OR11A1, OR11G2, OR11H1, OR11H2, OR11H4, OR11H6, OR11H12, OR11L1, OR12D2, OR12 D3, OR13A1, OR13C2, OR13C3, OR13C4, OR13C5, OR13C8, OR13C9, OR13D1, OR13F1, OR13G1, OR13H1, OR1 3J1, OR14A2, OR14A16, OR14C36, OR14I1, OR14J1, OR14K1, OR14L1P, OR51A1P, OR51A2, OR51A4, OR51 A7, OR51B2, OR51B4, OR51B5, OR51B6, OR51D1, OR51E1, OR51E2, OR51F1, OR51F2, OR51F5P, OR51G1, OR 51G2, OR51H1, OR51I1, OR51I2, OR51L1, OR51M1, OR51Q1, OR51S1, OR51T1, OR51V1, OR52A1, OR52A4, OR52A5, OR52B2, OR52B4, OR52B6, OR52D1, OR52E2, OR52E4, OR52E5, OR52E6, OR52E8, OR52H1, OR52I1 , OR52I2, OR52J3, OR52K1, OR52K2, OR52L1, OR52L2P, OR52M1, OR52N1, OR52N2, OR52N4, OR52N5, OR52P2P, OR52R1, OR52W1, OR52Z1P, OR56A1, OR56A3, OR56A4, OR56A5, OR56B1, OR56B2P, and OR56B4.

[0024] Olfactory receptors include, in particular, OR1A1, OR1A2, OR1B1, OR1C1, OR1D2, OR1D5, OR1E1, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1R1P, OR1S1, OR2A1, OR2A2, OR2A4, OR2A5, OR2A12, OR2A14, OR2A25, OR2AE1, OR2AG1, OR2AG2, OR2AJ1P, OR2AK2, and OR2A P1, OR2AT4, OR2B2, OR2B3, OR2B6, OR2B11, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2G2, OR2G3, OR2G6, OR2H1, OR2H2, OR2J2, OR2J3, OR2K2, OR2L2, OR2L8, OR2L13, OR2M2, OR2M4, OR2M7, OR2S2, OR2T1, OR2T2, OR2T5, OR2T6, OR2T8, OR2T10, OR2T11, OR2T27, OR2T34, OR2V2, OR2W1, OR2W3, OR2Y1, OR2Z1, OR3A1, OR3A2, OR3A3, OR3A4, OR4A5, OR4A15, OR4A16, OR4A47, OR4B1, OR4C3, OR4C5, OR4C6, OR4C11, OR4C12, OR4C13, OR4C15, OR4C16, OR4C46, OR4D1, OR4D2, OR 4D5, OR4D6, OR4D9, OR4D10, OR4D11, OR4E2, OR4F3, OR4F5, OR4F6, OR4F14P, OR4F15, OR4G11P, OR4H12P, OR4K1, OR4K2, OR4K5, OR4K13, OR4K14, OR4K15, O R4K17, OR4L1, OR4M1, OR4N2, OR4N4, OR4N5, OR4P4, OR4Q3, OR4S1, OR4S2, OR4X1, OR4X2, OR5A1, OR5A2, OR5AC2, OR5AK2, OR5AK3P, OR5AN1, OR5AP2, OR5AR 1, OR5AS1, OR5AU1, OR5B2, OR5B3, OR5B12, OR5B17, OR5B21, OR5C1, OR5D13, OR5D14, OR5D16, OR5D18, OR5F1, OR5H1, OR5H2, OR5H6, OR5H14, OR5I1, OR5J2,OR5K1、OR5K3、OR5K4、OR5L2、OR5M3、OR5M8、OR5M9、OR5M10、OR5M11、OR5P3、OR5R1、OR5T1、OR5T2、OR5T3、OR5V1、OR5W2、OR6A2、OR6B1、OR6B2、OR6C1、OR6C2、OR6C3、OR6C4、OR6C6、OR6C65、OR6C66P、OR6C68、OR6C70、OR6C74、OR6C75、OR6C76、OR6F1、OR6J1、OR6K2、OR6K3、OR6K6、OR6M1、OR6N1、OR6N2、OR6P1、OR6Q1、OR6S1、OR6T1、OR6V1、OR6X1、OR6Y1、OR7A3P、OR7A5、OR7A10、OR7A17、OR7C1、OR7C2、OR7D2、OR7D4、OR7E24、OR7G1、OR7G2、OR7G3、OR8A1、OR8B3、OR8B4、OR8B8、OR8B12、OR8D1、OR8D2、OR8D4、OR8G2、OR8G5、OR8H3、OR8I2、OR8J1、OR8J3、OR8K1、OR8K3、OR8K5、OR8S1、OR8U1、OR9A4、OR9G1、OR9G4、OR9I1、OR9K2、OR9Q1、OR9Q2、OR10A3、OR10A4、OR10A5、OR10A6、OR10A7、OR10AD1、OR10AG1、OR10C1、OR10D3、OR10D4P、OR10G2、OR10G3、OR10G4、OR10G6、OR10G7、OR10G9、OR10H2、OR10H4、OR10J1、OR10J3、OR10J5、OR10K1、OR10K2、OR10P1、OR10Q1、OR10R2、OR10S1、OR10T2、OR10V1、OR10W1、OR10X1、OR10Z1、OR11A1、OR11G2、OR11H4、OR11H6、OR11H12、OR11L1、OR12D2、OR12D3、OR13A1、OR13C2、OR13C3、OR13C4、OR13C8、OR13D1、OR13F1、OR13G1、OR13H1、OR13J1、OR14A2、OR14A16、OR14C36、OR14I1、OR14J1、OR14K1、OR14L1P、OR51A1P、OR51A4、OR51A7、OR51B2、OR51B4、OR51B5、OR51B6、OR51D1、OR51E1、OR51E2、OR51F1, OR51F2, OR51F5P, OR51G1, OR51G2, OR51H1, OR51I1, OR51I2, OR51L1, OR51M1, OR51Q1, OR51S 1, OR51T1, OR51V1, OR52A1, OR52A4, OR52A5, OR52B2, OR52B4, OR52B6, OR52D1, OR52E2, OR52E4, OR52 E5, OR52E8, OR52H1, OR52I2, OR52J3, OR52K2, OR52L2P, OR52M1, OR52N1, OR52N2, OR52N4, OR52N5, OR52P2P, OR52R1, OR52W1, OR52Z1P, OR56A1, OR56A3, OR56A4, OR56A5, OR56B1, OR56B2P, and OR56B4.

[0025] Genes that encode olfactory receptors are also called olfactory receptor genes. The olfactory receptor may be one type of olfactory receptor, or a combination of two or more types of olfactory receptors. In other words, "the presence or absence of olfactory receptor activation properties" may refer to the presence or absence of the property of activating any one type of olfactory receptor, or the presence or absence of the property of activating each of two or more types of olfactory receptors (i.e., the pattern of which olfactory receptors are activated and which olfactory receptors are not activated for two or more types of olfactory receptors).

[0026] Olfactory receptor genes and olfactory receptors include those of various organisms. Examples of organisms include animals such as mammals. Specific examples of animals such as mammals include Homo sapiens (humans), Mus musculus (mice), Rattus norvegicus (rat), Canis lupus familiaris (dogs), Felis catus (cats), Bos taurus (cattle), Sus scrofa (pigs), Pan troglodytes (chimpanzees), Macaca fascicularis (cyn-eating monkeys), and Equus caballus (horses). Examples of animals such as mammals include humans in particular. The nucleotide sequences of olfactory receptor genes and amino acid sequences of olfactory receptors of various organisms can be obtained from public databases such as NCBI and Ensembl.

[0027] The olfactory receptor may be, for example, a protein having a known or naturally occurring amino acid sequence of an olfactory receptor such as those described above. Furthermore, the olfactory receptor may be, for example, a variant of a protein having a known or naturally occurring amino acid sequence of an olfactory receptor such as those described above. That is, the olfactory receptors identified by the above names encompass, for example, proteins having the known or naturally occurring amino acid sequence of the olfactory receptor identified by the name, and variants thereof. The phrase "a protein having an amino acid sequence" means that the protein contains the amino acid sequence, unless otherwise specified, and also encompasses cases where the protein consists of the amino acid sequence. Examples of variants include proteins having an amino acid sequence in which one or several amino acids at one or several positions in a known or naturally occurring amino acid sequence have been substituted, deleted, inserted, and / or added. Specifically, "one or several" may mean, for example, 1 to 50, 1 to 40, 1 to 30, preferably 1 to 20, more preferably 1 to 10, even more preferably 1 to 5, and particularly preferably 1 to 3. Variants also include proteins having an amino acid sequence that is, for example, 50% or more, 65% or more, 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire amino acid sequence of a known or naturally occurring olfactory receptor. Note that olfactory receptors identified by their originating biological species are not limited to the olfactory receptors themselves found in the biological species, but also include proteins having the amino acid sequence of the olfactory receptors found in the biological species, and variants thereof. Variants may or may not be found in the biological species. For example, "human olfactory receptors" are not limited to the olfactory receptors themselves found in humans, but also include proteins having the amino acid sequence of the olfactory receptors found in humans, and variants thereof. Olfactory receptors may be, for example, chimeric proteins of two or more olfactory receptors of different origins. In other words, the olfactory receptors identified by the above names also include, for example, chimeric proteins of two or more olfactory receptors of different origins identified by the names.

[0028] The "identity" between amino acid sequences refers to the identity between amino acid sequences calculated using blastp with default scoring parameters (Matrix: BLOSUM62; Gap Costs: Existence = 11, Extension = 1; Compositional Adjustments: Conditional compositional score matrix adjustment).

[0029] <1-2> Test substance The term "test substance" refers to a substance that is the subject of prediction of the presence or absence of a desired property. In other words, the term "test substance" refers to a substance that is used as a candidate for a substance having a desired property in a method for screening for a substance having a desired property. The test substance is not particularly limited as long as its structure has been identified.

[0030] The structure of the test substance may be identified to the extent that multiple conformations of the test substance can be generated. The structure of the test substance may be identified, for example, as a chemical structural formula. The structure of the test substance may or may not be publicly known. If the structure of the test substance is not publicly known, the structure of the test substance may be identified as appropriate before generating multiple conformations. The method for identifying the structure of the test substance is not particularly limited. The structure of the test substance can be identified, for example, by a known method for identifying the structure of a substance. Such methods include nuclear magnetic resonance (NMR), electron spin resonance (ESR), ultraviolet-visible-near-infrared spectroscopy (UV-Vis-NIR), infrared spectroscopy (IR), Raman spectroscopy, and mass spectrometry (MS). These methods may be used alone or in appropriate combination.

[0031] The test substance may be a known substance or a novel substance. The test substance may be a natural product or an artificial product. For example, the test substance may be a compound library generated using combinatorial chemistry techniques. Examples of test substances include alcohols, ketones, aldehydes, ethers, esters, hydrocarbons, sugars, organic acids, nucleic acids, amino acids, peptides, and various other organic or inorganic compounds. In particular, the test substance may be an existing food additive. "Existing food additive" refers to a substance already approved for use as a food additive. The test substance may also be a hypothetical substance (i.e., a substance with a hypothetical structure). Examples of hypothetical substances include substances listed in compound databases such as GDB-11, GDB-13, GDB-17, ZINC15, FooDB, and VCF (Volatile Compounds in Food). A single test substance may be used, or two or more test substances may be used in combination. Test substances may be selected to include, for example, existing food additives, such as those listed above. That is, the test substance may be, for example, one existing food additive, two or more food additives in combination, or one or more food additives in combination with one or more other substances. Note that "using two or more test substances in combination" means predicting the presence or absence of the desired properties for each of the two or more test substances.

[0032] <1-3> Control substance The term "control substance" refers to a substance that can be used as an indicator of the presence or absence of a desired property. The control substance is not particularly limited as long as its structure has been identified and the presence or absence of the desired property has been identified.

[0033] The structure of the reference substance may be identified to the extent that multiple conformations of the reference substance can be generated. The structure of the reference substance may be identified, for example, as a chemical structural formula. The structure of the reference substance may be publicly known or not. If the structure of the reference substance is not publicly known, the structure of the reference substance may be identified as appropriate before generating multiple conformations. The method for identifying the structure of the reference substance is not particularly limited. The structure of the reference substance can be identified, for example, by a known method for identifying the structure of a substance. Such methods include nuclear magnetic resonance (NMR), electron spin resonance (ESR), ultraviolet-visible-near-infrared spectroscopy (UV-Vis-NIR), infrared spectroscopy (IR), Raman spectroscopy, and mass spectrometry (MS). These methods may be used alone or in appropriate combination.

[0034] The presence or absence of the target characteristic in the control substance may or may not be publicly known. If the presence or absence of the target characteristic in the control substance is not publicly known, the presence or absence of the target characteristic in the control substance may be identified as appropriate before the prediction step is performed. The method for identifying the presence or absence of the characteristic in the control substance is not particularly limited. The presence or absence of the characteristic in the control substance can be identified, for example, by a known method for identifying the presence or absence of a characteristic in a substance. The presence or absence of an aroma characteristic in the control substance can be identified, for example, by sensory evaluation by an expert panel. The presence or absence of the olfactory receptor activation characteristic in the control substance can be identified, for example, by contacting the olfactory receptor with the control substance and measuring the presence or absence of olfactory receptor activation due to contact with the control substance. Contact between the olfactory receptor and the control substance and measurement of the presence or absence of olfactory receptor activation due to this can be performed, for example, with reference to a screening method for a substance exhibiting a target odor using the olfactory receptor response as an indicator (e.g., JP 2019-037197 A). The olfactory receptor may be used by being supported on cells such as animal cells. The activation of olfactory receptors can be measured, for example, using an increase in the amount of intracellular calcium or the amount of intracellular cAMP as an index. Techniques for measuring the amount of intracellular cAMP include, for example, ELISA and reporter assays. Examples of reporter assays include luciferase assays. According to reporter assays, the amount of intracellular cAMP can be measured using a reporter gene (such as a luciferase gene) that is constructed so that its expression depends on the amount of cAMP. Examples of techniques for measuring the amount of intracellular calcium include calcium imaging.

[0035] Furthermore, the degree of the target characteristic of the control substance may be identified. In the case of an aroma characteristic, the "degree of the characteristic" may mean the intensity at which the substance exhibits the aroma. In the case of an olfactory receptor activation characteristic, the "degree of the characteristic" may mean the intensity at which the substance activates the olfactory receptor. The degree of the target characteristic in the control substance can be identified, for example, by a method similar to that used to identify the presence or absence of the target characteristic in the control substance.

[0036] Specifically, contact between an olfactory receptor and a control substance and measurement of the presence or absence or degree of activation of the olfactory receptor due to this can be carried out, for example, by the following procedure.

[0037] That is, the presence or absence or degree of activation of the olfactory receptor by a control substance can be determined by contacting the olfactory receptor with the control substance and using the degree of activation of the olfactory receptor (degree of activation D1) when the contact is carried out (i.e., under the conditions for contacting the olfactory receptor with the control substance) as an indicator. The concentration of the control substance contacted with the olfactory receptor can be appropriately set depending on various conditions, such as the type of olfactory receptor and the type of control substance. The concentration of the control substance contacted with the olfactory receptor may be, for example, 3 to 1000 μM. The concentration of the control substance contacted with the olfactory receptor may typically be 300 μM. Furthermore, for example, for a control substance that exhibits cytotoxicity at 300 μM, the concentration of the control substance contacted with the olfactory receptor may be 3 μM, 10 μM, 30 μM, or 100 μM.

[0038] The presence or absence or degree of activation of the olfactory receptor by the control substance can be determined by comparing the degree of activation D1 with the degree of activation of the olfactory receptor under control conditions (degree of activation D2). An example of a control condition is a condition in which the olfactory receptor is not brought into contact with the control substance.

[0039] The degrees of activation D1 and D2 can both be obtained and used as data reflecting parameters that serve as indicators of olfactory receptor activation. Examples of parameters that serve as indicators of olfactory receptor activation include the amount of intracellular calcium and the amount of intracellular cAMP. In the case of a luciferase assay, examples of data that reflect the amount of intracellular cAMP include luminescence intensity. Data that reflects parameters that serve as indicators of olfactory receptor activation can be used as is, or after being appropriately processed, such as by correction.

[0040] When the degree of activation D1 is high, it can be determined that the olfactory receptor has been activated by the control substance. For example, when the ratio of the degree of activation D1 to the degree of activation D2 (i.e., D1 / D2) is 1.5 or more, 2 or more, 3 or more, 5 or more, 10 or more, 20 or more, 50 or more, or 100 or more, it can be determined that the olfactory receptor has been activated by the control substance. Examples of the ratio of the degree of activation D1 to the degree of activation D2 include the normalized response values ​​described in the Examples.

[0041] Furthermore, the degree of activation of the olfactory receptor by the control substance can be determined by comparing the degree of activation D1 with the degree of activation D2 as an index. For example, the ratio of the degree of activation D1 to the degree of activation D2 (i.e., D1 / D2) can be considered to be the degree of activation of the olfactory receptor by the control substance. Examples of the ratio of the degree of activation D1 to the degree of activation D2 include the normalized response values ​​described in the Examples.

[0042] Control substances include positive controls and negative controls. A "positive control" refers to a substance that has the desired property. A "negative control" refers to a substance that does not have the desired property. Control substances may include at least a positive control.

[0043] The control substance may be a known substance or a novel substance. The control substance may be a natural product or an artificial product. The control substance may be, for example, a compound library created using combinatorial chemistry techniques. Examples of control substances include alcohols, ketones, aldehydes, ethers, esters, hydrocarbons, sugars, organic acids, nucleic acids, amino acids, peptides, and various other organic or inorganic components. Specific examples of control substances include substances whose presence or absence and / or degree of a desired property is known. Examples of substances whose presence or absence and / or degree of a desired property is known include substances listed on The Good Scents Company (http: / / www.thegoodscentscompany.com / ). That is, the control substance may include substances listed on The Good Scents Company. For example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more of the total number of control substances may be selected from substances listed on The Good Scents Company. Any substance listed in The Good Scents Company may be considered, for example, to exhibit the odor listed in its Odor Description (i.e., to be a positive control for the odor listed in its Odor Description). Any substance listed in The Good Scents Company may also be considered, for example, to not exhibit an odor not listed in its Odor Description (i.e., to be a negative control for the odor listed in its Odor Description). Substances for which the presence, absence, and / or degree of a desired characteristic are known also include substances listed in the Atlas of Odor Character Profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985). That is, the control substance may include a substance listed in the Atlas of Odor Character Profiles.Any substance listed in the Atlas of Odor Character Profiles may be considered a positive or negative control for a particular odor, depending on the odor's percentage of applicability. That is, any substance listed in the Atlas of Odor Character Profiles may be considered a positive control for a particular odor if the odor's percentage of applicability is high. Any substance listed in the Atlas of Odor Character Profiles may be considered a negative control for a particular odor if the odor's percentage of applicability is low. A "high percentage of applicability" may mean, for example, a percentage of applicability of 4 or more, 7 or more, 10 or more, 15 or more, or 20 or more. A "low percentage of applicability" may mean, for example, a percentage of applicability of less than 4, 3 or less, 2 or less, 1 or less, or 0.5 or less. A single control substance may be used, or two or more control substances may be used in combination.

[0044] The number of control substances, positive controls, and negative controls may each be, for example, 1 or more, 2 or more, 3 or more, 5 or more, 7 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more, or 10,000 or less, 5,000 or less, 2,000 or less, 1,000 or less, 500 or less, 200 or less, 150 or less, 100 or less, 70 or less, 50 or less, 40 or less, 30 or less, 25 or less, 20 or less, 15 or less, or 10 or less, or any compatible combination thereof. The number of control substances, the number of positive controls, and the number of negative controls may each specifically be, for example, 1 to 10,000, 1 to 1,000, 1 to 100, 1 to 10, 10 to 10,000, 10 to 1,000, 10 to 100, 100 to 10,000, 100 to 1,000, or 1,000 to 10,000. The number of control substances, the number of positive controls, and the number of negative controls may each specifically be, for example, 1 to 10, 10 to 100, 100 to 200, 200 to 500, 500 to 1,000, 1,000 to 2,000, 2,000 to 5,000, or 5,000 to 10,000.

[0045] The proportion of positive control in the control substance may be, for example, 1% or more, 3% or more, 5% or more, 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more, or 100% or less, 99% or less, 97% or less, 95% or less, 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, or 5% or less, or any compatible combination thereof. The ratio of positive controls in a control substance may be, for example, 1-100%, 1-50%, 1-20%, 1-10%, 1-5%, 5-100%, 5-50%, 5-20%, 5-10%, 10-100%, 10-50%, 10-20%, 20-100%, 20-50%, or 50-100%. The ratio of positive controls in a control substance may be, for example, 1-10%, 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, 70-80%, 80-90%, or 90-100%. "Ratio of positive controls in a control substance" refers to the ratio of the number of positive controls to the total number of control substances.

[0046] <1-4> Maximum similarity of stereochemical structure between substances The "maximum degree of stereochemical similarity between substances" refers to the maximum degree of similarity between the stereochemical structures of two or more substances (hereinafter, referred to as substance A and substance B). Specifically, the "maximum degree of stereochemical similarity between substances" refers to the maximum degree of similarity among all pairs of multiple conformations of substance A and multiple conformations of substance B. "Multiple conformations of a substance" refers to two or more conformations possessed by a substance, or in other words, the conformations of two or more conformer isomers possessed by a substance. In other words, the "maximum degree of stereochemical similarity between substances" refers to the maximum degree of similarity among n × m pairs (i.e., pairs A1 and B1 to An and Bm) when substance A has n conformations (A1 to An) and substance B has m conformations (B1 to Bm). The number of conformations possessed by the test substance and the control substance is not particularly limited, as long as they are two or more. The maximum degree of stereochemical similarity is also simply referred to as "maximum similarity."

[0047] The method for generating multiple conformations of a substance is not particularly limited. Multiple conformations of a substance can be generated, for example, by known methods. Specifically, multiple conformations of a substance can be generated using software such as conformation generation software OMEGA (OpenEye). That is, by using the software, multiple conformations of a substance can be generated from structural data of the substance. The software can be used, for example, in accordance with the manufacturer's manual. In the case of OMEGA, multiple conformations may be generated, for example, for macrocyclic compounds (e.g., cyclic compounds with 12 or more ring members) in OMEGA macrocyclic mode, and for other compounds in OMEGA classic mode.

[0048] "Structural data of a substance" refers to data that indicates the structure of a substance. The structural data of a substance is not particularly limited as long as it can generate multiple conformations. The structural data of a substance can be appropriately selected depending on various conditions, such as the type of software used to generate multiple conformations of a substance. The structural data of a substance may be, for example, existing data acquired and used, or data acquired by conversion from a chemical structural formula. Existing data can be acquired, for example, from chemical databases such as PubChem and ChemSpider or from the websites of reagent companies such as SigmaAldrich. Conversion from a chemical structural formula can be performed, for example, using software or a website such as ChemDraw. The acquired structural data of a substance may be used, for example, directly or after appropriate processing, to generate multiple conformations. For example, data in isomeric SMILES format may be canonicalized to absolute SMILES format, converted to 3D structural data in MOL format or SDF format, which includes MOL, and then processed, such as by hydrogen addition or optimization, before being used to generate multiple conformations. Canonicalization of SMILES data and conversion to 3D structural data can be performed using software such as the chemoinformatics software RDKit (http: / / www.rdkit.org). Processing of 3D structural data, such as hydrogenation and optimization, can be performed using software such as the integrated computational chemistry system MOE (CCG).

[0049] The maximum similarity between substances A and B can be obtained, for example, by calculating the similarity between each pair of multiple conformations of substance A and multiple conformations of substance B and obtaining the maximum value among the calculated similarities. The similarity between the pairs may be calculated for all pairs, or may be calculated only for some pairs that include at least the maximum value. For example, pairs with low similarity may be excluded in advance from the calculation of the similarity between the pairs based on an appropriate criterion. The similarity between the pairs may usually be calculated for all pairs.

[0050] The Tanimoto coefficient can be used as a measure of similarity in stereochemical structure. Examples of Tanimoto coefficients include the Shape Tanimoto score, which indicates the similarity of surface shape; the Color Tanimoto score, which indicates the similarity of surface chemical properties; and the Tanimoto Combo score, which indicates the similarity between surface shape and surface chemical properties. The Tanimoto Combo score is calculated as the sum of the Shape Tanimoto score and the Color Tanimoto score. The Tanimoto coefficient can be calculated using software such as the molecular surface shape similarity calculation software ROCS (OpenEye). When calculating the similarity of stereochemical structure using ROCS, the calculated similarity may vary depending on which of the compared substances is used as the query. In this case, any of the calculated similarities may be used to calculate the maximum similarity as long as prediction can be performed with the desired accuracy. For example, the lower or higher of the calculated similarities may be used to calculate the maximum similarity. Alternatively, for example, the average of the calculated similarities may be used to calculate the maximum similarity. The maximum similarity obtained as the maximum value of the Tanimoto coefficient is also referred to as the "maximum similarity based on the Tanimoto coefficient."

[0051] <1-5> Prediction process The prediction can be made based on the maximum similarity between the test substance and the control substance.

[0052] The prediction may be performed, for example, by directly assessing the maximum similarity between the test substance and the control substance. That is, "predicting the presence or absence of a property of interest for a test substance based on the maximum similarity of the stereochemical structures between the test substance and the control substance" may encompass performing the prediction by directly assessing the maximum similarity between the test substance and the control substance. Furthermore, the prediction step may include, for example, a step of directly assessing the maximum similarity between the test substance and the control substance.

[0053] That is, for example, if the maximum similarity between the test substance and the positive control is high, it may be predicted that the test substance has the desired property. "High maximum similarity between the test substance and the positive control" means, when the positive control is a single substance, that the maximum similarity between the test substance and the single positive control is high. "High maximum similarity between the test substance and the positive control" may mean, for example, that the average or maximum maximum similarity between the test substance and the two or more positive controls is high when the positive control is a combination of two or more substances. "High maximum similarity between the test substance and the positive control" may mean, for example, that the number or proportion of positive controls that show a high maximum similarity to the test substance is high when the positive control is a combination of two or more substances. Alternatively, for example, if the maximum similarity between the test substance and the positive control is not high, it may be predicted that the test substance does not have the desired property.

[0054] "High maximum similarity" may mean, for example, that the maximum similarity is equal to or greater than a predetermined value. The predetermined value is not particularly limited as long as prediction can be performed with the desired accuracy. "High maximum similarity" may mean, for example, that the maximum similarity normalized to 0 to 1 is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. "High maximum similarity" may specifically mean, for example, that the maximum similarity based on the Shape Tanimoto score is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. "High maximum similarity" may specifically mean, for example, that the maximum similarity based on the Color Tanimoto score is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. "High maximum similarity" may specifically mean, for example, that the maximum similarity based on the Tanimoto Combo score is 1 or greater, 1.2 or greater, 1.4 or greater, 1.6 or greater, or 1.8 or greater. When the maximum similarity between substances A and B is high, it can also be said that "substance A shows a high maximum similarity to substance B" or "substance B shows a high maximum similarity to substance A."

[0055] "A high average value of maximum similarities" or "a high maximum value of maximum similarities" may mean, for example, that the average value or maximum value of maximum similarities is equal to or greater than a predetermined value. The predetermined value is not particularly limited as long as prediction can be performed with the desired accuracy. "A high average value of maximum similarities" or "a high maximum value of maximum similarities" may mean, for example, that the average value or maximum value of maximum similarities normalized to 0 to 1 is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. "A high average value of maximum similarities" or "a high maximum value of maximum similarities" may specifically mean, for example, that the average value or maximum value of maximum similarities based on the Shape Tanimoto score is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. "A high average value of maximum similarities" or "a high maximum value of maximum similarities" may specifically mean, for example, that the average value or maximum value of maximum similarities based on the Color Tanimoto score is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. "A high average value of maximum similarity" or "a high maximum value of maximum similarity" may specifically mean, for example, that the average value or maximum value of maximum similarity based on the Tanimoto Combo score is 1 or more, 1.2 or more, 1.4 or more, 1.6 or more, or 1.8 or more.

[0056] "A large number of positive controls showing a high maximum similarity to the test substance" may mean, for example, that the number of positive controls showing a high maximum similarity to the test substance is 1 or more, 2 or more, 3 or more, 5 or more, 7 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, or 500 or more.

[0057] "A high proportion of positive controls showing a high degree of maximum similarity to the test substance" may mean, for example, that the proportion of positive controls showing a high degree of maximum similarity to the test substance is 1% or more, 3% or more, 5% or more, 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more. "Proportion of positive controls showing a high degree of maximum similarity to the test substance" refers to the ratio of the number of positive controls showing a high degree of maximum similarity to the test substance to the total number of control substances.

[0058] The prediction may be performed, for example, by clustering the test substance and the control substance based on the maximum similarity between the substances. That is, "predicting the presence or absence of a property of interest for a test substance based on the maximum similarity of the stereochemical structure between the test substance and the control substance" may include performing the prediction by clustering the substances based on the maximum similarity between the test substance and the control substance. The prediction step may also include, for example, clustering the substances based on the maximum similarity between the test substance and the control substance. Clustering may be performed particularly when two or more control substances are used in combination.

[0059] Clustering can be performed using the maximum similarity between the test substance and the control substance as a variable. The variable used for clustering may or may not be the maximum similarity between the test substance and the control substance alone. That is, in addition to the maximum similarity between the test substance and the control substance, other variables may also be used for clustering. The other variables are not particularly limited as long as prediction can be performed with the desired accuracy. Examples of other variables include the similarity (e.g., maximum similarity) of the stereochemical structure between the test substance and the control substance and other substances. In other words, in the prediction step, only the test substance and the control substance may be clustered, or other substances may be clustered in addition to the test substance and the control substance.

[0060] The maximum similarity between substances may be used alone or in combination with a structural similarity between substances other than the maximum similarity for prediction (e.g., clustering). The structural similarity between substances other than the maximum similarity is also referred to as an "additional structural similarity." The combination of the maximum similarity and the additional structural similarity is also referred to as a "mixed similarity." When prediction is performed based on mixed similarity, the "maximum similarity" in the above description of the prediction process may be read as "mixed similarity." That is, for example, when prediction is performed based on mixed similarity, "a high maximum similarity between the test substance and the positive control" may mean a high mixed similarity between the test substance and the positive control. Furthermore, for example, when prediction is performed based on mixed similarity, "a low maximum similarity between the test substance and the positive control" may mean a low mixed similarity between the test substance and the positive control. Examples of additional structural similarity include structural similarity between substances that does not take multiple conformations into account, such as molecular fingerprint similarity. When calculating the mixed similarity, the maximum similarity and the additional structural similarity may be scaled accordingly and then combined. The ratio of the maximum similarity in the mixed similarity is not particularly limited as long as prediction can be performed with the desired accuracy. The ratio of the maximum similarity in the mixed similarity may be, for example, 1% or more, 3% or more, 5% or more, 20% or more, 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, or 90% or more, or 99% or less, 97% or less, 95% or less, 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, or 30% or less, or a consistent combination thereof. Specifically, the ratio of the maximum similarity in the mixed similarity may be, for example, 1 to 99%, 10 to 99%, 30 to 99%, 50 to 99%, 60 to 95%, or 70 to 90%. By performing prediction based on the mixed similarity, the accuracy of prediction may be improved, for example, compared to when prediction is performed based only on the maximum similarity.

[0061] The ratio of the total number of test and control substances to the total number of clustered substances may be, for example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, 97% or more, or 99% or more. Also, the ratio of the number of control substances to the total number of clustered substances may be, for example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, 97% or more, or 99% or more.

[0062] Clustering may be performed once, or may be performed twice or more times, as long as prediction can be performed with the desired accuracy. For example, some substances may be clustered in advance, and the remaining substances may be further clustered based on the obtained clustering results. Specifically, for example, substances other than the test substance may be clustered in advance, and the test substance may be further clustered based on the obtained clustering results. That is, specifically, for example, it may be determined later which cluster of substances other than the test substance the test substance will be clustered into. When two or more test substances are used in combination, the test substances may be clustered together in one go, or may be clustered twice or more times.

[0063] The clustering method is not particularly limited. Clustering can be performed by, for example, a known method. Examples of such methods include hierarchical cluster analysis and dimensionality reduction. Examples of hierarchical cluster analysis include Ward's method, nearest neighbor method, furthest neighbor method, and group average method. Examples of hierarchical cluster analysis include Ward's method. Examples of distances between substances used in hierarchical cluster analysis include Euclidean distance, Mahalanobis distance, Manhattan distance, Chebyshev distance, Minkowski distance, Canberra distance, distance based on cosine similarity, angular distance, distance based on Pearson's correlation coefficient, and distance based on the extended Jaccard coefficient. Examples of distances between substances used in hierarchical cluster analysis include Euclidean distance. Specifically, hierarchical cluster analysis may be performed by Ward's method using Euclidean distance, for example.Examples of dimension reduction methods include random projection, principal component analysis (PCA), linear discriminant analysis (LDA), isometric mapping (Isomap), locally linear embedding (LLE), modified LLE (MLLE), Hessian-based LLE (HLLE), spectral embedding, local tangent space alignment (LTSA), multidimensional scaling (MDS), t-distributed stochastic neighbor embedding (t-SNE), random forest embedding, uniform manifold approximation and projection (UMAP), kernel PCA, and autoencoder. Examples of dimension reduction methods include t-SNE. These methods may be used alone or in combination.

[0064] The number of clusters is not particularly limited as long as prediction can be performed with the desired accuracy. The number of clusters may be, for example, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more, or may be 100 or less, 50 or less, 30 or less, 25 or less, 20 or less, 15 or less, 12 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less, or any consistent combination thereof. Specifically, the number of clusters may be, for example, 2 to 30, 3 to 20, or 4 to 15.

[0065] For example, if the test substance is clustered into a cluster that is likely to have the desired property, it may be predicted that the test substance has the desired property. For example, if the test substance is clustered into a cluster that is likely to have the desired property, it may be determined that the maximum similarity between the test substance and the positive control is high. For example, if the test substance is not clustered into a cluster that is likely to have the desired property, it may be predicted that the test substance does not have the desired property. For example, if the test substance is not clustered into a cluster that is likely to have the desired property, it may be determined that the maximum similarity between the test substance and the positive control is not high. A cluster that is likely to have the desired property is also referred to as a "positive cluster." As a result of clustering, only one positive cluster may be generated, or two or more positive clusters may be generated. Examples of clusters that are likely to have the desired property include clusters that include a positive control. A cluster that includes a positive control may include one or more positive controls. A cluster that includes a positive control may or may not include substances other than the positive control. A cluster that includes a positive control may or may not include, for example, a negative control. A cluster that includes a positive control may be, for example, a cluster with a high proportion of positive controls. A "cluster with a high proportion of positive controls" may refer to, for example, a cluster in which the proportion of positive controls is 1% or more, 3% or more, 5% or more, 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more. The "proportion of positive controls" in a cluster refers to the ratio of the number of positive controls contained in the cluster to the number of control substances contained in the cluster. A cluster containing a positive control may be, for example, a cluster with a high degree of a target characteristic. A "cluster with a high degree of a target characteristic" may refer to, for example, a cluster containing a positive control with the highest degree of a target characteristic. A "cluster with a high degree of a target characteristic" may refer to, for example, a cluster with the highest average degree of a target characteristic. When two or more positive clusters are generated, for example, clusters that satisfy the above-mentioned criteria may be selected as positive clusters in order.The "average degree of the characteristic of interest" in a cluster refers to the average degree of the characteristic of interest of all control substances included in the cluster.

[0066] The prediction method of the present invention may further include a step of evaluating the prediction results. That is, by evaluating the target property of the test substance, it is possible to confirm whether the target substance actually has the target property. Specifically, for example, by evaluating the target property of a test substance predicted to have the target property, it is possible to confirm whether the target substance actually has the target property. That is, the step of evaluating the prediction results may be, for example, a step of confirming the presence or absence of the target property for the test substance predicted to have the target property. The method of evaluating the prediction results is not particularly limited. The description regarding the identification of the presence or absence of the target property in a control substance can be applied mutatis mutandis to the method of evaluating the prediction results.

[0067] <2> The design method of the present invention according to the first aspect of the present invention The design method of the present invention is a method for designing a substance having desired properties. The terms "design of a substance" and "design of the structure of a substance" may be used interchangeably. Designing a substance having desired properties will hereinafter also be referred to simply as "design." Design can be carried out based on the maximum similarity of the stereochemical structure between the substance to be designed and a reference substance. In other words, the design method of the present invention may include a step of designing the substance to be designed based on the maximum similarity of the stereochemical structure between the substance to be designed and a reference substance. This step will also be referred to as the "design step."

[0068] Design can be performed, for example, so that the substance to be designed is predicted to have the desired properties based on the prediction method of the present invention. In other words, the substance to be designed can be designed to have a structure that is predicted to have the desired properties based on the prediction method of the present invention. For example, the structure of an existing substance may be modified so that it is predicted to have the desired properties based on the prediction method of the present invention. Alternatively, for example, the structures of a large number of compounds may be designed and those predicted to have the desired properties based on the prediction method of the present invention may be selected. Specifically, design can be performed, for example, so that the substance to be designed is clustered into a cluster that is likely to have the desired properties (e.g., a cluster including a positive control).

[0069] (B) Second Aspect of the Invention The second aspect of the present invention, specifically the method for producing a prediction model of the present invention and the prediction method of the present invention relating to the second aspect of the present invention, will be described below.

[0070] <1> A method for producing a prediction model according to a second aspect of the present invention The method for producing a predictive model of the present invention is a method for producing a model that predicts the presence or absence of a target component in a test substance. "Predicting the presence or absence of a target component in a test substance" means predicting whether or not the test substance has the target component. Predicting the presence or absence of a target component in a test substance will hereinafter also be referred to simply as "prediction." A model that predicts the presence or absence of a target component in a test substance will hereinafter also be referred to simply as a "prediction model."

[0071] <1-1> Prediction model The predictive model is a model that predicts the presence or absence of a target component in a test substance. That is, the predictive model can be used for prediction. Specifically, the predictive model can be used for prediction in the manner described in the prediction method of the present invention.

[0072] The prediction model may include a decision tree. A decision tree or a model including the same is also referred to as a "tree model." The decision tree is not particularly limited as long as it outputs a conclusion that serves as an indicator for prediction. Prediction can be performed based on test olfactory receptor activation data for a test substance. The test olfactory receptor activation data for a test substance will hereinafter also be referred to simply as "test olfactory receptor activation data." That is, the decision tree may output a conclusion that serves as an indicator for prediction based on the test olfactory receptor activation data (in other words, using the test olfactory receptor activation data as a variable). An example of a conclusion that serves as an indicator for prediction is a classification result regarding the presence or absence of a target component in a test substance. That is, the decision tree may output, for example, a classification result regarding the presence or absence of a target component in a test substance based on the test olfactory receptor activation data. The "classification result regarding the presence or absence of a target component in a test substance" refers to a classification result that suggests whether or not the test substance has the target component. The classification result regarding the presence or absence of a target component in a test substance is specifically obtained as a result of classifying the test substance into one of the leaf nodes included in the decision tree. That is, the decision tree may specifically classify the test substance into one of the leaf nodes contained in the decision tree based on the test olfactory receptor activation data.

[0073] <1-2> Components of the purpose The term "component of interest" refers to the component to be predicted, such as an aroma characteristic or a molecular structure.

[0074] The aroma, aroma characteristics, and the presence or absence of aroma characteristics are as described in the first embodiment of the present invention.

[0075] "Molecular structure" refers to parameters related to the structure of a substance. The type of molecular structure is not particularly limited. Examples of molecular structures include partial molecular structures. Examples of partial molecular structures include functional groups, skeletons, bonds, and atoms. Specific examples of molecular structures include carbonyl groups, acyl groups, aldehyde groups, ketone groups, carboxyl groups, carboxamide groups, alkanoyl groups, benzoyl groups, alkoxycarbonyl groups, phenoxycarbonyl groups, imide groups, enone groups, alkyl groups, alkenyl groups, hydroxyl groups, amino groups, imino groups, aryl groups, oxo groups, alkoxy groups, phenoxy groups, alkylenedioxy groups, thiol groups, sulfo groups, nitro groups, ester bonds, ether bonds, amide bonds, glycosidic bonds, nitrogen atoms, oxygen atoms, sulfur atoms, halogen atoms, monocyclic skeletons, heterocyclic skeletons, and terpenoid skeletons. Examples of heterocyclic skeletons include heterocyclic skeletons containing heteroatoms such as nitrogen, sulfur, and oxygen. The heterocyclic skeleton may contain one or more heteroatoms. Specific examples of heterocyclic skeletons include nitrogen-containing heterocyclic skeletons such as pyrazine skeletons and pyrrole skeletons, and nitrogen- and sulfur-containing heterocyclic skeletons such as thiazole skeletons. The molecular structure may be one type of molecular structure or a combination of two or more types of molecular structures. In other words, "presence or absence of a molecular structure" may mean the presence or absence of any one type of molecular structure, or the presence or absence of two or more types of molecular structures (i.e., a pattern of which molecular structures are present and which molecular structures are absent for two or more types of molecular structures).

[0076] <1-3> Test substance The term "test substance" refers to a substance that is the subject of prediction of the presence or absence of a target component. In other words, the term "test substance" refers to a substance used as a candidate for a substance having a target component in a method for screening for a substance having the target component. The test substance is not particularly limited as long as test olfactory receptor activation data can be used.

[0077] "Test olfactory receptor activation data by a test substance" refers to data regarding the activation of a test olfactory receptor by a test substance. "Activation of a test olfactory receptor by a test substance" may be used interchangeably with "response of a test olfactory receptor to a test substance." Test olfactory receptor activation data includes data indicating whether or not a test olfactory receptor is activated by a test substance, and data indicating the degree of activation of a test olfactory receptor by a test substance. Test olfactory receptor activation data particularly includes data indicating the degree of activation of a test olfactory receptor by a test substance. "The degree of activation of a test olfactory receptor by a test substance" may refer to the strength with which a test substance activates a test olfactory receptor. Test olfactory receptor activation data is specifically used in branches included in a decision tree.

[0078] "Test olfactory receptor" refers to an olfactory receptor used in a branch included in a decision tree. "An olfactory receptor is used in a branch included in a decision tree" may mean that olfactory receptor activation data for that olfactory receptor (i.e., data regarding the activation of that olfactory receptor by a test substance) is used in a branch included in the decision tree. Test olfactory receptors include the following olfactory receptors. The test olfactory receptor may be one type of olfactory receptor, or a combination of two or more types of olfactory receptors.

[0079] The olfactory receptors and the genes encoding them (olfactory receptor genes) are as described in the first embodiment of the present invention.

[0080] The test olfactory receptor activation data may or may not be publicly known. If the test olfactory receptor activation data is not publicly known, it may be obtained as appropriate before performing the prediction. The method for obtaining the test olfactory receptor activation data is not particularly limited. The test olfactory receptor activation data can be obtained, for example, by a known method for identifying the presence or absence, or the degree of activation of an olfactory receptor by a substance. Specifically, the test olfactory receptor activation data can be obtained, for example, by contacting the test olfactory receptor with a test substance and measuring the presence or absence, or the degree of activation of the test olfactory receptor due to contact with the test substance. The contact between the test olfactory receptor and the test substance and the measurement of the presence or absence, or the degree of activation of the test olfactory receptor due to this can be performed, for example, with reference to a screening method for a substance that exhibits a target odor using the response of the olfactory receptor as an indicator (e.g., JP 2019-037197 A). The test olfactory receptor may be used by being supported on cells such as animal cells. Activation of the test olfactory receptor can be measured, for example, using an increase in intracellular calcium or intracellular cAMP as an indicator. Techniques for measuring the amount of intracellular cAMP include, for example, ELISA and reporter assay. An example of a reporter assay is luciferase assay. According to the reporter assay, the amount of intracellular cAMP can be measured by utilizing a reporter gene (luciferase gene, etc.) that is constructed so that its expression depends on the amount of cAMP. An example of a technique for measuring the amount of intracellular calcium is calcium imaging.

[0081] Specifically, contact between the test olfactory receptor and the test substance and measurement of the presence or absence or degree of activation of the test olfactory receptor due to this can be carried out, for example, by the following procedure.

[0082] That is, the presence or absence or degree of activation of the test olfactory receptor by the test substance can be determined by contacting the test olfactory receptor with the test substance and using the degree of activation of the test olfactory receptor (degree of activation D1) when the contact is carried out (i.e., under the conditions for contacting the test olfactory receptor with the test substance) as an index. The concentration of the test substance contacted with the test olfactory receptor can be appropriately set depending on various conditions, such as the type of test olfactory receptor and the type of test substance. The concentration of the test substance contacted with the test olfactory receptor may be, for example, 3 to 1000 μM. The concentration of the test substance contacted with the test olfactory receptor may typically be 300 μM. Furthermore, for example, for a test substance that exhibits cytotoxicity at 300 μM, the concentration of the test substance contacted with the test olfactory receptor may be 3 μM, 10 μM, 30 μM, or 100 μM.

[0083] The presence or absence or degree of activation of the test olfactory receptor by the test substance can be determined by comparing the degree of activation D1 with the degree of activation of the test olfactory receptor under control conditions (degree of activation D2). Control conditions include conditions in which the test olfactory receptor is not brought into contact with the test substance.

[0084] The degrees of activation D1 and D2 can both be obtained and used as data reflecting parameters that serve as indicators of the activation of the test olfactory receptor. Examples of parameters that serve as indicators of the activation of the test olfactory receptor include the amount of intracellular calcium and the amount of intracellular cAMP. In the case of a luciferase assay, examples of data that reflect the amount of intracellular cAMP include luminescence intensity. Data that reflect parameters that serve as indicators of the activation of the test olfactory receptor can be used as is, or after being processed, such as by appropriate correction.

[0085] When the degree of activation D1 is high, it can be determined that the test olfactory receptor has been activated by the test substance. For example, when the ratio of the degree of activation D1 to the degree of activation D2 (i.e., D1 / D2) is 1.5 or more, 2 or more, 3 or more, 5 or more, 10 or more, 20 or more, 50 or more, or 100 or more, it can be determined that the test olfactory receptor has been activated by the test substance. Examples of the ratio of the degree of activation D1 to the degree of activation D2 include the normalized response values ​​described in the Examples.

[0086] Furthermore, the degree of activation of the test olfactory receptor by the test substance can be determined by using the comparison result between the degree of activation D1 and the degree of activation D2 as an index. For example, the ratio of the degree of activation D1 to the degree of activation D2 (i.e., D1 / D2) can be considered to be the degree of activation of the test olfactory receptor by the test substance. Examples of the ratio of the degree of activation D1 to the degree of activation D2 include the normalized response values ​​described in the Examples.

[0087] The test substance may be a known substance or a novel substance. The test substance may be a natural product or an artificial product. For example, the test substance may be a compound library generated using combinatorial chemistry techniques. Examples of test substances include alcohols, ketones, aldehydes, ethers, esters, hydrocarbons, sugars, organic acids, nucleic acids, amino acids, peptides, and various other organic or inorganic components. Test substances also include, in particular, existing food additives. "Existing food additives" refers to substances already approved for use as food additives. The test substance may be a single test substance or a combination of two or more test substances. The test substance may be selected to include, for example, substances such as those exemplified above, such as existing food additives. That is, the test substance may be, for example, a single existing food additive, a combination of two or more food additives, or a combination of one or more food additives with one or more other substances. The phrase "using two or more test substances in combination" means predicting the presence or absence of the target constituent element for each of two or more test substances.

[0088] In one embodiment, the test substance may be a mixture.

[0089] When the test substance is a mixture, "the presence or absence or degree of activation of the test olfactory receptor by the test substance" means the presence or absence or degree of activation of the test olfactory receptor by the entire mixture, regardless of the presence or absence or degree of activation of the test olfactory receptor by each substance that makes up the mixture.

[0090] Furthermore, when the test substance is a mixture, "the presence or absence of the target component in the test substance" refers to the presence or absence of the target component in the entire mixture, regardless of whether the target component is present in each of the substances that make up the mixture. That is, for example, when the test substance is a mixture, "the test substance has the target aroma characteristics" refers to the mixture as a whole having the target aroma characteristics, regardless of whether each of the substances that make up the mixture has the target aroma characteristics. Furthermore, "the test substance has the target molecular structure" refers to the mixture as a whole having the target molecular structure (i.e., at least one substance selected from the substances that make up the mixture has the target molecular structure), regardless of whether substances that make up the mixture other than the at least one substance also have the target molecular structure.

[0091] <1-4> Generation of decision trees The decision tree can be generated by machine learning. That is, the prediction model creation method of the present invention may include a step of generating a decision tree by machine learning. This step is also referred to as a "decision tree generation step."

[0092] The conditions for machine learning are not particularly limited as long as a decision tree that can perform prediction with the desired accuracy can be obtained.

[0093] Machine learning can be performed using a dataset including control substance component data and control olfactory receptor activation data. The control substance component data is hereinafter also referred to simply as "component data." The control substance control olfactory receptor activation data is hereinafter also referred to simply as "control olfactory receptor activation data."

[0094] Machine learning can be performed, for example, using the component data as the dependent variable and the control olfactory receptor activation data as the explanatory variable.

[0095] The machine learning method is not particularly limited as long as it can generate a decision tree. Examples of the machine learning method include CART (Classification and Regression Trees), CHAID (Chi-squared Automatic Interaction Detection), ID3 (Iterative Dichotomiser 3), and C4.5. In particular, CART is an example of the machine learning method.

[0096] Machine learning may be performed, for example, by ensemble learning. Examples of ensemble learning include bagging and boosting. Examples of bagging include random forests and extremely randomized trees (ExtraTrees). Examples of boosting include XGboost and LightGBM. When machine learning is performed by ensemble learning, the decision tree included in the prediction model may be a decision tree after ensemble learning. That is, for example, when bagging is performed, the prediction model may include multiple decision trees obtained by bagging. In this case, multiple decision trees can be used in combination in the prediction process. That is, bagging can generate multiple decision trees as weak learners, and a combination of these multiple weak learners can be used as a strong learner. Furthermore, for example, when boosting is performed, the prediction model may include a decision tree whose learning level has been improved by boosting. That is, boosting can generate and use a decision tree as a strong learner based on a decision tree generated as a weak learner.

[0097] The term "control substance" refers to a substance that can be used to generate a decision tree as an indicator of the presence or absence of a target component. The control substance is not particularly limited as long as its component data and control olfactory receptor activation data are available.

[0098] "Component data of a control substance" refers to data related to a target component in a control substance. When the target component is an aroma characteristic, the component data is also referred to as "aroma characteristic data." When the target component is a molecular structure, the component data is also referred to as "molecular structure data." Examples of component data include data indicating the presence or absence of a target component in a control substance.

[0099] "Control olfactory receptor activation data of a control substance" means data on the activation of a control olfactory receptor by a control substance. Examples of control olfactory receptor activation data include data showing whether or not a control olfactory receptor is activated by a control substance, and data showing the degree of activation of a control olfactory receptor by a control substance. Examples of control olfactory receptor activation data particularly include data showing the degree of activation of a control olfactory receptor by a control substance.

[0100] The term "control olfactory receptor" refers to an olfactory receptor used in generating a decision tree. "An olfactory receptor is used in generating a decision tree" may mean that olfactory receptor activation data for that olfactory receptor (i.e., data regarding the activation of that olfactory receptor by a control substance) is used in generating a decision tree. Examples of control olfactory receptors include the olfactory receptors described above. That is, the control olfactory receptor may include the olfactory receptor described above. For example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more of the total number of control olfactory receptors may be selected from the olfactory receptors described above. A combination of two or more olfactory receptors, including a test olfactory receptor, is used as the control olfactory receptor. The control olfactory receptor may consist of the test olfactory receptor, or may include other olfactory receptors in addition to the test olfactory receptor. In other words, some or all of the control olfactory receptors are selected as test olfactory receptors. That is, among the control olfactory receptors, an olfactory receptor used in a branch included in the decision tree is selected as the test olfactory receptor.

[0101] The number of control olfactory receptors is not particularly limited as long as a decision tree capable of performing prediction with the desired accuracy can be obtained. The number of control olfactory receptors can be appropriately set depending on various conditions, such as the type of target component and the machine learning method.

[0102] The number of control olfactory receptors may be, for example, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, or 500 or more, or 2000 or less, 1500 or less, 1000 or less, 500 or less, 400 or less, 300 or less, 200 or less, 150 or less, or 100 or less, or any compatible combination thereof. Specifically, the number of control olfactory receptors may be, for example, 50 to 2000, 100 to 1000, or 300 to 500.

[0103] The component data may or may not be publicly known. If the component data is not publicly known, it can be acquired as appropriate before generating the decision tree. The method for acquiring the component data is not particularly limited. The component data can be identified, for example, by a known method for identifying the presence or absence or level of a component of a substance. The presence or absence or level of a target aroma characteristic in a control substance can be identified, for example, by sensory evaluation by an expert panel. The presence or absence of a target molecular structure in a control substance can be identified, for example, by a known method for identifying the structure of a substance. Such methods include nuclear magnetic resonance (NMR), electron spin resonance (ESR), ultraviolet-visible-near-infrared spectroscopy (UV-Vis-NIR), infrared spectroscopy (IR), Raman spectroscopy, and mass spectrometry (MS). These methods may be used alone or in appropriate combination.

[0104] The control olfactory receptor activation data may or may not be publicly known. If the control olfactory receptor activation data is not publicly known, it is sufficient to obtain the control olfactory receptor activation data as appropriate before generating the decision tree. The method for obtaining the control olfactory receptor activation data is not particularly limited. The control olfactory receptor activation data can be identified, for example, by a known method for identifying the presence or absence, or the degree of activation of an olfactory receptor by a substance. Specifically, the control olfactory receptor activation data can be obtained, for example, by contacting the control olfactory receptor with a control substance and measuring the presence or absence, or the degree of activation of the control olfactory receptor due to contact with the control substance. The above-mentioned description of the contact of the test olfactory receptor with a test substance and the measurement of the presence or absence, or the degree of activation of the test olfactory receptor due to contact can be applied mutatis mutandis to the contact of the control olfactory receptor with the control substance and the measurement of the presence or absence, or the degree of activation of the test olfactory receptor.

[0105] The control substance may be a combination of two or more substances, including a positive control and a negative control. A "positive control" refers to a substance that contains the component of interest. A "negative control" refers to a substance that does not contain the component of interest.

[0106] The control substance may be a known substance or a novel substance. The control substance may be a natural product or an artificial product. The control substance may be, for example, a compound library created using combinatorial chemistry techniques. Examples of control substances include alcohols, ketones, aldehydes, ethers, esters, hydrocarbons, sugars, organic acids, nucleic acids, amino acids, peptides, and various other organic or inorganic components. Specific examples of control substances include substances in which the presence or absence and / or level of a target component is known. Examples of substances in which the presence or absence and / or level of a target component is known include substances listed on The Good Scents Company (http: / / www.thegoodscentscompany.com / ). That is, the control substance may include substances listed on The Good Scents Company. For example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more of the total number of control substances may be selected from substances listed on The Good Scents Company. Any substance listed in The Good Scents Company may be considered, for example, to exhibit the odor listed in its Odor Description (i.e., to be a positive control for the odor listed in its Odor Description). Any substance listed in The Good Scents Company may also be considered, for example, to not exhibit an odor not listed in its Odor Description (i.e., to be a negative control for the odor listed in its Odor Description). Substances for which the presence or absence and / or level of a target component are known also include substances listed in the Atlas of Odor Character Profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985). That is, the control substance may include a substance listed in the Atlas of Odor Character Profiles.Any substance listed in the Atlas of Odor Character Profiles may be considered a positive or negative control for a particular odor, depending on the odor's percentage of applicability. That is, any substance listed in the Atlas of Odor Character Profiles may be considered a positive control for a particular odor if the percentage of applicability for that odor is high. Any substance listed in the Atlas of Odor Character Profiles may be considered a negative control for a particular odor if the percentage of applicability for that odor is low. A "high percentage of applicability" may mean, for example, a percentage of applicability of 4 or more, 7 or more, 10 or more, 15 or more, or 20 or more. A "low percentage of applicability" may mean, for example, a percentage of applicability of less than 4, 3 or less, 2 or less, 1 or less, or 0.5 or less. Any of the substances listed above may be considered a positive control for the molecular structure of that substance. Furthermore, any of the substances exemplified above may be considered as a negative control for a molecular structure that the substance does not have.

[0107] In one embodiment, the control substance may be a mixture.

[0108] When the control substance is a mixture, "the presence or absence, or the degree of activation of the control olfactory receptor by the control substance" means the presence or absence, or the degree of activation of the control olfactory receptor by the entire mixture, and does not matter whether or not each substance constituting the mixture activates the control olfactory receptor.

[0109] Furthermore, when the control substance is a mixture, "the presence or absence of the target component in the control substance" refers to the presence or absence of the target component in the entire mixture, regardless of whether the target component is present in each of the substances constituting the mixture. That is, for example, when the control substance is a mixture, "the control substance has the target aroma characteristics" refers to the mixture as a whole having the target aroma characteristics, regardless of whether each of the substances constituting the mixture has the target aroma characteristics. Furthermore, "the control substance has the target molecular structure" refers to the mixture as a whole having the target molecular structure (i.e., at least one substance selected from the substances constituting the mixture has the target molecular structure), regardless of whether substances constituting the mixture other than the at least one substance also have the target molecular structure.

[0110] The number of control substances, the number of positive controls, the number of negative controls, and the ratio thereof are not particularly limited as long as a decision tree capable of performing prediction with the desired accuracy can be obtained. The number of control substances, the number of positive controls, the number of negative controls, and the ratio thereof can be appropriately set depending on various conditions, such as the type of target component and the machine learning method.

[0111] The number of control substances may be, for example, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, 2000 or more, 3000 or more, 5000 or more, 10,000 or more, 20,000 or more, 50,000 or more, or 100,000 or more; or 1,000,000 or less, 500,000 or less, 200,000 or less, 100,000 or less, 50,000 or less, 20,000 or less, 10,000 or less, 5000 or less, 3,000 or less, 2000 or less, 1500 or less, 1000 or less, or 500 or less, or any compatible combination thereof. The number of control substances may be, for example, 100 to 1,000,000, 200 to 500,000, 500 to 100,000, or 1,000 to 20,000. The number of control substances may be, for example, 100 to 200, 200 to 500, 500 to 1,000, 1,000 to 2,000, 2,000 to 5,000, 5,000 to 10,000, 10,000 to 20,000, 20,000 to 50,000, 50,000 to 100,000, or 100,000 to 200,000.

[0112] The number of positive controls and the number of negative controls can be, for example, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, 2000 or more, 3000 or more, 5000 or more, 10000 or more, 20000 or more, 5000 The number of positive controls may be 0 or more, or 100,000 or more, or 1,000,000 or less, 500,000 or less, 200,000 or less, 100,000 or less, 50,000 or less, 20,000 or less, 10,000 or less, 5,000 or less, 3,000 or less, 2,000 or less, 1,500 or less, 1,000 or less, 500 or less, 200 or less, 150 or less, 100 or less, 70 or less, or 50 or less, or a compatible combination thereof. The number of positive controls and the number of negative controls may be, specifically, for example, 5 to 1,000,000, 100 to 1,000,000, 200 to 500,000, 500 to 100,000, or 1,000 to 20,000. The number of positive controls and the number of negative controls may be, for example, 5 to 10, 10 to 100, 100 to 200, 200 to 500, 500 to 1000, 1000 to 2000, 2000 to 5000, 5000 to 10000, 10000 to 20000, 20000 to 50000, 50000 to 100000, or 100000 to 200000.

[0113] The positive control ratio and negative control ratio in the control substance may be, for example, more than 0%, 1% or more, 3% or more, 5% or more, 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more, or less than 100%, 99% or less, 97% or less, 95% or less, 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, or 5% or less, or any compatible combination thereof. The ratio of positive controls and the ratio of negative controls in a control substance may be, for example, 1-99%, 1-50%, 1-20%, 1-10%, 1-5%, 5-99%, 5-50%, 5-20%, 5-10%, 10-99%, 10-50%, 10-20%, 20-99%, 20-50%, or 50-99%. The ratio of positive controls and the ratio of negative controls in a control substance may be, for example, 1-10%, 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, 70-80%, 80-90%, or 90-99%. "The ratio of positive controls in a control substance" refers to the ratio of the number of positive controls to the total number of control substances. "Ratio of negative controls among control substances" means the ratio of the number of negative controls to the total number of control substances. The total number of control substances may be the sum of the number of positive controls and the number of negative controls.

[0114] By performing machine learning in this manner, a decision tree can be generated. The decision tree includes two or more leaf nodes. One or more of the leaf nodes included in the decision tree are positive leaf nodes. In other words, the decision tree includes one or more positive leaf nodes. A "positive leaf node" refers to a leaf node that is highly likely to contain a target component. Specifically, a "positive leaf node" refers to a leaf node where a substance classified into the leaf node is highly likely to contain the target component.

[0115] The number of leaf nodes included in a decision tree is not particularly limited as long as prediction can be performed with the desired accuracy. The number of leaf nodes included in a decision tree may be, for example, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more, or 100 or less, 50 or less, 30 or less, 25 or less, 20 or less, 15 or less, 12 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less, or any combination thereof that is not contradictory. Specifically, the number of leaf nodes included in a decision tree may be, for example, 2 to 30, 3 to 20, or 4 to 15.

[0116] The number of positive leaf nodes included in a decision tree is not particularly limited as long as prediction can be performed with the desired accuracy. A decision tree may include only one positive leaf node, or may include two or more. The number of positive leaf nodes included in a decision tree may be, for example, 1 or more, 2 or more, 3 or more, 4 or more, or 5 or more, or 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, 4 or less, 3 or less, or 2 or less, or any consistent combination thereof. The number of positive leaf nodes included in a decision tree may specifically be, for example, 1 to 10, 1 to 6, or 1 to 4.

[0117] There are no particular limitations on which leaf nodes are designated as positive leaf nodes, as long as prediction can be performed with the desired accuracy. Examples of positive leaf nodes include leaf nodes containing positive controls. Leaf nodes containing positive controls may contain one or more positive controls. Leaf nodes containing positive controls may or may not contain negative controls. Leaf nodes containing positive controls may be, for example, leaf nodes with a high ratio of positive controls. A "leaf node with a high ratio of positive controls" may refer to, for example, a leaf node with a ratio of positive controls of 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more. The "proportion of positive controls" in a certain leaf node refers to the ratio of the number of positive controls contained in the leaf node to the number of control substances contained in the leaf node. Alternatively, a desired number of leaf nodes may be designated as positive leaf nodes, for example, in descending order of the ratio of positive controls.

[0118] <2> The prediction method of the present invention according to the second aspect of the present invention The prediction method of the present invention is a method for predicting the presence or absence of a target component in a test substance. The prediction can be performed using the prediction model of the present invention. Specifically, the prediction can be performed based on the test olfactory receptor activation data of the test substance and the prediction model of the present invention. That is, the prediction method of the present invention may include a step of predicting the presence or absence of a target component in the test substance based on the test olfactory receptor activation data of the test substance and the prediction model of the present invention. This step is also referred to as a "prediction step."

[0119] Furthermore, by predicting the presence or absence of a target component in a test substance, it is possible to screen for a substance having the target component. That is, a test substance predicted to have the target component can be selected as a substance having the target component, thereby screening for a substance having the target component. That is, one aspect of the prediction method of the present invention may be a method of screening for a substance having the target component. That is, the prediction method of the present invention may further include a step of selecting a test substance predicted to have the target component as a substance having the target component. That is, the screening method may be a method of screening for a substance having the target component, including a step of predicting the presence or absence of the target component in the test substance based on the test olfactory receptor activation data of the test substance and a prediction model, and a step of selecting the test substance predicted to have the target component as a substance having the target component. In other words, the screening method may be a method of screening for a substance having the target component, including a step of predicting the presence or absence of the target component in the test substance using the prediction method of the present invention, and a step of selecting the test substance predicted to have the target component as a substance having the target component.

[0120] The prediction method of the present invention may further include, prior to the prediction step, a step of producing a prediction model by the prediction model production method of the present invention.

[0121] By applying the decision tree included in the prediction model to the test olfactory receptor activation data of the test substance, a conclusion that serves as an indicator of prediction, specifically, a classification result regarding the presence or absence of a target component in the test substance, can be output. Specifically, by applying the decision tree included in the prediction model to the test olfactory receptor activation data of the test substance, the test substance can be classified into one of the leaf nodes included in the decision tree.

[0122] For example, if the test substance is classified into a positive leaf node, it may be predicted that the test substance has the component of interest. Alternatively, if the test substance is not classified into a positive leaf node, it may be predicted that the test substance does not have the component of interest. Furthermore, when bagging is performed, classification results from multiple decision trees may be comprehensively evaluated. For example, if the ratio of the number of decision trees in which the test substance is classified into a positive leaf node to the total number of decision trees is high, it may be predicted that the test substance has the component of interest. "A high ratio of the number of decision trees in which the test substance is classified into a positive leaf node to the total number of decision trees" may mean, for example, that the ratio of the number of decision trees in which the test substance is classified into a positive leaf node to the total number of decision trees is greater than 50%, 60% or more, 70% or more, 80% or more, or 90% or more.

[0123] The prediction method of the present invention may further include a step of evaluating the prediction results. That is, by evaluating the target component of the test substance, it is possible to confirm whether the target substance actually contains the target component. Specifically, for example, by evaluating the target component of a test substance predicted to contain the target component, it is possible to confirm whether the target substance actually contains the target component. That is, the step of evaluating the prediction results may be, for example, a step of confirming the presence or absence of the target component in a test substance predicted to contain the target component. The method for evaluating the prediction results is not particularly limited. The description of the method for obtaining component data of a control substance can be applied mutatis mutandis to the method for evaluating the prediction results.

[0124] (C) Third Aspect of the Invention The third aspect of the present invention, specifically the method for producing a prediction model of the present invention and the prediction method of the present invention relating to the third aspect of the present invention, will be described below.

[0125] <1> A method for producing a prediction model according to a third aspect of the present invention. The prediction model production method of the present invention is a method for producing a model that predicts the degree of conformance of a test substance to a target aroma characteristic. Predicting the degree of conformance of a test substance to a target aroma characteristic will hereinafter also be referred to simply as "prediction." A model that predicts the degree of conformance of a test substance to a target aroma characteristic will hereinafter also be referred to simply as a "prediction model."

[0126] <1-1> Prediction model The predictive model is a model that predicts the degree of conformance of a test substance to a target odor characteristic. That is, the predictive model can be used for prediction. Specifically, the predictive model can be used for prediction in the manner described in the prediction method of the present invention.

[0127] The prediction model may include a regression equation. The regression equation is not particularly limited as long as it outputs a conclusion that serves as an indicator for prediction. Prediction can be performed based on test olfactory receptor activation data of a test substance. The test olfactory receptor activation data of a test substance will hereinafter also be referred to simply as "test olfactory receptor activation data." That is, the regression equation may output a conclusion that serves as an indicator for prediction based on the test olfactory receptor activation data (in other words, using the test olfactory receptor activation data as a variable). An example of a conclusion that serves as an indicator for prediction is a predicted value of the degree of conformance of the test substance to the target odor characteristic. That is, the regression equation may, for example, output a predicted value of the degree of conformance of the test substance to the target odor characteristic based on the test olfactory receptor activation data. The regression equation may, for example, be a linear regression equation.

[0128] <1-2> Degree of suitability for the desired aroma characteristics The term "target aroma characteristic" refers to the aroma characteristic that is the target of the matching prediction.

[0129] The aroma and aroma characteristics are as described in the first aspect of the present invention.

[0130] "Compatibility with an aroma characteristic" refers to qualitative closeness to the target aroma characteristic. In other words, "high compatibility with an aroma characteristic" means having the property of exhibiting an aroma close to the target aroma itself. For example, "high compatibility with the aroma characteristic 'STARWBERRY'" means having the property of exhibiting an aroma close to that of STARWBERRY itself. "High compatibility with an aroma characteristic" can also be referred to as "having a high compatibility with an aroma characteristic." The compatibility with an aroma characteristic can be expressed as a percentage of applicability calculated according to the criteria described in Atlas of odor character profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985). Specifically, the percentage of applicability value can be obtained as a score between 0 and 100 by having multiple expert panels rate the intensity of the target odor of the target substance on a six-point scale (0 to 5 points: 0, Absent; 1, Slightly; 3, Moderately; 5, Extremely), and then calculating the geometric mean of the "percentage (%) of expert panels that scored 1 or higher" and the "average score of all expert panels divided by 5."

[0131] The aroma may be one type of aroma or a combination of two or more types of aromas. That is, the "degree of suitability to an aroma characteristic" may mean the degree of suitability to any one type of aroma characteristic, or the degree of suitability to each of two or more types of aroma characteristics.

[0132] <1-3> Test substance The term "test substance" refers to a substance whose suitability for a desired aroma characteristic is predicted. In other words, the term "test substance" refers to a substance used as a candidate for a substance that is highly suitable for a desired aroma characteristic in a method for screening for a substance that is highly suitable for a desired aroma characteristic. There are no particular limitations on the test substance, as long as test olfactory receptor activation data is available.

[0133] "Test olfactory receptor activation data by a test substance" refers to data regarding the activation of a test olfactory receptor by a test substance. "Activation of a test olfactory receptor by a test substance" may be used interchangeably with "the response of a test olfactory receptor to a test substance." Test olfactory receptor activation data includes data indicating whether or not a test olfactory receptor is activated by a test substance, and data indicating the degree of activation of a test olfactory receptor by a test substance. Test olfactory receptor activation data particularly includes data indicating the degree of activation of a test olfactory receptor by a test substance. "The degree of activation of a test olfactory receptor by a test substance" may refer to the strength with which a test substance activates a test olfactory receptor. Specifically, the test olfactory receptor activation data is used by substituting it as a variable in a regression equation.

[0134] "Test olfactory receptor" refers to an olfactory receptor used in a regression equation. "An olfactory receptor is used in a regression equation" may mean that olfactory receptor activation data for that olfactory receptor (i.e., data regarding the activation of that olfactory receptor by a test substance) is substituted as a variable into the regression equation and used. Test olfactory receptors include the following olfactory receptors. The test olfactory receptor may be one type of olfactory receptor, or a combination of two or more types of olfactory receptors.

[0135] The olfactory receptors and the genes encoding them (olfactory receptor genes) are as described in the first embodiment of the present invention.

[0136] The test olfactory receptor activation data may or may not be publicly known. If the test olfactory receptor activation data is not publicly known, it may be obtained as appropriate before performing the prediction. The method for obtaining the test olfactory receptor activation data is not particularly limited. The test olfactory receptor activation data can be obtained, for example, by a known method for identifying the presence or absence, or the degree of activation of an olfactory receptor by a substance. Specifically, the test olfactory receptor activation data can be obtained, for example, by contacting the test olfactory receptor with a test substance and measuring the presence or absence, or the degree of activation of the test olfactory receptor due to contact with the test substance. The contact between the test olfactory receptor and the test substance and the measurement of the presence or absence, or the degree of activation of the test olfactory receptor due to this can be performed, for example, with reference to a screening method for a substance that exhibits a target odor using the response of the olfactory receptor as an indicator (e.g., JP 2019-037197 A). The test olfactory receptor may be used by being supported on cells such as animal cells. Activation of the test olfactory receptor can be measured, for example, using an increase in intracellular calcium or intracellular cAMP as an indicator. Techniques for measuring the amount of intracellular cAMP include, for example, ELISA and reporter assay. An example of a reporter assay is luciferase assay. According to the reporter assay, the amount of intracellular cAMP can be measured by utilizing a reporter gene (luciferase gene, etc.) that is constructed so that its expression depends on the amount of cAMP. An example of a technique for measuring the amount of intracellular calcium is calcium imaging.

[0137] Specifically, contact between the test olfactory receptor and the test substance and measurement of the presence or absence or degree of activation of the test olfactory receptor due to this can be carried out, for example, by the following procedure.

[0138] That is, the presence or absence or degree of activation of the test olfactory receptor by the test substance can be determined by contacting the test olfactory receptor with the test substance and using the degree of activation of the test olfactory receptor (degree of activation D1) when the contact is carried out (i.e., under the conditions for contacting the test olfactory receptor with the test substance) as an index. The concentration of the test substance contacted with the test olfactory receptor can be appropriately set depending on various conditions, such as the type of test olfactory receptor and the type of test substance. The concentration of the test substance contacted with the test olfactory receptor may be, for example, 3 to 1000 μM. The concentration of the test substance contacted with the test olfactory receptor may typically be 300 μM. Furthermore, for example, for a test substance that exhibits cytotoxicity at 300 μM, the concentration of the test substance contacted with the test olfactory receptor may be 3 μM, 10 μM, 30 μM, or 100 μM.

[0139] The presence or absence or degree of activation of the test olfactory receptor by the test substance can be determined by comparing the degree of activation D1 with the degree of activation of the test olfactory receptor under control conditions (degree of activation D2). Control conditions include conditions in which the test olfactory receptor is not brought into contact with the test substance.

[0140] The degrees of activation D1 and D2 can both be obtained and used as data reflecting parameters that serve as indicators of the activation of the test olfactory receptor. Examples of parameters that serve as indicators of the activation of the test olfactory receptor include the amount of intracellular calcium and the amount of intracellular cAMP. In the case of a luciferase assay, examples of data that reflect the amount of intracellular cAMP include luminescence intensity. Data that reflect parameters that serve as indicators of the activation of the test olfactory receptor can be used as is, or after being processed, such as by appropriate correction.

[0141] When the degree of activation D1 is high, it can be determined that the test olfactory receptor has been activated by the test substance. For example, when the ratio of the degree of activation D1 to the degree of activation D2 (i.e., D1 / D2) is 1.5 or more, 2 or more, 3 or more, 5 or more, 10 or more, 20 or more, 50 or more, or 100 or more, it can be determined that the test olfactory receptor has been activated by the test substance. Examples of the ratio of the degree of activation D1 to the degree of activation D2 include the normalized response values ​​described in the Examples.

[0142] Furthermore, the degree of activation of the test olfactory receptor by the test substance can be determined by using the comparison result between the degree of activation D1 and the degree of activation D2 as an index. For example, the ratio of the degree of activation D1 to the degree of activation D2 (i.e., D1 / D2) can be considered to be the degree of activation of the test olfactory receptor by the test substance. Examples of the ratio of the degree of activation D1 to the degree of activation D2 include the normalized response values ​​described in the Examples.

[0143] The test substance may be a known substance or a novel substance. The test substance may be a natural product or an artificial product. For example, the test substance may be a compound library generated using combinatorial chemistry techniques. Examples of test substances include alcohols, ketones, aldehydes, ethers, esters, hydrocarbons, sugars, organic acids, nucleic acids, amino acids, peptides, and various other organic or inorganic components. Test substances also include, in particular, existing food additives. "Existing food additives" refers to substances already approved for use as food additives. The test substance may be a single test substance or a combination of two or more test substances. The test substance may be selected to include, for example, substances such as those exemplified above, such as existing food additives. That is, the test substance may be, for example, a single existing food additive, a combination of two or more food additives, or a combination of one or more food additives with one or more other substances. The phrase "using two or more test substances in combination" means predicting the suitability of each of the two or more test substances for the desired aroma characteristics.

[0144] In one embodiment, the test substance may be a mixture.

[0145] When the test substance is a mixture, "the presence or absence or degree of activation of the test olfactory receptor by the test substance" means the presence or absence or degree of activation of the test olfactory receptor by the entire mixture, regardless of the presence or absence or degree of activation of the test olfactory receptor by each substance that makes up the mixture.

[0146] Furthermore, when the test substance is a mixture, the "degree of suitability of the test substance to the target aroma characteristics" refers to the degree of suitability of the mixture as a whole to the target aroma characteristics, regardless of the degree of suitability of each of the substances that make up the mixture to the target aroma characteristics. For example, when the test substance is a mixture, "the test substance has a high degree of suitability for the target aroma characteristics" means that the mixture as a whole has a high degree of suitability for the target aroma characteristics, regardless of whether each of the substances that make up the mixture has a high degree of suitability for the target aroma characteristics.

[0147] <1-4>Generating a regression equation The regression equation can be generated by machine learning. That is, the prediction model production method of the present invention may include a step of generating a regression equation by machine learning. This step is also referred to as a "regression equation generation step."

[0148] The conditions for machine learning are not particularly limited as long as a regression equation that can perform prediction with the desired accuracy can be obtained.

[0149] Machine learning can be performed using a dataset including odor profile data and control olfactory receptor activation data for a control substance. The odor profile data for a control substance is hereinafter also referred to simply as "odor profile data." The control olfactory receptor activation data for a control substance is hereinafter also referred to simply as "control olfactory receptor activation data."

[0150] Machine learning can be performed, for example, using the odor characteristic data as the response variable and the control olfactory receptor activation data as the explanatory variable.

[0151] The machine learning method is not particularly limited as long as it can generate a regression equation. An example of the machine learning method is regression analysis. Examples of regression analysis include simple regression analysis and multiple regression analysis. Examples of regression analysis include multiple regression analysis. Examples of regression analysis that can generate a linear regression equation include linear regression analysis. Examples of linear regression analysis include simple linear regression analysis and multiple linear regression analysis. Examples of linear regression analysis include multiple linear regression analysis.

[0152] Machine learning may be performed, for example, by ensemble learning. Examples of ensemble learning include bagging and boosting. When machine learning is performed by ensemble learning, the regression equation included in the prediction model may be the regression equation after ensemble learning. That is, for example, when bagging is performed, the prediction model may include multiple regression equations obtained by bagging. In this case, multiple regression equations can be used in combination in the prediction process. That is, bagging can generate multiple regression equations as weak learners, and a combination of these multiple weak learners can be used as a strong learner. Furthermore, for example, when boosting is performed, the prediction model may include a regression equation whose learning level has been improved by boosting. That is, boosting can generate and use a regression equation as a strong learner based on the regression equation generated as a weak learner.

[0153] The term "control substance" refers to a substance that can be used to generate a regression equation as an index of the degree of fit to the target odor characteristics. There are no particular limitations on the control substance, as long as its odor characteristic data and control olfactory receptor activation data are available.

[0154] "Odor characteristic data of a control substance" refers to data showing the degree of conformance of a control substance with a target odor characteristic. Examples of data showing the degree of conformance of a control substance with a target odor characteristic include percentage of applicability values ​​calculated according to the standards set forth in Atlas of odor character profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985).

[0155] "Control olfactory receptor activation data of a control substance" means data on the activation of a control olfactory receptor by a control substance. Examples of control olfactory receptor activation data include data showing whether or not a control olfactory receptor is activated by a control substance, and data showing the degree of activation of a control olfactory receptor by a control substance. Examples of control olfactory receptor activation data particularly include data showing the degree of activation of a control olfactory receptor by a control substance.

[0156] The term "control olfactory receptor" refers to an olfactory receptor used in generating the regression equation. "An olfactory receptor is used in generating the regression equation" may mean that olfactory receptor activation data for the olfactory receptor (i.e., data regarding the activation of the olfactory receptor by a control substance) is used in generating the regression equation. Examples of control olfactory receptors include the olfactory receptors described above. That is, the control olfactory receptor may include the olfactory receptors described above. For example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more of the total number of control olfactory receptors may be selected from the olfactory receptors described above. A combination of two or more olfactory receptors, including a test olfactory receptor, is used as the control olfactory receptor. The control olfactory receptor may consist of the test olfactory receptor, or may include other olfactory receptors in addition to the test olfactory receptor. In other words, some or all of the control olfactory receptors are selected as test olfactory receptors. That is, among the control olfactory receptors, the olfactory receptors used in the regression equation are selected as test olfactory receptors. In other words, machine learning may be performed using the control olfactory receptor activation data for some or all of the control olfactory receptors as explanatory variables. That is, "machine learning is performed using the control olfactory receptor activation data as explanatory variables" may mean that machine learning is performed using the control olfactory receptor activation data for some or all of the control olfactory receptors as explanatory variables. For example, among the control olfactory receptors, an olfactory receptor having a high correlation coefficient between the odor characteristic data and the control olfactory receptor activation data may be selected as the test olfactory receptor. In other words, among the control olfactory receptors, machine learning may be performed using the control olfactory receptor activation data for an olfactory receptor having a high correlation coefficient between the odor characteristic data and the control olfactory receptor activation data as an explanatory variable. "A high correlation coefficient between the odor characteristic data and the control olfactory receptor activation data" may mean, for example, that the absolute value of the correlation coefficient between the odor characteristic data and the control olfactory receptor activation data is greater than 0.1, greater than 0.15, greater than 0.2, greater than 0.25, or greater than 0.3. An olfactory receptor having a high correlation coefficient between the odor characteristic data and the control olfactory receptor activation data can be identified by calculating the correlation coefficient between the odor characteristic data and the control olfactory receptor activation data.That is, the regression equation generation step may include, for example, a step of calculating a correlation coefficient between the odor characteristic data and the control olfactory receptor activation data prior to machine learning.

[0157] The number of control olfactory receptors is not particularly limited as long as a regression equation that enables prediction with the desired accuracy can be obtained, and can be appropriately set depending on various conditions, such as the type of odor characteristic of interest and the machine learning method.

[0158] The number of control olfactory receptors may be, for example, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, or 500 or more, or 2000 or less, 1500 or less, 1000 or less, 500 or less, 400 or less, 300 or less, 200 or less, 150 or less, or 100 or less, or any compatible combination thereof. Specifically, the number of control olfactory receptors may be, for example, 50 to 2000, 100 to 1000, or 300 to 500.

[0159] The number of control olfactory receptors used in the regression equation (i.e., the number of test olfactory receptors) may be, for example, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, or 500 or more, or 2000 or less, 1500 or less, 1000 or less, 500 or less, 400 or less, 300 or less, 200 or less, 150 or less, 100 or less, 70 or less, 50 or less, 40 or less, 30 or less, 25 or less, or 20 or less, or any compatible combination thereof. Specifically, the number of control olfactory receptors may be, for example, 10 to 1000, 15 to 500, or 20 to 200.

[0160] The odor characteristic data may or may not be publicly known. If the odor characteristic data is not publicly known, it can be acquired as appropriate before generating the regression equation. The method for acquiring the odor characteristic data is not particularly limited. The odor characteristic data can be identified, for example, by a known method for identifying the degree of suitability of a substance to an odor characteristic. The degree of suitability of a control substance to a target odor characteristic can be identified, for example, by sensory evaluation by an expert panel. Specifically, for example, the percentage of applicability to a target odor characteristic can be calculated according to the standards described in Atlas of odor character profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985).

[0161] The control olfactory receptor activation data may or may not be publicly known. If the control olfactory receptor activation data is not publicly known, it is sufficient to obtain the control olfactory receptor activation data as appropriate before generating the regression equation. The method for obtaining the control olfactory receptor activation data is not particularly limited. The control olfactory receptor activation data can be identified, for example, by a known method for identifying the presence or absence, or the degree of activation of an olfactory receptor by a substance. Specifically, the control olfactory receptor activation data can be obtained, for example, by contacting the control olfactory receptor with a control substance and measuring the presence or absence, or the degree of activation of the control olfactory receptor due to contact with the control substance. The above-mentioned description of the contact of the test olfactory receptor with a test substance and the measurement of the presence or absence, or the degree of activation of the test olfactory receptor due to contact can be applied mutatis mutandis to the contact of the control olfactory receptor with the control substance and the measurement of the presence or absence, or the degree of activation of the test olfactory receptor.

[0162] As a control substance, a combination of two or more substances is used.

[0163] The control substance may be a known substance or a novel substance. The control substance may be a natural product or an artificial product. For example, the control substance may be a compound library created using combinatorial chemistry techniques. Examples of control substances include alcohols, ketones, aldehydes, ethers, esters, hydrocarbons, sugars, organic acids, nucleic acids, amino acids, peptides, and various other organic or inorganic components. Specific examples of control substances include substances whose degree of conformance to the desired odor characteristics is known. Examples of substances whose degree of conformance to the desired odor characteristics are substances listed in the Atlas of Odor Character Profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985). That is, the control substance may include substances listed in the Atlas of Odor Character Profiles. For example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more of the total number of control substances may be selected from substances listed in the Atlas of Odor Character Profiles.

[0164] In one embodiment, the control substance may be a mixture.

[0165] When the control substance is a mixture, "the presence or absence, or the degree of activation of the control olfactory receptor by the control substance" means the presence or absence, or the degree of activation of the control olfactory receptor by the entire mixture, and does not matter whether or not each substance constituting the mixture activates the control olfactory receptor.

[0166] Furthermore, when the control substance is a mixture, "the degree of suitability of the control substance with the target aroma characteristics" refers to the degree of suitability of the mixture as a whole with the target aroma characteristics, regardless of the degree of suitability of each of the substances that make up the mixture with the target aroma characteristics. That is, for example, when the control substance is a mixture, "the control substance has a high degree of suitability with the target aroma characteristics" means that the mixture as a whole has a high degree of suitability with the target aroma characteristics, regardless of whether each of the substances that make up the mixture has a high degree of suitability with the target aroma characteristics.

[0167] The number of control substances is not particularly limited as long as a regression equation that enables prediction with the desired accuracy can be obtained, and can be appropriately set depending on various conditions, such as the type of odor characteristic of interest and the machine learning method.

[0168] The number of control substances may be, for example, 30 or more, 40 or more, 50 or more, 70 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, 2000 or more, 3000 or more, 5000 or more, 10000 or more, 20000 or more, 50000 or more, or 100000 or more, The number of control substances may be, for example, 00,000 or less, 500,000 or less, 200,000 or less, 100,000 or less, 50,000 or less, 20,000 or less, 10,000 or less, 5,000 or less, 3,000 or less, 2,000 or less, 1,500 or less, 1,000 or less, 500 or less, 400 or less, 300 or less, 200 or less, 150 or less, 100 or less, 70 or less, or 50 or less, or any combination thereof that is compatible. The number of control substances may be, for example, 30 to 100, 100 to 200, 200 to 500, 500 to 1000, 1000 to 2000, 2000 to 5000, 5000 to 10000, 10000 to 20000, 20000 to 50000, 50000 to 100000, or 100000 to 200000. The number of control substances may be, for example, 30 to 1000, 50 to 500, or 100 to 200.

[0169] <2> The prediction method of the present invention according to the third aspect of the present invention The prediction method of the present invention is a method for predicting the degree of suitability of a test substance for a target odor characteristic. The prediction can be performed using the prediction model of the present invention. Specifically, the prediction can be performed based on the test olfactory receptor activation data of the test substance and the prediction model of the present invention. That is, the prediction method of the present invention may include a step of predicting the degree of suitability of the test substance for a target odor characteristic based on the test olfactory receptor activation data of the test substance and the prediction model of the present invention. This step is also referred to as a "prediction step."

[0170] Furthermore, by predicting the suitability of a test substance for a target aroma characteristic, it is possible to screen for substances that have a high suitability for the target aroma characteristic. That is, test substances predicted to have a high suitability for the target aroma characteristic can be selected as substances that have a high suitability for the target aroma characteristic, thereby screening for substances that have a high suitability for the aroma characteristic. That is, one embodiment of the prediction method of the present invention may be a method for screening for substances that have a high suitability for the target aroma characteristic. That is, the prediction method of the present invention may further include a step of selecting test substances predicted to have a high suitability for the target aroma characteristic as substances that have a high suitability for the target aroma characteristic. That is, the screening method may be a method for screening for substances that have a high suitability for the target aroma characteristic, including a step of predicting the suitability of the test substance for the target aroma characteristic based on the test olfactory receptor activation data of the test substance and a prediction model, and a step of selecting test substances predicted to have a high suitability for the target aroma characteristic as substances that have a high suitability for the target aroma characteristic. In other words, the screening method may be a method for screening for substances with high suitability for a target aroma characteristic, comprising the steps of predicting the suitability of a test substance for a target aroma characteristic using the prediction method of the present invention, and selecting test substances predicted to have high suitability for the target aroma characteristic as substances with high suitability for the target aroma characteristic.

[0171] The prediction method of the present invention may further include, prior to the prediction step, a step of producing a prediction model by the prediction model production method of the present invention.

[0172] By applying the regression equation included in the prediction model to the test olfactory receptor activation data of the test substance, a conclusion serving as an indicator of prediction, specifically, a predicted value of the test substance's conformance to the desired odor characteristics, can be output. Specifically, by applying the regression equation included in the prediction model to the test olfactory receptor activation data of the test substance, a predicted value of the test substance's conformance to the desired odor characteristics can be output. Furthermore, when bagging is performed, the output results from multiple regression equations may be comprehensively evaluated. When bagging is performed, the "predicted value of the test substance's conformance to the desired odor characteristics" may mean, for example, the average value of the predicted values ​​of the test substance's conformance to the desired odor characteristics output from multiple regression equations.

[0173] When the predicted value of the applicability of the test substance to the target aroma characteristic is high, it can be predicted that the test substance has a high applicability to the target aroma characteristic. "A high predicted value of applicability to the target aroma characteristic" may mean, for example, that the percentage of applicability to the target aroma characteristic is 4 or more, 7 or more, 10 or more, 15 or more, or 20 or more.

[0174] The prediction method of the present invention may further include a step of evaluating the prediction results. That is, by evaluating the degree of conformance of the test substance with the target aroma characteristics, it is possible to confirm whether the target substance actually has a high degree of conformance with the target aroma characteristics. Specifically, for example, by evaluating the degree of conformance of a test substance predicted to have a high degree of conformance with the target aroma characteristics, it is possible to confirm whether the target substance actually has a high degree of conformance with the target aroma characteristics. That is, the step of evaluating the prediction results may be, for example, a step of confirming the degree of conformance of a test substance predicted to have a high degree of conformance with the target aroma characteristics. The method for evaluating the prediction results is not particularly limited. The description of the method for obtaining aroma characteristic data of a control substance can be applied mutatis mutandis to the method for evaluating the prediction results. [Example]

[0175] Example A The present invention will be described in more detail below with reference to non-limiting examples according to the first aspect of the present invention. Although the substances used in the following examples are referred to as "test substances," these substances can also be used as control substances in the prediction method and design method of the present invention.

[0176] <1> Generation of cells expressing human olfactory receptors <1-1> Construction of expression vectors for human olfactory receptors The olfactory receptors are 352 types of human olfactory receptors (OR1A1, OR1A2, OR1B1, OR1C1, OR1D2, OR1D5, OR1E1, OR1F1, OR1F12, OR1G1, OR1I1, OR1J1, OR1J2, OR1J4, OR1K1, OR1L1, OR1L3, OR1L4, OR1L8, OR1M1, OR1N1, OR1N2, OR1Q1, OR1R1P, OR1S1, OR2A1, OR2A2, OR2A4, OR2A5, OR2A12, OR2A14, OR2A25, OR2AE1, OR2AG1, OR2AG2, OR2AJ1P, OR2AK2, OR2AP1, OR2AT4, OR2B2, OR2B3, OR2B6, OR2B11, OR2C1, OR2C3, OR2D2, OR2D3, OR2F1, OR2G2, OR2G3, OR2G6, OR2H1, OR2H2, OR2J2, OR2J3, OR2K2, O R2L2, OR2L8, OR2L13, OR2M2, OR2M4, OR2M7, OR2S2, OR2T1, OR2T2, OR2T5, OR2T6, OR2T8, OR2T10, OR2T11, OR2T27, OR2T34, OR2V2, OR2W1, OR2W3, OR2Y1, O R2Z1, OR3A1, OR3A2, OR3A3, OR3A4, OR4A5, OR4A15, OR4A16, OR4A47, OR4B1, OR4C3, OR4C5, OR4C6, OR4C11, OR4C12, OR4C13, OR4C15, OR4C16, OR4C46, OR4 D1, OR4D2, OR4D5, OR4D6, OR4D9, OR4D10, OR4D11, OR4E2, OR4F3, OR4F5, OR4F6, OR4F14P, OR4F15, OR4G11P, OR4H12P, OR4K1, OR4K2, OR4K5, OR4K13, OR4K 14, OR4K15, OR4K17, OR4L1, OR4M1, OR4N2, OR4N4, OR4N5, OR4P4, OR4Q3, OR4S1, OR4S2, OR4X1, OR4X2, OR5A1, OR5A2, OR5AC2, OR5AK2, OR5AK3P, OR5AN1, O R5AP2, OR5AR1, OR5AS1, OR5AU1, OR5B2, OR5B3, OR5B12, OR5B17, OR5B21, OR5C1, OR5D13, OR5D14, OR5D16, OR5D18, OR5F1, OR5H1, OR5H2, OR5H6, OR5H14,OR5I1、OR5J2、OR5K1、OR5K3、OR5K4、OR5L2、OR5M3、OR5M8、OR5M9、OR5M10、OR5M11、OR5P3、OR5R1、OR5T1、OR5T2、OR5T3、OR5V1、OR5W2、OR6A2、OR6B1、OR6B2、OR6C1、OR6C2、OR6C3、OR6C4、OR6C6、OR6C65、OR6C66P、OR6C68、OR6C70、OR6C74、OR6C75、OR6C76、OR6F1、OR6J1、OR6K2、OR6K3、OR6K6、OR6M1、OR6N1、OR6N2、OR6P1、OR6Q1、OR6S1、OR6T1、OR6V1、OR6X1、OR6Y1、OR7A3P、OR7A5、OR7A10、OR7A17、OR7C1、OR7C2、OR7D2、OR7D4、OR7E24、OR7G1、OR7G2、OR7G3、OR8A1、OR8B3、OR8B4、OR8B8、OR8B12、OR8D1、OR8D2、OR8D4、OR8G2、OR8G5、OR8H3、OR8I2、OR8J1、OR8J3、OR8K1、OR8K3、OR8K5、OR8S1、OR8U1、OR9A4、OR9G1、OR9G4、OR9I1、OR9K2、OR9Q1、OR9Q2、OR10A3、OR10A4、OR10A5、OR10A6、OR10A7、OR10AD1、OR10AG1、OR10C1、OR10D3、OR10D4P、OR10G2、OR10G3、OR10G4、OR10G6、OR10G7、OR10G9、OR10H2、OR10H4、OR10J1、OR10J3、OR10J5、OR10K1、OR10K2、OR10P1、OR10Q1、OR10R2、OR10S1、OR10T2、OR10V1、OR10W1、OR10X1、OR10Z1、OR11A1、OR11G2、OR11H4、OR11H6、OR11H12、OR11L1、OR12D2、OR12D3、OR13A1、OR13C2、OR13C3、OR13C4、OR13C8、OR13D1、OR13F1、OR13G1、OR13H1、OR13J1、OR14A2、OR14A16、OR14C36、OR14I1、OR14J1、OR14K1、OR14L1P、OR51A1P、OR51A4、OR51A7、OR51B2、OR51B4、OR51B5、OR51B6、OR51D1、OR51E1、OR51E2, OR51F1, OR51F2, OR51F5P, OR51G1, OR51G2, OR51H1, OR51I1, OR51I2, OR51L1, OR51M1, OR51Q1 , OR51S1, OR51T1, OR51V1, OR52A1, OR52A4, OR52A5, OR52B2, OR52B4, OR52B6, OR52D1, OR52E2, OR52E4, OR52E5, OR52E8, OR52H1, OR52I2, OR52J3, OR52K2, OR52L2P, OR52M1, OR52N1, OR52N2, OR52N4, OR52N5, OR52P2P, OR52R1, OR52W1, OR52Z1P, OR56A1, OR56A3, OR56A4, OR56A5, OR56B1, OR56B2P, and OR56B4 were adopted.

[0177] We purchased 352 human olfactory receptor genes from the TrueClone cDNA Clone Collection (OriGene). Using primers designed based on the sequence information registered in GenBank, we amplified subcloning fragments for each of the 352 human olfactory receptor genes by PCR using the purchased human olfactory receptor genes as templates. The amplified subcloning fragments for each gene were subcloned downstream of the Rho tag sequence in the Rho-pME18S vector (K. Kajiya et al., Journal of Neuroscience, 15 August 2001, 21 (16) 6018-6025) using the EcoRI and XhoI sites, yielding 352 expression vectors for human olfactory receptors.

[0178] <1-2> Preparation of olfactory receptor-expressing cells HEK293T cells expressing each of the 352 olfactory receptors were prepared using the following procedure. The gene mixture shown in Table 1 and the transfection reagent mixture shown in Table 2 were prepared and left to stand at room temperature for 5 minutes. pcDNA3.1-microbat RTP1s is an expression vector for bat RTP1s, pcDNA3.1-Golf is an expression vector for human Golf, and pcDNA3.1-Ric8B is an expression vector for rat Ric8B (JP Patent Publication No. 2019-037197). The gene mixture and transfection reagent mixture were mixed and dispensed in 12.5 μL aliquots into each well of a poly-D-lysine-coated 384-well plate and left to stand in a clean bench for 15 minutes. HEK293T cells (2.5 × 10 ) were seeded in a 10 cm dish the day before. 6 1.2 x 10 cells / 10 cm dish 5 The solution was adjusted to 1000 cells / mL, and 25 μL of each was seeded into each well of a 384-well plate and cultured overnight in an incubator maintained at 37°C and 5% CO2. In this way, 352 cultures of HEK293T cells transfected with the expression vectors shown in Table 1 and expressing the genes encoded by those expression vectors were obtained.

[0179] [Table 1]

[0180] [Table 2]

[0181] <2> Creation of a database of human olfactory receptor activity <2-1> Luciferase assay The olfactory receptor-expressing cells were used to measure the response of the olfactory receptor to the test substance.

[0182] The 352 olfactory receptors expressed in HEK293T cells conjugate with Golf to activate adenylate cyclase, thereby increasing intracellular cAMP levels. In this example, a luciferase reporter gene assay was used to measure the response of olfactory receptors to test substances. This assay monitors the increase in intracellular cAMP levels as an increase in luminescence intensity derived from firefly luciferase. The "luciferase reporter gene assay" is also referred to as the "luciferase assay." Firefly luciferase is expressed from the firefly luciferase gene carried by the pGL4.29[luc2P / CRE / Hygro] Vector in a manner dependent on the amount of intracellular cAMP. Additionally, the luminescence intensity derived from Renilla luciferase was used as an internal standard to correct for errors in gene transfer efficiency and cell number in each well. Renilla luciferase is constitutively expressed from the Renilla luciferase gene carried in the pGL4.74[hRluc / TK] Vector under the control of the TK promoter.

[0183] 941 test substances were selected from the list of substances listed in The Good Scents Company (http: / / www.thegoodscentscompany.com / ). The medium was removed from the 352 cultures obtained in <1-2> above, and 15 μL of the 941 test substance solutions was added to each, yielding 352 × 941 reaction solutions. Each test substance solution was prepared by dissolving the test substance in CD293 (Life Technologies, Inc.). The test substance concentration in the test substance solution was generally 300 μM. However, for test substances that showed cytotoxicity at 300 μM, the test substance concentration in the test substance solution was set to 3 μM, 10 μM, 30 μM, or 100 μM. For a small number of test substances, the test substance concentration in the test substance solution was set to 1000 μM. The reaction solution was placed in an incubator maintained at 37°C and 5% CO2, and the cells were cultured for 4 hours to allow sufficient intracellular expression of the firefly luciferase gene. The luminescence value derived from intracellular firefly luciferase was measured and designated as the "Luc value." The luminescence value derived from intracellular Renilla luciferase was also measured and designated as the "hRLuc value." The luminescence value derived from each luciferase was measured using Dual-Glo TM Measurement was performed using a luciferase assay system (Promega) according to the product's operating manual.

[0184] <2-2> Calculation of olfactory receptor activity The luminescence value (Luc value) derived from firefly luciferase induced by stimulation with the test substance was divided by the luminescence value (hRluc value) derived from Renilla luciferase in the same well to obtain the "Luc / hRluc value." The Luc / hRluc value in cells stimulated with the test substance was divided by the Luc / hRluc value in cells not stimulated with the test substance to obtain the "fold increase." Furthermore, the fold increase in cells transfected with an olfactory receptor expression vector was divided by the fold increase in cells transfected with the empty vector Rho-pME18S to obtain the "normalized response." The common logarithm of the normalized response was used as the "olfactory receptor activity," a quantitative index of the response strength of the olfactory receptor to the test substance. Hereinafter, when olfactory receptor activity is expressed as -1, 0, or 1, this means that the common logarithms of the normalized response are -1, 0, or 1, i.e., the normalized response is 0.1, 1, or 10, meaning that the response of olfactory receptor-introduced cells to test substance stimulation is 1 / 10, 1, or 10 times stronger, respectively, than the response of empty vector-introduced cells to test substance stimulation. For simplicity, the effect that differences in test substance concentration in the test substance solution may have on olfactory receptor activity was ignored.

[0185] <2-3> Molecular structure information of the test substance The isomeric SMILES of the test substance was obtained from PubChem (https: / / pubchem.ncbi.nlm.nih.gov / ). The isomeric SMILES was canonicalized using the open-source cheminformatics software RDKit (http: / / www.rdkit.org), converted to 3D structure data, and saved in SDF format.

[0186] <2-4> Information on the aroma characteristics of the test substance The information on the odor properties of the test substances was taken from the descriptors listed in Odor Description of Organoleptic Properties by The Good Scents Company (http: / / www.thegoodscentscompany.com).

[0187] <3> Scoring the similarity of stereochemical structures considering multiple conformations <3-1> Generation of multiple conformations of aroma compounds Using the SDF data obtained in <2-3> above, hydrogen addition and optimization of the structural data was performed using the integrated computational chemistry system MOE (CCG) under conditions of pH 7.0. Multiple conformations were generated using the conformation generation software OMEGA (OpenEye), with OMEGA macrocyclic for macrocyclic compounds and OMEGA classic for other compounds.

[0188] <3-2> Calculation of similarity of stereochemical structure Using the molecular surface shape similarity calculation software ROCS (OpenEye), similarities were calculated for all conformational pairs of all test substances generated in <3-1> above, focusing on surface shape and surface chemical properties. The stereochemical structure similarity between test substances was calculated by taking the maximum similarity among all conformational pairs between the substances. Due to the specifications of similarity calculations using ROCS, the similarity calculated does not necessarily match depending on which of the conformational pairs is used as the query. In such cases, a symmetric matrix was created by taking the average of the two values ​​as the similarity for the pair, and finally, a stereochemical structure similarity matrix that takes into account multiple conformations between all test substances was obtained.

[0189] <4> Molecular structure representation based on stereochemical similarity considering multiple conformations <4-1> Cluster analysis using stereochemical structural similarity considering multiple conformations The stereochemical similarity matrix, which took into account multiple conformations among all test substances, was considered to be a matrix consisting of multidimensional stereochemical information feature vectors for each test substance. Euclidean distances between each test substance were calculated, and hierarchical cluster analysis was performed using Ward's method. The stereochemical similarity matrix was sorted according to the results of the hierarchical cluster analysis, and a heat map showing the degree of similarity was created (Figure 1). The dendrogram generated by the hierarchical cluster analysis and the classification of all test substances into nine clusters based on the dendrogram are shown on the left side of the heat map, using different shades of color.

[0190] <4-2> Visualization of stereochemical structure similarity matrix by dimensional reduction considering multiple conformations The stereochemical similarity matrix, which takes into account multiple conformations among all test substances, was considered to be a matrix consisting of multidimensional stereochemical feature vectors for each test substance. A dimensionality reduction method was used to visualize the structural similarity relationships among the test substances. The dimensionality reduction method used was t-SNE (Van der Maaten et al., 2008, Visualizing Data Using t-SNE, Journal of Machine Learning Research 9: 2579-2605). Figure 2 shows the results of plotting all test substances in a 3D space. The shading of each point corresponds to the color shading of the clustering results in <4-1> above. In this 3D map (hereafter referred to as chemical structural similarity space), compounds with similar stereochemical structures are located close to each other, while compounds with distant stereochemical structures are located farther apart.

[0191] <4-3> Olfactory receptor activation characteristics in stereochemical structure similarity space The points representing each test substance in the chemical structure similarity space created in <4-2> above were color-coded according to the level of olfactory receptor activity calculated in <2-2> above and plotted as heat maps (Figures 3-5). In the figures, the points representing each test substance are darker for higher olfactory receptor activity and lighter for lower olfactory receptor activity. In the figures, "Response" indicates olfactory receptor activity. Figure 3 shows the results of color-coding according to the level of OR4S2 activity (i.e., olfactory receptor activity for the olfactory receptor OR4S2), Figure 4 shows the results of color-coding according to the level of OR5K1 activity (i.e., olfactory receptor activity for the olfactory receptor OR5K1), and Figure 5 shows the results of color-coding according to the level of OR10G4 activity (i.e., olfactory receptor activity for the olfactory receptor OR10G4). Figures 3-5 demonstrate that the black points are localized in a narrow range in the stereochemical structure similarity space, i.e., test substances exhibiting each olfactory receptor activation property are localized in a narrow range. Therefore, it became clear that the olfactory receptor activation properties of a substance can be predicted using the degree of stereochemical structural similarity that takes multiple conformations into account as an index.

[0192] <4-4> Aroma characteristics in stereochemical structure similarity space The points representing each test substance in the chemical structure similarity space created in <4-2> above are color-coded according to the presence or absence of the aroma characteristics obtained in <2-4> above (Figures 6-8). In the figures, the points representing each test substance are displayed darker for higher ranking descriptors representing aroma characteristics, and whiter for lower ranking descriptors representing aroma characteristics. The "order of appearance of descriptors representing aroma characteristics" refers to the listing order of the descriptors representing aroma characteristics in the Odor Description of Organoleptic Properties for each test substance from The Good Scents Company (http: / / www.thegoodscentscompany.com). In the figures, "Weight" indicates the square root of the reciprocal of the order of appearance of the descriptors representing aroma characteristics, except that if a descriptor representing an aroma characteristic does not appear, it is set to 0. Figure 6 shows the results of color-coding based on the presence or absence of the aroma characteristic "onion," Figure 7 shows the results of color-coding based on the presence or absence of the aroma characteristic "nutty," and Figure 8 shows the results of color-coding based on the presence or absence of the aroma characteristic "phenolic." 6 to 8 show that the black dots are localized in a narrow range in the stereochemical similarity space, i.e., the test substances exhibiting each odor characteristic are localized in a narrow range. This demonstrates that the odor characteristics of a substance can be predicted using stereochemical similarity that takes multiple conformations into account.

[0193] <5> Consideration This example demonstrates that even when aroma compounds appear to have dissimilar structural formulas or optimally stable conformations, they are likely to activate common olfactory receptors and exhibit a common odor if they share a conformation in terms of molecular surface shape and / or chemical properties. This is thought to be because, after dissolving in olfactory mucus, aroma compounds rotate around a single bond, adopting multiple conformations (multiple conformations), and activate the active sites of olfactory receptors in their appropriate conformations. Although many aroma compounds have been reported to activate multiple types of olfactory receptors, it is unlikely that a given aroma compound will necessarily adopt the same conformation when binding to the active site of one olfactory receptor and when binding to the active site of another. Information about the multiple conformations of aroma compounds is thought to be important for understanding the many-to-many combinatorial coding encoded by aroma compounds and olfactory receptors. In other words, the present invention is expected to enable highly accurate prediction of whether or not a substance has odor characteristics or olfactory receptor activation properties, something that was not possible with existing methods that ignore information about multiple conformations.

[0194] <6> Comparison of stereochemical structural similarity with existing methods and consideration of mixed methods The method of this example was compared with existing methods using data on the odor similarity between single substances from the results of the sensory evaluation conducted in Non-Patent Document 1. A combined method of both methods was also investigated.

[0195] <6-1> Odor similarity in sensory evaluation Non-Patent Document 1 contained 83 odor similarity data items between single substances. Excluding the results of similarity evaluations between identical substances, the number was 77. Of these, 9 items had a similarity score greater than 55 and 8 items had a similarity score less than 16, for a total of 17 items. The odor similarity evaluations in Non-Patent Document 1 were performed using the visual analog scale method, with similarity scores expressed on a scale of 0 (not at all similar) to 100 (very similar).

[0196] <6-2> Test substance The 25 compounds used in the 17 data narrowed down in <6-1> above were used.

[0197] <6-3> Calculation of similarity of stereochemical structure considering multiple conformations For the test substance of the above <6-2>, the stereochemical structural similarity was calculated taking into account multiple conformations in the same manner as in the above <3-1> and <3-2>.

[0198] <6-4> Calculation of molecular fingerprint (MACCS Keys) similarity MACCS Keys (155 bits) were generated using Canvas (Schroedinger) from the SDF data obtained in <2-3> above. The Tanimoto similarity of the MACCS Keys between the test substances was calculated and used as the molecular fingerprint similarity.

[0199] <6-5> Combination of stereochemical structure similarity and molecular fingerprint similarity The stereochemical similarity of <6-3> above was calculated in the range of 0 to 2, and the molecular fingerprint similarity of <6-4> above was calculated in the range of 0 to 1. Therefore, the stereochemical similarity was multiplied by 2 to align the range of similarities calculated by both methods. The similarities calculated by both methods were mixed in the ratios of 100:0, 90:10, 80:20, 70:30, 60:40, 50:50, 40:60, 30:70, 20:80, 10:90, or 0:100 to calculate a weighted average.

[0200] <6-6> Comparison of stereochemical structure similarity, molecular fingerprint similarity, and mixed methods Figure 9 shows a scatter plot of the odor similarity (referenced in <6-1> above) on the y-axis and the weighted average of the stereochemical similarity and molecular fingerprint similarity (calculated in <6-5> above) on the x-axis. Figure 10 shows the correlation coefficient between the odor similarity and the weighted average for each blend ratio. In the figure, "ROCS" indicates stereochemical similarity, and "MACCS" indicates molecular fingerprint similarity. A comparison of the correlation coefficients for each blend ratio showed that the method in which stereochemical similarity and molecular fingerprint similarity were blended at an 80:20 ratio (ROCS Ratio = 80) had the highest correlation with sensory perception (Figure 10). Furthermore, a comparison of the correlation coefficients for blend ratios of 100:0 and 0:100 showed that stereochemical similarity had a higher correlation with sensory perception than molecular fingerprint similarity (Figure 10).

[0201] Example B The present invention will be described in more detail below with reference to non-limiting examples according to the second aspect of the present invention. The substances used in the following examples are referred to as "test substances," but these substances can also be used as control substances in the prediction model production method and prediction method of the present invention.

[0202] <1> Generation of cells expressing human olfactory receptors <1-1> Construction of expression vectors for human olfactory receptors Using the same procedure as in Example A <1-1>, 352 expression vectors for human olfactory receptors were obtained.

[0203] <1-2> Preparation of olfactory receptor-expressing cells Using the same procedure as in Example A<1-2>, 352 HEK293T cell cultures expressing each of the 352 olfactory receptors were obtained.

[0204] <2> Creation of a database of human olfactory receptor activity <2-1> Luciferase assay A luciferase assay was carried out in the same manner as in Example A<2-1>, except that 1097 substances were selected as test substances from the substances listed in The Good Scents Company (http: / / www.thegoodscentscompany.com / ).

[0205] <2-2> Calculation of olfactory receptor activity The olfactory receptor activity was calculated using the same procedure as in Example A <2-2>.

[0206] <2-3> Aroma characteristics information The information on the odor properties of the test substances was taken from the descriptors listed in Odor Description of Organoleptic Properties by The Good Scents Company (http: / / www.thegoodscentscompany.com).

[0207] <3> Identifying olfactory receptor activity patterns characteristic of targeted odor properties A tree model was constructed using CART (L. Breiman, JH Friedman, R. A. Olshen and C. J. Stone, "Classification and Regression Trees", (Chapman and Hall, CRC, 1984)) with the presence or absence of descriptors representing the aroma characteristics obtained in <2-3> above flagged (present: 1, absent: 0) as the objective variable and the olfactory receptor activity calculated in <2-2> above as the explanatory variable.

[0208] A tree model is an algorithm that sequentially searches for the branching conditions that best divide the data in relation to the target variable from the explanatory variables. The analysis results return simple rules such as "If A, then B", and these rules can be illustrated in a tree structure, making the results easy to interpret. Gini impurity was used as the statistic used as the basis for division (an objective index showing how "cleanly" the data has been divided). For node t in the tree model, the number of samples in node t is N. tThe number of categories in node t is c, and the number of samples belonging to category i in node t is N i Then, the Gini impurity I(t) at node t is expressed as follows:

[0209]

number

[0210] At this time, the parent node D is determined based on the feature value f. p two child nodes D left and D right The information gain IG(Dp, f) obtained by dividing the p , N left , N right are node D p , D left , D right Let the number of samples contained in

[0211]

number

[0212] The feature f that maximizes this information gain IG(Dp, f) is designated as node D. p This process is repeated until a certain information gain cannot be obtained.

[0213] <3-1> Aroma characteristics: "burnt" By flagging compounds with the "burnt" or "roasted" descriptor in the Odor Description of The Good Scents Company's Organoleptic Properties, we constructed a tree model to identify olfactory receptor activity patterns characteristic of compounds with the "burnt" aroma characteristic (Figure 11). The tree model results are interpreted as follows (the same applies to subsequent experiments): the ellipses at the bottom are called leaves, and the other ellipses are called nodes. The numbers in square brackets above the ellipses identify the nodes and leaves. Of the two numbers above and below the ellipses, the lower number represents the proportion of compounds included in that node or leaf relative to the total number of compounds analyzed. The upper number represents the average value of the objective variable for the compounds included in that node or leaf. In this analysis, compounds with the "burnt" aroma characteristic (descriptor of "burnt" or "roasted") were assigned a value of 1, and all other compounds were assigned a value of 0. Therefore, the average value of the objective variable in Figure 11 represents the proportion of compounds with the "burnt" aroma characteristic. Branching conditions are shown below the ellipses for each node. Compounds that meet the condition are classified as the lower-left node or leaf, and compounds that do not meet the condition are classified as the lower-right node or leaf. Conditional branching is repeated for each compound until it reaches the leaf category. The main olfactory receptor activity pattern identified was "OR5K1 activity 4.10 or higher, OR6V1 activity 0.10 or higher, and OR1G1 activity less than 0.37" (identification number 7). In other words, it can be predicted that substances classified as leaf category with identification number 7 are likely to have the aroma characteristic "burnt."

[0214] <3-2> Aroma characteristics: "Sweet" By flagging compounds with the descriptor "sweet" in the Odor Description of Organoleptic Properties by The Good Scents Company and constructing a tree model, we identified olfactory receptor activity patterns characteristic of compounds with the scent characteristic "sweet" (Figure 12). The major olfactory receptor activity patterns identified were "OR8B3 activity ≥ 2.50 and OR5C1 activity ≥ -0.61" (identification number 15), "OR8B3 activity < 2.50, OR1D2 activity ≥ 1.40, and OR52A4 activity < -0.43" (identification number 12), "OR8B3 activity < 2.50, OR1D2 activity ≥ 1.40, OR52A4 activity ≥ -0.43, and OR1E1 activity < -0.13" (identification number 11), and "OR8B3 activity < 2.50, OR1D2 activity < 1.40, OR4S2 activity < 0.92, and OR2L8 activity ≥ 2.90" (identification number 7). In other words, substances classified as "leaf" with identification numbers 15, 12, 1, or 7 are predicted to have a high probability of possessing the odor characteristic "sweet."

[0215] <3-3> Aroma characteristics: "Nuts" By flagging compounds with the "nutty" descriptor in the Odor Description of The Good Scents Company's Organoleptic Properties, we constructed a tree model to identify olfactory receptor activity patterns characteristic of compounds with the "nutty" aroma characteristic (Figure 13). The main olfactory receptor activity patterns identified were "OR5K1 activity 3.80 or higher and OR1G1 activity less than 0.13" (identification number 7) and "OR5K1 activity 3.80 or higher, OR1G1 activity 0.13 or higher, and OR2AK2 activity 0.82 or higher" (identification number 6). In other words, substances classified as "Leaf" (identification number 7 or 6) are predicted to have a high probability of possessing the "nutty" aroma characteristic.

[0216] <4> Identification of olfactory receptor activity patterns characteristic of targeted molecular structures The presence or absence of the target molecular structure was flagged (present: 1, absent: 0) as the objective variable, and olfactory receptor activity was used as the explanatory variable. <3> A tree model was constructed using CART according to the procedure described in .

[0217] <4-1> Pyrazine skeleton By flagging compounds with a pyrazine skeleton and constructing a tree model, we identified olfactory receptor activity patterns characteristic of pyrazine-containing compounds (Figure 14). The main olfactory receptor activity patterns identified were "OR5K1 activity ≥ 3.90, OR13G1 activity ≥ -0.21, and OR5AR1 activity < 0.51" (identification number 13), "OR5K1 activity ≥ 3.90, OR13G1 activity ≥ -0.21, OR5AR1 activity ≥ 0.51, and OR2W1 activity < 1.00" (identification number 12), and "OR5K1 activity ≥ 2.30 but < 3.90, and OR8B3 activity < -1.2" (identification number 6). Therefore, substances classified under the leaf of identification number 13, 12, or 6 are predicted to have a high probability of having a pyrazine skeleton.

[0218] <4-2> Aldehyde group By flagging compounds with aldehyde groups and constructing a tree model, we identified olfactory receptor activity patterns characteristic of compounds with aldehyde groups (Figure 15). The main olfactory receptor activity patterns identified were "OR2J2 activity ≥ 2.10, OR2W1 activity < 0.83, and OR8B3 activity ≥ 0.40" (identification number 13), "OR2J2 activity ≥ 2.10, OR2W1 activity ≥ 0.83, OR6B1 activity < -1.60, and OR2Y1 activity ≥ -0.25" (identification number 10), and "OR2J2 activity ≥ 2.10, OR2W1 activity ≥ 0.83, OR6B1 activity ≥ -1.60, and OR1A1 activity < -0.15" (identification number 7). Therefore, substances classified under the leaf identification numbers 13, 10, or 7 are predicted to have a high probability of containing aldehyde groups.

[0219] <4-3> Ester bond By flagging compounds with ester bonds and constructing a tree model, we identified olfactory receptor activity patterns characteristic of compounds with ester bonds (Figure 16). The main olfactory receptor activity patterns identified were "OR2L8 activity ≥ 2.90, OR5K1 activity < 2.60, and OR4S2 activity < 0.80" (identification number 13) and "OR2L8 activity < 2.90, OR5P3 activity < 0.62, OR1D2 activity ≥ 0.74, and OR1G1 activity < 0.24" (identification number 8). In other words, substances classified under the leaf of identification number 13 or 8 are predicted to have a high probability of containing ester bonds.

[0220] Example C The present invention will be described in more detail below with reference to non-limiting examples according to the third aspect of the present invention. The substances used in the following examples are referred to as "test substances," but these substances can also be used as control substances in the prediction model production method and prediction method of the present invention.

[0221] <1> Generation of cells expressing human olfactory receptors <1-1> Construction of expression vectors for human olfactory receptors Using the same procedure as in Example A <1-1>, 352 expression vectors for human olfactory receptors were obtained.

[0222] <1-2> Preparation of olfactory receptor-expressing cells Using the same procedure as in Example A<1-2>, 352 HEK293T cell cultures expressing each of the 352 olfactory receptors were obtained.

[0223] <2> Creation of a database of human olfactory receptor activity <2-1> Luciferase assay A luciferase assay was performed in the same manner as in Example A<2-1>, except that all 144 substances listed in the Atlas of Odor Character Profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985) were selected as test substances.

[0224] <2-2> Calculation of olfactory receptor activity The olfactory receptor activity was calculated using the same procedure as in Example A <2-2>.

[0225] <2-3> Aroma characteristics information The odor property information of the test substance was quoted from the percentage of applicability (PA value) listed in Atlas of odor character profiles (Dravnieks, A., ASTM data series publication, DS 61, PCN 05-061000-36, 1985).

[0226] <3> Predicting the fit to aroma characteristics based on olfactory receptor activation data For each combination of test substance and olfactory receptor, the correlation coefficient between the PA value obtained in <2-3> above and the olfactory receptor activity calculated in <2-2> above was calculated. A linear regression model was constructed using machine learning, using the olfactory receptor activity for which the absolute value of the correlation coefficient exceeded a threshold (0.2) as the explanatory variable and the PA value as the response variable (Equations 1-3). In the equations, the olfactory receptor name (e.g., OR1F1) is substituted with the olfactory receptor activity for that olfactory receptor.

[0227] <3-1> Quantitative aroma characteristics of "STARWBERRY" There were 61 olfactory receptors for which the absolute value of the correlation coefficient between the PA value of "STARWBERRY" and olfactory receptor activity exceeded 0.2. Using the olfactory receptor activity of these 61 olfactory receptors, a linear regression model was constructed to predict the PA value of "STARWBERRY" (Equation 1). The correlation coefficient between the predicted PA value of the constructed regression model and the measured PA values ​​of 144 compounds was 0.932 (p < 0.001) (Figure 17).

[0228] Predicted PA value for "STARWBERRY" = 0.305 + 1.560 OR1F1 - 1.428 OR1I1 + 0.982 OR1J1 + 0.738 OR2B6 + 0.415 OR2B11 - 0.194 OR2C3 - 1.092 OR2G6 + 0.733 OR2K2 + 1.313 OR2L8 - 0.981 OR2T1 + 0.018 OR2T6 + 0.660 OR2W3 - 0.651 OR4A47 + 1.546 OR4B1 + 0.131 OR4C13 + 1.377 OR4D10 - 0.348 OR4F15 + 0.458 OR4K13 - 0.843 OR4P4 - 0.758 OR4Q3 + 1.342OR4X1 + 0.085OR5AK2 + 0.537OR5D14 - 2.108OR5H14 + 1.377OR5I1 + 0.265OR5J2 - 0.072OR5M3 - 0.332OR5M8 - 0.019OR6C2 - 0.695OR6C66P + 0.142OR6K2 - 1.751OR6T1 + 0.146OR8D4 + 3.164OR8K1 - 2.203OR8U1 - 0.562OR10A4 + 0.921OR10A7 - 1.501OR10D3 + 2.699OR10H2 - 1.454OR10J5 + 0.733OR10T2 - 2.356OR12D3 - 1.530OR13F1 - 1.186OR13G1 - 3.013OR13H1 + 0.075OR13J1 - 0.109OR14K1 + 0.587OR51B2 - 1.775OR51B4 - 0.116OR51M1 + 0.968OR51T1 + 1.253OR51V1 - 0.327OR52A4 - 1.189OR52B2 + 1.613OR52D1 - 1.592OR52H1 + 1.186OR52J3 + 0.560OR52N5 + 0.611OR52P2P - 2.133OR52R1 + 1.951OR56A5 ...(Formula 1)

[0229] <3-2> Quantitative aroma characteristics "ANISE (LICORICE)" There were 27 olfactory receptors for which the absolute value of the correlation coefficient between the PA value of "ANISE (LICORICE)" and olfactory receptor activity exceeded 0.2. Using the olfactory receptor activity of these 27 olfactory receptors, a linear regression model was constructed to predict the PA value of "ANISE (LICORICE)" (Equation 2). The correlation coefficient between the predicted PA value of the constructed regression model and the measured PA values ​​of 144 compounds was 0.823 (p < 0.001) (Figure 18).

[0230] Predicted PA value for "ANISE (LICORICE)" = 3.334 - 1.835 OR1J2 - 2.644 OR2A25 - 1.425 OR2G2 - 0.561 OR2L2 - 1.147 OR2T11 + 5.260 OR3A3 + 3.676 OR4C13 + 0.353 OR4D2 - 1.731 OR4P4 + 0.273 OR4X1 + 0.049 OR5AK2 - 0.645 OR6C6 + 1.418 OR6T1 - 0.144 OR7D4 - 2.990 OR8G5 + 0.613 OR9Q2 - 0.169 OR10A3 - 0.535 OR10J3 + 5.271 OR13C3 + 1.047OR13D1 - 2.075OR51A4 - 1.535OR51B6 + 0.880OR51G1 + 0.551OR51H1 - 0.467OR51M1 - 0.839OR52A4 - 2.291OR52N1 (Formula 2)

[0231] <3-3> Quantitative aroma characteristics of "NEW RUBBER" There were 56 olfactory receptors for which the absolute value of the correlation coefficient between the PA value of "NEW RUBBER" and olfactory receptor activity exceeded 0.2. Using the olfactory receptor activity of these 56 olfactory receptors, a linear regression model was constructed to predict the PA value of "NEW RUBBER" (Equation 3). The correlation coefficient between the predicted PA value of the constructed regression model and the measured PA values ​​of 144 compounds was 0.927 (p < 0.001) (Figure 19).

[0232] Predicted PA value for "NEW RUBBER" = 1.442 + 0.769 OR1G1 + 0.148 OR1J1 - 0.718 OR1L3 + 0.350 OR2A2 - 0.289 OR2AP1 + 0.184 OR2D2 + 0.041 OR2L8 + 0.211 OR2M2 - 0.060 OR2M4 - 0.077 OR4A16 + 0.470 OR4C6 - 0.712 OR4C12 - 1.136 OR4D9 + 0.786 OR4E2 + 0.019 OR4G11P + 0.248 OR4H12P - 0.262 OR4N2 + 0.512 OR4S1 + 1.233 OR5AU1 - 0.331 OR5B2 - 0.117OR5C1 - 0.869OR5L2 - 0.823OR5T2 + 0.126OR6B2 - 0.131OR6C70 + 0.463OR6K3 - 0.465OR6M1 + 0.004OR6Q1 + 0.108OR7A17 - 0.210OR8B3 - 0.471OR8G2 - 0.094OR8H3 - 0.978OR8K1 - 0.378OR9K2 + 0.658OR9Q2 - 1.111OR10A5 + 0.319OR10G3 + 0.183OR10G4 + 0.271OR10J3 - 0.038OR10K1 + 0.240OR10P1+ 0.584OR13D1 + 0.164OR14C36 - 0.772OR14I1 + 0.872OR51B2 + 0.179OR51H1 + 0.185OR51I2 - 0.936OR51L1 + 0.651OR51Q1 + 0.220OR52A5 - 0.001OR52B4 - 0.374OR52N2 + 0.344OR52W1 + 0.920OR56A3 - 0.786OR56A5 - 0.056OR56B1 (Formula 3) [Industrial Applicability]

[0233] According to one aspect of the present invention, it is possible to predict the presence or absence of an aroma characteristic or olfactory receptor activation characteristic in a substance. Also, according to one aspect of the present invention, it is possible to predict the presence or absence of components such as an aroma characteristic or a molecular structure in a substance. Also, according to one aspect of the present invention, it is possible to predict the compatibility of a substance with an aroma characteristic.

Claims

1. 1. A method for predicting the presence or absence of a property of interest for a test substance, comprising: predicting the presence or absence of the desired property for the test substance based on the maximum similarity of the stereochemical structure between the test substance and the reference substance; Including, The method, wherein the property is an odor property or an olfactory receptor activation property.

2. The method of claim 1 , wherein the control material comprises a positive control for the property of interest.

3. The method of claim 1 or 2, wherein the control substance is one substance.

4. The method of claim 1 or 2, wherein the control substance is a combination of two or more substances.

5. 5. The method of claim 1, wherein the prediction comprises clustering the test substance and the control substance based on the greatest similarity of stereochemical structure between the test substance and the control substance.

6. the control material comprises a positive control for the property of interest; 6. The method of claim 5, wherein the test substance is predicted to have the property of interest if the test substance is clustered into a cluster that includes the positive control.

7. The method according to any one of claims 1 to 6, further comprising the step of calculating the maximum similarity before the prediction.

8. 1. A method for screening for a substance having a desired property, comprising: predicting the presence or absence of said property of interest for a test substance by the method of any one of claims 1 to 7; and The test substance predicted to have the desired property is selected as a substance having the desired property. The process Including, The method, wherein the property is an odor property or an olfactory receptor activation property.

9. The method according to any one of claims 1 to 8, further comprising the step of confirming the presence or absence of the desired property for a test substance predicted to have the desired property.

10. The method according to any one of claims 1 to 9, wherein the maximum similarity is used in the prediction in combination with a structural similarity between the test substance and the reference substance other than the maximum similarity.

11. 1. A method for designing a material with desired properties, comprising: a step of designing a substance to be designed based on the maximum similarity of the stereochemical structure between the substance to be designed and a reference substance; Including, The method, wherein the property is an odor property or an olfactory receptor activation property.

12. the control material comprises a positive control for the property of interest; The design is performed such that the substance to be designed is clustered into a cluster that includes the positive control; The method of claim 11 , wherein the clustering comprises clustering the target substance and the control substance based on the maximum similarity of the stereochemical structures between the target substance and the control substance.

Citation Information

Patent Citations

  • Similarity calculation processing system, processing method and program of the same

    JP2008100918A

  • Roast-like perfume material screening method

    JP2019037197A

  • Method and Apparatus for Classification of Ligand by Using Molecular Vibrational Frequency Patterns

    KR101289948B1

  • Correlating olfactory perception with molecular structure

    US20180107803A1

Cited By

  • Method for evaluating and / or selecting agent for suppressing odor caused by aldehyde

    JP2023084478A