A method, system, device, and medium for predicting the aroma of a compound.

By collecting compound data in the FEMA Flavor system and constructing a causal graph, and then training it with a graph neural network, the problem of inaccurate compound aroma prediction was solved, achieving higher prediction accuracy and a deeper understanding of causal relationships.

CN120199361BActive Publication Date: 2025-11-14DADI HANKE BIOLOGICAL TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510290946.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-11-14
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The accuracy of compound aroma prediction in existing technologies is not high.

Method used

By collecting compound data from the FEMA Flavor system, preprocessing, feature extraction, and standardization were performed. A directed acyclic graph was constructed as a causal graph, and a graph neural network was used for training to predict the aroma characteristics of the compounds.

Benefits of technology

It improves the accuracy of compound aroma prediction, can more comprehensively capture various factors affecting aroma, and reveals the deep causal relationship between compounds and aroma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199361B_ABST
    Figure CN120199361B_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, device, and medium for predicting the aroma of compounds. The method includes: collecting target data of several compounds from several fragrance databases within the FEMA Flavor system; processing the target data of each compound using a preprocessing step, which includes cleaning, integration, feature extraction, standardization, and normalization; inputting the preprocessed target data of each compound into a preset algorithm to obtain a directed acyclic graph (DAG), and using the DAG as a causal graph, where the causal graph describes the causal relationship between chemicals and aroma; converting the molecular structures of the compounds in the causal graph into graph structures; and training a graph neural network based on the causal graph, where the graph neural network is used to predict the aroma characteristics of the compounds. This invention belongs to the field of compound aroma prediction. This invention can improve the accuracy of aroma prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of compound aroma prediction, and more particularly to a method, system, device, and medium for predicting compound aroma. Background Technology

[0002] The fragrance of a compound refers to the odor characteristic emitted by a specific chemical substance, which can be perceived through the human olfactory system. Different compounds produce a wide variety of fragrances or odors due to their unique molecular structures and physicochemical properties.

[0003] Currently, with advancements in analytical chemistry, sensory science, and machine learning, predicting the aroma of compounds is a popular topic. However, most current aroma prediction methods lack high accuracy. Therefore, a new method for predicting the aroma of compounds is urgently needed. Summary of the Invention

[0004] This invention provides a method, system, device, and medium for predicting the aroma of compounds, which solves the technical problem of inaccurate aroma prediction in the prior art and achieves the technical effect of accurately predicting the aroma of compounds.

[0005] In a first aspect, the present invention provides a method for predicting the aroma of a compound, the method comprising:

[0006] Target data for several compounds in the FEMAFlavor system were collected from several fragrance databases. These target data included molecular structure, physical properties, chemical properties, aroma description, and sensory scores.

[0007] The target data of each compound are processed by a preprocessing step, which includes cleaning, integration, feature extraction, standardization and normalization.

[0008] The preprocessed target data of each compound is input into the preset algorithm to obtain a directed acyclic graph, which is then used as a causal graph to describe the causal relationship between the chemical and the aroma.

[0009] Convert the molecular structure of the compound in the causal graph into a graph structure;

[0010] Based on causal graphs, graph neural networks are trained, which are used to predict the aroma characteristics of compounds.

[0011] Furthermore, the target data corresponding to each compound are processed through a preprocessing step, including:

[0012] Remove duplicate compounds and their corresponding target data;

[0013] Standardize the target data using a preset format;

[0014] Associate and match compounds with their corresponding aroma descriptions;

[0015] Feature extraction is performed on the target data to make it suitable for model training.

[0016] Furthermore, the preprocessed target data of each compound is input into the preset algorithm to obtain the directed acyclic graph of each compound, including:

[0017] Input the target data corresponding to each compound into the PC algorithm;

[0018] Initialize the preset parameters of the PC algorithm;

[0019] Based on the PC algorithm, the directed acyclic graphs corresponding to each compound are output.

[0020] Furthermore, after obtaining the causal diagrams of each compound, the following is also included:

[0021] Verify the cause-effect graph, including checking whether each causal edge in the cause-effect graph conforms to the preset logic;

[0022] If it does not conform, then the causal edge will be adjusted.

[0023] Furthermore, the molecular structures of compounds in the causal graph are converted into graph structures, including:

[0024] Convert the atoms in the molecular structure into nodes in the graph structure;

[0025] Convert chemical bonds in molecular structures into causal edges in graph structures.

[0026] Furthermore, based on several causal graphs, the graph neural network is trained, including:

[0027] Input the causal graph into the graph neural network;

[0028] When the graph neural network meets the preset training requirements, it will be used to predict the aroma characteristics of compounds.

[0029] Furthermore, sensory data includes:

[0030] Sensory data of various compounds are acquired using an electronic nose.

[0031] In a second aspect, the present invention provides a compound aroma prediction system, the system comprising:

[0032] The data acquisition module is used to collect target data of several compounds in the FEMA Flavor system from several flavor databases. The target data includes molecular structure, physical properties, chemical properties, flavor description and sensory score.

[0033] The data preprocessing module is used to process the target data of each compound through preprocessing steps, including cleaning, integration, feature extraction, standardization, and normalization.

[0034] The causal relationship module is used to input the preprocessed target data of each compound into the preset algorithm to obtain a directed acyclic graph, and use the directed acyclic graph as a causal graph. The causal graph is used to describe the causal relationship between the chemical and the aroma.

[0035] The conversion module is used to convert the molecular structure of compounds in a causal graph into a graph structure.

[0036] The training prediction module is used to train a graph neural network based on a causal graph, whereby the graph neural network is used to predict the aroma characteristics of compounds.

[0037] Thirdly, the present invention provides an electronic device, comprising:

[0038] processor;

[0039] Memory used to store processor-executable instructions;

[0040] The processor is configured to execute a compound aroma prediction method as provided in the first aspect.

[0041] Fourthly, the present invention provides a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform a compound aroma prediction method as provided in the first aspect.

[0042] One or more technical solutions provided in this invention have at least the following technical effects or advantages:

[0043] This invention, by collecting and integrating data from multiple sources (such as molecular structure, physical properties, and chemical properties), can more comprehensively capture various factors affecting aroma, thereby improving the accuracy of aroma prediction. Furthermore, by combining it with graph neural networks (GNNs) for training, and leveraging their ability to process graphical data, the model's ability to understand and predict complex relationships is further enhanced.

[0044] This invention converts the target data of each compound into a directed acyclic graph (DAG) and uses it as a causal graph to describe the causal relationship between the compound and the aroma. This helps to identify and understand how different chemical components work together to produce a specific aroma and reveals the deep causal relationships hidden in the data. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating a method for predicting the aroma of a compound provided by the present invention.

[0047] Figure 2 This is a schematic diagram of the structure of a compound aroma prediction system provided by the present invention. Detailed Implementation

[0048] This invention provides a method for predicting the aroma of compounds, thereby solving the technical problem of inaccurate aroma prediction in the prior art.

[0049] The technical solution of this invention is to solve the above-mentioned technical problems, and the overall idea is as follows:

[0050] A method for predicting the aroma of compounds includes: collecting target data of several compounds in the FEMA Flavor system from several fragrance databases, wherein the target data includes molecular structure, physical properties, chemical properties, aroma description, and sensory scores; processing the target data of each compound using a preprocessing step, wherein the preprocessing step includes cleaning, integration, feature extraction, standardization, and normalization; inputting the preprocessed target data of each compound into a preset algorithm to obtain a directed acyclic graph (DAG), and using the DAG as a causal graph, wherein the causal graph is used to describe the causal relationship between chemicals and aroma; converting the molecular structure of the compounds in the causal graph into a graph structure; and training a graph neural network based on the causal graph, wherein the graph neural network is used to predict the aroma characteristics of the compounds.

[0051] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0052] First, it should be clarified that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0053] This invention provides, for example Figure 1 The method for predicting the aroma of a compound, as shown, includes steps S11-S15:

[0054] Step S11: Collect target data of several compounds in the FEMA Flavor system from several fragrance databases. The target data includes molecular structure, physical properties, chemical properties, aroma description and sensory score.

[0055] The FEMA Flavor System, officially known as the Flavor and Extract Manufacturers Association's GRAS (Generally Recognized As Safe) program, is an internationally recognized and important standard system for assessing the safety of food flavorings.

[0056] Molecular structure refers to the arrangement of atoms in a compound and their interconnections. It is usually represented by a chemical formula, such as CHO for glucose. More detailed structures can be described using structural diagrams or SMILES (Simplified Molecular Input Line Entry System) strings.

[0057] Physical properties include melting point, boiling point, density, refractive index, etc. These characteristics are inherent properties of compounds and do not involve chemical changes.

[0058] Chemical properties refer to a compound’s ability to react with other substances, such as redox properties and acidity / alkalinity.

[0059] Fragrance descriptions are textual descriptions of the odor characteristics emitted by compounds, such as "fresh citrus" or "sweet woody." Fragrance descriptions can be collected in the FEMA Flavor system.

[0060] Sensory data for each compound can be acquired using an electronic nose. Sensory ratings can include olfactory intensity, persistence, irritation, or odor.

[0061] Sensory scores are collected through standardized testing procedures, relying on devices such as electronic noses and / or odor mass spectrometers to quantify fragrance characteristics. An electronic nose is a device that simulates the human olfactory system, using an array of sensors to detect and analyze odor molecules in the air. It can quantify fragrance characteristics and quickly detect and quantify odor concentration, intensity, and basic odor categories, assessing olfactory intensity and fragrance persistence. An odor mass spectrometer, on the other hand, is a device capable of detailed analysis of odor molecule components. It is used to accurately measure the complex composition of fragrances, precisely identifying fragrance components, types, and persistence, and assessing irritation or unpleasant odors.

[0062] Step S12 involves processing the target data of each compound using a preprocessing step, which includes cleaning, integration, feature extraction, standardization, and normalization.

[0063] The preprocessing steps involve processing the target data corresponding to each compound, including: removing duplicate compounds and their corresponding target data; standardizing the target data in a preset format; associating and matching the compounds with their corresponding aroma descriptions; and extracting features from the target data to make it suitable for model training.

[0064] Data fields from different sources (such as compound names and molecular formulas) can be standardized into a standard format. First, standards (such as IUPAC naming rules) need to be defined and followed. Then, data cleaning is performed to ensure consistent text formatting, remove redundant information, and correct errors. Furthermore, software tools and natural language processing techniques are used to transform and extract structured information (such as ChemDraw, ChemAxon JChem, PubChem, ChEBI, etc.), and mapping tables are created to standardize synonyms or near-synonyms into standardized terminology.

[0065] It can extract features from the molecular structure, physical properties, and chemical properties of the target data, and perform word segmentation and stop word removal on the fragrance description.

[0066] After data cleaning and integration, standardization or normalization can be performed to eliminate dimensional differences and improve model training efficiency. Standardization refers to converting data into a standard normal distribution with a mean of 0 and a standard deviation of 1 to speed up model training. Normalization refers to scaling data to a fixed range (usually [0,1] or [-1,1]) to prevent dimensional differences from affecting model training.

[0067] Step S13: Input the preprocessed target data of each compound into the preset algorithm to obtain a directed acyclic graph, and use the directed acyclic graph as a causal graph, wherein the causal graph is used to describe the causal relationship between the chemical and the aroma.

[0068] The purpose of step S13 is to process the relationship between compounds and aroma characteristics using a preset algorithm.

[0069] The preprocessed target data of each compound is input into the preset algorithm to obtain the directed acyclic graph of each compound, including:

[0070] The preset algorithm can be the PC algorithm, which is an algorithm for learning Bayesian network structures. The PC algorithm aims to discover conditional independence relationships between variables in data and construct a directed acyclic graph (DAG) based on these relationships to represent causal relationships or statistical dependencies between variables.

[0071] The target data corresponding to each compound can be input into the PC algorithm; the target data have all been standardized or normalized to ensure that the scale of all variables is consistent.

[0072] Initialize the preset parameters of the PC algorithm; based on the PC algorithm, output the directed acyclic graph corresponding to each compound.

[0073] Specifically:

[0074] In PC algorithms, choosing an appropriate conditional independence test method is crucial. Commonly used test methods include:

[0075] Chi-square test: Applicable to categorical variables, it determines whether two variables are independent by comparing observed frequencies and expected frequencies.

[0076] Mutual information: a measure of the degree of interdependence between two random variables, applicable to discrete or continuous variables, and capable of capturing nonlinear relationships.

[0077] The test method is used to determine whether two variables are conditionally independent under given conditions. If two variables remain correlated after controlling for other variables, there may be a direct causal relationship between them.

[0078] The significance level (usually denoted as α) is the criterion used in statistical hypothesis testing to determine whether to reject the null hypothesis. For conditional independence tests, the significance level determines the probability of committing a Type I error.

[0079] Use the selected conditional independence test method to analyze the data to identify conditional independence relationships between variables.

[0080] Based on the results of the conditional independence test, the PC algorithm constructs a directed acyclic graph (DAG), which is also known as a causal graph. In the causal graph, nodes represent variables, and edges represent direct causal relationships between variables.

[0081] It should be noted that the relationship between compounds and directed acyclic graphs is many-to-many, meaning that one compound may exist in multiple directed acyclic graphs, and multiple compounds may exist in a directed acyclic graph.

[0082] After obtaining the causal diagrams for each compound, the following is also included:

[0083] The causal graph is validated to ensure its rationality, including checking whether each causal edge in the causal graph conforms to the preset logic. The preset logic can be set according to the actual situation of each molecule. If it does not conform, the causal edge can be adjusted based on relevant domain knowledge or knowledge graphs of relevant domains.

[0084] Step S14: Convert the molecular structure of the compound in the causal diagram into a graph structure.

[0085] Specifically, this includes: converting atoms in a molecular structure into nodes in a graph structure; and converting chemical bonds in a molecular structure into causal edges in a graph structure.

[0086] Step S15: Based on the causal graph, train the graph neural network, which is used to predict the aroma characteristics of the compound.

[0087] The graph neural network is trained based on several causal graphs, including: inputting causal graphs into the graph neural network; and when the graph neural network meets the preset training requirements, using the graph neural network to predict the aroma characteristics of compounds.

[0088] Preset training requirements can include things like the maximum number of training sessions, but there are no restrictions here.

[0089] In summary, this invention provides a method for predicting the aroma of compounds. The method includes: collecting target data of several compounds from the FEMA Flavor system in several fragrance databases, wherein the target data includes molecular structure, physical properties, chemical properties, aroma description, and sensory scores; processing the target data of each compound using preprocessing steps, including cleaning, integration, feature extraction, standardization, and normalization; inputting the preprocessed target data of each compound into a preset algorithm to obtain a directed acyclic graph (DAG), and using the DAG as a causal graph, wherein the causal graph is used to describe the causal relationship between chemicals and aroma; converting the molecular structure of the compounds in the causal graph into a graph structure; and training a graph neural network based on the causal graph, wherein the graph neural network is used to predict the aroma characteristics of the compounds. This invention, by collecting and integrating data from multiple sources (such as molecular structure, physical properties, and chemical properties), can more comprehensively capture various factors affecting aroma, thereby improving the accuracy of aroma prediction. Combining the training with a graph neural network (GNN) utilizes its ability to process graphical data to further enhance the model's understanding and prediction capabilities of complex relationships. This invention converts the target data of each compound into a directed acyclic graph (DAG) and uses it as a causal graph to describe the causal relationship between the compound and the aroma. This helps to identify and understand how different chemical components work together to produce a specific aroma and reveals the deep causal relationships hidden in the data.

[0090] Based on the same inventive concept, the present invention provides, as follows: Figure 2 The system shown is a compound aroma prediction system, the system comprising:

[0091] The data acquisition module 21 is used to collect target data of several compounds in the FEMA Flavor system from several fragrance databases. The target data includes molecular structure, physical properties, chemical properties, aroma description and sensory score.

[0092] The data preprocessing module 22 is used to process the target data of each compound through preprocessing steps, including cleaning, integration, feature extraction, standardization and normalization.

[0093] The causal relationship module 23 is used to input the preprocessed target data of each compound into the preset algorithm to obtain a directed acyclic graph, and use the directed acyclic graph as a causal graph, wherein the causal graph is used to describe the causal relationship between the chemical and the aroma.

[0094] The conversion module 24 is used to convert the molecular structure of the compound in the causal diagram into a graph structure;

[0095] The training prediction module 25 is used to train a graph neural network based on a causal graph, wherein the graph neural network is used to predict the aroma characteristics of compounds.

[0096] Based on the same inventive concept, the present invention also provides an electronic device, comprising:

[0097] processor;

[0098] Memory used to store processor-executable instructions;

[0099] The processor is configured to execute a compound aroma prediction method as described above.

[0100] Based on the same inventive concept, the present invention also provides a non-transitory computer-readable storage medium, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform a compound aroma prediction method as described above.

[0101] Since the electronic device described in this embodiment is an electronic device used to implement the information processing method in the embodiments of the present invention, those skilled in the art can understand the specific implementation methods and various variations of the electronic device in this embodiment based on the information processing method described in the embodiments of the present invention. Therefore, how the electronic device implements the method in the embodiments of the present invention will not be described in detail here. Any electronic device used by those skilled in the art to implement the information processing method in the embodiments of the present invention falls within the scope of protection of the present invention.

[0102] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.

[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0106] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0107] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for predicting the aroma of a compound, characterized in that, The method includes: Target data for several compounds in the FEMA Flavor system were collected from several flavor databases. These target data included molecular structure, physical properties, chemical properties, flavor descriptions, and sensory scores. The target data of each compound are processed by a preprocessing step, which includes cleaning, integration, feature extraction, standardization and normalization. The preprocessed target data of each compound is input into a preset algorithm to obtain a directed acyclic graph (DAG), which is then used as a causal graph to describe the causal relationship between chemicals and aroma. The process includes: inputting the target data of each compound into the PC algorithm; initializing the preset parameters of the PC algorithm; and outputting the DAG corresponding to each compound based on the PC algorithm. Converting the molecular structure of a compound in a causal graph into a graph structure includes: converting atoms in the molecular structure into nodes in the graph structure; and converting chemical bonds in the molecular structure into causal edges in the graph structure. A graph neural network is trained based on a causal graph, wherein the graph neural network is used to predict the aroma characteristics of compounds.

2. The method for predicting the aroma of a compound as described in claim 1, characterized in that, The preprocessing steps involve processing the target data corresponding to each compound, including: Remove duplicate compounds and their corresponding target data; Standardize the target data using a preset format; Associate and match compounds with their corresponding aroma descriptions; Feature extraction is performed on the target data to make it suitable for model training.

3. The method for predicting the aroma of a compound as described in claim 1, characterized in that, After obtaining the causal diagrams for each compound, the following is also included: Verify the cause-effect graph, including checking whether each causal edge in the cause-effect graph conforms to the preset logic; If it does not conform, then the causal edge will be adjusted.

4. The method for predicting the aroma of a compound as described in claim 1, characterized in that, Based on several causal graphs, a graph neural network is trained, including: Input the causal graph into the graph neural network; When the graph neural network meets the preset training requirements, it is used to predict the aroma characteristics of compounds.

5. The method for predicting the aroma of a compound as described in claim 1, characterized in that, Sensory data, including: Sensory data of various compounds are acquired using an electronic nose.

6. A compound aroma prediction system, characterized in that, A method for predicting the aroma of a compound according to any one of claims 1-5, the system comprising: The data acquisition module is used to collect target data of several compounds in the FEMA Flavor system from several flavor databases. The target data includes molecular structure, physical properties, chemical properties, flavor description and sensory score. The data preprocessing module is used to process the target data of each compound through preprocessing steps, including cleaning, integration, feature extraction, standardization, and normalization. The causal relationship module is used to input the preprocessed target data of each compound into the preset algorithm to obtain a directed acyclic graph, and use the directed acyclic graph as a causal graph. The causal graph is used to describe the causal relationship between the chemical and the aroma. The conversion module is used to convert the molecular structure of compounds in a causal graph into a graph structure. A training prediction module is used to train a graph neural network based on a causal graph, wherein the graph neural network is used to predict the aroma characteristics of compounds.

7. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute a compound aroma prediction method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform a compound aroma prediction method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Identification method and device for miRNA causal regulation and control network, electronic equipment and storage medium

    CN111370062A

  • Machine learning for predicting attributes of chemical agents

    CN117223061A

  • Aromatic nitro compound toxicity prediction method based on graph attention network

    CN117238401A

  • Molecular force field modeling method and device, electronic equipment and storage medium

    CN119229948A