Molecular odor identification method, device, equipment, medium and product

By obtaining the material structure of gas molecules, determining the SMILES formula and generating target virtual encoding, and combining olfactory receptor information for odor label recognition, the problem that existing models are difficult to distinguish stereoisomeric information of odor molecules is solved, and the accuracy of odor space map generation and odor prediction is improved.

CN120254094APending Publication Date: 2025-07-04NANJING UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510248015.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing graph neural network-based models are difficult to distinguish stereoisomeric information of odor molecules, and AI odor prediction methods lack physiological mechanism interpretability, making it difficult to explain the differences in odor perception in isomers.

Method used

By obtaining the material structure of gas molecules, determining the SMILES formula and normalizing it, generating target virtual encoding, inputting it into the OR-POM model with olfactory receptor information, odor label recognition and odor space map generation are performed.

Benefits of technology

The analysis and encoding and identification of stereoisomeric information of odor molecules is realized, and an odor space map is generated, which improves the accuracy of odor prediction and physiological mechanism interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure HDA0005296134430000011
    Figure HDA0005296134430000011
Patent Text Reader

Abstract

The invention provides a molecular smell recognition method, device and equipment, a medium and a product, and the method comprises the steps: obtaining a plurality of gas molecules in a target region; performing substance identification on the plurality of gas molecules, and determining the substance structure of each gas molecule; determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule; carrying out standardization treatment on the SMILES formula of each gas molecule to obtain a target SMILES formula of each gas molecule; generating a target virtual code according to the plurality of target SMILES formulas and preset olfactory receptor information; inputting the target virtual code and a plurality of target SMILES into a trained OR-POM model, and determining a smell label of each gas molecule; and determining the odor in the target area and an odor space map in the target space according to the odor label of each gas molecule. According to the SMILES type of the gas molecules and olfactory receptor information, stereoisomerous information analysis of the odor molecules can be carried out, and the odor of the gas molecules can be coded and identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of gas molecule recognition. Specifically, it relates to a method, device, equipment, medium and product for identifying molecular odors. Background Technique

[0002] The generation of smell is a complex process involving multiple fields such as physics, chemistry, and neurobiology. From a physical perspective, characteristics such as the size, shape, and volatility of odor molecules affect their propagation in the air and the ability to be captured by the olfactory organs. From a chemical perspective, the interaction between odor molecules and receptor proteins on olfactory cells is based on intermolecular chemical bonding and affinity. Different odor molecules bind to different receptor proteins, thereby generating different olfactory sensations. From a neurobiological perspective, the processing and interpretation of olfactory signals by the brain are based on the connections and signal transduction between neurons. The brain analyzes and identifies olfactory signals based on past experiences and memories, thereby generating specific olfactory sensations and emotional responses. A fundamental problem in neuroscience is to map the physical characteristics of stimuli to perceptual features. In vision, wavelength maps to color; in audition, frequency maps to pitch. In contrast, little is known about the mapping from chemical structure to olfactory perception, and mapping molecular structure to odor perception is a key challenge in olfaction.

[0003] However, existing models based on graph neural networks are limited by the way of molecular characterization in the models and are difficult to distinguish the stereoisomeric information of odor molecules. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide a method, device, equipment, medium and product for identifying molecular odors, which can at least analyze the stereoisomeric information of odor molecules according to the SMILES formula and olfactory receptor information of gas molecules, encode and identify the odor of gas molecules, and can predict the area of the odor to generate an odor space map.

[0005] An embodiment of the present application provides a method for identifying molecular odors, including: obtaining a plurality of gas molecules in a target area; performing substance identification on the plurality of gas molecules to determine the substance structure of each gas molecule; determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule; performing normalization processing on the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule; generating a target virtual code according to a plurality of target SMILES formulas and pre-set olfactory receptor information, the target virtual code including a plurality of coding bits, and each coding bit representing the interaction probability between a gas molecule and an olfactory receptor; inputting the target virtual code and a plurality of target SMILES formulas into a trained OR-POM model to determine the odor label of each gas molecule; and determining the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0006] Optionally, the step of determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule includes: determining the SMILES formula of each gas molecule according to a pre-established list of sample identification results and the substance structure of each gas molecule; wherein, the list of sample identification results is established through the following steps: obtaining a variety of gas samples; analyzing and detecting each sample to obtain a sample data file, the sample data file including the substance structure of each gas sample; converting the format through Agilent unknown analysis software and performing non-targeted analysis and identification on the NIST2017 data set to obtain the list of sample identification results.

[0007] Optionally, the step of performing normalization processing on the SMILES formula of each gas molecule includes: based on the RDkit package, converting the SMILES formula of each gas molecule into a molecular object; performing standardization processing on each molecular object to obtain the target SMILES formula of each gas molecule, wherein the standardization processing includes: cleaning, removing metal bonds, neutralizing, removing isotopes, and retaining stereoisomer information.

[0008] Optionally, the pre-set olfactory receptor information includes a plurality of human olfactory receptors; wherein, generating a target virtual code according to a plurality of target SMILES formulas and pre-set olfactory receptor information includes: pairing each target SMILES formula and each human olfactory receptor according to a plurality of target SMILES formulas and pre-set olfactory receptor information to generate a plurality of scoring values of the interaction probability between a plurality of target SMILES formulas and a plurality of human olfactory receptors; and constructing a multi-dimensional target virtual code according to the plurality of scoring values of the interaction probability between a plurality of target SMILES formulas and a plurality of human olfactory receptors.

[0009] Optionally, the OR-POM model performs the following steps: extracting features from multiple molecular structures in multiple target SMILES formulas using a graph neural network to obtain structural feature vectors of multiple molecules, and fusing the interaction information of human olfactory receptors after splicing the structural feature vectors of multiple molecules with virtual encoding; wherein, the OR-POM model includes a graph neural network, a Layer Normalization layer, and a Feed-Forward Neural Network (FFN), the graph neural network is used to extract the structural features of molecules, the Layer Normalization (LN) layer is used to prevent model overfitting after fusing the interaction information of human olfactory receptors, and the Feed-Forward Neural Network (FFN) is used to adjust the vector dimension for predicting odor labels and modeling the embedding space.

[0010] Optionally, the steps of determining the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule include: obtaining the target embedding vector of the penultimate layer in the Feed-Forward Neural Network (FFN); based on the target embedding vector, using the kernel density estimation algorithm to perform odor label clustering statistics to obtain the odor label clustering statistics result; using the principal component analysis algorithm, according to the odor label clustering statistics result, performing data dimensionality reduction, and screening data points according to the odor label of each gas molecule, using Gaussian kernel density estimation to calculate the probability density of the odor label on the target area grid, and determining the density value of each odor label corresponding to the grid; determining the odor in the target area and the odor space map in the target space according to the density value of each odor label corresponding to the grid.

[0011] In a second aspect, the present application also provides a molecular odor recognition device, including: a gas molecule acquisition module for acquiring multiple gas molecules in a target area;

[0012] a substance structure determination module for identifying the substances of the multiple gas molecules to determine the substance structure of each gas molecule;

[0013] a SMILES formula determination module for determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule;

[0014] a normalization processing module for normalizing the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule;

[0015] A target virtual coding generation module, configured to generate target virtual coding according to a plurality of target SMILES formulas and preset olfactory receptor information, where the target virtual coding includes a plurality of coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0016] An odor label determination module, configured to input the target virtual coding and a plurality of target SMILES formulas into a trained OR-POM model to determine the odor label of each gas molecule;

[0017] An odor space map determination module, configured to determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0018] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0019] Obtain a plurality of gas molecules in the target area;

[0020] Perform material identification on the plurality of gas molecules to determine the material structure of each gas molecule;

[0021] Determine the SMILES formula of each gas molecule according to the material structure of each gas molecule;

[0022] Perform normalization processing on the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule;

[0023] Generate target virtual coding according to a plurality of target SMILES formulas and preset olfactory receptor information, where the target virtual coding includes a plurality of coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0024] Input the target virtual coding and a plurality of target SMILES formulas into a trained OR-POM model to determine the odor label of each gas molecule;

[0025] Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0026] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0027] Obtain a plurality of gas molecules in the target area;

[0028] Perform material identification on the plurality of gas molecules to determine the material structure of each gas molecule;

[0029] Determine the SMILES formula for each gas molecule according to the molecular structure of each gas molecule;

[0030] Normalize the SMILES formula of each gas molecule to obtain the target SMILES formula for each gas molecule;

[0031] Generate a target virtual code according to multiple target SMILES formulas and pre-set olfactory receptor information, where the target virtual code includes multiple coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0032] Input the target virtual code and multiple target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule;

[0033] Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0034] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0035] Obtain multiple gas molecules in the target area;

[0036] Conduct material identification on the multiple gas molecules to determine the molecular structure of each gas molecule;

[0037] Determine the SMILES formula for each gas molecule according to the molecular structure of each gas molecule;

[0038] Normalize the SMILES formula of each gas molecule to obtain the target SMILES formula for each gas molecule;

[0039] Generate a target virtual code according to multiple target SMILES formulas and pre-set olfactory receptor information, where the target virtual code includes multiple coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0040] Input the target virtual code and multiple target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule;

[0041] Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0042] The molecular odor recognition method, device, equipment, medium and product provided by the embodiments of the present application can encode and recognize the odor of gas molecules according to the SMILES formula and olfactory receptor information of gas molecules, and can predict the odor area to generate an odor space map.

[0043] To make the above objects, features and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0045] Figure 1 A flowchart of a molecular odor recognition method provided by an embodiment of the present application;

[0046] Figure 2 A flowchart of an odor substance recognition model provided by an embodiment of the present application;

[0047] Figure 3 An odor space map with retained label levels provided by an embodiment of the present application;

[0048] Figure 4 A performance comparison chart of the POM and OR-POM models on the test set provided by an embodiment of the present application;

[0049] Figure 5 A performance comparison chart of the application of POM and OR-POM on the stereoisomer dataset provided by an embodiment of the present application;

[0050] Figure 6 A comparison of the main odor maps constructed by POM and OR-POM provided by an embodiment of the present application;

[0051] Figure 7 A structural block diagram of a molecular odor recognition provided by an embodiment of the present application;

[0052] Figure 8 An internal structure diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are only some of the embodiments of this application, rather than all of them. The components of the embodiments of this application described and illustrated herein can generally be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without creative efforts belongs to the scope of protection of this application.

[0054] First, the applicable application scenarios of this application will be introduced. This application can be applied to the field of gas molecule recognition technology.

[0055] Through research, it has been found that the generation of smell is a complex process involving multiple fields such as physics, chemistry, and neurobiology. From a physical perspective, the size, shape, and volatility of odor molecules affect their propagation in the air and the ability to be captured by the olfactory organs. From a chemical perspective, the interaction between odor molecules and receptor proteins on olfactory cells is based on intermolecular chemical bonding and affinity. Different odor molecules bind to different receptor proteins, resulting in different olfactory sensations. From a neurobiological perspective, the processing and interpretation of olfactory signals by the brain are based on the connections and signal transduction between neurons. The brain analyzes and recognizes olfactory signals based on past experiences and memories, generating specific olfactory sensations and emotional responses. A fundamental problem in neuroscience is to map the physical characteristics of stimuli to perceptual features. In vision, wavelength maps to color; in audition, frequency maps to pitch. In contrast, little is known about the mapping from chemical structure to olfactory perception, and mapping molecular structure to odor perception is a key challenge in olfaction.

[0056] However, existing models based on graph neural networks are limited by the way of molecular characterization in their models and are difficult to distinguish the stereoisomer information of odor molecules.

[0057] Moreover, with the development of fields such as data analysis and artificial intelligence, researchers have proposed some AI algorithm models that can link structure and odor. However, the prediction of existing AI models is a data-driven black box model. Existing AI odor prediction methods only start from the molecular structure, ignoring the interaction information with olfactory receptors, often lacking research on the interpretability of physiological mechanisms, and also difficult to explain the differences in odor perception among isomers.

[0058] Based on this, the embodiments of the present application provide a method, device, equipment, medium and product for identifying molecular odors, which can analyze the stereoisomer information of odor molecules according to the SMILES formula and olfactory receptor information of gas molecules, encode and identify the odors of gas molecules, and can predict the regions of the odors to generate an odor space map.

[0059] And based on the results of substance identification, effective identification of molecular odor tags is carried out through the self-built OR-POM model.

[0060] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of a method for identifying molecular odors provided by the embodiments of the present application. As Figure 1 shown in

[0061] S101. Obtain a plurality of gas molecules in the target area.

[0062] S102. Conduct substance identification on the plurality of gas molecules to determine the substance structure of each gas molecule.

[0063] As an example, the steps for identifying the substance structure of each gas molecule include: sample pretreatment, and the storage and preparation of the sample are based on the prior art. Multiple solvents can be used for substance extraction, such as concentrated extracts of dichloromethane, n-hexane, etc., and finally volume is fixed for analysis on the machine.

[0064] When using instrumental analysis and substance structure identification, instrumental analysis can use a GC / -QTOF gas chromatography-mass spectrometry system (Agilent, USA) to analyze and detect the pretreated sample. The instrument setting parameters are: chromatograph: Agilent 8890 chromatograph; chromatographic column and flow rate: two 15m HP-5ms ultra-inert chromatographic columns are connected in series. The flow rate of the first chromatographic column is 1 mL / min, and the flow rate of the second chromatographic column is 1.2 mL / min.

[0065] In the use of gas, the quenching gas and auxiliary gas can be helium, and the collision gas can be nitrogen.

[0066] In the control of temperature and pressure, the inlet heater temperature is 250 °C, the pressure is 9.3508 psi, and the septum purge flow rate is 3 mL / min. Pulse splitless injection is adopted, the injection pulse pressure is 25 psi, and the time is 0.5 min. The initial temperature of the column oven is 60 °C, the equilibration time is 3 min, the method lasts for 46.5 min, and the gas chromatography programmed temperature program is shown in Table 1 below.

[0067] Table 1 Parameter setting results:

[0068]

[0069]

[0070] Specifically, the high-resolution mass spectrometry system of the mass spectrometer is an Agilent 7250 QTOF mass spectrometer. An electron impact ionization source (EI) is used, the ion source temperature is 200 °C, the electron energy is 70 eV, and the emission current is 5 μA. The acquisition mode is the MS mode, and ion data with a mass-to-charge ratio between 50 and 650 amu is acquired at an acquisition rate of 5 Hz.

[0071] Specifically, when performing substance structure identification, after the sample is analyzed and detected, a sample data file is obtained, and the format is converted through Agilent unknown substance analysis software for non-targeted analysis and identification on the NIST2017 dataset to obtain a sample identification result list.

[0072] S103. Determine the SMILES formula of each gas molecule according to the substance structure of each gas molecule.

[0073] Specifically, the steps of determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule include: determining the SMILES formula of each gas molecule according to the pre-established sample identification result list and the substance structure of each gas molecule.

[0074] Among them, the sample identification result list is established through the following steps: obtain a variety of gas samples; analyze and detect each sample to obtain a sample data file, and the sample data file includes the substance structure of each gas sample; convert the format through Agilent unknown substance analysis software and perform non-targeted analysis and identification on the NIST2017 dataset to obtain a sample identification result list.

[0075] S104. Normalize the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule.

[0076] Specifically, the steps of normalizing the SMILES formula of each gas molecule include: based on the RDkit package, convert the SMILES formula of each gas molecule into a molecular object; perform standardization processing on each molecular object to obtain the target SMILES formula of each gas molecule, where the standardization processing includes: removing metal bonds, neutralizing, removing isotopes, and retaining stereoisomer information.

[0077] S105. Generate a target virtual code according to multiple target SMILES formulas and pre-set olfactory receptor information.

[0078] Among them, the target virtual code includes multiple coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor.

[0079] Specifically, the preset olfactory receptor information includes multiple human olfactory receptors.

[0080] Among them, generating the target virtual code according to multiple target SMILES formulas and the preset olfactory receptor information includes: pairing each target SMILES formula with each human olfactory receptor according to the multiple target SMILES formulas and the preset olfactory receptor information to generate multiple scoring values of the interaction probability between multiple target SMILES formulas and multiple human olfactory receptors; constructing a multi-dimensional target virtual code according to the multiple scoring values of the interaction probability between multiple target SMILES formulas and multiple human olfactory receptors.

[0081] As an example, the Python language can be used to standardize the SMILES formula for substance identification based on the existing RDkit package. Specifically, the SMILES string can be first converted into a molecular object, and then a series of standardization steps can be carried out, including removing metal bonds, neutralization, removing isotopes, and retaining stereochemical information. Finally, the canonical target SMILES formula can be obtained. Among them, the canonical target SMILES formula can be saved as the standardized SMILES.xlsx file and stored on the hard disk.

[0082] Specifically, the Python language can be used to utilize a convolutional neural network model with an attention mechanism trained by the existing interaction data between human olfactory receptors and small molecules. The amino acid sequence information of all human olfactory receptor families and the trained weight information are built into the model script, which mainly includes three functions: data reading, interaction prediction, and data saving. The data reading function will parse the standardized SMILES.xlsx file previously stored on the hard disk to complete the data reading. Subsequently, in the interaction prediction, the model pairs each input SMILES formula with the built-in olfactory receptor information to generate a scoring value of the interaction probability for each molecular structure and each human olfactory receptor, completing the encoding of the human olfactory receptor interaction information, that is, the generation of the virtual code, and constructing a 382-dimensional virtual code. Each coding bit represents the interaction probability between the molecule and an olfactory receptor. Finally, the storage function of this module will save the virtual code corresponding to the molecular SMILES formula as the virtual code.xlsx file and store it on the hard disk.

[0083] S106. Input the target virtual code and multiple target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule.

[0084] Specifically, the OR-POM model performs the following steps: extracting features in multiple molecular structures in multiple target SMILES formulas by using a graph neural network, obtaining structural feature vectors of multiple molecules, and completing the fusion of human olfactory receptor interaction information after splicing the structural feature vectors of multiple molecules with virtual encoding.

[0085] Among them, the OR-POM model includes a graph neural network, a Layer Normalization layer, and a Feed-Forward Neural Network. The graph neural network is used to extract the structural features of molecules. The Layer Normalization (LN) layer is used to prevent model overfitting after fusing human olfactory receptor interaction information. The Feed-Forward Neural Network (FFN) is used to adjust the vector dimension, predict odor labels, and model the embedding space.

[0086] It should be noted that the determination of odor labels is based on a pre-established OR-POM model. This model uses the SMILES.xlsx file stored in the substance structure standardization module and the virtual encoding.xlsx file stored in the virtual encoding generation module as inputs, extracts features in the molecular structure by using a graph neural network to obtain the structural feature vector of the molecule, completes the fusion of human olfactory receptor interaction information after splicing the structural feature vector with virtual encoding, and uses techniques such as Layer Normalization (LN) layer standardization to prevent model overfitting. Subsequently, a Feed-Forward Neural Network (FFN) is used for fitting, adjusting the vector dimension, predicting odor labels, and modeling the embedding space.

[0087] Specifically, it can be based on a neural network model on the pytorch platform. The model training has been completed, and it contains the trained model weights and default model parameters, and the odor label recognition of the input molecule can be completed without other steps.

[0088] S107. Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0089] Specifically, the steps of determining the odor in the target area and the odor space map in the target space according to the odor labels of each gas molecule include: obtaining the target embedding vector of the penultimate layer in the Feed-Forward Neural Network; based on the target embedding vector, using the kernel density estimation algorithm to perform odor label clustering statistics to obtain the odor label clustering statistics result; using the principal component analysis algorithm, according to the odor label clustering statistics result, performing data dimensionality reduction, screening data points according to the odor labels of each gas molecule, using Gaussian kernel density estimation to calculate the probability density of the odor labels on the target area grid, and determining the density value of each odor label corresponding to the grid; determining the odor in the target area and the odor space map in the target space according to the density value of each odor label corresponding to the grid.

[0090] As an example, the identification process of odor substances is as Figure 2 shown.

[0091] Optionally, the main function of the prediction result output module is to select the presentation method of the output result based on the prediction of the OR-POM. As an example, the odor labels with the top 1, 3, and 5 prediction probabilities can be screened as the output respectively, or the embedding vectors of each molecule can be output to construct an odor space map.

[0092] Among them, the construction of the odor space is based on the embedding vector of the penultimate layer in the model FFN. The method of kernel density estimation (Kernel Density Estimation, abbreviated as KDE) is used to perform odor label clustering statistics, and the principal component analysis (Principal Component Analysis, abbreviated as PCA) method is used for data dimensionality reduction, analyzing the maximum difference shown by the principal components, highlighting the main features, and completing the visualization operation. The visualization diagram is as Figure 3 shown. This module has been compiled into a python function, and by adjusting the parameters of the OR-POM model, the output of the main odor map combined with human olfactory receptor information is automatically completed.

[0093] It should be noted that based on the same training data set and the same test set, the performance comparison of the POM model and the OR-POM model is as Figure 4 shown. The vertical axis of the legend represents the accuracy rate, and the horizontal axis represents that the model outputs the top 1, 3, and 5 labels with the highest prediction probability respectively. It can be seen from the figure that the OR-POM model combined with the information of interacting with human olfactory receptors has better accuracy rates in the above three cases.

[0094] In addition, using the pre-trained model parameters and weights, the performance of the POM model and the OR-POM model is tested again using a stereoisomer data set not in the training set. The performance comparison is asFigure 5 As shown in the figure, it can be seen that the OR-POM model still leads in all indicators, and the leading margin in the test set has increased, indicating that the addition of virtual coding enables the model to have better discrimination ability for stereoisomeric odor molecules.

[0095] In addition, both POM and OR-POM can construct a main odor map for odor molecules. A method comparison is carried out based on the same data set. For example, Figure 6 As shown, the odor map constructed by the embedding vectors of OR-POM can better retain the odor label perception level while better distinguishing different major categories of odor perception.

[0096] This application constructs virtual coding based on the interaction relationship between odor molecules and human olfactory receptors and integrates it into the feature vectors of the graph neural network to construct the OR-POM model, achieving performance improvement. At the same time, it also proposes a new solution idea for the major challenge of distinguishing stereoisomeric odor molecules, and can accurately complete odor recognition based on substance identification.

[0097] The molecular odor recognition method, device, equipment, medium and product provided by the embodiments of this application can encode and recognize the odor of gas molecules according to the SMILES formula and olfactory receptor information of gas molecules, and can predict the area of the odor to generate an odor space map.

[0098] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0099] Based on the same inventive concept, the embodiments of this application also provide a sub-odor recognition control device for implementing the above-mentioned sub-odor recognition method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following sub-odor recognition device embodiments can refer to the limitations on the molecular odor recognition method in the above text, and will not be repeated here.

[0100] Please refer to Figure 7, in an exemplary embodiment, a molecular odor recognition device is provided, including:

[0101] A gas molecule acquisition module 10, configured to acquire a plurality of gas molecules in a target area;

[0102] A substance structure determination module 20, configured to identify substances of the plurality of gas molecules and determine the substance structure of each gas molecule;

[0103] A SMILES formula determination module 30, configured to determine the SMILES formula of each gas molecule according to the substance structure of each gas molecule;

[0104] A normalization processing module 40, configured to perform normalization processing on the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule;

[0105] A target virtual code generation module 50, configured to generate a target virtual code according to a plurality of target SMILES formulas and preset olfactory receptor information, where the target virtual code includes a plurality of code bits, and each code bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0106] An odor label determination module 60, configured to input the target virtual code and a plurality of target SMILES formulas into a trained OR-POM model to determine the odor label of each gas molecule;

[0107] An odor space map determination module 70, configured to determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0108] Each module in the above-mentioned molecular odor recognition device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0109] In an exemplary embodiment, a computer device is provided, which may be a terminal. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a method for identifying molecular odors. The display unit of the computer device is used to form a visually visible picture, which may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0110] Those skilled in the art can understand that Figure 8 the structure shown in

[0111] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0112] Obtain a plurality of gas molecules within a target area;

[0113] Perform substance identification on the plurality of gas molecules to determine the substance structure of each gas molecule;

[0114] According to the substance structure of each gas molecule, determine the SMILES formula of each gas molecule;

[0115] Perform normalization processing on the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule;

[0116] Generate a target virtual code according to multiple target SMILES formulas and preset olfactory receptor information, where the target virtual code includes multiple coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0117] Input the target virtual code and multiple target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule;

[0118] Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0119] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0120] Obtain multiple gas molecules in the target area;

[0121] Perform material identification on the multiple gas molecules to determine the material structure of each gas molecule;

[0122] Determine the SMILES formula of each gas molecule according to the material structure of each gas molecule;

[0123] Normalize the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule;

[0124] Generate a target virtual code according to multiple target SMILES formulas and preset olfactory receptor information, where the target virtual code includes multiple coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0125] Input the target virtual code and multiple target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule;

[0126] Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0127] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0128] Obtain multiple gas molecules in the target area;

[0129] Perform material identification on the multiple gas molecules to determine the material structure of each gas molecule;

[0130] Determine the SMILES formula of each gas molecule according to the material structure of each gas molecule;

[0131] Normalize the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule;

[0132] Generate a target virtual code according to multiple target SMILES formulas and preset olfactory receptor information, where the target virtual code includes multiple coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor;

[0133] Input the target virtual code and multiple target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule;

[0134] Determine the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0137] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application

[0138] Under the premise, several modifications and improvements can also be made, and these all fall within the protection scope of this application.

[0139] Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A method for identifying molecular odors, characterized in that, The method includes: Obtaining a plurality of gas molecules within a target area; Performing substance identification on the plurality of gas molecules to determine the substance structure of each gas molecule; Determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule; Performing normalization processing on the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule; Generating a target virtual code according to a plurality of target SMILES formulas and pre-set olfactory receptor information, where the target virtual code includes a plurality of coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor; Inputting the target virtual code and a plurality of target SMILES formulas into a trained OR-POM model to determine the odor label of each gas molecule; Determining the odor within the target area and the odor space map within the target space according to the odor label of each gas molecule.

2. The method according to claim 1, wherein The step of determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule includes: Determining the SMILES formula of each gas molecule according to a pre-established list of sample identification results and the substance structure of each gas molecule; Among them, the list of sample identification results is established through the following steps: Obtaining a variety of gas samples; Performing analysis and detection on each sample to obtain a mass spectrometry data file corresponding to the sample, where the mass spectrometry data file includes the mass spectrum information of each gas sample; Converting the format through Agilent unknown analysis software and performing non-targeted analysis and identification on the NIST2017 dataset to obtain a list of sample identification results.

3. The method according to claim 1, wherein The step of performing normalization processing on the SMILES formula of each gas molecule includes: Based on the RDkit package, converting the SMILES formula of each gas molecule into a molecular object; Performing standardization processing on each molecular object to obtain the target SMILES formula of each gas molecule, where the standardization processing includes: cleaning, removing metal bonds, neutralization, removing isotopes, and retaining stereoisomer information.

4. The method according to claim 1, wherein The pre-set olfactory receptor information includes multiple human olfactory receptors; Among them, generating a target virtual code according to a plurality of target SMILES formulas and pre-set olfactory receptor information includes: Pairing each target SMILES formula and each human olfactory receptor according to a plurality of target SMILES formulas and pre-set olfactory receptor information to generate a plurality of scoring values of the interaction probability between a plurality of target SMILES formulas and a plurality of human olfactory receptors; Constructing a multi-dimensional target virtual code according to a plurality of scoring values of the interaction probability between a plurality of target SMILES formulas and a plurality of human olfactory receptors.

5. The method according to claim 1, characterized in that, The OR-POM model performs the following steps: Using a graph neural network to extract the features in the molecular structures of a plurality of target SMILES formulas to obtain the structural feature vectors of a plurality of molecules, and splicing the structural feature vectors of a plurality of molecules with the virtual code to complete the fusion of human olfactory receptor interaction information; Among them, the OR-POM model includes a graph neural network, a Layer Normalization layer, and a Feed-Forward Neural Network (FFN). The graph neural network is used to extract the structural features of molecules. The Layer Normalization (LN) layer is used to fuse the interaction information of human olfactory receptors to prevent the model from overfitting. The Feed-Forward Neural Network (FFN) is used to adjust the vector dimension for predicting odor labels and modeling the embedding space.

6. The method according to claim 5, wherein The steps of determining the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule include: Obtaining the target embedding vector of the penultimate layer in the Feed-Forward Neural Network; Based on the target embedding vector, using the kernel density estimation algorithm to perform odor label clustering statistics to obtain the odor label clustering statistics result; Using the principal component analysis algorithm, according to the odor label clustering statistics result, performing data dimensionality reduction, screening data points according to the odor label of each gas molecule, using Gaussian kernel density estimation to calculate the probability density of the odor label on the target area grid, and determining the density value of each odor label corresponding to the grid; Determining the odor in the target area and the odor space map in the target space according to the density value of each odor label corresponding to the grid.

7. An apparatus for identifying molecular odors, characterized in that, The device includes: A gas molecule acquisition module for acquiring a plurality of gas molecules in the target area; A substance structure determination module for identifying the substances of the plurality of gas molecules and determining the substance structure of each gas molecule; A SMILES formula determination module for determining the SMILES formula of each gas molecule according to the substance structure of each gas molecule; A normalization processing module for normalizing the SMILES formula of each gas molecule to obtain the target SMILES formula of each gas molecule; A target virtual code generation module for generating a target virtual code according to a plurality of target SMILES formulas and preset olfactory receptor information. The target virtual code includes a plurality of coding bits, and each coding bit represents the interaction probability between a gas molecule and an olfactory receptor; An odor label determination module for inputting the target virtual code and a plurality of target SMILES formulas into the trained OR-POM model to determine the odor label of each gas molecule; An odor space map determination module for determining the odor in the target area and the odor space map in the target space according to the odor label of each gas molecule.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Odor prediction method, device and equipment, readable storage medium and program product

    CN121747762A

  • Digital olfaction analysis method based on deep learning

    CN121838940A