Hybrid multi-molecular intelligent sensing analytical deep learning method and device based on high-dimensional interaction correlation modeling
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
然而,解析和建模嗅觉感知仍然面临着重大的挑战需要解决
本发明提出的混合嗅觉感知深度学习框架适用范围广、实用价值高,可有效推动多相关产业发展,显著降低研发与应用试错成本。在具身智能机器人方面,本框架填补了机器人多模态感知的关键技术空白,赋予机器人可解释嗅觉感知能力,能够嵌入感知-决策回路实时处理气体传感器阵列信号,与视觉、听觉信息融合构建完整环境认知,拓展了机器人在安全监测、医疗健康、环境交互等场景的应用深度,使其突破预设任务的被动工作模式,可依托实时嗅觉感知主动识别异常、预判风险、自主调整行为策略,助力具身智能向多模态主动感知与自主决策方向升级。在食品与风味产业中,本发明可精准预测不同分子组合及浓度配比的气味效果,解决传统风味配方研发依赖经验、周期长、成本高的问题,可辅助快速筛选配方、节约原料损耗,同时具备气味逆向解析能力,可反推气味分子组成与配比,加速天然风味替代品开发,优化健康食品风味表现,具备良好科研与商业价值。在环境监测领域,本发明适配复杂混合气体与环境分子波动场景,弥补传统单一传感器检测的不足,实现气体质量精准实时评估,可构建贴合人类嗅觉标准的监测体系,为化工泄漏、环境异味扰动提供智能预警。在数字嗅觉领域,本发明搭建了分子到气味感知的高精度映射模型,攻克人机交互嗅觉缺失难题,可实现高保真气味合成与数字编码远程传输,在虚拟现实沉浸式交互、远程医疗等场景具备重要应用前景。
Smart Images

Figure CN122548477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of olfactory perception technology, and in particular to a hybrid multi-molecule intelligent perception and analysis deep learning method and device based on high-dimensional interactive association modeling. Background Technology
[0002] The sense of smell has accompanied humanity through millions of years of adaptive challenges—from hunting and gathering, danger recognition to reproduction—it has always been an essential sensory modality for individual survival and social cohesion (A. Keller, RC Gerkin, Y. Guan, A. Dhurandhar, G. Turu, B. Szalai, JD Mainland, Y. Ihara, CW Yu, R. Wolfinger. Predicting human olfactory perception from chemical features of odor molecules [J]. Science, 2017, 355(6327): 820-6.). Even in modern civilization, the sensory function of smell permeates various fields. In industrial production, olfactory monitoring is applied in scenarios such as food safety assurance and chemical leak early warning. In both Eastern and Western medical practices, olfaction has been used to aid in the diagnosis of pathological odors (KA Fulton, D. Zimmerman, A. Samuel, K. Vogt, SR Datta. Common principles for odor coding across vertebrates and invertebrates [J]. Nature Reviews Neuroscience, 2024, 25(7): 453-72.). Meanwhile, with the advent of the era of embodied intelligence, vision- and language-based models have successfully empowered industrial applications, enabling robots to perform complex tasks. Integrating olfactory intelligence into robotic systems, endowing them with the ability to assess safety and danger, will significantly enhance the robots' autonomous decision-making capabilities in different scenarios. However, the analysis and modeling of olfactory perception still face significant challenges that need to be addressed.
[0003] Current research primarily focuses on exploring the relationship between single molecules and odor perception, a significant limitation because most natural odors exist as mixtures of multiple molecules, rather than isolated compounds (G. Bratman, C. Bembibre, G. Daily, R. Doty, T. Hummel, L. Jacobs, P. Kahn Jr, C. Lashus, A. Majid, J. Miller. Nature and human well-being: the olfactory pathway. Sci Adv10: eadn3028 [Z]. 2024). The relationship between mixtures of multiple molecules and odor perception is exceptionally complex, exhibiting high-dimensional interaction mappings, such as... Figure 1As shown, the composition and concentration of molecules in a mixture, the complex interactions between molecules, and the spatiotemporal integration of receptor activation, among other factors, collectively form a multi-level, cross-scale sensory coding network that can significantly shape the final odor perception. Therefore, olfactory decoding of mixtures involves not only the combinatorial explosion of chemical space, but also fundamental neural mechanisms such as receptor competition and sensory emergence (M. Zhang, L. Zhu, J. He, Y. Liu, S. Ding, X. Lin. Clinical study on the application of a high-sensitivity electronic nose onthin-film gas sensor array technology combined with deep learning algorithm for early non-invasive diagnosis of chronic atrophic gastritis [J]. Biomedical Signal Processing and Control, 2025, 107: 107851.;L. Aziz, H. Adil, R. Sarwar. Artificial sensing: AI-driven electronic nose for real-time gas leak detection and food spoilage monitoring [J]. Sir Syed University Research Journal of Engineering & Technology, 2025, 15(1): 71-82.;M. Aleixandre, D. Prasetyawan, T. Nakamoto. Automatic scent creation by cheminformatics method [J]. Scientific Reports, 2024, 14(1): 31284.). Overcoming this bottleneck will enable us to build interpretable intelligent perception systems and give robots the ability to make more autonomous decisions in complex scenarios.
[0004] This invention aims to elucidate the core transduction mechanisms of multi-molecular mixtures in the biological olfactory system and establish a complete computational pathway from chemical mixing to neural encoding and then to perception formation. To this end, this invention innovatively proposes a deep learning framework that integrates biomimetic olfactory principles for accurate odor perception identification. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a hybrid multi-molecule intelligent perception and analysis deep learning method and device based on high-dimensional interactive association modeling.
[0006] The objective of this invention is achieved through the following technical solution: Firstly, this invention provides a hybrid multi-molecule intelligent perception analytical deep learning method based on high-dimensional interactive association modeling, the method comprising: (1) Deep neural decision forest is used to extract odor molecule features, and hypergraph neural network is used to extract olfactory receptor features. The two features are fused and input into a fully connected network to predict the decay sinusoidal response curve of a single odor molecule-single receptor. (2) For multiple receptors, the decay sinusoidal response curves of multiple corresponding single odor molecules-single receptors are obtained by using the method of step (1), and weighted fusion is performed to obtain the fusion response curve of single odor molecules-multiple receptors; based on the principle of odor perception difference, the fusion response curve between each pair of odor molecules is aligned with the odor perception difference to optimize the model, and the alignment deviation is used as the loss function to guide the generation of odor molecule-receptor response curves for training, thereby obtaining the mapping relationship between odor perception and fusion response curves, and obtaining the weight parameters of each receptor relative to a single odor molecule; (3) Based on the concentration weighting of odor molecules, the fusion response curve of multiple odor molecules-single receptor is obtained, and then multiple corresponding fusion response curves of multiple odor molecules-single receptor are obtained for multiple receptors. The fusion is performed based on the weight parameters of each receptor relative to the single odor molecule to obtain the fusion response curve of multiple odor molecules-multiple receptors. (4) Find the fusion response curves of multiple odor molecules and multiple receptors that are similar to the fusion response curves of multiple odor molecules and multiple receptors obtained in step (2), obtain the corresponding similarity weights as the odor perception weights of single odor molecules, perform weighted fusion of odor perception of multiple single odor molecules, obtain the odor probability distribution of multi-molecule mixture, and obtain the set of perceived odors of multi-molecule mixture based on threshold screening.
[0007] Secondly, the present invention also provides a hybrid multi-molecule intelligent perception and analysis deep learning device based on high-dimensional interaction and correlation modeling, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it implements the hybrid multi-molecule intelligent perception and analysis deep learning method based on high-dimensional interaction and correlation modeling.
[0008] Thirdly, the present invention also provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the aforementioned hybrid multi-molecule intelligent perception and analytical deep learning method based on high-dimensional interactive association modeling.
[0009] Fourthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned hybrid multi-molecule intelligent perception and analytical deep learning method based on high-dimensional interactive association modeling.
[0010] The beneficial effects of this invention are: The hybrid olfactory perception deep learning framework proposed in this invention has a wide range of applications and high practical value, which can effectively promote the development of many related industries and significantly reduce the cost of R&D and application trial and error. In the field of embodied intelligent robots, this framework fills the key technological gap in multimodal perception, endowing robots with interpretable olfactory perception capabilities. It can embed a perception-decision loop to process gas sensor array signals in real time, and fuse them with visual and auditory information to construct a complete environmental cognition. This expands the application depth of robots in scenarios such as safety monitoring, medical health, and environmental interaction, enabling them to break through the passive working mode of preset tasks. They can actively identify anomalies, predict risks, and autonomously adjust behavioral strategies based on real-time olfactory perception, helping embodied intelligence upgrade towards multimodal active perception and autonomous decision-making. In the food and flavor industry, this invention can accurately predict the odor effects of different molecular combinations and concentration ratios, solving the problems of traditional flavor formula development relying on experience, long cycles, and high costs. It can assist in rapid formula screening, save raw material losses, and also has odor reverse analysis capabilities, which can reverse-engineer the composition and ratio of odor molecules, accelerate the development of natural flavor substitutes, optimize the flavor performance of healthy foods, and possess good scientific research and commercial value. In the field of environmental monitoring, this invention is adaptable to complex mixed gas and environmental molecular fluctuation scenarios, overcoming the shortcomings of traditional single-sensor detection, and achieving accurate real-time assessment of gas quality. It can construct a monitoring system that conforms to human olfactory standards, providing intelligent early warning for chemical leaks and environmental odor disturbances. In the field of digital olfaction, this invention establishes a high-precision mapping model from molecules to odor perception, overcoming the challenge of olfactory loss in human-computer interaction. It can achieve high-fidelity odor synthesis and remote transmission of digital encoding, showing significant application prospects in scenarios such as virtual reality immersive interaction and telemedicine. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1This is a schematic diagram of the explicit and implicit mappings between single molecules and odor perception; the left side represents the explicit mapping and the right side represents the implicit mapping. The complex composition of multi-molecule mixtures, with varying types and concentrations of molecules, significantly affects odor perception, leading to implicit mappings.
[0013] Figure 2 This is a schematic diagram of the dataset in an embodiment of the present invention.
[0014] Figure 3 This is a flowchart of a hybrid multi-molecule intelligent perception and analysis deep learning method based on high-dimensional interactive association modeling provided by the present invention.
[0015] Figure 4 This is a schematic diagram of the attenuated sine wave response curve.
[0016] Figure 5 This is a schematic diagram of the model optimization through calibration differences.
[0017] Figure 6 This is a schematic diagram of the attention-weighted response curve fusion strategy.
[0018] Figure 7 This is a schematic diagram of the fusion of multiple molecules into multiple receptors.
[0019] Figure 8 This is a schematic diagram illustrating the overall network model design improvement.
[0020] Figure 9 This diagram illustrates the comparison between response curves estimated by different methods and the actual curves. (a) shows the neural response curves predicted by different methods; the method of this invention best matches the actual curve. (b) shows a comparison of the matching accuracy and peak alignment of the response curves obtained by different methods. (c) shows the matching accuracy for each multi-molecule mixture obtained by different methods; the method of this invention achieves the highest accuracy.
[0021] Figure 10 This is a schematic diagram comparing the Pearson correlation coefficients between molecular odor perception and response curves obtained by different methods.
[0022] Figure 11 This is a schematic diagram comparing the similarity of odor perception results for multi-molecular mixtures obtained by different methods.
[0023] Figure 12 This is a diagram comparing the odor perception prediction performance of multi-molecular mixtures obtained from different models.
[0024] Figure 13 This is a structural diagram of a hybrid multi-molecule intelligent sensing and analytical deep learning device based on high-dimensional interactive correlation modeling provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are merely illustrative and not intended to limit the invention.
[0026] Due to the lack of sufficient mixture-sensory datasets for model training, this invention does not employ an end-to-end mixture-sensory mapping method, such as... Figure 1 As shown. However, a large number of single-molecule sensing datasets are available. This invention uses single-molecule sensing data to guide multi-molecule sensing recognition, therefore, single-molecule receptor datasets and single-molecule sensing datasets were collected. To evaluate the performance of the method, real-world mixture similarity scoring datasets and mixture sensing datasets were also obtained. Figure 2 As shown, each dataset is described below: Figure 2In section a: Distribution of molecular-receptor dataset characteristics: To explore the interaction mechanism between odor molecules and olfactory receptors—a crucial step in olfactory perception—this invention collected a set of single-molecule receptor datasets. This dataset contains individual molecules and their corresponding response receptors. The distribution of molecules associated with each receptor is severely imbalanced; 18 receptors are associated with only 100-200 molecules, while 7 receptors are associated with 500-800 molecules. The molecular-receptor relationship data are sourced from several authoritative datasets, including ODORDB, ODORactor, and OlfactionDB (ODORactor: a web server for deciphering olfactory coding. Bioinformatics (Oxford, England) (2011). DOI:https: / / doi.org / 10.1093 / bioinformatics / btr385; Modena, D., et al. OlfactionDB: A Database of Olfactory Receptors and Their Ligands. Advances in Life Sciences 1.1(2011):1-5. DOI: https: / / doi.org / 10.1093 / bioinformatics / btr385). These resources have been empirically validated and have broad application value and reliability. Furthermore, this invention also references data compiled from other studies (Achebouche R, Tromelin A, Audouze K, et al. Application of artificial intelligence to decode the relationships between smell, olfactory receptors and small molecules[J]. Scientific reports, 2022, 12(1): 18817. DOI: https: / / doi.org / 10.1038 / s41598-022-23176-y), which contains a wealth of molecular-receptor interaction information. These datasets describe in detail the receptor characteristics of each molecular compound and label one or more molecules associated with olfactory receptors. The correspondences in the datasets are mainly derived from published academic papers, recording the interactions between specific molecules and olfactory receptors. This invention collects data manually, integrates and filters molecular-receptor associations from multiple databases, and eliminates duplicate entries to ensure the uniqueness of each data point.Furthermore, sample data containing missing values and mislabeled data were directly removed to avoid bias, thereby improving overall data quality. Redundant molecule-receptor pairs were removed based on unique identifiers. To ensure data reliability, experimental conditions and quantitative parameters were cross-checked. When inconsistencies existed in the dataset, a comprehensive literature search was conducted to resolve the conflicts. These integration strategies effectively eliminated data bias, and the final molecule-receptor correspondences were all verified and reliable. The final integrated dataset contained 5,023 unique single molecules and 104 olfactory receptors, ensuring high fidelity and reliability. This dataset will be used to subsequently construct a mixture-based olfactory perception biological model within the proposed framework.
[0027] Figure 2 b: Distribution of features in the molecular-odor dataset: Given the scarcity of large-scale mixed odor perception data, this invention innovatively uses single-molecule data to guide mixture odor prediction. This invention compiles a large amount of single-molecule perception data for model optimization. This dataset also exhibits a skewed distribution, with 2,160 molecules possessing 3-4 odors and 672 molecules possessing 6-10 odors. Furthermore, the frequencies of different odors vary. The single-molecule odor perception dataset originates from the Good Scents and Lefngwell databases (The Good Scents Company, Available online: http: / / www.thegoodscentscompany.com / ; Lefngwell & Associates. Flavor-Base. 9th Edition. Available online: http: / / www.lefngwell.com / favbase.htm), revealing the interaction between molecules and odors. These molecules are labeled with one or more odor descriptors. Both databases are publicly available and have undergone rigorous validation through biochemical experiments. Professional olfactory experts labeled odor molecules based on empirical test results, providing detailed odor characteristic descriptions for each compound according to actual olfactory properties. Therefore, the obtained molecule-odor correspondences are both verified and reliable. This invention ensures the consistency of data points across the dataset by removing duplicate entries and employs cross-dataset cross-validation to minimize differences and inconsistencies. Retaining consistent and reliable data significantly improves the overall accuracy and robustness of the final dataset. Furthermore, this invention removes samples with missing values and mislabeled entries to avoid introducing bias, while discarding entries lacking key information to prevent noise interference. These processes effectively improve the overall quality of the data required for subsequent analysis and model training. Finally, an accurate molecule-odor association dataset was constructed, containing 3,015 individual molecules and 151 unique odors.
[0028] Figure 2c in the table: Mixture similarity rating dataset: To rigorously verify the accuracy of our method in identifying the ability of mixtures to perceive, we base our data on existing literature (Snitz K, Yablonka A, Weiss T, et al. Predicting odor perceptual similarity from odor structure[J]. PLoS computational biology,2013, 9(9): e1003184. https: / / doi.org / 10.1371 / journal.pcbi.1003184;AmitDhurandhar, Hongyang Li, Guillermo A Cecchi, Pablo Meyer, Expansivelinguistic representations to predict interpretable odor mixture discriminability, Chemical Senses, Volume 48, 2023, bjad018. https: / / doi.org / 10.1093 / chemse / bjad018;Kowalewski J, Ray A. Predicting human olfactory perception from activities of odorant receptors[J]. IScience, (2020, 23(8)) A mixture similarity dataset was compiled, which contains odor perception similarity ratings between mixture pairs evaluated through human olfactory psychophysical experiments. It contains 524 unique single molecules forming 77 different mixtures and includes 381 pairs of similarity comparisons. The similarity scores for each mixture were obtained through multiple human perception experiments, and data collection followed strict olfactory psychophysical experimental standards to ensure data reliability and scientific validity. The complexity of the mixtures varied significantly, ranging from 4 to 43 constituent molecules. The mixtures in the dataset cover everything from simple combinations of a few molecules to complex mixtures consisting of dozens of molecules, reflecting the complexity and diversity of multi-molecular interactions in odor perception. In each experiment, participants were exposed to two mixture stimuli and asked to rate their perceived similarity. The similarity score was based on the participant's subjective judgment of the similarity between the two odor mixtures, ranging from 0 to 100, where 0 indicates completely different and 100 indicates completely identical. Each subject underwent rigorous training, and each mixture in the dataset was independently scored multiple times to ensure the stability and reliability of the scores.This method enables the present invention to accurately capture the perceptual similarity of different mixtures, laying a solid data foundation for research on the odor perception of mixtures. The present invention utilizes this dataset to evaluate the performance of the method, focusing on its ability to accurately predict the perceptual similarity of mixtures.
[0029] Figure 2 d: Mixture Perception Dataset: To evaluate the performance of the method in real-world scenarios, particularly its ability to accurately identify the odor perception of multi-molecular mixtures, this invention integrates two authoritative sources to construct a validation dataset: the Essential Oil Safety Database (Tisserand R, Young R. Essential oil safety: a guide for healthcare professionals [M]. Elsevier Health Sciences, 2013.) and the Good Scents Company Database (The Good Scents Company. (2024). The Good Scents Company InformationSystem. Retrieved from https: / / www.thegoodscentscompany.com). The Essential Oil Safety Database, as an authoritative reference in the field of essential oils and fragrances, provides detailed chemical composition data based on comprehensive chromatographic analysis, covering the constituent molecules and their concentration ratios of various essential oils. The Good Scents Company Database is a standard resource in the fragrance and flavor industry, providing standardized odor perception data for thousands of raw materials and natural mixtures. This invention extracts the chemical composition of natural mixtures from the Essential Oil Safety Database, recording the constituent molecules and their corresponding concentration ratios. Subsequently, odor perception labels for these mixtures are obtained from the Good Scents Company Database. This process enables the present invention to establish a precise mapping relationship between multi-molecular mixtures and odor perception. Finally, the present invention compiles a dataset containing 172 real-world test samples as an independent benchmark for validating model performance. Through this dataset of real-world mixtures and their corresponding odor perception datasets (including the specific chemical composition and concentration ratios of molecules in the mixtures), the present invention can effectively evaluate the model's ability to recognize the odor perception of multi-molecular mixtures in real-world environments, thereby further validating the effectiveness and generalization ability of the method.
[0030] This invention leverages abundant existing single-molecule perception data to guide mixture perception prediction by establishing semantic associations between molecule and mixture response patterns. This allows the model to transfer knowledge from the rich semantic space of single molecules, achieving robust recognition of odor features in complex mixtures even with limited labeled data. Furthermore, this invention develops a concentration-dependent multi-molecule curve fusion strategy for mixture multi-receptor response curves to simulate competitive activation and cooperative integration of mixture components. The model is optimized by minimizing the difference between molecular odor perception and response curves. This fusion strategy effectively reflects the influence of the composition and concentration of individual molecules in the mixture on the final odor perception. Finally, by comparing the consistency of single-molecule and multi-molecule response curves and dynamically assigning differentiated weights to odor features based on pattern similarity, accurate identification of odor perception in multi-molecule mixtures can be achieved. The overall workflow is as follows: Figure 3 As shown.
[0031] First, this invention develops a deep learning-based molecule-to-receptor response curve prediction model. Second, this invention utilizes the consistency between molecular odor perception and response curves, optimizing the model by minimizing and aligning the differences between the two, enabling the model to learn a stable mapping relationship between perception and response curves. Third, based on biological response principles, this invention establishes an attention-weighted multi-receptor response curve fusion strategy to construct response curves at the multi-receptor level. Fourth, considering the significant impact of molecular concentration on odor perception, this invention designs a concentration-perceived mixture response curve fusion strategy. Fifth, this invention uses single-molecule odor perception data as prior knowledge to guide the recognition of mixture odor perception, thereby achieving a precise mapping from the neural response domain to the odor perception space. The specific steps are as follows: Step 1 ( Figure 3 (a) ): Prediction of Molecular-Receptor Response Curves Based on Deep Learning. Based on molecular sequence features and the three-dimensional structure of olfactory receptors, this invention specifically designs a deep learning model to capture their interaction patterns. Deep Neural Decision Forest (DNDF) and Hypergraph Neural Network (HGNN) are used to extract feature representations of molecules and receptors, respectively. The fused feature vectors are input into a fully connected network to predict key features of the neural response curve, thereby reflecting the dynamic interaction between each pair of molecules and receptors. Furthermore, inspired by mature olfactory biological mechanisms, this invention uses a decaying sine wave model to construct neural response curves, simulating the characteristic temporal discharge dynamics of olfactory receptors after odor stimulation. This step lays the foundation for constructing hybrid response curves and simulating overall olfactory perception.
[0032] In this step, improved DNDF and HGNN models are used to describe the interaction between molecules and olfactory receptors. Structural features are extracted and fully connected neural networks are designed to predict the neural response curves characterizing their interaction. Furthermore, based on biological mechanisms, a decaying sine wave function is used to mathematically simulate the dynamic neural response characteristics induced by molecule-receptor interactions. The specific methods are as follows: (1) Data characteristics of molecules and receptors Molecular sequence data and receptor structure data were used as features, and a deep learning model was employed for modeling. The data features of molecules and receptors are described below: Each molecule M Through high-dimensional feature vectors This indicates that the molecular sequence data is composed of two types of information:
[0033] This represents physicochemical descriptor data, containing 1,444 features such as molecular weight, topological polar surface area, number of atoms, and molecular refractive index. These features are used to describe the physicochemical properties and biological activities of molecules.
[0034] Representing molecular fingerprint data, this feature vector maps molecular structures to binary vectors, reflecting their atom types, chemical bonding patterns, functional groups, and molecular substructures. This feature vector integrates physicochemical properties with molecular topology, providing rich input for subsequent deep learning modeling.
[0035] Secondly, to accurately capture the three-dimensional conformational features of olfactory receptors, we use a graph structure to represent the receptor conformation. RR Represented as ,in: Node set : Each atom in the vector corresponds to a node, whose attribute vector... It includes properties such as atom type (e.g., C, N, O), partial charge, atomic mass, and van der Waals radius.
[0036] hyper-edge set Each chemical bond (including covalent bonds, hydrogen bonds, etc.) has a superedge. This indicates that its attribute vector It includes geometric and energy parameters, such as bond type (single bond, double bond, etc.), bond length, bond angle, and dihedral angle.
[0037] Molecular sequence data were processed using a hybrid feature set that combined physicochemical descriptors (Moriwaki H, Tian YS, Kawashita N, et al. Mordred: a molecular descriptor calculator[J]. Journal of cheminformatics, 2018, 10(1): 4. https: / / doi.org / 10.1186 / s13321-018-0258-y) with molecular fingerprints (Konda R, Reddy ST, Moon SA, et al. AI-Driven Drug Discovery: Leveraging Machine Learning for Predictive Molecular Design and Accelerated Pharmaceutical Innovation[C] / / 2025 IEEE 4th World Conference on Applied Intelligence and Computing (AIC).0[2026-01-22]. https: / / doi.org / 10.1109 / AIC66080.2025.11212021). The two vectors have dimensions of 1444 and 473 respectively, and when concatenated, they form a feature vector with a total dimension of 1917. This fusion ensures robust characterization of molecular physicochemical properties and topological structure. Specifically, physicochemical descriptors are used to characterize molecular properties that are crucial for analyzing the correlation between molecular structure and biological activity. Parameters such as molecular weight, topological polar surface area, number of atoms, and molecular refractive index help characterize molecular behavior, interaction patterns, and biological activity. Furthermore, molecular fingerprints, as a numerical representation, capture molecular structural information by encoding atomic and bond-level features. These fingerprints contain information such as functional groups, atomic connections, and molecular substructures, and can be used for comparative analysis and similarity assessment between molecules.
[0038] To obtain the three-dimensional conformational information of olfactory receptors, a comprehensive multi-attribute encoding scheme was used to characterize each receptor. This scheme covers atomic coordinates, charge, mass, bonding parameters, and other key physical features (Shaw DE, Maragakis P, Lindorff-Larsen K, et al. Atomic-level characterization of the structural dynamics of proteins[J]. Science, 2010, 330(6002): 341-346.https: / / doi.org / 10.1126 / science.1187409; Billesbølle CB, de March CA, vander Velden WJC, et al. Structural basis of odorant recognition by a human odorant receptor[J]. Nature, 2023, 615(7953): 742-749. https: / / doi.org / 10.1038 / s41586-023-05798-y). Data extraction was achieved by integrating multiple structural file formats: residue and atom identification information was derived from PDB files; spatial coordinates were taken from GRO files; and force field parameters (including charge, mass, and bond / dihedral constraints) were extracted from TOP files. These integrated features reflect both the physicochemical interactions between atoms and the spatial topological arrangement of residues, providing a rigorous biophysical basis for modeling the three-dimensional structural features of receptors.
[0039] (2) Deep learning modeling of molecules and receptors This invention develops a deep learning model to simultaneously model molecules and receptors, aiming to accurately extract structural features. For molecular sequence data, we prioritize algorithms that can efficiently handle high-dimensional numerical feature vectors. Given the effectiveness of the Random Forest (RF) algorithm in processing such data, we employ the Deep Neural Decision Forest (DNDF) model. DNDF effectively bridges the gap between the interpretability of RF-based methods and the representation learning capabilities of deep neural networks. By implementing a stochastic differentiable decision tree, DNDF can achieve global optimization of the splitting node parameters through backpropagation. This mechanism enables DNDF to effectively capture complex nonlinear patterns in molecular sequence data.
[0040] During the forward propagation process, each sample It needs to be processed through multiple decision trees, and the final output is the weighted sum of the embedding vectors of all leaf nodes:
[0041] in Indicates that the sample belongs to the first The probability of a leaf node. Let be the embedding vector of this leaf node. The splitting function for each tree node is modeled using a differentiable sigmoid function:
[0042] All tree node parameters Both methods employ backpropagation and gradient descent for global optimization, enabling the model to adaptively capture complex nonlinear patterns in molecular features.
[0043] To address the complex three-dimensional topology of olfactory receptors using receptor structure data, we employed a Hypergraph Neural Network (HGNN) framework. This architecture is particularly well-suited for capturing key spatial dependencies in ligand binding. We represent the receptor structure as a hypergraph: atoms as nodes (with associated properties such as type, charge, and mass), and chemical bonds as hyperedges (with associated properties such as type, length, and dihedral angle). Through a multi-layer graph convolution strategy, the model iteratively aggregates information from neighboring nodes to generate a representation reflecting the spatial conformation of the receptor.
[0044] The structural features of the receptor are modeled using an enhanced hypergraph neural network (HGNN). The core operation of HGNN is hypergraph convolution, which... The update rule for layer node representation is defined as follows:
[0045] Presentation layer Feature matrix of all nodes. and These represent the node degree matrix and the hyperedge degree matrix, respectively, used for normalization. Let be the adjacency matrix of the hypergraph, where Represents a node Belongs to superedge . It is a learnable hyperedge weight matrix. It is a layer The trainable weight matrix. This represents a non-linear activation function.
[0046] go through After each convolutional layer, global average pooling is used to obtain a comprehensive characterization of the receptor:
[0047] Therefore, deep neural networks (DNDF) and high-dimensional neural networks (HGNN) were selected as the core networks for molecular and receptor modeling to accurately extract structural features.
[0048] (3) Single-molecule-single-receptor response curve This invention designs a fully connected neural network that predicts neural response curves by taking molecular and receptor features as input. Furthermore, this invention employs a decaying sine wave function to mathematically simulate the dynamic neural response characteristics induced by the interaction between a single molecule and a single receptor. Molecular features With receptor characteristics After splicing, the parameters of the attenuated sine wave are predicted using a fully connected neural network (FCN):
[0049] in These represent amplitude, damping coefficient, angular frequency, and phase shift, respectively. This fully connected neural network employs multiple hidden layers and uses the ReLU activation function, while also applying dropout regularization to mitigate overfitting.
[0050] Based on established findings in olfactory neurophysiology—that neurons exhibit a characteristic pattern of rapid activation, oscillatory firing, and gradual decay in response to odor stimuli—we model the neural response curve using a decaying sinusoidal function:
[0051] Each parameter in this formula has a unique biological interpretation: amplitude : Reflects the peak discharge intensity during the initial activation phase and is related to ligand-receptor binding affinity. Damping coefficient : Controlling the rate of exponential decay to baseline, simulating the receptor desensitization process. Angular frequency : Determines the oscillation rhythm, corresponding to the synchronous firing frequency of the ensemble of neurons in the olfactory cortex. Phase shift : Represents the time delay of the response initiation, capturing the initial state of signal transduction and neuron activation.
[0052] To ensure physiological rationality, based on experimental observations of olfactory receptor neurons, the response function must satisfy the following constraints:
[0053]
[0054]
[0055]
[0056] These boundaries limit the prediction curve to a biologically feasible range, thereby improving the reliability of the model in simulating real neural dynamics.
[0057] The mathematical model used in this step is based on its ability to accurately reproduce the fundamental neurophysiological characteristics of the biological olfactory system. First, the model aims to capture the mechanisms of neural adaptation and receptor desensitization. Physiological studies (Kim WK, Choi K, Hyeon C, et al. General chemical reaction network theory for olfactory sensing based ong-protein-coupled receptors: Elucidation of odorant mixture effects and agonist–synergist threshold[J]. The Journal of Physical Chemistry Letters, 2023, 14(38): 8412-8420. https: / / doi.org / 10.1021 / acs.jpclett.3c02310) have shown that olfactory receptor neurons do not maintain a constant firing rate under continuous odor stimulation. Instead, they typically exhibit a characteristic response pattern: a brief, high-intensity phase response upon initial stimulation, followed by a rapid decline back to baseline levels at an exponential decay rate. This dynamic decay behavior reflects the receptor's adaptation to continuous stimulation. Therefore, we model this biophysical phenomenon using the exponential decay term in the wavefunction. Secondly, the model introduces a periodic term to reflect the core intrinsic temporal oscillations of olfactory coding (Nagel KI, Wilson R I. Biophysical mechanisms underlying olfactory receptor neuron dynamics[J]. Nature neuroscience, 2011, 14(2): 208-216. https: / / doi.org / 10.1038 / nn.2725. Su CY, Menuz K, Reisert J, et al. Non-synaptic inhibition between grouped neurons in an olfactory circuit[J]. Nature, 2012,492(7427): 66-71. https: / / doi.org / 10.1038 / nature11712). The olfactory system does not rely on static coding, but rather generates rhythmic oscillatory activity through the olfactory bulb and the collection of neurons in the cortex. These oscillations act like an internal clock, synchronizing the firing sequence of neurons, thus providing a key time window for accurate odor discrimination.The sine term in the wave function effectively characterizes the rhythmic fluctuations of neuronal excitability and its periodic changes over time.
[0058] Step 2 ( Figure 3 (b) in the paper: Optimizing the model by aligning response curves with odor perception differences. The differences in odor perception between different molecules exhibit a high degree of consistency with the differences in their neural response curves: the greater the perception difference, the more significant the deviation in the response curve, and vice versa. Utilizing this inherent consistency, this invention uses the deviation between the differences in intermolecular response curves and the perception differences as the loss function for model optimization, while simultaneously utilizing single-molecule odor perception data and the corresponding multi-receptor neural response curves. By minimizing this deviation, the model can learn a stable and coherent mapping relationship between odor perception and neural response curves, thereby improving the prediction accuracy of molecule-to-receptor neural response curves and establishing a reliable foundation for the perception of multi-molecule mixtures.
[0059] This invention uses a deep learning model to predict the neural response curves of single molecules and single olfactory receptors, and can derive the fusion response curve of a single molecule on multiple receptors—that is, the final molecular response curve. However, a key challenge remains: the lack of rich real-world neural response records as direct supervisory labels for training and optimizing the model. To address this, this invention innovatively proposes a perception-guided contrastive learning unsupervised training strategy. This method guides the generation of molecule-receptor response curves by coordinating the differences between molecular neural response curves and corresponding odor perceptions.
[0060] The goal of this invention is to reconcile the discrepancy between neural response curves and differences in odor perception. Specifically, there is a strict mapping relationship between the perceptual properties of odor molecules and the neural dynamic patterns they induce. When there is a significant difference in the odor perception of two molecules, the corresponding neural response curves also show obvious morphological differences, and vice versa (Stopfer M, Jayaraman V, Laurent G. Intensity versus identity coding in an olfactorysystem[J]. Neuron, 2003, 39(6): 991-1004. https: / / doi.org / 10.1016 / j.neuron.2003.08.011). This indicates that the divergence in neural response trajectories is not simply due to the feedback effect of different chemical stimuli, but constitutes the basic biological mechanism by which the brain distinguishes odor perception. Therefore, we use a single-molecule odor tag set as a discriminant to align and minimize the deviation between the response curves and odor perception. For each pair of molecules, we calculate the difference in response curves using the root mean square error (MAE) and measure the difference in odor perception using the cosine distance. These two indicators are aligned using the Pearson correlation coefficient. The deviation between the differences in molecular response curves and the perceived differences is used as a loss function, and the deep learning model is optimized by minimizing this loss. The dataset contains 5,023 single molecules, and this invention calculates the pairwise deviations of all molecular pairs, generating approximately 12 million deviation values, which is sufficient to meet the model optimization requirements.
[0061] like Figure 5 As shown, the model is optimized by aligning the neural response curves between each pair of molecules with the differences in odor perception. This alignment bias is used as a loss function to guide the generation of molecule-receptor response curves. The ultimate optimization goal is to ensure that the difference between the neural response curves of any two molecules is isomorphic to their differences in odor perception. In other words, a large difference in odor perception corresponds to a large difference in neural response curves, and vice versa.
[0062] By leveraging existing single-molecule odor-sensing tags, we guide the generation of response curves, thereby optimizing the neural network model developed in step 1. Based on the model introduced in step 1 and the attention-weighted multi-receptor response curve fusion strategy, when the molecule... The integrated neural response curve when interacting with the olfactory receptor set ℛ The prediction formula is as follows:
[0063] in , They represent molecules respectively Assigned to receptor The attention weight reflects the relative contribution of the receptor to the overall perceptual response.
[0064] This method aims to ensure that the generated response curves align with perceptual differences, even in the absence of real neural response labels. Specifically, the model optimizes the correlation between predicted response curves and the differences in odor perception characteristics among different molecules. This perception-guided approach ensures that molecules with different odor perception characteristics produce dissimilar response curves, while molecules with similar perceptions exhibit similar neural response patterns.
[0065] (1) Difference between quantitative response curve and odor perception For any two molecules and Its neural response curve and Differences through observation period T Quantization of root mean square error (MAE) within:
[0066] Perceptual difference is achieved through perceptual label vectors and Cosine distance measurement:
[0067] To ensure that the differences in response curves are consistent with the perceived differences, we calculated all molecular pairs Pearson correlation coefficient between :
[0068] in and Representing all unique molecular pairs and The average value.
[0069] The goal is to maximize This ensures a strong positive correlation between differences in neural responses and perceptual differences. Therefore, alignment loss is defined as a negative correlation coefficient:
[0070] Minimizing this loss can prompt the model to generate response curves whose pairwise differences are consistent with the corresponding differences in odor perception.
[0071] (2) Model optimization based on contrastive learning To enhance the model's ability to distinguish between perceived similar molecules, a contrastive learning component is introduced. For each anchored molecule... Its positive samples Defined as the molecule with the most similar odor perception characteristics, while negative samples This corresponds to the molecule with the greatest perceived difference. The contrast loss formula is as follows:
[0072] The total loss function is defined as a weighted combination of alignment loss and contrast loss:
[0073] Parameters of deep learning models This is achieved by optimizing through backpropagation and minimizing the total loss:
[0074] The dataset contains FF = 5,023 molecules, forming multiple unique molecular pairs:
[0075] This equates to approximately 12.6 million difference pairs, providing a strong and comprehensive supervisory signal for model optimization even in the absence of explicit neural response labels.
[0076] Step 3 ( Figure 3 (c) : Attention-weighted multi-receptor response curve fusion strategy. Olfactory perception originates from the co-activation of multiple olfactory receptors, each contributing differently to the final perception. Based on this biological principle, this invention incorporates an attention mechanism into the proposed framework, dynamically assigning unique weights to different olfactory receptors to capture the relative importance in the perception formation process. Based on this mechanism, this invention designs a multi-receptor response curve fusion strategy to effectively integrate the collective responses of multiple receptors. Through this strategy, this invention can construct single-molecule to multi-receptor response curves for the second-step model training, and further generate multi-molecule to multi-receptor response curves, which serve as key inputs for subsequent mixed odor perception modeling.
[0077] To accurately capture the differences in the influence of different receptors on the perception of specific odor molecules, this model integrates an attention module into a deep learning framework. This is because the degree of influence of each receptor on the final odor perception varies (Si G, Kanwal JK, Hu Y, et al. Structured odorant response patterns across a completeolfactory receptor neuron population[J]. Neuron, 2019, 101(5): 950-962. e7.https: / / doi.org / 10.1016 / j.neuron.2018.12.030). For a specific odor molecule, not all receptors contribute equally to the formation of perception—only a small subset of key receptors determine the final olfactory characteristics. Therefore, the attention module aims to simulate this biological selection mechanism: through adaptive learning and weight allocation based on molecular-receptor interaction features, this module can effectively amplify the salience of key receptors while suppressing the interference of redundant receptors. Specifically, the attention module of this invention uses the concatenation of molecular features and receptor features as input and predicts the relevance score through nonlinear mapping. To ensure that the weights satisfy the probability distribution characteristics, we use the softmax function to normalize the scores, making the sum of the weights of all receptor sets always equal to 1, thus obtaining the final attention weights. This module has been integrated into the main framework for joint training and optimization.
[0078] Since olfactory perception relies on the collective activity of the entire receptor spectrum rather than the isolated signal of a single receptor, integrating the neural responses generated by all receptors is crucial. For multi-receptor response curves, this invention employs an attention-weighted multi-receptor response curve fusion strategy. We simulate the biological signal integration process by superimposing the response curves of each receptor in the time domain, thereby preserving the dynamic temporal characteristics of neural activity (Chong E, Moroni M, Wilson C, et al. Manipulating synthetic optogenetic odors reveals the coding logic of olfactory perception[J]. Science, 2020, 368(6497): eaba2357. https: / / doi.org / 10.1126 / science.aba2357). Furthermore, this invention develops a time-point weighted linear superposition method: for the response curves induced by molecules in different receptor spectrums, the response curves are weighted and averaged according to the corresponding attention weights to calculate the response value at each time point, ultimately generating a global response curve for all receptors, such as... Figure 6 As shown.
[0079] Because different olfactory receptors contribute unevenly to odor perception, a simple unweighted average of response curves would dilute key response features. Therefore, an attention-based weighting mechanism is introduced to enhance the contribution of key receptors. The attention module dynamically assigns different weights to each active receptor and performs a weighted average of individual response curves based on these weights, thereby generating a global response curve for all receptors.
[0080] Step 3 is implemented as follows: (1) Receptor Contribution Scoring Attention Module: To adaptively quantify the differences in the influence of each olfactory receptor on odor perception, this framework integrates the attention mechanism into the deep learning system. Let: This represents the feature vector of the input molecule. This represents the feature vector of the olfactory receptor.
[0081] First, the molecule-receptor pairs are spliced together, and then the original correlation score is calculated by projecting a nonlinear transformation:
[0082] in For learnable weight matrix, For bias vectors, Represents the ReLU activation function. To output the projection vector, This indicates vector concatenation.
[0083] The original scores of all receptors are then normalized using the softmax function to obtain the final attention weights.
[0084] These weights satisfy and , represents the probability distribution on the receiver set.
[0085] (2) Weighted fusion of multi-receptor response curves: for a given molecule receptor The induced individual neural response curve is defined as:
[0086] in These are the parameters predicted in step 1.
[0087] Global response curve The contributions of all receptors are integrated through time-point attention-weighted summation calculation:
[0088] This formula ensures that receptors with higher attention weights have a greater impact on the fusion response, while preserving the time-domain dynamics of each individual curve.
[0089] Step 4 (Figure 3, (d)): Concentration-Perceived Mixed Response Curve Fusion Strategy. Given the significant differences in molecular concentration among the components in the mixture, and the crucial influence of molecular concentration on shaping neural response curve characteristics and odor perception, this invention proposes a concentration-perceived mixed response curve fusion strategy. Specifically, for molecules under different concentration conditions, this invention first constructs individual response curves between the molecule and each olfactory receptor based on Step 1. Then, a concentration-perceived weighted curve fusion mechanism is employed to capture the moderating effect of concentration changes on receptor activation patterns. Finally, for multi-molecule mixtures, specific interactive response curves between the mixture and each olfactory receptor can be derived. Subsequently, these receptor-level response curves are integrated using the attention-weighted multi-receptor response curve fusion strategy introduced in Step 3 to generate a unified response curve, effectively characterizing the overall odor perception features of multi-molecule mixtures.
[0090] Based on the deep learning model designed above, this invention can accurately predict the neural response curves induced by a single molecule to multiple receptors. Step 4 aims to generate a fusion response curve of the entire olfactory receptor library triggered by a mixture of multiple molecules. This process consists of two stages: Stage 1: Multiple molecules interact with a single receptor to generate the response curve of the single receptor to the mixture. Stage 2: Multi-receptor integration, i.e., multiple molecules interact with multiple receptors to generate the overall response curve of multiple receptors to the mixture. Multi-molecule and single-receptor interaction: This involves the interaction between multiple molecules and a single olfactory receptor to generate the response curve of that receptor to the mixture. Stage 2: Multi-receptor integration: Through the proposed attention-weighted response curve fusion strategy, the response curves generated by all receptors are globally integrated to obtain the collective neural response characteristics triggered by the mixture.
[0091] In complex mixture environments, individual olfactory receptors are subject to competitive binding by multiple molecules. Therefore, this step aims to analyze the synergistic effects of multiple molecules, particularly by constructing fusion neural response curves induced by mixtures at the individual receptor level (Singh V, Murphy NR, Balasubramanian V, et al. Competitive binding predicts nonlinear responses of olfactory receptors to complex mixtures[J]. Proceedings of the National Academy of Sciences, 2019, 116(19):9598-9603. https: / / doi.org / 10.1073 / pnas.1813230116). Given the significant influence of molecule type and concentration in mixtures on odor perception, this invention proposes a concentration-based fusion strategy for mixture response curves. Figure 7 As shown, the concentration ratio of each molecule in the mixture is first calculated and normalized. Then, the deep learning models from steps 2 and 3 are used to predict the neural response curve of a single molecule to a single olfactory receptor. Subsequently, concentration is used as the weight of the molecular response curve to measure its competitive binding advantage at the receptor binding site. Finally, the weighted response curves of all molecules are fused through temporal linear superposition: at each time point, the response values of all molecules are summed and averaged according to their concentration weights, thereby generating a global response curve of a single receptor to a multi-molecule mixture. After obtaining the single-receptor response curve, the attention-weighted multi-receptor response curve fusion strategy described in step 2 can be used for integration. Through this process, a global neural response curve of a multi-molecule mixture covering the entire receptor spectrum can be obtained.
[0092] Step 4 is implemented as follows: Concentration-weighted fusion at the single-receptor level: The concentration of each component molecule in a mixture significantly affects odor perception. To incorporate this effect into neural response modeling, we introduce a concentration-proportion weighting mechanism during response curve generation. Considering the... A mixture of different odor molecules The concentrations of each molecule are respectively The total concentration of the mixture is:
[0093] molecular The normalized concentration ratio is given by the following formula: , in =1 in Representative based on concentration to endow molecules The weight of the ligand reflects its relative abundance, and thus reflects its competitive binding advantage at the receptor binding site under multi-ligand conditions.
[0094] For each molecule The model described in Sections B and C predicts its interaction with the receptor. Single-molecule-single-receptor response curve The response curves were then weighted according to their respective concentration ratios. (Receptor) Mixed response curve The weighting factor was obtained by linearly superimposing the response curves of each single molecule over time. :
[0095] This generates a concentration-modulated neural response curve of a mixture at the level of a single receptor.
[0096] (2) Integration of the olfactory receptor library: In obtaining each receptor Mixed response curve Then, we applied the attention-weighted multi-receptor fusion strategy outlined in step 2 to integrate the responses from the entire olfactory receptor repertoire. Global mixed response curve. The calculation method is as follows:
[0097] in Receptor The attention weights are learned by the attention module described in step 2. These weights reflect the receptor's... Perceptual relevance to a specific mixture.
[0098] Through this two-stage fusion process, we obtained the integrated neural response curves induced by multi-molecular mixtures across the entire receptor spectrum, effectively simulating the integrated olfactory signals behind complex odor perception.
[0099] Step 5 (Figure 3(e)): Response curve similarity metric for mixture odor perception determination. This invention utilizes single-molecule odor perception as prior information to guide the odor perception recognition of multi-molecule mixtures. Specifically, by calculating the similarity index between the mixture response curve and the individual molecule response curve, this invention adaptively assigns weights to the odor perception vectors associated with different molecules, thereby reflecting their differences in influence on overall odor perception. These weighted odor vectors are aggregated and then filtered through a threshold to determine the final perception output. Ultimately, this model can robustly predict the multi-label odor perception of the target mixture, achieving accurate mapping from the neural response domain to the odor perception space, and completing the final determination of the mixture odor perception.
[0100] Step 5 determines odor perception by comparing the similarity between the neural response curve of the mixture and the well-annotated single-molecule response curve. Let... The global neural response curve representing the target mixture is generated using a concentration-aware fusion strategy combined with an attention-weighted multi-receptor response curve fusion strategy. A well-annotated single-molecule reference dataset is also provided.
[0101] in It is the first The neural response curves of each reference molecule are accurately predicted by the aforementioned model. It is the corresponding odor perception vector, with each dimension representing the intensity of the perceived attribute (such as floral, fruity, woody, etc.).
[0102] Subsequently, the mixed response curve was calculated. With each reference curve Cosine similarity between To convert similarity scores into interpretable contribution weights, soft maximization normalization is applied:
[0103] in This is a temperature parameter that controls the sharpness of the weight distribution. The weights satisfy:
[0104]
[0105] therefore, Essentially, it is a probability distribution of perceived attributes.
[0106] To extract a discrete set of dominant odor tags, a confidence threshold needs to be applied. :
[0107] in express of Number of elements, threshold It can be adjusted according to the required specificity and recall rate.
[0108] After generating global neural response curves for multi-molecule mixtures, the final task is to decode the corresponding perceived odors from these response curves. We employ a similarity-based inference strategy, which is based on the isomorphic truth between response curves and odor perception—the theory that molecular stimuli evoking similar neural response patterns often have similar odor perceptions (Pashkovski SL, Iurilli G, Brann D, et al. Structure and flexibility incortical representations of odor space[J]. Nature, 2020, 583(7815): 253-258.https: / / doi.org / 10.1038 / s41586-020-2451-1). Therefore, we utilize a complete labeled dataset containing information on single molecules and their corresponding odor perceptions as a reference library. By measuring the similarity between the neural responses of mixtures and single molecules, we achieve the transfer of odor perception from the single-molecule domain to the mixture domain.
[0109] The specific reasoning process is as follows: 1) Similarity measurement calculation: First, the root mean square error (MAE) is used to calculate the geometric similarity between the response curve of the mixture and the response curve of each single molecule, and the similarity score is quantified.
[0110] 2) Weight Normalization: The obtained similarity scores are normalized and used as the contribution weight of each individual molecule's odor perception. The higher the similarity, the closer the odor perception of that individual molecule is to that of the mixture.
[0111] 3) Weighted semantic aggregation: Based on the calculated weights, the odor perception of all single molecules is aggregated by weighted average, thereby generating the odor probability distribution of the mixture in the perception space.
[0112] 4) Threshold Filtering: Finally, a confidence threshold is applied for filtering. Odor tags with a polymerization probability exceeding the threshold are retained, collectively forming the set of perceived odors for multi-molecular mixtures.
[0113] Through the above steps, the design improvement of the overall network model constructed by this invention is as follows: Figure 8 As shown.
[0114] like Figure 8As shown in (a), given the varying importance and inherent high sparsity of features in molecular sequence data, we replace the convolutional neural network (CNN) module in the original DNDF with a feature extraction module composed of a multi-head self-attention mechanism and a fully connected network (FCN). This module treats the molecular sequence as a high-dimensional sparse numerical input, effectively extracting information-rich feature vectors.
[0115] 1) Multi-head Self-Attention Mechanism. We introduce a multi-head self-attention mechanism to capture diverse key features embedded in molecular sequences (such as hydrophobicity, steric effects, and electronic properties), while dynamically assigning weights to each descriptor. Molecular sequence features (such as molecular weight, degree of polymerization, Zagreb index, and MACCS fingerprint) originate from heterogeneous dimensions and exhibit complex interdependencies. The self-attention mechanism automatically learns these associations and generates context-aware representations, enabling the model to construct meaningful feature embeddings. Unlike traditional neural networks that treat all input features equally, the attention mechanism allows the model to focus on informative features while mitigating the influence of irrelevant or noisy features. This is particularly important for molecular modeling tasks—accurately identifying key features such as aromaticity and polarity is crucial. Multiple attention heads enhance the model's representational capabilities by collaboratively focusing on different subspaces of physicochemical information (such as logP and molar refractive index). Furthermore, the model can dynamically adjust the importance of each descriptor based on the molecular structure, thereby capturing nonlinear and higher-order interactions between molecular sequences and significantly improving the accuracy of molecular property predictions.
[0116] 2) Fully Connected Network (FCN). To efficiently extract and transform high-dimensional molecular sequence features with input dimensions up to 2000, we employ a multi-layer fully connected network (FCN) architecture. FCN transforms sparse high-dimensional fingerprint data (such as MACCSFP150 and SubFP304) into compact, dense feature representations. Each layer implements a nonlinear transformation using the ReLU activation function, highlighting informative features while suppressing irrelevant noise, effectively identifying fingerprint locations associated with specific pharmacophores or structural motifs. Compared to directly using the original sparse input features, this method significantly improves training efficiency and accelerates the convergence process. Furthermore, when the attention mechanism highlights key descriptors such as molecular weight (MW) and polarity, FCN can further capture their higher-order interactions with other potential or weak features, thereby improving model prediction performance. Given that molecular sequences typically have hundreds of dimensions, FCN plays a crucial role in dimensionality reduction, generating unified, compact, and discriminative molecular representations suitable for downstream decision-making modules.
[0117] 3) Adaptive Feature Weighted Decision Tree. To improve the model's decision-making accuracy, we introduce a feature-weighted decision tree module in the output layer of the fully connected network. This module dynamically learns feature-specific weights and biases during training, transforming high-dimensional molecular representations into probabilistic routing decisions. The decision tree adaptively adjusts the feature contribution at each internal node, guiding samples along the optimized path. The hierarchical weighting mechanism refines feature importance at different levels of the tree, ultimately generating the final prediction at the leaf node through weighted aggregation of molecular features. This architecture supports both accurate classification of molecular datasets and interpretable analysis of feature importance.
[0118] like Figure 8 As shown in (b), this invention implements three key improvements over traditional HGNN: 1) Hierarchical Multi-Scale Hypergraph Neural Network. Traditional hypergraph neural networks typically construct hypergraphs at only a single scale (e.g., the atomic level), which limits their ability to represent multi-level features such as atoms, residues, and domains. Furthermore, the hyperedges in traditional hypergraphs are mainly based on general topological relationships (e.g., spatial distance), lacking biophysical constraints. To address these limitations, we propose an enhanced model—the hierarchical multi-scale hypergraph neural network. This model constructs hypergraphs at multiple scales and incorporates prior biological knowledge, achieving comprehensive modeling of the three-dimensional structure of olfactory receptors. We designed a three-layer hierarchical hypergraph structure: First, an atomic-level hypergraph, where atoms serve as nodes, and hyperedges are constructed based on chemical bonds, hydrogen bonds, hydrophobic interactions, etc., primarily predicting interaction energies at the atomic level; second, a residue-level hypergraph, where amino acid residues serve as nodes, constructing hyperedges based on spatial proximity to capture the spatial arrangement and functional synergistic relationships between residues, used for identifying functional sites and binding residues; and third, a domain-level hypergraph. This layer treats functional domains (e.g., transmembrane helices, loops, etc.) as nodes, constructing hyperedges based on spatial connections and functional coupling between domains. This hierarchy reflects the interactions between different functional modules of the receptor and infers the synergistic relationships between receptor structures. This hierarchical design enables the model to learn structural features at different scales, avoiding information loss that may occur with single-scale modeling. Finally, the model outputs a receptor-level representation vector by globally integrating features from all levels.
[0119] 2) Transformer-based receptor feature enhancement strategy. This strategy, building upon the HGNN-based extraction of the receptor's three-dimensional topology, further introduces a Transformer architecture to fully exploit the biological information contained within the receptor's amino acid sequence. This module treats the receptor sequence as a hierarchical semantic biological language, employing a multi-head self-attention mechanism to model the global interactions between all residues and dynamically calculating the association weights between any two pairs of amino acid residues. Through adaptive allocation of attention weights, this module effectively captures long-range dependencies between residues. The feature enhancement strategy enables the model to automatically identify key residue clusters—residues that, although discretely distributed in the receptor sequence, are spatially adjacent in the three-dimensional folded structure, collectively forming functional regions such as ligand binding sites. Through the stacking and integration of multiple attention layers, the model gradually focuses on key structural regions—active centers and allosteric regulatory sites. Furthermore, the enhanced Transformer mechanism simulates the complex residue interaction network within the receptor at the feature learning level, thereby achieving functional conformation reconstruction based on semantic information. Finally, the introduction of the Transformer module enhances the model's ability to represent the global sequence features of the receptor and improves its ability to capture cross-sequence synergistic effects and function-structure associations. This provides more distinctive and biologically interpretable feature representations for downstream tasks.
[0120] 3) Adaptive Cross-Modal Attention Fusion Module. The first two improvements extract receptor features from 3D geometric topology and 1D sequence semantics, respectively. However, simple feature concatenation cannot effectively bridge the semantic gap between 1D sequences and 3D structures. To address this multimodal heterogeneity issue, we propose an adaptive cross-modal attention fusion mechanism. This module comprises two components: First, a cross-modal interactive attention mechanism. A bidirectional cross-modal attention mechanism is introduced to establish a dynamic interaction relationship between structural features and sequence features. Specifically, the sequence features extracted by the Transformer are used as query vectors to focus on the structural feature key-value pairs extracted by the HGNN; conversely, structural features can also be used as query vectors to interact with sequence features. This mechanism promotes cross-modal semantic alignment. For example, when a sequence branch identifies a key residue, the attention mechanism automatically guides the model to focus on the local chemical environment of that residue (such as hydrogen bond networks and hydrophobic interactions) in the atomic-level hypergraph, thereby achieving a precise association between sequence features and structural constraints. Second, dynamic gating fusion. To address the uneven quality of receptor structural data, a dynamic gating network with feature confidence awareness is designed. This network adaptively adjusts the fusion weights of multi-source information based on the inherent uncertainties of each input modality. For example, when noise is detected in structural features, the gating unit automatically reduces the contribution weight of structural branches while enhancing the influence of more reliable sequence features. This design significantly improves the model's robustness in scenarios with imbalanced data quality. Through this mechanism, the model ultimately outputs a highly complete receptor representation, providing more robust and interpretable multimodal features for downstream tasks.
[0121] like Figure 8 As shown in (c), based on the improved DNDF for extracting molecular structure features and the enhanced HGNN for extracting receptor structure features, we further designed a fully connected network module to effectively fuse bimodal features, thereby predicting the molecule-receptor response curve. To improve the performance of this module, we introduced a cross-modal attention mechanism and a residual connection feature fusion strategy, the details of which are as follows: 1) Cross-modal attention fusion. To accurately model the interaction between molecules and receptors, we introduce a feature fusion mechanism that integrates cross-modal attention into a fully connected network. This module first models the molecular structural features extracted by Deep Neural Decision Forest (DNDF) and the 3D topological features of the receptor extracted by Hypergraph Neural Network (HGNN) based on a bidirectional attention mechanism. Using molecular features as query vectors and receptor features as key-value pairs, the module calculates attention weights to focus the model on key binding regions in the receptor structure (such as key residue clusters within the active site). Simultaneously, receptor features can also be used as query vectors to focus on key functional groups or atomic arrangements in the molecule, thus establishing a bidirectional cross-modal perception and alignment mechanism. This mechanism strengthens the semantic association between molecules and receptors in the feature space, enabling the model to adaptively identify structural features closely related to binding affinity, such as hydrogen bonds of specific amino acid residues or aromatic ring characteristics within the molecule. Through attention weight integration, this module reveals the key atomic and residue interaction patterns in the molecule-receptor binding process, providing an interpretable structural basis for model prediction.
[0122] 2) Residual Connections and Gated Feature Fusion. To further enhance the depth and stability of feature extraction, we introduce a residual connection and gated multi-scale fusion mechanism into the fully connected network. This module consists of stacked layers of residual modules, each containing two fully connected layers with skip connections, effectively alleviating the gradient vanishing problem in deep network training. This enables the model to progressively learn high-order interaction features, from atomic-level physicochemical compatibility (such as charge complementarity and hydrophobic effects) to residue-level conformational synergies. Simultaneously, the network employs a gating mechanism to adaptively fuse feature representations at different abstraction levels: lower-level features capture fine-grained local interactions between atoms, while higher-level features encode the global relationship between the overall receptor conformation and molecular binding patterns, thus achieving multi-granularity modeling of molecular-receptor interactions. Finally, the constrained output layer maps the fused features to dynamic parameters (amplitude, damping coefficient, frequency, and phase) of the response curve. By integrating physiological constraints (such as firing frequency limits and decay timescales), we ensure that the prediction results conform to the biological rationale of neural responses. This design significantly enhances the model's ability to represent complex interaction patterns while improving the efficiency and generalization performance of cross-receptor and molecular training.
[0123] This invention conducted a series of comprehensive experiments to evaluate the performance of the proposed method. First, the robustness of the neural response curves was verified by comparing the predicted molecular-receptor response curves with actually recorded curves to verify the accuracy of the response curve modeling. Specifically: To verify the accuracy and biological plausibility of the predicted molecular-receptor neural response curves, this invention collected real neural response curve data induced by multi-molecule mixtures from existing experiments. These mixtures consist of multiple individual molecules with significantly different concentration ratios, thus reflecting the complexity of the real olfactory stimulation environment. Finally, this invention constructed a dataset containing 35 pairs of mixture-response curves for external validation.
[0124] This invention uses curve similarity to assess the closeness between the actual neural response curve and the predicted curve. Curve similarity is defined as the similarity of the response values of two curves at corresponding time points. Experimental results show that the matching score of the method of this invention reaches 0.925, indicating that the deviation between the two is minimal and that they have a high degree of consistency. Furthermore, the method of this invention was compared with other models, including traditional models (RF, CNN) (HA Salman, A. Kalakech, A. Steiti. Randomforest algorithm overview [J]. Babylonian Journal of Machine Learning, 2024, 2024: 69-79.; Z. Li, F. Liu, W. Yang, S. Peng, J. Zhou. A survey of convolutional neural networks: analysis, applications, and prospects [J]. IEEE transactions on neural networks and learning systems, 2021, 33(12):6999-7019.) and benchmark models (DKNN, GNN) (N. Papernot, P. Mcdaniel. Deep k-nearestneighbors: Towards confident, interpretable and robust deep learning [J]. arXiv preprint arXiv:180304765, 2018. Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, S. Xie. A convnet for the 2020s; proceedings of the Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, F, 2022 [C].) and advanced pre-trained molecular models (POM, MolFormer) (J. Ross, B. Belgodere, V. Chenthamarakshan, I. Padhi, Y. Mroueh, P. Das.Large-scalechemical language representations capture molecular structure and properties[J]. Nature Machine Intelligence, 2022, 4(12): 1256-64.;BK Lee, EJMayhew, B. Sanchez-Lengeling, JN Wei, WW Qian, KA Little, M. Andres,BB Nguyen, T. Moloy, J. Yasonik. A principal odor map unifies diversetasks in olfactory perception [J]. Science, 2023, 381(6661): 999-1006.). . Figure 9 Figure (a) shows the response curves predicted by each model, and the results show that the method of the present invention best matches the true curve. Figure 9 As shown in (b), the curve similarities for RF and CNN are 0.712 and 0.732, respectively; the curve similarities for DKNN and GNN are 0.765 and 0.802, respectively; and the curve similarities for POM and MolFormer are 0.835 and 0.864, respectively. Furthermore, this invention also evaluates the alignment of response peaks to measure the performance of the method, i.e., the degree of matching between the predicted response peak and the actual response peak. The method of this invention achieves an accuracy of 0.945, while other methods all yield results worse than this method. This demonstrates that the method of this invention can accurately predict the peak value of the curve. Figure 9 Figure (c) shows the matching accuracy of each of the 35 mixture cases obtained by different methods, and the method of this invention achieved the highest performance in all cases. Compared with the other six models, the method of this invention significantly outperforms them in response curve prediction accuracy. Experimental results show that the method of this invention can accurately reconstruct neural response curves under complex and diverse real-world mixture conditions, and maintains stable and superior performance on external datasets, thus demonstrating the robustness, practicality, and strong generalization ability of this method.
[0125] Secondly, the alignment performance between the response curves and odor perception was examined. This invention tested whether the differences between neural response curves were consistent with the differences in odor perception, thereby verifying the proposed isomorphic relationship. Specifically, as follows: This invention verifies the isomorphic relationship between neural response curves and differences in odor perception. Specifically, this means that when the perceptual differences between molecules are large, the corresponding neural response curves also differ significantly, and vice versa. Using a dataset containing 5,023 molecules, this invention constructed all permutations of molecular pairs, generating approximately 12.6 million sample pairs. For each pair of molecules, the neural response curves induced by these molecules were predicted using the method of this invention. Then, this invention calculated pairwise curve similarity and pairwise perceptual similarity. Furthermore, this invention used the Pearson correlation coefficient r to quantify the consistency between these two indicators. The final Pearson correlation coefficient was r = 0.921. This significant positive correlation indicates that the model of this invention successfully captures the topological structure of the perceptual space: the greater the perceptual differences between molecules, the greater the differences in the neural response curves predicted by the model. Figure 10 As shown in (a), the Pearson density plot demonstrates a linear positive correlation and isomorphism between the response curves and differences in odor perception. Furthermore, this invention evaluates the performance of other methods in assessing the isomorphism between curves and perceived similarity. The obtained Pearson correlation coefficients are: RF (0.521), CNN (0.558), DKNN (0.635), GNN (0.721), POM (0.815), and MolFormer (0.821), all of which are significantly inferior to the performance of the method of this invention. Therefore, the method of this invention demonstrates superior discriminative power, effectively aligning differences in molecular response curves with differences in odor perception.
[0126] exist Figure 10 In (b) of this invention, 70 of the most challenging and representative molecules were selected, capable of capturing the key features of the molecular dataset. The selection criteria were as follows: functional group diversity—the selected molecules covered common functional groups related to perception, including carboxyl, hydroxyl, amino, and ether groups; chemical category diversity—the selected molecules covered major chemical categories, including hydrocarbons, alcohols, esters, and aromatic compounds; and molecular structure diversity—the selected molecules varied in molecular weight (from simple low-molecular-weight compounds to complex high-molecular-weight compounds) and shape (straight-chain, cyclic, and branched). The method of this invention achieved optimal Pearson correlation coefficients on these molecules, with all values remaining above 0.8, indicating a high degree of consistency between the predicted response curves and their corresponding odor perceptions. In contrast, the performance of the other six comparative methods was significantly inferior. Figure 10In section (c), this invention further analyzes the effect of the number of odor perceptions corresponding to these representative molecules on the Pearson coefficient. The results show that as the number of odor perceptions increases from 1 to 15, the Pearson coefficient obtained by the method of this invention remains high, demonstrating good stability. This indicates that even if the odor perceptions corresponding to a molecule are complex and diverse, the method of this invention can still accurately predict the response curve that the molecule can excite. Figure 10 The gray dots represent the average results of the other six models, whose Pearson coefficient gradually decreases as the number of odor perceptions increases, indicating poor predictive performance.
[0127] Finally, the similarity and recognition of odor perception in the mixture are evaluated to further demonstrate the superior performance of the proposed method. Specifically, the method is compared with existing mixture perception similarity recognition methods and other deep learning models for mixture odor perception recognition, as follows: To evaluate the ability of the method of the present invention to identify the odor perception of real-world multi-molecular mixtures, the present invention compiled and constructed a real dataset containing 360 mixtures. Figure 2(c) This dataset provides detailed annotations of individual molecular components in each mixture, as well as true perceived similarity scores between each pair of mixtures obtained through human sensory testing. The method of this invention identifies the odor perception of each mixture, and quantifies the similarity of odor perception between mixtures by calculating cosine similarity. Finally, this invention applies the Pearson correlation coefficient to evaluate whether the calculated odor perception similarity matches the true human perception similarity. Furthermore, this invention uses the mean absolute error (MAE) to evaluate the absolute deviation between the two similarities, i.e., the difference between the perceived similarity between mixtures calculated by the method of this invention and the actual perceived similarity. The MAE value ranges from [0, 1], with a value closer to 0 indicating a smaller deviation and better model performance. Furthermore, the method of this invention was compared with other existing recognition models, namely those of Amit et al. (A. Dhurandhar, H. Li, GA Cecchi, P. Meyer. Expansive linguistic representations to predict interpretable odor mixture discriminability [J]. Chemical senses, 2023, 48:bjad018.) and Ravia et al. (A. Ravia, K. Snitz, D. Honigstein, M. Finkel, R. Zirler, O. Perl, L. Secundo, C. Laudamiel, D. Harel, N. Sobel. A measure of smell enables the creation of olfactory metamers [J]. Nature, 2020, 588(7836): 118-23.). Both of these models are specifically designed for recognizing the perceptual similarity between mixtures, and the interpretable metric learning framework proposed by Amit et al. is based on semantic descriptors. This method constructs a structure-to-perception model, predicting intensity, pleasantness, and semantic descriptors through molecular structure, and obtaining the perceptual representation of mixtures through averaging. This method achieves excellent root mean square error on multiple datasets. Ravia et al. represented each multi-molecule mixture as a vector composed of physicochemical descriptors and quantified odor distance by calculating the angular distance between vectors. This model incorporates perceptual intensity weights during vector aggregation by converting single-molecule intensity scores into weights using a sigmoid function, thus reflecting the concentration contribution. This method demonstrates strong predictive consistency on multiple datasets.
[0128] The Pearson correlation coefficient between the perceived similarity and actual similarity of odors in mixtures calculated by the method of this invention is r = 0.931. The result obtained using the method of Amit et al. is 0.811, and the result obtained using the method of Ravia et al. is 0.834, both of which are significantly inferior to the method of this invention. The distribution of perceived similarity for 360 mixtures is as follows: Figure 11 As shown in (a) of the figure, the method of the present invention achieves better correlation. Furthermore, the MAE obtained by the method of the present invention (… Figure 11 In (b) of the above, the MAE is 0.072, while the MAE obtained by the methods of Amit et al. and Ravia et al. are 0.162 and 0.134, respectively. The method of the present invention produces a lower MAE, indicating higher accuracy. Simultaneously, the method of the present invention achieves a Spearman correlation coefficient of 0.962 (an improvement of 0.211 and 0.200 compared to Amit et al. and Ravia et al., respectively), an RMSE of 0.084 (an improvement of 0.089 and 0.081, respectively), an F1 score of 0.945 (an improvement of 0.141 and 0.110, respectively), and an R-squared value of 0.942. 2 (Increased by 0.166 and 0.127 respectively). These results demonstrate that the method of this invention is a significant improvement over the methods of Amit et al. and Ravia et al. Furthermore, the method of this invention was compared with advanced models POM and MolFormer, which produced Pearson correlation coefficients of 0.721 and 0.755, and MAE values of 0.415 and 0.381, respectively—both significantly inferior to the method of this invention. Figure 11 (c)). Therefore, the method of the present invention accurately aligns odor perception similarity with actual similarity. This high correlation further demonstrates that the method of the present invention can accurately distinguish odor perception differences between mixtures. Figure 11 In section (d), this invention selected the 16 most frequently occurring odors in the dataset, presented the perceptual distributions of multi-molecular mixtures on these odors predicted by different methods, and compared them with the actual odor perceptual distributions. The perceptual distributions obtained by the method of this invention are in best agreement with the actual results, while the results obtained by other methods all have significant deviations.
[0129] To evaluate the robustness of the method of this invention in identifying specific odor perceptions of real-world mixtures, this invention collected a dataset containing odor perception data for 172 real-world mixtures. This dataset is extremely challenging because these mixtures have complex chemical compositions, including diverse single molecular components and precise concentration ratios. Figure 2(d) In this invention, cosine similarity is used as a quantitative indicator to calculate the similarity between predicted odor perception and real human odor perception.
[0130] Furthermore, the method of the present invention was compared with traditional classification models (RF and CNN), benchmark models (DKNN and GNN), and advanced pre-trained molecular models (POM and MolFormer).
[0131] First, this invention evaluated the cosine similarity between the odor perception of mixtures calculated by different models and the actual odor perception. The cosine similarity obtained by the method of this invention is 0.942, while RF and CNN are 0.657 and 0.661, respectively. The results of DKNN and GNN are 0.742 and 0.774, respectively, which are significantly worse than the method of this invention. The results of POM and MolFormer are 0.814 and 0.832, respectively. Therefore, the evaluation based on cosine similarity shows that the method of this invention is significantly superior to other models and can accurately identify the odor perception corresponding to multi-molecular mixtures.
[0132] Table 1 details the results of the six evaluation indicators. Figure 12 (a) further illustrates the performance distribution of each model across various metrics using box plots, while Figure 12 Figure (b) shows the ROC curves of different models and the method of this invention. Experimental results demonstrate the superior performance of the method of this invention, with the following specific metrics: AUROC = 0.924 ± 0.031, AUPRC = 0.919 ± 0.018, precision = 0.932 ± 0.028, recall = 0.916 ± 0.052, specificity = 0.915 ± 0.021, and accuracy = 0.922 ± 0.016. Notably, the small standard deviations of these metrics provide strong evidence for the superior stability and generalization ability of the method of this invention in the odor perception prediction task. Comparison with other models further reinforces the advantages of the method of this invention. First, compared with traditional models, the accuracy of RF and CNN is only 0.665 and 0.684, respectively, which is significantly lower than the method of this invention. Second, compared with benchmark models (DKNN and GNN), the method of this invention leads in all six evaluation metrics. Specifically, in terms of the AUROC metric, the method of this invention improves upon DKNN (0.745) by 0.179 and GNN (0.835) by 0.089; in terms of the AUPRC metric, the method of this invention improves upon DKNN (0.751) by 0.168 and GNN (0.724) by 0.195; and in terms of accuracy, the method of this invention is 0.197 and 0.158 higher than DKNN (0.725) and GNN (0.764), respectively.
[0133] Table 1 Performance comparison of different methods
[0134] Subsequently, the method of this invention was compared with advanced models POM and MolFormer. Although POM achieved an accuracy of 0.844 AUROC, 0.836 AUPRC, and 0.802, the method of this invention significantly outperformed these results on all metrics, improving accuracy by 0.080, 0.093, and 0.120, respectively. Compared to MolFormer, the method of this invention also achieved performance improvements of 0.073, 0.067, and 0.096 on the same three metrics. Notably, the method of this invention performed particularly well in terms of recall and specificity: recall improved from 0.847 for POM and 0.875 for MolFormer to 0.916, while specificity significantly improved from 0.834 and 0.872 to 0.915, respectively. Therefore, the method of this invention can accurately predict the odor perception of multi-molecular mixtures, demonstrating its strong stability. Figure 12 In (c), the present invention evaluated and ranked the accuracy of 16 of the most common odor perceptions. The results showed that all accuracy values were above 0.9, indicating that the method of the present invention can reliably identify these odor perceptions corresponding to the mixture. Figure 12 Figure (d) shows the connection strength between these most common odors and their heatmap distribution with other odors, where the connection strength is calculated based on the number of molecules shared between odors. The results show that these most common odors have significant correlations with other odors, with a large number of shared molecules. This demonstrates that predicting frequent odors in a mixture is an extremely challenging task, yet the method of this invention maintains high accuracy.
[0135] Experimental results demonstrate that the method of this invention achieves a significant improvement in performance metrics, reaching an accuracy of 92.2% on a newly compiled real-world dataset. The method of this invention can accurately identify the odor perception of mixtures. Therefore, this work elucidates the formation process of olfactory perception in mixtures, providing a computationally feasible solution to the long-standing challenge in olfactory science of moving from chemical mixing to sensory emergence.
[0136] This invention addresses the long-standing challenge of decoding the olfaction of multi-molecular mixtures—a fundamental problem in human sensory science—by developing a novel method for recognizing the odor perception of multi-molecular mixtures through a biologically based olfactory coding model. First, this invention designs a deep learning network to characterize the topological features of molecules and olfactory receptors, and utilizes a biomimetic mechanism to construct neural response curves representing their interactions. Furthermore, this invention proposes an attention-weighted multi-receptor response curve fusion strategy, combined with a concentration-perceived mixture response curve fusion scheme, to reflect the influence of different receptor and molecule concentrations on odor perception. Simultaneously, addressing the scarcity of real-world data on the association between multi-molecular mixtures and odor perception, this invention designs an unsupervised pre-training paradigm based on contrastive learning. Finally, a transfer learning strategy is employed, utilizing abundant single-molecule perception data as prior knowledge to guide the inference of complex mixture perception. To verify the performance of the proposed method, this invention collects a dataset of real-world multi-molecular mixtures containing different molecule and concentration ratios, and verifies the robustness of the method by comparing it with traditional models, benchmark models, and state-of-the-art models. Experimental results show that the method of this invention achieves an accuracy of 92.2%, accurately recognizing the odor perception of multi-molecular mixtures. Therefore, the method proposed in this invention provides a high-precision computational framework for olfactory perception and recognition, and establishes an interpretable pathway from chemical stimulation to neural coding, and ultimately to perception formation. This work lays the technical foundation for the digital modeling of the olfactory system and has broad application prospects in future embodied intelligence scenarios.
[0137] Corresponding to the aforementioned embodiment of a hybrid multi-molecule intelligent perception and analytical deep learning method based on high-dimensional interaction association modeling, the present invention also provides an embodiment of a hybrid multi-molecule intelligent perception and analytical deep learning device based on high-dimensional interaction association modeling.
[0138] See Figure 13 The present invention provides a hybrid multi-molecule intelligent perception and analysis deep learning device based on high-dimensional interaction association modeling, comprising a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a hybrid multi-molecule intelligent perception and analysis deep learning method based on high-dimensional interaction association modeling as described in the above embodiment.
[0139] The embodiment of the hybrid multi-molecule intelligent sensing and analytical deep learning device based on high-dimensional interactive correlation modeling provided by this invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 13 The diagram shown is a hardware structure diagram of any device with data processing capabilities, which is a hybrid multi-molecule intelligent sensing and analytical deep learning device based on high-dimensional interactive correlation modeling provided by the present invention. (Except for...) Figure 13 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0140] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0141] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0142] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a hybrid multi-molecule intelligent perception and analytical deep learning method based on high-dimensional interactive association modeling as described in the above embodiments.
[0143] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0144] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned hybrid multi-molecule intelligent perception and analytical deep learning method based on high-dimensional interactive association modeling.
[0145] The above embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
Claims
1. A hybrid multi-molecular intelligent sensing analytical deep learning method based on high-dimensional interaction correlation modeling, characterized in that, The method includes: (1) Deep neural decision forest is used to extract odor molecule features, and hypergraph neural network is used to extract olfactory receptor features. The two features are fused and input into a fully connected network to predict the decay sinusoidal response curve of a single odor molecule-single receptor. (2) For multiple receptors, the decay sinusoidal response curves of multiple corresponding single odor molecules-single receptors are obtained by using the method of step (1), and weighted fusion is performed to obtain the fusion response curve of single odor molecules-multiple receptors; based on the principle of odor perception difference, the fusion response curve between each pair of odor molecules is aligned with the odor perception difference to optimize the model, and the alignment deviation is used as the loss function to guide the generation of odor molecule-receptor response curves for training, thereby obtaining the mapping relationship between odor perception and fusion response curves, and obtaining the weight parameters of each receptor relative to a single odor molecule; (3) Based on the concentration weighting of odor molecules, the fusion response curve of multiple odor molecules-single receptor is obtained, and then multiple corresponding fusion response curves of multiple odor molecules-single receptor are obtained for multiple receptors. The fusion is performed based on the weight parameters of each receptor relative to the single odor molecule to obtain the fusion response curve of multiple odor molecules-multiple receptors. (4) Find the fusion response curves of multiple odor molecules and multiple receptors that are similar to the fusion response curves of multiple odor molecules and multiple receptors obtained in step (2), obtain the corresponding similarity weights as the odor perception weights of single odor molecules, perform weighted fusion of odor perception of multiple single odor molecules, obtain the odor probability distribution of the multi-molecule mixture, and obtain the set of perceived odors of the multi-molecule mixture based on threshold screening. 2.The hybrid multi-molecular intelligent sensing analytical deep learning method based on high-dimensional interaction correlation modeling according to claim 1, wherein, In step (1), a decaying sine wave model is used to construct a neural response curve to simulate the characteristic temporal discharge dynamics of olfactory receptors after odor stimulation.
3. The hybrid multi-molecule intelligent perception analytical deep learning method based on high-dimensional interactive association modeling according to claim 1, characterized in that, In step (2), the differences in odor perception between different odor molecules are highly consistent with the differences in their neural response curves. The greater the difference in perception, the more significant the deviation of the response curve usually is. Single odor molecule odor labels are used as discrimination indicators to align and minimize the deviation between the response curve and odor perception. For each pair of odor molecules, the difference in response curves is calculated by the root mean square error, and the difference in odor perception is measured by the cosine distance.
4. The hybrid multi-molecule intelligent perception analytical deep learning method based on high-dimensional interaction association modeling according to claim 1, characterized in that, In step (3), the decay sinusoidal response curves of a single odor molecule and a single olfactory receptor are predicted based on a deep learning model. Then, the concentration is used as the weight of the decay sinusoidal response curve of the odor molecule to measure its competitive binding advantage at the single receptor binding site. Finally, the weighted response curves of all odor molecules are fused by linear superposition in the time domain: at each time point, the response values of all odor molecules are summed and averaged according to the concentration weight to generate the fusion response curve of multiple odor molecules-single receptor.
5. The hybrid multi-molecule intelligent perception analytical deep learning method based on high-dimensional interactive association modeling according to claim 1, characterized in that, In step (3), for the fusion response curves induced by odor molecules of different concentrations in different receptor spectrums, the fusion response curves are weighted and averaged according to the corresponding attention weights, and the response value at each time point is calculated, and finally the fusion response curves of multiple odor molecules-multiple receptors are generated.
6. The deep learning method for decoding olfactory perception of multi-molecular mixtures based on receptor interaction modeling according to claim 5, characterized in that, For each olfactory receptor, the average weight of each single odor molecule in its odor perception is taken to obtain the weight of the olfactory receptor's contribution to odor perception, which is used as the weight parameter of the receptor relative to the single odor molecule.
7. The hybrid multi-molecule intelligent perception analytical deep learning method based on high-dimensional interactive association modeling according to claim 1, characterized in that, In step (4), by calculating the similarity index between the fusion response curve of multiple odor molecules and multiple receptors and the fusion response curve of single odor molecules and multiple receptors, the odor perception vectors associated with different odor molecules are adaptively weighted to reflect their differences in influence on the overall odor perception; after the weighted odor vectors are aggregated, the final perception output is determined by threshold screening.
8. A hybrid multi-molecule intelligent sensing and analytical deep learning device based on high-dimensional interactive correlation modeling, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the processor executes the executable code, it implements a hybrid multi-molecule intelligent perception parsing deep learning method based on high-dimensional interactive association modeling as described in any one of claims 1-7.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a hybrid multi-molecule intelligent perception and parsing deep learning method based on high-dimensional interactive association modeling as described in any one of claims 1-8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a hybrid multi-molecule intelligent perception and parsing deep learning method based on high-dimensional interactive association modeling as described in any one of claims 1-9.