Aroma prediction method and system based on olfactory receptor and dynamic simulation
By integrating a public olfactory receptor database with biomimetic reactor simulation environmental parameters, an aroma contribution prediction model is generated using an aroma prediction method based on olfactory receptors and dynamic simulation. This solves the problems of deviation between prediction results and real olfactory experience and low efficiency in traditional aroma analysis techniques, and achieves efficient and accurate aroma analysis.
Patent Information
- Application Number
- CN202511036112.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-26
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional aroma analysis techniques suffer from several drawbacks: a disconnect between chemical instrument detection and human sensory evaluation; static extraction techniques cannot capture dynamic release processes; and the high cost of repetitive sensory experiments leads to low analytical efficiency, making it difficult to meet the high-throughput demands of industrial applications.
The aroma prediction method based on olfactory receptors and dynamic simulation integrates data through a public olfactory receptor database, uses a biomimetic reactor to simulate environmental parameters, generates dynamic release data, trains an aroma contribution prediction model using a machine learning model, and conducts aroma reconstruction experiments and sensory evaluations to verify the accuracy of the prediction.
It improves the consistency between aroma prediction and real olfactory experience, reduces the correlation between dynamic release process and sensory intensity, improves analysis efficiency, and meets the high-throughput requirements of industrial applications.
Smart Images

Figure CN120929752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of food flavor analysis and intelligent sensing technology, and in particular to an aroma prediction method and system based on olfactory receptors and dynamic simulation. Background Technology
[0002] With the rapid development of the food industry and flavor science, accurately analyzing the sensory contribution mechanisms of aroma compounds has become a core requirement for improving product quality. Especially in food processing, flavor design, and environmental odor monitoring, efficient prediction of key aroma compounds is crucial for optimizing product flavor experiences.
[0003] However, traditional aroma analysis techniques have the following problems: chemical instrument detection (such as GC-MS) and human sensory evaluation are disconnected from the receptor physiological response mechanism, resulting in a significant deviation between the predicted results and the actual olfactory experience; static extraction techniques (such as headspace solid-phase microextraction) cannot capture the dynamic release process in real scenarios, causing a misalignment between the aroma release peak and sensory intensity; repetitive sensory experiments and reliance on high-cost equipment lead to low analysis efficiency, making it difficult to meet the high-throughput requirements of industrialization. Summary of the Invention
[0004] Therefore, it is necessary to provide aroma prediction methods and systems based on olfactory receptors and dynamic simulation to address the aforementioned technical problems, so as to improve the consistency between aroma prediction and real olfactory experience, reduce the correlation deviation between dynamic release process and sensory intensity, and improve analysis efficiency to meet the high-throughput requirements of industrialization.
[0005] In a first aspect, this application provides an aroma prediction method based on olfactory receptors and dynamic simulation, the method comprising:
[0006] Based on a public olfactory receptor database, ligand-specific data of various human olfactory receptors and binding experimental data of target aroma substances are integrated and processed to generate a binding characteristic database.
[0007] By simulating the environmental parameters of the target scenario using a biomimetic reactor, dynamic release data is obtained.
[0008] Based on the combination of characteristic database and dynamic release data, a machine learning model is trained to generate an aroma contribution prediction model.
[0009] Based on the ranking of key aroma substances output by the aroma contribution prediction model, aroma reconstruction experiments and sensory evaluation were conducted to verify the accuracy of the prediction.
[0010] Furthermore, by simulating the environmental parameters of the target scenario using a biomimetic reactor, dynamic release data is obtained, including:
[0011] The oral processing environment parameters of the target scenario are simulated by a biomimetic oral reactor to generate simulated environmental data. The oral processing environment parameters include chewing frequency, saliva secretion rate, temperature and pH value.
[0012] Based on simulated environmental data, dynamic release simulation processing of target samples is performed to obtain volatile release process data;
[0013] Based on the data of volatile release process, volatile substances are collected and processed at multiple time points using online volatile capture technology to generate time-series aroma substance concentration data;
[0014] Based on the time-series aroma substance concentration data, aroma feature annotation is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive word data;
[0015] Based on time-series data of aroma substance concentration and aroma characteristic descriptor data, dynamic release data is generated by integrating and processing the data.
[0016] Furthermore, based on time-series aroma substance concentration data and aroma characteristic descriptor data, data integration and processing are performed to generate dynamic release data, including:
[0017] Time-point alignment processing is performed on the time series aroma substance concentration data and aroma feature descriptor data to generate time-aligned data;
[0018] Based on time-aligned data, aroma release pattern recognition processing is performed to generate dynamic release pattern data.
[0019] Aroma perception encoding conversion processing is performed on the dynamic release mode data to generate dynamic release data.
[0020] Furthermore, based on the time-series aroma substance concentration data, aroma feature annotation is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive term data, including:
[0021] Based on the time-series aroma substance concentration data at each time point, sensory evaluation technology is used to process volatile substances through olfaction and generate sensory response data.
[0022] Aroma feature descriptor mapping is performed on sensory response data to generate aroma feature descriptor data;
[0023] The aroma feature descriptor data is subjected to key aroma contribution quantification processing to generate key aroma contribution data;
[0024] Aroma feature descriptor data and key aroma contribution data are integrated to generate aroma feature descriptor data.
[0025] Furthermore, based on a public olfactory receptor database, ligand-specific data of various human olfactory receptors and binding experimental data of target aroma substances are integrated and processed to generate a binding characteristic database, including:
[0026] Based on a public olfactory receptor database, protein sequences of various human olfactory receptors were acquired and processed to obtain protein sequence data;
[0027] A three-dimensional structural model of the olfactory receptor is generated by performing structural prediction processing on protein sequence data using homology modeling methods.
[0028] The target aroma substance is processed to obtain its three-dimensional structure data.
[0029] Molecular docking simulation was performed on the three-dimensional structural models of olfactory receptors and the three-dimensional structural data of aroma substances to generate binding parameter data.
[0030] Integrate ligand-specific data with binding parameter data to generate a binding characteristic database.
[0031] Furthermore, based on the ranking of key aroma compounds output by the aroma contribution prediction model, aroma reconstruction experiments and sensory evaluation were conducted to verify the prediction accuracy, including:
[0032] Key aroma compounds are screened based on the contribution ranking output by the aroma contribution prediction model.
[0033] Key aroma compounds were mixed in the predicted proportions to prepare aroma reconstruction samples;
[0034] Sensory evaluation technology was used to compare and evaluate the aroma reconstruction samples with the original samples to generate sensory score data.
[0035] Based on sensory rating data, a significance test was performed to generate the test results;
[0036] If the test results meet the set threshold, the prediction is considered valid.
[0037] Furthermore, based on the combination of a feature database and dynamic release data, a machine learning model is trained to generate an aroma contribution prediction model, including:
[0038] Based on the binding characteristic database, olfactory receptor binding parameter data were extracted;
[0039] Based on dynamic release data, time series data and substance concentration data are extracted;
[0040] Feature fusion processing is performed on olfactory receptor binding parameter data, time series data, and substance concentration data to generate multidimensional input features;
[0041] Based on sensory evaluation data, obtain key aroma contribution labels;
[0042] An ensemble learning algorithm is used to train the model on multidimensional input features and key aroma contribution labels to generate an initial aroma contribution prediction model.
[0043] The initial aroma contribution prediction model was optimized using cross-validation techniques to generate an optimized aroma contribution prediction model.
[0044] Secondly, this application also provides an aroma prediction system based on olfactory receptors and dynamic simulation, the system comprising:
[0045] The database construction module is used to integrate and process ligand-specific data of various human olfactory receptors with experimental data on the binding of target aroma substances based on a public olfactory receptor database, and generate a binding characteristic database.
[0046] The dynamic simulation module is used to simulate the environmental parameters of the target scene using a biomimetic reactor to obtain dynamic release data;
[0047] The intelligent prediction module is used to train machine learning models based on the combined characteristic database and dynamic release data, and generate aroma contribution prediction models.
[0048] The closed-loop verification module is used to rank the key aroma substances based on the output of the aroma contribution prediction model, conduct aroma reconstruction experiments and sensory evaluation processing, and verify the accuracy of the prediction.
[0049] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods in the first aspect of this application.
[0050] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods in the first aspect of this application.
[0051] This application provides an aroma prediction method and system based on olfactory receptors and dynamic simulation. The method includes: integrating ligand-specific data of various human olfactory receptors with binding experimental data of target aroma substances based on a public olfactory receptor database to generate a binding characteristic database; simulating environmental parameters of the target scene using a biomimetic reactor to obtain dynamic release data; training a machine learning model based on the binding characteristic database and dynamic release data to generate an aroma contribution prediction model; and conducting aroma reconstruction experiments and sensory evaluation based on the ranking of key aroma substances output by the aroma contribution prediction model to verify the prediction accuracy. The aim is to improve the consistency between aroma prediction and real olfactory experience, reduce the correlation deviation between dynamic release process and sensory intensity, and improve analysis efficiency to meet the high-throughput requirements of industrial applications. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 This is a flowchart of an aroma prediction method based on olfactory receptors and dynamic simulation in one embodiment of the present invention;
[0054] Figure 2 This is a flowchart of generating aroma feature descriptive word data by simultaneously performing aroma feature annotation processing through sensory evaluation technology based on time-series aroma substance concentration data in one embodiment of the present invention.
[0055] Figure 3 This is a structural diagram of an aroma prediction system based on olfactory receptors and dynamic simulation in one embodiment of the present invention. Detailed Implementation
[0056] To make the above-mentioned objects, features, and advantages of this application more apparent and understandable, the specific implementation methods of this application will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0057] like Figure 1 As shown, this application provides an aroma prediction method based on olfactory receptors and dynamic simulation, the method comprising:
[0058] S101: Based on a public olfactory receptor database, ligand-specific data of various human olfactory receptors are integrated and processed with experimental data on the binding of target aroma substances to generate a binding characteristic database.
[0059] Specifically, ligand-specific data for various human olfactory receptors were obtained from a public olfactory receptor database. This data included specific information on the binding of olfactory receptors to different ligands, providing fundamental olfactory receptor characteristic information for subsequent analysis. Simultaneously, binding experiments were conducted on target aroma compounds to collect corresponding experimental data, which reflect the actual binding of the target aroma compounds to olfactory receptors.
[0060] Next, the acquired ligand-specific data of various human olfactory receptors were integrated with the binding experimental data of the target aroma substances. This process included data format standardization, data cleaning, and screening to ensure data quality and consistency. Subsequently, a binding characteristic database was generated, containing binding characteristic information of various human olfactory receptors and target aroma substances, providing crucial data support and basis for subsequent aroma prediction and analysis steps.
[0061] S102: The environmental parameters of the target scene are simulated and processed by a biomimetic reactor to obtain dynamic release data.
[0062] Specifically, the environmental parameters of the target scenario are clearly defined, including chewing frequency, saliva secretion rate, temperature, and pH value, which significantly influence the release of aroma compounds. A biomimetic reactor is then used to simulate these environmental parameters. This reactor accurately reproduces the oral processing environment of the target scenario, making the experimental conditions closer to reality. Under this simulated environment, the target sample undergoes dynamic release simulation treatment, capturing the release of aroma compounds at different time points. Simultaneously, volatile substances are collected at multiple time points using online volatile matter capture technology. This technology efficiently collects released aroma compounds, providing samples for subsequent analysis.
[0063] The collected samples were then analyzed to obtain time-series aroma compound concentration data, reflecting the concentration changes of aroma compounds during the dynamic process. To more comprehensively describe the aroma characteristics, aroma features were simultaneously labeled using sensory evaluation techniques based on the time-series aroma compound concentration data, resulting in aroma feature descriptive term data. This process combines chemical analysis results with human sensory experience, enriching our understanding of aroma compounds. Subsequently, the time-series aroma compound concentration data and aroma feature descriptive term data were integrated to generate dynamic release data. This data integrates chemical and sensory information, providing a more comprehensive and accurate input for subsequent aroma prediction model training, thus helping to improve the model's prediction accuracy and reliability.
[0064] S103: Based on the combination of characteristic database and dynamic release data, machine learning model training is performed to generate an aroma contribution prediction model.
[0065] Specifically, parameters related to the binding of olfactory receptors and aroma compounds are extracted from the binding characteristic database. These parameters reflect the binding characteristics between the two and provide key feature data for model training. Simultaneously, time-series information and substance concentration information are extracted from dynamic release data. The time-series information reflects the release of aroma compounds at different time points, while the substance concentration information reflects the concentration changes of aroma compounds. The extracted olfactory receptor binding parameters, time-series information, and substance concentration information are then fused using relevant algorithms and methods. This integrates these different types of feature data to generate multi-dimensional input features that comprehensively describe the characteristics of aroma compounds, enabling the model to learn and understand the characteristics of aroma compounds from multiple perspectives.
[0066] Based on sensory evaluation data, key aroma contribution labels are obtained. These labels reflect the importance of aroma substances in actual sensory experience, providing target guidance for model training. An ensemble learning algorithm is used, with multi-dimensional input features and key aroma contribution labels as input and output, to train the model. The ensemble learning algorithm improves the model's generalization ability and prediction accuracy by combining the prediction results of multiple learners. During training, the model's parameters and structure are continuously adjusted to better fit the training data.
[0067] The initial aroma contribution prediction model was optimized using cross-validation. Cross-validation divided the dataset into multiple subsets, using one subset as the training set and the remainder as the validation set, repeating the training and validation process multiple times to evaluate the model's performance and stability. Based on the cross-validation results, the model was further adjusted and optimized, resulting in an optimized aroma contribution prediction model. This model can accurately predict the contribution of aroma substances based on the input olfactory receptor combined with characteristic data and dynamic release data, providing a powerful tool for aroma prediction and analysis.
[0068] S104: Based on the ranking of key aroma substances output by the aroma contribution prediction model, aroma reconstruction experiments and sensory evaluation processing were carried out to verify the accuracy of the prediction.
[0069] Specifically, based on the ranking of key aroma compounds output by the aroma contribution prediction model, the top-ranking key aroma compounds are selected. These compounds have a high contribution to the model's prediction and a significant impact on the overall aroma characteristics. Then, according to the proportions predicted by the model, the selected key aroma compounds are precisely proportioned and mixed to prepare an aroma reconstruction sample, which aims to simulate the aroma characteristics of the original sample.
[0070] Subsequently, sensory evaluation techniques were used to conduct a professional sensory comparison evaluation between the aroma reconstruction samples and the original samples. During the evaluation process, evaluators smelled and perceived the aroma characteristics of the samples, generating sensory score data including dimensions such as aroma intensity and feature similarity. Based on the above sensory score data, a significance test analysis was further conducted to assess whether the sensory differences between the aroma reconstruction samples and the original samples were within a set threshold range using statistical methods. If the test results showed that the sensory differences between the two met the pre-set significance level requirements, the prediction results of the aroma contribution prediction model could be determined to be effective, thus verifying the accuracy of the model's predictions.
[0071] One embodiment of this application provides an aroma prediction method based on olfactory receptors and dynamic simulation, comprising: integrating ligand-specific data of various human olfactory receptors with binding experimental data of target aroma substances based on a public olfactory receptor database to generate a binding characteristic database; simulating environmental parameters of a target scenario using a biomimetic reactor to obtain dynamic release data; training a machine learning model based on the binding characteristic database and dynamic release data to generate an aroma contribution prediction model; and performing aroma reconstruction experiments and sensory evaluation based on the ranking of key aroma substances output by the aroma contribution prediction model to verify the prediction accuracy, thereby achieving the technical effects of improving the consistency between aroma prediction and real olfactory experience, reducing the correlation deviation between dynamic release process and sensory intensity, and improving analysis efficiency to meet the high-throughput requirements of industrialization.
[0072] Furthermore, by simulating the environmental parameters of the target scenario using a biomimetic reactor, dynamic release data is obtained, including:
[0073] The oral processing environment parameters of the target scenario are simulated by a biomimetic oral reactor to generate simulated environmental data. The oral processing environment parameters include chewing frequency, saliva secretion rate, temperature and pH value.
[0074] Based on simulated environmental data, dynamic release simulation processing of target samples is performed to obtain volatile release process data;
[0075] Based on the data of volatile release process, volatile substances are collected and processed at multiple time points using online volatile capture technology to generate time-series aroma substance concentration data;
[0076] Based on the time-series aroma substance concentration data, aroma feature annotation is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive word data;
[0077] Based on time-series data of aroma substance concentration and aroma characteristic descriptor data, dynamic release data is generated by integrating and processing the data.
[0078] Specifically, a biomimetic oral reactor is used to simulate oral processing environment parameters in the target scenario, including chewing frequency, saliva secretion rate, temperature, and pH value, generating simulated environmental data to provide near-realistic oral processing conditions for subsequent experiments. Then, based on the generated simulated environmental data, the target sample undergoes dynamic release simulation treatment, processing the sample in a simulated oral environment to promote the gradual release of volatile substances, thereby obtaining volatile release process data to reflect the dynamic release of aroma substances during oral processing.
[0079] Based on data from the volatile release process, online volatile capture technology was used to collect and process the released volatile substances at multiple different time points, obtaining time-series aroma substance concentration data. This data reflects the concentration changes of aroma substances at different time points. Simultaneously, based on the time-series aroma substance concentration data, sensory evaluation technology was used at each time point to smell and perceive the released volatile substances, generating aroma characteristic descriptive terminology data. By combining chemical detection results with actual human sensory experience, specific olfactory characteristic descriptions were assigned to the aroma substances.
[0080] Subsequently, the time-series aroma substance concentration data and aroma feature descriptor data are integrated and processed. Through relevant algorithms and methods, the two are aligned in time and feature fusion to generate dynamic release data containing aroma substance concentration information and sensory feature information. This provides more comprehensive and accurate data support for subsequent aroma prediction model training and helps improve the model's ability to predict the contribution of aroma substances.
[0081] Furthermore, based on time-series aroma substance concentration data and aroma characteristic descriptor data, data integration and processing are performed to generate dynamic release data, including:
[0082] Time-point alignment processing is performed on the time series aroma substance concentration data and aroma feature descriptor data to generate time-aligned data;
[0083] Based on time-aligned data, aroma release pattern recognition processing is performed to generate dynamic release pattern data.
[0084] Aroma perception encoding conversion processing is performed on the dynamic release mode data to generate dynamic release data.
[0085] Specifically, time-series aroma compound concentration data and aroma feature descriptor data are aligned at specific points in time. Concentration data at the same time point are matched with corresponding aroma feature descriptors to ensure data consistency over time, thereby generating time-aligned data. Then, based on this time-aligned data, the changes in aroma compound concentrations at different time points and their corresponding aroma feature descriptors are analyzed to identify patterns in the release process of aroma compounds, including the start time, peak time, and decay trend, thus generating dynamic release pattern data.
[0086] Subsequently, the dynamic release pattern data is transformed from physical concentration changes and feature descriptions into a coding form that can reflect human olfactory perception. Taking into account factors such as the binding characteristics of olfactory receptors and aroma substances, dynamic release data is generated.
[0087] like Figure 2 As shown, based on the time-series aroma substance concentration data, aroma feature annotation is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive term data, including:
[0088] S201: Based on the time-series aroma substance concentration data at each time point, sensory evaluation technology is used to process the volatile substances through olfaction and generate sensory response data.
[0089] S202: Perform aroma feature descriptor mapping on the sensory response data to generate aroma feature descriptor data;
[0090] S203: Perform key aroma contribution quantification on aroma feature descriptor data to generate key aroma contribution data;
[0091] S204: Integrate aroma feature descriptor data with key aroma contribution data to generate aroma feature descriptor data.
[0092] Specifically, for each time point in the time-series aroma substance concentration data, sensory evaluation techniques are used to organize professionals to smell the volatile substances and record the sensory feedback generated during the smelling process, thereby generating sensory response data. Then, the obtained sensory response data is matched with a pre-established aroma feature descriptor library to transform abstract sensory perceptions into specific aroma feature descriptors, thus generating aroma feature descriptor data.
[0093] Subsequently, based on certain evaluation criteria and methods, the importance of each aroma substance in the aroma feature descriptor data is quantitatively assessed to determine its contribution to the overall aroma, generating key aroma contribution data. Then, the aroma feature descriptor data and key aroma contribution data are integrated to associate the descriptors with their corresponding contributions, generating aroma feature descriptor data containing both aroma feature descriptions and contribution information.
[0094] Furthermore, based on a public olfactory receptor database, ligand-specific data of various human olfactory receptors and binding experimental data of target aroma substances are integrated and processed to generate a binding characteristic database, including:
[0095] Based on a public olfactory receptor database, protein sequences of various human olfactory receptors were acquired and processed to obtain protein sequence data;
[0096] A three-dimensional structural model of the olfactory receptor is generated by performing structural prediction processing on protein sequence data using homology modeling methods.
[0097] The target aroma substance is processed to obtain its three-dimensional structure data.
[0098] Molecular docking simulation was performed on the three-dimensional structural models of olfactory receptors and the three-dimensional structural data of aroma substances to generate binding parameter data.
[0099] Integrate ligand-specific data with binding parameter data to generate a binding characteristic database.
[0100] Specifically, protein sequences of various human olfactory receptors are retrieved from public olfactory receptor databases. After format verification and integrity screening, standardized protein sequence data is generated. Then, using homology modeling methods and proteins with known structures as templates, the three-dimensional structure of the protein sequence data is predicted. Through structure optimization and validation steps, a three-dimensional structural model of the olfactory receptor with spatial conformation is generated. Next, for target aroma substances, their three-dimensional structures are extracted from chemical databases or generated using molecular construction software. After geometric optimization, relatively accurate three-dimensional structural data of the aroma substances are obtained.
[0101] The three-dimensional structural models of olfactory receptors and the three-dimensional structural data of aroma substances are imported into molecular docking software to simulate the binding process, including the calculation of binding parameters such as binding affinity and binding sites. Then, ligand-specific data (such as binding affinity and activation threshold) obtained from a database are integrated with the binding parameter data generated by molecular docking, stored and indexed according to a unified data format, generating a database containing the binding characteristics of olfactory receptors and target aroma substances.
[0102] Furthermore, based on the ranking of key aroma compounds output by the aroma contribution prediction model, aroma reconstruction experiments and sensory evaluation were conducted to verify the prediction accuracy, including:
[0103] Key aroma compounds are screened based on the contribution ranking output by the aroma contribution prediction model.
[0104] Key aroma compounds were mixed in the predicted proportions to prepare aroma reconstruction samples;
[0105] Sensory evaluation technology was used to compare and evaluate the aroma reconstruction samples with the original samples to generate sensory score data.
[0106] Based on sensory rating data, a significance test was performed to generate the test results;
[0107] If the test results meet the set threshold, the prediction is considered valid.
[0108] Specifically, based on the aroma contribution ranking of aroma substances output by the aroma contribution prediction model, aroma substances that play a key role in the overall aroma are selected according to preset screening rules (such as contribution threshold or top N rankings). Then, based on the concentration ratio or relative content of each key aroma substance in the original sample given by the prediction model, these substances are accurately weighed and mixed to prepare an aroma reconstruction sample.
[0109] Subsequently, professionally trained evaluators used sensory evaluation techniques (such as quantitative descriptive analysis, QDA) to compare and evaluate the aroma characteristics of the reconstructed aroma samples with those of the original samples, recording and organizing the sensory score data. Then, statistical methods were used to test the significance of differences in the sensory score data, analyzing the degree of difference in perceived aroma between the reconstructed and original samples, and generating test results. Finally, the test results were compared with pre-set thresholds (such as a significance level of P > 0.05 or a sensory score difference of less than 10%). When the test results met the set thresholds, the predictive model was deemed effective in predicting key aroma substances.
[0110] Furthermore, based on the combination of a feature database and dynamic release data, a machine learning model is trained to generate an aroma contribution prediction model, including:
[0111] Based on the binding characteristic database, olfactory receptor binding parameter data were extracted;
[0112] Based on dynamic release data, time series data and substance concentration data are extracted;
[0113] Feature fusion processing is performed on olfactory receptor binding parameter data, time series data, and substance concentration data to generate multidimensional input features;
[0114] Based on sensory evaluation data, obtain key aroma contribution labels;
[0115] An ensemble learning algorithm is used to train the model on multidimensional input features and key aroma contribution labels to generate an initial aroma contribution prediction model.
[0116] The initial aroma contribution prediction model was optimized using cross-validation techniques to generate an optimized aroma contribution prediction model.
[0117] Specifically, binding parameter data between olfactory receptors and aroma compounds are extracted from a binding characteristic database, including key parameters reflecting the binding characteristics of both, such as binding affinity and activation threshold. Then, time-series data and corresponding substance concentration data are extracted from dynamic release data. The time-series data covers various time points of aroma compound release, while the substance concentration data reflects the concentration of aroma compounds at each time point. Subsequently, feature fusion processing is performed on the extracted olfactory receptor binding parameter data, time-series data, and substance concentration data to integrate different types of data into input features containing multi-dimensional information, thus comprehensively reflecting the release and effects of aroma compounds.
[0118] Based on sensory evaluation data, key aroma contribution labels are obtained for model training. These labels identify the contribution of each aroma compound to the overall aroma at different time points. Then, ensemble learning algorithms, such as random forest or XGBoost, are used to train the model on the multi-dimensional input features and key aroma contribution labels. The algorithm learns the mapping relationship between input features and labels, generating an initial aroma contribution prediction model. Next, cross-validation is used to divide the training data into multiple subsets, and training and validation are performed on different subsets. The hyperparameters of the initial model are then optimized and adjusted to generate an optimized aroma contribution prediction model.
[0119] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0120] In one embodiment, such as Figure 3 As shown, this application also provides an aroma prediction system 300 based on olfactory receptors and dynamic simulation, the system 300 comprising:
[0121] The database construction module 301 is used to integrate and process ligand-specific data of various human olfactory receptors and binding experimental data of target aroma substances based on a public olfactory receptor database to generate a binding characteristic database.
[0122] The dynamic simulation module 302 is used to simulate and process the environmental parameters of the target scene through a biomimetic reactor to obtain dynamic release data;
[0123] The intelligent prediction module 303 is used to train a machine learning model based on the combined characteristic database and dynamic release data to generate an aroma contribution prediction model.
[0124] The closed-loop verification module 304 is used to sort the key aroma substances based on the output of the aroma contribution prediction model, conduct aroma reconstruction experiments and sensory evaluation processing, and verify the prediction accuracy.
[0125] Specifically, the database construction module 301 obtains ligand-specific data of various human olfactory receptors from a public olfactory receptor database, and combines it with binding experimental data of target aroma substances. After integration and processing, a binding characteristic database is generated to provide basic data support for subsequent analysis.
[0126] The dynamic simulation module 302 uses a biomimetic reactor to simulate the environmental parameters of the target scene more accurately. During the simulation, it collects and processes data on the release process of volatile substances from the target sample to obtain dynamic release data, which reflects the release of aroma substances in the real scene.
[0127] The intelligent prediction module 303 extracts and fuses relevant features based on the combination of characteristic database and dynamic release data, and then combines the labels obtained from sensory evaluation data. It uses an ensemble learning algorithm to train the model and optimizes it through cross-validation to generate an aroma contribution prediction model to predict the aroma contribution.
[0128] The verification closed-loop module 304 performs aroma reconstruction experiments based on the ranking of key aroma substances output by the prediction model. After preparing the reconstructed sample, it conducts sensory comparison and evaluation with the original sample. The accuracy of the prediction is verified through processes such as difference significance test, thus forming a complete technical closed loop.
[0129] The dynamic simulation module 302 is also used for:
[0130] The oral processing environment parameters of the target scenario are simulated by a biomimetic oral reactor to generate simulated environmental data. The oral processing environment parameters include chewing frequency, saliva secretion rate, temperature and pH value.
[0131] Based on simulated environmental data, dynamic release simulation processing of target samples is performed to obtain volatile release process data;
[0132] Based on the data of volatile release process, volatile substances are collected and processed at multiple time points using online volatile capture technology to generate time-series aroma substance concentration data;
[0133] Based on the time-series aroma substance concentration data, aroma feature annotation is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive word data;
[0134] Based on time-series data of aroma substance concentration and aroma characteristic descriptor data, dynamic release data is generated by integrating and processing the data.
[0135] The dynamic simulation module 302 is also used for:
[0136] Time-point alignment processing is performed on the time series aroma substance concentration data and aroma feature descriptor data to generate time-aligned data;
[0137] Based on time-aligned data, aroma release pattern recognition processing is performed to generate dynamic release pattern data.
[0138] Aroma perception encoding conversion processing is performed on the dynamic release mode data to generate dynamic release data.
[0139] The dynamic simulation module 302 is also used for:
[0140] Based on the time-series aroma substance concentration data at each time point, sensory evaluation technology is used to process volatile substances through olfaction and generate sensory response data.
[0141] Aroma feature descriptor mapping is performed on sensory response data to generate aroma feature descriptor data;
[0142] The aroma feature descriptor data is subjected to key aroma contribution quantification processing to generate key aroma contribution data;
[0143] Aroma feature descriptor data and key aroma contribution data are integrated to generate aroma feature descriptor data.
[0144] Database building module 301 is also used for:
[0145] Based on a public olfactory receptor database, protein sequences of various human olfactory receptors were acquired and processed to obtain protein sequence data;
[0146] A three-dimensional structural model of the olfactory receptor is generated by performing structural prediction processing on protein sequence data using homology modeling methods.
[0147] The target aroma substance is processed to obtain its three-dimensional structure data.
[0148] Molecular docking simulation was performed on the three-dimensional structural models of olfactory receptors and the three-dimensional structural data of aroma substances to generate binding parameter data.
[0149] Integrate ligand-specific data with binding parameter data to generate a binding characteristic database.
[0150] The verification closed-loop module 304 is also used for:
[0151] Key aroma compounds are screened based on the contribution ranking output by the aroma contribution prediction model.
[0152] Key aroma compounds were mixed in the predicted proportions to prepare aroma reconstruction samples;
[0153] Sensory evaluation technology was used to compare and evaluate the aroma reconstruction samples with the original samples to generate sensory score data.
[0154] Based on sensory rating data, a significance test was performed to generate the test results;
[0155] If the test results meet the set threshold, the prediction is considered valid.
[0156] The intelligent prediction module 303 is also used for:
[0157] Based on the binding characteristic database, olfactory receptor binding parameter data were extracted;
[0158] Based on dynamic release data, time series data and substance concentration data are extracted;
[0159] Feature fusion processing is performed on olfactory receptor binding parameter data, time series data, and substance concentration data to generate multidimensional input features;
[0160] Based on sensory evaluation data, obtain key aroma contribution labels;
[0161] An ensemble learning algorithm is used to train the model on multidimensional input features and key aroma contribution labels to generate an initial aroma contribution prediction model.
[0162] The initial aroma contribution prediction model was optimized using cross-validation techniques to generate an optimized aroma contribution prediction model.
[0163] In one embodiment, (1) a database of human ORs-aroma substance binding characteristics is established:
[0164] Based on public ORs databases (such as ORDB), ligand-specific data (binding affinity Kd, activation threshold EC50) of 403 human ORs were integrated; for characteristic volatile substances of japonica rice (such as 2-acetyl-1-pyrrolline, nonanal, and phenylacetaldehyde), molecular docking simulation data (using AutoDock Vina software to calculate the binding free energy of ligand-ORs) or cell experiment data (HEK293 cells expressing ORs, detecting the intensity of calcium signal response) were added.
[0165] (2) Data collection of dynamic aroma release during simulated oral processing of japonica rice:
[0166] A biomimetic oral reactor was used, with the following parameters: chewing frequency 100 rpm, saliva secretion rate 1 mL / min (artificial saliva composition: NaCl 0.9 g / L, KCl 0.2 g / L, CaCl 0.1 g / L, α-amylase 100 U / mL), temperature 37℃, and pH 6.8. 5 g of cooked japonica rice was added to the reactor, and at 0, 30, 60, 90, and 120 seconds, the dynamic concentration curves of each substance were obtained through online HS-SPME (65 μm PDMS / DVB fiber, maintained at 40℃ for 3 min, then increased to 250℃ at 5℃ / min and maintained for 5 min). Simultaneously, GC-O technology was used (3 trainees smelled the rice and labeled the characteristics and intensity of "rice aroma," "sweet aroma," and "grass aroma").
[0167] (3) Machine learning model construction and training:
[0168] Input characteristics: ORs binding free energy (n=403), time (0-120s), substance concentration (μg / g);
[0169] Output label: "Key aroma contribution" as indicated by GC-O (1-5 points, 5 points being the highest contribution);
[0170] Model selection: The XGBoost algorithm was adopted (parameters: learning rate 0.1, maximum depth 6, number of iterations 100), and the hyperparameters were optimized with 10-fold cross-validation;
[0171] Training data: Includes dynamic release data from 100 groups of japonica rice samples (covering 5 varieties).
[0172] (4) Prediction and verification of key aroma compounds:
[0173] Prediction: Input the dynamic release data and ORs binding characteristics of the japonica rice to be analyzed, and the model outputs the ranking of the "key contribution" of each aroma substance (e.g., the top 3 are 2-acetyl-1-pyrrolline, nonanal, and phenylacetaldehyde).
[0174] Verification: The predicted key substances were mixed in proportion (e.g., 50 μg / g of 2-acetyl-1-pyrrolline, 20 μg / g of nonanal, and 15 μg / g of phenylacetaldehyde) to prepare a reconstructed sample. The reconstructed sample was evaluated by 10 trainers using QDA sensory evaluation (scores of "rice aroma" and "sweet aroma" intensity). If the score difference between the reconstructed sample and the original japonica rice was less than 10% (P>0.05), the prediction was considered valid.
[0175] In one embodiment, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0176] In one embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.
[0177] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0178] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. An aroma prediction method based on olfactory receptors and dynamic simulation, characterized in that, The method includes: Based on a public olfactory receptor database, ligand-specific data of various human olfactory receptors and binding experimental data of target aroma substances are integrated and processed to generate a binding characteristic database. By simulating the environmental parameters of the target scenario using a biomimetic reactor, dynamic release data is obtained. Based on the combined characteristic database and the dynamic release data, a machine learning model is trained to generate an aroma contribution prediction model. Based on the ranking of key aroma substances output by the aroma contribution prediction model, aroma reconstruction experiments and sensory evaluation were conducted to verify the accuracy of the prediction.
2. The aroma prediction method based on olfactory receptors and dynamic simulation according to claim 1, characterized in that, The process of simulating environmental parameters of the target scenario using a biomimetic reactor to obtain dynamic release data includes: The oral processing environment parameters of the target scenario are simulated by a biomimetic oral reactor to generate simulated environmental data. The oral processing environment parameters include chewing frequency, saliva secretion rate, temperature and pH value. Based on the simulated environment data, the target sample is subjected to dynamic release simulation processing to obtain volatile release process data; Based on the data of the volatile release process, volatile substances are collected and processed at multiple time points using online volatile capture technology to generate time-series aroma substance concentration data. Based on the aroma substance concentration data of the time series, aroma feature annotation is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive word data; Based on the time series aroma substance concentration data and the aroma feature descriptor data, the data is integrated and processed to generate the dynamic release data.
3. The aroma prediction method based on olfactory receptors and dynamic simulation according to claim 2, characterized in that, The aroma substance concentration data based on the time series and the aroma feature descriptor data are integrated and processed to generate the dynamic release data, including: The aroma substance concentration data and aroma feature descriptor data of the time series are aligned at time points to generate time-aligned data; Based on the time-aligned data, aroma release pattern recognition processing is performed to generate dynamic release pattern data; The dynamic release mode data is processed by aroma perception encoding conversion to generate the dynamic release data.
4. The aroma prediction method based on olfactory receptors and dynamic simulation according to claim 2, characterized in that, Based on the aroma substance concentration data of the time series, aroma feature annotation processing is performed simultaneously using sensory evaluation technology to generate aroma feature descriptive word data, including: Based on each time point in the aroma substance concentration data of the time series, volatile substance olfaction processing is performed using sensory evaluation technology to generate sensory response data. The sensory response data is processed by aroma feature descriptor mapping to generate aroma feature descriptor data; The aroma feature descriptor data is subjected to key aroma contribution quantification processing to generate key aroma contribution data; The aroma feature descriptor data is generated by integrating the aroma feature descriptor data with the key aroma contribution data.
5. The aroma prediction method based on olfactory receptors and dynamic simulation according to claim 1, characterized in that, The method, based on a public olfactory receptor database, integrates ligand-specific data of various human olfactory receptors with binding experimental data of target aroma substances to generate a binding characteristic database, including: Based on a public olfactory receptor database, protein sequences of various human olfactory receptors were acquired and processed to obtain protein sequence data. The protein sequence data were processed by homology modeling to predict the structure and generate a three-dimensional structural model of the olfactory receptor. The target aroma substance is processed to obtain its three-dimensional structure data. Molecular docking simulation was performed on the three-dimensional structural model of the olfactory receptor and the three-dimensional structural data of the aroma substance to generate binding parameter data. The binding characteristic database is generated by integrating the ligand specificity data and the binding parameter data.
6. The aroma prediction method based on olfactory receptors and dynamic simulation according to claim 1, characterized in that, The ranking of key aroma substances based on the aroma contribution prediction model, followed by aroma reconstruction experiments and sensory evaluation processing to verify prediction accuracy, includes: Based on the contribution ranking output by the aroma contribution prediction model, key aroma substances are screened. The key aroma substances were mixed in the predicted proportions to prepare an aroma reconstruction sample; Sensory evaluation technology is used to compare and evaluate the aroma reconstruction sample with the original sample to generate sensory score data. Based on the sensory rating data, a significant difference test is performed to generate the test results; When the test result meets the set threshold, the prediction is deemed valid.
7. The aroma prediction method based on olfactory receptors and dynamic simulation according to claim 1, characterized in that, The step of training a machine learning model based on the combined characteristic database and the dynamic release data to generate an aroma contribution prediction model includes: Based on the aforementioned binding characteristic database, olfactory receptor binding parameter data are extracted; Based on the dynamic release data, time series data and substance concentration data are extracted; The olfactory receptor binding parameter data, the time series data, and the substance concentration data are subjected to feature fusion processing to generate multidimensional input features; Based on sensory evaluation data, obtain key aroma contribution labels; An ensemble learning algorithm is used to train the model on the multidimensional input features and the key aroma contribution labels to generate an initial aroma contribution prediction model. The initial aroma contribution prediction model is optimized using cross-validation technology to generate the optimized aroma contribution prediction model.
8. An aroma prediction system based on olfactory receptors and dynamic simulation, characterized in that, The system includes: The database construction module is used to integrate and process ligand-specific data of various human olfactory receptors with experimental data on the binding of target aroma substances based on a public olfactory receptor database, and generate a binding characteristic database. The dynamic simulation module is used to simulate the environmental parameters of the target scene using a biomimetic reactor to obtain dynamic release data; The intelligent prediction module is used to train a machine learning model based on the combined characteristic database and the dynamic release data to generate an aroma contribution prediction model. The closed-loop verification module is used to perform aroma reconstruction experiments and sensory evaluation processing based on the ranking of key aroma substances output by the aroma contribution prediction model, and to verify the accuracy of the prediction.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the aroma prediction method based on olfactory receptors and dynamic simulation as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the aroma prediction method based on olfactory receptors and dynamic simulation as described in any one of claims 1 to 7.