Method, equipment and medium for designing precursor perfume and predicting release efficiency
By using multi-model collaborative design of latent fragrance body structures and combining chemical bond breaking mechanisms, stable preservation and precise release of fragrances in the daily chemical industry have been achieved, solving the problem of insufficient controlled-release performance of latent fragrance bodies and improving design efficiency and release accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, there is insufficient research on the controlled release performance of latent fragrance components, and a lack of release efficiency prediction models and on-demand reverse design, which limits the application of fragrances in the daily chemical industry.
A multi-model collaborative approach was adopted, using message passing neural network (MPNN), Retro-MTGR, BoltzGen and MACE models, combined with chemical bond breaking mechanism, to design latent aroma body structures and predict their release efficiency, achieving accurate matching from aroma to molecular structure and prediction of release rate.
It improves the efficiency and accuracy of latent fragrance design, ensuring stable preservation and precise release of fragrances under different environmental conditions, and meeting the application needs of different matrices.
Smart Images

Figure CN121789828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fragrance and flavor technology, and in particular to a method, device, and medium for designing latent fragrance components and predicting their release efficiency. Background Technology
[0002] The volatility and poor stability of most fragrances limit their application in the daily chemical industry. Taking aldehydes and ketones as an example, these are highly active substances with distinctive aromas but are volatile and unstable. The main methods for controlled release of fragrances are physical microcapsules and chemical latent fragrance bodies. Microcapsules release their core material in a concentrated manner after the wall material ruptures, and their morphology is often spherical, making it difficult to meet the requirements of different matrices for fragrance materials. To address this industry pain point, latent fragrance body technology has emerged. It precisely anchors highly volatile fragrance molecules onto a non-volatile matrix through characteristic chemical bonds, forming a structurally stable "fragrance precursor." Environmental factors then gradually trigger the controlled breaking of these chemical bonds. This overcomes the passive protection limitations of physical encapsulation, preventing premature loss of active ingredients and ensuring continuous and stable fragrance release. It represents a breakthrough in the technological upgrading of high-end fragrance products in my country.
[0003] Currently, research on the application of artificial intelligence in the field of latent aroma compounds is still in its early stages, both domestically and internationally. The controlled-release performance of latent aroma compounds is the core foundation for their application, but there is currently no research on predictive models for latent aroma compound release efficiency, nor is there any exploration of on-demand reverse engineering of latent aroma compounds. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a method, device, and medium for latent aroma body design and release efficiency prediction. This method features multi-model collaborative driving, which improves design efficiency and accuracy. The release performance design is precise and controllable, and has reliability and robustness.
[0005] The objective of this invention can be achieved through the following technical solutions: The core mechanism of this invention lies in the fact that latent fragrance compounds connect aroma molecules to modifying groups through responsive chemical bonds. These chemical bonds can break under specific environmental stimuli. The essence of the breaking mechanism is the precise matching between environmental signals and chemical bond characteristics: under acidic conditions, proton attack causes polar chemical bonds to polarize and hydrolyze; increased temperature overcomes bond energy barriers by intensifying molecular thermal motion; and enzymes catalyze targeted breaking through specific substrate conformations, ultimately releasing free aroma molecules. This achieves stable storage and precise release during use. Based on the formation and release mechanism of latent fragrance compounds, this invention focuses on the formation and breaking of chemical bonds in the fragrance field. The developed model derives the fragrance structure from the aroma, designs a latent fragrance compound model, and predicts its release rate. The model's mechanism is a cross-scale structure-activity relationship (SCR) modeling driven by a responsive chemical bond breaking mechanism. Using the core mechanism of latent aroma compounds—"environmental stimulus-chemical bond breaking-aroma release"—as its logical anchor, it visualizes the breaking mechanism as quantifiable characteristics and integrates a full-chain prediction framework from the molecular structure of synthetic fragrance molecules to the design of latent aroma compound molecules and the prediction of their release efficiency. Through the synergistic adaptation of multiple models, it achieves matching from aroma to molecular structure, and from the molecular design of latent aroma compounds to the prediction of their release rate. With the chemical nature of chemical bond formation and breaking as the underlying constraint, the model prediction not only relies on data correlation but also aligns more closely with the responsive release mechanism of latent aroma compounds.
[0006] This invention provides a method for designing latent fragrance components and predicting their release efficiency, comprising the following steps: S1. Using the message passing neural network (MPNN) mechanism, the input aroma description is mapped to the corresponding aroma molecular structure and output as the string SMILES; Furthermore, the Message Passing Neural Network (MPNN) mechanism includes the following steps: integrating the aroma tag-molecular structure corresponding dataset and performing standardization and data augmentation; designing a four-module architecture of aroma embedding, feature mapping, molecular graph generation, and rationality verification, and converting aroma tags into structured embedding vectors through a Transformer encoder; realizing the mapping from aroma features to molecular structure features through the MPN encoder and a multilayer perceptron (MLP); combining chemical bonding rules and syntheticity constraints, using a generative graph neural network (GNN) to gradually generate molecular graphs and output the string "SMILES"; evaluating the model performance based on database and sensory evaluation, using multi-dimensional indicators such as structural rationality and aroma matching degree, and optimizing the model according to the evaluation results, ultimately achieving accurate prediction from the target aroma to a highly compatible aroma molecular structure.
[0007] S2. The aroma molecule structure is analyzed using the Retro-MTGR model. Inference is performed using an atom embedding enhancer and a reaction center predictor to filter out a list of atoms with the highest reactivity. Further, in S2, the Retro-MTGR model's workflow includes: receiving the SMILES string of the aroma molecule and converting it into a molecular diagram; extracting atom type, hybridization mode, and bond type features; and performing inference using the atom embedding enhancer and reaction center predictor to finally output a list of atoms with the highest reactivity.
[0008] S3. Based on the BoltzGen model, combined with the site atom list and the user-defined target release conditions, a candidate ligand molecular map containing controllable bond breaking is generated. Further, in S3, the BoltzGen model is a conditional variational graph autoencoder, and its workflow includes: encoding the target release conditions and inputting them into the model along with the features of the site atom list; generating a 2D molecular map of ligands containing environmentally responsive chemical bonds; and performing chemical rationality filtering on the generated molecular map to obtain a candidate ligand library.
[0009] Furthermore, the environmentally responsive chemical bonds include at least one of acetal bonds, ester bonds, or glycosidic bonds. The target release conditions include at least one of pH value, temperature, oxygen concentration, enzyme concentration, and light intensity.
[0010] S4. Using the RDKit tool, the aroma molecules and candidate ligands are docked according to the reaction rules to generate a latent aroma body structure; S5. The MACE model is used to predict the bond dissociation energy of the latent fragrance body structure under normal and triggered conditions, and candidate latent fragrance bodies that meet the requirements of stability and responsiveness are screened out. Further, in S5, the MACE model is used to predict the bond dissociation energy of the latent fragrance body. The workflow includes: receiving the latent fragrance body 3D structure generated by the docking of aroma molecules and candidate ligands; predicting the bond dissociation energy of the latent fragrance body 3D structure under normal environment and specific triggered conditions; and screening out latent fragrance bodies with high bond dissociation energy under normal conditions and significantly reduced bond dissociation energy under triggered conditions as core candidates.
[0011] S2-S5 designs latent aroma compounds that are stably encapsulated under normal conditions and precisely released under specific conditions. The implementation involves a phased model collaboration: Retro-MTGR first receives aroma molecules (SMILES), converts them into molecular graphs, and extracts atomic and chemical bond features. Through inference using an atom embedding enhancer and a reaction center predictor, a list of sites with the highest reactivity is ultimately selected. Subsequently, BoltzGen integrates this site list with the target release conditions. The conditions are encoded and combined with site features, then input into a conditional variational graph autoencoder to generate a 2D molecular graph of ligands containing controllable bond breaking. Molecules with poor chemical plausibility are then filtered to obtain candidate ligands. Using the original synthesized aroma molecule structure, binding sites, ligands, and given reaction rules, the RDKit tool generates latent aroma compound structures. Further, MACE is used to receive possible latent aroma compound structures generated by aroma-ligand docking, converts the format, and predicts their bond dissociation energies under normal and triggered conditions, selecting candidate latent aroma compounds. Finally, the corresponding latent aroma compounds are synthesized, the generated structural model is validated, and the performance of the latent aroma compounds is tested using the previously set trigger conditions.
[0012] S6. Using a gradient boosting model, predict the release efficiency of the potential fragrance compounds based on their molecular structure characteristics and environmental conditions.
[0013] Furthermore, in S6, the gradient boosting model is one of the following models: XGBoost, LightGBM, GBRT, CatBoost, Histogram-Based GBDT, NGBoost, etc. The steps for predicting release efficiency using the gradient boosting model include: collecting a dataset containing the molecular structure, environmental conditions, and release rate of the latent fragrance; extracting molecular descriptors using RDKit and standardizing the environmental parameters to form a feature matrix; dividing the dataset into training, validation, and test sets; performing feature importance analysis to select key features; setting hyperparameters and training the model; and finally inputting the feature data of the new latent fragrance into the trained model and outputting the predicted release rate value.
[0014] The specific implementation steps are as follows: First, collect a dataset containing information on the molecular structure of the latent fragrance, environmental conditions, and corresponding measured release rates. Use RDKit to extract the structural features of the molecules, and standardize or normalize the environmental conditions and process parameters to form a complete feature matrix. Then, divide the dataset into a training set: validation set: test set ratio of 7:2:1, handle missing and outlier values, and select key features based on feature importance analysis to reduce dimensionality. Next, use a gradient boosting framework to build a prediction model, set hyperparameters such as learning rate, tree depth, and number of iterations, and fine-tune the hyperparameters through grid search or Bayesian optimization. Use mean squared error as the loss function, train the model using the training set, and monitor overfitting using the validation set, terminating training with an early stopping strategy. Then, evaluate the model's generalization ability on the test set using metrics such as root mean square error, mean absolute error, and coefficient of determination. Finally, input the feature data of the newly designed latent fragrance into the trained model, output the predicted release rate, and optimize the latent fragrance molecular structure or application conditions based on the prediction results.
[0015] Finally, the user interface settings clearly define the input and output. First, the user needs to input a scent description. After providing the desired scent, the system matches the corresponding molecular structure and then introduces the molecular structure into specific atomic or chemical bond sites of the environmental response group. In this step, the user needs to input specific environmental response conditions, and the system provides the response group and optimal site selection. Finally, the user needs to provide specific application data to display a predicted release rate curve for the latent fragrance molecules.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method for latent aroma body design and release efficiency prediction.
[0017] The present invention also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, a method for designing latent aroma compounds and predicting release efficiency is implemented.
[0018] Compared with the prior art, the present invention has the following advantages: (1) Multi-model collaborative driving improves design efficiency and accuracy. Multiple artificial intelligence models (message passing neural network, Retro-MTGR, BoltzGen, MACE, and gradient boosting model) are applied in series to achieve an integrated process from aroma description to aroma molecule structure generation, latent aroma body design, and release efficiency prediction. Aroma-molecule matching stage: Abstract aroma labels are mapped to specific molecular structures using a message passing neural network. Combined with chemical bonding rules and syntheticity constraints, the generated molecules are ensured to have both aroma compatibility and chemical rationality. Latent aroma body design stage: Retro-MTGR is used to locate the optimal modification sites for fragrance molecules, BoltzGen generates environmentally responsive ligands, and MACE predicts bond dissociation energies. Finally, latent aroma body structures that are stable under normal conditions and release efficiently under specific conditions are selected. Release efficiency prediction stage: Based on gradient boosting models (such as XGBoost), molecular structure features and environmental parameters are integrated to achieve quantitative prediction of release rates, providing data support for practical applications.
[0019] (2) Accurate prediction of release efficiency. Key features such as bond energy and logP are screened in the release prediction stage to avoid the blindness of pure black box models. A complete chain of "theoretical prediction - experimental verification - optimization iteration" is formed by combining DFT calculation, molecular dynamics simulation, organic synthesis experiments and sensory evaluation. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method for designing and predicting the release efficiency of Qingzixiang latent fragrance components in Example 1; Figure 2 This is a flowchart of the method for designing and predicting the release efficiency of honey-sweet fragrance latent body in Example 2; Figure 3 This is a flowchart of the method for designing animal-derived fragrance components and predicting release efficiency in Example 3; Figure 4 A schematic diagram of a fragrance wheel representing aldehydes and ketones with different aromas. Detailed Implementation
[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Component models, material names, connection structures, control methods, algorithms, and other features not explicitly described in this technical solution are considered common technical features disclosed in the prior art.
[0022] Examples 1, 2, and 3 respectively use Figure 4 Take, for example, the subtle aromas of green spice, honey spice, and animalic spice in the fragrance wheel.
[0023] Example 1 This embodiment provides a method for designing latent aroma components and predicting release efficiency, such as... Figure 1As shown, it includes the following steps: S1. Using the message passing neural network (MPNN) mechanism, the input aroma description is mapped to the corresponding aroma molecular structure and output as the string SMILES; This embodiment takes the Qingzixiang aroma profile as an example, with "a strong, refreshing minty scent with a slightly sweet herbal note" as the target aroma input. First, it enters the aroma-molecule generation stage, integrating public databases and proprietary measured aroma-structure data to construct a standardized dataset containing hierarchical labels such as "minty, refreshing, herbal, and sweet." RDKit is used to perform molecular structure deduplication, standardization, and data augmentation. Then, an MPNN architecture with four modules—aroma embedding, feature mapping, molecular graph generation, and plausibility verification—is employed, using a Transformer encoder to process the target aroma. The labels are converted into structured embedding vectors, which are then mapped from aroma features to molecular structure features via an MPN encoder and an MLP. The generative GNN gradually constructs a molecular graph under the constraints of chemical bonding rules and syntheticity. The model is then optimized through "pre-training + fine-tuning". After closed-loop verification through structural rationality verification, aroma matching degree evaluation and DFT calculation, organic synthesis experiment and sensory evaluation, the target aroma molecule menthone (SMILES: CC1(C)CCCC(C1=O)C) is output. Its ketone group matches the cycloalkane skeleton to meet the requirements of cool aroma, and the side chain methyl matches the characteristics of sweet and herbal aroma.
[0024] S2. The aroma molecule structure was analyzed using the Retro-MTGR model, and a list of sites with the highest reactivity was selected. S3. Based on the BoltzGen model, combined with the site atom list and the user-defined target release conditions, a candidate ligand molecule map containing controllable bond breaking is generated. S2-S3 is the multi-model collaborative stage: Retro-MTGR receives menthone smiles and converts them into molecular maps, extracting features such as atom type, hybridization mode, and bond type. Through inference by an atom embedding enhancer and a reaction center predictor, the most reactive C atom on the ketone group is selected as the optimal modification site. BoltzGen integrates this site list with the target condition of "precise release at pH 3.5," encodes the environmental conditions, and inputs them along with the site features into a conditional variational graph autoencoder to generate a 2D molecular map of ligands containing pH-responsive acetal bonds. After filtering out molecules with poor chemical rationality, candidate molecules are obtained. The ligand library was selected; the MACE receiver obtained the 3D structure of the complex formed by docking menthone with the candidate ligands. After converting the format, the bond dissociation energy under neutral environment, 25℃ and trigger conditions of pH 3.5 and 25℃ was predicted. The core candidate was screened out as "high bond dissociation energy under normal conditions and significantly reduced bond dissociation energy under trigger conditions". The acetal bond breaking probability under acidic conditions was confirmed to be ≥90% through simulation. The release kinetics met the requirements of "slow storage-rapid release". Finally, a pH-responsive menthone latent aroma body (SMILES: CC1(C)CCCC(C1(OCC2CO2)=O)C) was formed.
[0025] S4. Using the RDKit tool, the aroma molecules and candidate ligands are docked according to the reaction rules to generate a latent aroma body structure; S5. The MACE model is used to predict the bond dissociation energy of the latent fragrance body structure under normal and triggered conditions, and candidate latent fragrance bodies that meet the requirements of stability and responsiveness are screened out. In a specific implementation, in S5, the MACE model is used to predict the bond dissociation energy of the latent fragrance body. The workflow includes: receiving the latent fragrance body 3D structure generated by the docking of aroma molecules and candidate ligands; predicting the bond dissociation energy of the latent fragrance body 3D structure under normal environment and specific triggered conditions; and screening out latent fragrance bodies with high bond dissociation energy under normal conditions and significantly reduced bond dissociation energy under triggered conditions as core candidates.
[0026] S6. Using a gradient boosting model, predict the release efficiency of the potential fragrance compounds based on their molecular structure characteristics and environmental conditions.
[0027] S3-S6 represent the release rate prediction stage, using the XGBoost gradient boosting model. The prediction model was constructed using XGBoost by collecting molecular structure information, environmental parameters, and corresponding measured release rate data for this type of pH-responsive latent aroma compound. Molecular structure features were extracted using RDKit, and environmental parameters were standardized and merged to form a feature matrix. The training, validation, and test sets were divided in a 7:2:1 ratio. After handling missing and outlier values, key features such as acetal bond energy, logP, and pH value were selected based on mutual information. Hyperparameters such as a learning rate of 0.05, tree depth of 6, and number of iterations of 200 were set and optimized using Bayesian optimization. The model was trained using mean squared error as the loss function, employing an early stopping strategy to avoid overfitting. On the test set, the model achieved RMSE=0.032, MAE=0.025, and R0.032. 2 After verifying the generalization ability of 0.93, the newly designed latent aroma features were input into the model to predict their release rate in beverage application scenarios. The results showed that the release amount within the shelf life was ≤10%, and the release amount within 30 minutes of drinking was ≥85%, ultimately achieving a precise match between the target aroma and the controllable release performance.
[0028] Finally, the user interface settings clearly define the input and output. First, the user needs to input a scent description, specifying the desired scent. The interface then matches the corresponding molecular structure and introduces this structure into specific atomic or chemical bond sites of the environmental response group. Next, the user selects environmental response conditions, specifying the response group and optimal site selection. Finally, the user provides specific application data to display a predicted release rate curve for the latent fragrance molecules.
[0029] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for designing latent fragrance bodies and predicting release efficiency.
[0030] This embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for latent aroma body design and release efficiency prediction.
[0031] Example 2 This embodiment provides a method for designing latent aroma components and predicting release efficiency, such as... Figure 2 As shown, it includes the following steps: S1. Using the message passing neural network (MPNN) mechanism, the input aroma description is mapped to the corresponding aroma molecular structure and output as the string SMILES; This embodiment takes the honey-sweet aroma profile as an example, with "rich honey-sweet aroma, accompanied by soft floral notes and a slight woody undertone" as the target aroma input. First, it enters the aroma-molecule generation stage, integrating the publicly available FlavorDB database and proprietary measured aroma-structure data to construct a standardized hierarchical label dataset containing "honey sweetness, floral notes, woody notes, and softness." RDKit is used to perform molecular structure deduplication, standardization, and data augmentation. Subsequently, an MPNN architecture with four modules—aroma embedding, feature mapping, molecular graph generation, and plausibility verification—is employed. A Transformer encoder is used to convert the target aroma labels into structural data. The structured embedding vector is mapped from aroma features to molecular structure features through MPN encoder and MLP. Generative GNN gradually constructs molecular graph under chemical bonding rules and syntheticity constraints. After the model is optimized by "pre-training + fine-tuning", it is verified by structural rationality check, aroma matching degree evaluation and DFT calculation, organic synthesis experiment and sensory evaluation closed loop verification to output the target aroma molecule irisone (SMILES: CC1=CC(=O)C(C)(C)C1). Its cyclohexenone skeleton and isopropyl side chain are the core source of honey sweetness. The double bond structure gives it a soft floral quality, and the ring structure carries a slight woody background.
[0032] S2. The aroma molecule structure was analyzed using the Retro-MTGR model, and a list of sites with the highest reactivity was selected. S3. Based on the BoltzGen model, combined with the site atom list and the user-defined target release conditions, a candidate ligand molecule map containing controllable bond breaking is generated. S2-S3 is the multi-model collaborative stage: Retro-MTGR receives iris smiles and converts them into molecular maps, extracting features such as atom type, hybridization mode, and bond type. Through inference by the atom embedding enhancer and reaction center predictor, the most reactive C atom on the cyclohexenone ring is selected as the optimal modification site. BoltzGen integrates this site list with the target condition of "precise release at 32℃," encodes the environmental conditions, and inputs them along with the site features into the conditional variational map autoencoder to generate a 2D molecular map of ligands containing temperature-responsive ester bonds. After filtering out molecules with poor chemical rationality, candidate molecules are obtained. The ligand library was selected; MACE received the 3D structure of the complex formed by docking irisone with the candidate ligand, and after format conversion, predicted its bond dissociation energy at 25℃, dry environment and triggering condition of 32℃ and 60% humidity. The core candidate with "high bond dissociation energy under normal conditions and significantly reduced bond dissociation energy under triggering conditions" was screened out. The simulation verification confirmed that the ester bond breakage probability at 32℃ was ≥92%, and the release kinetics met the requirements of "slow storage-efficient release". Finally, a temperature-responsive irisone latent aroma body (SMILES: CC1=CC(=O)C(C)(C)C1OOC(CCCCCCCCC)) was formed.
[0033] S4. Using the RDKit tool, the aroma molecules and candidate ligands are docked according to the reaction rules to generate a latent aroma body structure; S5. The MACE model is used to predict the bond dissociation energy of the latent fragrance body structure under normal and triggered conditions, and candidate latent fragrance bodies that meet the requirements of stability and responsiveness are screened out. In a specific implementation, in S5, the MACE model is used to predict the bond dissociation energy of the latent fragrance body. The workflow includes: receiving the latent fragrance body 3D structure generated by the docking of aroma molecules and candidate ligands; predicting the bond dissociation energy of the latent fragrance body 3D structure under normal environment and specific triggered conditions; and screening out latent fragrance bodies with high bond dissociation energy under normal conditions and significantly reduced bond dissociation energy under triggered conditions as core candidates.
[0034] S6. Using a gradient boosting model, predict the release efficiency of the potential fragrance compounds based on their molecular structure characteristics and environmental conditions.
[0035] S3-S6 represent the release rate prediction stage, using XGBoost as the gradient boosting model. The prediction model was constructed using XGBoost by collecting molecular structure information, environmental parameters (25℃ and 32℃, 60% humidity), and corresponding measured release rate data for this type of temperature-responsive latent aroma compound. Molecular structure features were extracted using RDKit, and the environmental parameters were standardized and merged to form a feature matrix. The training, validation, and test sets were divided in a 7:2:1 ratio. After handling missing and outlier values, key features such as ester bond energy, logP, and temperature were selected based on feature importance analysis. Hyperparameters such as a learning rate of 0.04, tree depth of 5, and number of iterations of 180 were set and optimized using grid search. The model was trained using mean squared error as the loss function, employing an early stopping strategy to avoid overfitting. On the test set, the model achieved RMSE=0.028, MAE=0.021, and R0.021. 2 After verifying the generalization ability of 0.94, the newly designed latent fragrance characteristics were input into the model to predict its release rate in cosmetic application scenarios (32℃, 60% humidity). The results showed that the release amount during storage was ≤8%, and the release amount at 32℃ for 1 hour during use was ≥90%, ultimately achieving a precise match between the target fragrance and the controllable release performance.
[0036] Finally, the user interface settings clearly define the input and output. First, the user needs to input a scent description, specifying the desired scent. The interface then matches the corresponding molecular structure and introduces this structure into specific atomic or chemical bond sites of the environmental response group. Next, the user selects environmental response conditions, specifying the response group and optimal site selection. Finally, the user provides specific application data to display a predicted release rate curve for the latent fragrance molecules.
[0037] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for designing latent fragrance bodies and predicting release efficiency.
[0038] This embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for latent aroma body design and release efficiency prediction.
[0039] Example 3 This embodiment provides a method for designing latent aroma components and predicting release efficiency, such as... Figure 3 As shown, it includes the following steps: S1. Using the message passing neural network (MPNN) mechanism, the input aroma description is mapped to the corresponding aroma molecular structure and output as the string SMILES; This embodiment takes animalic aroma profiles as an example, with "rich animalic musk, accompanied by a warm, oily feel and a slight leathery undertone" as the target aroma input. First, it enters the aroma-molecule generation stage, integrating the public database PubChem and proprietary measured aroma-structure data to construct a standardized hierarchical label dataset containing "musky feel, oiliness, leathery notes, and warmth." RDKit is used to perform molecular structure deduplication, standardization, and data augmentation using structural rotation. Subsequently, an MPNN architecture with four modules—aroma embedding, feature mapping, molecular graph generation, and plausibility verification—is employed, and a Transformer encoder is used to transform the target aroma labels. The 256-dimensional structured embedding vector is mapped from aroma features to molecular structure features through MPN encoder and MLP. Generative GNN gradually constructs molecular graphs under chemical bonding rules and syntheticity constraints. After the model is optimized through "pre-training + fine-tuning", it is verified through structural rationality check, aroma matching degree evaluation and DFT calculation, organic synthesis experiment and sensory evaluation closed loop verification to output the target aroma molecule civetone (SMILES: CC1=CC(=O)CCCC1). Its cycloheptenone skeleton and methyl substituent are the core source of rich animal musk. The cyclic structure gives a warm and greasy feeling, while the side chain alkyl chain carries a slight leather background.
[0040] S2. The aroma molecule structure was analyzed using the Retro-MTGR model, and a list of sites with the highest reactivity was selected. S3. Based on the BoltzGen model, combined with the site atom list and the user-defined target release conditions, a candidate ligand molecule map containing controllable bond breaking is generated. S2-S3 is the multi-model collaborative stage: Retro-MTGR receives civet smiles and converts them into molecular diagrams, extracting features such as atom type, hybridization mode, and bond type. Through inference by the atom embedding enhancer and reaction center predictor, the C atom with the highest reactivity on the cycloheptenone ring is selected as the optimal modification site. BoltzGen integrates this site list with the target conditions of "enzyme response in daily chemical scenarios," encodes the environmental conditions into a 64-dimensional vector, and inputs it along with the site features into a conditional variational graph autoencoder to generate 2D molecular diagrams and 3D structures of ligands containing enzyme-responsive glycosidic bonds. Molecules with unreasonable bond types and excessive ring strain are filtered using RDKit to obtain a candidate ligand library; MA CE received the 3D structure of the complex formed by the docking of civetone and the candidate ligand, converted it into the QM9 format, and predicted its bond dissociation energy at normal temperature (25℃), without enzyme and under triggered conditions (25℃) and 0.1 U / mL glycosidase. Core candidates with "bond dissociation energy ≥380 kJ / mol for stable storage under normal conditions and bond dissociation energy ≤220 kJ / mol under triggered conditions" were screened. The GROMACS molecular dynamics simulation confirmed that the probability of glycosidic bond breakage was ≥91% in the presence of enzyme, and the release kinetics met the requirements of "slow release from storage - rapid release triggered by enzyme". Finally, an enzyme-responsive civetone latent aroma body (SMILES: CC1=CC(=O)CCCC1OGlc) was formed.
[0041] S4. Using the RDKit tool, the aroma molecules and candidate ligands are docked according to the reaction rules to generate a latent aroma body structure; S5. The MACE model is used to predict the bond dissociation energy of the latent fragrance body structure under normal and triggered conditions, and candidate latent fragrance bodies that meet the requirements of stability and responsiveness are screened out. In a specific implementation, in S5, the MACE model is used to predict the bond dissociation energy of the latent fragrance body. The workflow includes: receiving the latent fragrance body 3D structure generated by the docking of aroma molecules and candidate ligands; predicting the bond dissociation energy of the latent fragrance body 3D structure under normal environment and specific triggered conditions; and screening out latent fragrance bodies with high bond dissociation energy under normal conditions and significantly reduced bond dissociation energy under triggered conditions as core candidates.
[0042] S6. Using a gradient boosting model, predict the release efficiency of the potential fragrance compounds based on their molecular structure characteristics and environmental conditions.
[0043] S3-S6 represent the release rate prediction stage, using the XGBoost gradient boosting model. The prediction model was constructed using XGBoost by collecting molecular structure information, environmental parameters, and corresponding measured release rate data for this type of enzyme-responsive latent aromasome. 200+ dimensional molecular structure features were extracted using RDKit. Environmental parameters were Z-score standardized and then merged to form a feature matrix. The training, validation, and test sets were divided in a 7:2:1 ratio. Missing values were handled using KNN interpolation, and outliers were removed based on the IQR criterion. Eight key features, including glycosidic bond energy, logP, and enzyme concentration, were selected through SHAP value analysis. Hyperparameters such as a learning rate of 0.05, tree depth of 6, number of iterations of 200, and subsample ratio of 0.8 were set and optimized using Bayesian optimization. The model was trained using mean squared error as the loss function, employing an early stopping strategy to avoid overfitting. On the test set, the model achieved RMSE=0.030, MAE=0.023, and R0.05. 2 After verifying the generalization ability of 0.93, the newly designed latent fragrance features were input into the model to predict its release rate in daily chemical product application scenarios. The results showed that the release amount during storage was ≤7%, and the release amount within 45 minutes of use was ≥88%, ultimately achieving a precise match between the target fragrance and the enzyme-responsive controllable release performance.
[0044] Finally, the user interface settings clearly define the input and output. First, the user needs to input a scent description, specifying the desired scent. The interface then matches the corresponding molecular structure and introduces this structure into specific atomic or chemical bond sites of the environmental response group. Next, the user selects environmental response conditions, specifying the response group and optimal site selection. Finally, the user provides specific application data to display a predicted release rate curve for the latent fragrance molecules.
[0045] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for designing latent fragrance bodies and predicting release efficiency.
[0046] This embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for latent aroma body design and release efficiency prediction.
[0047] Components not described in detail in this embodiment are all existing components that can be purchased through public channels.
[0048] The above description of the embodiments is provided to enable those skilled in the art to understand and use the invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without inventive effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the invention should be within the protection scope of the present invention.
Claims
1. A method for designing latent aroma compounds and predicting their release efficiency, characterized in that, Includes the following steps: S1. Using a message-passing neural network mechanism, the input aroma description is mapped to the corresponding aroma molecular structure and output as the string SMILES. S2. The aroma molecule structure is analyzed using the Retro-MTGR model. The list of sites with the highest reactivity is selected by reasoning through the atom embedding enhancer and reaction center predictor. S3. Based on the BoltzGen model, combined with the site atom list and the user-defined target release conditions, a candidate ligand molecule map containing controllable bond breaking is generated. S4. Using the RDKit tool, the aroma molecules and candidate ligands are docked according to the reaction rules to generate a latent aroma body structure; S5. The MACE model is used to predict the bond dissociation energy of the latent fragrance body structure under normal and triggered conditions, and candidate latent fragrance bodies that meet the requirements of stability and responsiveness are screened out. S6. Using a gradient boosting model, predict the release efficiency of the potential fragrance compounds based on their molecular structure characteristics and environmental conditions.
2. The method for designing latent aroma components and predicting release efficiency according to claim 1, characterized in that, In S1, the message passing neural network mechanism includes the following steps: integrating the aroma tag-molecular structure corresponding dataset and performing standardization and data augmentation; converting aroma tags into structured embedding vectors through a Transformer encoder; mapping aroma features to molecular structure features through an MPN encoder and a multilayer perceptron (MLP); combining chemical bonding rules and syntheticity constraints, using a generative graph neural network (GNN) to gradually generate a molecular graph and output the string "SMILES"; evaluating and optimizing the model based on structural rationality and aroma matching index.
3. The method for designing latent aroma compounds and predicting release efficiency according to claim 1, characterized in that, In S2, the workflow of the Retro-MTGR model includes: receiving the SMILES string of aroma molecules and converting the SMILES string into a molecular diagram; extracting atom type, hybridization mode and bond type features; performing inference through an atom embedding enhancer and a reaction center predictor, and finally outputting a list of the most reactive site atoms.
4. The method for designing latent aroma compounds and predicting release efficiency according to claim 1, characterized in that, In S3, the BoltzGen model is a conditional variational graph autoencoder. The workflow includes: encoding the target release conditions and inputting them into the model along with the features of the site atom list; generating a 2D molecular graph of ligands containing environmentally responsive chemical bonds; and performing chemical rationality filtering on the generated molecular graph to obtain a candidate ligand library.
5. The method for designing latent aroma components and predicting release efficiency according to claim 4, characterized in that, The environmentally responsive chemical bonds include at least one of acetal bonds, ester bonds, or glycosidic bonds.
6. The method for designing latent aroma compounds and predicting release efficiency according to claim 1, characterized in that, In S5, the MACE model is used to predict the bond dissociation energy of latent fragrance compounds. The workflow includes: receiving the 3D structure of latent fragrance compounds generated by docking aroma molecules and candidate ligands; predicting the bond dissociation energy of the 3D structure of latent fragrance compounds under normal conditions and specific triggering conditions; and screening latent fragrance compounds with high bond dissociation energy under normal conditions and reduced bond dissociation energy under triggering conditions as core candidates.
7. The method for designing latent aroma components and predicting release efficiency according to claim 1, characterized in that, In S6, the gradient boosting model is one of XGBoost, LightGBM, GBRT, CatBoost, Histogram-Based GBDT, and NGBoost. The steps of the gradient boosting model to predict release efficiency include: collecting a dataset containing the molecular structure, environmental conditions, and release rate of the latent fragrance; extracting molecular descriptors using RDKit and standardizing the environmental parameters to form a feature matrix; dividing the dataset into training, validation, and test sets; performing feature importance analysis to select key features; setting hyperparameters and training the model; and finally inputting the feature data of the new latent fragrance into the trained model and outputting the predicted release rate value.
8. The method for designing latent aroma compounds and predicting release efficiency according to claim 1, characterized in that, In S3, the target release conditions include at least one of pH value, temperature, oxygen concentration, enzyme concentration, and light intensity.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the method for designing latent fragrance bodies and predicting release efficiency as described in any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method for designing latent aroma compounds and predicting release efficiency as described in any one of claims 1 to 8.