Olfactory threshold prediction method based on multi-dimensional molecular descriptors and multi-layer stacked ensemble model

By constructing a multidimensional molecular descriptor and a multi-layer stacked integrated model, the problems of olfactory threshold determination relying on human experience and the weak generalization ability of traditional models are solved. This achieves high-precision and high-stability prediction of olfactory threshold, improving the accuracy and interpretability of olfactory assessment.

CN122224332APending Publication Date: 2026-06-16TIANJIN ACAD OF ECOLOGICAL & ENVIRONMENTAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610321776.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-17
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Current methods for measuring olfactory threshold rely on human experience, are highly subjective and have poor repeatability, and cannot cover high-boiling-point, difficult-to-vaporize substances and highly toxic substances. Traditional molecular descriptors fail to systematically integrate multi-dimensional information related to olfaction, and existing prediction models have weak generalization ability, poor interpretability, and insufficient predictive consistency.

Method used

A multidimensional molecular descriptor system is constructed and combined with a multi-layer stacked ensemble model. Through hierarchical optimization and weight adaptation mechanisms of heterogeneous base learners and meta-learners, the ability to characterize olfactory response relationships is improved. An olfactory threshold prediction method based on multidimensional molecular descriptors and multi-layer stacked ensemble models is adopted.

Benefits of technology

It significantly improves the accuracy and stability of olfactory threshold prediction, increasing the R2 value from below 60% to above 85%, and greatly reducing RMSE and MAE, providing a reliable and interpretable means of olfactory assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224332A_ABST
    Figure CN122224332A_ABST
Patent Text Reader

Abstract

The application discloses an olfactory threshold prediction method based on a multi-dimensional molecular descriptor and a multi-layer stacked integrated model. It relates to the technical field of olfactory threshold prediction. The method comprises the following steps: step 1, establishing a multi-dimensional molecular descriptor system to extract original features of the odor characteristics of a substance; step 2, pre-processing the original features to obtain a feature subset; step 3, constructing a multi-layer stacked integrated model and training the multi-layer stacked integrated model using the feature subset; and step 4, inputting the original features of the odor characteristics of the substance into the trained multi-layer stacked integrated model to obtain a final olfactory threshold prediction result. The application combines a mechanism-oriented multi-dimensional molecular descriptor system with a multi-layer stacked integrated model suitable for data characteristics, and forms a complete innovative solution from feature engineering to model architecture, thereby providing strong technical support for high-precision and high-reliability olfactory threshold calculation and prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of olfactory threshold prediction technology, and more specifically to an olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model. Background Technology

[0002] Odor pollution, a typical nuisance problem, is increasingly attracting public and environmental protection department attention. Currently, odor activity value analysis based on olfactory thresholds is the main method for identifying key odor-causing substances in complex odors. However, obtaining olfactory thresholds has long relied on manual measurement, and the results are significantly affected by subjective factors such as individual emotions, diet, and environment. Data measured by different institutions often vary considerably. In addition, the difficulty in vaporizing high-boiling-point and highly toxic substances, and the health risks involved, further increase the complexity and uncertainty of experimental measurements. This results in a lack of reliable olfactory threshold data for many odor substances, seriously affecting the accuracy and practical application value of odor activity value analysis.

[0003] To overcome the limitations of experimental measurements, existing research has attempted to establish quantitative prediction models for olfactory thresholds from a molecular structure perspective, such as those based on stereostructure theory or molecular vibrational theory, using partial least squares regression and principal component analysis for fitting. However, these methods often rely on general chemical features in the selection of molecular descriptors, failing to systematically extract multidimensional feature combinations closely related to the olfactory perception mechanism, resulting in the ineffective characterization of key olfactory influencing factors. Furthermore, existing machine learning methods, when applied to olfactory threshold prediction, typically directly use conventional regression models without optimizing the model structure for the high dimensionality, small sample size, and strong nonlinear response characteristics of olfactory data, leading to limited prediction accuracy and a low coefficient of determination (R²) between the true and predicted values. 2 The accuracy is often below 60%, which cannot support the actual needs of high-reliability olfactory assessment.

[0004] Therefore, there is an urgent need to propose an olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model to solve the following problems existing in the prior art: The determination of olfactory threshold relies on human experience, which is highly subjective, has poor repeatability, and cannot cover high-boiling-point non-vaporizable substances and highly toxic substances. Traditional molecular descriptor systems fail to systematically integrate multi-dimensional information such as olfactory physicochemical, topological structure and electronic properties, resulting in the omission of key features. Existing prediction models have not been structurally optimized for the characteristics of olfactory data, and generally suffer from weak generalization ability, poor interpretability, and insufficient prediction consistency. Summary of the Invention

[0005] In view of this, the present invention provides an olfactory threshold prediction method based on multidimensional molecular descriptors and a multi-layer stacked ensemble model. By constructing a multidimensional molecular descriptor system oriented towards the olfactory perception mechanism and combining it with a multi-layer stacked ensemble modeling method adapted to the distribution characteristics of olfactory data, high-precision and high-stability prediction of olfactory thresholds is achieved. The method integrates olfactory-related multi-source descriptors at the feature level and enhances the ability to characterize complex olfactory response relationships at the model level through hierarchical optimization and weight adaptation mechanisms of heterogeneous base learners and meta-learners. This provides a reliable and interpretable computational means for odor identification and control.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: An olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked ensemble model includes: Step 1: Establish a multidimensional molecular descriptor system to extract the original characteristics of the substance's odor properties; Step 2: Preprocess the original features to obtain a feature subset; Step 3: Construct a multi-layer stacked ensemble model and train it using a subset of features; Step 4: Input the original features of the substance's odor characteristics into the trained multi-layer stacked ensemble model to obtain the final olfactory threshold prediction result.

[0007] Preferably, step 1 specifically includes: Based on stereochemistry theory and molecular vibration theory, key molecular behavioral stages that affect olfactory perception are identified, and the scope of descriptor calculation is defined. Based on the scope of descriptor calculation, the descriptor is divided into multiple dimensions, the determining factors of the description are extracted, and the relationship between descriptors in multiple dimensions is established. The original features are obtained by filtering the feature set that is significantly related to the olfactory threshold based on the correlation between descriptors.

[0008] Preferably, multi-dimensional descriptors include one-dimensional descriptors, two-dimensional descriptors, three-dimensional descriptors, quantum chemical descriptors, and derived data descriptors. One-dimensional descriptors are derived directly from molecular formulas and atomic composition; Two-dimensional descriptors are based on molecular topology theory and characterize structural connection patterns through molecular linear input systems, molecular fingerprints, extended connectivity fingerprint symbols, and numerical methods. The three-dimensional descriptor uses molecular mechanics and computational chemistry methods to obtain spatial and interaction parameters, including molar refractive index and the number of hydrogen bond donors and acceptors. Quantum chemical descriptors are derived from density functional theory quantum chemical calculations and are used to quantify the electronic distribution and response characteristics of molecules. Derivative data descriptors are obtained through experiments or predictive models and are directly related to the volatility, solubility, and olfactory behavior of substances.

[0009] Preferably, step 2 specifically includes: The original features are Z-score standardized to eliminate the influence of dimensions; Principal component analysis is applied to perform a linear transformation on the standardized feature matrix to extract principal components. Set a threshold for cumulative variance contribution rate to automatically determine the number of principal components to retain; The original features are projected onto a new orthogonal feature space composed of selected principal components, forming a low-dimensional, redundant, and highly information-condensed feature subset.

[0010] Preferably, the multi-layer stacked ensemble model includes a heterogeneous multi-base learner set and adopts a two-layer stacked ensemble strategy to adaptively fuse it with the multi-base learner set.

[0011] Preferably, the heterogeneous multi-basis learner set includes: a random forest model, a gradient boosting framework, a regularization term, and a kernel function; Random forest models are used to provide global fit and feature importance assessment; gradient boosting frameworks focus on reducing prediction bias through sequential optimization; support vector regression models in kernel functions are used to maximize the interval of small samples.

[0012] Preferably, the process of adaptively fusing a two-layer stacked integration strategy with a multi-base learner set includes: The first layer employs a K-fold cross-validation strategy, using a subset of features as training data and dividing it into K parts. K-1 parts are used sequentially to train a heterogeneous multi-based base learner set, while the remaining 1 part is used to generate the original olfactory prediction values ​​on the validation set. After traversing all folds, each base learner obtains a set of unbiased olfactory prediction values ​​that correspond one-to-one with the original samples on the training set. The olfactory prediction values ​​generated by the heterogeneous multi-based base learner set are used as new features and combined to form a new training dataset. Each row of the new training dataset matrix corresponds to an original sample, and each column corresponds to the prediction output of a base learner. The second layer introduces a lightweight and highly interpretable linear model as a meta-learner, using a new training dataset as input and the original olfactory predictions as output to train the meta-learner.

[0013] Preferably, it also includes: comprehensively evaluating the performance of the multilayer stacked ensemble model on an independent test set by fitting the coefficient of determination, root mean square error, and mean absolute error indices.

[0014] As can be seen from the above technical solution, compared with the prior art, this invention discloses an olfactory threshold prediction method based on multidimensional molecular descriptors and multi-layer stacked ensemble models. By constructing a multidimensional molecular descriptor system oriented towards the olfactory perception mechanism and combining it with a multi-layer stacked ensemble modeling method adapted to the distribution characteristics of olfactory data, high-precision and high-stability prediction of olfactory thresholds is achieved. The method integrates olfactory-related multi-source descriptors at the feature level and improves the ability to characterize complex olfactory response relationships at the model level through hierarchical optimization and weight adaptation mechanisms of heterogeneous base learners and meta-learners, thereby providing a reliable and interpretable computational means for odor identification and control.

[0015] Compared to traditional single regression models (such as PLS, PCR) and simple ensemble methods (such as Bagging, Boosting), this method can predict the R-value of the olfactory threshold. 2 The value was significantly improved from usually below 60% to above 85%, while RMSE and MAE were greatly reduced, which fully verified the effectiveness and innovation of the method in solving the problems of low accuracy and poor consistency in olfactory threshold prediction. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 This invention provides a partial feature importance ranking diagram based on random forest.

[0018] Figure 2 This is a correlation analysis diagram of some features based on random forest provided by the present invention.

[0019] Figure 3 The evaluation results of the three depth models provided by this invention are shown in the figure.

[0020] Figure 4 A comparison chart of true and predicted values ​​provided for this invention.

[0021] Figure 5 The first schematic diagram is a basic data preparation table provided for this invention.

[0022] Figure 6 The second schematic diagram shows the basic data preparation table provided for this invention.

[0023] Figure 7 The third schematic diagram shows the basic data preparation table provided for this invention.

[0024] Figure 8 The fourth schematic diagram shows the basic data preparation table provided for this invention.

[0025] Figure 9 The fifth schematic diagram shows the basic data preparation table provided for this invention.

[0026] Figure 10 The sixth schematic diagram shows the basic data preparation table provided for this invention.

[0027] Figure 11 The seventh schematic diagram shows the basic data preparation table provided for this invention.

[0028] Figure 12 The method flowchart provided by the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] Example 1

[0031] like Figure 12 As shown in the figure, this invention discloses an olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked ensemble model, comprising: Step 1: Establish a multidimensional molecular descriptor system to extract the original characteristics of the substance's odor properties; Step 2: Preprocess the original features to obtain a feature subset; Step 3: Construct a multi-layer stacked ensemble model and train it using a subset of features; Step 4: Input the original features of the substance's odor characteristics into the trained multi-layer stacked ensemble model to obtain the final olfactory threshold prediction result.

[0032] Specifically, step 1 includes: Based on stereochemistry theory and molecular vibration theory, key molecular behavioral stages that affect olfactory perception are identified, and the scope of descriptor calculation is defined. Based on the scope of descriptor calculation, the descriptor is divided into multiple dimensions, the determining factors of the description are extracted, and the relationship between descriptors in multiple dimensions is established. The original features are obtained by filtering the feature set that is significantly related to the olfactory threshold based on the correlation between descriptors.

[0033] Specifically, multi-dimensional descriptors include one-dimensional descriptors, two-dimensional descriptors, three-dimensional descriptors, quantum chemical descriptors, and derived data descriptors. One-dimensional descriptors are derived directly from molecular formulas and atomic composition; Two-dimensional descriptors are based on molecular topology theory and characterize structural connection patterns through molecular linear input systems, molecular fingerprints, extended connectivity fingerprint symbols, and numerical methods. The three-dimensional descriptor uses molecular mechanics and computational chemistry methods to obtain spatial and interaction parameters, including molar refractive index and the number of hydrogen bond donors and acceptors. Quantum chemical descriptors are derived from density functional theory quantum chemical calculations and are used to quantify the electronic distribution and response characteristics of molecules. Derivative data descriptors are obtained through experiments or predictive models and are directly related to the volatility, solubility, and olfactory behavior of substances.

[0034] In a specific embodiment of the present invention, to achieve high-precision and interpretable prediction of olfactory thresholds, the present invention provides a method for constructing a multidimensional molecular descriptor system that integrates olfactory perception mechanisms and computable chemical information. This method establishes an interpretable mapping from molecular structure to olfactory response, and systematically designs, calculates, screens, and optimizes the features of molecular descriptors representing the odor characteristics of substances, providing a well-defined and comprehensive input foundation for subsequent high-performance prediction models. The specific steps are as follows: The first step is to establish the computational goals and scope of descriptors based on olfactory theory. Based on stereochemistry and molecular vibration theory, key molecular behavioral stages affecting olfactory perception are identified, including four stages: environmental release, mediator transport, receptor recognition, and signal triggering. Based on this, molecular characteristic categories to be calculated are established, ensuring that each category directly corresponds to one or more olfactory processes.

[0035] In the environmental release phase, volatility-related characteristics are considered, which determine the initial concentration of molecules in the air and are a prerequisite for olfactory perception. In the medium transport phase, polarity, lipophilicity, and solubility-related characteristics are considered, which are key indicators affecting the diffusion efficiency of molecules in nasal mucus and their interaction with hydrophilic / lipophilic receptor environments. In the receptor recognition phase, three-dimensional spatial and topological characteristics are considered, which determine whether molecules can match the binding pocket of olfactory receptors in terms of spatial shape and surface properties. In the signal triggering phase, electronic structure and chemical bonding characteristics are considered, which affect the strength and specificity of non-covalent interactions (such as hydrogen bonds, dipole-dipole, and π-π stacking) between molecules and receptors.

[0036] This step defines the scope of descriptor computation from the perspective of biological and physicochemical principles, avoiding the randomness of feature selection.

[0037] The second step is to establish a multidimensional descriptor classification framework, referring to Table 1.

[0038] Based on the above property analysis, the descriptor is divided into multiple dimensions, from shallow to deep and from general to detailed.

[0039] One-dimensional descriptors are used to reflect the overall composition and size of molecules, such as molecular weight, number of atoms, and number of functional groups; two-dimensional descriptors are used to characterize the molecular skeleton and connection mode, such as topological index and number of bond types; three-dimensional descriptors are used to describe the spatial configuration and interaction potential of molecules, such as molecular volume, surface area, and electrostatic potential distribution; quantum chemical descriptors are used to reveal electronic structure and reactivity, such as HOMO / LUMO energy and polarizability; and derived data descriptors are used to extract physicochemical parameters related to olfaction, such as saturated vapor pressure and lipophilicity parameters.

[0040] This framework conforms to the cognitive logic from geometric topology to electronic structure and from static parameters to dynamic response, ensuring the systematicity and integrity of the descriptor system in structural representation.

[0041] The third step involves progressively deepening the extraction of elements and the modeling of relationships within the descriptor.

[0042] For each level of descriptor, its determining factors are further extracted, and the relationships between descriptors at each level are established: one-dimensional descriptors are directly derived from molecular formulas and atomic compositions, providing a foundation for subsequent descriptor calculations; two-dimensional descriptors are based on molecular topology theory, using symbolic and numerical methods such as Simplified Molecular Input Line EntrySystem (SMILES), Molecular Access System (MACCS) fingerprints, and Extended-Connectivity Fingerprints (ECFPs) to characterize structural connection patterns; three-dimensional descriptors are based on molecular mechanics and computational chemistry methods to obtain spatial and interaction parameters, including molar refractive index and the number of hydrogen bond donors and acceptors; quantum chemical descriptors are based on quantum chemical calculations such as density functional theory to quantify the electronic distribution and response characteristics of molecules; and derived data descriptors are obtained through experiments or predictive models, directly relating to olfactory-related behaviors such as volatility and solubility of substances.

[0043] The fourth step is to filter and determine the set of key descriptors for olfactory prediction.

[0044] By combining the physicochemical mechanisms of olfactory perception, and considering the availability of descriptors while reducing redundancy, a subset of features significantly related to the olfactory threshold is selected from the above multi-layer descriptors. This set not only covers the multi-scale features of molecular structure, but also focuses on introducing parameters closely related to key processes such as odor molecule volatilization, adsorption, and receptor binding, providing a comprehensive and focused information foundation for subsequent model construction.

[0045] Table 1 Multi-level descriptors

[0046] Through the above four steps, the multidimensional molecular descriptor system constructed in this embodiment not only has systematicity and theoretical basis, but also closely fits the mechanism of olfactory perception, providing a reliable data foundation for overcoming the bottlenecks of insufficient feature representation and poor interpretability in existing prediction models.

[0047] Furthermore, to address the inherent challenges of olfactory data being "high-dimensional, small-sample, and with strong nonlinear responses," a multi-layer stacked ensemble modeling method adapted to the characteristics of olfactory data is proposed. This method significantly improves the prediction accuracy, stability, and generalization ability of the model through a three-level optimization strategy of feature dimensionality reduction, heterogeneous base learner ensemble, and meta-learner adaptive fusion.

[0048] Specifically, step 2 includes: The original features are Z-score standardized to eliminate the influence of dimensions; Principal component analysis is applied to perform a linear transformation on the standardized feature matrix to extract principal components. Set a threshold for cumulative variance contribution rate to automatically determine the number of principal components to retain; The original features are projected onto a new orthogonal feature space composed of selected principal components, forming a low-dimensional, redundant, and highly information-condensed feature subset.

[0049] In a specific embodiment of the present invention, feature dimensionality reduction and information preservation for high-dimensional small sample data includes: To address the "curse of dimensionality" and overfitting issues that high-dimensional descriptors may cause, while preserving olfactory-related feature information to the maximum extent, this invention introduces a dimensionality reduction strategy combining principal component analysis (PCA) with an adaptive threshold for variance contribution rate. Specifically: 1. The original features extracted from the multidimensional molecular descriptor system are Z-score normalized to eliminate the influence of dimensions.

[0050] 2. Apply principal component analysis to perform a linear transformation on the standardized feature matrix and extract principal components.

[0051] 3. Set a cumulative variance contribution rate threshold of ≥95% to automatically determine the number of principal components to retain. This threshold ensures that the feature space after dimensionality reduction can cover most (over 95%) of the variation information in the original data, significantly reducing data dimensionality while avoiding the loss of key olfactory feature information due to excessive dimensionality reduction.

[0052] 4. Project the original high-dimensional features onto a new orthogonal feature space composed of the selected principal components to form a low-dimensional, redundant, and highly information-condensed feature subset, which serves as the input for subsequent model training.

[0053] This step specifically addresses the modeling dilemma caused by the coexistence of high-dimensional and small-sample olfactory data, laying a stable and efficient input foundation for subsequent ensemble learning.

[0054] Specifically, the multi-layer stacked ensemble model includes a heterogeneous set of multi-base learners, and adopts a two-layer stacked ensemble strategy to adaptively fuse it with the set of multi-base learners.

[0055] Specifically, the heterogeneous multi-basis learner set includes: a random forest model, a gradient boosting framework, a regularization term, and a kernel function; Random forest models are used to provide global fit and feature importance assessment; gradient boosting frameworks focus on reducing prediction bias through sequential optimization; support vector regression models in kernel functions are used to maximize the interval of small samples.

[0056] In a specific embodiment of the present invention, in order to fully characterize the complex nonlinear mapping relationship between the olfactory threshold and molecular features, and to avoid the inherent biases and limitations of a single model, a heterogeneous multi-based learner ensemble (StackingEnsemble) is designed, including: By leveraging the feature interaction capture capability and anti-overfitting properties of Random Forest (RF), nonlinearity and coupling relationships in high-dimensional features are addressed. By employing the Gradient Boosting (XGBoost) framework and regularization terms, it efficiently learns complex patterns and exhibits excellent fitting and generalization performance on small sample data. Data is mapped to a high-dimensional space based on kernel functions, and the optimal separating hyperplane is found using support vector regression (SVR).

[0057] The innovation lies in the fact that this embodiment does not simply use multiple models in parallel, but rather strategically combines them based on the complementarity of their principles. Random forest provides global fitting and feature importance assessment; gradient boosting focuses on reducing prediction bias through sequential optimization; and support vector regression maximizes the margin in small sample sizes to improve the model's generalization ability. These three elements work together to model from multiple perspectives, including bias-variance tradeoffs, global-local fitting, and linear-nonlinear mapping, achieving a more complete coverage of the complex shapes of the olfactory response surface.

[0058] Specifically, the process of adaptively fusing a two-layer stacked ensemble strategy with a multi-base learner set includes: The first layer employs a K-fold cross-validation strategy, using a subset of features as training data and dividing it into K parts. K-1 parts are used sequentially to train a heterogeneous multi-based base learner set, while the remaining 1 part is used to generate the original olfactory prediction values ​​on the validation set. After traversing all folds, each base learner obtains a set of unbiased olfactory prediction values ​​that correspond one-to-one with the original samples on the training set. The olfactory prediction values ​​generated by the heterogeneous multi-based base learner set are used as new features and combined to form a new training dataset. Each row of the new training dataset matrix corresponds to an original sample, and each column corresponds to the prediction output of a base learner. The second layer introduces a lightweight and highly interpretable linear model as a meta-learner, using a new training dataset (meta-feature matrix) as input and the original olfactory prediction values ​​as output to train the meta-learner.

[0059] In a specific embodiment of the present invention, the design of a two-layer stacked architecture and adaptive fusion of meta-learners includes: To further integrate the advantages of each base learner and generate a final unified prediction, this invention employs a two-layer stacked integration strategy: The first layer uses a K-fold cross-validation strategy to divide the training data into K parts. K-1 parts are then used to train the three heterogeneous base learners mentioned above, and predictions are generated on the remaining validation set. After iterating through all folds, each base learner will obtain a set of unbiased predictions that correspond one-to-one with the original samples on the training set.

[0060] The predicted values ​​generated by the three base learners are used as new features and combined to form a new training dataset. Each row of this matrix corresponds to an original sample, and each column corresponds to the predicted output of a base learner.

[0061] The second layer introduces a lightweight and highly interpretable linear model (such as ridge regression or elastic networks) as a meta-learner. This meta-learner is trained using the meta-feature matrix as input and the original olfactory threshold as output. The meta-learner's task is to learn how to optimally weight and combine the predictions from the three base learners.

[0062] The core innovation and advantages of this design lie in the fact that the entire stacked architecture is specifically designed for "small sample" data. The K-fold cross-validation method for generating meta-features maximizes the utilization of limited data and effectively prevents the second-layer model from overfitting the predictions of the first layer. The meta-learner automatically learns the prediction reliability of each base learner in different data subspaces through training and assigns optimal weights to achieve adaptive and non-linear fusion of prediction results, which is generally superior to any single base learner or simple voting / averaging strategy. The stacked structure, through the error correction mechanism of the two-layer model, can significantly improve the generalization ability and prediction stability of the final model on unknown samples, effectively addressing the problems of high noise and complex distribution of olfactory data.

[0063] Specifically, this also includes: comprehensively evaluating the performance of multilayer stacked ensemble models on independent test sets by fitting deterministic coefficients, root mean square error, and mean absolute error indices.

[0064] In a specific embodiment of the present invention, after completing the above-described architecture design, a dataset containing multidimensional molecular descriptors and corresponding experimental olfactory thresholds is used to train the entire multi-layer stacked ensemble model according to the aforementioned steps. The model is then fitted with deterministic coefficients (R²). 2 The model performance is comprehensively evaluated on independent test sets using metrics such as root mean square error (RMSE) and mean absolute error (MAE).

[0065] In summary, this embodiment organically combines a mechanism-oriented multidimensional molecular descriptor system with a multi-layer stacked integrated model adapted to data characteristics, forming a complete and innovative solution from feature engineering to model architecture, providing strong technical support for high-precision and high-reliability olfactory threshold calculation and prediction.

[0066] Example 2

[0067] 1. Basic data preparation, a total of 294 sets of data, see [link / reference] Figures 5-11 ; 2. Model development environment selection: Python 3.9.7, Scikit-learn 1.6.1, NumPy 1.20.3, Pandas 1.3.4, Matplotlib 3.4.3; 3. Data missing value imputation: Missing values ​​of LOG P and saturated vapor pressure are imputed using random forest regression. That is, continuous features with missing values ​​are treated as "target variables", other complete features are used as inputs, a random forest regression model is trained, and the model's predicted values ​​are used to fill in the missing values, achieving unbiased imputation; 4. Molecular fingerprint generation and feature engineering enhancement to obtain 167-bit MACCS fingerprints and 1024-bit ECFP4 fingerprints; 5. Combine all features and apply them to the target variable (see olfactory threshold). Figures 5-11 Perform logarithmic transformation; 6. Data segmentation and feature preprocessing: 70% of all data is used as the training set, 15% as the validation set, and 15% as the test set. 7. All continuous features except for molecular fingerprints were standardized, and the final number of original features was determined to be 1205; 8. Random forest was used to rank the features by importance, and 603 features were ultimately selected; after dimensionality reduction using PCA (0.95), 91 features were selected, as shown below. Figures 1-2 As shown; 9. Input the dimensionality-reduced features and target values ​​into the following three models: optimized boosting, deep neural, and stacking ensemble, respectively, and evaluate them. The stacking ensemble model was found to be the best, with the following evaluation results: r2 = 0.73, MAE = 0.23, and RMSE = 0.45. Figure 3 The evaluation results are for the three models.

[0068] 10. For example Figure 4 As shown in Table 2, the true values ​​are the results measured according to the "Three-Point Comparison Odor Bag Method for Determination of Odor in Ambient Air and Exhaust Gas" (HJ 1262—2022), and the predicted values ​​are the model results. The comparison shows that the predicted results trend is consistent with the true values, with an overall deviation of no more than one order of magnitude. Since the results of manual testing of the olfactory threshold inherently have a certain error (determined by human subjective perception), this invention also compared the difference between the predicted value and the three-point comparison odor bag method result with the difference between the three-point comparison odor bag method result and the Japanese test result. It was found that there are also certain differences in performance among different species, but the difference is controlled within one order of magnitude.

[0069] Table 2 Comparison of Test Results

[0070] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0071] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for predicting olfactory thresholds based on a multidimensional molecular descriptor and a multi-layer stacked ensemble model, characterized in that, include: Step 1: Establish a multidimensional molecular descriptor system to extract the original characteristics of the substance's odor properties; Step 2: Preprocess the original features to obtain a feature subset; Step 3: Construct a multi-layer stacked ensemble model and train it using a subset of features; Step 4: Input the original features of the substance's odor characteristics into the trained multi-layer stacked ensemble model to obtain the final olfactory threshold prediction result.

2. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked ensemble model according to claim 1, characterized in that, Step 1 specifically includes: Based on stereochemistry theory and molecular vibration theory, key molecular behavioral stages that affect olfactory perception are identified, and the scope of descriptor calculation is defined. Based on the scope of descriptor calculation, the descriptor is divided into multiple dimensions, the determining factors of the description are extracted, and the relationship between descriptors in multiple dimensions is established. The original features are obtained by filtering the feature set that is significantly related to the olfactory threshold based on the correlation between descriptors.

3. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model according to claim 2, characterized in that, Multi-dimensional descriptors include one-dimensional descriptors, two-dimensional descriptors, three-dimensional descriptors, quantum chemical descriptors, and derived data descriptors. One-dimensional descriptors are derived directly from molecular formulas and atomic composition; Two-dimensional descriptors are based on molecular topology theory and characterize structural connection patterns through molecular linear input systems, molecular fingerprints, extended connectivity fingerprint symbols, and numerical methods. The three-dimensional descriptor uses molecular mechanics and computational chemistry methods to obtain spatial and interaction parameters, including molar refractive index and the number of hydrogen bond donors and acceptors. Quantum chemical descriptors are derived from density functional theory quantum chemical calculations and are used to quantify the electronic distribution and response characteristics of molecules. Derivative data descriptors are obtained through experiments or predictive models and are directly related to the volatility, solubility, and olfactory behavior of substances.

4. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model according to claim 1, characterized in that, Step 2 specifically includes: The original features are Z-score standardized to eliminate the influence of dimensions; Principal component analysis is applied to perform a linear transformation on the standardized feature matrix to extract principal components. Set a threshold for cumulative variance contribution rate to automatically determine the number of principal components to retain; The original features are projected onto a new orthogonal feature space composed of selected principal components, forming a low-dimensional, redundant, and highly information-condensed feature subset.

5. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model according to claim 1, characterized in that, The multi-layer stacked ensemble model includes a heterogeneous set of multi-base learners and adopts a two-layer stacked ensemble strategy to adaptively fuse it with the set of multi-base learners.

6. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model according to claim 5, characterized in that, The heterogeneous multi-base learner set includes: a random forest model, a gradient boosting framework, a regularization term, and a kernel function; Random forest models are used to provide global fit and feature importance assessment; gradient boosting frameworks focus on reducing prediction bias through sequential optimization; support vector regression models in kernel functions are used to maximize the interval of small samples.

7. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model according to claim 6, characterized in that, The process of adaptively fusing a two-layer stacked ensemble strategy with a multi-base learner set includes: The first layer employs a K-fold cross-validation strategy, using a subset of features as training data and dividing it into K parts. K-1 parts are used sequentially to train a heterogeneous multi-based base learner set, while the remaining 1 part is used to generate the original olfactory prediction values ​​on the validation set. After traversing all folds, each base learner obtains a set of unbiased olfactory prediction values ​​that correspond one-to-one with the original samples on the training set. The olfactory prediction values ​​generated by the heterogeneous multi-based base learner set are used as new features and combined to form a new training dataset. Each row of the new training dataset matrix corresponds to an original sample, and each column corresponds to the prediction output of a base learner. The second layer introduces a lightweight and highly interpretable linear model as a meta-learner, using a new training dataset as input and the original olfactory predictions as output to train the meta-learner.

8. The olfactory threshold prediction method based on a multidimensional molecular descriptor and a multi-layer stacked integrated model according to claim 7, characterized in that, Also includes: The performance of the multilayer stacked ensemble model is comprehensively evaluated on an independent test set by fitting deterministic coefficients, root mean square error, and mean absolute error indices.