Ferulic acid synthesis yield prediction and process optimization method based on machine learning

By constructing a ferulic acid synthesis yield prediction and process optimization model through machine learning, the problem of nonlinear interaction of process parameters in traditional chemical synthesis methods is solved, realizing efficient and low-cost ferulic acid synthesis process optimization, achieving high yield and high efficiency.

CN121983160APending Publication Date: 2026-05-05XIAN UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN UNIV OF SCI & TECH
Filing Date
2026-04-08
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional chemical synthesis methods for optimizing ferulic acid processes rely on time-consuming trial-and-error experiments, which make it difficult to analyze the complex nonlinear interactions between multiple process parameters, resulting in low R&D efficiency and high costs.

Method used

A machine learning approach was used to construct a model for predicting the synthesis yield and optimizing the process of ferulic acid. By collecting experimental data, extracting features, training with various machine learning algorithms, and selecting the optimal model using a Bayesian optimization framework, the SHAP algorithm was combined to analyze the contribution weights of variables, thereby achieving efficient optimization of process parameters.

Benefits of technology

Precisely pinpoint the optimal process range for high yield (>80%), shorten the R&D cycle, improve the efficiency of ferulic acid synthesis, and reduce costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983160A_ABST
    Figure CN121983160A_ABST
Patent Text Reader

Abstract

The invention provides a ferulic acid synthesis yield prediction and process optimization method based on machine learning, and the method comprises the steps: fusing a Morga chemical fingerprint of a molecule with physical process parameters (temperature, time and the like), and constructing a unified feature matrix; bayesian optimization is introduced to automatically search the optimal configuration of 10 algorithms, and the subjectivity of manual parameter adjustment is eliminated; rapid optimization from small sample data to a global optimal point is realized through data flow directions of prediction, verification and feedback; an interpretability analysis is used for revealing a collaborative evolution rule among variables instead of linear superposition of a single variable; according to the method, intelligent optimization normal form transformation of a ferulic acid synthesis process from experience driving to data driving is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of cheminformatics and machine learning, and relates to a method for predicting the synthesis yield and optimizing the process of ferulic acid based on machine learning. Background Technology

[0002] Ferulic acid (FA) is a phenolic compound widely found in plant cell walls. Due to its remarkable antioxidant, anti-inflammatory, antibacterial, and anticancer bioactivities, it has significant applications in the pharmaceutical, food, and cosmetic fields. For example, FA is in high demand in the food industry as a natural preservative, antiseptic, and precursor for vanillin synthesis. With its effects of inhibiting melanin production, combating photoaging, and promoting wound healing, FA is increasingly favored in the research and development of high-end cosmetics such as sunscreens and whitening products. Given its broad industrial demand and diverse bioactivities, developing efficient pathways for obtaining ferulic acid is of significant practical importance.

[0003] Currently, ferulic acid (FA) is mainly obtained through natural extraction, biosynthesis, and chemical synthesis. Natural extraction is limited by long cycles, low yields, and high costs; while biosynthesis has the potential to be green and efficient, it is currently difficult to achieve large-scale industrial application. In contrast, chemical synthesis methods (such as Knoevenagel condensation, Perkin, and Wittig-Horner reactions) have become the main approach for large-scale production due to their advantages of short production cycles and controllable costs. However, traditional chemical process optimization relies heavily on time-consuming trial-and-error experiments, making it difficult to analyze the complex nonlinear interactions between multiple process parameters, resulting in low R&D efficiency. Therefore, there is an urgent need to introduce advanced optimization strategies. Although machine learning-driven methods are receiving increasing attention in multidisciplinary fields, their application in the synthesis process optimization of fine chemicals such as ferulic acid (FA) has not yet been reported.

[0004] The synthesis of ferulic acid involves multiple variables, including the ratio of raw materials (vanillin, malonic acid), the type and amount of catalyst (such as piperazine, aniline), solvent (toluene), reaction temperature, and time. These factors exhibit complex nonlinear coupling relationships, which traditional statistical methods or simple univariate analysis struggle to accurately characterize. Traditional process optimization relies heavily on trial and error based on the experience of experimenters. Finding the optimal yield often requires numerous repetitive experiments, resulting in extremely long development cycles, high consumption of chemical reagents, and high production costs.

[0005] Therefore, in order to effectively address the challenges of multivariate coupling in complex reaction systems such as Knoevenagel condensation, it is urgent to develop a machine learning-based intelligent optimization paradigm for ferulic acid. Summary of the Invention

[0006] To address the technical problems existing in the prior art, this invention provides a method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning.

[0007] According to a first aspect of the present invention, a method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning is provided.

[0008] Specifically, the method includes the following steps: S1. Collect experimental data on ferulic acid synthesis, construct an initial yield dataset, preprocess the experimental data, and extract features to obtain a unified feature matrix; S2. Based on a unified feature matrix, the data is trained using various machine learning algorithms, and parameters are tuned using a Bayesian optimization framework. Multi-dimensional evaluation metrics are used to comprehensively evaluate the model performance, and the model with the best fitting accuracy and generalization ability is selected. S3. Based on the optimal prediction model, predict candidate formulations in the design space, generate high-yield solutions and conduct experimental verification, and feed new data back to the model for iterative retraining. S4. The SHAP algorithm and partial dependency graph are used to analyze the contribution weights of each reaction condition to the yield and the interaction between variables, so as to obtain the optimal experimental parameter range.

[0009] Based on the above scheme, the initial yield dataset mentioned in step S1 includes reactants and process parameters as input features and synthesis yield as the target variable.

[0010] Based on the above scheme, step S1, collecting experimental data on ferulic acid synthesis, constructing an initial yield dataset, preprocessing the experimental data, and extracting features to obtain a unified feature matrix, specifically includes: S101. The system collects experimental sample data covering raw material ratios and process conditions, and uses the target performance index as a monitoring signal. S102. Eliminate dimensional differences in physical parameters through standardization, and at the same time use cheminformatics tools to transform molecular structures into digital representations. S103. Construct a unified multimodal feature matrix by integrating physical parameters and chemical characteristics.

[0011] Based on the above scheme, step S2, which involves training the data using multiple machine learning algorithms based on a unified feature matrix, fine-tuning parameters using a Bayesian optimization framework, comprehensively evaluating model performance using multi-dimensional evaluation metrics, and selecting the model with the optimal fitting accuracy and generalization ability, specifically includes: S201. By comparing the predictive performance of various typical regression algorithms, a basic model performance evaluation framework is established. S202. Use Bayesian optimization strategy to automatically optimize hyperparameters for various algorithms; S203, using the coefficient of determination R 2 The optimal prediction model is selected by using the mean absolute error (MAE) and root mean square error (RMSE).

[0012] Based on the above scheme, step S3, which involves predicting candidate formulations within the design space using the optimal prediction model, generating high-yield schemes and conducting experimental verification, and feeding new data back to the model for iterative retraining, specifically includes: S301. Screening high-yield potential formulations in the preset candidate process space based on the optimal prediction model; S302. Verify the actual synthesis data through experiments and feed it back to the database; S303, rapidly converges to the high-efficiency process range within a limited number of experiments.

[0013] Based on the above scheme, 80 ferulic acid synthesis experimental samples were collected in step S1. In each subsequent round of active learning, 4 new data samples were added. Each sample was described by 8 input features. The experimental parameters included 6 reaction raw materials and 2 process parameters. The target variable was the yield. The reaction raw materials are vanillin, malonic acid, piperazine, aniline, toluene, and potassium carbonate, and the process parameters are reaction temperature and reaction time.

[0014] Based on the above scheme, S102, the dimensional differences of physical parameters are eliminated through standardization, and the molecular structure is transformed into a digital characterization using cheminformatics tools, specifically including: Eight experimental parameters were Z-score normalized and mapped to the [0, 1] interval. The RDKit toolkit was used to convert the SMILES string of the raw materials into a 1024-bit Morgan molecular fingerprint, which was then merged with the normalized reaction features to construct an initial feature matrix containing 1032-dimensional information.

[0015] Based on the above scheme, step S2, which involves training the data using multiple machine learning algorithms based on a unified feature matrix, fine-tuning parameters using a Bayesian optimization framework, comprehensively evaluating model performance using multi-dimensional evaluation metrics, and selecting the model with the optimal fitting accuracy and generalization ability, specifically includes: The study evaluated and compared gradient boosting decision trees, random forests, extreme gradient boosting, support vector regression, adaptive boosting algorithms, category feature boosting algorithms, K-nearest neighbor algorithms, linear regression, decision trees, and ridge regression models. Using the Optuna Bayesian optimization framework, with the goal of minimizing the MAE of 5-fold cross-validation, the study automatically searched for the optimal configuration within the preset parameter space through 50 rounds of iterative search, and selected random forest as the optimal yield prediction model.

[0016] According to a second aspect of the present invention, a computer-readable storage medium is provided for storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the machine learning-based method for predicting the yield and optimizing the process of ferulic acid synthesis.

[0017] According to a third aspect of the present invention, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the machine learning-based method for predicting the yield of ferulic acid synthesis and optimizing the process.

[0018] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention proposes a machine learning-based method for predicting the yield and optimizing the process of ferulic acid synthesis. By constructing an interpretable active learning closed-loop framework, it accurately identifies the optimal process range for achieving high yield (>80%), providing an efficient intelligent driving paradigm for the precise optimization of complex organic synthesis processes.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0021] Figure 1 This is a flowchart illustrating a machine learning-based method for predicting the yield and optimizing the process of ferulic acid synthesis, according to an exemplary embodiment. Figure 2 This is a schematic diagram of the correlation matrix between the output (ferulic acid yield) and the input (vanillin, malonic acid) according to an exemplary embodiment; Figure 3 This is a comparison graph of the predicted and actual yields of ferulic acid using the Random Forest model on the training and test sets, according to an exemplary embodiment. Figure 4A comparison graph showing the predicted and actual yields of ferulic acid using an XGBoost model on the training and test sets, according to an exemplary embodiment. Figure 5 A comparison graph showing the predicted and actual yields of ferulic acid using the AdaBoost model on the training and test sets, according to an exemplary embodiment. Figure 6 A comparison graph showing the predicted and actual yields of ferulic acid using a CatBoost model on the training and test sets, according to an exemplary embodiment. Figure 7 A graph showing a comparison between the predicted and actual yields of ferulic acid using a GBDT model on the training and test sets, according to an exemplary embodiment. Figure 8 A graph showing a comparison between the predicted and actual yields of ferulic acid using an SVR model on the training and test sets, according to an exemplary embodiment. Figure 9 A comparison graph showing the predicted and actual yields of ferulic acid using a Decision Tree model illustrated according to an exemplary embodiment on the training and test sets. Figure 10 A comparison graph showing the predicted and actual yields of ferulic acid using a KNN model on the training and test sets, according to an exemplary embodiment. Figure 11 A graph showing a comparison between the predicted and actual yields of ferulic acid using a Linear Regression model on the training and test sets, according to an exemplary embodiment. Figure 12 A graph showing a comparison between the predicted and actual yields of ferulic acid using the Ridge model on the training and test sets, according to an exemplary embodiment. Figure 13 The models shown in the exemplary embodiment are in the training set R. 2 Comparison chart; Figure 14 The models shown in the test set R are based on an exemplary embodiment. 2 Comparison chart; Figure 15 This is a comparison graph of the various models on the training set MAE according to an exemplary embodiment; Figure 16 This is a comparison graph of the various models on the MAE test set, according to an exemplary embodiment. Figure 17 This is a comparison graph of the RMSE of each model on the training set, according to an exemplary embodiment. Figure 18 This is a comparison graph of the RMSE of each model on the test set, according to an exemplary embodiment. Figure 19 This is a graph showing the comparison between the predicted yield and the actual yield of ferulic acid synthesis through two rounds of active learning, according to an exemplary embodiment. Figure 20 This is a graph showing the comparison between the actual yield value obtained through two rounds of active learning and the initial experimental data in the synthesis of ferulic acid according to an exemplary embodiment; Figure 21 This is a SHAP bar chart of an RF model illustrated according to an exemplary embodiment; Figure 22 This is a SHAP beehive graph of an RF model illustrated according to an exemplary embodiment; Figure 23 This is a partial correlation curve of the two-factor interaction between aniline and piperazine, illustrated according to an exemplary embodiment. Figure 24 This is a partial correlation curve of the two-factor interaction between time and piperazine, according to an exemplary embodiment; Figure 25 This is a partial correlation curve of the two-factor interaction between time and aniline, according to an exemplary embodiment. Figure 26 This is a partial correlation curve of the two-factor interaction between potassium carbonate and piperazine, as illustrated in an exemplary embodiment. Figure 27 This is a partial correlation curve illustrating the two-factor interaction between piperazine and malonic acid, according to an exemplary embodiment. Figure 28 This is a partial correlation curve of the two-factor interaction between toluene and aniline, illustrated according to an exemplary embodiment. Figure 29 This is a partial correlation curve of the two-factor interaction between piperazine and vanillin, as illustrated in an exemplary embodiment. Figure 30 This is a partial correlation curve of the two-factor interaction between aniline and vanillin, as illustrated in an exemplary embodiment. Figure 31 This is a partial correlation curve of the two-factor interaction between toluene and piperazine, illustrated according to an exemplary embodiment. Figure 32 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment. Detailed Implementation

[0022] The following description and accompanying drawings fully illustrate specific embodiments of this application to enable those skilled in the art to practice them. Some parts and features of some embodiments may be included in or replace parts and features of other embodiments.

[0023] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0024] Figure 1 This invention illustrates a machine learning-based method for predicting the yield and optimizing the process of ferulic acid synthesis. Specifically, the method includes the following steps: S1. Collect experimental data on ferulic acid synthesis, construct an initial yield dataset, preprocess the experimental data, and extract features to obtain a unified feature matrix; The initial yield dataset uses reactants and process parameters as input features and synthesis yield as the target variable. Specifically, S101, the system collects experimental sample data covering raw material ratios and process conditions, and uses the target performance index as a monitoring signal; S102. Eliminate dimensional differences in physical parameters through standardization, and at the same time use cheminformatics tools to transform molecular structures into digital representations. S103. Construct a unified multimodal feature matrix by integrating physical parameters and chemical characteristics.

[0025] S2. Based on the multimodal feature matrix, the data is trained using various machine learning algorithms, and the parameters are tuned using a Bayesian optimization framework. Multidimensional evaluation metrics are used to comprehensively evaluate the model performance, and the model with the best fitting accuracy and generalization ability is selected. Specifically, S201, by comparing the predictive performance of various typical regression algorithms, a basic model performance evaluation framework is established; S202. Use Bayesian optimization strategy to automatically optimize hyperparameters for various algorithms; S203, using the coefficient of determination R 2 The optimal prediction model is selected by using the mean absolute error (MAE) and root mean square error (RMSE).

[0026] S3. Based on the optimal prediction model, predict candidate formulations in the design space, generate high-yield solutions and conduct experimental verification, and feed new data back to the model for iterative retraining. Specifically, S301, based on the optimal prediction model, high-yield potential formulations are screened in the preset candidate process space; S302. Verify the actual synthesis data through experiments and feed it back to the database; S303, rapidly converges to the high-efficiency process range within a limited number of experiments.

[0027] S4. The SHAP algorithm and partial dependency graph are used to analyze the contribution weights of each reaction condition to the yield and the interaction between variables, so as to obtain the optimal experimental parameter range.

[0028] This embodiment uses 80 ferulic acid synthesis experimental samples as an example to illustrate the method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning.

[0029] This invention collected a total of 80 ferulic acid synthesis experimental samples (with 4 new samples added in each subsequent round of active learning). Each sample is described by 8 input features, including 6 reaction raw materials (vanillin, malonic acid, piperazine, aniline, toluene, and potassium carbonate) and 2 process parameters (reaction temperature and reaction time). The target variable is the yield (%).

[0030] Frequency distribution analysis of the yields in the experimental samples was conducted to examine the statistical characteristics of the target variable. The results showed that the yield values ​​in the dataset covered a wide range of approximately 30%–100%, indicating that changes in experimental conditions significantly affected the response results. The data distribution exhibited an approximately normal distribution, with most experimental results concentrated in the 60%–90% range, and the peak occurring between 70% and 80%. This continuous and widely distributed yield data provided a solid foundation for constructing a regression model.

[0031] To eliminate the impact of dimensional differences on prediction accuracy, the eight experimental parameters were first Z-score normalized and mapped to the [0, 1] interval to ensure the balance of feature contributions. For molecular characterization, the RDKit toolkit was used to convert the SMILES string of the raw materials into a 1024-bit Morgan molecular fingerprint (radius = 3), which was then merged with the normalized reaction features to construct an initial feature matrix containing 1032 dimensions of information.

[0032] To extract key descriptors from the high-dimensional feature space, this application implements a two-step dimensionality reduction strategy using variance filtering and mutual information screening. First, redundant features with low variance (threshold < 0.01) are removed. Then, mutual information is used to identify the 300 core features most closely related to yield. To further overcome modeling bias caused by uneven data distribution, a hierarchical resampling mechanism is introduced for four yield intervals (<50%, 50%-70%, 70%-85%, 85%-100%). This preprocessing effectively improves data quality, laying a solid foundation for accurate prediction and optimization of the ferulic acid synthesis process.

[0033] After completing feature engineering and dimensionality reduction, such as Figure 2As shown, this invention uses Pearson correlation coefficient (PCC) to analyze the intrinsic relationships between variables and their impact on yield. The results show a significant and strong positive correlation between reactants (vanillin and malonic acid) and the catalyst aniline (r > 0.9), with vanillin and malonic acid exhibiting a completely linear correlation (r = 1.00). This high coupling indicates that the design of experimental parameters (DoE) for vanillin and malonic acid is based on a specific molar ratio. Regarding the correlation between variables and the target yield, piperazine showed the highest positive correlation (r = 0.23), while reaction temperature showed a weak negative correlation (r = -0.17). Further analysis of the interaction effects among key variables revealed a high correlation coefficient (0.95) between 25% K₂CO₃ solution and the catalyst piperazine, indicating a strong synergistic effect between the two in the reaction system. This means that subsequent reaction optimization requires simultaneous adjustment of the ratio of these two factors to avoid interference from single-factor variables. Furthermore, the negative correlation (-0.13) between aniline, as one of the catalysts, and the yield suggests that there may be side reactions interfering with the reaction, which needs to be further verified by subsequent interpretability analysis.

[0034] Based on the aforementioned feature matrix, to select the most suitable prediction algorithm for this reaction system, this invention systematically compared ten classic machine learning regression models. In the Python environment, this invention implemented ten machine learning algorithms using the Scikit-Learn package, including GBDT (Gradient Boosting Decision Tree), RF (Random Forest), XGBoost (Extreme Gradient Boosting), SVR (Support Vector Regression), Adaboost (Adaptive Boosting), CatBoost (Categorical Boosting), KNN (K-Nearest Neighbors), LR (Linear Regression), DT (Decision Tree), and Ridge Regression.

[0035] To fully exploit the predictive potential of each algorithm and eliminate the impact of hyperparameters on model performance, this invention employs the Optuna Bayesian optimization framework for parameter tuning. For each machine learning algorithm, 50 iterations are conducted within a predefined parameter space to automatically search for the optimal hyperparameter configuration, aiming to minimize the mean absolute error (MAE) of 5-fold cross-validation. Compared to traditional grid search methods, Bayesian optimization can explore high-dimensional parameter spaces more efficiently, thus obtaining a better model configuration with limited computational resources. After parameter optimization, to rigorously verify the model's generalization performance and avoid overfitting, the dataset is divided into training and test sets in an 8:2 ratio. Subsequently, this invention uses the coefficient of determination (R²)... 2 The mean absolute error (MAE) and root mean square error (RMSE) were used as performance evaluation metrics. After completing hyperparameter optimization, this invention compared the predictive performance of ten machine learning algorithms, including GBDT, XGBoost, and CatBoost. The fitting results of each model for the yield of ferulic acid synthesis are shown below. Figures 3-12 As shown.

[0036] like Figures 3-12 As shown, the models exhibit significant differences in their performance on the ferulic acid synthesis yield prediction task. Among all evaluated algorithms, Random Forest (RF) demonstrates the best overall performance and robustness, with its test set R0... 2 The R-value is as high as 0.875. In contrast, decision trees (DT) and KNN achieve extremely high fit on the training set (R²). 2 >0.97), but the performance on the test set is significantly reduced (R 2 The value is <0.77, indicating a clear tendency towards overfitting. To further quantify the model performance from a statistical perspective, such as... Figures 13-18 As shown, the system compared the R-values ​​of each model. 2 The results showed that the RF model demonstrated excellent prediction accuracy (R²) on both the training and test sets, along with MAE and RMSE metrics. 2 The approximate values ​​for yield prediction were 0.954 and 0.875, respectively, and the MAE (0.052) and RMSE (0.064) on the test set were the lowest among all models. This indicates that the RF model not only possesses excellent fitting ability but also exhibits superior predictive stability and generalization ability. Based on the above multi-dimensional statistical evaluation, this invention ultimately selected RF as the optimal yield prediction model and used it for subsequent process optimization and interpretability analysis.

[0037] The optimal prediction model, RF, was combined with an active learning strategy. Seven limiting process rules were established (vanillin mass: 1-30 g; vanillin to malonic acid ratio: 1:1.1-1:1.5; reaction temperature: 80-85℃; reaction time: 4-8 h; piperazine amount: 6%-10% of vanillin mass; aniline amount: 4%-6% of vanillin mass; toluene volume: 8-15 times of vanillin mass; 25% potassium carbonate solution volume: 8-12 times of vanillin mass). A design space containing 500 candidate formulations was constructed for iterative screening (iteration terminated after two rounds). High-yield potential formulations were experimentally validated, and the results were fed back into the database for model retraining.

[0038] In the above approach, Active Learning (AL) is a general term for a set of optimal experimental design strategies. It is typically deployed iteratively in conjunction with machine learning (ML) models, with the core objective of improving the model's predictive accuracy. In fields such as chemistry, where data acquisition is costly and limited, Active Learning has proven to be a highly valuable strategy for overcoming data constraints. By combining machine learning-based reaction yield prediction models with experimental design techniques from the Active Learning domain, the model can predict the results of experiments that have not yet been performed, thereby significantly reducing the overall experimental burden.

[0039] This invention utilizes a top-performing random forest (RF) model combined with an active learning (AL) strategy to iteratively search a design space containing 500 candidate formulations to optimize the ferulic acid synthesis process. Two rounds of independent sample experiments (Table 1) validated the results: in the first round, two of the top four recommended schemes achieved yields (87.00% and 85.17%) exceeding the predicted values, confirming the efficiency of the AL strategy in small-sample optimization; after updating the model with the data from the first round, the highest yield reached 82.64% in the second round of validation. Predictive performance analysis (e.g.) Figure 19 As shown in the figure, the RF model has extremely high confidence in the high-yield range (>80%), such as the relative error of only 5% for sample 6. Although the model tends to overestimate in the <75% range due to the imbalance of the dataset classes, the yield distribution plot ( Figure 20 This study confirms that the experimental parameters optimized by AL converge significantly within the efficient range of 70%-88%, effectively avoiding the inefficient region. To further reveal the basis of model decision-making from a chemical mechanism perspective, this invention employs the SHAP algorithm to perform interpretability analysis on the optimized RF model. SHAP interpretability analysis (…) Figure 21 and Figure 22 As shown in the figure, reaction time and piperazine dosage are key factors determining yield, while reaction temperature is insensitive to yield fluctuations within the range under investigation.

[0040] Table 1. Prediction results and experimental verification results of the two rounds of active learning This invention utilizes partial dependence plots (PDPs) to analyze the nonlinear response behavior and optimal operating window of key process parameters. Analysis shows that reaction time exhibits a significant "threshold effect": the SHAP value rises sharply to its peak (>0.25) after 3.5 h, indicating this is the critical point for complete reaction. Piperazine dosage exhibits a non-monotonic "bimodal" positive contribution characteristic (0.4-0.6 g and 1.0-1.2 g), revealing optimal catalytic efficiency within a specific concentration range. Furthermore, aniline content shows extremely high sensitivity in the low-value region (<0.5), with the SHAP value rapidly recovering from -2.0 to 0, corresponding to a significant increase in yield; conversely, the SHAP fluctuation across the reaction temperature range is weak (<0.01), confirming it is not a major controlling factor within the investigated range.

[0041] To reveal the complex coupling mechanisms in nonlinear systems, such as Figures 23-31 As shown, this invention further employs a bivariate partial dependency graph (Bi-PDP) to analyze the interaction effects between features. The results indicate that reaction time and piperazine exhibit the most significant synergistic effect (e.g., Figure 24 As shown in the figure): the high-yield region (>0.75) exhibits a typical "L-shaped" distribution, revealing that a qualitative leap in yield can only be triggered when both simultaneously cross a specific threshold (time > 3.5 h, piperazine > 1.3 g). Furthermore, piperazine demonstrates absolute dominance in multiple interactions (…). Figure 23 and Figure 26 As shown in the figure, the presence of sufficient amounts (>1.5 g) significantly broadens the operating window (0.6-1.4 g) of minor variables such as aniline, establishing the high tolerance of the process; while the interaction gradient of toluene is gentle, indicating that the system is not sensitive to solvent fluctuations. Finally, the high consistency of the sensory properties of the products macroscopically confirms the excellent selectivity and stability of the ML-optimized process.

[0042] This invention successfully constructs an interpretable closed-loop active learning framework, effectively capturing and modeling the complex nonlinear relationships between various variables and yield in the ferulic acid synthesis process. Among ten commonly used machine learning algorithms, the RF model exhibits excellent prediction accuracy and stability, with a test set determination coefficient R0. 2 With a significance level as high as 0.875, SHAP (Shapley Additive explanations) characterization analysis indicated that reaction time and the amount of piperazine catalyst were key control variables affecting the yield of ferulic acid synthesis. Further PDP and Bi-PDP analyses not only quantitatively revealed the synergistic mechanism among the variables but also precisely identified the optimal operating range for achieving high yields (>80%) (e.g., reaction time > 3.5 h, piperazine dosage > 1.3 g).

[0043] In summary, this invention not only efficiently optimizes the experimental synthesis scheme of ferulic acid, but also utilizes artificial intelligence (AI) to assist in the synthesis of traditional fine chemicals, establishing a practical and interpretable new research and development paradigm. Through active learning, it can quickly converge to the high-efficiency reaction zone (70%-88%) with a small number of experimental samples, significantly shortening the research and development cycle.

[0044] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 32 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0045] Those skilled in the art will understand that Figure 32 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0046] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0047] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the method embodiments described above.

[0048] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0049] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.

[0050] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0051] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning, characterized in that, Includes the following steps: S1. Collect experimental data on ferulic acid synthesis, construct an initial yield dataset, preprocess the experimental data, and extract features to obtain a unified feature matrix; S2. Based on a unified feature matrix, the data is trained using various machine learning algorithms, and parameters are tuned using a Bayesian optimization framework. Multi-dimensional evaluation metrics are used to comprehensively evaluate the model performance, and the model with the best fitting accuracy and generalization ability is selected. S3. Based on the optimal prediction model, predict candidate formulations in the design space, generate high-yield solutions and conduct experimental verification, and feed new data back to the model for iterative retraining. S4. The SHAP algorithm and partial dependency graph are used to analyze the contribution weights of each reaction condition to the yield and the interaction between variables, so as to obtain the optimal experimental parameter range.

2. The method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning according to claim 1, characterized in that, The initial yield dataset mentioned in step S1 includes reactants and process parameters as input features, and synthesis yield as the target variable.

3. The method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning according to claim 2, characterized in that, The steps S1, which involve collecting experimental data on ferulic acid synthesis, constructing an initial yield dataset, preprocessing the experimental data, and extracting features to obtain a unified feature matrix, specifically include: S101. The system collects experimental sample data covering raw material ratios and process conditions, and uses the target performance index as a monitoring signal. S102. Eliminate dimensional differences in physical parameters through standardization, and at the same time use cheminformatics tools to transform molecular structures into digital representations. S103. Construct a unified multimodal feature matrix by integrating physical parameters and chemical characteristics.

4. The method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning according to claim 1, characterized in that, The steps in S2, which involve training the data using multiple machine learning algorithms based on a unified feature matrix, fine-tuning parameters using a Bayesian optimization framework, comprehensively evaluating model performance using multi-dimensional evaluation metrics, and selecting the model with the optimal fitting accuracy and generalization ability, specifically include: S201. By comparing the predictive performance of various typical regression algorithms, a basic model performance evaluation framework is established. S202. Use Bayesian optimization strategy to automatically optimize hyperparameters for various algorithms; S203, using the coefficient of determination R 2 The optimal prediction model is selected by using the mean absolute error (MAE) and root mean square error (RMSE).

5. The method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning according to claim 3, characterized in that, The steps S3, which involve predicting candidate formulations within the design space based on the optimal prediction model, generating high-yield solutions and conducting experimental verification, and feeding new data back to the model for iterative retraining, specifically include: S301. Screening high-yield potential formulations in the preset candidate process space based on the optimal prediction model; S302. Verify the actual synthesis data through experiments and feed it back to the database; S303, rapidly converges to the high-efficiency process range within a limited number of experiments.

6. The method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning according to claim 3, characterized in that, In step S1, 80 ferulic acid synthesis experimental samples were collected. In each subsequent round of active learning, 4 new data samples were added. Each sample was described by 8 input features. The experimental parameters included 6 reaction raw materials and 2 process parameters. The target variable was the yield. The reaction raw materials are vanillin, malonic acid, piperazine, aniline, toluene, and potassium carbonate, and the process parameters are reaction temperature and reaction time.

7. The method for predicting the synthesis yield and optimizing the process of ferulic acid based on machine learning according to claim 6, S102, eliminates the dimensional differences of physical parameters through standardization, and simultaneously uses cheminformatics tools to transform the molecular structure into a digital characterization, specifically including: Eight experimental parameters were Z-score normalized and mapped to the [0, 1] interval. The RDKit toolkit was used to convert the SMILES string of the raw materials into a 1024-bit Morgan molecular fingerprint, which was then merged with the normalized reaction features to construct an initial feature matrix containing 1032-dimensional information.

8. The method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning according to claim 4, characterized in that, The steps in S2, which involve training the data using multiple machine learning algorithms based on a unified feature matrix, fine-tuning parameters using a Bayesian optimization framework, comprehensively evaluating model performance using multi-dimensional evaluation metrics, and selecting the model with the optimal fitting accuracy and generalization ability, specifically include: The study evaluated and compared gradient boosting decision trees, random forests, extreme gradient boosting, support vector regression, adaptive boosting algorithms, category feature boosting algorithms, K-nearest neighbor algorithms, linear regression, decision trees, and ridge regression models. Using the Optuna Bayesian optimization framework, with the goal of minimizing the MAE of 5-fold cross-validation, the study automatically searched for the optimal configuration within the preset parameter space through 50 rounds of iterative search, and selected random forest as the optimal yield prediction model.

9. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the steps of the machine learning-based method for predicting the yield and optimizing the process of ferulic acid synthesis as described in any one of claims 1-8.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for predicting the yield and optimizing the process of ferulic acid synthesis based on machine learning, as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent polypeptide synthesis process optimization method, equipment and medium

    CN120278025A

  • Lignocellulose biomass glucose production prediction method based on machine learning

    CN120564867A

  • Biopharmaceutical downstream process optimization method and system based on machine learning

    CN121115680A

  • Method and system for optimizing biomass catalytic pyrolysis process

    CN121483440A