Method for screening des based on molecular simulation and machine learning to extract tea saponin
By combining molecular simulation and machine learning, molecular models of tea saponin and cellulose were constructed, binding energy and selectivity index were calculated, and machine learning models were trained. This solved the problem of low extraction rate of tea saponin and achieved efficient and economical extraction of tea saponin.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ACAD OF NAT FOOD & STRATEGIC RESERVES ADMINISTRATION
- Filing Date
- 2025-12-25
- Publication Date
- 2026-07-24
Smart Images

Figure CN121768522B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tea saponin extraction technology, and more particularly to a method for screening DES for tea saponin extraction based on molecular simulation and machine learning. Background Technology
[0002] Tea saponin is one of the most valuable functional components in camellia seed cake, widely used in the daily chemical, pharmaceutical, and food industries. Traditional extraction methods often use water, ethanol, or methanol as solvents, supplemented by ultrasound or reflux heating. However, these methods suffer from drawbacks such as low extraction rates, solvent volatility, high energy consumption in subsequent concentration, and difficulty in breaking the cellulose-tea saponin interaction. In recent years, deep eutectic solvents (DES) have emerged as a green alternative solvent due to their advantages, including tunable hydrogen bond networks, wide polarity range, virtually no vapor pressure, and biodegradability. They have already shown potential in the extraction of natural products such as polyphenols and flavonoids.
[0003] However, the designability of the DES system brings many problems: there are many types of hydrogen bond donors and acceptors, their molar ratios are adjustable, and process parameters such as extraction temperature, water content, and stirring rate are interdependent. Relying entirely on experiments is not only time-consuming and material-intensive, but also makes it difficult to reveal the competitive binding mechanism of tea saponin-cellulose. Some studies have attempted to use COSMO-RS or single molecular dynamics (MD) simulations to pre-evaluate solvent performance, but these only remain at the level of static or single-point energy comparison, lacking high-throughput prediction methods for the actual extraction rate, and even more so, failing to directly map the microscopic interaction information obtained from the simulation to the macroscopic extraction effect.
[0004] Furthermore, existing machine learning (MD) work mainly focuses on proving the feasibility of a specific DES; while the application of machine learning in the field of natural product extraction is still in its infancy, generally using shallow models with limited input feature dimensions, making it difficult to fully utilize multi-source heterogeneous information such as DES molecular structure, interaction energy and process conditions, resulting in insufficient prediction accuracy and difficulty in guiding precise formulation design.
[0005] In summary, there is an urgent need in this field for a systematic method to quantitatively describe the competitive binding between tea saponins and cellulose while reducing reagents and experimental quantities, so as to achieve high-throughput screening and optimization of DES formulations and process conditions, thereby breaking through the bottlenecks of low extraction efficiency and high screening costs in existing technologies. Summary of the Invention
[0006] The technical problem to be solved by this invention is to address the shortcomings of existing technologies. Specifically, it provides a method for screening and extracting tea saponins using DES based on molecular simulation and machine learning, as detailed below: 1) In a first aspect, the present invention provides a method for screening and extracting DES for tea saponins based on molecular simulation and machine learning, the specific technical solution of which is as follows: Molecular models of tea saponin, cellulose, hydrogen bond donor, hydrogen bond acceptor, and DES formed by pairing the hydrogen bond donor and the hydrogen bond acceptor were constructed respectively, and the corresponding three-dimensional structure files were obtained. Quantum chemical optimization was performed on all three-dimensional structure files to obtain the optimized structure, molecular polarity index and RESP charge of each molecular model; Based on all optimized structures and the molecular dynamics topology files corresponding to RESP charge generation, a simulation box containing tea saponin molecular models, cellulose molecular models, and DES molecular models was constructed. Based on the initial coordinates and topological information of the simulated box, and combined with molecular dynamics simulation experiments, the trajectory and energy information of the simulated box are determined. Based on the energy information, the binding energy of tea saponin-DES and the binding energy of tea saponin-cellulose are calculated. Based on the binding energy of tea saponin-DES and the binding energy of tea saponin-cellulose, the selectivity index is determined. The measured extraction rate of tea saponin was obtained. Based on the selectivity index, the molecular polarity index corresponding to all molecular models, and the process conditions, an original dataset was constructed. A machine learning model was trained based on the original dataset. The machine learning model was used to predict the extraction rate of tea saponin under different DES and different process conditions. Based on the extraction rate of tea saponin, the optimal DES and the corresponding process conditions were determined, and the extraction was completed.
[0007] The beneficial effects of the method for screening and extracting tea saponins using DES based on molecular simulation and machine learning provided by this invention are as follows: The molecular polarity index, tea saponin-DES binding energy, tea saponin-cellulose binding energy, and selectivity index obtained by joint calculations using quantum chemistry and molecular dynamics, together with the measured extraction rate, form the original dataset. The trained machine learning model can predict the tea saponin extraction rate under unknown DES and process conditions based on the selectivity index and molecular polarity index given by the dataset, thereby completing the selection of the optimal DES and process conditions without increasing the number of experiments.
[0008] Based on the above solution, the present invention can be further improved as follows.
[0009] Furthermore, the measured tea saponin extraction rate was obtained as follows: The extraction of tea saponins from camellia seed cake was performed experimentally using DES solution, and the actual extraction rate of tea saponins was determined. The DES solution is prepared by mixing hydrogen bond donors and hydrogen bond acceptors in a molar ratio of 1:1 to 1:4.
[0010] Furthermore, the hydrogen bond donor is at least one of ethylene glycol, methylurea, glycerol, acetic acid, ammonium acetate, lactic acid, glucose, urea, N,N-dimethylurea, acetamide, sorbitol, or citric acid; and the hydrogen bond acceptor is at least one of choline chloride, betaine, or L-proline.
[0011] Furthermore, the quantum chemical optimization process includes: performing geometry and energy optimization in Gaussian 09 software and at the BLYP / 6-311++G** level to obtain the optimized structure and analytical wave function, and using the Multiwfn program to analyze the wave function to obtain the molecular polarity index and RESP charge.
[0012] Furthermore, the machine learning model is a GMM-DNN coupled model: in the machine learning model, a Gaussian mixture model is first used to pre-cluster the original dataset to obtain the soft assignment probability ω. k The original dataset is then standardized to obtain a standardized feature vector. This standardized feature vector is then compared with the soft assignment probability ω. k The parallel input deep neural network is weighted and fused with K expert networks through a shared layer to output the predicted tea saponin extraction rate.
[0013] Furthermore, the process of standardizing the original dataset includes: The original dataset is divided into a training set and a test set by K-fold cross-validation. The mean and standard deviation of the training set are calculated column by column. Then, the training set and the test set are standardized simultaneously according to the formula X=(x-μ) / σ to obtain the standardized feature vector. Where x is the original value of any column in the original dataset; μ is the arithmetic mean of any column calculated on the training set; σ is the standard deviation of any column calculated on the training set; and X is the value at the corresponding position in the standardized feature vector obtained after standardization.
[0014] Furthermore, the process of determining the optimal DES and the corresponding process conditions is as follows: Among all the DES and corresponding process condition combinations, the one with the highest tea saponin extraction rate predicted by the machine learning model and the smallest prediction error MAE is selected as the optimal DES and corresponding process conditions.
[0015] 2) Secondly, the present invention also provides a system for screening and extracting tea saponins using molecular simulation and machine learning, the specific technical solution of which is as follows: The first construction module is used to: construct molecular models of tea saponin, cellulose, hydrogen bond donor, hydrogen bond acceptor, and DES formed by pairing the hydrogen bond donor and the hydrogen bond acceptor, respectively, and obtain the corresponding three-dimensional structure files; The optimization module is used to perform quantum chemical optimization on all 3D structure files to obtain the optimized structure, molecular polarity index and RESP charge of each molecular model. The second building module is used to: generate corresponding molecular dynamics topology files based on all optimized structures and RESP charges, and build simulation boxes containing tea saponin molecular models, cellulose molecular models and DES molecular models; The determination module is used to: determine the trajectory and energy information of the simulation box based on the initial coordinates and topological information of the simulation box, combined with molecular dynamics simulation experiments; calculate the tea saponin-DES binding energy and tea saponin-cellulose binding energy based on the energy information; and determine the selectivity index based on the tea saponin-DES binding energy and the tea saponin-cellulose binding energy. The extraction module is used to: obtain the measured extraction rate of tea saponin, construct an original dataset by combining the selectivity index, the molecular polarity index corresponding to all molecular models and the process conditions, and train a machine learning model based on the original dataset. The machine learning model is used to predict the extraction rate of tea saponin under different DES and different process conditions. Based on the extraction rate of tea saponin, the optimal DES and the corresponding process conditions are determined to complete the extraction.
[0016] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to perform any of the above methods.
[0017] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above methods.
[0018] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0019] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic flowchart of a method for screening and extracting tea saponins using DES based on molecular simulation and machine learning, according to an embodiment of the present invention. Figure 2(a) is an initial structural schematic diagram of a molecular simulation structure of a method for screening and extracting tea saponins based on molecular simulation and machine learning according to an embodiment of the present invention. Figure 2 (b) is a schematic diagram of the dispersion structure of a molecular simulation structure diagram of a method for screening and extracting tea saponins based on molecular simulation and machine learning according to an embodiment of the present invention. Figure 3 The image shows an HPLC chromatogram of tea saponins from an embodiment of the present invention, which is a method for screening and extracting tea saponins using DES based on molecular simulation and machine learning. Figure 4 This is a schematic diagram of the GMM-DNN model framework for a method of screening and extracting tea saponins based on molecular simulation and machine learning according to an embodiment of the present invention. Figure 5 This is a model representation diagram of a method for screening and extracting tea saponins using molecular simulation and machine learning, according to an embodiment of the present invention. Figure 6 This is a structural framework diagram of an electronic device. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0021] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for screening and extracting tea saponins using DES based on molecular simulation and machine learning, comprising the following steps: S1. Molecular models of tea saponin, cellulose, hydrogen bond donor, hydrogen bond acceptor, and DES formed by pairing the hydrogen bond donor and the hydrogen bond acceptor are constructed respectively to obtain the corresponding three-dimensional structure files. S2, perform quantum chemical optimization on all three-dimensional structure files to obtain the optimized structure, molecular polarity index and RESP charge of each molecular model; S3 generates corresponding molecular dynamics topology files based on all optimized structures and RESP charges, and constructs simulation boxes containing tea saponin molecular models, cellulose molecular models and DES molecular models. S4. Based on the initial coordinates and topological information of the simulation box, and combined with molecular dynamics simulation experiments, determine the trajectory and energy information of the simulation box, calculate the tea saponin-DES binding energy and the tea saponin-cellulose binding energy based on the energy information, and determine the selectivity index based on the tea saponin-DES binding energy and the tea saponin-cellulose binding energy. S5. Obtain the measured tea saponin extraction rate, and construct an original dataset by combining the selectivity index, the molecular polarity index corresponding to all molecular models, and the process conditions. Train a machine learning model based on the original dataset. The machine learning model is used to predict the tea saponin extraction rate under different DES and different process conditions. Based on the tea saponin extraction rate, determine the optimal DES and the corresponding process conditions to complete the extraction.
[0022] The beneficial effects of the method for screening and extracting tea saponins using DES based on molecular simulation and machine learning provided by this invention are as follows: The molecular polarity index, tea saponin-DES binding energy, tea saponin-cellulose binding energy, and selectivity index obtained by joint calculations using quantum chemistry and molecular dynamics, together with process conditions and measured extraction rates, form the original dataset. The trained machine learning model can predict the extraction rate of tea saponin under unknown DES and process conditions based on the selectivity index and molecular polarity index given by the dataset, thereby completing the selection of the optimal DES and process conditions without increasing the number of experiments.
[0023] Example 1: Molecular models of tea saponin (TeaS), cellulose (Cells), hydrogen bond donors, hydrogen bond acceptors, and DES were constructed, and three-dimensional structure files corresponding to each molecular model were generated. The three-dimensional structure file refers to a text file recording the Cartesian coordinates of each atom in the molecule, element type, bond relationships, and residue information. DES hydrogen bond acceptors included choline chloride (ChCl), betaine (Bet), and L-proline (L-Pro), etc., while hydrogen bond donors included ethylene glycol (Eg), methylurea (Met), glycerol (Gly), acetic acid (Ac), ammonium acetate (Aa), lactic acid (La), glucose (Glu), urea (Ur), N,N-dimethylurea (Dme), acetamide (Ace), sorbitol (Sor), and citric acid (CA), etc.
[0024] The molecular model was optimized geometrically and energetically using Gaussian 09 at the BLYP / 6-311++G** level. The calculated wavefunction file was analyzed using the Multiwfn analysis tool to obtain the molecular polarity index (MPI) and RESP charge; the specific process is as follows: After performing geometric and energy optimization, a .chk checkpoint file containing the final optimized structure is output. Then, the formchk tool is used to convert the .chk file into a .fchk text-based wavefunction file, which contains both the optimized structure coordinates and wavefunction information. The Multiwfn analysis tool is then used to read the .fchk file and calculate the molecular polarity index and RESP charge based on the optimized structure.
[0025] Gaussian09, a quantum chemical calculation package released by Gaussian Inc., is used to perform geometric and energy optimization on molecular models of tea saponin, cellulose, hydrogen bond donors, hydrogen bond acceptors, and DES at the BLYP / 6-311++G** level, and outputs wavefunction files required for subsequent Multiwfn analysis. The wavefunction files record the combination coefficients, orbital energy levels, electron densities, and overall wavefunction information for each basis function in the molecular system, which are read by the Multiwfn program for further calculation of the molecular polarity index and RESP charge. BLYP / 6-311++G** refers to the use of BLYP density functional theory combined with 6-311++G** basis sets for geometric optimization and energy calculation.
[0026] The results of geometric optimization were processed using Ambertools to obtain a molecular dynamics topology file, and the charges in the topology file were replaced with RESP charges. An initial molecular dynamics simulation model was constructed using Packmol. Five to ten tea saponin molecules were placed in the center of the initial simulation box, and five to ten cellulose molecules were placed on one side of the tea saponin molecules. Each cellulose molecule contained four repeating units. The initial structures of tea saponin and cellulose are shown below. Figure 2 As shown in the figure. 2000~5000 pairs of DES molecules were added to the above box as a solvent. The size of the box was 10 nm × 10 nm × 10 nm.
[0027] It should be noted that the process of constructing a molecular dynamics topology file includes: Force field type matching is performed on all optimized structures to generate an initial topology. The missing force field parameters in the initial topology are filled in using parmchk2. The GAFF force field is loaded using tleap and the optimized structure and RESP charge are imported in sequence. The RESP charge is written to the corresponding atom using the setCharge command. The prmtop file containing atom type, bond, angle, dihedral angle, LJ parameter and RESP charge is output. This prmtop file is the molecular dynamics topology file.
[0028] The initial coordinates of the initial simulation box and the molecular dynamics topology file (topology information) were used as input files. A molecular dynamics simulation of the tea saponin extraction process was performed using the Gromacs program. The steep descent method was employed to minimize the system energy until the interatomic force was less than 100 kJ / mol / nm. Simulations were performed for 2 ns in the NVT ensemble and 50 ns in the NPT ensemble to bring the system to complete equilibrium. Calculations were then performed for 200 ns in the NPT ensemble at a temperature of 333–373 K and a pressure of 1 bar. The simulation step size was 0.002 ps, and atomic trajectories, energy, and interaction forces were output every 2.0 ps. The steep descent method, also known as the steepest descent algorithm, is an algorithm in Gromacs used for minimizing system energy. It iteratively moves atoms along the negative potential gradient with a variable step size until the interatomic force is less than 100 kJ / mol / nm. -1 ·nm -1 This eliminates poor contact in the initial structure and provides reasonable starting coordinates for subsequent NVT and NPT balancing.
[0029] By analyzing atomic trajectories, energy, and forces, the binding energy E(TeaS-DES) of tea saponin to DES and the binding energy E(TeaS-Cells) of tea saponin to cellulose were obtained. The selective binding index SI=E(TeaS-DES) / E(TeaS-Cells) was calculated. The SI value was used to judge the effect of different DES on separating tea saponin and cellulose. The larger the SI value, the better the extraction effect of DES on tea saponin.
[0030] It should be further explained that the process of obtaining the binding energy E (TeaS-DES) of tea saponin and the binding energy E (TeaS-Cells) of tea saponin is as follows: The tea saponin-DES interaction energies LJ-SR and Coul-SR were extracted separately, and the average tea saponin-DES interaction energy was obtained by averaging the 200 ns trajectory frame by frame. The tea saponin-DES binding energy E(TeaS-DES) was then calculated using the formula ΔE_bind = E_complex - E_TeaS - E_DES, where E_complex is the total energy of the simulation box containing tea saponin and DES, and E_TeaS and E_DES are obtained through independent energy integration. Similarly, the tea saponin-cellulose interaction energy was extracted and the same calculation was performed to obtain the tea saponin-cellulose binding energy E(TeaS-Cells). The resulting E(TeaS-DES) and E(TeaS-Cells) are the binding energies. LJ-SR represents the short-range Lennard-Jones interaction energy, i.e., the contribution of van der Waals interactions between atoms within the system within the cutoff distance. Coul-SR is the Gromacs energy terminology, representing the short-range Coulomb interaction energy, i.e., the contribution of charge-charge interactions between atoms within the system within the cutoff distance. ΔE_bind: The binding energy to be determined; E_complex: The total energy of the simulation box containing tea saponin DES (or tea saponin and cellulose) output by Gromacs; E_TeaS: The total energy obtained by integrating the Gromacs energy when tea saponin is placed alone in the same box; E_DES: The total energy obtained by integrating the Gromacs energy when DES is placed alone in the same box; E_Cells: The total energy obtained by integrating the Gromacs energy when cellulose is placed alone in the same box; LJ-SR: The short-range Lennard-Jones interaction energy given in the Gromacs energy information; Coul-SR: The short-range Coulomb interaction energy given in the Gromacs energy information; 200 ns: The total duration of the molecular dynamics simulation.
[0031] The process of obtaining the measured tea saponin extraction rate includes: To prepare the DES solution, hydrogen bond donors and acceptors were mixed at a molar ratio of 1:1 to 1:4 and stirred at 100 °C and 300 r / min for 120–180 min. 1 g of camellia seed cake sample was weighed and added to 15 mL of DES solution. The mixture was stirred at 60–100 °C and 300 r / min for 30–60 min. The resulting extract was centrifuged at 4000 r / min, and the supernatant was collected. The tea saponin content was determined by HPLC. Figure 3 The figure shows the HPLC chromatogram of tea saponin.
[0032] The original dataset includes: molecular descriptors of DES (molar ratio of hydrogen bond acceptors and hydrogen bond donors, as well as their number of amino groups, number of hydroxyl groups, molecular polarity index, and molecular polar surface area), binding selectivity of tea saponins (ratio of tea saponin-DES binding energy to tea saponin-cellulose binding energy), extraction process conditions (water content, stirring rate, extraction temperature), and the percentage of tea saponins extracted under different extraction conditions (measured tea saponin extraction rate), totaling 13 independent variables and 1 dependent variable.
[0033] Based on machine learning methods, a deep neural network model GMM-DNN is constructed, and the model structure is as follows. Figure 4 As shown. The model was trained using the original dataset to predict the percentage content of tea saponins extracted under different processing conditions, and the most promising DES were screened out, such as... Figure 5 As shown. To give the DNN a better starting point, accelerate convergence, and potentially avoid getting trapped in poor local optima, a Gaussian Mixture Model (GMM) is used as a pre-training tool for the deep neural network (DNN). The input data is first clustered using the GMM, and then the posterior probabilities generated by the clustering results are used to initialize the input and hidden layers of the DNN, providing the DNN with supplementary information about the data structure.
[0034] ①GMM pre-training: The original dataset is split into training and test sets using K-Fold cross-validation. The training data at each fold is standardized, and the same standardization parameters are applied to the test data. The purpose of data standardization is to eliminate the influence of data magnitude and units on the results, thereby improving model accuracy. The formula for data standardization is: Where X represents the standardized data, x represents the original data, and μ and σ are the mean and standard deviation of the original data, respectively.
[0035] GMM calculates the probability that a sample belongs to the k-th Gaussian component. The prior weight of the k-th component, It follows a multivariate Gaussian distribution.
[0036] ②Construction of GMM-DNN model: The model structure was created using Keras, with the input layer containing feature inputs and GMM soft-assigned probability inputs. Shared layers and an expert network were added, the model was compiled, and the optimizer, loss function, and evaluation metric were set. The model was trained on the training set, using early stopping and learning rate decay to prevent overfitting.
[0037] The input layer consists of two tensors: the standardized feature vector X and the soft weights ω output by the GMM.k : In soft weight ω k In the formula, z is the latent variable component label assigned to the input sample x by the Gaussian mixture model, and all inputs extract common features through a shared layer: Where h1 is the output vector of the first hidden layer of the shared layer, W1 is the weight matrix from the input features to the first hidden layer, b1 is the bias vector of the first hidden layer, h2 is the output vector of the second hidden layer of the shared layer, W2 is the weight matrix from the first hidden layer to the second hidden layer, and b2 is the bias vector of the second hidden layer.
[0038] Using K expert networks (ENNs), receiving the shared feature h2 as input, the output of the nth expert network is E. k , in, Let be the output vector of the first hidden layer of the k-th expert network. This is the weight matrix from the second hidden layer of the shared layer to the first hidden layer of the k-th expert network. The weighted fusion output is the bias vector of the first hidden layer of the k-th expert network and the GMM weights. Let this be the output vector of the second hidden layer of the k-th expert network. Let be the weight matrix from the first hidden layer to the second hidden layer of the k-th expert network. Let be the bias vector of the second hidden layer of the k-th expert network; This is the unstandardized predicted value output by the k-th expert network. Let be the weight vector from the second hidden layer to the output layer of the k-th expert network. Let be the bias scalar of the output layer of the k-th expert network; Model's final prediction: in, The standardized prediction extraction rate is the final output of the model. Let x be the predicted value generated by the k-th expert network for input x; The original scale prediction is obtained through inverse standardization: ; The standard deviation of the extraction rate from the training set. This represents the average extraction rate of the training set.
[0039] ③ Model training and optimization The mean squared error loss function is minimized to reduce the difference between the model's predicted values and the actual values.
[0040] This represents the measured true tea saponin extraction rate for the i-th sample. This is the standardized output value of the tea saponin extraction rate predicted by the model for the i-th sample.
[0041] Total loss with regularization: λ is the L2 regularization strength, W l Let be the weight matrix of the l-th layer.
[0042] ④ Cross-validation Make predictions on the test set at each fold and return the unstandardized results. Calculate the MAE and R² of the model at each fold. Evaluate the model performance based on the results of all folds to obtain the optimal model parameters. The model performance is as follows: Figure 4 As shown.
[0043] Based on the above solution, the present invention can be further improved as follows.
[0044] Furthermore, the measured tea saponin extraction rate was obtained as follows: The extraction of tea saponins from camellia seed cake was performed experimentally using DES solution, and the actual extraction rate of tea saponins was determined. The DES solution is prepared by mixing hydrogen bond donors and hydrogen bond acceptors in a molar ratio of 1:1 to 1:4.
[0045] Furthermore, the hydrogen bond donor is at least one of ethylene glycol, methylurea, glycerol, acetic acid, ammonium acetate, lactic acid, glucose, urea, N,N-dimethylurea, acetamide, sorbitol, or citric acid; and the hydrogen bond acceptor is at least one of choline chloride, betaine, or L-proline.
[0046] Furthermore, the quantum chemical optimization process includes: performing geometry and energy optimization in Gaussian 09 software and at the BLYP / 6-311++G** level to obtain the optimized structure and analytical wave function, and using the Multiwfn program to analyze the wave function to obtain the molecular polarity index and RESP charge.
[0047] Furthermore, the machine learning model is a GMM-DNN coupled model: in the machine learning model, a Gaussian mixture model is first used to pre-cluster the original dataset to obtain the soft assignment probability ω. k The original dataset is then standardized to obtain a standardized feature vector. This standardized feature vector is then compared with the soft assignment probability ω. k The parallel input deep neural network is weighted and fused with K expert networks through a shared layer to output the predicted tea saponin extraction rate.
[0048] Furthermore, the process of standardizing the original dataset includes: The original dataset is divided into a training set and a test set by K-fold cross-validation. The mean and standard deviation of the training set are calculated column by column. Then, the training set and the test set are standardized simultaneously according to the formula X=(x-μ) / σ to obtain the standardized feature vector. Where x is the original value of any column in the original dataset; μ is the arithmetic mean of any column calculated on the training set; σ is the standard deviation of any column calculated on the training set; and X is the value at the corresponding position in the standardized feature vector obtained after standardization.
[0049] Furthermore, the process of determining the optimal DES and the corresponding process conditions is as follows: Among all the DES and corresponding process condition combinations, the one with the highest tea saponin extraction rate predicted by the machine learning model and the smallest prediction error MAE is selected as the optimal DES and corresponding process conditions.
[0050] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, and these situations are also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.
[0051] This invention also provides a system for screening and extracting tea saponins using DES based on molecular simulation and machine learning, the specific technical solution of which is as follows: The first construction module is used to: construct molecular models of tea saponin, cellulose, hydrogen bond donor, hydrogen bond acceptor, and DES formed by pairing the hydrogen bond donor and the hydrogen bond acceptor, respectively, and obtain the corresponding three-dimensional structure files; The optimization module is used to perform quantum chemical optimization on all 3D structure files to obtain the optimized structure, molecular polarity index and RESP charge of each molecular model. The second building module is used to: generate corresponding molecular dynamics topology files based on all optimized structures and RESP charges, and build simulation boxes containing tea saponin molecular models, cellulose molecular models and DES molecular models; The determination module is used to: determine the trajectory and energy information of the simulation box based on the initial coordinates and topological information of the simulation box, combined with molecular dynamics simulation experiments; calculate the tea saponin-DES binding energy and tea saponin-cellulose binding energy based on the energy information; and determine the selectivity index based on the tea saponin-DES binding energy and the tea saponin-cellulose binding energy. The extraction module is used to: obtain the measured extraction rate of tea saponin, construct an original dataset by combining the selectivity index, the molecular polarity index corresponding to all molecular models and the process conditions, and train a machine learning model based on the original dataset. The machine learning model is used to predict the extraction rate of tea saponin under different DES and different process conditions. Based on the extraction rate of tea saponin, the optimal DES and the corresponding process conditions are determined to complete the extraction.
[0052] It should be noted that the beneficial effects of the system for screening and extracting tea saponins using DES based on molecular simulation and machine learning provided in the above embodiments are the same as those of the method for screening and extracting tea saponins using molecular simulation and machine learning, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.
[0053] like Figure 6 As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330, which is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned methods. Specifically: The electronic device 300 can vary considerably due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. The memories 310 store at least one computer program 330, which is loaded and executed by the processors 320 to enable the electronic device 300 to implement the method for screening and extracting tea saponins using molecular simulation and machine learning, as described in the above embodiment. Of course, the electronic device 300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. It may also include other components for implementing device functions, which will not be elaborated upon here.
[0054] An embodiment of the present invention provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-described methods.
[0055] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.
[0056] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the methods described above.
[0057] It should be noted that the terms "first" and "second" in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.
[0058] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.
[0059] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0060] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for screening and extracting tea saponins using DES based on molecular simulation and machine learning, characterized in that, include: Molecular models of tea saponin, cellulose, hydrogen bond donor, hydrogen bond acceptor, and DES formed by pairing the hydrogen bond donor and the hydrogen bond acceptor were constructed respectively, and the corresponding three-dimensional structure files were obtained. Quantum chemical optimization was performed on all three-dimensional structure files to obtain the optimized structure, molecular polarity index and RESP charge of each molecular model; Based on all optimized structures and the molecular dynamics topology files corresponding to RESP charge generation, a simulation box containing tea saponin molecular models, cellulose molecular models, and DES molecular models was constructed. Based on the initial coordinates and topological information of the simulated box, and combined with molecular dynamics simulation experiments, the trajectory and energy information of the simulated box are determined. Based on the energy information, the binding energy of tea saponin-DES and the binding energy of tea saponin-cellulose are calculated. Based on the binding energy of tea saponin-DES and the binding energy of tea saponin-cellulose, the selectivity index is determined. The measured extraction rate of tea saponin was obtained. Based on the selectivity index, the molecular polarity index corresponding to all molecular models, and the process conditions, an original dataset was constructed. A machine learning model was trained based on the original dataset. The machine learning model was used to predict the extraction rate of tea saponin under different DES and different process conditions. Based on the extraction rate of tea saponin, the optimal DES and the corresponding process conditions were determined to complete the extraction. The machine learning model is a GMM-DNN coupled model: In this machine learning model, a Gaussian mixture model is first used to pre-cluster the original dataset to obtain the soft assignment probability ω. k The original dataset is then standardized to obtain a standardized feature vector. This standardized feature vector is then compared with the soft assignment probability ω. k The parallel input deep neural network is weighted and fused with K expert networks through a shared layer to output the predicted tea saponin extraction rate.
2. The method for screening and extracting tea saponins using DES based on molecular simulation and machine learning according to claim 1, characterized in that, The measured tea saponin extraction rate was obtained as follows: The extraction of tea saponins from camellia seed cake was performed experimentally using DES solution, and the actual extraction rate of tea saponins was determined. The DES solution is prepared by mixing hydrogen bond donors and hydrogen bond acceptors in a molar ratio of 1:1 to 1:
4.
3. The method for screening and extracting tea saponins using DES based on molecular simulation and machine learning according to claim 1, characterized in that, The hydrogen bond donor is at least one of ethylene glycol, methylurea, glycerol, acetic acid, ammonium acetate, lactic acid, glucose, urea, N,N-dimethylurea, acetamide, sorbitol, or citric acid; the hydrogen bond acceptor is at least one of choline chloride, betaine, or L-proline.
4. The method for screening and extracting tea saponins using DES based on molecular simulation and machine learning according to claim 1, characterized in that, The quantum chemical optimization process includes: using Gaussian 09 software and BLYP / 6-311++G The optimized structure and analytical wave function are obtained by performing geometry and energy optimization at the horizontal level. The molecular polarity index and RESP charge are obtained by analyzing the wave function using the Multiwfn program.
5. The method for screening and extracting tea saponins using DES based on molecular simulation and machine learning according to claim 1, characterized in that, The process of standardizing the original dataset includes: The original dataset is divided into a training set and a test set by K-fold cross-validation. The mean and standard deviation of the training set are calculated column by column. Then, the training set and the test set are standardized simultaneously according to the formula X=(x-μ) / σ to obtain the standardized feature vector. Where x is the original value of any column in the original dataset; μ is the arithmetic mean of any column calculated on the training set; σ is the standard deviation of any column calculated on the training set; and X is the value at the corresponding position in the standardized feature vector obtained after standardization.
6. The method for screening and extracting tea saponins using DES based on molecular simulation and machine learning according to claim 1, characterized in that, The process of determining the optimal DES and the corresponding process conditions is as follows: Among all the DES and corresponding process condition combinations, the one with the highest tea saponin extraction rate predicted by the machine learning model and the smallest prediction error MAE is selected as the optimal DES and corresponding process conditions.
7. A system for screening and extracting tea saponins using molecular simulation and machine learning, employing the method for screening and extracting tea saponins using molecular simulation and machine learning as described in claim 1, characterized in that... include: The first construction module is used to: construct molecular models of tea saponin, cellulose, hydrogen bond donor, hydrogen bond acceptor, and DES formed by pairing the hydrogen bond donor and the hydrogen bond acceptor, respectively, and obtain the corresponding three-dimensional structure files; The optimization module is used to perform quantum chemical optimization on all 3D structure files to obtain the optimized structure, molecular polarity index and RESP charge of each molecular model. The second building module is used to: generate corresponding molecular dynamics topology files based on all optimized structures and RESP charges, and build simulation boxes containing tea saponin molecular models, cellulose molecular models and DES molecular models; The determination module is used to: determine the trajectory and energy information of the simulation box based on the initial coordinates and topological information of the simulation box, combined with molecular dynamics simulation experiments; calculate the tea saponin-DES binding energy and tea saponin-cellulose binding energy based on the energy information; and determine the selectivity index based on the tea saponin-DES binding energy and the tea saponin-cellulose binding energy. The extraction module is used to: obtain the measured extraction rate of tea saponin, construct an original dataset by combining the selectivity index, the molecular polarity index corresponding to all molecular models and the process conditions, and train a machine learning model based on the original dataset. The machine learning model is used to predict the extraction rate of tea saponin under different DES and different process conditions. Based on the extraction rate of tea saponin, the optimal DES and the corresponding process conditions are determined to complete the extraction.
8. An electronic device, characterized in that, The electronic device includes a processor coupled to a memory storing at least one computer program, which is loaded and executed by the processor to enable the electronic device to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the computer to perform the method as described in any one of claims 1 to 6.