A method and system for screening anthocyanin stabilizers based on the integration of quantum chemistry calculation and machine learning

By integrating quantum chemical computation with machine learning, the problems of high computational resource consumption and low efficiency in anthocyanin stabilizer screening were solved, enabling rapid and efficient screening of stabilizers and providing a basis for targeted design.

CN121963908BActive Publication Date: 2026-05-29PEKING UNIV INST OF ADVANCED AGRI SCI +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIV INST OF ADVANCED AGRI SCI
Filing Date
2026-03-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for screening anthocyanin stabilizers rely on experimental techniques, which suffer from low screening throughput and high computational resource requirements, making it difficult to achieve rapid screening of large-scale candidate polyphenol molecule libraries, resulting in low efficiency in the discovery of anthocyanin stabilizers.

Method used

By employing a method that integrates quantum chemical computation and machine learning, we collected structural data of anthocyanin and polyphenol molecules, performed geometric structure optimization and multidimensional molecular descriptor calculations, constructed a combined free energy prediction model, and screened out candidate molecules for stabilizers.

Benefits of technology

This study enabled rapid screening of anthocyanin stabilizers, reduced computational costs, improved screening efficiency, and revealed key factors in the binding of polyphenols and anthocyanins through feature importance analysis, providing a theoretical basis for the targeted design of stabilizers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963908B_ABST
    Figure CN121963908B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electric data processing, and provides a screening method and system for anthocyanin stabilizer based on quantum chemistry calculation and machine learning fusion. The method comprises the following steps: constructing a molecular structure dataset; performing geometric structure optimization on anthocyanin molecules and candidate polyphenol molecules to obtain ground state optimized structures; performing multidimensional molecular descriptor calculation on the candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix; performing structure optimization, single-point energy calculation and free energy calculation on the anthocyanin polyphenol complex system to obtain a binding free energy numerical set; integrating the polyphenol molecule descriptor matrix and the binding free energy numerical set, constructing a training dataset and performing feature screening; performing model training through a machine learning algorithm to obtain a binding free energy prediction model; and predicting the binding free energy of the anthocyanin polyphenol complex through the binding free energy prediction model to obtain candidate molecules for anthocyanin stabilizer. The present application improves the screening efficiency of anthocyanin stabilizer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electrical data processing technology, and in particular to a method and system for screening anthocyanin stabilizers based on the fusion of quantum chemical computation and machine learning. Background Technology

[0002] Anthocyanins are a class of natural polyphenolic pigments composed of anthocyanins and sugar molecules linked by glycosidic bonds. They are widely found in fruits, flowers, and vegetables and are extensively used as natural pigments in food and beverages due to their excellent coloring properties and nutritional value. However, anthocyanins are easily degraded under the influence of environmental factors such as oxygen, pH, temperature, light, and metal ions, leading to a decline in color and quality. Therefore, finding co-pigment molecules that can stabilize the color of anthocyanins is currently a research hotspot in the field of food science. Studies have shown that polyphenolic compounds can form complexes with anthocyanins through π-π stacking and intermolecular hydrogen bonding, thereby exerting a co-coloring effect. The binding free energy between anthocyanins and polyphenols is considered an important indicator for evaluating the effect of polyphenols in stabilizing anthocyanins.

[0003] Existing methods for screening anthocyanin stabilizers mainly rely on experimental techniques, evaluating polyphenol molecules one by one based on indicators such as color enhancement and thermal stability. These methods suffer from drawbacks such as long experimental cycles and low screening throughput. Although molecular dynamics simulations and quantum chemical calculations can predict binding free energies to some extent, these methods require significant computational resources and are difficult to rapidly screen large-scale candidate polyphenol libraries, resulting in low efficiency in the discovery of anthocyanin stabilizers. Summary of the Invention

[0004] This invention provides a method and system for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning, in order to overcome the shortcomings of existing technologies.

[0005] This invention provides a method for screening anthocyanin stabilizers based on the fusion of quantum chemical computation and machine learning, comprising:

[0006] S1: Collect structural data of anthocyanin molecules and candidate polyphenol molecules, and construct a molecular structure dataset;

[0007] S2: Based on the molecular structure dataset, the geometric structure of anthocyanin molecules and multiple candidate polyphenol molecules is optimized to obtain the first ground state optimized structure and monomer free energy thermal correction for each molecule.

[0008] S3: Based on the descriptor extraction tool, multidimensional molecular descriptor calculation is performed on multiple candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix;

[0009] S4: The second ground-state optimized structure of the anthocyanin polyphenol complex is obtained by configuration search and optimization screening of the molecular structure dataset. The first ground-state optimized structure and the second ground-state optimized structure are calculated by single-point energy calculation. The free energy difference is calculated by combining the free energy thermal correction of the monomer to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins.

[0010] S5: Integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct a training dataset, and perform feature filtering on the training dataset to obtain a preferred feature subset;

[0011] S6: Using the preferred feature subset as input variables and the set of combined free energy values ​​as output variables, a model is trained using a machine learning algorithm to obtain a combined free energy prediction model;

[0012] S7: Input the molecular descriptor of the polyphenol molecule to be screened into the binding free energy prediction model to predict the binding free energy of the anthocyanin polyphenol complex, and screen anthocyanin stabilizer candidate molecules based on the predicted binding free energy.

[0013] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S1 further includes:

[0014] S11: Collect the first SMILES representation of anthocyanin molecules and the second SMILES representation of various candidate polyphenol molecules;

[0015] S12: The first SMILES representation and the second SMILES representation are converted into three-dimensional coordinate files using a molecular format conversion tool, which are then used as structural input files for quantum chemical calculations to construct and obtain a molecular structure dataset.

[0016] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S2 further includes:

[0017] S21: Convert multiple molecular structures in the molecular structure dataset into the input file format of quantum chemical calculation software;

[0018] S22: The IEFPCM solvation model is used to simulate the aqueous environment. Based on preset quantum chemical calculation parameters, the ground-state geometry of anthocyanin molecules is optimized. Under vacuum conditions, the ground-state geometry of multiple candidate polyphenol molecules is optimized based on preset quantum chemical calculation parameters to obtain the first ground-state optimized structure that converges to the minimum point and the corresponding structure checkpoint file. The first ground-state optimized structure includes the ground-state optimized structure of anthocyanin monomers and the ground-state optimized structures of multiple candidate polyphenol molecule monomers.

[0019] S23: Read the frequency calculation output file corresponding to the first ground state optimized structure, perform thermodynamic statistical analysis, and obtain the monomer free energy thermal correction of multiple molecules.

[0020] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S3 further includes:

[0021] S31: Using the second SMILES representation of multiple candidate polyphenol molecules as input, calculate the physicochemical and topological parameters of multiple candidate polyphenol molecules using the RDKit toolkit; the physicochemical parameters include the lipid-water partition coefficient, the exact molecular weight, and the topological polar surface area, and the topological parameters include the number of hydrogen bond acceptors, the number of hydrogen bond donors, the number of rings, the number of aromatic rings, and the number of aliphatic rings.

[0022] S32: The structure checkpoint file is converted into an fchk format file to obtain a quantum chemical structure file. The quantum chemical structure file is then input into the Multiwfn program to perform electrostatic potential surface analysis on multiple candidate polyphenol molecules to obtain an electrostatic surface descriptor. The electrostatic surface descriptor includes the molecular polarity index, molecular volume, minimum electrostatic potential, and maximum electrostatic potential.

[0023] S33: Integrate the physicochemical parameters, the topological parameters, and the electrostatic surface descriptors to obtain a polyphenol molecule descriptor matrix containing multidimensional features.

[0024] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S4 further includes:

[0025] S41: Based on the molecular structures of anthocyanin molecules and candidate polyphenol molecules in the molecular structure dataset, the genmer component in Molclus is used to generate multiple initial configurations of anthocyanin-polyphenol complexes.

[0026] S42: The initial configuration of the anthocyanin polyphenol complex is pre-optimized using the GFNO-xtb method to obtain a first isomer information file; the structures in the first isomer information file are deduplicated and sorted by energy to obtain a first deduplicated and sorted structure set;

[0027] S43: Under the GFN1-xtb combined with the implicit water model, the first deduplication and sorting structure set is batch optimized to obtain the second isomer information file; the structures in the second isomer information file are deduplicated and sorted by energy to obtain the second deduplication and sorting structure set.

[0028] S44: For the preferred anthocyanin polyphenol complex structure in the second deduplication and sorting structure set, the geometric structure optimization and vibration analysis are performed in an aqueous environment by calling Gaussian through Molclus to obtain the second ground state optimized structure that converges to the minimum point. Thermodynamic statistical analysis is then performed on the second ground state optimized structure to obtain the thermal correction of the complex free energy.

[0029] S45: Based on the second ground state optimized structure, perform high-precision single-point energy calculation, sort the calculation results by energy, take the anthocyanin polyphenol complex structure with the lowest energy, and perform free energy calculation in combination with the free energy thermal correction of the complex to obtain the free energy data of the complex.

[0030] S46: Based on the first ground state optimized structure, single-point energy calculations are performed on anthocyanin monomers and multiple candidate polyphenol molecules respectively, and free energy calculations are performed in combination with the corresponding monomer free energy thermal correction to obtain monomer free energy data.

[0031] S47: Calculate the free energy difference based on the monomer free energy data corresponding to the free energy data of the complex to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins.

[0032] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S5 further includes:

[0033] S51: Align and integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct an original dataset for machine learning;

[0034] S52: Calculate the Pearson correlation coefficient between multiple feature descriptors in the original dataset and between each feature descriptor and the binding free energy to generate a correlation heatmap in order to remove redundant features and obtain a clean dataset.

[0035] S53: The Z-score normalization method is used to normalize all feature variables in the cleaned dataset to obtain a standardized dataset;

[0036] S54: Divide the standardized dataset into a training set and a test set. For the training set, use recursive feature elimination, feature importance ranking based on random forest, and SelectKBest method based on statistical test to select features. Then, use principal component analysis to reduce the dimensionality of features and obtain the preferred feature subset.

[0037] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S6 further includes:

[0038] S61: Using the preferred feature subset as input and the combined free energy numerical set as output, various machine learning algorithms are introduced to obtain multiple algorithm models to be trained.

[0039] S62: The hyperparameters of multiple algorithm models are continuously optimized by using grid search combined with 5-fold cross-validation to obtain the optimal hyperparameter configurations for multiple algorithm models;

[0040] S63: Select evaluation metrics, evaluate the performance of multiple algorithm models and combinations of multiple algorithm models on the test set based on multiple optimal hyperparameter configurations, and select the combined free energy prediction model.

[0041] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided. In step S61, the machine learning algorithm includes: linear regression, ridge regression, lasso regression, elastic network, decision tree, random forest, gradient boosting, support vector regression, K-nearest neighbors and neural network.

[0042] According to the present invention, a method for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning is provided, wherein step S7 further includes:

[0043] S71: Calculate multidimensional molecular descriptors for the polyphenol molecules to be screened to obtain descriptor data for the samples to be predicted;

[0044] S72: The descriptor data of the sample to be predicted is processed by Z-score normalization and feature processing, and then input into the binding free energy prediction model to obtain the predicted binding free energy between multiple polyphenol molecules to be screened and anthocyanins.

[0045] S73: Based on the magnitude of the negative value of the predicted binding free energy, sort the polyphenol molecules to be screened, and select anthocyanin stabilizer candidate molecules from the polyphenol molecules with the largest negative binding free energy.

[0046] This invention also provides an anthocyanin stabilizer screening system based on the fusion of quantum chemical computation and machine learning, for performing an anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning as described in any of the above claims, comprising:

[0047] Collection module: Used to collect structural data of anthocyanin molecules and candidate polyphenol molecules, and construct a molecular structure dataset;

[0048] Optimization module: Based on the molecular structure dataset, it is used to optimize the geometric structure of anthocyanin molecules and multiple candidate polyphenol molecules respectively, and obtain the first ground state optimized structure and monomer free energy thermal correction amount corresponding to each molecule;

[0049] Extraction module: Used to perform multidimensional molecular descriptor calculations on multiple candidate polyphenol molecules based on descriptor extraction tools, and obtain a polyphenol molecule descriptor matrix;

[0050] The calculation module is used to obtain the second ground-state optimized structure of the anthocyanin-polyphenol complex from the molecular structure dataset through configuration search and optimization screening, and to perform single-point energy calculations on the first ground-state optimized structure and the second ground-state optimized structure respectively. Combined with the free energy thermal correction of the monomer, the free energy difference is calculated to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins.

[0051] The filtering module is used to integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct a training dataset, and to perform feature filtering on the training dataset to obtain a preferred feature subset.

[0052] Training module: used to train the model using the preferred feature subset as input variable and the set of combined free energy values ​​as output variable, through machine learning algorithms, to obtain a combined free energy prediction model;

[0053] The prediction module is configured with the binding free energy prediction model obtained by the training module. It is used to receive molecular descriptors of polyphenol molecules to be screened, predict the binding free energy of anthocyanin polyphenol complexes, and screen anthocyanin stabilizer candidate molecules based on the predicted binding free energy.

[0054] This invention deeply integrates quantum chemical computation with machine learning algorithms to construct a prediction system that takes a polyphenol molecule descriptor matrix as input and binding free energy as output. This eliminates the need to repeatedly perform the computationally expensive quantum chemical process for each anthocyanin-polyphenol pair. Instead, the descriptor of the candidate molecule is extracted and input into the binding free energy prediction model to complete the screening and judgment. This fundamentally breaks through the bottleneck of traditional one-by-one experimental screening and the excessive consumption of large-scale quantum chemical computational resources.

[0055] The present invention also provides an anthocyanin stabilizer screening device based on the fusion of quantum chemical calculation and machine learning, comprising: a memory and at least one processor, wherein the memory stores instructions; the processor invokes the instructions in the memory to cause the anthocyanin stabilizer screening device based on the fusion of quantum chemical calculation and machine learning to perform an anthocyanin stabilizer screening method as described above.

[0056] The present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement an anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning as described in any of the preceding claims.

[0057] At the algorithmic level, this invention introduces various feature engineering techniques, such as recursive feature elimination, feature importance ranking based on random forest, SelectKBest statistical test, and principal component analysis, to extract optimal feature subsets from the polyphenol molecule descriptor matrix. Furthermore, it combines grid search and 5-fold cross-validation to optimize the hyperparameters of ten machine learning algorithms. This results in a final binding free energy prediction model that achieves a good balance between generalization ability and prediction accuracy, fully demonstrating the contribution of the selection and combination of algorithm features to prediction performance. Especially in the specific application scenario of anthocyanin stabilizer screening, the multidimensional combination of physicochemical parameters, topological parameters, and electrostatic surface descriptors serves as the feature input for the machine learning model. This enables the model not only to output predicted binding free energy values ​​but also to reveal key molecular structural factors affecting the binding of polyphenols and anthocyanins through feature importance analysis. This provides an interpretable theoretical basis for the targeted molecular design of anthocyanin stabilizers, effectively implementing artificial intelligence algorithms in the field of food natural pigment stabilizer screening and fully releasing their application value. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0059] Figure 1 A schematic diagram of a screening method for anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning, provided in an embodiment of the present invention;

[0060] Figure 2 A schematic diagram of an anthocyanin stabilizer screening system based on the fusion of quantum chemical computation and machine learning is provided in an embodiment of the present invention.

[0061] Figure 3 This is a schematic diagram of the distribution of binding free energy provided in an embodiment of the present invention.

[0062] Figure 4 The heatmap obtained by applying the Pearson correlation coefficient to data sample values ​​is provided in an embodiment of the present invention.

[0063] Figure 5 A schematic diagram of the model performance obtained by the anthocyanin stabilizer screening method based on the fusion of quantum chemical calculation and machine learning provided in the embodiments of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0065] The embodiments of the present invention are described below with reference to the figures.

[0066] like Figure 1 As shown, this invention provides a method for screening anthocyanin stabilizers based on the fusion of quantum chemical computation and machine learning, comprising:

[0067] S1: Collect structural data of anthocyanin molecules and candidate polyphenol molecules to construct a molecular structure dataset.

[0068] Step S1 further includes:

[0069] S11: Collects the first SMILES representation of anthocyanin molecules and the second SMILES representation of various candidate polyphenol molecules.

[0070] Furthermore, in a specific embodiment, the present invention uses cyanidin-3-O-glucoside chloride as the target anthocyanin molecule and 20 polyphenol molecules, including ferulic acid, syringic acid, p-coumaric acid, cinnamic acid, phenylpropionic acid, isochlorogenic acid, veracic acid, mycophenolic acid, kaempferol, daidzein, catechin, baicalein, quercetin, rutin, isoquercetin, tannic acid, caffeic acid, phenylacetic acid, salvianolic acid, and quercetin, as candidate stabilizers. In step S11, the present invention first extracts the SMILES (Simplified Molecular Input Line Entry System) of the above molecules, that is, the molecular structure is encoded as a string representation. The SMILES of the anthocyanin molecule is denoted as the first SMILES, and the SMILES of each of the 20 candidate polyphenol molecules is denoted as the second SMILES.

[0071] S12: The first SMILES representation and the second SMILES representation are converted into three-dimensional coordinate files using a molecular format conversion tool, which are then used as structural input files for quantum chemical calculations to construct and obtain a molecular structure dataset.

[0072] In step S12, the present invention uses a molecular format conversion tool to parse the first SMILES representation and the second SMILES representation into .xyz format files containing the three-dimensional spatial coordinates of each atom. The .xyz file records the initial three-dimensional conformation of the molecule in a row-by-row manner with atomic element symbols and their x, y, and z coordinates. Each of the 21 molecules corresponds to an independent .xyz coordinate file, and all files are compiled to construct a molecular structure dataset.

[0073] S2: Based on the molecular structure dataset, the geometric structure of anthocyanin molecules and multiple candidate polyphenol molecules is optimized to obtain the first ground state optimized structure and monomer free energy thermal correction for each molecule.

[0074] Step S2 further includes:

[0075] S21: Convert multiple molecular structures in the molecular structure dataset into the input file format of quantum chemistry calculation software.

[0076] Furthermore, in step S21, the present invention converts the .xyz coordinate files of each molecule in the molecular structure dataset into the .gjf input file format required by the Gaussian quantum chemistry calculation software. In addition to containing atomic coordinate information, the .gjf file also contains key parameters such as the calculation task type, functional, basis set, and solvent model.

[0077] S22: The IEFPCM solvation model is used to simulate the aqueous environment. Based on preset quantum chemical calculation parameters, the ground-state geometry of anthocyanin molecules is optimized. Under vacuum conditions, the ground-state geometry of multiple candidate polyphenol molecules is optimized based on preset quantum chemical calculation parameters to obtain the first ground-state optimized structure that converges to the minimum point and the corresponding structure checkpoint file. The first ground-state optimized structure includes the ground-state optimized structure of anthocyanin monomers and the ground-state optimized structures of multiple candidate polyphenol molecule monomers.

[0078] After the input file format conversion is completed in S21, the present invention performs ground-state geometric structure optimization on anthocyanin molecules and candidate polyphenol molecules using different solvent environment settings.

[0079] For anthocyanin molecules, this invention uses the IEFPCM solvation model to simulate the aqueous environment. IEFPCM treats the solvent as a continuous polarization medium and describes the effect of the solvent on the electronic structure of anthocyanins by establishing a polarization charge distribution on the surface of the anthocyanin molecule. Under this solvation condition, Gaussian software performs ground-state geometry optimization on the anthocyanin molecule with preset quantum chemical calculation parameters. The electronic structure is solved by SCF iteration and the forces on each atom are calculated. The atomic coordinates are iteratively updated along the negative gradient direction until the energy change and atomic forces converge to below the set threshold, thus obtaining the ground-state optimized structure of the anthocyanin monomer and the corresponding structure checkpoint file.

[0080] For 20 candidate polyphenol molecules, this invention performs ground-state geometric structure optimization on each candidate polyphenol molecule under vacuum conditions, i.e. without introducing any solvent model, using the same preset quantum chemical calculation parameters. The optimization process also drives the atomic coordinates to converge to the potential energy surface minimum point through SCF iteration and gradient descent. Each candidate polyphenol molecule outputs a candidate polyphenol molecule monomer ground-state optimized structure and the corresponding structure checkpoint file.

[0081] Meanwhile, the Gaussian software generates a binary-formatted structure checkpoint file (.chk file) for each optimization task, recording all quantum chemical information of the molecule at the minimum point, such as wave function, electron density, and energy, which corresponds to the subsequent step S32 and can be read and used by the subsequent S32 step.

[0082] S23: Read the frequency calculation output file corresponding to the first ground state optimized structure, perform thermodynamic statistical analysis, and obtain the monomer free energy thermal correction of multiple molecules.

[0083] After optimizing the ground-state geometry of each monomer in step S22, the frequency calculation output file (.out file) from the Gaussian software is input into the Shermo program. After input, the Shermo program reads the vibrational frequency data, zero-point vibrational energy, and thermodynamic statistical information from the file. Under standard temperature and pressure conditions, the translational, rotational, and vibrational partition functions of each molecule are calculated according to statistical thermodynamics. The corresponding free energy thermal correction values ​​are extracted from the frequency calculation output files of each monomer. One free energy thermal correction value is output for each monomer. Each of the 20 candidate polyphenol monomers and anthocyanin monomers corresponds to one free energy thermal correction value, which is used for the monomer free energy calculation in step S4.

[0084] S3: Based on the descriptor extraction tool, multidimensional molecular descriptor calculations are performed on multiple candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix.

[0085] Step S3 further includes:

[0086] S31: Using the second SMILES representation of multiple candidate polyphenol molecules as input, calculate the physicochemical and topological parameters of multiple candidate polyphenol molecules using the RDKit toolkit; the physicochemical parameters include the lipid-water partition coefficient, the exact molecular weight, and the topological polar surface area, and the topological parameters include the number of hydrogen bond acceptors, the number of hydrogen bond donors, the number of rings, the number of aromatic rings, and the number of aliphatic rings.

[0087] In step S21, this invention uses the second SMILES representation of each of the 20 candidate polyphenol molecules as direct input. Using the RDKit toolkit, the molecular topology is parsed from the SMILES string, and molecular properties are calculated. Specifically, physicochemical and topological parameters are calculated for each polyphenol molecule. The physicochemical parameters include the lipid-water partition coefficient (logP, reflecting the molecule's hydrophilicity / hydrophobicity), exact molecular weight, and topological polar surface area (TPSA, reflecting the total surface area of ​​polar atoms). The topological parameters include the number of hydrogen bond acceptors, hydrogen bond donors, rings, aromatic rings, and aliphatic rings. All these parameters are directly calculated from the molecular bonding relationships described in the SMILES string. Each polyphenol molecule outputs a parameter record containing eight values ​​on a single line.

[0088] S32: The structure checkpoint file is converted into an fchk format file to obtain a quantum chemical structure file. The quantum chemical structure file is then input into the Multiwfn program to perform electrostatic potential surface analysis on multiple candidate polyphenol molecules to obtain electrostatic surface descriptors. The electrostatic surface descriptors include molecular polarity index, molecular volume, minimum electrostatic potential, and maximum electrostatic potential.

[0089] Furthermore, this invention converts the .chk format structure checkpoint files corresponding to each polyphenol molecule in S22 into readable fchk (formatted checkpoint) format files using the formchk command. The fchk file stores the electron density matrix and wave function coefficients of the molecule at the minimum point in text form. Subsequently, this invention inputs the fchk file into the Multiwfn program, and inputs the commands 12 and 0 sequentially in the program interface. Multiwfn calculates the electrostatic potential based on the distribution of electron density on the van der Waals surface of the molecule, and extracts the electrostatic surface descriptor from the electrostatic potential distribution, including the molecular polarity index (MPI), molecular volume, minimum electrostatic potential, and maximum electrostatic potential. Each polyphenol molecule outputs a record of the electrostatic surface descriptor containing four values ​​on a single line.

[0090] S33: Integrate the physicochemical parameters, the topological parameters, and the electrostatic surface descriptors to obtain a polyphenol molecule descriptor matrix containing multidimensional features.

[0091] In step S33, the present invention merges the 8 physicochemical parameters and topological parameters of each candidate polyphenol molecule and the 4 electrostatic surface descriptor values ​​by column to form a two-dimensional numerical matrix in which each row corresponds to one polyphenol molecule and each column corresponds to one descriptor dimension, namely the polyphenol molecule descriptor matrix.

[0092] S4: Based on the molecular structure dataset, a second ground-state optimized structure of the anthocyanin-polyphenol complex is obtained through configuration search and optimization screening. Single-point energies are then calculated for both the first and second ground-state optimized structures. Combined with the monomer free energy thermal correction, the free energy difference is calculated to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins. Step S4 further includes:

[0093] S41: The molecular structures of anthocyanin molecules and candidate polyphenol molecules in the molecular structure dataset are used to generate multiple initial configurations of anthocyanin-polyphenol complexes using the genmer component in Molclus.

[0094] Furthermore, in step S41, this invention uses the molecular structures and corresponding three-dimensional geometric coordinates of anthocyanin molecules and candidate polyphenol molecules in the molecular structure dataset as input, and calls the genmer component in the Molclus software to perform spatial combination sampling of the geometric coordinates of anthocyanins and each candidate polyphenol molecule. Subsequently, the genmer component generates anthocyanin-polyphenol complex configurations covering various spatial arrangements by systematically rotating and translating the relative positions and orientations of the two molecules. Finally, it generates a batch of initial anthocyanin-polyphenol complex configurations for each of the 20 candidate polyphenol molecules. Each batch of initial configurations is output as a set of coordinate files as input for the pre-optimization task in S42.

[0095] S42: Under the GFN0-xtb method, the initial configuration of the anthocyanin polyphenol complex is pre-optimized to obtain the first isomer information file; the structures in the first isomer information file are deduplicated and sorted by energy to obtain the first deduplicated and sorted structure set.

[0096] In step S42, the present invention uses Molclus software to call the xtb program in batches to pre-optimize the initial configurations of all anthocyanin polyphenol complexes obtained in S41 one by one using the GFNO-xtb semi-empirical quantum chemistry method. The GFNO-xtb method rapidly optimizes molecular configurations through parameterized Hamiltonian and is suitable for rapid energy screening of a large number of initial configurations.

[0097] During optimization, the GFN0-xtb method uses a parameterized zero-order Hamiltonian to describe the electronic structure of the molecular system. It does not explicitly calculate inter-electron interactions, but instead rapidly estimates the system energy and drives geometric optimization through analytical functions of inter-atomic distances and bond order parameters. Specifically, for each initial configuration, the xtb program calculates the forces acting on each atom under that configuration, iteratively adjusting the coordinates of each atom along the negative gradient direction until all force components acting on all atoms are below a set convergence threshold. At this point, optimization convergence is considered complete, and the optimized coordinates and corresponding energy of the configuration are output. Because GFN0-xtb has extremely low computational cost, it is suitable for handling a large number of initial configurations generated by S41, quickly eliminating high-energy, unreasonable configurations.

[0098] Then, Molclus sequentially writes the structure coordinates and corresponding energies of each initial configuration after optimization into the same first isomer.xyz file, i.e., the first isomer information file. This file stores all optimized structures in a multi-frame xyz format. Subsequently, the present invention uses the isostat program to read the first isomer.xyz file, performs deduplication on all frames according to the structure similarity threshold, removes structures with duplicate configurations, and sorts the remaining structures in order of energy from low to high, outputting the first deduplicated sorted structure set.

[0099] S43: Under the GFN1-xtb combined with the implicit water model, the first deduplication and sorting structure set is batch optimized to obtain the second isomer information file; the structures in the second isomer information file are deduplicated and sorted by energy to obtain the second deduplication and sorting structure set.

[0100] In step S43, the present invention uses Molclus software to call the xtb program in batches, and uses the GFN1-xtb method combined with the implicit water model to optimize all the structures in the first deduplication and sorting structure set obtained in S42 one by one. The GFN1-xtb is more accurate in describing electronic structure than the aforementioned GFN0-xtb. Combined with the implicit water model, the influence of solvent effect on the configuration of the complex can be introduced with low computational cost.

[0101] In S43, the GFN1-xtb method used in this invention introduces a more complete electronic structure description based on the GFN0-xtb method, including a self-consistent charge (SCC) iterative process. This involves iteratively solving for the system's charge distribution until self-consistent convergence, then calculating the system's energy and atomic forces based on the converged charge distribution. Subsequently, geometry optimization is driven to convergence using the same gradient descent iterative method as in S42. Combined with the implicit water model, the xtb program introduces an additional solvation contribution term in each energy and gradient calculation step, ensuring that the optimized configuration reflects the influence of solvation effects in aqueous solution on the complex's geometry. After two rounds of deduplication and sorting in S42 and S43, the number of configurations has been significantly reduced from an initial large batch to at least a few low-energy representative structures.

[0102] Subsequently, Molclus writes the optimized structure coordinates and corresponding energies of each frame into the second isomer.xyz file. Then, the isostat program is used to deduplicate and sort the structures in the second isomer.xyz file again, outputting the second deduplicated and sorted structure set. The structures in this set have undergone two rounds of progressive refinement and screening, with lower energies and non-repeating configurations.

[0103] S44: For the preferred anthocyanin polyphenol complex structure in the second deduplication and sorting structure set, geometric structure optimization and vibration analysis are performed in an aqueous environment using Molclus and Gaussian to obtain the second ground state optimized structure that converges to the minimum point. Thermodynamic statistical analysis is then performed on the second ground state optimized structure to obtain the thermal correction of the complex free energy.

[0104] Furthermore, for the second deduplication and sorting structure set obtained in S43, the present invention selects the top 10 anthocyanin polyphenol complex structures with the lowest energy, and uses Molclus to call Gaussian software to perform ground state geometry optimization and vibration analysis calculations on the above 10 structures in an aqueous environment. The vibration analysis is used to confirm that the optimized structure converges to the potential energy surface minimum point rather than the saddle point.

[0105] Specifically, the Gaussian software uses density functional theory (DFT) to perform high-precision geometric optimization on the first 10 low-energy complex structures. The DFT method solves the Kohn-Sham equation iteratively through self-consistent field (SCF). After each SCF convergence, the force on each atom is calculated, driving the atomic coordinates to be updated along the negative gradient direction. The above SCF calculation and coordinate update process is repeated until the energy change and force components meet the convergence criteria, thus obtaining the second ground state optimized structure.

[0106] After geometric optimization convergence, the Gaussian software performs vibrational analysis on the converged structure in the same task. The vibrational analysis calculates the second derivative matrix of the system's energy with respect to atomic coordinates (i.e., the Hessian matrix) and diagonalizes it to obtain the vibrational frequencies of each normal vibrational mode. If the imaginary frequency is 0, the structure is considered meaningful.

[0107] Subsequently, after each structure optimization is completed, a second ground-state optimized structure converged to the minimum point and the corresponding frequency calculation output file are output. Then, the present invention inputs the above 10 frequency calculation output files into the Shermo program, and following the same statistical thermodynamic calculation process as S23, extracts the complex free energy thermal correction value corresponding to each anthocyanin polyphenol complex. A complex free energy thermal correction value is output for each complex structure, which is then used for the free energy calculation in S45.

[0108] S45: Based on the second ground-state optimized structure, perform high-precision single-point energy calculation, sort the calculation results by energy, select the anthocyanin polyphenol complex structure with the lowest energy, and perform free energy calculation in combination with the free energy thermal correction of the complex to obtain the free energy data of the complex.

[0109] In step S45, the present invention uses the 10 optimized second ground state structures obtained in S44 as input, and performs single-point energy calculations on each complex structure under preset high-precision quantum chemical calculation parameters to obtain high-precision electronic energy values ​​of each complex under the current geometric configuration. Subsequently, the present invention uses the isostat program to sort the single-point energy values ​​of the above 10 complexes, selects the anthocyanin polyphenol complex structure with the lowest single-point energy, adds the high-precision single-point energy value of this structure to the corresponding complex free energy thermal correction amount in S44, and obtains the complex free energy data of the anthocyanin polyphenol complex. This value represents the free energy corresponding to the configuration of the anthocyanin polyphenol complex with the most stable energy in an aqueous solution environment.

[0110] S46: Based on the first ground-state optimized structure, perform single-point energy calculations on anthocyanin monomers and multiple candidate polyphenol molecules respectively, and perform free energy calculations in combination with the corresponding monomer free energy thermal correction to obtain monomer free energy data; S47: Perform free energy difference calculations on the monomer free energy data corresponding to the free energy data of the complex to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanin.

[0111] In steps S46-S47, the present invention uses the ground-state optimized structure of anthocyanin monomers obtained in S22 and the ground-state optimized structures of 20 candidate polyphenol monomers as inputs. Under the same preset high-precision quantum chemical calculation parameters as in S45, high-precision single-point energy calculations are performed on the anthocyanin monomers and each candidate polyphenol monomer to obtain the high-precision electronic energy values ​​of each monomer.

[0112] Subsequently, the single-point energy value of the anthocyanin monomer is added to the thermal correction amount of the free energy of the anthocyanin monomer obtained in S23 to obtain the free energy data of the anthocyanin monomer. Similarly, the single-point energy value of each candidate polyphenol monomer is added to the thermal correction amount of the corresponding polyphenol monomer free energy obtained in S23 to obtain the free energy data of each candidate polyphenol monomer.

[0113] Subsequently, for each anthocyanin-polyphenol complex, the present invention calculates the free energy difference by subtracting the sum of the free energy data of the corresponding anthocyanin monomer and the free energy data of the candidate polyphenol molecule monomer from the free energy data of the complex obtained in S45, thereby obtaining the binding free energy value between the candidate polyphenol molecule and the anthocyanin. Each of the 20 candidate polyphenol molecules corresponds to a binding free energy value, which are then summarized to form a set of binding free energy values.

[0114] S5: Integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct a training dataset, and perform feature filtering on the training dataset to obtain a preferred feature subset.

[0115] Step S5 further includes:

[0116] S51: Align and integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct the original dataset for machine learning.

[0117] In step S51, the present invention aligns the polyphenol molecule descriptor matrix obtained in S33 with the binding free energy value set obtained in S43 according to the polyphenol molecule number, and adds the binding free energy value as a new column to the right side of the descriptor matrix to construct an original dataset of 20 rows × 13 columns, where each row corresponds to a candidate polyphenol molecule, the first 12 columns are molecule descriptors, and the 13th column is the corresponding binding free energy value.

[0118] S52: Calculate the Pearson correlation coefficient between multiple feature descriptors in the original dataset and between each feature descriptor and the binding free energy to generate a correlation heatmap, so as to remove redundant features and obtain a clean dataset.

[0119] In step S52, this invention calculates the Pearson correlation coefficient between all pairs of columns in the original dataset. The Pearson correlation coefficient is a statistic that measures the degree of linear correlation between two variables, ranging from -1 to 1. The closer the absolute value is to 1, the stronger the linear correlation between the two variables. This invention displays all correlation coefficients as a correlation heatmap using color coding. By identifying feature pairs with excessively high correlation (i.e., redundancy) between descriptors through the heatmap, redundant features are removed from the original dataset, resulting in a cleaned dataset.

[0120] S53: The Z-score normalization method is used to normalize all feature variables in the cleaned dataset to obtain a standardized dataset.

[0121] In step S53, the present invention uses the Z-score standardization method to process all feature columns in the clean dataset. Z-score standardization subtracts the mean of each column from the value of each column and divides it by the standard deviation of the column, so that the data of each column is transformed into a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating the influence of the difference in the dimensions of different descriptors on model training and obtaining a standardized dataset.

[0122] S54: Divide the standardized dataset into a training set and a test set. For the training set, use recursive feature elimination, feature importance ranking based on random forest, and SelectKBest method based on statistical test to select features. Then, use principal component analysis to reduce the dimensionality of features and obtain the preferred feature subset.

[0123] In step S54, the standardized dataset is divided into a training set and a test set at a ratio of 80% and 20%, respectively. For the training set, the invention performs three feature selection methods and one feature dimensionality reduction method: Recursive Feature Elimination (RFE) starts from all features and gradually eliminates the features with the smallest contribution based on the model's weight coefficients for each feature until the number of remaining features meets a set threshold; Random Forest-based Feature Importance Ranking constructs multiple decision trees and calculates the average decrease in impurity of each feature during tree node splits, then sorts all features in descending order of importance and selects the top N; SelectKBest scores the linear correlation significance between each descriptor and binding free energy based on the F-statistic, retaining the K features with the highest scores; Principal Component Analysis (PCA) projects the original feature space onto the directions of several orthogonal principal components with the largest variance through a linear transformation, compressing the 12-dimensional descriptor into 5 principal components, resulting in a new feature representation after dimensionality reduction. Each of the above four processing methods outputs a preferred feature subset for use by the various machine learning algorithms in S61.

[0124] S6: Using the preferred feature subset as input variables and the set of combined free energy values ​​as output variables, a model is trained using a machine learning algorithm to obtain a combined free energy prediction model.

[0125] Step S6 further includes:

[0126] S61: Using the preferred feature subset as input and the set of combined free energy values ​​as output, various machine learning algorithms are introduced to obtain multiple algorithm models to be trained; wherein, the machine learning algorithms include: linear regression, ridge regression, lasso regression, elastic network, decision tree, random forest, gradient boosting, support vector regression, K nearest neighbors and neural network.

[0127] In step S61, the present invention uses the preferred feature subset obtained in S54 as input variables and the set of free energy values ​​as output variables. It introduces ten machine learning algorithms, namely linear regression, ridge regression, lasso regression, elastic network, decision tree, random forest, gradient boosting, support vector regression, K nearest neighbors and neural network, and combines them with the four feature processing methods in S54 to construct multiple algorithm models to be trained.

[0128] Among them, linear regression directly fits the linear relationship between the descriptor and the binding free energy; ridge regression and lasso regression introduce L2 and L1 regularization terms respectively to suppress overfitting on the basis of linear regression; elastic network introduces both L1 and L2 regularization; decision tree builds a classification and regression tree by recursively dividing the feature space; random forest constructs multiple decision trees and takes the average of the prediction results; gradient boosting gradually stacks weak learners with the residual as the optimization objective; support vector regression finds the optimal regression hyperplane in the high-dimensional feature space; K-nearest neighbors use the average binding free energy of the K nearest samples in the feature space as the prediction value; and neural network fits the nonlinear relationship between the descriptor and the binding free energy through a multi-layer fully connected structure.

[0129] S62: The hyperparameters of multiple algorithm models are continuously optimized by using grid search combined with 5-fold cross-validation to obtain the optimal hyperparameter configurations for multiple algorithm models.

[0130] After establishing multiple models, in step S62 of this invention, a grid search method is used to enumerate the list of candidate hyperparameter values ​​for each algorithm. Combined with 5-fold cross-validation, the training set is divided into 5 equal parts, and 4 parts are used for training and 1 part for validation. The cycle is repeated 5 times and the average validation error is taken to evaluate each combination of hyperparameters and determine the optimal hyperparameter configuration for each algorithm model.

[0131] S63: Select evaluation metrics, evaluate the performance of multiple algorithm models and combinations of multiple algorithm models on the test set based on multiple optimal hyperparameter configurations, and select the combined free energy prediction model.

[0132] Furthermore, this invention uses the coefficient of determination R², root mean square error (RMSE), and mean absolute error (MAE) as evaluation indicators. The algorithm model under each optimal hyperparameter configuration is used to predict on the test set. The values ​​of the above three indicators are calculated. After calculation, the feature processing method and algorithm combination with the highest R² and lowest MAE on the test set are selected to obtain the combined free energy prediction model.

[0133] S7: Input the molecular descriptor of the polyphenol molecule to be screened into the binding free energy prediction model to predict the binding free energy of the anthocyanin polyphenol complex, and screen anthocyanin stabilizer candidate molecules based on the predicted binding free energy.

[0134] Step S7 further includes:

[0135] S71: Calculate multidimensional molecular descriptors for the polyphenol molecules to be screened to obtain sample descriptor data to be predicted; S72: Process the sample descriptor data to be predicted using Z-score normalization and feature processing, and input it into the binding free energy prediction model to obtain the predicted binding free energy between multiple polyphenol molecules to be screened and anthocyanins; S73: Sort the polyphenol molecules to be screened according to the magnitude of the negative value of the predicted binding free energy, and select anthocyanin stabilizer candidate molecules from the polyphenol molecules with the largest negative value of binding free energy.

[0136] In the prediction phase, the present invention calculates the physicochemical and topological parameters of the polyphenol molecules to be screened using the methods described in S31 and S32, and extracts electrostatic surface descriptors using the Multiwfn program. The two sets of data are then merged to obtain the descriptor data for the samples to be predicted. This data format is completely consistent with the column structure of each row in the polyphenol molecule descriptor matrix. Subsequently, the present invention applies Z-score standardization to the descriptor data for the samples to be predicted, using the same mean and standard deviation parameters as in S53. Then, according to the optimal feature processing method in S54, feature selection or dimensionality reduction is performed on the standardized data to obtain a feature vector consistent with the input dimension of the binding free energy prediction model. This feature vector is input into the binding free energy prediction model, which outputs the predicted binding free energy value corresponding to each polyphenol molecule to be screened. The present invention sorts all the predicted binding free energy values ​​of the polyphenol molecules to be screened from largest to smallest negative value. The polyphenol molecule with the largest negative binding free energy value has the strongest driving force and the most stable binding with anthocyanins, thus obtaining candidate molecules for anthocyanin stabilizers.

[0137] like Figure 2 As shown, the present invention also provides an anthocyanin stabilizer screening system based on the fusion of quantum chemical computation and machine learning, comprising:

[0138] Collection module 100: Used to collect structural data of anthocyanin molecules and candidate polyphenol molecules, and construct a molecular structure dataset;

[0139] Optimization module 200: Based on the molecular structure dataset, it performs geometric structure optimization on anthocyanin molecules and multiple candidate polyphenol molecules respectively, to obtain the first ground state optimized structure and monomer free energy thermal correction amount corresponding to each molecule;

[0140] Extraction module 300: Used to perform multidimensional molecular descriptor calculations on multiple candidate polyphenol molecules based on descriptor extraction tools, and obtain a polyphenol molecule descriptor matrix;

[0141] Calculation module 400: is used to obtain the second ground state optimized structure of the anthocyanin polyphenol complex from the molecular structure dataset through configuration search and optimization screening, and to perform single-point energy calculation on the first ground state optimized structure and the second ground state optimized structure respectively, and to perform free energy difference calculation in combination with the monomer free energy thermal correction amount, so as to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins.

[0142] Filtering module 500: used to integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct a training dataset, and to perform feature filtering on the training dataset to obtain a preferred feature subset;

[0143] Training module 600: used to train the model using the preferred feature subset as input variables and the set of combined free energy values ​​as output variables through machine learning algorithms to obtain a combined free energy prediction model;

[0144] The prediction module 700 is configured as the binding free energy prediction model obtained by the training module 600. It is used to receive molecular descriptors of polyphenol molecules to be screened, so as to predict the binding free energy of anthocyanin polyphenol complexes, and to screen anthocyanin stabilizer candidate molecules by predicting the binding free energy.

[0145] The present invention also provides an anthocyanin stabilizer screening device based on the fusion of quantum chemical calculation and machine learning, comprising: a memory and at least one processor, wherein the memory stores instructions; the processor invokes the instructions in the memory to cause the anthocyanin stabilizer screening device based on the fusion of quantum chemical calculation and machine learning to perform an anthocyanin stabilizer screening method as described above.

[0146] The present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement an anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning as described in any of the preceding claims.

[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0148] The following describes a method and system for screening anthocyanin stabilizers based on the fusion of quantum chemical calculation and machine learning, using specific embodiments of the present invention.

[0149] Step S101: Collect the SMILES numbers of cyanidin-3-O-glucoside chloroanthocyanin and 20 polyphenol molecules, and convert them into .xyz coordinate files.

[0150] Step S102: Based on the SMILES number of the polyphenol molecule, the physicochemical and topological parameters of each polyphenol molecule are calculated using RDkit.

[0151] Step S103: The polyphenol molecular structure is converted into an input file for a quantum chemical calculation program, and the ground-state geometry of the polyphenol molecular structure is optimized. Based on the optimized ground-state structure, the electrostatic potential surface descriptor is calculated using the Multiwfn program.

[0152] Step S104: Calculate the single-point energy and free energy based on the ground-state structure optimized in step S103;

[0153] The molecular structure of the polyphenol-anthocyanin complex was constructed and converted into an input file for a quantum chemical calculation program. The geometry of the complex's ground state was optimized. Based on the optimized ground state structure, single-point energies and free energies were calculated.

[0154] Step S105: Calculate the binding free energy of anthocyanin-polyphenol molecules, and obtain the following results: Figure 3 As shown.

[0155] Figure 3 This is a histogram showing the distribution of binding free energies between 20 candidate polyphenol molecules and cyanidin-3-O-glucoside chloroanthocyanin in this embodiment of the invention. The horizontal axis represents the numerical range of binding free energy in kcal / mol, from left to right: [-20.28, -17.28], [-17.28, -14.28], [-14.28, -11.28], [-11.28, -8.28], [-8.28, -5.28], and [-5.28, -2.28], a total of six ranges. The vertical axis represents the number of polyphenol molecules falling into each range, with values ​​ranging from 0 to 8. Figure 2The binding free energy values ​​are mainly concentrated in the range of [-17.28, -14.28], with the highest number of polyphenol molecules (up to 7) within this range. This is followed by the range of [-20.28, -17.28] (5 molecules), and the ranges of [-8.28, -5.28] and [-5.28, -2.28] (3 molecules each). The ranges of [-14.28, -11.28] and [-11.28, -8.28] have relatively fewer molecules. This distribution indicates that the binding free energy between the 20 candidate polyphenol molecules and anthocyanins spans approximately 18 kcal / mol, and there are significant differences in the binding affinity between different polyphenol molecules and anthocyanins. The dataset covers a variety of samples, from strong to weak binding, providing sufficiently discriminative labeled data for subsequent machine learning model training.

[0156] Step S106: Use the descriptors extracted in step S103 and the binding free energy calculated in step S105 as the dataset for machine learning. The heatmap of the Pearson correlation coefficients between feature descriptors and between descriptors and binding free energies in the dataset is shown below. Figure 4 As shown.

[0157] Figure 4 This is a Pearson correlation coefficient heatmap between the polyphenol molecule descriptor matrix and the set of binding free energy values ​​in this embodiment of the invention. The heatmap's rows and columns each correspond to 13 variables, from top to bottom and left to right: hydrogen bond donor, hydrogen bond acceptor, number of rings, number of aromatic rings, number of aliphatic rings, lipid-water partition coefficient, molecular weight, topological polar surface area, molecular polarity index, minimum electrostatic potential, maximum electrostatic potential, volume, and binding free energy. Color coding is based on the right-hand color bar: yellow-green represents a correlation coefficient close to 1.0, dark blue-purple represents a correlation coefficient close to -1.0, and intermediate shades correspond to a correlation coefficient close to 0. Figure 4 The following key numerical relationships can be obtained: the correlation coefficient between binding free energy and ring number is -0.93, with hydrogen bond donors -0.77, with hydrogen bond acceptors -0.74, with aromatic ring number -0.87, with volume -0.67, and with lipid-water partition coefficient 0.28. This indicates that the number of rings, aromatic rings, and hydrogen bond donors have the strongest negative correlation with binding free energy; that is, polyphenol molecules containing more ring structures and hydrogen bond donors have a larger negative binding free energy and more stable binding with anthocyanins. Furthermore, the correlation coefficient between hydrogen bond donors and hydrogen bond acceptors is 0.96, and the correlation coefficient between ring number and molecular weight is 0.82, indicating strong collinearity among some descriptors.

[0158] Step S107: Using the feature descriptors from the dataset in Step S106 as input independent variables for machine learning, and combining them with free energy as the output dependent variable, all independent variables are standardized using Z-score. The dataset is then divided into training and test sets, with the training set comprising 80% and the test set comprising 20%. Feature engineering is performed using feature selection or feature dimensionality reduction methods to improve model performance, reduce the risk of overfitting, and enhance computational efficiency. Specific feature selection methods include Recursive Feature Elimination (RFE), Random Forest-based importance ranking, and SelectKBest-based feature selection. Specific feature dimensionality reduction methods include Principal Component Analysis (PCA). Ten algorithms were introduced: Linear Regression, Ridge Regression, Lasso Regression, Elastic Net, Decision Tree, Random Forest, Gradient Boosting, Support Vector Regression (SVR), K-Nearest Neighbors (KNN), and Neural Network. Grid search and 5-fold cross-validation were used to determine the hyperparameters of the algorithms, and the coefficient of determination (R²) was used to... 2 The model tested included root mean square error (RMSE) and mean absolute error (MAE). Training revealed that partial feature selection / feature dimensionality reduction combined with the model exhibited superior performance. The final model performance and analysis are shown in Table 1 and [Table data would be inserted here]. Figure 5 .

[0159] Figure 5 This is a bar chart comparing the R² values ​​of four feature processing and algorithm combinations on the training and test sets in this embodiment of the invention. The horizontal axis lists the four model combinations, from left to right: Principal Component Analysis-Decision Tree (PCA), Principal Component Analysis-Random Forest (PCA), Principal Component Analysis-Gradient Boosting (PCA), and Feature Selection Based on Statistical Tests-Support Vector Regression (SelectKBest). The vertical axis represents the coefficient of determination R², ranging from 0 to 1.0. A closer R² to 1.0 indicates a better fit between the model's prediction and the actual bounding free energy. Each group contains two bars: the orange bar represents the R² value on the training set, and the green bar represents the R² value on the test set. Figure 5The specific values ​​are as follows: PCA-Decision Tree has a training set R² of 0.82 and a test set R² of 0.79; PCA-Random Forest has a training set R² of 0.90 and a test set R² of 0.89; PCA-GradientBoosting has a training set R² of 0.86 and a test set R² of 0.83; and SelectKBest-Support Vector Regression has a training set R² of 0.97 and a test set R² of 0.93. The training set R² and test set R² values ​​of the four models are close, indicating that the models do not exhibit significant overfitting. The SelectKBest-Support Vector Regression combination achieves the highest R² value of 0.93 on the test set. Combined with the MAE of 0.72 kcal / mol for this combination in Table 1, it is determined to be the final bound free energy prediction model adopted in this invention.

[0160] Table 1 shows the performance scores of some models.

[0161]

[0162] Step S108: After calculating and extracting the polyphenol molecule-related feature descriptors using quantum chemical methods, they are imported into the prediction model in this embodiment as machine learning input variables to predict the binding free energy.

[0163] To better demonstrate the specific implementation and technical effects of the present invention, the method for screening anthocyanin stabilizers based on quantum chemical methods and machine learning, as shown in steps S101-S108 of the preferred implementation described above, is applied to the following specific example.

[0164] The specific implementation process of the method for screening anthocyanin stabilizers based on quantum chemistry and machine learning used in this embodiment is as described above and will not be repeated here.

[0165] Table 2 shows the prediction of the binding free energy of the cyanidin-3-O-glucoside chloro-polyphenol complex using the PCA-Decision Treee algorithm of this invention.

[0166] Table 2 shows the predicted values ​​using the PCA-Decision Tree algorithm.

[0167]

[0168] As shown in Table 3, this table demonstrates the prediction of the binding free energy of the cyanidin-3-O-glucoside chloro-polyphenol complex using the PCA-Random Forest algorithm of this invention.

[0169] Table 3 shows the predicted values ​​using the PCA-Random Forest algorithm.

[0170]

[0171] As shown in Table 4, this table demonstrates the prediction of the binding free energy of the cyanidin-3-O-glucoside chloro-polyphenol complex using the PCA-Gradient Boosting algorithm of this invention.

[0172] Table 4. Predictions using the PCA-Gradient Boosting algorithm

[0173]

[0174] As shown in Table 5, this table demonstrates the prediction of the binding free energy of the cyanidin-3-O-glucoside chloro-polyphenol complex using the SelectKBest-support Vector Regression algorithm of this invention.

[0175] Table 5 shows the predicted values ​​using the SelectKBest-support Vector Regression algorithm.

[0176]

[0177] This invention breaks away from the traditional method of using trial and error to discover ideal anthocyanin stabilizers. Based on the model of this invention, it is only necessary to calculate the descriptor of the anthocyanin stabilizer-polyphenol molecule and input it into the machine learning model to predict the binding free energy between the molecule and the anthocyanin molecule. It has the advantages of low computational cost, high accuracy and high efficiency.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for screening anthocyanin stabilizers based on the fusion of quantum chemical computation and machine learning, characterized in that, include: S1: Collect structural data of anthocyanin molecules and candidate polyphenol molecules, and construct a molecular structure dataset; S2: Based on the molecular structure dataset, the geometric structure of anthocyanin molecules and multiple candidate polyphenol molecules is optimized to obtain the first ground state optimized structure and monomer free energy thermal correction for each molecule. S3: Based on the descriptor extraction tool, multidimensional molecular descriptor calculation is performed on multiple candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix; S4: The second ground-state optimized structure of the anthocyanin polyphenol complex is obtained by configuration search and optimization screening of the molecular structure dataset. The first ground-state optimized structure and the second ground-state optimized structure are calculated by single-point energy calculation. The free energy difference is calculated by combining the free energy thermal correction of the monomer to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins. S5: Integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct a training dataset, and perform feature filtering on the training dataset to obtain a preferred feature subset; S6: Using the preferred feature subset as input variables and the set of combined free energy values ​​as output variables, a model is trained using a machine learning algorithm to obtain a combined free energy prediction model; S7: Input the molecular descriptor of the polyphenol molecule to be screened into the binding free energy prediction model to predict the binding free energy of the anthocyanin polyphenol complex, and screen anthocyanin stabilizer candidate molecules based on the predicted binding free energy.

2. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 1, characterized in that, Step S1 further includes: S11: Collect the first SMILES representation of anthocyanin molecules and the second SMILES representation of various candidate polyphenol molecules; S12: The first SMILES representation and the second SMILES representation are converted into three-dimensional coordinate files using a molecular format conversion tool, which are then used as structural input files for quantum chemical calculations to construct and obtain a molecular structure dataset.

3. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 2, characterized in that, Step S2 further includes: S21: Convert multiple molecular structures in the molecular structure dataset into the input file format of quantum chemical calculation software; S22: The IEFPCM solvation model is used to simulate the aqueous environment. Based on preset quantum chemical calculation parameters, the ground-state geometry of anthocyanin molecules is optimized. Under vacuum conditions, the ground-state geometry of multiple candidate polyphenol molecules is optimized based on preset quantum chemical calculation parameters to obtain the first ground-state optimized structure that converges to the minimum point and the corresponding structure checkpoint file. The first ground-state optimized structure includes the ground-state optimized structure of anthocyanin monomers and the ground-state optimized structures of multiple candidate polyphenol molecule monomers. S23: Read the frequency calculation output file corresponding to the first ground state optimized structure, perform thermodynamic statistical analysis, and obtain the monomer free energy thermal correction of multiple molecules.

4. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 3, characterized in that, Step S3 further includes: S31: Using the second SMILES representation of multiple candidate polyphenol molecules as input, calculate the physicochemical and topological parameters of multiple candidate polyphenol molecules using the RDKit toolkit; the physicochemical parameters include the lipid-water partition coefficient, the exact molecular weight, and the topological polar surface area, and the topological parameters include the number of hydrogen bond acceptors, the number of hydrogen bond donors, the number of rings, the number of aromatic rings, and the number of aliphatic rings. S32: The structure checkpoint file is converted into an fchk format file to obtain a quantum chemical structure file. The quantum chemical structure file is then input into the Multiwfn program to perform electrostatic potential surface analysis on multiple candidate polyphenol molecules to obtain an electrostatic surface descriptor. The electrostatic surface descriptor includes the molecular polarity index, molecular volume, minimum electrostatic potential, and maximum electrostatic potential. S33: Integrate the physicochemical parameters, the topological parameters, and the electrostatic surface descriptors to obtain a polyphenol molecule descriptor matrix containing multidimensional features.

5. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 1, characterized in that, Step S4 further includes: S41: Based on the molecular structures of anthocyanin molecules and candidate polyphenol molecules in the molecular structure dataset, the genmer component in Molclus is used to generate multiple initial configurations of anthocyanin-polyphenol complexes. S42: The initial configuration of the anthocyanin polyphenol complex is pre-optimized using the GFNO-xtb method to obtain a first isomer information file; the structures in the first isomer information file are deduplicated and sorted by energy to obtain a first deduplicated and sorted structure set; S43: Under the GFN1-xtb combined with the implicit water model, the first deduplication and sorting structure set is batch optimized to obtain the second isomer information file; the structures in the second isomer information file are deduplicated and sorted by energy to obtain the second deduplication and sorting structure set. S44: For the preferred anthocyanin polyphenol complex structure in the second deduplication and sorting structure set, the geometric structure optimization and vibration analysis are performed in an aqueous environment by calling Gaussian through Molclus to obtain the second ground state optimized structure that converges to the minimum point. Thermodynamic statistical analysis is then performed on the second ground state optimized structure to obtain the thermal correction of the complex free energy. S45: Based on the second ground state optimized structure, perform high-precision single-point energy calculation, sort the calculation results by energy, take the anthocyanin polyphenol complex structure with the lowest energy, and perform free energy calculation in combination with the free energy thermal correction of the complex to obtain the free energy data of the complex. S46: Based on the first ground state optimized structure, single-point energy calculations are performed on anthocyanin monomers and multiple candidate polyphenol molecules respectively, and free energy calculations are performed in combination with the corresponding monomer free energy thermal correction to obtain monomer free energy data. S47: Calculate the free energy difference based on the monomer free energy data corresponding to the free energy data of the complex to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins.

6. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 1, characterized in that, Step S5 further includes: S51: Align and integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct an original dataset for machine learning; S52: Calculate the Pearson correlation coefficient between multiple feature descriptors in the original dataset and between each feature descriptor and the binding free energy to generate a correlation heatmap in order to remove redundant features and obtain a clean dataset. S53: The Z-score normalization method is used to normalize all feature variables in the cleaned dataset to obtain a standardized dataset; S54: Divide the standardized dataset into a training set and a test set. For the training set, use recursive feature elimination, feature importance ranking based on random forest, and SelectKBest method based on statistical test to select features. Then, use principal component analysis to reduce the dimensionality of features and obtain the preferred feature subset.

7. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 1, characterized in that, Step S6 further includes: S61: Using the preferred feature subset as input and the combined free energy numerical set as output, various machine learning algorithms are introduced to obtain multiple algorithm models to be trained. S62: The hyperparameters of multiple algorithm models are continuously optimized by using grid search combined with 5-fold cross-validation to obtain the optimal hyperparameter configurations for multiple algorithm models; S63: Select evaluation metrics, evaluate the performance of multiple algorithm models and combinations of multiple algorithm models on the test set based on multiple optimal hyperparameter configurations, and select the combined free energy prediction model.

8. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 7, characterized in that, In step S61, the machine learning algorithms include: linear regression, ridge regression, lasso regression, elastic network, decision tree, random forest, gradient boosting, support vector regression, K-nearest neighbors, and neural network.

9. The anthocyanin stabilizer screening method based on the fusion of quantum chemical computation and machine learning according to claim 1, characterized in that, Step S7 further includes: S71: Calculate multidimensional molecular descriptors for the polyphenol molecules to be screened to obtain descriptor data for the samples to be predicted; S72: The descriptor data of the sample to be predicted is processed by Z-score normalization and feature processing, and then input into the binding free energy prediction model to obtain the predicted binding free energy between multiple polyphenol molecules to be screened and anthocyanins. S73: Based on the magnitude of the negative value of the predicted binding free energy, sort the polyphenol molecules to be screened, and select anthocyanin stabilizer candidate molecules from the polyphenol molecules with the largest negative binding free energy.

10. A screening system for anthocyanin stabilizers based on the fusion of quantum chemical computation and machine learning, used to perform the screening method for anthocyanin stabilizers based on the fusion of quantum chemical computation and machine learning as described in any one of claims 1 to 9, characterized in that, include: Collection module: Used to collect structural data of anthocyanin molecules and candidate polyphenol molecules, and construct a molecular structure dataset; Optimization module: Based on the molecular structure dataset, it is used to optimize the geometric structure of anthocyanin molecules and multiple candidate polyphenol molecules respectively, and obtain the first ground state optimized structure and monomer free energy thermal correction amount corresponding to each molecule; Extraction module: Used to perform multidimensional molecular descriptor calculations on multiple candidate polyphenol molecules based on descriptor extraction tools, and obtain a polyphenol molecule descriptor matrix; The calculation module is used to obtain the second ground-state optimized structure of the anthocyanin-polyphenol complex from the molecular structure dataset through configuration search and optimization screening, and to perform single-point energy calculations on the first ground-state optimized structure and the second ground-state optimized structure respectively. Combined with the free energy thermal correction of the monomer, the free energy difference is calculated to obtain a set of binding free energy values ​​between multiple candidate polyphenol molecules and anthocyanins. The filtering module is used to integrate the polyphenol molecule descriptor matrix with the set of binding free energy values ​​to construct a training dataset, and to perform feature filtering on the training dataset to obtain a preferred feature subset. Training module: used to train the model using the preferred feature subset as input variable and the set of combined free energy values ​​as output variable, through machine learning algorithms, to obtain a combined free energy prediction model; The prediction module is configured with the binding free energy prediction model obtained by the training module. It is used to receive molecular descriptors of polyphenol molecules to be screened, predict the binding free energy of anthocyanin polyphenol complexes, and screen anthocyanin stabilizer candidate molecules based on the predicted binding free energy.