Method and equipment for predicting surface tension of nonionic surfactant and medium
By calculating the molecular descriptor of nonionic surfactant and using pre-trained models, the problem of low prediction accuracy of nonionic surfactant surface tension in the prior art is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202510667904.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The method of predicting the surface tension of nonionic surfactants based on QSPR in the prior art has low prediction accuracy.
The surface tension is determined by calculating the molecular descriptor of the nonionic surfactant to be predicted, including structural descriptors, empirical descriptors, topological descriptors, and potential descriptors, and using the trained surface tension prediction model.
The accuracy of nonionic surfactant surface tension prediction is improved, and the prediction can be accurately completed.
Smart Images

Figure CN120183518A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of surface tension prediction, and in particular to a method, device and medium for predicting the surface tension of a nonionic surfactant. Background Art
[0002] As a representative product in the field of fine chemicals, surfactants play an important role in the national economy. Their development level has become one of the important signs of the progress of the chemical industry in various countries. They are widely used in many fields such as daily chemicals, textiles, papermaking, pesticides, leather and petrochemicals. People also give it the reputation of industrial monosodium glutamate. For example, in the surfactant industry, surfactants for pesticides are an important field. The final application form of pesticides is pesticide preparations. Pesticide surfactants can prepare pesticide raw materials that cannot be used directly into pesticide preparations that can be used directly. Most pesticide preparations need to be diluted with water before use. Pesticide surfactants play a decisive role in the technical indicators of the emulsion stability, dispersion stability, wettability, persistent foaming, suspension rate and other technical indicators of the pesticide preparation diluent. They are also closely related to the fumigation, stomach poison, systemic absorption and contact killing effects of the liquid during application, as well as the spreading, penetration, spreading and deposition effects on the target, thus playing an important role in fully exerting the efficacy of pesticides. For example, in pesticide spraying, surfactants can promote the wetting and spreading of droplets on plant leaves by reducing the surface tension of the pesticide solution, thereby improving the deposition efficiency and utilization rate of the pesticide.
[0003] Surfactants are a class of compounds with unique amphiphilic molecular structures. Their molecules are composed of hydrophilic groups and hydrophobic groups. They can adsorb on the interface and significantly reduce the surface tension of liquids. However, the performance of surfactants is closely related to their molecular structure. Traditional experimental methods for determining their physical and chemical properties are often time-consuming and labor-intensive, and it is difficult to meet the needs of efficient molecular design. The quantitative structure-property relationship (QSPR) method provides an effective way to solve this problem. QSPR can achieve rapid prediction of physical and chemical properties by establishing a mathematical model between molecular structure and its physical and chemical properties. Its core idea is that the physical and chemical properties of molecules are determined by their molecular structure. By building a model through statistical or machine learning algorithms, quantitative prediction of physical and chemical properties can be achieved. In recent years, with the rapid development of machine learning algorithms, QSPR has been significantly improved in prediction accuracy and application scope, and has become an important tool in fields such as drug design, materials science, and environmental chemistry.
[0004] However, the current method of predicting the surface tension of non-ionic surfactants based on QSPR still has the problem of low prediction accuracy. Summary of the invention
[0005] The purpose of this application is to provide a method, device and medium for predicting the surface tension of non-ionic surfactants, which can accurately complete the prediction of the surface tension of non-ionic surfactants and improve the prediction accuracy.
[0006] To achieve the above object, the present application provides the following solutions.
[0007] In a first aspect, the present application provides a method for predicting the surface tension of non-ionic surfactants, and the method for predicting the surface tension of non-ionic surfactants includes: Based on the molecular structure of the non-ionic surfactant to be predicted, molecular descriptors are calculated; the molecular descriptors include structural descriptors, empirical descriptors, topological descriptors and potential descriptors; Using the molecular descriptors as inputs, the surface tension of the non-ionic surfactant to be predicted is determined by using a trained surface tension prediction model.
[0008] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above method for predicting the surface tension of non-ionic surfactants.
[0009] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method for predicting the surface tension of non-ionic surfactants is implemented.
[0010] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application provides a method, device and medium for predicting the surface tension of non-ionic surfactants. Based on the molecular structure of the non-ionic surfactant to be predicted, molecular descriptors are calculated. The molecular descriptors include structural descriptors, empirical descriptors, topological descriptors and potential descriptors. Using the molecular descriptors as inputs, the surface tension of the non-ionic surfactant to be predicted is determined by using a trained surface tension prediction model. By designing the molecular descriptors to include structural descriptors, empirical descriptors, topological descriptors and potential descriptors, and then directly using the molecular descriptors as inputs and predicting the surface tension by using a trained surface tension prediction model, since more comprehensive and suitable molecular descriptors are selected, the surface tension prediction of non-ionic surfactants can be accurately completed and the prediction accuracy can be improved. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0012] Figure 1 It is an application environment diagram of a method for predicting the surface tension of a non-ionic surfactant provided in Embodiment 1 of the present application.
[0013] Figure 2 It is a schematic flowchart of a method for predicting the surface tension of a non-ionic surfactant provided in Embodiment 1 of the present application.
[0014] Figure 3 It is a schematic diagram of the QSPR research process provided in Embodiment 1 of the present application.
[0015] Figure 4 It is a schematic diagram of 7 typical structures of non-ionic surfactants provided in Embodiment 1 of the present application; among them, Figure 4 in (1) is branched alkyl ethoxylate, Figure 4 in (2) is linear alkyl ethoxylate, Figure 4 in (3) is octylphenol ethoxylate, Figure 4 in (4) is alkanediol, Figure 4 in (5) is alkyl monosaccharide and disaccharide ethers and esters, Figure 4 in (6a) is ethoxylated alkylamine, Figure 4 in (6b) is ethoxylated fluoroalkylamine, Figure 4 in (7a) is fluorinated linear alkyl ethoxylate, Figure 4 in (7b) is fluorinated ether ethoxylate.
[0016] Figure 5 It is a comparison schematic diagram of the predicted value and the true value of the surface tension of a non-ionic surfactant predicted by the SVR (Support Vector Regression) model provided in Embodiment 1 of the present application.
[0017] Figure 6 It is a residual schematic diagram of the surface tension of a non-ionic surfactant predicted by the SVR model provided in Embodiment 1 of the present application.
[0018] Figure 7 It is a Williams diagram schematic diagram of the surface tension of a non-ionic surfactant predicted by the SVR model provided in Embodiment 1 of the present application.
[0019] Figure 8A schematic structural diagram of a computer device provided in Embodiment 2 of this application. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0021] Embodiment 1.
[0022] The method for predicting the surface tension of a nonionic surfactant provided in the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set separately, integrated on the server, placed in the cloud or on other servers. The terminal can send a prediction request to be processed to the server. After receiving the prediction request to be processed, for the prediction request to be processed, the server calculates molecular descriptors based on the molecular structure of the nonionic surfactant to be predicted. The molecular descriptors include structural descriptors, empirical descriptors, topological descriptors, and potential descriptors. Using the molecular descriptors as input, the trained surface tension prediction model is used to determine the surface tension of the nonionic surfactant to be predicted. The server can feedback the prediction result of the surface tension for the prediction request to the terminal.
[0023] In addition, in some embodiments, the method for predicting the surface tension of a nonionic surfactant can also be implemented independently by the server or the terminal. For example, the terminal can directly process the prediction request to be processed, or the server can obtain the prediction request to be processed from the data storage system and process the prediction request to be processed.
[0024] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0025] In an exemplary embodiment, as Figure 2As shown, a method for predicting the surface tension of a non-ionic surfactant is provided. This method is executed by a computer device, which can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the application of this method to Figure 1 the server in it as an example for illustration, it includes the following steps.
[0026] Step S1, based on the molecular structure of the non-ionic surfactant to be predicted, calculate molecular descriptors; the molecular descriptors include structural descriptors, empirical descriptors, topological descriptors, and potential descriptors.
[0027] Step S2, using the molecular descriptors as inputs, and using the trained surface tension prediction model to determine the surface tension of the non-ionic surfactant to be predicted.
[0028] Implementing the above steps S1 to S2, in this embodiment, for the application scenario of predicting the surface tension of non-ionic surfactants, a method for predicting the surface tension of non-ionic surfactants using QSPR is provided. Corresponding molecular descriptors are creatively designed, specifically including structural descriptors, empirical descriptors, topological descriptors, and potential descriptors. Subsequently, after calculating the molecular descriptors based on the molecular structure of the non-ionic surfactant to be predicted, directly using the molecular descriptors as inputs, and using the trained surface tension prediction model to determine the surface tension of the non-ionic surfactant to be predicted. Since more comprehensive and suitable molecular descriptors are selected, the surface tension prediction of non-ionic surfactants can be accurately completed, improving the prediction accuracy.
[0029] The calculation of molecular descriptors is an essential step in QSPR research. In order to accurately represent various characteristics of molecules, in this embodiment, RDKit is used to calculate a large number of molecular descriptors according to the molecular structure of non-ionic surfactants. The surface tension prediction model established in this embodiment later is a 2D-QSPR model. Therefore, a large number of molecular descriptors calculated at this time are all 2D molecular descriptors, mainly including structural descriptors, empirical descriptors, geometric descriptors, quantum chemical descriptors, topological descriptors, potential descriptors, etc. It should be noted that RDKit is an open-source toolkit for chemoinformatics. The core data structures and algorithms are written in C++ to improve performance and achieve maximum portability. It provides wrappers for programming languages such as Python, R, C#, and Java to allow the use of this open-source toolkit in many different environments. This open-source toolkit provides powerful functions, including support for conformation generation, chemical reactions, a series of chemical file formats, etc. RDKit has a large and active user community and has been integrated into many other open-source and commercial projects.
[0030] The structural descriptors comprehensively describe the overall configuration of a molecule based on fundamental features such as atom types, bond types, and ring structures in the molecule. Although the expression forms of structural descriptors are relatively simple, they play an irreplaceable role in QSPR studies. Especially in the study of physical properties of small molecules, they can significantly improve the fitting effect of models and provide efficient support for the research.
[0031] Empirical descriptors such as octanol / water partition coefficient and molecular refractive index are usually determined experimentally or calculated by atom contribution methods. For example, an innovative atom contribution model was proposed and successfully applied to the calculation of octanol / water partition coefficient and molecular refractive index. This atom contribution model has now been integrated into the descriptor library of RDKit. The octanol / water partition coefficient, as an important indicator for measuring the hydrophobicity of molecules, is often used to predict the distribution behavior of small molecules in organisms and provides key references for drug design and environmental chemistry research. The combined use of these descriptors further expands the depth and breadth of molecular property characterization.
[0032] Topological descriptors are extracted from the structure of molecular graphs through mathematical methods. These descriptors can capture the connection patterns and branching characteristics of molecules. As one of the most widely used descriptor categories, topological descriptors cover a variety of classical parameters such as Wiener index, Randic index, and Kier-Hall shape index. These descriptors are highly sensitive to the connectivity of molecules, especially to the arrangement and distribution of different branches in the molecule. Although they are not as direct as structural descriptors in describing molecular composition, some topological descriptors (such as Kier-Hall index) can still reflect the molecular composition information to a certain extent.
[0033] Electrostatic potential descriptors are descriptors calculated based on the molecular charge distribution and electrostatic potential field, used to characterize the electrostatic properties of molecules, including partial charges, electrostatic potential, polar surface area (PSA), and extreme values of electrostatic potential, etc. These descriptors are calculated by quantum chemistry or empirical methods, can reflect the charge distribution and electrostatic interaction ability on the molecular surface, and are widely used in drug design, research on intermolecular interactions, and QSPR models. Electrostatic potential descriptors not only reveal the reactive sites and solvation effects of molecules, but also provide an important basis for predicting the physicochemical properties and biological activities of molecules, and are indispensable tools in chemical and biological research. The molecular electrostatic potential (MEP) in electrostatic potential descriptors was initially used to study the electrophilicity of molecules, and later gradually extended to analyze atomic and molecular properties related to nucleophilicity such as energy, covalent radius, ionic radius, and electronegativity. The distribution of electrostatic potential is affected by the electronic effects and steric structures of groups around specific atoms. Therefore, understanding the complementary relationship between the overall electrostatic potential and the local ionization potential is of great significance for calculating electrostatic potential descriptors and can effectively analyze and predict the reaction behavior of molecular systems. A certain study shows that the MEP surface value can be used to characterize the relative strengths of hydrogen bond donors and acceptors. For example, the hydrogen bond interaction between 4-hydroxybenzoic acid and isoniazide can be revealed by the relationship between MEP and ionization constant. Common electrostatic potential descriptors include the maximum and minimum partial charges in the molecule, polarity parameters, and the surface area of the charged part, etc. These descriptors provide an important quantitative basis for studying intermolecular interactions and reaction activities.
[0034] Therefore, the specific molecular descriptors designed in this embodiment include structural descriptors, empirical descriptors, topological descriptors, and electrostatic potential descriptors. Structural descriptors include Labute approximate surface area, average molecular weight, exact molecular weight, number of heteroatoms, number of heavy atoms, total number of atoms, number of chiral centers, number of unspecified chiral centers, number of aromatic rings, total number of rings, number of aromatic heterocycles, number of aliphatic carbocycles, and the proportion of SP3 hybridized carbon in the total carbon. Empirical descriptors include the logarithm of the octanol / water partition coefficient and molecular refractive index. Topological descriptors include the number of rotatable bonds, number of hydrogen bond donors, number of hydrogen bond acceptors, molecular connectivity index, symmetry and bonds, chain saturation topological index, and quantum topological electronic index. Electrostatic potential descriptors include the polar surface area of each molecule, number of Lipinski hydrogen bond acceptors, number of Lipinski hydrogen bond donors, hydrophilic / hydrophobic interaction, polarization, and electrostatic interaction. Hydrophilic / hydrophobic interaction, polarization, and electrostatic interaction can be calculated by MOE (Molecular Operating Environment). The molecular descriptors adopted in this embodiment are specifically shown in Table 1 below.
[0035] Table 1 Molecular Descriptors
[0036] The molecular structure of the non-ionic surfactant to be predicted is input into RDKit for calculation to obtain molecular descriptors, which mainly include 2D molecular descriptors such as structural parameters, electrical parameters, and topological parameters, as shown in Table 1 specifically. Based on the molecular structure of the non-ionic surfactant to be predicted, molecular descriptors are calculated, and subsequently, the molecular descriptors are used as inputs to determine the surface tension of the non-ionic surfactant to be predicted by using the trained surface tension prediction model.
[0037] At this time, in this embodiment, the molecular structure can adopt the SMILES (Simplified molecular input line entry specification) format. At this time, based on the molecular structure of the non-ionic surfactant to be predicted, molecular descriptors are calculated, specifically including: using the molecular structure of the non-ionic surfactant to be predicted as an input and calculating molecular descriptors by using RDKit.
[0038] This embodiment takes the surface tension of non-ionic surfactants as the research object and systematically elaborates the construction process of the QSPR model. Specifically, KNIME (Konstanz Information Miner) is used to complete the construction. KNIME is a data analysis platform, and its core advantage lies in the graphical interface design. Users do not need to write code and can build a complete data science workflow through intuitive node dragging and dropping, covering the entire process from data acquisition, cleaning, analysis, modeling to visualization. It has more than 4,000 built-in functional nodes, supports various data formats from text, databases to images, and can be seamlessly integrated with tools such as Python, R, and Spark, and is compatible with distributed computing environments such as Hadoop, easily handling large-scale data sets of millions or even billions of levels. A quantitative regression model is constructed using a data set composed of the molecular structure in the SMILES format of non-ionic surfactants and the molecular compound properties (specifically referring to surface tension). The quantitative regression model can use various machine learning algorithms, and the methods of leave-one-out cross-validation and external testing are used to test the quantitative regression model, aiming to construct a reliable QSPR model for non-ionic surfactants to provide theoretical support for the molecular design and performance optimization of surfactants.
[0039] The following specifically introduces the training process of the trained surface tension prediction model in this embodiment. The QSPR modeling is mainly divided into five tasks: (1) data collection and collation; (2) calculation of molecular descriptors; (3) feature selection; (4) modeling; (5) model testing. The basic steps include: establishing a data set, calculating molecular descriptors, feature selection, dividing into a training set and a validation set, constructing a QSPR model using various machine learning algorithms, model testing, domain representation, etc., as Figure 3 shown.
[0040] (I) Data collection and collation.
[0041] The collection and collation of data is the first step in QSPR modeling, which is related to the scope of application and fitting degree of the model. The relevant data such as the activity and properties (i.e., surface tension) of the compounds used in modeling mainly come from various published literatures and established chemical substance databases.
[0042] In this embodiment, 91 sample non-ionic surfactants are selected. Non-ionic surfactants can be divided into 7 categories according to the structural characteristics of the hydrophilic part and the hydrophobic part. As Figure 4 shown, Figure 4 (6a) and (6b) in Figure 4 are of one category, denoted as ethoxylated (fluoro)alkylamines,
[0043] (7a) and (7b) in
[0044] Table 2 Sample non-ionic surfactants numbered 1-25
[0045] Table 3 Sample non-ionic surfactants numbered 26-50
[0046] Table 4 Non-ionic surfactants of samples numbered 51 - 75
[0047] Table 5 Non-ionic surfactants of samples numbered 76 - 91
[0048] In Tables 2 - 5, It represents the non-ionic surfactants of the samples used to form the validation set. That is, in this embodiment, the non-ionic surfactants of samples numbered 1 - 73 are used as the first sample non-ionic surfactants to form the training set, and the non-ionic surfactants of samples numbered 74 - 91 are used as the second sample non-ionic surfactants to form the validation set. EST represents the true value of the surface tension, GPR represents the predicted value of the surface tension predicted by using the GPR (Gaussian Process Regression) model, and SVR represents the predicted value of the surface tension predicted by using the SVR model.
[0049] Data preparation and management workflows are crucial in the development of QSPR models. After obtaining the original dataset, further data management is carried out. The first step is to tidy up the data to identify and correct errors in chemical structures. The second step is duplicate analysis to evaluate the quality of experimental data and remove: chemical substances related to duplicate records with conflicting experimental results and the results of the same compounds for duplicate experiments. The presence of duplicates directly affects the quality of the model. That is, the existence of duplicates with the same activity in the training set and test set will lead to an overestimation of the model quality. Manual inspection is required at the end of the process to ensure that all structures are correct. The third step must identify and delete unreliable sources, convert all molecules into SMILES format, save them as.smi format files, input them into RDKit to generate molecular 2D coordinates, and then perform molecular structure checking (MolecularSanitization): convert all molecules into the Kekulé form, then apply the Hueckel rule to determine the aromaticity and turn on the Keep Hydrogens option to prevent RDKit from removing hydrogens from molecules constructed from SDF (Structure Data File).
[0050] When implementing the identification and deletion of unreliable sources based on the KNIME platform, a cheminformatics analysis environment was constructed based on the KNIME platform. First, the KNIME 4.14 version software and functional plug-ins such as Weka, Chemistry Development Kit (CDK), RDKit, Erl Wood, and Indigo were downloaded and installed from the official KNIME website (https: / / www.knime.com / ). Among them, RDKit, as the core open-source toolkit in the field of cheminformatics, provides key technical support through its built-in molecular descriptor calculation module. In QSPR research, the molecular structure is imported in SDF format through the SDF Reader module. SDF is a file format used to store and exchange molecular structure information, including 2D or 3D coordinates of molecules, atomic types, bond information, etc. It can also contain additional information such as the physicochemical properties and bioactivity data of molecules. Then, the SDF molecules are converted into RDKit molecules, and the molecular 2D coordinates are generated. Then, the molecular structure is checked, the molecule is converted into the Kekulé form, and the aromaticity is accurately identified. At the same time, the hydrogen atom information is retained through the Keep Hydrogens option.
[0051] If the true value of the surface tension cannot be directly obtained, the surface tension can be tested. The non-ionic surfactant is dissolved in deionized water, and the concentration gradients are set to be 10 -2 、10 -3 、10 -4 、10 -5 、10 -6 、10 -7 。A clean platinum plate is taken with clips and the platinum plate is burned with an alcohol lamp until it turns red. When burning, the flame of the alcohol lamp is at a 45-degree angle to the horizontal plane. The burned platinum plate is hung on the hook. Diluted liquid (i.e., diluted non-ionic surfactant) is added to the sample dish, and the sample dish is placed on the sample stage. The DCAT21 surface tensiometer is used to measure the static surface tension of the diluted liquid by the Wilhelmy plate method (Wielandt plate method). Each diluted liquid is measured 3 times repeatedly and the average value is taken. As the concentration of the diluted liquid decreases, the decrease in surface tension gradually tends to be stable. The value when the surface tension decreases to stability is selected as the surface tension of this non-ionic surfactant. Among them, the value when the surface tension decreases to stability is the surface tension when the non-ionic surfactant reaches the CMC (critical micelle concentration) value.
[0052] (2) Feature dimensionality reduction and feature selection.
[0053] If the model is overfitted, it will lead to a poor model, so feature dimensionality reduction and feature selection need to be carried out.
[0054] Principal Component Analysis (PCA) is a feature dimensionality reduction method, which is equivalent to redefining new metrics. The finally obtained new features contain most of the information of the data. PCA projects higher-dimensional data onto a lower dimension, maximizing the variance in the dimension. The greater the variance, the greater the amount of information. Each eigenterm of the covariance matrix is a projection plane. The advantage is that the principal components obtained by this transformation method are orthogonal, so they are linearly uncorrelated and can eliminate the mutual influence between data. The disadvantage is that the reduced features are not as complete as the original features, the interpretation is relatively complex, and it is not applicable to machine learning algorithms based on classification learners because it selects features with larger variances rather than features that are more meaningful to the overall model. When performing dimensionality reduction in the PCA process, only the dimensionality features with larger variances are retained, while the features with variances close to 0 are ignored or deleted. The input feature dimensionality can be determined, and its quantity can be directly selected or the minimum amount of information to be saved can be specified. In this embodiment, both choose to save 99.99% of the information. Since PCA has weak interpretability and may not be the best feature dimension for some non-linear algorithms based on optimal plane classification, the original molecular descriptors are still selected for subsequent screening calculations.
[0055] In this embodiment, the Forward Selection (FS) method is selected. It is a feature screening method that sequentially adds a molecular descriptor to the QSPR model. The first molecular descriptor included in the regression is the most relevant feature, such as the feature with the highest correlation or the smallest residual sum of squares with the compound or pharmacophore properties. The first selected molecular descriptor is forced into all further QSPR models, and new molecular descriptors are gradually added to the regression. Equivalent to considering local optimality for each additional variable on the basis of the optimal subset, various rules can be used as the stopping criterion for the FS algorithm. The specific process is to sequentially eliminate the variables with no significant influence in the equation containing all variables until no variable can be eliminated.
[0056] The purpose of feature selection is to select the most suitable feature subset for model construction from all features in the input dataset (including the original molecular descriptors and the molecular descriptors obtained after dimensionality reduction), removing the descriptors with high collinearity. The FS algorithm is selected and variables are added. In each loop, the data is randomly divided into a training set and a test set in a ratio of 4:1. The training set data is used to construct a multiple linear regression model, while the test set data is used to test the model. Some model evaluation metrics will be obtained in each loop. The model evaluation metrics are set to minimize the mean square error, and the feature combinations with smaller mean square error and fewer features are manually selected from the results for the next model construction.
[0057] The specific process of feature selection using the KNIME platform is as follows: In the first step, descriptors with collinearity > 0.8 are removed. In the second step, the feature selection loop starts. The compound properties are set as static columns, and the molecular descriptor data is used as variable columns to participate in the screening process. In each loop, the data is randomly divided into a training set and a test set at a ratio of 4:1. The feature selection uses the FS algorithm, which conducts the feature selection work in a step-by-step manner. Initially, the model feature set is empty. Subsequently, in each round of iteration, the change in model performance after adding the remaining unselected features to the current feature set is evaluated one by one. The model performance evaluation metrics can be reasonably selected according to the specific research questions and data characteristics. For example, the mean squared error is commonly used in regression problems. The algorithm will select the feature that can most significantly improve the model performance and incorporate it into the current feature set. This process is repeated until the improvement in model performance is lower than the preset threshold or the upper limit of the preset number of features is reached, at which point the algorithm terminates. The determined feature set at this time is the finally selected feature subset. By using the FS algorithm, this embodiment can effectively identify and extract features that are closely related to the target variable and crucial for model construction from a complex and high-dimensional data feature space.
[0058] (III) Machine learning modeling.
[0059] Multiple algorithms can be used to establish the QSPR model. In this embodiment, the support vector regression model is specifically selected. The core idea of the support vector regression model is to construct a regression hyperplane so that all data points are as much as possible within the ε-insensitive band of the hyperplane. The ε-insensitive band is a banded region centered on the regression hyperplane with a width of 2ε. As long as the deviation between the predicted value and the true value of the data point does not exceed ε, the prediction is considered accurate. The goal of SVR is to find a regression hyperplane so that as many data points as possible fall within the ε-insensitive band while minimizing the total deviation between the regression hyperplane and the data points.
[0060] (IV) Evaluation and testing of the model.
[0061] The evaluation and testing of the model is the last step in the QSPR modeling process, mainly including the analysis of the model's robustness and predictive ability. This embodiment adopts methods such as leave-one-out cross-validation and external testing, and evaluates the quality of this model according to the results. In addition, the application domain range of the model is further discussed, the reliability of the model is checked, and whether there is over extrapolation is determined.
[0062] After completing the QSPR model construction process, it is necessary to analyze the fitting, stability, and predictive ability of the model to screen out the best model. The commonly used model evaluation parameters are as follows.
[0063] (1) Coefficient of determination (R-Square, R2 ): ; wherein, R 2 is the coefficient of determination; is the total number of samples; is the true value of the surface tension of the th sample; is the predicted value of the surface tension of the th sample;
[0064] (2) Mean Squared Error (MSE): ; wherein, MSE is the mean squared error.
[0065] (3) Mean Absolute Error (MAE): ; wherein, MAE is the mean absolute error.
[0066] (4) Root Mean Squard Error (RMSE): ; wherein, RMSE is the root mean squared error.
[0067] Internal model validation uses leave-one-out cross-validation, that is, each time a model is built, only one sample is used as the test set, and the remaining samples are used as the training set. The quality of the obtained model is closest to the true result, and at the same time, the internal validation accuracy is obtained. External model validation uses the validation set. Usually, the samples in the validation set account for 20%-40% of all the samples in the original data set, and the external validation accuracy is obtained. If the coefficient of determination R 2 of the finally obtained internal validation accuracy > 0.6, and the coefficient of determination R 2 pred of the external validation accuracy > 0.5, it proves that the model has statistical significance. If the difference between R 2 and R 2 pred is less than 0.3, it can be considered that the model does not have overfitting.
[0068] The application domain (AD) is an important concept in QSPR research. The AD characterization can be used to detect the similarity of the compounds used to build the model to estimate the uncertainty in prediction, that is, to check whether the model is reliable. If there are data points outside the AD, it means that the predicted value of this data point is a meaningless extrapolation point and can be considered for deletion in future models. Methods for determining the AD of a model include the leverage method, Euclidean distance method, Mahalanobis distance method, etc., and there are also integrated methods, such as the nearest neighbor algorithm combined with the Euclidean distance. The leverage method is used in this embodiment, that is, the Williams chart is used to characterize the domain range. In the Williams chart, the abscissa is the leverage value (h), and the ordinate is the standard residual (δ). δ is the z-score normalization of the difference between the true value and the predicted value (z-score Normalization). The domain range in the Williams chart is a rectangular area, and the abscissa is guaranteed to be less than , and the ordinate is guaranteed to be between [-3δ, 3δ], generally between [-3, 3]. This data point is within the domain range. If the data point exceeds the domain, the similarity with other data points in the dataset needs to be discussed to confirm whether this data point has excessive extrapolation, and the impact on the model needs to be considered comprehensively.
[0069] ; ; where h i is the leverage value of the i-th observation; x i is a vector representing the independent variable values of the i-th observation; X is the design matrix, and each row of it corresponds to the independent variable values of an observation; is a critical value for judging high leverage points; k is the number of independent variables; n is the number of observations. After calculation, the critical value in the SVR model is 0.62.
[0070] Table 6 below shows the results of internal test accuracy and external test accuracy. It is observed that the adjusted determination coefficient R 2 adjust of the training set and the adjusted determination coefficient R 2 PREadjust of the validation set both reach above 0.90, indicating that the model has a certain interpretability, and the difference between the adjusted determination coefficient R 2 adjust of the training set and the adjusted determination coefficient R 2 PREadjust of the validation set is less than 0.3, indicating that the model has a good fitting degree.
[0071] Table 6 Internal test accuracy and external test accuracy
[0072] In Table 6, MSD is Mean Square Displacement, that is, mean square displacement, and MAPE is Mean Absolute Percentage Error, that is, mean absolute percentage error.
[0073] From Table 6, it can be seen that the adjusted coefficient of determination R of the internal test accuracy of SVR 2 adjust is 0.917. Similar to the GPR model, the adjusted coefficient of determination R of the external test accuracy of SVR 2 PREadjust reaches 0.920, and the RMSE is also small.
[0074] Figure 5 Fig. shows the comparison diagram of the predicted values and the true values of non-ionic surfactants under the SVR model. In the SVR model, all the predicted values are evenly distributed on both sides of the true value, and the model is stable.
[0075] Figure 6 is the residual plot corresponding to the SVR model. It can be seen that most of the data points in the training set and the validation set are evenly distributed on both sides of the 0 point, and there is no objective law. It can be considered that there is no systematic error in this model, but it can be observed that the residual values corresponding to the values with larger surface tension are relatively large.
[0076] Figure 7 is the Williams plot corresponding to the SVR model. Figure 7 In it, the blue squares are the molecular samples of the training set, and the red triangles are the molecular samples of the validation set. The data shows that 98.9% of the molecular samples are within the application domain range, indicating that the model has a certain reliable extrapolation prediction ability. Further analysis shows that the molecular sample numbered 35 in the training set is defined as an outlier through statistical tests, with a high residual value and a lever value meeting the requirements. Among them, the reason for the outlier of the molecular sample numbered 35 in the training set may be that the existing descriptors fail to fully capture the non-linear relationship between this molecule and EST. However, the outlier shows a positive contribution to the model evaluation in the residual analysis, indicating that the model still maintains a high reliability as a whole and cannot be directly removed.
[0077] To sum up, this embodiment believes that the prediction model of the surface tension of non-ionic surfactants established by this SVR model has good stability, robustness and reliability. This model has passed the training set test and the validation set test, and the application domain, out-of-domain points and residual value distribution are discussed, comprehensively proving the stability, robustness and reliability of the model.
[0078] At this time, in this embodiment, before using the trained surface tension prediction model to determine the surface tension of the non-ionic surfactant to be predicted with molecular descriptors as the input, the surface tension prediction method for non-ionic surfactants in this embodiment further includes the following steps.
[0079] (1) Obtain a training set, where the training set includes the sample molecular structure and sample surface tension corresponding to each of the multiple first-sample non-ionic surfactants.
[0080] (2) Based on the sample molecular structure, calculate the sample molecular descriptors, and perform feature dimensionality reduction and feature selection on the sample molecular descriptors to obtain the selected sample molecular descriptors.
[0081] The sample molecular structure is in SMILES format. At this time, based on the sample molecular structure, calculate the sample molecular descriptors, specifically including: using RDKit to calculate the sample molecular descriptors with the sample molecular structure as the input.
[0082] Perform feature dimensionality reduction and feature selection on the sample molecular descriptors to obtain the selected sample molecular descriptors, specifically including: using the principal component analysis method to perform feature dimensionality reduction on the sample molecular descriptors to obtain the dimensionality-reduced sample molecular descriptors; using the forward selection method to perform feature selection on the sample molecular descriptors and the dimensionality-reduced sample molecular descriptors to obtain the selected sample molecular descriptors.
[0083] (3) Use the selected sample molecular descriptors as the input and the sample surface tension as the label to train the initial surface tension prediction model to obtain the trained surface tension prediction model.
[0084] The initial surface tension prediction model is a support vector regression model. At this time, use the selected sample molecular descriptors as the input and the sample surface tension as the label to train the initial surface tension prediction model to obtain the trained surface tension prediction model, specifically including: using the selected sample molecular descriptors as the input and the sample surface tension as the label, and using the leave-one-out cross-validation method to train the initial surface tension prediction model to obtain the trained surface tension prediction model.
[0085] Since feature dimensionality reduction and feature selection are performed during training, therefore, using the molecular descriptors as the input and using the trained surface tension prediction model to determine the surface tension of the non-ionic surfactant to be predicted specifically includes: screening the molecular descriptors based on the selected sample molecular descriptors to obtain the selected molecular descriptors; using the selected molecular descriptors as the input and using the trained surface tension prediction model to determine the surface tension of the non-ionic surfactant to be predicted.
[0086] After obtaining the trained surface tension prediction model, the surface tension prediction method for non-ionic surfactants in this embodiment further includes the following steps.
[0087] (1) Obtain a validation set, where the validation set includes the selected sample molecular descriptors and sample surface tensions corresponding to each of the multiple second sample non-ionic surfactants.
[0088] (2) Use the training set to test the trained surface tension prediction model to obtain the internal test accuracy of the trained surface tension prediction model.
[0089] (3) Use the validation set to test the trained surface tension prediction model to obtain the external test accuracy of the trained surface tension prediction model.
[0090] (4) Use the leverage method, Euclidean distance method, or Mahalanobis distance method to determine the application domain of the trained surface tension prediction model.
[0091] This application also provides an application scenario that applies the above-mentioned surface tension prediction method for non-ionic surfactants. Specifically, the surface tension prediction method for non-ionic surfactants provided in this embodiment can be applied in an agricultural spraying scenario. The agricultural spraying scenario includes a prediction link and a spraying link. The prediction link is used to predict the surface tension of the non-ionic surfactant, and the spraying link is used to prepare a pesticide formulation based on the predicted surface tension using the non-ionic surfactant and to perform spraying using the pesticide formulation. The surface tension prediction method for non-ionic surfactants provided in this embodiment belongs to the prediction link.
[0092] Embodiment 2.
[0093] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for predicting the surface tension of non-ionic surfactants.
[0094] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0095] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the method for predicting the surface tension of non-ionic surfactants in Embodiment 1.
[0096] Embodiment 3.
[0097] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, it implements the method for predicting the surface tension of non-ionic surfactants in Embodiment 1.
[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0100] In this text, specific examples are used to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for predicting the surface tension of a non-ionic surfactant, characterized in that, The method for predicting the surface tension of the non-ionic surfactant includes: Based on the molecular structure of the non-ionic surfactant to be predicted, molecular descriptors are calculated; the molecular descriptors include structural descriptors, empirical descriptors, topological descriptors, and electrostatic potential descriptors; Using the molecular descriptors as inputs, the surface tension of the non-ionic surfactant to be predicted is determined by using the trained surface tension prediction model.
2. The method for predicting the surface tension of a non-ionic surfactant according to claim 1, characterized in that, The structural descriptors include Labute approximate surface area, average molecular weight, exact molecular weight, number of heteroatoms, number of heavy atoms, total number of atoms, number of chiral centers, number of unspecified chiral centers, number of aromatic rings, total number of rings, number of aromatic heterocycles, number of aliphatic carbocycles, and proportion of SP3 hybridized carbon in total carbon; the empirical descriptors include logarithm of octanol / water partition coefficient and molecular refractive index; the topological descriptors include number of rotatable bonds, number of hydrogen bond donors, number of hydrogen bond acceptors, molecular connectivity index, symmetry and bond, chain saturation topological index, and quantum topological electronic index; the electrostatic potential descriptors include polar surface area of each molecule, Lipinski number of hydrogen bond acceptors, Lipinski number of hydrogen bond donors, hydrophilic / hydrophobic interaction, polarization, and electrostatic interaction.
3. The method for predicting the surface tension of a non-ionic surfactant according to claim 1, characterized in that, Before using the molecular descriptors as inputs and determining the surface tension of the non-ionic surfactant to be predicted by using the trained surface tension prediction model, the method for predicting the surface tension of the non-ionic surfactant further includes: Obtaining a training set; the training set includes the sample molecular structure and sample surface tension corresponding to each of a plurality of first sample non-ionic surfactants; Based on the sample molecular structure, sample molecular descriptors are calculated, and feature dimensionality reduction and feature selection are performed on the sample molecular descriptors to obtain the selected sample molecular descriptors; Using the selected sample molecular descriptors as inputs and the sample surface tension as labels, an initial surface tension prediction model is trained to obtain a trained surface tension prediction model.
4. The method for predicting the surface tension of a non-ionic surfactant according to claim 3, characterized in that, The molecular structure is in SMILES format. At this time, based on the molecular structure of the non-ionic surfactant to be predicted, molecular descriptors are calculated, specifically including: using the molecular structure of the non-ionic surfactant to be predicted as an input, and calculating molecular descriptors by using RDKit; The sample molecular structure is in SMILES format. At this time, based on the sample molecular structure, sample molecular descriptors are calculated, specifically including: using the sample molecular structure as an input, and calculating sample molecular descriptors by using RDKit.
5. The method for predicting the surface tension of a non-ionic surfactant according to claim 3, characterized in that, Performing feature dimensionality reduction and feature selection on the sample molecular descriptors to obtain the selected sample molecular descriptors specifically includes: Using the principal component analysis method to perform feature dimensionality reduction on the sample molecular descriptors to obtain the dimensionality-reduced sample molecular descriptors; Using the forward selection method to perform feature selection on the sample molecular descriptors and the dimensionality-reduced sample molecular descriptors to obtain the selected sample molecular descriptors.
6. The method for predicting the surface tension of a non-ionic surfactant according to claim 3, characterized in that, Using the molecular descriptors as inputs, the surface tension of the non-ionic surfactant to be predicted is determined by using the trained surface tension prediction model, specifically including: Based on the selected sample molecular descriptors, the molecular descriptors are screened to obtain the selected molecular descriptors; Using the selected molecular descriptors as inputs, the surface tension of the non-ionic surfactant to be predicted is determined by using the trained surface tension prediction model.
7. The method for predicting the surface tension of a non-ionic surfactant according to claim 3, characterized in that, The initial surface tension prediction model is a support vector regression model. At this time, using the selected sample molecular descriptors as inputs and the sample surface tension as labels, the initial surface tension prediction model is trained to obtain the trained surface tension prediction model, specifically including: Using the selected sample molecular descriptors as inputs and the sample surface tension as labels, the initial surface tension prediction model is trained by using the leave-one-out cross-validation method to obtain the trained surface tension prediction model.
8. The method for predicting the surface tension of a non-ionic surfactant according to claim 3, characterized in that, After obtaining the trained surface tension prediction model, the surface tension prediction method for the non-ionic surfactant further includes: Obtaining a validation set; the validation set includes the selected sample molecular descriptors and sample surface tensions corresponding to each of the second sample non-ionic surfactants in a plurality of second sample non-ionic surfactants; Using the training set to test the trained surface tension prediction model to obtain the internal test accuracy of the trained surface tension prediction model; Using the validation set to test the trained surface tension prediction model to obtain the external test accuracy of the trained surface tension prediction model; Using the leverage method, Euclidean distance method or Mahalanobis distance method to determine the application domain of the trained surface tension prediction model.
9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the surface tension prediction method for the non-ionic surfactant according to any one of claims 1-8.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the surface tension prediction method for the non-ionic surfactant according to any one of claims 1-8.
Citation Information
Patent Citations
Method for predicting surface tension of ionic liquid by inputting features into generalized machine learning model based on SMILES character string and group contribution
CN117747013A
Molecular antibacterial activity prediction method, system and device and storage medium
CN119724411A
QSPR Model Predicting Surface Tension of Liquid of Pure Organic Compound
KR1020120085173A
Methods and systems for fuel design using machine learning
US20240363203A1
Machine learning-based polymer surface energy prediction system
US20250087308A1