Method for predicting solubility parameter of pure compound by using multiple linear regression model
By screening molecular descriptors through a multiple linear regression model and establishing a compound solubility parameter prediction model, the problems of long calculation time and high cost in the existing technology are solved, and a fast and accurate solubility parameter prediction is achieved, which is applicable to compounds composed of a wide range of elements.
Patent Information
- Application Number
- CN202411354389.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-09-19
AI Technical Summary
Existing compound solubility parameter prediction methods have problems such as long calculation time, high cost and limited applicability. In particular, for compounds composed of specific elements, it is difficult to obtain solubility parameters quickly and accurately.
A multiple linear regression model is used to screen out suitable molecular descriptors and establish a solubility parameter prediction model for the compound. The optimal multiple linear regression model is gradually selected using the experimental data training set and test set, avoiding quantum chemical calculations, shortening the prediction time and improving accuracy.
The solubility parameters of compounds can be predicted with high accuracy in a short time, which is applicable to compounds composed of a wide range of elements, saving experimental costs and time, and supporting the smooth progress of industrial and academic research.
Smart Images

Figure CN120673899A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of physical chemistry called physical property prediction, and relates to a method for predicting the solubility parameter of a pure compound with high accuracy. Background Art
[0002] Today, humanity relies on a vast array of chemical compounds, including plastics, fibers, rubber, paints, fertilizers, pharmaceuticals, and fuels, and this trend is expected to intensify. According to the American Chemical Society (ACS), the total number of registered chemical substances as of 2023 is over 204 million. In comparison, the number of compounds for which even a single physical property is experimentally known is only in the tens of thousands, a tiny fraction of the total number of chemical substances. However, the physical properties of chemical compounds are essential for a better material life, including the development of new substances and drugs, the optimal design of chemical equipment, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment.
[0003] Specifically, understanding the exact values of various physical properties of a compound plays a decisive role in various decision-making matters throughout the production and consumption process, such as determining the rationality of the substance's use or designing synthesis and purification processes, setting methods and conditions for storage, transportation, use, and disposal. Therefore, it is of great significance both industrially and academically.
[0004] Currently, the most accurate way to obtain the relevant physical property values of the compound of interest is through experimental measurement. However, this requires considerable cost and time in various aspects, such as preparing purified samples and constructing an environment for accurate measurement. Sometimes, it may not be possible depending on the actual situation.
[0005] Therefore, as an alternative, many researchers have long been working on predicting the accurate values of various physical properties of compounds. Since physical property prediction has a long history and new prediction methods continue to emerge, various prediction models that differ from each other currently coexist according to physical properties, accuracy, and application range.
[0006] As a model for predicting physical properties, the main method currently widely known and used is the group contribution method determined by statistical methods. This method has the best performance for predicting physical properties of compounds with experimental values.
[0007] This group contribution method has previously achieved some impressive results, but due to a lack of theoretical basis, it can sometimes fail to calculate numerical values due to non-unique or even non-existent segmentation methods based on fragment form. Furthermore, as models are refined to improve predictive performance, they eventually become more complex and difficult to handle.
[0008] In the process of building prediction models, one of the alternatives to the group contribution method is the quantitative structure-property relationship (QSPR) method. The basis of this method is the assumption that the physical properties of a compound are functions of the structural characteristics of the molecule, and it uses a variety of molecular descriptors that reflect a variety of different structural characteristics. The types of molecular descriptors proposed so far in this method have reached thousands, ranging from simple molecular descriptors such as the number of carbon or hydrogen atoms in a molecule to complex molecular descriptors such as the shape, connection state, and electrochemical properties of the molecule. In addition, many types of calculation methods for molecular descriptors have been developed [Todeschini R., V. Consonni V., Molecular Descriptors for Chemoinformatics: Second, Revised and Enlarged Edition: Volume I / II, Wiley-VCH, 2009]. The QSPR prediction model is presented in the form of a function that includes these molecular descriptors and, sometimes, other physicochemical properties of the predicted compound (which are also functions of structural characteristics) as independent variables.
[0009] In this case, the most common function is the random variable (representer) X shown below: i The linear combination function, the coefficients c0 and c i It was mainly determined by multiple linear regression analysis based on experimental data.
[0010]
[0011] Another way to create a QSPR model is to use artificial neural networks. Artificial neural network technology is a technology that uses human neural cells to model intelligent machines and is currently a widely used information processing technology.
[0012] By training an artificial neural network using a sample set that binds various input values to the output values corresponding to those input values, the artificial neural network can establish general rules through learning, even without the rules or knowledge required to solve the problem. This allows it to output appropriate outputs even for unknown inputs. Therefore, artificial neural networks are widely used as a very useful tool in fields that lack basic theory, such as predicting the physical properties of compounds.
[0013] When using computers to uniformly calculate the values of more than 3,779 molecular descriptors of compounds required for the QSPR model for predicting physical properties, in order to calculate the electronic structure of the molecule, the electron energy solution is usually obtained by solving the Schrödinger equation. However, for systems with a large number of electrons, the calculation time required is very long, and the following attempts have been made in the prior art.
[0014] The calculation theories tried to select the best quantum mechanical calculation method are Hartree-Fock method [CCJ Roothan, Rev. Mod. Phys. 23, 69 (1951)], various post-Hartree-Fock methods [C. Moller and MSPlesset, Phys. Rev. 46, 618 (1934)], Gauss method combining Hartree-Fock and post-Hartree-Fock methods [LA Curtiss, K. Raghavachari, GW Trucks, and JA Popple, J. Chem. Phys. 94, 7221 (1991); LA Curtiss, K. Raghavachari, PC Redfern, V. Rassolov, and J.A.Pople, J.Chem.Phys.109,7764(1998)], density functional theory (DFT) [R.Seeger and J.A.Pople, J.Chem.Phys.66,3045(1977)] which uses electron density function instead of wave function with multi-dimensional perturbation term to consider the correlation between electrons in molecules composed of many electrons and uses total energy functional to find the ground state. However, these methods still require a very large amount of calculation, so when calculating large molecules, there are still difficulties in cost and time.
[0015] Furthermore, the solubility parameter, a physical property of interest in the present invention, can be explained as follows: in order for a solute to dissolve in a solvent, an equal solute-solvent attraction must be applied to the attraction between solute molecules or the attraction between solvent molecules. Such attraction between solute molecules and attraction between solvent molecules is the energy required to separate one molecule from its respective molecular cluster. This energy is called cohesive energy, and the cohesive energy per unit volume is called cohesive energy density (CED). The square root of the cohesive energy density is the solubility parameter (δ).
[0016] Two substances with similar solubility parameters have sufficient energy to disperse and mix. Conversely, two substances with different solubility parameters require more energy to disperse than they have to mix, resulting in the inability to mix. Various prediction models have been proposed to predict solubility parameters.
[0017] Currently, the Hildebrand formula is widely known and commonly used as a model for predicting solubility parameters [JH Hildebrand, RL Scott, The Solubility of Nonelectrolytes, 3rd ed., Reinhold Publishing (1950); Dover Publications (1964), Chap. VII, p. 129; Chap. XXIII, for the original definition, theory, and extensive examples Poling BE, Prausnitz JM, O'Connell JP].
[0018] The cohesive energy (U coh ) is usually determined by the heat of vaporization (ΔH vap ) to obtain.
[0019] U coh (T,P sat )=ΔH vap (T)-RT
[0020] Solubility parameter,
[0021] R is the ideal gas constant, ΔH vap= is the heat of vaporization, and LMV is the liquid molar volume. As a method for indirectly calculating the solubility parameter, the heat of vaporization and liquid molar volume in the mathematical formula are calculated using QSPR simulation, T = 298.15K.
[0022] In the literature [James E. Code, Andrew J. Holder and J. David Eick, Direct and Indirect Quantum Mechanical QSPR Hildebrand Solubility Parameter Models, QSARComb. Sci. 27, 2008, No. 7, 841-849], data from 56 compounds were reported using a method with a standard error of 0.69 (J / cm 3 ) 1 / 2 , a model with four molecular descriptors with a coefficient of determination of 0.97. However, few QSPR models for solubility parameters have been proposed so far, and the number of sample compounds proposed is limited or their applicability is limited to specific types of compounds.
[0023] Therefore, the present invention provides a method for predicting solubility parameters as follows: in order to shorten the calculation time that is a problem of the above-mentioned existing methods and improve the accuracy of predicting physical properties, molecular descriptors that can serve as independent variables are screened out from molecular descriptors, and without adding the use of artificial neural networks, an optimal multiple linear regression model for the solubility parameters of the compound is established to accurately predict the solubility parameters of the pure compound. Summary of the Invention
[0024] The technical problem to be solved by the present invention is to provide a more reliable QSPR model for the solubility parameters of compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As), based on more experimental data.
[0025] This model is expected to overcome the shortcomings of the model constructed by the group contribution method described above and to show fast and accurate prediction performance for a wider range of elements. The range of compounds to which the prediction model can be applied is limited to compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As). This is mainly because the model is based on experimental values and is therefore limited to the constituent elements of the compound for which experimental values exist, but the number of each atom is not limited.
[0026] As a predictive model capable of matching the composition of the constituent elements or molecular sizes of compounds with experimental values, it can be expanded to encompass a wider range of constituent elements and a greater number of atoms. Even within the aforementioned limited range of compounds, there are a vast number of compounds, including a significant number of industrially important ones. Based on the molecular structure of a compound, physical properties can be predicted in a very short time, facilitating the processing and analysis of physical property information on a large amount of compounds. Therefore, the present invention is expected to have significant beneficial effects on human society.
[0027] To achieve this purpose, an example of the present invention is a method for predicting the solubility parameters of pure compounds using a multiple linear regression model, the method comprising the following steps: (1) inputting the collected experimental data of the liquid sample compound; (2) preparing molecular descriptors of the liquid sample compound inputted above; (3) screening out the best molecular descriptors from step (2); (4) separating the experimental data of step (1) into a training set and a test set; (5) using the molecular descriptors of step (3) to find the best multiple linear regression model that meets the training set; (6) judging the rationality of the multiple linear regression model; (7) if the multiple linear regression model is not rational in step (6), repeating steps (5) and (6); if it is rational, testing the prediction performance of the model using the test set; (8) in the test of step (7), if the performance does not meet the standard, repeating steps (4) to (7); if it meets the standard, using the predicted value of the solubility parameter obtained by the multiple linear regression model as the solubility parameter value.
[0028] The liquid sample compound in step (1) is a compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
[0029] The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to solubility parameters.
[0030] In the step (3), the optimal molecular descriptor is an independent molecular descriptor having different values for all sample compounds.
[0031] The optimal molecular descriptor comprises one or more of the following: P1: C atomic percentage; P2: N atomic percentage; P3: O atomic percentage; P4: normalized spectral positive sum from Laplace matrix; P5: fifth-order spectral moment from chi matrix.
[0032] In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.
[0033] In step (5), the multiple linear regression model is explored by applying a stepwise selection method to the training set.
[0034] The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.
[0035] The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.
[0036] The test in step (7) is to test the multiple linear regression model using the test set.
[0037] The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.
[0038] The physical properties of compounds are essential for the development of new substances and drugs, the optimal design of chemical equipment, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment. Solubility parameters are also valuable in characterizing the intermolecular relationships within complex substances such as asphalt and crude oil, enabling prediction of solubility, which determines whether molecules are miscible with each other.
[0039] However, the number of compounds for which experimental values are currently known does not exceed several thousand at most, and obtaining data through experiments is sometimes extremely difficult due to the toxicity, instability, and difficulty in purification of the compounds.
[0040] From this perspective, the present invention, which can obtain highly accurate solubility parameter values for many compounds without conducting experiments, based solely on molecular information, not only saves the cost and time required for experiments, but also allows the values to be estimated even when experiments are not possible, thereby facilitating research and development activities in related industries. Furthermore, appropriate information can be provided to all places such as academia and government that require such values, enabling such activities to be carried out more smoothly.
[0041] Furthermore, although three-dimensional molecular descriptors are insufficient because quantum chemical calculations are not performed, the physical properties of unknown compounds can be predicted in a very short time. Furthermore, the prediction time can be further shortened without using additional models such as artificial neural networks, thereby enabling real-time prediction of physical properties. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart showing a process of constructing a solubility parameter QSPR model provided by the present invention.
[0043] Figure 2 is a scatter plot showing the prediction performance of the multiple linear regression model.
[0044] Figure 3 It is a bar chart of 1514 experimental data of the QSPR model. DETAILED DESCRIPTION
[0045] Hereinafter, the method for predicting the solubility parameters of pure compounds using the multiple linear regression model according to the present invention will be described in detail with reference to the accompanying drawings so that persons with general knowledge in the technical field to which the present invention belongs can easily implement the present invention.
[0046] In the process of describing the present invention, if it is determined that the detailed description of the related known technology may cause unnecessary obscurity to the gist of the present invention, the detailed description will be omitted.
[0047] When using the words "including," "having," "forming," "comprising," and the like in this specification, other parts may be added unless "only" is used. When describing a component in the singular, the plural is included unless otherwise specified.
[0048] Figure 1 A flow chart schematically illustrating the process of constructing a QSPR model of solubility parameters.
[0049] The first task in building the model is to collect experimental data and analyze and classify them as specified in the first step.
[0050] For the present invention, we conducted an extensive search of all available reference materials and documents, including various papers, monographs, and websites, to collect experimental data on solubility parameters for compounds meeting the criteria of the present invention. We conducted a comprehensive review to determine whether the collected data represented truly reasonable values for model construction. We then carefully analyzed and modified or deleted data for compounds with the following conditions: (i) non-experimental values; (ii) data labeling errors; (iii) values for the same compound that differed significantly; (iv) values that deviated unreliably from those of other similar compounds; or (v) values for which molecular descriptors were difficult to readily generate. Finally, we selected data for a total of 2,037 compounds.
[0051] In addition, when constructing the physical property prediction model, the sample compounds were divided into gaseous compounds, liquid compounds, and solid compounds including one or more of the elements C, H, N, O, S, F, Cl, Br, I, Si, P, and As, and explored separately, but only the model for the liquid compounds was established.
[0052] Since the above three classifications showed more effective prediction performance, they were divided into 104 gaseous compounds, 1514 liquid compounds, and 419 solid compounds and explored separately. However, the solubility parameters of gaseous compounds can be calculated using mathematical formulas. The solid compound model has better prediction performance locally because it does not conform to the independent variables, so only the liquid compound model was established. In addition, in the present invention, "compound" refers to a substance formed by molecules composed of no more than 12 elements such as hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As).
[0053] The second step is to prepare molecular descriptor values for these compounds. Molecular descriptors suitable for use as model variables are calculated using the following method for the values of over 3,779 molecular descriptors: a file (MOL or SDF file) containing structural information, including the X, Y, and Z coordinates of each molecule's constituent elements and the bonding patterns between them, is input into a computer program. Based on this input, molecular descriptors representing the diversity of molecules are calculated using various theoretical mathematical formulas, moving beyond simple element and bond counts. Furthermore, the process of extracting these 3,779 molecular descriptors is repeated for all sample compounds used to construct the physical property prediction model.
[0054] The theoretical mathematical formulas used in the calculation are known in the art and are used in various forms, for example, a large number of theoretical mathematical formulas including Chi Connectivity Indices, Topological Descriptors, Information Indices, Walk and Path Counts, 2D Autocorrelation Descriptors, etc.
[0055] In the best results, not only physical property information can be obtained, but also molecular descriptors that reflect the characteristics of the molecules and are expressed in a variety of meaningful numerical values can be obtained.
[0056] For pure substances, there are molecular descriptors that can characterize both two-dimensional and three-dimensional structures. However, three-dimensional molecular descriptors are affected by molecular structure and are more complex to compute. Therefore, in this invention, in order to rapidly process physical property predictions without performing quantum mechanical calculations for structure optimization, three-dimensional molecular descriptors are not used. This approach significantly reduces the computational time and cost of quantum mechanical calculations.
[0057] Furthermore, in the third step, after calculating the molecular descriptor values, it is necessary to screen out inappropriate molecular descriptor values. Specifically, by applying a theoretical mathematical formula to the molecular descriptors of the sample compounds, the system screens out molecular descriptor values that yield the same value for all sample compounds and therefore cannot serve as independent variables in the model. This screening of molecular descriptors prevents the inclusion of irrelevant molecular descriptors in the prediction model, thereby improving model reliability while reducing the number of molecular descriptors, the number of calculations, and the computational effort, ultimately shortening the computational time required to find the optimal model.
[0058] The best molecular descriptors obtained in the present invention include C atomic percentage, N atomic percentage, O atomic percentage, normalized spectral sum from Laplacian matrix and fifth-order spectral moment from chi matrix.
[0059] In the fourth step, the sample compounds are divided into a training set and a test set. The training set will be used to find the prediction model, and the test set will be used to test the predictive performance of the determined model. While taking care not to bias the distribution of similar molecules to one side, the training set and the test set are divided in a ratio of 5-8:2-5, preferably 7:3.
[0060] Then, in the fifth step, the optimal multiple linear regression model is found based on the training set. Here, "optimal" is used in a relative sense, meaning that it can be found in a relatively short time while having performance very close to the optimal solution in an absolute sense.
[0061] The reason for not directly finding the optimal solution is that the calculation time is long. For example, when the number of suitable molecular descriptors among 3779 molecular descriptors is 1700, the total number of multiple linear regression models that can be made by extracting 5 different molecular descriptors is It is therefore practically impossible to investigate them all.
[0062] In order to obtain useful results within a limited time, a stepwise selection method is adopted in the present invention.
[0063] The stepwise selection method is a common method for selecting variables from data with many independent variables. During each step of creating a regression equation, the following steps are repeated: Among the independent variables not already in the equation, variables with high correlations with the dependent variable are added by entering a benchmark, and among the independent variables already in the equation, variables with low correlations with the dependent variable are removed by removing the benchmark. If no further variables are added or removed during this process, the selection of independent variables ends, and the model based on the stepwise selection method is finally completed.
[0064] In other words, a method of using the so-called hypothesis test to determine whether the hypothesis about the population is statistically correct. The significance level is generally set to 0.05 (5%), and the p-value (p-value) as the probability of significance is calculated. When the p-value is less than 0.05, the null hypothesis is rejected and the alternative hypothesis is adopted. By performing hypothesis testing on the selected variables, the process of adopting newly added variables or removing existing variables is repeated. Ultimately, when there are no new variables to be added or removed, the stepwise selection process ends. After the stepwise selection is successful, the molecular descriptor information obtained will perform a more advanced search, and a model with reliable accuracy will be obtained after the statistical test is passed, as shown below:
[0065]
[0066] Where P represents the molecular descriptor after stepwise selection, Y is the property to be predicted, and C is the parameter estimated from the experimental data by the multivariate regression method.
[0067] Preferably, the least square residual is used to estimate the parameter C. The specific calculation process is as follows:
[0068]
[0069] Where n is the number of substances to be predicted and Pn is the number of optimal molecular descriptors.
[0070] Once the best multiple linear regression model is selected, the rationality of the model is determined in the sixth step, which is the next step. Specifically, the rationality of the model can be tested using the t-test method, and the calculation process is as follows:
[0071]
[0072] If the t-test value for a molecular descriptor included in the model is found to be poor according to the above method, the process returns to the previous step and searches for another model. For example, if the number of sample compounds is 1005 and the selected model consists of five molecular descriptors, if the t-test value for one molecular descriptor is 3.3 or higher, the probability that the molecular descriptor is unrelated to the corresponding physical property is less than 0.1%.
[0073] In the present invention, when there is a molecular descriptor with a t-test value less than about 3, the selected model is discarded and other models are sought. In addition, the situation that the value of a molecular descriptor for the sample compound is the same except for a few compounds cannot be regarded as a model with reliability, so measures are also taken. Usually, if the number of molecular descriptors included in the model is increased, the prediction performance will improve, but the above-mentioned problems will occur. Therefore, usually, the final model is obtained by repeatedly performing these steps by changing the number of molecular descriptors and through multiple trial and error. If the selected model no longer has problems, then proceed to the next step.
[0074] Next, in the seventh step, the prediction performance of the found model is evaluated using a test set that does not participate in forming the model.
[0075] If significant performance degradation or significant deviations in prediction are observed in the training set, the fourth step is to readjust the training and test sets before proceeding to the subsequent steps. If the difference between the training and test sets does not exceed 20% of the mean absolute error (AAE) obtained for the training set, the prediction performance is considered satisfactory and a multiple linear regression model for the solubility parameters is established.
[0076] The results of the model finally established through this process are briefly shown in Table 1 below.
[0077] Table 1
[0078] Main contents of the QSPR prediction model for compound solubility parameters
[0079]
[0080]
[0081]
[0082]
[0083]
[0084] Example
[0085] This example is used to further illustrate the accuracy of the present invention in predicting the solubility parameters of pure compounds, but the present invention is not limited thereto.
[0086] Table 2 below shows some of the results of predicting the solubility parameters of pure compounds using the model of the present invention:
[0087] Table 2
[0088]
[0089]
[0090] In order to confirm the excellence of the present invention, the multiple linear regression model of the present invention was used to predict the solubility parameters of 1514 liquid compounds with known experimental values, and a determination coefficient value of 0.956 and a solubility coefficient value of 250.9078 (cal / m 3 ) 1 / 2 The average absolute error value of , which shows that the prediction performance is excellent.
[0091] Figure 2 This is a scatter plot showing the predictive performance of a multiple linear regression model. It evaluates how closely experimental values are aligned with predicted values. The X-axis represents the experimental values, and the Y-axis represents the predicted values. The experimental and predicted values are plotted at the (X, Y) coordinates with dots (hollow blue circles), and the baseline representing Y=X is marked with a diagonal line (solid red line).
[0092] Points located on the baseline indicate that the predicted value is the same as the experimental value. Points located below the baseline indicate that the predicted value is smaller than the experimental value. Points located above the baseline indicate that the predicted value is larger than the experimental value. For example, if the experimental and predicted values are (100, 100), they are on the baseline; if the experimental and predicted values are (100, 90), they are below the baseline; and if the experimental and predicted values are (100, 10), they are above the baseline.
[0093] exist Figure 2In the parity plot, we can observe that there are some points that are relatively separated from the baseline in the middle. However, in most areas, the points are concentrated on the baseline. The points are distributed on the baseline in a spindle shape and spread thinner upward and downward rather than widely. Therefore, the predicted value can be said to be close to the experimental value as a whole.
[0094] In addition, in the experimental data of 1514 liquid compounds, the average error of the known experimental error is about 7.82%, Figure 3 The error between the experimental and predicted values is plotted in a bar graph centered around this value. In this graph, the multiple linear regression model predicts solubility parameter values within the average experimental error with a probability of 94.78%, demonstrating its high predictive performance.
[0095] The present invention is not limited to the above-mentioned embodiments. Without departing from the gist of the invention claimed for protection in the claims, anyone with general knowledge in the technical field to which the present invention belongs can make various modified implementations, and the above-mentioned changes fall within the scope of the claims.
Claims
1. A method for predicting the solubility parameters of pure compounds using a multiple linear regression model, the method comprising the following steps: (1) Input the collected experimental data of the liquid sample compound; (2) preparing a molecular descriptor of the liquid sample compound input above; (3) Select the best molecular descriptor from step (2); (4) separating the experimental data in step (1) into a training set and a test set; (5) constructing an optimal multiple linear regression model that matches the training set using the molecular descriptors described in step (3); (6) Determine the rationality of the above multiple linear regression model; (7) If the multiple linear regression model is not reasonable in step (6), repeat steps (5) and (6). If it is reasonable, test the prediction performance of the model using the test set; (8) In the test of step (7), if the performance does not meet the standard, steps (4) to (7) are repeated. If the standard is met, the predicted value of the solubility parameter obtained by the above-mentioned multiple linear regression model is used as the solubility parameter value.
2. The method according to claim 1, wherein The liquid sample compound in step (1) is a compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
3. The method according to any one of claims 1 to 2, wherein The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to solubility parameters.
4. The method according to any one of claims 1 to 2, wherein In the step (3), the optimal molecular descriptor is an independent molecular descriptor having different values for all sample compounds.
5. The method of claim 4, wherein the optimal molecular descriptor comprises one or more of the following: P1: C atomic percentage; P2: N atomic percentage; P3: O atomic percentage; P4: normalized spectral positive sum from the Laplace matrix; P5: fifth-order spectral moment from the chi matrix.
6. The method according to any one of claims 1 to 2, wherein In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.
7. The method according to any one of claims 1 to 2, wherein In step (5), the multiple linear regression model is explored by applying a stepwise selection method to the training set.
8. The method of claim 7, wherein: The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.
9. The method according to any one of claims 1 to 2, wherein the rationality of step (6) is determined by a t-test value, wherein: When the t-test value is greater than or equal to 3, the model is reasonable, otherwise it is not reasonable.
10. The method according to any one of claims 1 to 2, wherein: The test in step (7) is to test the multiple linear regression model using the test set.
11. The method according to any one of claims 1 to 2, wherein: The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.