Method for predicting critical temperature of pure compound by using multiple linear regression model
By screening molecular descriptors and establishing a multiple linear regression model, the problems of long calculation time and low accuracy in compound critical temperature prediction are solved, and fast and accurate compound critical temperature prediction is achieved, especially for compounds composed of specific elements.
Patent Information
- Application Number
- CN202411354078.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies for predicting the critical temperature of compounds suffer from long calculation time, high cost, and low accuracy. This is especially true for compounds composed of 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As). Effective models and methods are lacking.
By screening suitable molecular descriptors, establishing a multiple linear regression model, and using experimental data to construct a critical temperature prediction model for compounds, the use of artificial neural networks is avoided, the calculation time is shortened, and the accuracy is improved.
The critical temperature of compounds can be efficiently predicted in a short time, the computational cost is reduced, and it is applicable to compounds composed of a wide range of elements. The prediction performance and accuracy are improved, especially the coefficient of determination of 0.9849 and the mean absolute error of 10.649K are achieved through the multiple linear regression model.
Smart Images

Figure CN120673894A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of physical property prediction, which is one of the fields of physical chemistry, and relates to a method for predicting the critical temperature, which is one of the various physical properties of a compound, with high accuracy. Background Art
[0002] Today, humanity relies on a vast array of chemical compounds, including plastics, fibers, rubber, paints, fertilizers, pharmaceuticals, and fuels, and this trend is expected to intensify. According to the American Chemical Society (ACS), the total number of registered chemical substances as of 2023 is over 204 million. In comparison, the number of compounds for which even a single physical property is experimentally known is only in the tens of thousands, a tiny fraction of the total number of chemical substances. However, the physical properties of chemical compounds are essential for a better material life, including the development of new substances and drugs, the optimal design of chemical equipment, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment.
[0003] Specifically, understanding the exact values of various physical properties of a compound plays a decisive role in various decision-making matters throughout the production and consumption process, such as determining the rationality of the substance's use or designing synthesis and purification processes, setting methods and conditions for storage, transportation, use, and disposal. Therefore, it is of great significance both industrially and academically.
[0004] Currently, the most accurate way to obtain the relevant physical property values of the compound of interest is through experimental measurement. However, this requires considerable cost and time in various aspects, such as preparing purified samples and constructing an environment for accurate measurement, and may sometimes be impossible to achieve depending on the actual situation.
[0005] Therefore, as an alternative, many researchers have long been working on predicting the accurate values of various physical properties of compounds. Since physical property prediction has a long history and new prediction methods continue to emerge, various prediction models that differ from each other currently coexist according to physical properties, accuracy, and application range.
[0006] As a model for predicting physical properties, the main method currently widely known and used is the group contribution method determined by statistical methods. This method has the best performance for predicting physical properties of compounds with experimental values.
[0007] The group contribution method has previously achieved some impressive results. However, due to its lack of theoretical basis, it sometimes encounters problems with segmentation based on fragment form, such as non-unique or even non-existent segmentation, resulting in inability to calculate numerical values. Furthermore, the need to continuously refine the model to improve predictive performance ultimately leads to increased complexity and difficulty in handling.
[0008] In the process of building prediction models, one of the alternatives to the group contribution method is the quantitative structure-property relationship (QSPR) method. The basis of this method is the assumption that the physical properties of a compound are functions of the structural characteristics of the molecule, and it uses a variety of molecular descriptors that reflect a variety of different structural characteristics. The types of molecular descriptors proposed so far in this method have reached thousands, ranging from simple molecular descriptors such as the number of carbon or hydrogen atoms in a molecule to complex molecular descriptors such as the shape, connection state, and electrochemical properties of the molecule. In addition, many types of calculation methods for molecular descriptors have been developed [Todeschini R., V. Consonni V., Molecular Descriptors for Chemoinformatics: Second, Revised and Enlarged Edition: Volume I / II, Wiley-VCH, 2009]. The QSPR prediction model is presented in the form of a function that includes these molecular descriptors and, sometimes, other physicochemical properties of the predicted compound (which are also functions of structural characteristics) as independent variables.
[0009] In this case, the most common function is the random variable (representer) X shown below: i The linear combination function, the coefficients c0 and c i It was mainly determined by multiple linear regression analysis based on experimental data.
[0010]
[0011] Another way to create a QSPR model is to use artificial neural networks. Artificial neural network technology is a widely used information processing technology that uses models based on human neural cells to create machines with artificial intelligence.
[0012] By training an artificial neural network using a sample set that binds various input values to the output values corresponding to those input values, the artificial neural network can establish general rules through learning, even without the rules or knowledge required to solve the problem. This allows it to output appropriate outputs even for unknown inputs. Therefore, artificial neural networks are widely used as a very useful tool in fields that lack basic theory, such as predicting the physical properties of compounds.
[0013] When using computers to uniformly calculate the values of more than 3,779 molecular descriptors of compounds required for the QSPR model for predicting physical properties, in order to calculate the electronic structure of the molecule, the electron energy solution is usually obtained by solving the Schrödinger equation. However, for systems with a large number of electrons, the calculation time required is very long, and the following attempts have been made in the prior art.
[0014] The calculation theories tried to select the best quantum mechanical calculation method are Hartree-Fock method [CCJ Roothan, Rev. Mod. Phys. 23, 69 (1951)], various post-Hartree-Fock methods [C. Moller and MSPlesset, Phys. Rev. 46, 618 (1934)], Gauss method combining Hartree-Fock and post-Hartree-Fock methods [LA Curtiss, K. Raghavachari, GW Trucks, and JA Popple, J. Chem. Phys. 94, 7221 (1991); LA Curtiss, K. Raghavachari, PC Redfern, V. Rassolov, and J.A.Pople, J.Chem.Phys.109,7764(1998)], density functional theory (DFT) [R.Seeger and J.A.Pople, J.Chem.Phys.66,3045(1977)] which uses electron density function instead of wave function with multi-dimensional perturbation term to consider the correlation between electrons in molecules composed of many electrons and uses total energy functional to find the ground state. However, these methods still require a very large amount of calculation, so when calculating large molecules, there are still difficulties in cost and time.
[0015] Furthermore, few QSPR prediction models for the critical temperature of interest to the present invention have been proposed, and those that have been proposed have the disadvantages of being limited to a small number of sample compounds or being applicable to specific types of compounds. A linear regression analysis model using 12 molecular descriptors for 1230 compounds was proposed in the literature [Srinivasa S. Godavarthy, Robert L. Robinson Jr., Khaled A. M. Gaset, Improved structure-property relationship models for prediction of critical properties, Fluid Phase Equilibria 264 (2008) 122–136], with a coefficient of determination of 0.913 and a mean absolute error of 16.1 K, indicating unsatisfactory prediction performance.
[0016] While some QSPR models using artificial neural networks have been proposed for critical temperature prediction, most use a small number of sample compounds or are limited to specific series of compounds. Some of these models and their results were presented in [Srinivasa S. Godavarthy, Robert L. Robinson Jr., Khaled A. M. Gaset, Improved structure-property relationship models for prediction of critical properties, Fluid Phase Equilibria 264 (2008) 122–136]. The nonlinear analysis model showed improved predictive performance, with a coefficient of determination of 0.995 and a mean absolute error of 3.7 K.
[0017] Therefore, the present invention provides a method for predicting critical temperature as follows: in order to shorten the calculation time which is a problem of the above-mentioned existing methods and improve the accuracy of predicting physical properties, molecular descriptors that can serve as independent variables are screened out from molecular descriptors, and without adding the use of artificial neural networks, the accurate critical temperature of the pure compound is predicted by establishing an optimal multiple linear regression model for the critical temperature of the compound. Summary of the Invention
[0018] The technical problem to be solved by the present invention is to provide a more reliable QSPR model for the critical temperature of compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As), based on more experimental data.
[0019] This model is expected to overcome the shortcomings of the model constructed by the group contribution method described above and to show fast and accurate prediction performance for a wider range of elements. The range of compounds to which the prediction model can be applied is limited to compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As). This is mainly because the model is based on experimental values and is therefore limited to the constituent elements of the compound for which experimental values exist, but the number of each atom is not limited.
[0020] As a predictive model capable of matching the composition of the constituent elements or molecular sizes of compounds with experimental values, it can be expanded to encompass a wider range of constituent elements and a greater number of atoms. Even within the aforementioned limited range of compounds, there are a vast number of compounds, including a significant number of industrially important ones. Based on the molecular structure of a compound, physical properties can be predicted in a very short time, facilitating the processing and analysis of physical property information on a large amount of compounds. Therefore, the present invention is expected to have significant beneficial effects on human society.
[0021] In order to achieve such a purpose, an example of the present invention is a method for calculating the critical temperature of a compound by a multiple linear regression model, comprising the following steps: (1) inputting the collected experimental data of the sample compound; (2) preparing the molecular descriptors of the sample compound inputted above; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data into a training set and a test set; (5) constructing the best multiple linear regression model that conforms to the training set using the molecular descriptors described in step (3); (6) judging the rationality of the multiple linear regression model; (7) if the multiple linear regression model is not rational in step (6), repeating steps (5) and (6); if it is rational, testing the prediction performance of the model using the test set; (8) in the test of step (7), if the performance does not meet the standard, repeating steps (4) to (7); if it meets the standard, using the predicted value of the critical temperature obtained by using the multiple linear regression model as the actual critical temperature value.
[0022] The sample compound in step (1) is a pure compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
[0023] The sample compound in step (1) is one of the following: a. a hydrocarbon compound containing carbon and hydrogen elements; b. a non-hydrocarbon compound containing one or more of carbon, hydrogen, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
[0024] The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to the critical temperature.
[0025] The optimal molecular descriptor of step (3) includes one or more of the following: P1: number of double bonds; P2: number of sp hybridized carbon atoms; P3: number of circuits; P4: number of eleven-membered rings; P5: distance / detour ring index of order 10; P6: average atomic Sanderson electronegativity based on the proportion of carbon atoms; P7: average first ionization potential based on the proportion of carbon atoms; P8: number of rotatable bonds; P9: percentage of halogen atoms.
[0026] The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values for all sample compounds.
[0027] In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.
[0028] The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to the training set to construct a multiple linear regression model.
[0029] The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.
[0030] The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.
[0031] The test in step (7) is to test the multiple linear regression model using the test set.
[0032] The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.
[0033] The physical properties of compounds are essential for the development of new substances and drugs, the optimization of chemical equipment design, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment. In particular, the critical temperature and critical pressure, both of which are important benchmarks for the supercritical state, are the points at which the gas-liquid equilibrium curve reaches its maximum. These properties are required for precise values in commonly used programs such as Aspen Plus and Pro / II, widely known for their optimization of chemical equipment design. They also provide important reference points when attempting to predict the values of various physical properties through correlations based on the principle of corresponding states. However, the number of compounds for which experimental values are currently known does not exceed several thousand, and obtaining these data experimentally can be extremely difficult due to factors such as the compound's toxicity, instability, and difficulty in purification.
[0034] From this point of view, the present invention, which can obtain highly accurate critical temperature values for many compounds using only information on molecules without conducting experiments, not only saves the cost and time required for experiments, but also makes it possible to estimate their values even when experiments are not possible, thereby facilitating research and development activities in related industries. Furthermore, appropriate information can be provided to all places such as academia and government that require their values, allowing such activities to be carried out more smoothly.
[0035] Furthermore, although three-dimensional molecular descriptors are insufficient because quantum chemical calculations are not performed, the physical properties of unknown compounds can be predicted in a very short time. Furthermore, the prediction time can be further shortened without using additional models such as artificial neural networks, thereby enabling real-time prediction of physical properties. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart schematically illustrating a process of constructing a QSPR model of critical temperature.
[0037] Figure 2 is a scatter plot showing the prediction performance of the multiple linear regression model.
[0038] Figure 3 It is a bar chart of 2403 experimental data of the QSPR model. DETAILED DESCRIPTION
[0039] Hereinafter, a method for predicting the critical temperature of a pure compound using a multiple linear regression model according to the present invention will be described in detail with reference to the accompanying drawings so that persons having ordinary knowledge in the technical field to which the present invention pertains can easily implement the present invention.
[0040] In the process of describing the present invention, if it is determined that the detailed description of the related known technology may cause unnecessary obscurity to the gist of the present invention, the detailed description will be omitted.
[0041] When using the words "including," "having," "forming," "comprising," and the like in this specification, other parts may be added unless "only" is used. When describing a component in the singular, the plural is included unless otherwise specified.
[0042] Figure 1 A flowchart schematically illustrates the process of constructing a QSPR model of critical temperature.
[0043] The first task in building the model is to collect experimental data and analyze and classify them as specified in the first step.
[0044] For the present invention, we conducted an extensive search of all available references and materials, including various papers, monographs, and websites, to collect experimental data on critical temperatures for compounds meeting the criteria of the present invention. We conducted a comprehensive review to determine whether the collected data represented truly reasonable values for model construction. We carefully analyzed and modified or deleted data for compounds with the following conditions: (i) non-experimental values; (ii) data labeling errors; (iii) values for the same compound that differed significantly; (iv) values that deviated unreliably from those of other similar compounds; or (v) values for which molecular descriptors were difficult to readily generate. Finally, we selected data for a total of 2,403 compounds.
[0045] In addition, when constructing the physical property prediction model, the sample compounds were divided into (a) hydrocarbon compounds composed only of C and H and (b) non-hydrocarbon compounds containing one or more of the elements C, H, N, O, S, F, Cl, Br, I, Si, P, and As, and models were established for each type.
[0046] Because the two aforementioned classifications showed more effective prediction performance, they were divided into 793 hydrocarbon compounds and 1,610 non-hydrocarbon compounds, and models were established for each. Furthermore, in this invention, "compound" refers to a substance formed by molecules composed of no more than 12 elements: hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As).
[0047] The second step is to prepare molecular descriptor values for these compounds. Molecular descriptors suitable for use as model variables are calculated using the following method for the values of over 3,779 molecular descriptors: a file (MOL or SDF file) containing structural information, including the X, Y, and Z coordinates of each molecule's constituent elements and the bonding patterns between them, is input into a computer program. Based on this input, molecular descriptors representing the diversity of molecules are calculated using various theoretical mathematical formulas, moving beyond simple element and bond counts. Furthermore, the process of extracting these 3,779 molecular descriptors is repeated for all sample compounds used to construct the physical property prediction model.
[0048] The theoretical mathematical formulas used in the calculation are known in the art and are used in various forms, for example, a large number of theoretical mathematical formulas including Chi Connectivity Indices, Topological Descriptors, Information Indices, Walk and Path Counts, 2D Autocorrelation Descriptors, etc.
[0049] In the best results, not only physical property information can be obtained, but also molecular descriptors that reflect the characteristics of the molecules and are expressed in a variety of meaningful numerical values can be obtained.
[0050] For pure substances, there are molecular descriptors that can characterize both two-dimensional and three-dimensional structures. However, three-dimensional molecular descriptors are affected by molecular structure and are more complex to compute. Therefore, in this invention, in order to rapidly process physical property predictions without performing quantum mechanical calculations for structure optimization, three-dimensional molecular descriptors are not used. This approach significantly reduces the computational time and cost of quantum mechanical calculations.
[0051] Furthermore, in the third step, after calculating the molecular descriptor values, it is necessary to screen out inappropriate molecular descriptor values. Specifically, by applying a theoretical mathematical formula to the molecular descriptors of the sample compounds, the system screens out molecular descriptor values that yield the same value for all sample compounds and therefore cannot serve as model independent variables. This screening of molecular descriptors prevents the inclusion of irrelevant molecular descriptors in the prediction model, thereby improving model reliability while reducing the number of molecular descriptors, the number of calculations, and the computational effort required to find the optimal model.
[0052] In the present invention, when the sample compound is (a) a hydrocarbon composed only of C and H, the optimal molecular descriptors obtained include the number of double bonds, the number of sp hybridized carbon atoms, the number of rings, the number of eleven-membered rings, and the 10th-order distance / detour ring index; when the sample compound is (b) a non-hydrocarbon including one or more of the elements C, H, N, O, SF, Cl, Br, I, Si, P, and As, the optimal molecular descriptors obtained include the number of double bonds, the number of rotatable bonds, the average atomic Sanderson electronegativity based on the proportion of carbon atoms, the average number of rotatable bonds with the first ionization potential based on the proportion of carbon atoms, and the percentage of halogen atoms.
[0053] In the fourth step, the sample compounds are divided into a training set and a test set. The training set will be used to find the prediction model, and the test set will be used to test the predictive performance of the determined model. While taking care not to bias the distribution of similar molecules to one side, the training set and test set are divided in a ratio of 5-8:2-5, preferably 7:3.
[0054] Then, in the fifth step, the optimal multiple linear regression model is found based on the training set. Here, "optimal" is used in a relative sense, meaning that it can be found in a relatively short time while having performance very close to the optimal solution in an absolute sense.
[0055] The reason for not directly finding the optimal solution is that the calculation time is long. For example, when the number of suitable molecular descriptors among 3779 molecular descriptors is 1700, the total number of multiple linear regression models that can be made by extracting 5 different molecular descriptors is It is therefore practically impossible to investigate them all.
[0056] In order to obtain useful results within a limited time, a stepwise selection method is adopted in the present invention.
[0057] The stepwise selection method is a common method for selecting variables from data with many independent variables. During each step of creating a regression equation, the following steps are repeated: Among the independent variables not already in the equation, variables with high correlations with the dependent variable are added by entering a benchmark, and among the independent variables already in the equation, variables with low correlations with the dependent variable are removed by removing the benchmark. If no further variables are added or removed during this process, the selection of independent variables ends, and the model based on the stepwise selection method is finally completed.
[0058] In other words, a method of using the so-called hypothesis test to determine whether the hypothesis about the population is statistically correct. The significance level is generally set to 0.05 (5%), and the p-value (p-value) as the probability of significance is calculated. When the p-value is less than 0.05, the null hypothesis is rejected and the alternative hypothesis is adopted. By performing hypothesis testing on the selected variables, the process of adopting newly added variables or removing existing variables is repeated. Ultimately, when there are no new variables to be added or removed, the stepwise selection process ends. After the stepwise selection is successful, the molecular descriptor information obtained will perform a more advanced search, and a model with reliable accuracy will be obtained after the statistical test is passed, as shown below:
[0059]
[0060] Where P represents the molecular descriptor after stepwise selection, Y is the property to be predicted, and C is the parameter estimated from the experimental data by the multivariate regression method.
[0061] Preferably, the least square residual is used to estimate the parameter C. The specific calculation process is as follows:
[0062]
[0063] Where n is the number of substances to be predicted and Pn is the number of optimal molecular descriptors.
[0064] Once the best multiple linear regression model is selected, the rationality of the model is determined in the sixth step, which is the next step. Specifically, the rationality of the model can be tested using the t-test method, and the calculation process is as follows:
[0065]
[0066]
[0067] If the t-test value for a molecular descriptor included in the model is found to be poor according to the above method, the process returns to the previous step and searches for another model. For example, if the number of sample compounds is 1005 and the selected model consists of five molecular descriptors, if the t-test value for one molecular descriptor is 3.3 or higher, the probability that the molecular descriptor is unrelated to the corresponding physical property is less than 0.1%.
[0068] In the present invention, when there is a molecular descriptor with a t-test value less than about 3, the selected model is discarded and other models are sought. In addition, the value of a molecular descriptor for a sample compound is the same except for a few compounds, which cannot be considered a model with reliability, so measures are also taken. Generally, if the number of molecular descriptors included in the model is increased, the prediction performance will improve, but the problems mentioned above will occur. Therefore, generally, the final model is obtained by changing the number of molecular descriptors and repeatedly performing these steps through multiple trial and error. If the selected model no longer presents problems, the next step is entered.
[0069] Next, in the seventh step, the prediction performance of the found model is evaluated using a test set that does not participate in forming the model.
[0070] If problems such as a significant decrease in prediction performance or significant deviations in prediction are observed in the training set, the fourth step is to readjust the training and test sets before proceeding to the subsequent steps. If the difference between the training and test sets does not exceed 20% of the mean absolute error (AAE) obtained for the training set, the prediction performance is considered satisfactory and a multiple linear regression model for the critical temperature is established.
[0071] The results of the model finally established through this process are briefly shown in Tables 1 and 2 below.
[0072] Table 1
[0073] Main contents of the QSPR prediction model for the critical temperature of hydrocarbon compounds
[0074]
[0075]
[0076] Table 2 Main contents of the QSPR prediction model for the critical temperature of non-hydrocarbon compounds
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084] Example
[0085] This example is used to further illustrate the accuracy of the present invention in predicting the critical temperature of pure compounds, but the present invention is not limited thereto.
[0086] Table 3 below shows some results of predicting the critical temperature of pure compounds using the model of the present invention:
[0087] Table 3
[0088]
[0089]
[0090]
[0091] To confirm the excellence of the present invention, the multiple linear regression model of the present invention was used to predict the critical temperatures of 2403 compounds with known experimental values, resulting in a coefficient of determination value of 0.9849 and a mean absolute error value of 10.649 K, indicating excellent predictive performance.
[0092] Figure 2 This is a scatter plot showing the predictive performance of a multiple linear regression model. It evaluates how closely experimental values are aligned with predicted values. The X-axis represents the experimental values, and the Y-axis represents the predicted values. The experimental and predicted values are plotted at the (X, Y) coordinates with dots (hollow blue circles), and the baseline representing Y=X is marked with a diagonal line (solid red line).
[0093] Points located on the baseline indicate that the predicted value is the same as the experimental value. Points located below the baseline indicate that the predicted value is smaller than the experimental value. Points located above the baseline indicate that the predicted value is larger than the experimental value. For example, if the experimental and predicted values are (100, 100), they are on the baseline; if the experimental and predicted values are (100, 90), they are below the baseline; and if the experimental and predicted values are (100, 10), they are above the baseline.
[0094] exist Figure 2 In the parity plot, we can observe that there are some points that are relatively separated from the baseline in the middle. However, in most areas, the points are concentrated on the baseline. The points are distributed on the baseline in a spindle shape and spread thinner upward and downward rather than widely. Therefore, the predicted value can be said to be close to the experimental value as a whole.
[0095] In addition, in the experimental data of 2403 compounds, the average error of the known experimental error is about 3.8%, Figure 3A bar graph plots the error between the experimental and predicted values, centered around this value. In this graph, the multiple linear regression model predicts critical temperatures within the average experimental error with 88.09% probability, demonstrating its high predictive performance.
[0096] The present invention is not limited to the above-mentioned embodiments. Without departing from the gist of the invention claimed for protection in the claims, anyone with general knowledge in the technical field to which the present invention belongs can make various modified implementations, and the above-mentioned changes fall within the scope of the claims.
Claims
1. A method for predicting the critical temperature of a pure compound using a multiple linear regression model, comprising the following steps: (1) Input the collected experimental data of the sample compound; (2) preparing a molecular descriptor for the sample compound input above; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data in step (1) into a training set and a test set; (5) constructing an optimal multiple linear regression model that matches the training set using the molecular descriptors described in step (3); (6) Determine the rationality of the above multiple linear regression model; (7) If the multiple linear regression model is not reasonable in step (6), repeat steps (5) and (6). If it is reasonable, test the prediction performance of the model using the test set; (8) In the test of step (7), if the performance does not meet the standard, steps (4) to (7) are repeated; if the standard is met, the predicted value of the critical temperature obtained by using the multiple linear regression model is used as the actual critical temperature value.
2. The method according to claim 1, wherein The sample compound in step (1) is a pure compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
3. The method according to claim 2, wherein: The sample compound in step (1) is one of the following: a. Hydrocarbon compounds containing carbon and hydrogen elements; b. Non-hydrocarbon compounds containing one or more of carbon, hydrogen, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
4. The method according to any one of claims 1 to 3, wherein The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to the critical temperature.
5. The method according to any one of claims 1 to 3, wherein The optimal molecular descriptor in step (3) includes one or more of the following: P1: number of double bonds; P2: number of sp hybridized carbon atoms; P3: number of loops; P4: number of eleven-membered rings; P5: 10th order distance / detour index; P6: Average atomic Sanderson electronegativity based on the proportion of carbon atoms; P7: average first ionization potential based on the proportion of carbon atoms; P8: number of rotatable bonds; P9: Halogen atomic percentage.
6. The method according to claim 5, wherein: The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values for all sample compounds.
7. The method according to any one of claims 1 to 3, wherein In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.
8. The method according to any one of claims 1 to 3, wherein The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to the training set to construct a multiple linear regression model.
9. The method according to claim 8, wherein The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.
10. The method according to any one of claims 1 to 3, wherein The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.
11. The method according to any one of claims 1 to 3, wherein The test in step (7) is to test the multiple linear regression model using the test set.
12. The method according to any one of claims 1 to 3, wherein The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.