Method for predicting critical volume of pure compound by using multiple linear regression model
By screening molecular descriptors through a multiple linear regression model and establishing a compound critical volume prediction model, the problems of computational complexity and high cost in existing technologies are solved, and efficient, rapid and accurate prediction of specific element compounds is achieved, which is suitable for compound physical property analysis and industrial design.
Patent Information
- Application Number
- CN202411354451.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-09-19
AI Technical Summary
Existing methods for predicting compound properties, such as group contribution methods and artificial neural network models, have problems such as complex calculations, high costs, long time periods, and limited applicability. In particular, when predicting the critical volume of compounds, the accuracy and efficiency are difficult to meet industrial needs.
A multiple linear regression model is used to screen out suitable molecular descriptors and establish a critical volume prediction model for compounds. A fast and accurate multiple linear regression model is constructed using experimental data training and test sets, avoiding quantum mechanical calculations and shortening prediction time.
It achieves highly accurate prediction of the critical volume of compounds composed of specific elements, shortens calculation time, reduces costs, expands the scope of application, and is suitable for physical property analysis and industrial applications of more compounds.
Smart Images

Figure CN120673901A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of physical property prediction, which is one of the fields of physical chemistry, and relates to a method for predicting the critical volume, which is one of the various physical properties of a compound, with high accuracy. Background Art
[0002] Today, humanity relies on a vast array of chemical compounds, including plastics, fibers, rubber, paints, fertilizers, pharmaceuticals, and fuels, and this trend is expected to intensify. According to the American Chemical Society (ACS), the total number of registered chemical substances as of 2023 is over 204 million. In comparison, the number of compounds for which even a single physical property is experimentally known is only in the tens of thousands, a tiny fraction of the total number of chemical substances. However, the physical properties of chemical compounds are essential for a better material life, including the development of new substances and drugs, the optimal design of chemical equipment, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment.
[0003] Specifically, knowing the exact values of various physical properties of a compound plays a decisive role in various decision-making matters throughout the production and consumption process, such as determining the rationality of the substance's use or designing synthesis and purification processes, setting methods and conditions for storage, transportation, use, and disposal. Therefore, it is of great significance both industrially and academically.
[0004] Currently, the most accurate way to obtain the relevant physical property values of the compound of interest is through experimental measurement. However, this requires considerable cost and time in various aspects, such as preparing purified samples and constructing an environment for accurate measurement, and may sometimes be impossible to achieve depending on the actual situation.
[0005] Therefore, as an alternative, many researchers have long been working on predicting the accurate values of various physical properties of compounds. Since physical property prediction has a long history and new prediction methods continue to emerge, various prediction models that differ from each other currently coexist according to physical properties, accuracy, and application range.
[0006] As a model for predicting physical properties, the main method currently widely known and used is the group contribution method determined by statistical methods. This method has the best performance for predicting physical properties of compounds with experimental values.
[0007] The group contribution method has previously achieved some impressive results, but due to its lack of theoretical basis, it sometimes encounters problems with segmentation based on fragment form, sometimes even when no segmentation exists, making it impossible to calculate numerical values. Furthermore, as models are continually refined to improve predictive performance, they become increasingly complex and difficult to manage.
[0008] In the process of building prediction models, one of the alternatives to the group contribution method is the quantitative structure-property relationship (QSPR) method. The basis of this method is the assumption that the physical properties of a compound are functions of the structural characteristics of the molecule, and it uses a variety of molecular descriptors that reflect a variety of different structural characteristics. The types of molecular descriptors proposed so far in this method have reached thousands, ranging from simple molecular descriptors such as the number of carbon or hydrogen atoms in a molecule to complex molecular descriptors such as the shape, connection state, and electrochemical properties of the molecule. In addition, many types of calculation methods for molecular descriptors have been developed [Todeschini R., V. Consonni V., Molecular Descriptors for Chemoinformatics: Second, Revised and Enlarged Edition: Volume I / II, Wiley-VCH, 2009]. The QSPR prediction model is presented in the form of a function that includes these molecular descriptors and, sometimes, other physicochemical properties of the predicted compound (which are also functions of structural characteristics) as independent variables.
[0009] In this case, the most common function is the random variable (representer) X shown below: i The linear combination function, the coefficients c0 and c i It was mainly determined by multiple linear regression analysis based on experimental data.
[0010]
[0011] Another way to create a QSPR model is to use artificial neural networks. Artificial neural network technology is a widely used information processing technology that uses models based on human neural cells to create machines with artificial intelligence.
[0012] By training an artificial neural network using a sample set that binds various input values to the output values corresponding to those input values, the artificial neural network can establish general rules through learning, even without the rules or knowledge required to solve the problem. This allows it to output appropriate outputs even for unknown inputs. Therefore, artificial neural networks are widely used as a very useful tool in fields that lack basic theory, such as predicting the physical properties of compounds.
[0013] When using computers to uniformly calculate the values of more than 3,779 molecular descriptors of compounds required for the QSPR model for predicting physical properties, in order to calculate the electronic structure of the molecule, the electron energy solution is usually obtained by solving the Schrödinger equation. However, for systems with a large number of electrons, the calculation time required is very long, and the following attempts have been made in the prior art.
[0014] The calculation theories tried to select the best quantum mechanical calculation method are Hartree-Fock method [CCJ Roothan, Rev. Mod. Phys. 23, 69 (1951)], various post-Hartree-Fock methods [C. Moller and MSPlesset, Phys. Rev. 46, 618 (1934)], Gauss method combining Hartree-Fock and post-Hartree-Fock methods [LA Curtiss, K. Raghavachari, GW Trucks, and JA Popple, J. Chem. Phys. 94, 7221 (1991); LA Curtiss, K. Raghavachari, PC Redfern, V. Rassolov, and J.A.Pople, J.Chem.Phys.109,7764(1998)], density functional theory (DFT) [R.Seeger and J.A.Pople, J.Chem.Phys.66,3045(1977)] which uses electron density function instead of wave function with multi-dimensional perturbation term to consider the correlation between electrons in molecules composed of many electrons and uses total energy functional to find the ground state. However, these methods still require a very large amount of calculation, so when calculating large molecules, there are still difficulties in cost and time.
[0015] In addition, there are few prediction models for the critical volume that the present invention is concerned with. In the literature [Yash Nannoolal, Jürgen Rarey, Deresh Ramjugernath, Estimation of pure component properties Part 2. Estimation of critical property data by group contribution, Fluid Phase Equilibria 252 (2007) 1–27], a model for 348 compounds using the group contribution method is described. Although the prediction performance is reported to be a mean absolute error of 0.0000064 m 3 / mol, but the number of sample compounds is small, and the problems of the group contribution method still exist. In addition, there is a problem of lack of diversity due to the small number of sample compounds. In the literature [Srinivasa S. Godavarthy, Robert L. Robinson Jr., Khaled AMGasem, Improved structure-property relationship models for prediction of critical properties, Fluid Phase Equilibria 264 (2008) 122–136], a linear regression analysis model using 6 molecular descriptors for 1230 compounds was recorded, and the coefficient of determination was 0.99 and the mean absolute error was 0.0000142m 3 / mol, the prediction performance is not satisfactory.
[0016] Although some QSPR models using artificial neural networks to predict critical volumes have been proposed, most of them are limited to a small number of sample compounds or are limited to a specific series of compounds. The model and results are described in the literature [Srinivasa S. Godavarthy, Robert L. Robinson Jr., Khaled A. M. Gasem, Improved structure-property relationship models for prediction of critical properties, Fluid Phase Equilibria 264 (2008) 122–136]. The prediction performance of the nonlinear analysis model was improved, with a coefficient of determination of 0.999 and a mean absolute error of 0.0000052 m. 3 / mol results.
[0017] Therefore, the present invention provides a method for predicting critical volume as follows: in order to shorten the calculation time that is a problem of the above-mentioned existing methods and improve the accuracy of predicting physical properties, molecular descriptors that can serve as independent variables are screened out from molecular descriptors, and without adding the use of artificial neural networks, the accurate critical volume of the pure compound is predicted by establishing an optimal multiple linear regression model for the critical volume of the compound. Summary of the Invention
[0018] The technical problem to be solved by the present invention is to provide a more reliable QSPR model for the critical volume of compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As), based on more experimental data.
[0019] This model is expected to overcome the shortcomings of the model constructed by the group contribution method described above and to show fast and accurate prediction performance for a wider range of elements. The range of compounds to which the prediction model can be applied is limited to compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As). This is mainly because the model is based on experimental values and is therefore limited to the constituent elements of the compound for which experimental values exist, but the number of each atom is not limited.
[0020] As a predictive model capable of matching the composition of the constituent elements or molecular sizes of compounds with experimental values, it can be expanded to encompass a wider range of constituent elements and a greater number of atoms. Even within the aforementioned limited range of compounds, there are a vast number of compounds, including a significant number of industrially important ones. Based on the molecular structure of a compound, physical properties can be predicted in a very short time, facilitating the processing and analysis of physical property information on a large amount of compounds. Therefore, the present invention is expected to have significant beneficial effects on human society.
[0021] To achieve this purpose, an example of the present invention is a method for calculating the critical volume of a compound using a multiple linear regression model, comprising the following steps: (1) inputting the collected experimental data of the sample compound; (2) preparing the molecular descriptors of the sample compound inputted above; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data into a training set and a test set; (5) constructing the best multiple linear regression model that conforms to the training set using the molecular descriptors described in step (3); (6) judging the rationality of the multiple linear regression model; (7) if the multiple linear regression model is not rational in step (6), repeating steps (5) and (6); if it is rational, testing the prediction performance of the model using the test set; (8) in the test of step (7), if the performance does not meet the standard, repeating steps (4) to (7); if it meets the standard, using the predicted value of the critical volume obtained using the multiple linear regression model as the critical volume value.
[0022] The sample compound in step (1) is a pure compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
[0023] The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to the critical volume.
[0024] The optimal molecular descriptor in step (3) includes one or more of the following: P1: the number of nine-membered rings; P2: the tenth-order distance / detour ring index; P3: Balaban's mean square distance index; P4: the Hosoya-like index from the Barysz matrix weighted by ionization potential, wherein the Hosoya-like index is a logarithmic function; P5: the sum of the logarithmic coefficients of the last eigenvector of the Burden matrix weighted by I-states.
[0025] The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values for all sample compounds.
[0026] In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.
[0027] The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to the training set to construct a multiple linear regression model.
[0028] The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.
[0029] The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.
[0030] The test in step (7) is to test the multiple linear regression model using the test set.
[0031] The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.
[0032] The physical properties of compounds are essential for the development of new substances and drugs, the optimization of chemical equipment design, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment. In particular, critical volume, an important indicator of the critical point along with critical temperature and critical pressure, is a physical property whose accurate value is essential in commonly used programs such as Aspen Plus and Pro / II, widely known as optimization programs for chemical equipment design. It is an important physical property that provides a reference point when predicting the values of various physical properties through correlations based on equations of state or the corresponding states principle. However, the number of compounds for which experimental values are currently known does not exceed a few thousand, and obtaining data experimentally can be extremely difficult due to factors such as the compound's toxicity, instability, and difficulty in purification.
[0033] From this point of view, the present invention, which can obtain highly accurate critical volume values for many compounds using only information on the molecules without conducting experiments, not only saves the cost and time required for experiments, but also makes it possible to estimate the values even when experiments are not possible, thereby facilitating research and development activities in related industries. Furthermore, appropriate information can be provided to all places such as academia and government that require the values, allowing such activities to be carried out more smoothly.
[0034] Furthermore, although three-dimensional molecular descriptors are insufficient because quantum chemical calculations are not performed, the physical properties of unknown compounds can be predicted in a very short time. Furthermore, the prediction time can be further shortened without using additional models such as artificial neural networks, thereby enabling real-time prediction of physical properties. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 FIG. 4 is a flow chart illustrating a process of constructing a critical volume QSPR model provided by the present invention.
[0036] Figure 2is a scatter plot showing the prediction performance of the multiple linear regression model.
[0037] Figure 3 It is a bar chart of 2422 experimental data of the QSPR model. DETAILED DESCRIPTION
[0038] Hereinafter, a method for predicting the critical volume of a pure compound using a multiple linear regression model according to the present invention will be described in detail with reference to the accompanying drawings so that persons having ordinary knowledge in the technical field to which the present invention pertains can easily implement the present invention.
[0039] In the process of describing the present invention, if it is determined that the detailed description of the related known technology may cause unnecessary obscurity to the gist of the present invention, the detailed description will be omitted.
[0040] When using the words "including," "having," "forming," "comprising," and the like in this specification, other parts may be added unless "only" is used. When describing a component in the singular, the plural is included unless otherwise specified.
[0041] Figure 1 A flowchart schematically illustrates the process of constructing a QSPR model of critical volume.
[0042] The first task in building the model is to collect experimental data and analyze and classify them as specified in the first step.
[0043] For the present invention, we conducted an extensive search of all available references and materials, including various papers, monographs, and websites, to collect experimental data on critical volumes for compounds meeting the criteria of the present invention. We conducted a comprehensive review to determine whether the collected data represented truly reasonable values that could be used to construct a model. We then carefully analyzed and modified or deleted data for compounds with the following conditions: (i) non-experimental values; (ii) data with erroneous labeling; (iii) values that differed significantly even for the same compound; (iv) values that deviated unreliably from those of other similar compounds; or (v) values for which molecular descriptors were difficult to immediately generate. Finally, we selected data for a total of 2,422 compounds.
[0044] In addition, when constructing the physical property prediction model, all compounds containing one or more of the elements C, H, N, O, S, F, Cl, Br, I, Si, P, and As were integrated to establish a model for the sample compound.
[0045] Classifying compounds into a single type shows more efficient prediction performance for compounds composed of multiple elements, so all 2,422 compounds were combined into a single model. In this invention, "compound" refers to substances formed from molecules composed of no more than 12 elements: hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As).
[0046] The second step is to prepare molecular descriptor values for these compounds. Molecular descriptors suitable for use as model variables are calculated using the following method for the values of over 3,779 molecular descriptors: a file (MOL or SDF file) containing structural information, including the X, Y, and Z coordinates of each molecule's constituent elements and the bonding patterns between them, is input into a computer program. Based on this input, molecular descriptors representing the diversity of molecules are calculated using various theoretical mathematical formulas, moving beyond simple element and bond counts. Furthermore, the process of extracting these 3,779 molecular descriptors is repeated for all sample compounds used to construct the physical property prediction model.
[0047] The theoretical mathematical formulas used in the calculation are known in the art and are used in various forms, for example, a large number of theoretical mathematical formulas including Chi Connectivity Indices, Topological Descriptors, Information Indices, Walk and Path Counts, 2D Autocorrelation Descriptors, etc.
[0048] In the best results, not only physical property information but also molecular descriptors expressed as a variety of meaningful numerical values reflecting the characteristics of the molecules can be obtained.
[0049] For pure substances, there are molecular descriptors that can characterize both two-dimensional and three-dimensional structures. However, three-dimensional molecular descriptors are affected by molecular structure and are more complex to compute. Therefore, in this invention, in order to rapidly process physical property predictions without performing quantum mechanical calculations for structure optimization, three-dimensional molecular descriptors are not used. This approach significantly reduces the computational time and cost of quantum mechanical calculations.
[0050] Furthermore, in the third step, after calculating the molecular descriptor values, it is necessary to screen out inappropriate molecular descriptor values. Specifically, by applying a theoretical mathematical formula to the molecular descriptors of the sample compounds, the system screens out molecular descriptor values that yield the same value for all sample compounds and therefore cannot serve as independent variables in the model. This screening of molecular descriptors prevents the inclusion of irrelevant molecular descriptors in the prediction model, thereby improving model reliability while reducing the number of molecular descriptors, the number of calculations, and the computational effort, ultimately shortening the computational time required to find the optimal model.
[0051] In the present invention, the obtained optimal molecular descriptors include the number of nine-membered rings, the tenth-order distance / detour ring index, the Balaban mean square distance index, the Hosoya-like index from the Barysz matrix weighted by ionization potential, the Hosoya-like index is a logarithmic function, and the logarithmic coefficient of the last eigenvector of the Burden matrix weighted by I-state.
[0052] In the fourth step, the sample compounds are divided into a training set and a test set. The training set will be used to find the prediction model, and the test set will be used to test the predictive performance of the determined model. While taking care not to bias the distribution of similar molecules to one side, the training set and test set are divided in a ratio of 5-8:2-5, preferably 7:3.
[0053] Then, in the fifth step, the optimal multiple linear regression model is found based on the training set. Here, "optimal" is used in a relative sense, meaning that it can be found in a relatively short time while having performance very close to the optimal solution in an absolute sense.
[0054] The reason for not directly finding the optimal solution is that the calculation time is long. For example, when the number of suitable molecular descriptors among 3779 molecular descriptors is 1700, the total number of multiple linear regression models that can be made by extracting 5 different molecular descriptors is It is therefore practically impossible to investigate them all.
[0055] In order to obtain useful results within a limited time, a stepwise selection method is adopted in the present invention.
[0056] The stepwise selection method is a common method for selecting variables from data with many independent variables. During each step of creating a regression equation, the following steps are repeated: Among the independent variables not already in the equation, variables with high correlations with the dependent variable are added by entering a benchmark, and among the independent variables already in the equation, variables with low correlations with the dependent variable are removed by removing the benchmark. If no further variables are added or removed during this process, the selection of independent variables ends, and the model based on the stepwise selection method is finally completed.
[0057] In other words, a method called hypothesis testing is used to determine whether a hypothesis about a population is statistically correct. The significance level is typically set at 0.05 (5%), and a p-value (probability of significance) is calculated. When the p-value is less than 0.05, the null hypothesis is rejected and the alternative hypothesis is adopted. Hypothesis testing is performed on selected variables, and the process of adding new variables or removing existing ones is repeated. Ultimately, when no new variables are added or removed, the stepwise selection process ends.
[0058] After the stepwise selection is successful, the molecular descriptor information obtained will be used to perform a more advanced search. After the statistical test is passed, a model with reliable accuracy is obtained, as shown below:
[0059]
[0060] Where P represents the molecular descriptor after stepwise selection, Y is the property to be predicted, and C is the parameter estimated from the experimental data by the multivariate regression method.
[0061] Preferably, the least square residual is used to estimate the parameter C. The specific calculation process is as follows:
[0062]
[0063] Where n is the number of substances to be predicted and Pn is the number of optimal molecular descriptors.
[0064] Once the best multiple linear regression model is selected, the rationality of the model is determined in the sixth step, which is the next step. Specifically, the rationality of the model can be tested using the t-test method, and the calculation process is as follows:
[0065]
[0066]
[0067] If the t-test value for a molecular descriptor included in the model is found to be poor according to the above method, the process returns to the previous step and searches for another model. For example, if the number of sample compounds is 1005 and the selected model consists of five molecular descriptors, if the t-test value for one molecular descriptor is 3.3 or higher, the probability that the molecular descriptor is unrelated to the corresponding physical property is less than 0.1%.
[0068] In the present invention, when there is a molecular descriptor with a t-test value less than about 3, the selected model is discarded and other models are sought. In addition, the situation that the value of a molecular descriptor for the sample compound is the same except for a few compounds cannot be regarded as a model with reliability, so measures are also taken. Usually, if the number of molecular descriptors included in the model is increased, the prediction performance will improve, but the above-mentioned problems will occur. Therefore, usually, the final model is obtained by repeatedly performing these steps by changing the number of molecular descriptors and through multiple trial and error. If the selected model no longer has problems, then proceed to the next step.
[0069] Next, in the seventh step, the prediction performance of the found model is evaluated using a test set that does not participate in forming the model.
[0070] If significant performance degradation or significant deviations in prediction are observed in the training set, the fourth step is to readjust the training and test sets before proceeding to the subsequent steps. If the difference between the training and test sets does not exceed 20% of the mean absolute error (AAE) obtained for the training set, the prediction performance is considered satisfactory, and a multiple linear regression model for the critical volume is established.
[0071] The results of the model finally established through this process are briefly shown in Table 1 below.
[0072] Table 1
[0073] Main contents of the QSPR prediction model for the critical volume of all compounds
[0074]
[0075]
[0076] Example
[0077] This example is used to further illustrate the accuracy of the present invention in predicting the critical volume of a pure compound, but the present invention is not limited thereto.
[0078] Table 2 below shows some results of predicting the critical volume of pure compounds using the model of the present invention:
[0079] Table 2
[0080]
[0081]
[0082]
[0083] In order to confirm the excellence of the present invention, the multiple linear regression model of the present invention was used to predict the critical volumes of 2422 compounds with known experimental values, and a coefficient of determination value of 0.9918 and a value of 0.000012605 m 3 The average absolute error value of / mol indicates that the prediction performance is excellent.
[0084] Figure 2 This is a scatter plot showing the predictive performance of a multiple linear regression model. It evaluates how closely experimental values are aligned with predicted values. The X-axis represents the experimental values, and the Y-axis represents the predicted values. The experimental and predicted values are plotted at the (X, Y) coordinates with dots (hollow blue circles), and the baseline representing Y=X is marked with a diagonal line (solid red line).
[0085] Points located on the baseline indicate that the predicted value is the same as the experimental value. Points located below the baseline indicate that the predicted value is smaller than the experimental value. Points located above the baseline indicate that the predicted value is larger than the experimental value. For example, if the experimental and predicted values are (100, 100), they are on the baseline; if the experimental and predicted values are (100, 90), they are below the baseline; and if the experimental and predicted values are (100, 10), they are above the baseline.
[0086] exist Figure 2 In the parity plot, we can observe that there are some points that are relatively separated from the baseline in the middle. However, in most areas, the points are concentrated on the baseline. The points are distributed on the baseline in a spindle shape and spread thinner upward and downward rather than widely. Therefore, the predicted value can be said to be close to the experimental value as a whole.
[0087] In addition, in the experimental data of 2422 compounds, the average error of the known experimental error is about 15.62%, Figure 3 The error between the experimental and predicted values is plotted in a bar graph centered around this value. In this graph, the multiple linear regression model predicts critical volume values within the average experimental error with a probability of 98.01%, demonstrating its high predictive performance.
[0088] The present invention is not limited to the above-mentioned embodiments. Without departing from the gist of the invention claimed for protection in the claims, anyone with general knowledge in the technical field to which the present invention belongs can make various modified implementations, and the above-mentioned changes fall within the scope of the claims.
Claims
1. A method for predicting the critical volume of a pure compound using a multiple linear regression model, comprising the following steps: (1) Input the collected experimental data of the sample compound; (2) preparing a molecular descriptor for the sample compound input above; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data in step (1) into a training set and a test set; (5) constructing an optimal multiple linear regression model that matches the training set using the molecular descriptors described in step (3); (6) Determine the rationality of the above multiple linear regression model; (7) If the multiple linear regression model is not reasonable in step (6), repeat steps (5) and (6). If it is reasonable, test the prediction performance of the model using the test set; (8) In the test of step (7), if the performance does not meet the standard, repeat steps (4) to (7); if the standard is met, the predicted value of the critical volume obtained using the multiple linear regression model is used as the critical volume value.
2. The method according to claim 1, wherein The sample compound in step (1) is a pure compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.
3. The method according to any one of claims 1 to 2, wherein The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to the critical volume.
4. The method according to any one of claims 1 to 3, wherein The optimal molecular descriptor in step (3) includes one or more of the following: P1: number of nine-membered rings; P2: distance / detour ring index of tenth order; P3: Balaban's mean square distance index; P4: Hosoya-like index from the Barysz matrix weighted by ionization potential, which is a logarithmic function; P5: The sum of the logarithmic coefficients of the last eigenvector of the Burden matrix from the I-state weights.
5. The method according to claim 4, wherein The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values for all sample compounds.
6. The method according to any one of claims 1 to 2, wherein In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.
7. The method according to any one of claims 1 to 2, wherein The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to the training set to construct a multiple linear regression model.
8. The method according to claim 7, wherein: The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.
9. The method according to any one of claims 1 to 2, wherein: The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.
10. The method according to any one of claims 1 to 2, wherein: The test in step (7) is to test the multiple linear regression model using the test set.
11. The method according to any one of claims 1 to 2, wherein: The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.