Method for predicting absolute entropy of ideal gas in standard state

By screening molecular descriptors through a multiple linear regression model, a prediction model for the absolute entropy of ideal gas under standard conditions of compounds was established, which solved the problems existing in the existing technology in predicting the absolute entropy of ideal gas under standard conditions of compounds, achieved fast and accurate prediction results, and supported the optimization of chemical equipment and the development of new substances.

CN120673898APending Publication Date: 2025-09-19BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411354384.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have problems with long calculation time and low accuracy when predicting the absolute entropy of ideal gas under standard conditions of compounds. In particular, high-dimensional calculations require high resources, and the group contribution method lacks a theoretical basis, resulting in a complex and difficult-to-handle model.

Method used

A multiple linear regression model was used to screen out suitable molecular descriptors and establish a prediction model for the absolute entropy of ideal gas under standard conditions. Experimental data were used to construct training and test sets, and the stepwise selection method was used to find the optimal model, avoiding quantum chemical calculations and artificial neural networks, shortening the calculation time and improving accuracy.

Benefits of technology

The absolute entropy of ideal gas under standard state of a compound can be predicted with high accuracy in a short time, saving experimental cost and time. It is applicable to compounds composed of a wide range of elements, supports chemical equipment optimization and new material development, and improves production efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673898A_ABST
    Figure CN120673898A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting ideal gas absolute entropy of a pure compound in a standard state. According to the method, a mathematical model for predicting the ideal gas absolute entropy of a pure compound in a standard state with high accuracy, wherein the pure compound is composed of 12 or less elements of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, arsenic and the like, and the atom number of the pure compound is 25 or less (not containing hydrogen). The model is used as a universal quantitative structure-property relation model, after an optimal model is obtained from a plurality of multiple linear regression models by means of a stepwise selection method, values of molecular descriptors included in the model are received in a short time, and an ideal gas absolute entropy in a standard state is output. As long as the specific value of the molecule descriptor included in the model is known, the ideal absolute gas entropy in the standard state of the compound purely composed of the molecule is predicted for any molecule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of physical property prediction, which is one of the fields of physical chemistry, and relates to a method for predicting the absolute entropy of an ideal gas under standard conditions, which is one of various physical properties of a compound, with high accuracy. Background Art

[0002] Today, humanity relies on a vast array of chemical compounds, including plastics, fibers, rubber, paints, fertilizers, pharmaceuticals, and fuels, and this trend is expected to intensify. According to the American Chemical Society (ACS), the total number of registered chemical substances as of 2023 is over 204 million. In comparison, the number of compounds for which even a single physical property is experimentally known is only in the tens of thousands, a tiny fraction of the total number of chemical substances. However, the physical properties of chemical compounds are essential for a better material life, including the development of new substances and drugs, the optimal design of chemical equipment, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment.

[0003] Specifically, understanding the exact values ​​of various physical properties of a compound plays a decisive role in various decision-making matters throughout the production and consumption process, such as determining the rationality of the substance's use or designing synthesis and purification processes, setting methods and conditions for storage, transportation, use, and disposal. Therefore, it is of great significance both industrially and academically.

[0004] Currently, the most accurate way to obtain the relevant physical property values ​​of the compound of interest is through experimental measurement. However, this requires considerable cost and time in various aspects, such as preparing purified samples and constructing an environment for accurate measurement. Sometimes, it may not be possible depending on the actual situation.

[0005] Therefore, as an alternative, many researchers have long been working on predicting the accurate values ​​of various physical properties of compounds. Since physical property prediction has a long history and new prediction methods continue to emerge, various prediction models that differ from each other currently coexist according to physical properties, accuracy, and application range.

[0006] As a model for predicting physical properties, the main method currently widely known and used is the group contribution method determined by statistical methods. This method has the best performance for predicting physical properties of compounds with experimental values.

[0007] This group contribution method has previously achieved some impressive results. However, due to its lack of theoretical basis, it sometimes encounters problems with segmentation based on fragment form, such as non-uniqueness or even the inability to calculate numerical values ​​when the group does not exist after segmentation. Furthermore, as models are refined to improve predictive performance, they become increasingly complex and difficult to handle.

[0008] In the process of building prediction models, one of the alternatives to the group contribution method is the quantitative structure-property relationship (QSPR) method. The basis of this method is the assumption that the physical properties of a compound are functions of the structural characteristics of the molecule, and it uses a variety of molecular descriptors that reflect a variety of different structural characteristics. The types of molecular descriptors proposed so far in this method have reached thousands, ranging from simple molecular descriptors such as the number of carbon or hydrogen atoms in a molecule to complex molecular descriptors such as the shape, connection state, and electrochemical properties of the molecule. In addition, many types of calculation methods for molecular descriptors have been developed [Todeschini R., V. Consonni V., Molecular Descriptors for Chemoinformatics: Second, Revised and Enlarged Edition: Volume I / II, Wiley-VCH, 2009]. The QSPR prediction model is presented in the form of a function that includes these molecular descriptors and, sometimes, other physicochemical properties of the predicted compound (which are also functions of structural characteristics) as independent variables.

[0009] In this case, the most common function is the random variable (representer) X shown below: i The linear combination function, the coefficients c0 and c i It was mainly determined by multiple linear regression analysis based on experimental data.

[0010]

[0011] Another way to create a QSPR model is to use artificial neural networks. Artificial neural network technology is a technology that uses human neural cells to model intelligent machines and is currently a widely used information processing technology.

[0012] By training an artificial neural network using a sample set that binds various input values ​​to the output values ​​corresponding to those input values, the artificial neural network can establish general rules through learning, even without the rules or knowledge required to solve the problem. This allows it to output appropriate outputs even for unknown inputs. Therefore, artificial neural networks are widely used as a very useful tool in fields that lack basic theory, such as predicting the physical properties of compounds.

[0013] Using a computer, we uniformly calculated the values ​​of the 3,779 or so molecular descriptors of compounds required for the QSPR model for predicting physical properties. To calculate the electronic structure of a molecule, we typically solve the Schrödinger equation to determine the electron energy. However, this requires a considerable amount of computation time for systems with a large number of electrons. Therefore, we attempted the following experiment.

[0014] The calculation theories tried to select the best quantum mechanical calculation method are Hartree-Fock method [CCJ Roothan, Rev. Mod. Phys. 23, 69 (1951)], various post-Hartree-Fock methods [C. Moller and MSPlesset, Phys. Rev. 46, 618 (1934)], Gauss method combining Hartree-Fock and post-Hartree-Fock methods [LA Curtiss, K. Raghavachari, GW Trucks, and JA Popple, J. Chem. Phys. 94, 7221 (1991); LA Curtiss, K. Raghavachari, PC Redfern, V. Rassolov, and J.A.Pople, J.Chem.Phys.109,7764(1998)], density functional theory (DFT) [R.Seeger and J.A.Pople, J.Chem.Phys.66,3045(1977)] which uses electron density function instead of wave function with multi-dimensional perturbation term to consider the correlation between electrons in molecules composed of many electrons and uses total energy functional to find the ground state. However, these methods still require a very large amount of calculation, so when calculating large molecules, there are still difficulties in cost and time.

[0015] In addition, the standard state absolute entropy of ideal gas, which is the physical property of interest in the present invention, is the absolute entropy of ideal gas when a pure substance has a standard state (i.e., 298.15K (absolute temperature) and 1 bar), and various prediction models have been proposed so far.

[0016] The following table shows the main group contribution models proposed previously for the prediction of the absolute entropy of ideal gas under standard conditions by year.

[0017] years Proposer 1968 Benson 1969 Benson&Buss 1971 Thinh, T.-P. & T.K. Trong 1984 Joback 1994 Constantinou&Gani 1994 Domalski 1998 Benson

[0018] Previous research results on the prediction of the absolute entropy of ideal gas under standard conditions are briefly introduced in the literature [Poling BE, Prausnitz JM, O'Connell JP The Properties of Gases and Liquids, (5 ed.). New York: McGraw Hill. (2000).].

[0019] Currently, the most widely known and used models for predicting the absolute entropy of an ideal gas under standard conditions are those that utilize quantum mechanical calculations or the group contribution method. While quantum mechanical calculations improve accuracy when performing higher-dimensional calculations, they are difficult to perform due to the excessive time and resources required. Therefore, lower-dimensional quantum mechanical calculations are often used. While these can shorten calculation time and conserve resources, they suffer from lower accuracy.

[0020] Therefore, the present invention provides a method for calculating the absolute entropy of an ideal gas under standard conditions as follows: in order to shorten the calculation time problem of the above-mentioned existing methods and improve the accuracy of predicting physical properties, molecular descriptors that can become independent variables are screened out from the molecular descriptors, and without adding the use of artificial neural networks, the accurate absolute entropy of an ideal gas under standard conditions of a pure compound is calculated by establishing an optimal multiple linear regression model for the absolute entropy of an ideal gas under standard conditions of the compound. Summary of the Invention

[0021] The technical problem to be solved by the present invention is to provide a more reliable QSPR model of the absolute entropy of ideal gas under standard conditions for compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As), based on more experimental data.

[0022] This model is expected to overcome the shortcomings of the model constructed by the group contribution method as described above and to present fast and accurate prediction performance for a wider range of elements. In the present invention, the scope of compounds to which the prediction model can be applied is limited to compounds composed of 12 elements or less, such as hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As). The main reason is that the model is based on experimental values ​​to form the prediction model, and is therefore limited to the constituent elements of the compound for which there are experimental values, but the number of various atoms is not limited.

[0023] The model described in the present invention, as a predictive model for the constituent elements or molecules of compounds that can match experimental values, can be expanded to encompass a wider range of constituent elements and a greater number of atoms. Even within the range of compounds limited by the aforementioned elements, a vast number of compounds exist, including a considerable number of industrially important compounds. Based on the molecular structure of a compound, physical properties can be predicted in a very short time, thereby facilitating the processing and analysis of physical property information for a large amount of compounds. Therefore, the present invention is expected to have significant beneficial effects on human society.

[0024] In order to achieve such a purpose, an example of the present invention is a method for predicting the absolute entropy of an ideal gas under standard conditions using a multiple linear regression model, the method comprising the following steps: (1) inputting the collected experimental data of the sample compound; (2) preparing the molecular descriptors of the sample compound; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data of step (1) into a training set and a test set; (5) using the molecular descriptors of step (3) to find the best multiple linear regression model that meets the training set; (6) judging the rationality of the multiple linear regression model; (7) if the model is not rational in step (6), repeating steps (5) and (6); if it is rational, testing the prediction performance of the model using the test set; (8) in the test of step (7), if the performance does not meet the standard, repeating steps (4) to (7); if it meets the standard, predicting the absolute entropy of the ideal gas under standard conditions using the multiple linear regression model, and the obtained value is used as the absolute entropy value of the ideal gas under standard conditions.

[0025] The sample compound in step (1) is a compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

[0026] The sample compound in step (1) is one or more of the following: a. a liquid type 1 compound containing one or more of carbon, hydrogen, nitrogen, oxygen, and sulfur; b. a liquid type 2 compound containing one or more of carbon, hydrogen, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

[0027] The preparation of molecular descriptors in step (2) refers to the preparation of molecular descriptors related to the absolute entropy of ideal gas under standard conditions.

[0028] The optimal molecular descriptor in step (3) includes one or more of the following: P1: number of rotatable bonds; P2: number of sulfur atoms; P3: number of nine-membered rings; P4: distance / detour ring index of order 4; P5: distance / detour ring index of order 9; P6: number of four-membered rings; P7: pseudo-connectivity index of electrotopological state - type 0; P8: lopping center index; P9: average information index of atomic composition.

[0029] The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values ​​from all sample compounds.

[0030] In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.

[0031] The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to explore the multiple linear regression model for the training set.

[0032] The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.

[0033] The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.

[0034] The test in step (7) is to test the multiple linear regression model using the test set.

[0035] The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.

[0036] The physical properties of compounds are crucial for the development of new substances and drugs, the optimization of chemical equipment design, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment. In particular, the absolute entropy of an ideal gas under standard conditions is a property for which accurate values ​​are crucial in commonly used software such as Aspen Plus and Pro / II, which are used for optimization of chemical equipment design.

[0037] However, the number of compounds for which experimental values ​​are currently known does not exceed several thousand at most, and obtaining data through experiments is sometimes extremely difficult due to constraints such as the toxicity, instability, and difficulty in purification of the compounds.

[0038] The present invention can obtain the absolute entropy values ​​of ideal gases under standard conditions for many compounds with high accuracy using only information on molecules without conducting experiments. The method described in the present invention not only saves the cost and time required for experiments, but also can infer its value even when experiments are not possible, thereby facilitating research and development activities in related industries. Furthermore, appropriate information can be provided to all places such as academia and government where its value is needed, allowing such activities to be carried out more smoothly.

[0039] In addition, although the present invention lacks three-dimensional molecular descriptors because it does not perform quantum chemical calculations, it can predict the physical properties of unknown compounds in a very short time, and the prediction time is further shortened by not using additional models such as artificial neural networks, thereby enabling real-time prediction of physical properties. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart showing a process of constructing a QSPR model of absolute entropy of an ideal gas under standard conditions provided by the present invention.

[0041] Figure 2 It is a scatter plot (parity) of 1886 experimental data of the QSPR model of absolute entropy of ideal gas under standard conditions provided by the present invention.

[0042] Figure 3 It is a bar chart of 1886 experimental data of the QSPR model. DETAILED DESCRIPTION

[0043] Hereinafter, a method for predicting the absolute entropy of an ideal gas under standard conditions of a pure compound according to the multiple linear regression model of the present invention will be described in detail with reference to the accompanying drawings, so that persons with general knowledge in the technical field to which the present invention belongs can easily implement the present invention.

[0044] In the process of describing the present invention, if it is determined that the detailed description of the related known technology may cause unnecessary obscurity to the gist of the present invention, the detailed description will be omitted.

[0045] When using the words "including," "having," "forming," "comprising," and the like in this specification, other parts may be added unless "only" is used. When describing a component in the singular, the plural is included unless otherwise specified.

[0046] Figure 1 A flowchart schematically illustrates the process of constructing a QSPR model of the absolute entropy of an ideal gas under standard conditions.

[0047] The first step in building a model is to collect experimental data and analyze and classify them as mentioned in the first step.

[0048] For the present invention, we conducted an extensive search of all available reference materials and documents, including various papers, monographs, and websites, to collect experimental data on the absolute entropy of ideal gases under standard conditions for compounds meeting the requirements of the present invention. We conducted a comprehensive review to determine whether the collected data were truly reasonable values ​​for model construction. We carefully analyzed and modified or deleted data for compounds that (i) were non-experimental values, (ii) had data labeling errors, (iii) had values ​​that differed significantly even for the same compound, (iv) deviated unreliably from the values ​​of other similar compounds, or (v) were difficult to immediately prepare molecular descriptors for. Finally, we selected data for 1,886 compounds.

[0049] In addition, when constructing the physical property prediction model, the sample compounds were divided into (a) a class of compounds (compound 1) composed of one or more of the elements C, H, N, O, and S, and (b) a class of compounds (compound 2) composed of one or more of the elements C, H, F, Cl, Br, I, Si, P, and As, and models were established for each.

[0050] Since the above two classifications showed more effective prediction performance, they were divided into 1272 compounds in Class I and 614 compounds in Class II, and models were established for each. In addition, in the present invention, "compound" refers to a substance formed by molecules composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As).

[0051] The second step is to prepare molecular descriptor values ​​for these compounds. Molecular descriptors suitable for use as model variables are calculated using the following method for the values ​​of over 3,779 molecular descriptors: a file (MOL or SDF file) containing structural information, including the X, Y, and Z coordinates of each molecule's constituent elements and the bonding patterns between them, is input into a computer program. Based on this input, molecular descriptors representing the diversity of molecules are calculated using various theoretical mathematical formulas, moving beyond simple element and bond counts. Furthermore, the process of extracting these 3,779 molecular descriptors is repeated for all sample compounds used to construct the physical property prediction model.

[0052] The theoretical mathematical formulas used in the calculation are known in the art, for example, a variety of theoretical mathematical formulas including Chi Connectivity Indices, topological descriptors, information indices, walk and path counts, 2D autocorrelation descriptors, etc., which can be used in various forms.

[0053] In the best results, not only physical property information but also molecular descriptors expressed as a variety of meaningful numerical values ​​reflecting the characteristics of the molecules can be obtained.

[0054] For pure substances, there are molecular descriptors that can characterize both two-dimensional and three-dimensional structures. However, three-dimensional molecular descriptors are affected by molecular structure and are more complex to compute. Therefore, in this invention, in order to rapidly process physical property predictions without performing quantum mechanical calculations for structure optimization, three-dimensional molecular descriptors are not used. This approach significantly reduces the computational time and cost of quantum mechanical calculations required to obtain molecular descriptors.

[0055] Furthermore, in the third step, after calculating the molecular descriptor values, it is necessary to screen out inappropriate molecular descriptor values. Specifically, by applying a theoretical mathematical formula to the molecular descriptors of the sample compounds, the inappropriate molecular descriptor values ​​that yield the same value for all sample compounds and therefore cannot serve as independent variables in the model are screened out. This screening of molecular descriptors prevents the inclusion of irrelevant molecular descriptors in the prediction model, thereby improving model reliability. This reduces the number of molecular descriptors, the number of calculations, and the computational effort, thereby shortening the computational time required to find the optimal model.

[0056] In the present invention, when the sample compound is (a) a class of compounds (compound 1) composed of one or more elements including C, H, N, O, and S, the optimal molecular descriptors obtained include the number of rotatable bonds, the number of sulfur atoms, the number of nine-membered rings, the 4th-order distance / detour ring index, and the 9th-order distance / detour ring index; when the sample compound is (b) a class of compounds (compound 2) composed of one or more elements including C, H, F, Cl, Br, I, Si, P, and As, the optimal molecular descriptors obtained include the number of four-membered rings, the number of nine-membered rings, the electrotopological state pseudo-connectivity index-type 0, the lopping center index, and the average information index of the atomic composition.

[0057] In the fourth step, the sample compounds are divided into a training set and a test set. The training set will be used to find the prediction model, and the test set will be used to test the predictive performance of the determined model. While taking care not to bias the distribution of similar molecules to one side, the training set and test set are divided in a ratio of 5-8:2-5, preferably 7:3.

[0058] Then, in the fifth step, the optimal multiple linear regression model is found based on the training set. Here, "optimal" is used in a relative sense, meaning that it can be found in a relatively short time while having performance very close to the optimal solution in an absolute sense.

[0059] The reason for not directly finding the optimal solution is that the calculation time is long. For example, when the number of suitable molecular descriptors among 3779 molecular descriptors is 1700, the total number of multiple linear regression models that can be made by extracting 5 different molecular descriptors is It is therefore practically impossible to investigate them all.

[0060] In order to obtain useful results within a limited time, a stepwise selection method is adopted in the present invention.

[0061] The stepwise selection method is a common method for selecting variables from data with many independent variables. During each step of creating a regression equation, the following steps are repeated: Among the independent variables not already in the equation, variables with high correlations with the dependent variable are added by entering a benchmark, and among the independent variables already in the equation, variables with low correlations with the dependent variable are removed by removing the benchmark. If no further variables are added or removed during this process, the selection of independent variables ends, and the model based on the stepwise selection method is finally completed.

[0062] In other words, a method of using the so-called hypothesis test to determine whether the hypothesis about the population is statistically correct. The significance level is generally set to 0.05 (5%), and the p-value (p-value) as the probability of significance is calculated. When the p-value is less than 0.05, the null hypothesis is rejected and the alternative hypothesis is adopted. By performing hypothesis testing on the selected variables, the process of adopting newly added variables or removing existing variables is repeated. Ultimately, when there are no new variables to be added or removed, the stepwise selection process ends. After the stepwise selection is successful, the molecular descriptor information obtained will perform a more advanced search, and a model with reliable accuracy will be obtained after the statistical test is passed, as shown below:

[0063]

[0064] Where P represents the molecular descriptor after stepwise selection, Y is the property to be predicted, and C is the parameter estimated from the experimental data by the multivariate regression method.

[0065] Preferably, the least square residual is used to estimate the parameter C. The specific calculation process is as follows:

[0066]

[0067] Where n is the number of substances to be predicted and Pn is the number of optimal molecular descriptors.

[0068] Once the best multiple linear regression model is selected, the rationality of the model is determined in the sixth step, which is the next step. Specifically, the rationality of the model can be tested using the t-test method, and the calculation process is as follows:

[0069]

[0070] If a poor t-test value is found for a molecular descriptor included in the model, return to the previous step and search for another model. For example, if there are 1005 sample compounds and the selected model consists of five molecular descriptors, if the t-test value for one molecular descriptor is 3.3 or higher, the probability that the molecular descriptor is unrelated to the corresponding physical property is less than 0.1%.

[0071] In the present invention, when there is a molecular descriptor with a t-test value less than about 3, the selected model is discarded and other models are sought. In addition, the value of a molecular descriptor for the sample compound is all the same except for a few compounds and cannot be regarded as a model with reliability, so measures are also taken to discard. Usually, if the quantity of the molecular descriptors included in the model is increased, the prediction performance will improve, but problems as mentioned above will occur. Therefore, the final model is obtained by repeatedly performing these steps through multiple trial and error by changing the quantity of the molecular descriptors. If the selected model no longer has problems, then proceed to the next step.

[0072] Next, in the seventh step, the prediction performance of the found model is evaluated using a test set that does not participate in forming the model.

[0073] If significant performance degradation or significant deviations are observed in the training set, the training and test sets are readjusted in step 4 before proceeding to the subsequent steps. If the difference between the training and test sets does not exceed 20% of the mean absolute error (AAE) obtained for the training set, the prediction performance is considered satisfactory, and a multiple linear regression model for the absolute entropy of an ideal gas under standard conditions is established.

[0074] The results of the model finally established through this process are briefly shown in Tables 1 and 2 below.

[0075] Table 1

[0076] The main contents of the QSPR prediction model for the absolute entropy of ideal gas under standard conditions for a class of compounds

[0077]

[0078]

[0079]

[0080] Table 2

[0081] The main contents of the QSPR prediction model for the absolute entropy of ideal gas under standard conditions for two types of compounds

[0082]

[0083]

[0084]

[0085] Example

[0086] This example is used to further illustrate the accuracy of the present invention in predicting the absolute entropy of an ideal gas under standard conditions of a pure compound, but the present invention is not limited thereto.

[0087] Table 3 below shows some results of predicting the absolute entropy of pure compound ideal gas using the model of the present invention:

[0088] Table 3

[0089]

[0090]

[0091] To confirm the excellence of the present invention, the multiple linear regression model of the present invention was used to predict the absolute entropy of ideal gases in the standard state of 1886 compounds whose experimental values ​​were known, resulting in a coefficient of determination value of 0.9974 and an average absolute error value of 1.3735 cal / mol·K, indicating excellent predictive performance.

[0092] Figure 2 This is a scatter plot showing the predictive performance of a multiple linear regression model. It evaluates how closely experimental values ​​are aligned with predicted values. The X-axis represents the experimental values, and the Y-axis represents the predicted values. The experimental and predicted values ​​are plotted at the (X, Y) coordinates with dots (hollow blue circles), and the baseline representing Y=X is marked with a diagonal line (solid red line).

[0093] Points located on the baseline indicate that the predicted value is the same as the experimental value. Points located below the baseline indicate that the predicted value is smaller than the experimental value. Points located above the baseline indicate that the predicted value is larger than the experimental value. For example, if the experimental and predicted values ​​are (100, 100), they are on the baseline; if the experimental and predicted values ​​are (100, 90), they are below the baseline; and if the experimental and predicted values ​​are (100, 10), they are above the baseline.

[0094] exist Figure 2 In the parity plot, we can observe that there are some points that are relatively separated from the baseline in the middle. However, in most areas, the points are concentrated on the baseline. The points are distributed on the baseline in a spindle shape and spread thinner upward and downward rather than widely. Therefore, the predicted value can be said to be close to the experimental value as a whole.

[0095] In addition, in the experimental data of 1886 compounds, the average error of the known experimental error is about 4.23 cal / mol·K. Figure 3The error between the experimental and predicted values ​​is plotted in a bar graph centered on this value. In this graph, the multiple linear regression model predicts the absolute entropy of the standard state ideal gas within the average experimental error with a probability of 95.44%, demonstrating its high predictive performance.

[0096] The present invention is not limited to the above-mentioned embodiments. Without departing from the gist of the invention claimed for protection in the claims, anyone with general knowledge in the technical field to which the present invention belongs can make various modified implementations, and the above-mentioned changes fall within the scope of the claims.

Claims

1. A method for predicting the absolute entropy of an ideal gas under standard conditions using a multiple linear regression model, the method comprising the following steps: (1) Input the collected experimental data of the sample compound; (2) preparing a molecular descriptor for the sample compound; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data in step (1) into a training set and a test set; (5) using the molecular descriptors described in step (3) to find the best multiple linear regression model that matches the training set; (6) Determine the rationality of the above multiple linear regression model; (7) If the result in step (6) is not reasonable, repeat steps (5) and (6). If the result is reasonable, use the test set to test the prediction performance of the model. (8) In the test of step (7), if the performance does not meet the standard, repeat steps (4) to (7); if the standard is met, the absolute entropy of the ideal gas under standard conditions is predicted using the above-mentioned multiple linear regression model, and the obtained value is used as the absolute entropy value of the ideal gas under standard conditions.

2. The method according to claim 1, wherein The sample compound in step (1) is a compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

3. The method according to claim 2, wherein: The sample compound in step (1) is one or more of the following: a. Compounds containing one or more of carbon, hydrogen, nitrogen, oxygen, and sulfur; b. Class II compounds containing one or more of carbon, hydrogen, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

4. The method according to any one of claims 1 to 3, wherein The preparation of molecular descriptors in step (2) refers to the preparation of molecular descriptors related to the absolute entropy of ideal gas under standard conditions.

5. The method according to any one of claims 1 to 3, wherein The optimal molecular descriptor in step (3) includes one or more of the following: P1: number of rotatable bonds; P2: number of sulfur atoms; P3: number of nine-membered rings; P4: 4th order distance / detour index; P5: 9th order distance / detour index; P6: number of four-membered rings; P7: Pseudo-connectivity index of electrical topological state - type 0; P8: lopping center index; P9: Average information index of atomic composition.

6. The method according to claim 5, wherein: The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values ​​from all sample compounds.

7. The method according to any one of claims 1 to 3, wherein In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.

8. The method according to any one of claims 1 to 3, wherein The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to explore the multiple linear regression model for the training set.

9. The method according to claim 8, wherein The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.

10. The method according to any one of claims 1 to 3, wherein: The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.

11. The method according to any one of claims 1 to 3, wherein: The test in step (7) is to test the multiple linear regression model using the test set.

12. The method according to any one of claims 1 to 3, wherein: The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.