Method for predicting flash point of pure compound by using multiple linear regression model

By screening molecular descriptors through a multiple linear regression model and establishing a compound flash point prediction model, the problems of long calculation time and low accuracy in the existing technology are solved, and fast and highly accurate flash point prediction is achieved, which is suitable for compounds composed of a wide range of elements.

CN120708754APending Publication Date: 2025-09-26BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411354322.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies have problems with predicting the flash point of compounds, such as long calculation time and low accuracy. In particular, the group contribution method lacks a theoretical basis and increases complexity, and the artificial neural network has a large amount of calculation, making it difficult to predict the flash point quickly and accurately.

Method used

A multiple linear regression model is used to screen out suitable molecular descriptors and establish a prediction model for the flash point of compounds. The optimal model is gradually selected using experimental data training and test sets, avoiding quantum mechanical calculations, shortening the prediction time and improving accuracy.

Benefits of technology

It can efficiently predict the flash point of compounds in a short time, improve prediction accuracy, save experimental costs and time, and is applicable to compounds composed of a wide range of elements, especially compounds composed of 12 elements such as hydrogen, carbon, nitrogen, and oxygen, and is suitable for industrial and academic research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708754A_ABST
    Figure CN120708754A_ABST
Patent Text Reader

Abstract

The invention provides a method for predicting a flash point of a pure compound by using a multiple linear regression model. According to the method, a mathematical model for predicting the flash point of the pure compound with high accuracy is established to predict a flash point value. Specifically, the model is used as a quantitative structure-property relationship model, after the optimal model is obtained from a plurality of multiple linear regression models by using a stepwise selection method for most compounds with known flash point experimental values, the molecular descriptors included in the model are received in a short time, and the flash point values are output. As long as the specific values of the molecule descriptors included in the model are known, the flash point of the compound composed of the molecule is predicted for any molecule. As described above, the present invention provides a method and model capable of predicting a reliable flash point value even for many compounds having unknown experimental values, thereby saving the cost and time for physical performance measurement experiments or chemical structure prediction, and enabling research and development activities of related industries to become easier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of physical property prediction, which is one of the fields of physical chemistry, and relates to a method for predicting the flash point, which is one of the various physical properties of a compound, with high accuracy. Background Art

[0002] Today, humanity relies on a vast array of organic compounds, including plastics, fibers, rubber, paints, fertilizers, pharmaceuticals, and fuels, and this trend is expected to intensify. According to the American Chemical Society (ACS), the total number of registered chemical substances as of 2023 exceeds 204 million. In comparison, the number of compounds for which even a single physical property is experimentally known is only in the tens of thousands, a tiny fraction of the total number of chemical substances. However, the physical properties of compounds are essential for a better material life, including the development of new substances and drugs, the optimal design of chemical equipment, improving the productivity of existing equipment, developing and conserving resources, ensuring safety, and protecting the environment.

[0003] Specifically, understanding the exact values ​​of various physical properties of a compound plays a decisive role in various decision-making matters throughout the production and consumption process, such as determining the rationality of the substance's use or designing synthesis and purification processes, setting methods and conditions for storage, transportation, use, and disposal. Therefore, it is of great significance both industrially and academically.

[0004] Currently, the most accurate way to obtain the relevant physical property values ​​of the compound of interest is through experimental measurement. However, this requires considerable cost and time in various aspects, such as preparing purified samples and constructing an environment for accurate measurement. Sometimes, it may not be possible depending on the actual situation.

[0005] Therefore, as an alternative, many researchers have long been working on predicting the accurate values ​​of various physical properties of organic compounds. Since physical property prediction has a long history and new prediction methods continue to emerge, various prediction models that differ from each other currently coexist depending on the physical properties, accuracy, and scope of application.

[0006] As a model for predicting physical properties, the main method currently widely known and used is the group contribution method determined by statistical methods. This method has the best performance for predicting physical properties of compounds with experimental values.

[0007] This group contribution method has previously achieved some impressive results, but due to its lack of theoretical basis, it sometimes encounters problems with segmentation based on fragment form, sometimes with non-unique or even non-existent segmentation, resulting in inability to calculate numerical values. Furthermore, as models are refined to improve predictive performance, they become increasingly complex and difficult to handle.

[0008] In the process of building prediction models, one of the alternatives to the group contribution method is the quantitative structure-property relationship (QSPR) method. The basis of this method is the assumption that the physical properties of a compound are functions of the structural characteristics of the molecule, and it uses a variety of molecular descriptors that reflect a variety of different structural characteristics. The types of molecular descriptors proposed so far in this method have reached thousands, ranging from simple molecular descriptors such as the number of carbon or hydrogen atoms in a molecule to complex molecular descriptors such as the shape, connection state, and electrochemical properties of the molecule. In addition, many types of calculation methods for molecular descriptors have been developed [Todeschini R., V. Consonni V., Molecular Descriptors for Chemoinformatics: Second, Revised and Enlarged Edition: Volume I / II, Wiley-VCH, 2009]. The QSPR prediction model is presented in the form of a function that includes these molecular descriptors and, sometimes, other physicochemical properties of the predicted compound (which are also functions of structural characteristics) as independent variables.

[0009] In this case, the most common function is the random variable (representer) X shown below: i The linear combination function, the coefficients c0 and c i It was mainly determined by multiple linear regression analysis based on experimental data.

[0010]

[0011] Another way to create a QSPR model is to use artificial neural networks. Artificial neural network technology is a technology that uses human neural cells to model intelligent machines and is currently a widely used information processing technology.

[0012] By training artificial neural networks using sample sets that bind various input values ​​to corresponding output values, artificial neural networks can establish general rules through learning, even without the rules or knowledge required to solve the problem. This allows them to output appropriate outputs even for unknown inputs. Therefore, artificial neural networks are widely used as a very useful tool in fields that lack basic theory, such as predicting the physical properties of compounds.

[0013] When using computers to uniformly calculate the values ​​of more than 3,779 molecular descriptors of compounds required for the QSPR model for predicting physical properties, in order to calculate the electronic structure of the molecule, the electron energy solution is usually obtained by solving the Schrödinger equation. However, for systems with a large number of electrons, the calculation time required is very long, and the following attempts have been made in the prior art.

[0014] The calculation theories tried to select the best quantum mechanical calculation method are Hartree-Fock method [CCJ Roothan, Rev. Mod. Phys. 23, 69 (1951)], various post-Hartree-Fock methods [C. Moller and MSPlesset, Phys. Rev. 46, 618 (1934)], Gauss method combining Hartree-Fock and post-Hartree-Fock methods [LA Curtiss, K. Raghavachari, GW Trucks, and JA Popple, J. Chem. Phys. 94, 7221 (1991); LA Curtiss, K. Raghavachari, PC Redfern, V. Rassolov, and J.A.Pople, J.Chem.Phys.109,7764(1998)], density functional theory (DFT) [R.Seeger and J.A.Pople, J.Chem.Phys.66,3045(1977)] which uses electron density functions instead of wave functions with multidimensional perturbations to consider the correlation between electrons in molecules composed of many electrons and uses a functional of the total energy to find the ground state. However, these methods still require a very large amount of calculation, so when calculating large molecules, there are difficulties in terms of cost and time.

[0015] Furthermore, various prediction models have been proposed for the flash point, a physical property of interest in the present invention. The flash point is the lowest temperature at which ignition can occur in a flammable liquid or solid when a flammable liquid or solid has an ignition point.

[0016] The flash point prediction models reported in the literature so far can be classified as follows.

[0017] The first type of prediction model uses a prediction method that utilizes correlations with physical properties such as boiling point, density, vapor pressure, critical properties, and heat of vaporization. However, the accuracy of the correlation is directly related to the required physical properties or the accuracy of the method. Since the required physical properties are not fully available, there are many difficulties in performing the calculations used for prediction. The method proposed by Patil showed an average absolute deviation (AAD) of 19.7 K for 102 alkanes, indicating poor prediction performance. [Patil, GS Fire Mat. 1988, 12, 127]

[0018] The second type of prediction model is to use the QSPR method.

[0019] Suhani J. Patel et al. [Ind. Eng. Chem. Res., 2010, 49, 8282] classified 236 compounds into five functional group groups and proposed a model for flash point prediction using one or two core molecular descriptors. The model achieved the highest prediction performance, with a coefficient of determination of 0.855 for monohydric alcohols and a less satisfactory coefficient of determination of 0.370 for polyhydric alcohols.

[0020] In addition, as another model, predictions were performed using a combination of the group contribution method and an artificial neural network [Gharagheizi, F., Alamdari, RF, Angaji MT, Energy & Fuels 2008, 22, 1628]. The model used 1,378 compounds, treating a wide range of molecules. The coefficient of determination was 0.9757, and the absolute average error (AAE) was 8.101 K, showing a relatively good performance. However, the factors used as the input layer of the artificial neural network were in the form of fragments used for the group contribution method, so there is a concern that problems with the group contribution method may occur.

[0021] Therefore, the present invention provides a method for predicting flash point as follows: in order to shorten the calculation time that is a problem of the above-mentioned conventional methods and improve the accuracy of predicting physical properties, molecular descriptors that can serve as independent variables are screened out from molecular descriptors, and the accurate flash point of the pure compound is predicted by establishing an optimal multiple linear regression model for the flash point of the compound without additionally using an artificial neural network. Summary of the Invention

[0022] The technical problem to be solved by the present invention is to provide a more reliable QSPR model for the flash point of compounds composed of no more than 12 elements, including hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As), based on more experimental data.

[0023] This model is expected to overcome the shortcomings of the model constructed by the group contribution method as described above and to present fast and accurate prediction performance for a wider range of elements. The reason why the scope of compounds to which the prediction model can be applied in the present invention is limited to compounds composed of 12 or fewer elements such as hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As) is mainly because the model is based on experimental values ​​to form the prediction model, and is therefore limited to the constituent elements of the compound for which there are experimental values, but the number of various atoms is not limited.

[0024] As a predictive model capable of matching the composition of the constituent elements or molecular sizes of compounds with experimental values, it can be expanded to encompass a wider range of constituent elements and a greater number of atoms. Even within the aforementioned limited range of compounds, there are a vast number of compounds, including a significant number of industrially important ones. Based on the molecular structure of a compound, physical properties can be predicted in a very short time, facilitating the processing and analysis of physical property information on a large amount of compounds. Therefore, the present invention is expected to have significant beneficial effects on human society.

[0025] In order to achieve such a purpose, an example of the present invention is a method for predicting the flash point of a pure compound by a multiple linear regression model, the method comprising the following steps: (1) inputting the collected experimental data of the sample compound; (2) preparing the molecular descriptors of the sample compound inputted above; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data of step (1) into a training set and a test set; (5) using the molecular descriptors of step (3) to find the best multiple linear regression model that meets the above training set; (6) judging the rationality of the above multiple linear regression model; (7) if the above multiple linear regression model is not rational in step (6), repeating steps (5) and (6); if it is rational, testing the prediction performance of the model using the test set; (8) in the test of step (7), if the prediction performance does not meet the standard, repeating steps (4) to (7); if it meets the standard, predicting the flash point of the compound using the above multiple linear regression model, and the flash point prediction value obtained is the flash point value.

[0026] Wherein, the sample compound in step (1) is a pure compound composed of one or more elements selected from hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus and arsenic.

[0027] The sample compound in step (1) is one or more of the following: a. a hydrocarbon compound containing carbon and hydrogen elements; b. a non-hydrocarbon type I compound containing one or more of carbon, hydrogen, nitrogen, oxygen, and sulfur; c. a non-hydrocarbon type II compound containing one or more of carbon, hydrogen, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

[0028] Wherein, the preparation of molecular descriptors in step (2) refers to the preparation of molecular descriptors related to flash point.

[0029] Wherein, the optimal molecular descriptor in step (3) is one or more of the following: P1: rotatable bond ratio; P2: ring complexity index; P3: aromatic ratio; P4: 3rd order autoregressive path count; P5: conventional bond sequence ID number; P6: number of double bonds; P7: 8th order distance / detour ring index; P8: overall structural connectivity index; P9: structural information content index based on first-order neighborhood symmetry index; P 10 : Narumi harmonic topology index; P 11 : Total path count; P 12 : 0-order solvation connectivity index; P 13 : information content index based on the first-order neighborhood symmetry index; P 14 : The first-order spectral moment from the rate-weighted Barysz matrix.

[0030] The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values ​​from all sample compounds.

[0031] In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.

[0032] Wherein, the step (5) of searching for the best multiple linear regression model of the training set is to apply a stepwise selection method to explore the multiple linear regression model for the training set.

[0033] The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or the mean absolute error of the regression model.

[0034] The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.

[0035] The test in step (7) is to test the multiple linear regression model using a test set.

[0036] Among them, the satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.

[0037] Volatile substances can cause explosions or fires if not handled with care. To prevent this, VOC regulations are established and minimum flash points are determined through testing. Since vapors generated by volatile liquids become flammable gases when mixed with air, flash points provide useful information for the safe storage, handling, and transportation of volatile substances.

[0038] However, the number of compounds for which experimental values ​​are currently known does not exceed several thousand at most, and obtaining data through experiments is sometimes extremely difficult due to the toxicity, instability, and difficulty in purification of the compounds.

[0039] From this point of view, the present invention, which can obtain highly accurate flash point values ​​for many compounds using only information on the molecules without conducting experiments, not only saves the cost and time required for experiments, but also makes it possible to estimate the values ​​even when experiments are not possible, thereby facilitating research and development activities in related industries. Furthermore, appropriate information can be provided to all places such as academia and government that require the values, allowing such activities to be carried out more smoothly.

[0040] Furthermore, although three-dimensional molecular descriptors are insufficient because quantum chemical calculations are not performed, the physical properties of unknown compounds can be predicted in a very short time. Furthermore, the prediction time can be further shortened without using additional models such as artificial neural networks, thereby enabling real-time prediction of physical properties. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 4 is a flow chart showing a process of constructing a flash point QSPR model provided by the present invention.

[0042] Figure 2 It is a scatter plot (parity) of 3745 experimental data for the QSPR model of flash point provided by the present invention.

[0043] Figure 3 It is a bar chart of 3745 experimental data of the QSPR model. DETAILED DESCRIPTION

[0044] Hereinafter, a method for predicting the flash point of a pure compound using a multiple linear regression model according to the present invention will be described in detail with reference to the accompanying drawings so that persons having ordinary knowledge in the technical field to which the present invention pertains can easily implement the present invention.

[0045] In the process of describing the present invention, if it is determined that the detailed description of the related known technology may cause unnecessary obscurity to the gist of the present invention, the detailed description will be omitted.

[0046] When using the words "including," "having," "forming," "comprising," and the like in this specification, other parts may be added unless "only" is used. When describing a component in the singular, the plural is included unless otherwise specified.

[0047] Figure 1 A flow chart schematically illustrating the process of constructing a QSPR model of a flash point.

[0048] The first task in building the model is to collect experimental data and analyze and classify them as specified in the first step.

[0049] For the present invention, we conducted an extensive search of all available references and materials, including various papers, monographs, and websites, to collect experimental flash point data for compounds meeting the criteria of the present invention. We extensively examined whether the collected data represented truly reasonable values ​​for model construction. We carefully analyzed and modified or deleted data for compounds that (i) were non-experimental values, (ii) had data labeling errors, (iii) had values ​​that differed significantly even for the same compound, (iv) were unreliably deviated from the values ​​of other similar compounds, or (v) had values ​​for which molecular descriptors were difficult to readily generate. Finally, we selected data for a total of 3,745 compounds.

[0050] In addition, when constructing the physical property prediction model, the sample compounds were divided into (a) hydrocarbons (hydrocarbon) consisting only of C and H, (b) non-hydrocarbons (nonhydrocarbon1) including one or more of the elements C, H, N, O, and S, and (c) non-hydrocarbons (nonhydrocarbon2) including one or more of the elements C, H, F, Cl, Br, I, Si, P, and As, and models were established for each type.

[0051] Because the three aforementioned classifications demonstrated more effective prediction performance, they were divided into 409 hydrocarbons, 889 non-hydrocarbons (I), and 2,447 non-hydrocarbons (II), and models were established for each. Furthermore, in this invention, "compound" refers to substances formed from molecules composed of no more than 12 elements: hydrogen (H), carbon (C), nitrogen (N), oxygen (O), sulfur (S), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), silicon (Si), phosphorus (P), and arsenic (As).

[0052] The second step is to prepare molecular descriptor values ​​for these compounds. Molecular descriptors suitable for use as model variables are calculated using the following method for the values ​​of over 3,779 molecular descriptors: a file (MOL or SDF file) containing structural information, including the X, Y, and Z coordinates of each molecule's constituent elements and the bonding patterns between them, is input into a computer program. Based on this input, molecular descriptors representing the diversity of molecules are calculated using various theoretical mathematical formulas, moving beyond simple element and bond counts. Furthermore, the process of extracting these 3,779 molecular descriptors is repeated for all sample compounds used to construct the physical property prediction model.

[0053] The theoretical mathematical formulas used in the calculations are known in the art and are used in various forms, including, for example, Chi Connectivity Indices, Topological Descriptors, Information Indices, Walk and Path Counts, 2D Autocorrelation Descriptors, and the like.

[0054] In the best results, not only physical property information can be obtained, but also molecular descriptors that reflect the characteristics of the molecules and are expressed in a variety of meaningful numerical values ​​can be obtained.

[0055] For pure substances, there are molecular descriptors that can characterize both two-dimensional and three-dimensional structures. However, three-dimensional molecular descriptors are affected by molecular structure and are more complex to compute. Therefore, in this invention, in order to rapidly process physical property predictions without performing quantum mechanical calculations for structure optimization, three-dimensional molecular descriptors are not used. This approach significantly reduces the computational time and cost of quantum mechanical calculations.

[0056] Furthermore, in the third step, after calculating the molecular descriptor values, it is necessary to screen out inappropriate molecular descriptor values. Specifically, by applying a theoretical mathematical formula to the molecular descriptors of the sample compounds, the system screens out molecular descriptor values ​​that yield the same value for all sample compounds and therefore cannot serve as independent variables in the model. This screening of molecular descriptors prevents the inclusion of irrelevant molecular descriptors in the prediction model, thereby improving model reliability while reducing the number of molecular descriptors, the number of calculations, and the computational effort, ultimately shortening the computational time required to find the optimal model.

[0057] In the present invention, when the sample compound is (a) a hydrocarbon composed only of C and H, the optimal molecular descriptors obtained include the proportion of rotatable bonds, the ring complexity index, the aromatic ratio, the 3rd-order autoregressive path count and the conventional bond sequence ID number; when the sample compound is (b) a non-hydrocarbon 1 including one or more of the elements C, H, N, O, and S, the optimal molecular descriptors obtained include the conventional bond sequence ID number, the number of double bonds, the 8th-order distance / detour ring index, the overall structural connectivity index and the structural information content index based on the first-order neighborhood symmetry index; when the sample compound is (c) a non-hydrocarbon 2 including one or more of the elements C, H, F, Cl, Br, I, Si, P, and As, the optimal molecular descriptors obtained include the Narumi harmonic topology index; the total path count, the 0th-order solvation connectivity index, the information content index based on the first-order neighborhood symmetry index and the first-order spectral moment from the rate-weighted Barysz matrix.

[0058] In the fourth step, the sample compounds are divided into a training set and a test set. The training set will be used to find the prediction model, and the test set will be used to test the predictive performance of the determined model. While taking care not to bias the distribution of similar molecules to one side, the training set and the test set are divided in a ratio of 5-8:2-5, preferably 7:3.

[0059] In the fifth step, the optimal multiple linear regression model is found based on the training set. Here, "optimal" is used in a relative sense, indicating that it can be found in a relatively short time while having performance very close to the optimal solution in an absolute sense.

[0060] The reason for not directly finding the optimal solution is that the calculation time is long. For example, when the number of suitable molecular descriptors among 3779 molecular descriptors is 1700, the total number of multiple linear regression models that can be made by extracting 5 different molecular descriptors is It is therefore practically impossible to investigate them all.

[0061] In order to obtain useful results within a limited time, a stepwise selection method is adopted in the present invention.

[0062] The stepwise selection method is a common method for selecting variables from data with many independent variables. During each step of creating a regression equation, the following steps are repeated: Among the independent variables not already in the equation, variables with high correlations with the dependent variable are added by entering a benchmark, and among the independent variables already in the equation, variables with low correlations with the dependent variable are removed by removing the benchmark. If no further variables are added or removed during this process, the selection of independent variables ends, and the model based on the stepwise selection method is finally completed.

[0063] In other words, a method of using the so-called hypothesis test to determine whether the hypothesis about the population is statistically correct. The significance level is generally set to 0.05 (5%), and the p-value (p-value) as the probability of significance is calculated. When the p-value is less than 0.05, the null hypothesis is rejected and the alternative hypothesis is adopted. By performing hypothesis testing on the selected variables, the process of adopting newly added variables or removing existing variables is repeated. Ultimately, when there are no new variables to be added or removed, the stepwise selection process ends. After the stepwise selection is successful, the molecular descriptor information obtained will perform a more advanced search, and a model with reliable accuracy will be obtained after the statistical test is passed, as shown below:

[0064]

[0065] Among them, P represents the molecular descriptor after stepwise selection, Y is the property to be predicted, and C is the parameter estimated by the experimental data through the multivariate regression method.

[0066] Preferably, the least square residual is used to estimate the parameter C. The specific calculation process is as follows:

[0067]

[0068] Where n is the number of substances to be predicted and Pn is the number of optimal molecular descriptors.

[0069] Once the best multiple linear regression model is selected, the rationality of the model is determined in the sixth step, which is the next step. Specifically, the rationality of the model can be tested using the t-test method, and the calculation process is as follows:

[0070]

[0071] According to the above method, if the t-test value of a molecular descriptor included in the model does not meet the requirements, the process returns to the previous step to search for another model. For example, if the number of sample compounds is 1005 and the selected model consists of five molecular descriptors, if the t-test value for one molecular descriptor is 3.3 or higher, the probability that the molecular descriptor is not related to the corresponding physical property is less than 0.1%.

[0072] In the present invention, when there is a molecular descriptor with a t-test value less than about 3, the selected model is discarded and other models are sought. In addition, the situation that the value of a molecular descriptor for the sample compound is the same except for a few compounds cannot be regarded as a model with reliability, so measures are also taken. Usually, if the number of molecular descriptors included in the model is increased, the prediction performance will improve, but the above-mentioned problems will occur. Therefore, usually, the final model is obtained by repeatedly performing these steps by changing the number of molecular descriptors and through multiple trial and error. If the selected model no longer has problems, then proceed to the next step.

[0073] Next, in the seventh step, the prediction performance of the found model is evaluated using a test set that does not participate in forming the model.

[0074] If significant performance degradation or deviations are observed in the training set, the fourth step involves realigning the training and test sets before proceeding to the next steps. If the difference between the training and test sets does not exceed 20% of the mean absolute error (AAE) obtained for the training set, the prediction performance is considered satisfactory and a multiple linear regression model for flash point is established.

[0075] The results of the model finally established through this process are briefly shown in Table 1, Table 2, and Table 3 below.

[0076] Table 1

[0077] Main contents of the QSPR prediction model for hydrocarbon flash point

[0078]

[0079]

[0080] Table 2 Main contents of the QSPR prediction model for the flash point of non-hydrocarbons

[0081]

[0082]

[0083]

[0084]

[0085] Table 3 Main contents of the QSPR prediction model for the flash point of non-hydrocarbons

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] Example

[0092] This example is used to further illustrate the accuracy of the present invention in predicting the flash point of a pure compound, but the present invention is not limited thereto.

[0093] Table 4 below shows some of the results of predicting the flash point values ​​of pure compounds using the model described in the present invention:

[0094] Table 4

[0095]

[0096]

[0097]

[0098] To confirm the excellence of the present invention, the flash points of 3745 compounds with known experimental values ​​were predicted using the multiple linear regression model of the present invention, resulting in a coefficient of determination value of 0.94011 and a mean absolute error value of 13.13261 K, indicating excellent prediction performance.

[0099] Figure 2 This is a scatter plot showing the prediction performance of the multiple linear regression model. It can be visually confirmed that the comparison data with the experimental values ​​are concentrated on the diagonal line, indicating excellent performance.

[0100] Figure 2This is a scatter plot showing the predictive performance of a multiple linear regression model. It evaluates how closely experimental values ​​are aligned with predicted values. The X-axis represents the experimental values, and the Y-axis represents the predicted values. The experimental and predicted values ​​are plotted at the (X, Y) coordinates with dots (hollow blue circles), and the baseline representing Y=X is marked with a diagonal line (solid red line).

[0101] Points located on the baseline indicate that the predicted value is the same as the experimental value. Points located below the baseline indicate that the predicted value is smaller than the experimental value. Points located above the baseline indicate that the predicted value is larger than the experimental value. For example, if the experimental and predicted values ​​are (100, 100), they are on the baseline; if the experimental and predicted values ​​are (100, 90), they are below the baseline; and if the experimental and predicted values ​​are (100, 10), they are above the baseline.

[0102] exist Figure 2 In the parity plot, we can observe that there are some points that are relatively separated from the baseline in the middle. However, in most areas, the points are concentrated on the baseline. The points are distributed on the baseline in a spindle shape and spread thinner upward and downward rather than widely. Therefore, the predicted value can be said to be close to the experimental value as a whole.

[0103] In addition, in the experimental data of 3745 compounds, the average percentage error of the known experimental error is about 9.13%, Figure 3 The error between the experimental and predicted values ​​is plotted in a bar graph centered around this value. In this graph, the multiple linear regression model predicts flash point values ​​within the average experimental error with a probability of 93.43%, demonstrating its high predictive performance.

[0104] The present invention is not limited to the above-mentioned embodiments. Without departing from the gist of the invention claimed for protection in the claims, anyone with general knowledge in the technical field to which the present invention belongs can make various modified implementations, and the above-mentioned changes fall within the scope of the claims.

Claims

1. A method for predicting the flash point of a pure compound using a multiple linear regression model, the method comprising the following steps: (1) Input the collected experimental data of the sample compound; (2) preparing a molecular descriptor for the sample compound input above; (3) screening the best molecular descriptors from step (2); (4) separating the experimental data in step (1) into a training set and a test set; (5) using the molecular descriptors described in step (3) to find the best multiple linear regression model that matches the training set; (6) Determine the rationality of the above multiple linear regression model; (7) If the multiple linear regression model is not reasonable in step (6), repeat steps (5) and (6). If it is reasonable, test the prediction performance of the model using the test set; (8) In the test of step (7), if the above prediction performance does not meet the standard, repeat steps (4) to (7); if the standard is met, the flash point of the compound will be predicted using the above multiple linear regression model, and the flash point prediction value obtained is the flash point value.

2. The method according to claim 1, wherein The sample compound in step (1) is a pure compound composed of one or more elements selected from the group consisting of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

3. The method according to claim 2, wherein: The sample compound in step (1) is one or more of the following: a. Hydrocarbon compounds containing carbon and hydrogen elements; b. Non-hydrocarbon compounds containing one or more of carbon, hydrogen, nitrogen, oxygen and sulfur; c. Non-hydrocarbon compounds containing one or more elements of carbon, hydrogen, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic.

4. The method according to any one of claims 1 to 3, wherein The step (2) of preparing molecular descriptors refers to preparing molecular descriptors related to flash point.

5. The method according to any one of claims 1 to 3, wherein The optimal molecular descriptor in step (3) is one or more of the following: P1: proportion of rotatable bonds; P2: Cyclomatic complexity index; P3: aromatic ratio; P4: 3rd order autoregressive path count; P5: Conventional key sequence ID number; P6: number of double bonds; P7: 8th order distance / detour index; P8: overall structural connectivity index; P9: Structural information content index based on the first-order neighborhood symmetry index; P 10 : Narumi harmonic topology index; P 11 : total path count; P 12 : 0-order solvation connectivity index; P 13 : information content index based on the first-order neighborhood symmetry index; P 14 : The first-order spectral moment from the rate-weighted Barysz matrix.

6. The method according to claim 5, wherein: The optimal molecular descriptor in step (3) is an independent molecular descriptor having different values ​​from all sample compounds.

7. The method according to any one of claims 1 to 3, wherein In step (4), the training set and the test set are divided in a ratio of 5-8:2-5.

8. The method according to any one of claims 1 to 3, wherein The step (5) of finding the best multiple linear regression model of the training set is to apply a stepwise selection method to explore the multiple linear regression model for the training set.

9. The method according to claim 8, wherein The optimal multiple linear regression model in step (5) is found by judging the prediction performance by the coefficient of determination or mean absolute error of the regression model.

10. The method according to any one of claims 1 to 3, wherein The rationality of step (6) is determined by the t-test value, wherein the model is rational when the t-test value is greater than or equal to 3, otherwise it is not rational.

11. The method according to any one of claims 1 to 3, wherein The test in step (7) is to test the multiple linear regression model using the test set.

12. The method according to any one of claims 1 to 3, wherein The satisfaction standard in step (8) is achieved by comparing the difference between the training set and the test set. If the difference between the training set and the test set does not exceed 20% of the mean absolute error obtained for the training set, it is judged that the performance meets the standard; otherwise, the performance does not meet the standard.