Intelligent prediction method and system for phase equilibrium of natural gas hydrate in organic solution system
By building a machine learning model based on intelligent optimization and fusing organic solute molecular characteristics and functional group information, the problem of complexity and insufficient accuracy in the prediction of natural gas hydrate phase equilibrium is solved, and high-precision and stable prediction effects are achieved.
Patent Information
- Application Number
- CN202510457439.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional natural gas hydrate phase equilibrium calculation methods require a large amount of experimental data to calibrate parameters, the calculations are complex and cannot be accurately predicted, especially when organic chemical inhibitor concentration changes.
By fusing organic solute molecular characteristics and functional group information, a machine learning model based on intelligent optimization is constructed. The specific steps include experimental data set construction, feature extraction and selection, model construction and training, and prediction is performed using the CatBoost-Optuna model.
It improves the accuracy and stability of the phase equilibrium prediction of natural gas hydrate, has excellent generalization ability and practical value, and can accurately predict the phase equilibrium temperature of new samples.
Smart Images

Figure CN119993318A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of petroleum engineering, and in particular relates to an intelligent prediction method and system for natural gas hydrate phase equilibrium in an organic solution system. Background Art
[0002] Energy shortage is an important issue facing the rapid economic development of the world today. Natural gas hydrates, as a highly efficient and clean resource, are widely distributed in permafrost, shallow strata in the sea and other areas around the world. According to statistics, its total resource volume even exceeds the sum of traditional fossil energy, which has attracted widespread attention. Natural gas hydrates are a non-stoichiometric crystalline inclusion complex formed by water and gas under low temperature and high pressure conditions. As one of the most important properties, phase equilibrium conditions are the thermodynamic basis that must be met for the growth and decomposition of hydrates. In the process of oil and gas extraction and gathering, the formation of natural gas hydrates may block pipelines and equipment. In addition, natural gas hydrates can also be used in gas extraction and separation, energy storage, and CO 2 Sealing, etc. Therefore, in many application scenarios in the above-mentioned energy and environmental fields, the prediction of natural gas hydrate phase equilibrium has important research value.
[0003] Traditional methods for calculating the phase equilibrium of natural gas hydrates are mainly based on thermodynamic theory and empirical models. The formation or decomposition conditions of hydrates are determined by describing the chemical potential equilibrium relationship of different components in each phase (gas phase, water phase and hydrate phase) under different conditions. Among them, the thermodynamic model is the core method for calculating the phase equilibrium of natural gas hydrates. It is based on the principles of classical thermodynamics, assumes that the chemical potential of the system is equal in the equilibrium state, and combines the multiphase system characteristics of hydrate formation to predict the phase equilibrium conditions. Although traditional thermodynamic models have been successfully applied to predict hydrate phase equilibrium, they require a large amount of experimental data to calibrate parameters (such as chemical potential, activity coefficient, interaction parameters, etc.). When there are multiple gas components or complex solute types (electrolytes, organic matter, etc.), the computational complexity of the model is greatly increased, and existing thermodynamic models often cannot accurately predict. In particular, for different organic chemical inhibitor concentrations, the relationship between the hydrate equilibrium temperature and pressure is not linear. Relying on experimental methods to measure one by one in the multidimensional space of component-component concentration-temperature and pressure requires a lot of experimental costs, and the complexity of the interaction between solute molecules, water, and guest molecules makes it extremely difficult to accurately construct traditional thermodynamic models.
[0004] With the development of computer technology, artificial intelligence algorithms have become an important hot topic, and machine learning is currently the most important method for implementing artificial intelligence. It can automatically learn potential laws in complex data. In other words, machine learning can be used to simulate complex systems. At present, many studies have applied machine learning to the field of hydrate phase equilibrium prediction. Existing studies have shown that under the conditions of high complexity of gas-solution systems, machine learning models can effectively make up for the shortcomings of traditional thermodynamic models that are difficult to characterize the interactions between components and accurately predict the phase equilibrium state, and achieve fast and high-precision phase equilibrium prediction. Summary of the invention
[0005] In view of the deficiencies of the prior art, the present invention provides an intelligent prediction method and system for natural gas hydrate phase equilibrium in an organic solution system; The present invention integrates the molecular characteristics of organic solutes and functional group information to construct a machine learning model based on intelligent optimization, which can avoid model overfitting and improve the prediction accuracy of natural gas hydrate phase equilibrium.
[0006] The technical solution of the present invention is: An intelligent prediction method for natural gas hydrate phase equilibrium in an organic solution system includes: Step 1: Experimental data set construction; including: creating experimental data set and performing data preprocessing; Step 2: Feature extraction; including: extracting organic molecular features; extracting functional group features; Step 3: Feature selection; including: removing redundant features based on correlation heat map; Boruta feature screening; Step 4: Build and train the CatBoost-Optuna model; Step 5: After preprocessing, feature extraction and feature selection, the data to be predicted is input into the trained CatBoost-Optuna model to output the prediction results.
[0007] Preferably, according to the present invention, creating an experimental data set includes: Each sample data includes: organic matter name, mole fraction concentration (Mol_fr), equilibrium pressure (Mpa), equilibrium temperature (K). By changing the sample type and mole fraction concentration, measuring the equilibrium pressure and equilibrium temperature under different pressure conditions, and summarizing the data, an experimental data set is formed.
[0008] Preferably, according to the present invention, data preprocessing includes: ① Eliminate outliers; ②Use the data normalization function to normalize the data; ③ Use random oversampling to find sparse samples through the quantile of the target value, concatenate the sparse samples with the original data, and generate an oversampled data set.
[0009] Preferably, according to the present invention, extracting the molecular characteristics of organic matter comprises: RDKit was used to extract and quantify the features of 34 water-soluble organic substances contained in the dataset.
[0010] Preferably, according to the present invention, extracting functional group characteristics comprises: The frequency of occurrence of each functional group in the statistical data set was counted, and the 12 most frequently occurring functional groups were obtained, including: Hydroxyl, Methyl, Alcohol, Ketone, Carboxyl, Amino, Ethyl, Epoxide, Propyl, Alkene, Imine and Ester.
[0011] Preferably, according to the present invention, redundant features are removed; comprising: Use the matplotlib library in Python to draw feature heat maps based on the extracted features; The strongly correlated features were excluded, and the final excluded features were: Alcohol, Carboxyl, Amino, NumAtom, Aspa, NumHetero, Psa; The final heat map is obtained after removing the strongly correlated features.
[0012] Preferably, according to the present invention, the Boruta feature selection algorithm is used for feature screening; comprising: The Boruta feature selection algorithm was used for feature screening, and the input features of the experiment were determined to be: compound molar fraction concentration Mol_fr, ratio of polar solvent accessible surface area to total solvent accessible surface area Psa_div_sa, total solvent accessible surface area Sa, molecular weight Mw, propyl Propyl, ethyl Ethyl, logarithmic octanol partition coefficient to water Logp, methyl Methyl, hydroxyl Hydroxyl, equilibrium pressure P_Mpa, ketone Ketone; the output feature was equilibrium temperature (T_K). The feature correlation heat map after feature screening is shown in the figure below. Figure 4 shown.
[0013] Preferably, according to the present invention, a CatBoost-Optuna model is constructed; comprising: The CatBoost-Optuna model consists of the CatBoost model and the Optuna framework; Suppose the sample data set D has n features and Y iis the label value of the target variable, randomly sort the samples in D and generate multiple sets of sequences, one of which is a random sequence σ=[σ 1 ,σ 2 ,…,σ n ], the CatBoost model replaces the random sequence with the following equation: ; in, X σp,q is the characteristic σ p No. q training values; X σj,q is the characteristic σ j No. q training values; Y σj is the partial target value of the corresponding feature; P is the prior value; β is the prior weight value; for regression tasks, the method for calculating the prior value is to take the average value of the data set; Further preferably, the training process of the CatBoost-Optuna model specifically includes: (1) Input data: The input data is the number of iterations I and the characteristic matrix X of natural gas hydrate. i and the corresponding target value y i ; (2) Initialize the model: Before training the CatBoost-Optuna model, initialize the prediction values of all samples; (3) Optuna hyperparameter optimization: Use the Optuna optimizer to optimize the hyperparameters of the CatBoost-Optuna model; (4) Model training and iteration: ① Calculate the gradient: For each iteration iter=1,2,3,…,I, calculate the gradient g between the current CatBoost-Optuna model prediction value and the true target value i ; ② Update the predicted value: For each sample j, use the CatBoost-Optuna model to learn the current gradient g i and feature X j , and update the predicted value M j , and then the updated prediction value M j Accumulated to the predicted value M of the current sample i middle; ③ Iteration process: Repeat the above steps of calculating gradients and updating prediction values until the predetermined number of iterations I is completed; (5) Output model: After the CatBoost-Optuna model training and iteration are completed, the final CatBoost-Optuna model M is output.
[0014] A computer device comprises a memory and a processor, wherein the memory stores a computer program and the processor implements the steps of an intelligent prediction method for phase equilibrium of natural gas hydrate in an organic solution system when executing the computer program.
[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of an intelligent prediction method for natural gas hydrate phase equilibrium in an organic solution system.
[0016] The intelligent prediction system for natural gas hydrate phase equilibrium in organic solution system includes: The experimental data set building module is configured to: create an experimental data set and perform data preprocessing; The feature extraction module is configured to: include: extracting organic molecular features; extracting functional group features; The feature selection module is configured to include: removing redundant features according to the correlation heat map; Boruta feature screening; The model building and training module is configured to: build and train the CatBoost-Optuna model; The prediction module is configured to: input the data to be predicted into the trained CatBoost-Optuna model after preprocessing, feature extraction and feature selection, and output the prediction result.
[0017] The beneficial effects of the present invention are: The CatBoost-Optuna model has better prediction ability on known data sets than the traditional phase equilibrium model. By testing the phase equilibrium experimental data of 34 organic solvents in the data set, the CatBoost-Optuna prediction process has smaller errors and higher prediction accuracy. In addition, the CatBoost-Optuna model performs very well in extrapolation experiments, has excellent generalization ability, strong stability and reliability. The model has high practical value in practical applications and can accurately predict the phase equilibrium temperature of new samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Heatmap for raw correlation analysis; Figure 2 This is the correlation heat map after removing the strongly correlated features; Figure 3 This is a schematic diagram of the feature selection results of the Boruta algorithm; Figure 4This is the correlation analysis heat map after Boruta feature screening; Figure 5 This is a comparison chart of the root mean square error (RMSE) results of each model predicting the natural gas hydrate formation temperature; Figure 6 This is a comparison chart of the mean absolute percentage error (MAPE) results of each model predicting the natural gas hydrate formation temperature; Figure 7 Comparison chart of the coefficient of determination (R²) of each model predicting the gas hydrate formation temperature; Figure 8 This is a schematic diagram comparing the prediction performance of the CatBoost-optuna model and the CatBoost model; Fig. 9 It is the residual graph of CatBoost-Optuna model; Fig.10 This is the scatter plot of the CatBoost-Optuna model; Fig.11 It is the scatter plot of THERM model; Fig.12 This is the scatter plot of the Catboost-Optuna model; Fig.13 is the residual graph of the THERM model; Fig.14 It is the residual graph of CatBoost-Optuna model; Fig.15 is the error histogram of THERM model; Fig.16 is the Catboost-Optuna model error histogram; Fig.17 This is the residual plot of the CatBoost-Optuna model extrapolation experiment; Fig.18 This is a scatter plot of the CatBoost-Optuna model extrapolation experiment; Fig.19 Flowchart for CatBoost-Optuna model training. DETAILED DESCRIPTION
[0019] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.
[0020] Example 1 An intelligent prediction method for natural gas hydrate phase equilibrium in an organic solution system includes: Step 1: Experimental data set construction; including: creating experimental data set and performing data preprocessing; Step 2: Feature extraction; including: extracting organic molecular features; extracting functional group features; Step 3: Feature selection; including: removing redundant features based on correlation heat map; Boruta feature screening; Step 4: Build and train the CatBoost-Optuna model; Step 5: After preprocessing, feature extraction and feature selection, the data to be predicted is input into the trained CatBoost-Optuna model to output the prediction results.
[0021] Example 2 The difference between the method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system described in Example 1 is that: Create an experimental dataset, including: Construct a phase equilibrium experiment measurement system, which includes a sample pool, a temperature control system, a pressure control system, a data acquisition system, a stirring device, and a safety device. The temperature control system accurately controls the temperature of the sample pool through a heating or cooling device; the pressure control system accurately controls the pressure of the sample pool through a pressure sensor and a regulating valve; the data acquisition system records temperature, pressure, concentration and other data in real time through a temperature sensor and a pressure sensor; the stirring device ensures that the sample is evenly mixed to avoid local concentration differences; the safety device includes a pressure relief valve and an emergency stop device to ensure experimental safety.
[0022] The experimental steps include: (1) Sample preparation: Prepare organic samples with different mole fraction concentrations; (2) System calibration: calibrate temperature and pressure sensors; (3) Experimental operation: ① inject the sample into the sample pool, ② set the initial pressure, ③ wait for the system to reach equilibrium, ④ wait for the system to reach equilibrium; (4) Data recording: Record the name, mole fraction concentration, equilibrium pressure and equilibrium temperature of each sample.
[0023] Each sample data includes: organic matter name, mole fraction concentration (Mol_fr), equilibrium pressure (Mpa), equilibrium temperature (K). By changing the sample type and mole fraction concentration, measuring the equilibrium pressure and equilibrium temperature under different pressure conditions, and summarizing the data, an experimental data set is formed.
[0024] Data preprocessing, including: ① Eliminate outliers: While retaining the original data as much as possible, we compared the experimental data under similar conditions with the experimental researchers and the simulation results under the same conditions, and eliminated the obvious outliers to obtain a data set containing 788 samples. The data is summarized in Table 1.
[0025] Table 1 The original data and the data table after removing outliers;
[0026] ②Use the data normalization function to normalize the data; ③Use random oversampling to find sparse samples through the quantile of the target value, and simply splice the sparse samples with the original data to generate an oversampled data set. Specifically, it includes: loading data and separating the majority class and the minority class. In this experiment, samples whose target value equilibrium temperature is less than the 10% quantile of the target value equilibrium temperature are regarded as sparse samples (minority class samples). Then, the minority class samples are copied using the random sampling method, and the copied samples are spliced with the original data to generate an oversampled data set.
[0027] Extraction of molecular characteristics of organic matter; including: RDKit was used to extract and quantify the features of 34 water-soluble organic substances in the dataset. The results are shown in Table 2.
[0028] Table 2 Characteristics of water-soluble organic matter for extraction and quantification;
[0029] Extract functional group features; including: Several common functional group features were extracted from the previous studies of various polymers, including 34 common functional groups (such as hydroxyl, methyl, ester, etc.). The frequency of occurrence of each functional group in the statistical data set was calculated, and the 12 most frequently occurring functional groups were obtained, including: Hydroxyl, Methyl, Alcohol, Ketone, Carboxyl, Amino, Ethyl, Epoxide, Propyl, Alkene, Imine and Ester. The results are organized into a table, as shown in Table 3.
[0030] Table 3 Table of the 12 most frequently occurring functional groups;
[0031] Remove redundant features; including: Based on the extracted features, use the matplotlib library in Python to draw a feature heat map; Figure 1 The horizontal and vertical axes are the molecular features and functional groups extracted from organic matter, and the specific information is shown in Table 2 and Table 3; The strongly correlated features are excluded. Figure 1The correlation between the features can be seen in the figure. (1) Hydroxyl and Alcohol are strongly correlated, because hydroxyl is the core functional group of alcohol molecules, which means that they express the same chemical properties to some extent. If a molecule is detected to contain Hydroxyl, it can usually be considered as Alcohol, so Alcohol is excluded. (2) Ester, Carboxyl and Amino are strongly correlated. The reason for the strong correlation is the coexistence of these three functional groups in the molecular structure. In some compounds, carboxyl, ester and amino groups may coexist, resulting in Statistical correlation, so Carboxyl and Amino are excluded; (3) Mw, NumAtom and Aspa are strongly correlated, mainly because they are jointly dependent on the size and composition of the molecule. Molecules composed of more atoms are usually heavier and have a larger surface area, so NumAtom and Aspa are excluded; (4) NumHacc and NumHetero are strongly correlated. Heteroatoms usually act as hydrogen bond receptors, so the number of heteroatoms directly affects the number of hydrogen bond receptors, so NumHetero is excluded; Psa is a component of Sa and is strongly correlated in many cases, so Psa is excluded. Therefore, the final excluded features are: Alcohol, Carboxyl, Amino, NumAtom, Aspa, NumHetero, Psa; After removing the strongly correlated features, the final heat map is obtained. Figure 2 shown.
[0032] The Boruta feature selection algorithm is used for feature screening; including: The Boruta feature selection algorithm is used. The Boruta feature selection algorithm is a feature selection algorithm based on random forest. The algorithm can screen out variables that are highly important to the target variable from many feature variables, obtain the importance ranking of feature variables, and provide the best classification accuracy for the selected feature variables.
[0033] The Boruta algorithm screens features by generating randomly arranged shadow features and inputting them into the random forest model together with the original features, comparing their importance scores. If the importance of the original feature is significantly higher than the maximum importance of the shadow feature, it is considered an important feature; if it is significantly lower, it is an unimportant feature; if they are similar, it is temporarily uncertain and further iteration is required. Finally, the Boruta algorithm outputs a list of important features to help identify features that contribute to model predictions and exclude redundant or irrelevant features. It is suitable for feature selection of high-dimensional data sets.
[0034] The feature screening results of Boruta algorithm are as follows Figure 3The vertical axis is the molecular features and functional groups extracted from organic matter, the specific information is shown in Table 2 and Table 3, and the horizontal axis is the order of feature importance; Figure 3 The results show that 11 out of 19 features are defined as important features, namely: compound mole fraction concentration (Mol_fr), ratio of polar solvent accessible surface area to total solvent accessible surface area (Psa_div_sa), total solvent accessible surface area (Sa), molecular weight (Mw), propyl (Propyl), ethyl (Ethyl), logarithmic octanol partition coefficient for water (Logp), methyl (Methyl), hydroxyl (Hydroxyl), equilibrium pressure (P_Mpa), ketone (Ketone). 8 features are defined as unimportant features: number of rotatable bonds (NumRT), number of hydrogen bond donors (NumHdon), number of hydrogen bond acceptors (NumHacc), olefins (Alkene), epoxides (Epoxide), imines (Imine), esters (Ester), ratio of heteroatoms to total atoms (HeteroFr). The Boruta feature selection algorithm was used for feature screening, and the final input features of the experiment were determined as follows: compound molar fraction concentration Mol_fr, ratio of polar solvent accessible surface area to total solvent accessible surface area Psa_div_sa, total solvent accessible surface area Sa, molecular weight Mw, propyl Propyl, ethyl Ethyl, logarithmic octanol partition coefficient to water Logp, methyl Methyl, hydroxyl Hydroxyl, equilibrium pressure P_Mpa, ketone Ketone; the output feature was equilibrium temperature (T_K). The feature correlation heat map after feature screening is shown in the figure below. Figure 4 shown.
[0035] Build a CatBoost-Optuna model; including: The CatBoost-Optuna model consists of the CatBoost model and the Optuna framework; The Optuna framework is a software framework that can automatically optimize hyperparameters based on the Bayesian algorithm. It finds the optimal parameter value for the best performance through trial and error. Optuna mainly determines the combination of hyperparameter values that need to be tested next based on historical running data. The optimization method of this software framework belongs to the tree Parzen optimizer in the improved Bayesian optimization algorithm. The tree structured Parzen optimizer (TPE) simulates p(x|y) by transforming the generation process and replacing the previously configured distribution with a non-parametric density.
[0036] ; in, is the best value found after observation, is the value of different observations Observe the density formed, so that the corresponding loss Less than . is the density formed by the remaining observations.
[0037] The CatBoost model is an improved ensemble learning algorithm based on the gradient boosting tree framework, using the symmetric decision tree (Oblivious Tree) as a weak learner, Ordered Boosting as a strong learner, and integrating categorical feature processing. The CatBoost algorithm mainly proposes two key methods, one is an algorithm based on statistical learning for processing categorical features, and the other is a ranking boosting algorithm.
[0038] Suppose the sample data set D has n features and Y i is the label value of the target variable, randomly sort the samples in D and generate multiple sets of sequences, one of which is a random sequence σ=[σ 1 ,σ 2 ,…,σ n ], the CatBoost model replaces the random sequence with the following equation: ; in, X σp,q is the characteristic σ p No. q training values; X σj,q is the characteristic σ j No. q training values; Y σj is the partial target value of the corresponding feature; P is the prior value; β is the prior weight value; for regression tasks, the method for calculating the prior value is to take the average value of the data set; The traditional gradient boosting tree calculates the gradient of the loss function for the current model for the same data set at each iteration, and then trains the base learner based on this gradient. However, this method will cause estimation bias in the point-by-point gradient, which will eventually lead to model overfitting. The CatBoost algorithm proposes to use the Order boosting method to change the gradient estimation method in the traditional algorithm, thereby obtaining an unbiased estimate of the gradient, reducing the impact of estimation bias, and improving the generalization ability of the model.
[0039] The input of the CatBoost model is the training set and the number of iterations; First, the initial prediction values of all samples are set to 0; Then, start iterative training; In each round of training, each sample is processed: all samples before the current sample are traversed, and a sub-model is trained using the features and gradients of all samples before the current sample; and the prediction value of the sub-model is accumulated to the prediction value of the current sample; after traversing all samples before the current sample, the prediction value of the current sample is updated to the cumulative result of the prediction values of all previous sub-models; Finally, the final prediction value of all samples is returned. This algorithm gradually updates the prediction value of the current sample by using the information of the previous sample. It is suitable for tasks that are sensitive to data order and has high flexibility and optimization capabilities.
[0040] like Fig.19 As shown in the figure, the training process of the CatBoost-Optuna model specifically includes: (1) Input data: The input data is the number of iterations I and the characteristic matrix X of natural gas hydrate. i and the corresponding target value y i ; Among them, X i Represents the feature vector of the sample, y i Represents the target value of the sample. These data are used for model training and hyperparameter optimization.
[0041] (2) Initialize the model: Before training the CatBoost-Optuna model, initialize the prediction values of all samples; specifically, initialize the initial prediction value M of each sample. i Set to 0, that is, Mi←0, where i=1,2,…,n, and n is the total number of samples.
[0042] (3) Optuna hyperparameter optimization: Use the Optuna optimizer to optimize the hyperparameters of the CatBoost-Optuna model. The specific steps include: ① Define the hyperparameter search space, including key parameters such as learning rate and tree depth, ②Use the loss function L and the input data set {X i ,y i} as the optimization target, and search in the hyperparameter space through Optuna’s Bayesian optimization algorithm. ③Finally obtain the optimal hyperparameter set H.
[0043] (4) Model training and iteration: ① Calculate the gradient: For each iteration iter=1,2,3,…,I, calculate the gradient g between the current CatBoost-Optuna model prediction value and the true target value i ; ② Update the predicted value: For each sample j, use the CatBoost-Optuna model to learn the current gradient g iand feature X j , and update the predicted value M j , and then the updated prediction value M j Accumulated to the predicted value M of the current sample i middle; ③ Iteration process: Repeat the above steps of calculating gradients and updating prediction values until the predetermined number of iterations I is completed; (5) Output model: After the CatBoost-Optuna model training and iteration are completed, the final CatBoost-Optuna model M is output. The CatBoost-Optuna model can be used to predict new samples.
[0044] Select the algorithm evaluation indicators as follows: The coefficient of determination (R²), root mean square error (RMSE) and mean absolute percentage error (MAPE, %) are selected as the analysis indicators of the model's predictive ability. When evaluating the model, R² is used to view the overall fit of the model, and RMSE and MAPE are used to evaluate the actual predictive performance of the model.
[0045] R² is an indicator for evaluating the goodness of fit of a model, reflecting the proportion of the independent variable's explanation of the total variation of the dependent variable. The closer it is to 1, the better the model fit.
[0046] ; in, represents the actual value of equilibrium temperature, represents the predicted equilibrium temperature, Represents the actual mean equilibrium temperature.
[0047] RMSE is an indicator that measures the difference between the model's predicted value and the true value. It represents the standard deviation between the predicted value and the true value, and is used to reflect the absolute amount of the model's prediction error.
[0048] ; MAPE is an indicator of the relative error between the predicted value and the true value, given as a percentage, reflecting the average prediction error rate of the model. The following table is used to calculate the metric.
[0049] ; Experimental example: The following is the experimental data set constructed according to step 1, and the results obtained through the above steps: 1) According to step 1, the experimental data set was constructed and data preprocessed to obtain a data set containing 788 samples. Each sample data includes information such as organic name, mole fraction concentration (Mol_fr), equilibrium pressure (Mpa), and equilibrium temperature (K).
[0050] 2) According to step 2, extract the molecular features and functional group features of organic matter and remove redundant features. The extracted molecular features of organic matter are Mw, NumAtom, NumHacc, NumHdon, NumHetero, NumRT, Logp, Sa, Psa, Apsa, Psa_div_sa, HeteroFr, and the extracted functional group features are Hydroxyl, Methyl, Alcohol, Ketone, Carboxyl, Amino, Ethyl, Epoxide, Propyl, Alkene, Imine, Ester. The original feature heat map is as follows: Figure 1 The feature heat map after removing redundant features is shown in Figure 2 shown.
[0051] 3) According to step 3, the extracted features are screened, and the features finally excluded are: Alcohol, Carboxyl, Amino, NumAtom, Aspa, NumHetero, Psa. The feature screening results of Boruta algorithm are as follows Figure 3 The feature correlation heat map after feature screening is shown in Figure 4 Shown 4) Build the CatBoost-Optuna model according to step 4, and input the preprocessed data into the model for training. Finally, predict the experimental effect through the test data set and analyze the experimental results.
[0052] 5. The experimental results were compared and analyzed among 6 models, including linear regression model (LR), support vector machine model (SVM), decision tree model (DT), random forest model (RF), XGBoost model, and CatBoost model. Finally, Optuna was used to optimize the hyperparameters of the models with better results. The prediction performance of each model in predicting the formation temperature of natural gas hydrate is shown in the figure below. Figure 5 , Figure 6 , Figure 7 The variable axis is the name of the machine learning model, including linear regression model (LR), support vector machine model (SVM), decision tree model (DT), random forest model (RF), XGBoost model, and CatBoost model; the comparison results of CatBoost optimized by Optuna and CatBoost without Optuna optimization are shown in the figure Figure 8As shown, in the variable axis, RMSE is the root mean square error, MAPE is the mean absolute percentage error, and R² is the coefficient of determination; the CatBoost-optuna model effect is as follows Fig. 9 , Fig.10 shown.
[0053] 6) Compare the CatBoost-Optuna model with the classical thermodynamic model (THERM). By comparing the predicted and actual values of the phase equilibrium temperature under experimental conditions, the MAPE, RMSE and R² of the two models are calculated. The results are shown in Table 4. The scatter plot of the THERM model and the CatBoost-Optuna model is shown in Fig.11 , Fig.12 The residual comparison of the THERM model and the CatBoost-Optuna model is shown in Fig.13 , Fig.14 The error histograms of the THERM model and the CatBoost-Optuna model are shown in Fig.15 , Fig.16 shown.
[0054] Table 4 MAPE, RMSE and R² of the two models;
[0055] 7) Design an extrapolation test experiment and select a set of solvents that are not included in the data set (but whose functional group composition has appeared in the training data). Keep the experimental pressure conditions consistent, predict the phase equilibrium temperature of the extrapolated solvent, and compare it with the actual value. The test results are shown in the figure Fig.17 , Fig.18 shown.
[0056] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system described in Example 1 or 2 are implemented.
[0057] Example 4 A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system described in Example 1 or 2.
[0058] Example 5 The intelligent prediction system for natural gas hydrate phase equilibrium in organic solution system includes: The experimental data set building module is configured to: create an experimental data set and perform data preprocessing; The feature extraction module is configured to: include: extracting organic molecular features; extracting functional group features; The feature selection module is configured to include: removing redundant features according to the correlation heat map; Boruta feature screening; The model building and training module is configured to: build and train the CatBoost-Optuna model; The prediction module is configured to: input the data to be predicted into the trained CatBoost-Optuna model after preprocessing, feature extraction and feature selection, and output the prediction result.
Claims
1. An intelligent prediction method for natural gas hydrate phase equilibrium in an organic solution system, characterized in that: include: Step 1: Experimental data set construction; including: creating experimental data set and performing data preprocessing; Step 2: Feature extraction; including: extracting organic molecular features; extracting functional group features; Step 3: Feature selection; including: removing redundant features based on correlation heat map; Boruta feature screening; Step 4: Build and train the CatBoost-Optuna model; Step 5: After preprocessing, feature extraction and feature selection, the data to be predicted is input into the trained CatBoost-Optuna model to output the prediction results.
2. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: Create an experimental dataset, including: Each sample data includes: organic matter name, mole fraction concentration, equilibrium pressure, equilibrium temperature. By changing the sample type and mole fraction concentration, measuring the equilibrium pressure and equilibrium temperature under different pressure conditions, and summarizing the data, an experimental data set is formed.
3. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: Data preprocessing, including: ① Eliminate outliers; ②Use the data normalization function to normalize the data; ③ Use random oversampling to find sparse samples through the quantile of the target value, concatenate the sparse samples with the original data, and generate an oversampled data set.
4. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: Extraction of molecular characteristics of organic matter; including: RDKit was used to extract and quantify the features of 34 water-soluble organic substances contained in the dataset.
5. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: Extract functional group features; including: The frequency of occurrence of each functional group in the statistical data set was counted, and the 12 most frequently occurring functional groups were obtained, including: Hydroxyl, Methyl, Alcohol, Ketone, Carboxyl, Amino, Ethyl, Epoxide, Propyl, Alkene, Imine and Ester.
6. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: Remove redundant features; including: Use the matplotlib library in Python to draw feature heat maps based on the extracted features; The strongly correlated features were excluded, and the final excluded features were: Alcohol, Carboxyl, Amino, NumAtom, Aspa, NumHetero, Psa; The final heat map is obtained after removing the strongly correlated features.
7. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: The Boruta feature selection algorithm is used for feature screening; including: The Boruta feature selection algorithm was used for feature screening, and the input features of the experiment were determined to be: molar fraction concentration of the compound, ratio of polar solvent accessible surface area to total solvent accessible surface area, total solvent accessible surface area, molecular weight, propyl, ethyl, logarithmic octanol partition coefficient for water, methyl, hydroxyl, equilibrium pressure, and ketone; the output feature was equilibrium temperature.
8. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: Build a CatBoost-Optuna model; including: The CatBoost-Optuna model consists of the CatBoost model and the Optuna framework; Suppose the sample data set D has n features and Y i is the label value of the target variable, randomly sort the samples in D and generate multiple sets of sequences, one of which is a random sequence σ=[σ1,σ2,…,σ n ], the CatBoost model replaces the random sequence with the following equation: ; in, X σp,q is the characteristic σ p No. q training values; X σj,q is the characteristic σ j No. q training values; Y σj is the partial target value of the corresponding feature; P is the prior value; β is the prior weight value; for regression tasks, the method for calculating the prior value is to take the average value of the data set.
9. The method for intelligent prediction of natural gas hydrate phase equilibrium in an organic solution system according to claim 1, characterized in that: The training process of the CatBoost-Optuna model specifically includes: (1) Input data: The input data is the number of iterations I and the characteristic matrix X of natural gas hydrate. i and the corresponding target value y i ; (2) Initialize the model: Before training the CatBoost-Optuna model, initialize the prediction values of all samples; (3) Optuna hyperparameter optimization: Use the Optuna optimizer to optimize the hyperparameters of the CatBoost-Optuna model; (4) Model training and iteration: ① Calculate the gradient: For each iteration iter=1,2,3,…,I, calculate the gradient g between the current CatBoost-Optuna model prediction value and the true target value i ; ② Update the predicted value: For each sample j, use the CatBoost-Optuna model to learn the current gradient g i and feature X j , and update the predicted value M j , and then the updated prediction value M j Accumulated to the predicted value M of the current sample i middle; ③ Iteration process: Repeat the above steps of calculating gradients and updating prediction values until the predetermined number of iterations I is completed; (5) Output model: After the CatBoost-Optuna model training and iteration are completed, the final CatBoost-Optuna model M is output.
10. An intelligent prediction system for natural gas hydrate phase equilibrium in an organic solution system, used to implement the intelligent prediction method for natural gas hydrate phase equilibrium in an organic solution system according to any one of claims 1 to 9, characterized in that: include: The experimental dataset building module,is configured as; Create experimental data sets and perform data preprocessing; The feature extraction module is configured to: include: extracting organic molecular features; extracting functional group features; The feature selection module is configured to include: removing redundant features according to the correlation heat map; Boruta feature screening; The model building and training module is configured to: build and train the CatBoost-Optuna model; The prediction module is configured to: input the data to be predicted into the trained CatBoost-Optuna model after preprocessing, feature extraction and feature selection, and output the prediction result.
Citation Information
Patent Citations
Method for predicting tribological performance of polymer composite material based on machine learning
CN118609729A