Method for predicting adsorption of modified magnesium oxide to carbon dioxide based on hybrid driving

Through a hybrid-driven method, computer simulation and machine learning technology are used to predict the adsorption performance of modified magnesium oxide, which solves the problems of high cost, low efficiency and poor generalization capabilities in the existing technology, and achieves efficient, economical and widely applicable adsorption performance prediction.

CN120148684APending Publication Date: 2025-06-13CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201256.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has problems of high cost, low efficiency and poor generalization ability when studying the adsorption properties of modified magnesium oxide on carbon dioxide, and it is difficult to fully cover the performance of all adsorption sites and doping structures.

Method used

Using a hybrid-driven method, data sets are constructed through computer simulation and machine learning technology and data augmentation and adsorption energy prediction using diffusion models and neural network models to achieve efficient, economical and widely applicable adsorption performance prediction of modified magnesium oxide.

Benefits of technology

It significantly reduces the research cost and time, improves the material screening speed and research efficiency, enhances the applicability to new materials and unknown structures, and can generate a large number of doped structures with potential application value in a short period of time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148684A_ABST
    Figure CN120148684A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting carbon dioxide adsorption by modified magnesium oxide based on hybrid driving, and relates to the technical field of carbon dioxide adsorption. According to the method, key attributes with high contribution to adsorption energy are accurately screened out through a clustering analysis and correlation analysis method, and the generation process of the diffusion model is controlled by utilizing the key attributes with high contribution, so that the diversity and coverage range of samples are remarkably improved by the generated new data, and the accuracy of the diffusion model is improved. And potential efficient material structures which are difficult to cover by traditional experiments or model methods can be excavated. And fitting prediction is performed on the extended data in combination with a neural network model, so that a large number of expensive experiments and time-consuming theoretical calculation are avoided. Compared with a method depending on hardware equipment (such as a laboratory mass spectrometer or adsorption test equipment) in the market, the method is more economical and efficient, the resource demand is remarkably reduced, and a low-cost solution is provided for material research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of carbon dioxide adsorption, and particularly to a prediction method for carbon dioxide adsorption by modified magnesium oxide based on hybrid drive. Background Art

[0002] Magnesium oxide is a solid adsorbent for carbon dioxide that has attracted much attention in the scientific research and industrial circles and has market prospects;

[0003] The research on the adsorption mechanism of carbon dioxide is the key to the development and efficient utilization of this adsorbent. However, the research on the mechanism of modified magnesium oxide and carbon dioxide adsorption is still not clear and lacking. At present, the technologies for studying the carbon dioxide gas adsorption capacity of adsorption materials are mainly divided into two categories: experimental measurement and theoretical calculation.

[0004] 1. Experimental Measurement:

[0005] Experimental measurement is a traditional and direct technical method for studying the performance of adsorption materials. In this process, researchers usually synthesize candidate materials under specific conditions and directly evaluate the adsorption capacity of the materials by measuring the change in the concentration of carbon dioxide gas in the materials before and after adsorption.

[0006] The advantage of this method is that real experimental data can be obtained, providing an intuitive verification of the material performance. However, its disadvantages are also relatively obvious. First, the experiment requires precise control of reaction conditions, including parameters such as temperature, pressure, and pH value. The operation is complex and the requirements for experimental equipment are high. Second, the experimental period is long. The entire process from material synthesis to performance evaluation may take several days or even weeks. In addition, the preparation of experimental materials, the use of testing instruments, and the consumption of reagents also make the cost high, especially for research that needs to test a large number of candidate materials, the economic and time costs are particularly significant.

[0007] 2. Theoretical Calculation:

[0008] In order to overcome the problems of high cost and high difficulty in experimental measurement, more and more researchers have begun to use computer simulation calculation software such as Material studio and Vasp as important tools for assisting in the development of adsorption materials. Theoretical calculation can predict the adsorption performance of materials without actual experiments by simulating the interactions at the molecular level. Common methods include first-principles calculation, density functional theory (DFT), and molecular dynamics simulation, etc.

[0009] The advantages of theoretical calculations lie in high efficiency and low cost. For example, by calculating the adsorption energy between candidate materials and carbon dioxide gas molecules, materials with excellent performance can be quickly screened out, significantly reducing the trial-and-error cost of experiments. In addition, theoretical calculations can also provide microscopic mechanism information that is difficult to directly obtain through experiments, such as changes in the electron distribution of adsorption sites, the formation and breakage of chemical bonds, etc., thus providing more in-depth guidance for material design.

[0010] However, theoretical calculations also have their limitations. First of all, accurately simulating the adsorption process requires a large amount of computing resources. Especially for complex systems, the computing time and hardware requirements may be very high. Therefore, most theoretical calculations actually only consider some of the main adsorption sites for research. Secondly, the results of theoretical calculations depend on the models and assumptions adopted, and their accuracy depends to a certain extent on the degree of matching with the actual situation. Therefore, theoretical calculations usually need to be combined with experimental measurements to verify their reliability and improve the credibility of the results.

[0011] Even when combined, existing technologies face a problem of lacking generalization ability. Limited by function and time, researchers often can only select some representative modified materials for research, and cannot cover materials with other doping atoms and modification methods.

[0012] In summary, the objective disadvantages of existing technologies are as follows:

[0013] Complicated, time-consuming and low in accuracy. Complicated, time-consuming and low in accuracy. Existing technologies require expensive chemical material costs when using experiments to prepare adsorption materials and measure adsorption amounts. Moreover, the cumbersome steps of experimental preparation, synthesis, measurement, etc. usually require precise instruments such as XPS for operation, and the test cycle is long.

[0014] The experimental results have a small coverage area and cannot cover the adsorption of all adsorption sites. In laboratory experiments, it is somewhat difficult to accurately determine the specific sites where carbon dioxide gas is adsorbed on the adsorption material. And theoretical experimental research usually focuses on the optimal adsorption sites, lacking a comprehensive understanding of the overall adsorption performance of the adsorption material. And calculating the adsorption energy of any adsorption site through theoretical research usually requires a large number of first-principles (DFT) calculations, which is not only time-consuming but also costly.

[0015] Poor generalization ability. The generalization ability of existing technologies is weak. Usually, it can only conduct research and design ideal adsorption materials for magnesium oxide adsorption materials doped with specific atoms, and the applicability to magnesium oxide adsorption materials doped with other atoms is limited.

[0016] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0017] The object of the present invention is to provide a prediction method for carbon dioxide adsorption by modified magnesium oxide based on hybrid drive, aiming to overcome the deficiencies in the prior art such as high cost, low efficiency, and poor generalization ability. By integrating the advantages of data-driven and model-driven, it provides an efficient, economical and widely applicable solution for predicting carbon dioxide adsorption performance to solve the technical problems raised in the background art.

[0018] To achieve the above object, the present invention provides the following technical solutions: A prediction method for carbon dioxide adsorption by modified magnesium oxide based on hybrid drive, at least including the following steps:

[0019] S1: Optimize the structures of various modified magnesium oxides through computer simulation software to ensure their stability and reliability as materials;

[0020] S2: Consider various possible adsorption sites, place carbon dioxide gas molecules on the adsorption sites, and calculate the adsorption energy and electrical properties of the corresponding structures, etc., to construct a data set for machine learning;

[0021] S3: Considering the complexity of the data set, use clustering to divide the data set into multiple clusters for research;

[0022] S4: Construct a correlation analysis model for any machine learning, whose main function is to screen out the attributes with high correlation to the adsorption energy from the data set;

[0023] S5: Judge whether the existing data set needs to be expanded;

[0024] S6: Apply the attributes highly correlated with the adsorption energy to the machine learning model for fitting and predicting the adsorption energy, thereby constructing a prediction model for the adsorption energy;

[0025] S7: Learn the existing data set through a diffusion model, which can expand the data set and increase the diversity of data. Moreover, the diffusion model supplements more diverse training data for the artificial intelligence system by generating more random adsorption situations and carbon dioxide adsorption results under different conditions;

[0026] S8: Use a neural network model to fit and predict the adsorption energy of the expanded data set, so that the relationship between "adsorption energy - structure" is mapped;

[0027] S9: Statistically analyze the statistical data of all trained adsorption energy prediction models to evaluate the differences in the carbon dioxide gas adsorption capabilities of different modified magnesium oxides, and construct a mapping relationship between "adsorption energy - structure".

[0028] Furthermore, the S1 at least includes the following steps:

[0029] Dope simulation of magnesium oxide (MgO) is carried out using Material Studio computer simulation software (or other simulation software);

[0030] During the simulation process, 17 elements from six different groups in the periodic table are selected;

[0031] These 17 elements will be doped into magnesium oxide at doping concentrations of 1%, 2% and 3% respectively to obtain the simulation results after doping;

[0032] The six different groups include alkali metals, alkaline earth metals, transition metals, post-transition metals, metal compounds and non-metals;

[0033] The alkali metals at least include Li, Na and K, the transition metals at least include Fe, Co, Ir and Pt, the post-transition metals at least include Al and Ga, the metal compounds at least include Si and Ge, and the non-metals at least include N, P, S and Cl.

[0034] Further, the S2 at least includes the following steps:

[0035] For each magnesium oxide system doped with an element, there are structures with three doping concentrations and a total of 12 adsorption sites;

[0036] A total of 204 systems are established through 17 elements;

[0037] Subsequently, after screening and excluding the systems without stable structures, a total of 141 systems are obtained;

[0038] The screening and exclusion method is to analyze properties, and the properties at least include energy band, density of states, differential charge density and charge population;

[0039] Each system is named with E-C-A, that is, Element-Concentration-Adsorption sites, where E represents the doped atom, C represents the doping concentration, and A represents the adsorption site;

[0040] These 141 systems correspond to a dataset of 68 various types of properties, and these properties can be divided into the following four major categories: free state of the doped atom, system before carbon dioxide adsorption, system after carbon dioxide adsorption, and property differences of the system before and after carbon dioxide adsorption.

[0041] Further, the dataset in the S3 contains data from 17 different groups of elements at different doping concentrations and adsorption sites, and the situation is relatively complex;

[0042] Although the selected elements represent six main element families in the periodic table, the order of the elements is not determined by the adsorption energy. Therefore, it is more reliable to use cluster analysis to divide the dataset into different clusters and perform correlation analysis on each cluster than to directly study by family;

[0043] For complex attribute data, the spectral clustering method is used to divide the data into multiple clusters, and the data within each cluster has similar adsorption behavior characteristics;

[0044] Then, Spearman correlation analysis is used to screen out the high - contribution attributes that have a significant impact on the adsorption energy, which provides an important reference for subsequent expansion of the diffusion model dataset.

[0045] Furthermore, since there is a certain probability that the adsorption energy contains outliers, and the electrical properties of compounds do not conform to the normal distribution but are fixed properties determined by structure and composition, Spearman correlation analysis was selected for the correlation analysis in S4 instead of Pearson correlation analysis. The Spearman correlation does not depend on the distribution assumption of the data and is more suitable for dealing with non - normally distributed data.

[0046] Furthermore, S6 at least includes the following steps:

[0047] Use multiple fitting algorithms to perform regression prediction on the dataset. The fitting algorithms include but are not limited to decision trees, gradients, Lasso regression, and MLP;

[0048] The objects of algorithm fitting are not only the simple entire dataset, but also each data in the dataset marked with the doping element family, doping element period, and cluster label after clustering;

[0049] Further combine the four major attribute categories in S2;

[0050] In this way, the performance of different algorithms under specific conditions is studied separately.

[0051] Furthermore, the design of the diffusion model in S7 at least includes the following steps:

[0052] Based on the analysis results of S4, screen out the key attributes that contribute highly to the adsorption energy. The key attributes include but are not limited to the number of electrons in each orbital of the doping atom and the change in charge distribution before and after carbon dioxide adsorption;

[0053] Combine these attributes with the corresponding structure data (based on one - hot encoding) to form the input of the diffusion model, and introduce the diffusion model to learn these data. The diffusion model generates new structures through a process of gradually adding noise and removing noise, and ensures that the newly generated data is consistent with the real data in terms of physical and chemical properties;

[0054] To adapt to the magnesium oxide structure, a voxel grid diffusion model is introduced. The voxel grid diffusion model cuts the entire structure into multiple grids through grids, introduces noise into the grids, and each grid is 0.25 angstroms, which ensures that it will maintain a hierarchical structure regardless of how noise is added or removed, thus better adapting to the magnesium oxide molecular structure to expand the dataset;

[0055] The process of the diffusion model is divided into noise addition and denoising. First, a denoising neural network is trained during the noise addition process, and this denoising neural network learns the mapping from the noise molecular distribution to the real molecular distribution;

[0056] The process of noise addition at least includes the following steps:

[0057] Noise addition: Gaussian noise is added to each voxel grid to generate the noise sample at the next moment;

[0058] Denoising neural network: The goal of this neural network is to recover the original clean molecule from the noise sample during the denoising process. Therefore, the denoising neural network should start training during the noise addition process, and the training goal is to minimize the mean square error (MSE) between the voxel grid at time t+1, that is, the voxel grid after noise addition, and the voxel grid at time t after denoising;

[0059] The denoising process is different from the traditional diffusion model. After noise addition, in fact, the entire voxel grid has become smooth. Here, the denoising method of the VoxMol voxel grid diffusion model is introduced, and the core steps of the denoising method of the VoxMol voxel grid diffusion model include walking and jumping.

[0060] The goal of the walking step is to sample the noise molecule yk from the smooth noise distribution p(y) through the MCMC (Markov chain Monte Carlo) method, and gradually reverse through MCMC in the smooth noise distribution to generate a noise molecule that conforms to the distribution. Because the noise distribution p(y) is the result of the convolution of the original molecular distribution p(x) and Gaussian noise (smoother and easier to sample), and by iteratively updating the noise molecule yk, the walking step can better approximate the distribution characteristics of p(y) and provide high-quality initial noise samples for subsequent denoising. The formula is as follows:

[0061]

[0062] where, is the score function of the smooth distribution (indirectly learned through the denoising network), δ is the step size, and ∈k is Gaussian noise;

[0063] The goal of the walking step is to map the noisy molecule yk to the true molecule distribution p(x) through single-step denoising. Using the denoising network that has been trained during the previous noise addition process, a clean molecule xk is directly generated from the noisy molecule yk in one step. This step skips the multi-step iterative denoising process of traditional diffusion models. The formula is as follows:

[0064] xk = yk + σ2gθ(yk)

[0065] Where gθ(yk) is the score function learned by the denoising network, and σ2 is the fixed noise level;

[0066] In summary, it can be considered that the walking step is responsible for efficiently sampling in the smooth noise distribution p(y) to provide "rough" noise samples (because p(y) is easier to sample), while the jumping step is responsible for mapping the noise samples to the true molecule distribution p(x) in one step and refining the true molecule distribution p(x) through a neural network;

[0067] This enables the VoxMol voxel format diffusion model to avoid multi-step iterations (e.g., 1000 steps) of other diffusion models and significantly accelerates generation.

[0068] Meanwhile, in addition, various strategies can also be adopted during the training process to optimize the generated data, including introducing physical constraints (such as stability and rationality checks) and property constraints (such as bias adjustment of high-contribution properties), thereby expanding the diversity of the original dataset.

[0069] Furthermore, the S8 at least includes the following steps:

[0070] Use a neural network model to fit and predict the adsorption energy for the dataset expanded in S7;

[0071] According to the scale of the dataset, design a multi-layer graph neural network; the input layer includes all the attributes of the expanded dataset; the number of layers and the number of neurons in each layer of the hidden layer are determined by cross-validation, and the activation function uses ReLU or other non-linear functions; the output layer is a single node for predicting the adsorption energy; use stratified cross-validation to divide the dataset into a training set, a validation set, and a test set to ensure that each stratification evenly covers different categories of structures and adsorption energy values;

[0072] Select a suitable optimizer and loss function for training and adjust the learning rate to avoid overfitting or underfitting;

[0073] In addition, use an early stopping strategy to monitor the performance of the validation set and stop training before the model overfits;

[0074] During the training process, a multi-layer perceptron (MLP) is introduced, and the high-contribution attributes are preferentially learned in combination with the results of feature selection. Residual connections are introduced to facilitate increasing the number of layers of the graph neural network to extract data features. After optimizing the model hyperparameters and obtaining good results on the public dataset, it is applied to the augmented dataset to ensure that the model can predict the adsorption energy with high accuracy.

[0075] Further, the execution process of the neural network model is as follows:

[0076] First is the input layer, which assigns initial feature vectors to each node and edge;

[0077] Then, information is passed through the graph convolutional layer. In each layer, the feature vectors of the nodes are aggregated and updated by the feature vectors of the neighbor nodes to capture local structural information;

[0078] The methods of the aggregation at least include summation, averaging, and max pooling;

[0079] To ensure the equivariance of the model to graph isomorphism, a symmetric aggregation function can be adopted, which can avoid the sensitivity of the model to the node order and prevent isomorphism non-invariance;

[0080] After multi-layer graph convolution, a global pooling layer is used to aggregate the node features into a fixed-length global feature vector, representing the features of the entire crystal structure;

[0081] Finally, the global feature vector is input into the fully connected layer to perform regression prediction of the adsorption energy of magnesium oxide. At the same time, in addition to the adsorption energy, according to the task requirements, the prediction results of other attributes are output to achieve the result of the study and discussion on the adsorption mechanism of magnesium oxide.

[0082] Further, when judging whether the dataset needs to be augmented in S5, a systematic evaluation is carried out based on multi-dimensional criteria;

[0083] The systematic evaluation based on multi-dimensional criteria at least includes plotting the learning curve of the machine learning model, adopting statistical power analysis, and the feature space coverage of the dataset;

[0084] Plot the learning curve (Learning Curve) of the machine learning model. If the training error and the validation error do not reach the convergence platform, it indicates that increasing the data volume may improve the performance;

[0085] Adopt statistical power analysis (Statistical Power Analysis) to calculate the confidence interval coverage ability of the current sample size for the target prediction accuracy;

[0086] Feature space coverage of the dataset: Visualize the high-dimensional feature distribution through t-SNE, identify the sparse regions in the feature space, and supplement the uncovered adsorption site configurations (such as edge sites, defect sites, etc.). However, considering that some compounds with certain features may not exist in reality, or the stable structures are very likely to be distributed in the same threshold space, this solution is only an alternative. The main judgment method still depends on whether the training error and validation error mentioned above reach convergence, and the confidence interval coverage ability of the current sample size for the target prediction accuracy to determine whether to supplement the dataset.

[0087] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0088] 1. Low cost: Compared with the traditional method for predicting carbon dioxide adsorption performance that relies on expensive hardware equipment, the present invention is based on a data and model hybrid-driven strategy, which can be completed only through computer simulation and common machine learning algorithms. The whole process does not require high-cost experimental consumables and large-scale equipment support, significantly reducing the economic cost of research and development, and at the same time improving the utilization rate of laboratory resources;

[0089] 2. Enhanced exploration potential and improved efficiency: Traditional adsorption performance research usually requires complex experimental measurements and physical model derivations, which are time-consuming and limited. The present invention uses machine learning to screen highly correlated attributes, expand the dataset with a diffusion model, and fit the adsorption energy with a neural network model, which not only speeds up the material screening process but also generates a large number of doped structures with potential application value in a short time, quickly predicting their adsorption performance. It is especially suitable for evaluating doped structures that have not been fully studied and has broad application potential in material design, greatly improving the research efficiency;

[0090] 3. Enhanced generalization ability: Compared with the existing technologies based on single data-driven or physical model-driven approaches, the present invention combines the high efficiency of data-driven and the reliability of physical model-driven through a hybrid-driven strategy, greatly enhancing the applicability to new materials and unknown structures. Especially after introducing the diffusion model, the data diversity and the generalization performance of the model are further improved;

[0091] 4. Through clustering analysis and correlation analysis methods, the present invention accurately screens out the key attributes with high contribution to the adsorption energy, and uses these key attributes with high contribution to control the generation process of the diffusion model. This enables the generated new data to significantly improve the diversity and coverage of the samples, and potential high-efficiency material structures that are difficult to cover by traditional experimental or model methods can be mined. Combining with the neural network model to fit and predict the extended data, in this way, a large number of expensive experiments and time-consuming theoretical calculations are avoided. Compared with the methods on the market that rely on hardware devices (such as laboratory mass spectrometers or adsorption test devices), the present invention is more economical and efficient, and the resource requirements are significantly reduced, providing a low-cost solution for material research. All in all, the present invention realizes the prediction of carbon dioxide adsorption performance with low cost, high efficiency and high precision, providing technical support and theoretical basis for the further development of high-efficiency carbon capture materials. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0093] Figure 1 is the flowchart of the present invention;

[0094] Figure 2 is the explanatory diagram of the voxel grid diffusion model; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0095] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.

[0096] Please refer to Figure 1 , the method for predicting carbon dioxide adsorption by modified magnesium oxide based on hybrid drive includes at least the following steps:

[0097] S1: Optimize the structures of various modified magnesium oxides through computer simulation software to ensure their stability and reliability as materials;

[0098] S2: Consider various possible adsorption sites, place carbon dioxide gas molecules on the adsorption sites, and calculate the adsorption energy and electrical properties of the corresponding structures, etc., to construct a data set for machine learning;

[0099] S3: Considering the complexity of the data set, use clustering to divide the data set into multiple clusters for research;

[0100] S4: Construct a correlation analysis model for any machine learning, whose main function is to screen out the attributes with high correlation with the adsorption energy from the dataset;

[0101] S5: Determine whether the existing dataset needs to be expanded;

[0102] S6: Apply the attributes highly correlated with the adsorption energy to the machine learning model for fitting and predicting the adsorption energy, thereby constructing a prediction model for the adsorption energy;

[0103] S7: Learn the existing dataset through the diffusion model, which can expand the dataset and increase the diversity of data. Moreover, the diffusion model supplements more diverse training data for the artificial intelligence system by generating more random adsorption situations and carbon dioxide adsorption results under different conditions;

[0104] S8: Use the neural network model to fit and predict the adsorption energy of the expanded dataset, so as to map the relationship between "adsorption energy - structure";

[0105] S9: Statistically analyze the statistical data of all trained adsorption energy prediction models to evaluate the differences in the adsorption capacity of different modified magnesium oxides for carbon dioxide gas, and construct the mapping relationship of "adsorption energy - structure";

[0106] Summarize the model fitting results of S8 through common model evaluation methods such as R2, RMSE, MSE, and MAE. Obtain rules such as: The structures of magnesium oxides doped with other non-metals often have stronger adsorption energy. Obtain rules such as: The adsorption energy of most modified magnesium oxides is affected by the electrical properties of the structure (such as the PDOS peaks of doped atoms and the changes in the PDOS peaks of C).

[0107] S1 includes at least the following steps:

[0108] Use the Material Studio computer simulation software (or other simulation software) to perform doping simulation on magnesium oxide (MgO);

[0109] Select 17 elements from six different groups in the periodic table during the simulation process;

[0110] These 17 elements will be doped into magnesium oxide at doping concentrations of 1%, 2%, and 3% respectively to obtain the simulation results after doping;

[0111] The six different groups include alkali metals, alkaline earth metals, transition metals, post-transition metals, metal compounds, and non-metals;

[0112] The alkali metals include at least Li, Na, and K, the transition metals include at least Fe, Co, Ir, and Pt, the post-transition metals include at least Al and Ga, the metal compounds include at least Si and Ge, and the non-metals include at least N, P, S, and Cl.

[0113] S2 includes at least the following steps:

[0114] For each magnesium oxide system doped with an element, there is a structure with three doping concentrations and a total of 12 adsorption sites;

[0115] A total of 204 systems were established with 17 elements;

[0116] Subsequently, after screening and excluding the systems without stable structures, a total of 141 systems were obtained;

[0117] The screening and exclusion method is to analyze properties, which at least include energy band, density of states, differential charge density, and charge population;

[0118] Each system is named with E-C-A, that is, Element-Concentration-Adsorption sites, where E represents the doped atom, C represents the doping concentration, and A represents the adsorption site. For example, for a magnesium oxide system with a doping concentration of 1, the doped atom is aluminum, and the carbon dioxide adsorption site is on the oxygen atom of magnesium oxide, its name is: EAl-C1-AO;

[0119] These 141 systems correspond to a dataset of 68 various properties, which can be divided into the following four major categories: free state of the doped atom, system before carbon dioxide adsorption, system after carbon dioxide adsorption, and property differences of the system before and after carbon dioxide adsorption.

[0120] The dataset in S3 contains data from 17 different groups of elements at different doping concentrations and adsorption sites, and the situation is relatively complex;

[0121] Although the selected elements represent 6 major element groups in the periodic table, the arrangement order of the elements is not determined by the adsorption energy. Therefore, using cluster analysis to divide the dataset into different clusters and performing correlation analysis on each cluster is more reliable than directly studying by group;

[0122] For the complex property data, the spectral clustering method is used to divide the data into multiple clusters, and the data within each cluster has similar adsorption behavior characteristics;

[0123] Then, Spearman correlation analysis is used to screen out the high-contribution properties that have a significant impact on the adsorption energy. For example, some clusters show a strong influence on the property of "charge transfer in the system before and after adsorption", which provides an important reference for subsequent expansion of the diffusion model dataset.

[0124] Since there is a certain probability that the adsorption energy contains outliers, and the electrical properties of compounds do not conform to the normal distribution but are fixed properties determined by the structure and composition, Spearman correlation analysis was selected instead of Pearson correlation analysis for the correlation analysis in S4. Spearman correlation does not depend on the distribution assumption of the data and is more suitable for dealing with non-normally distributed data.

[0125] When judging whether the data set needs to be expanded in S5, a systematic evaluation is carried out based on multi-dimensional criteria;

[0126] The systematic evaluation based on multi-dimensional criteria at least includes plotting the learning curve of the machine learning model, performing statistical power analysis, and the feature space coverage of the data set;

[0127] Plot the learning curve of the machine learning model (Learning Curve). If the training error and validation error do not reach the convergence platform, it indicates that increasing the data volume may improve the performance;

[0128] Perform statistical power analysis (Statistical Power Analysis) to calculate the confidence interval coverage ability of the current sample size for the target prediction accuracy;

[0129] The feature space coverage of the data set: Visualize the high-dimensional feature distribution through t-SNE, identify the sparse regions in the feature space, and supplement the uncovered adsorption site configurations (such as edge sites, defect sites, etc.). However, considering that some compounds with certain features may not exist in reality, or stable structures are very likely to be distributed in the same threshold space, so this scheme is only used as an alternative. The main judgment method still depends on whether the training error and validation error mentioned above reach convergence, and the confidence interval coverage ability of the current sample size for the target prediction accuracy to determine whether to supplement the data set.

[0130] S6 at least includes the following steps:

[0131] Use multiple fitting algorithms to perform regression prediction on the data set. The fitting algorithms include but are not limited to decision trees, gradients, Lasso regression, and MLP;

[0132] The objects of algorithm fitting are not only the simple entire data set, but also each data in the data set is labeled with the doping element family, doping element period, and the cluster label after clustering;

[0133] Further combine the four major attribute categories in S2;

[0134] In this way, the performance of different algorithms under specific conditions is studied separately. For example, when the dataset is divided by doping element families and the major attribute category is "the system after carbon dioxide adsorption", it means that the current discussion is about the comparison of the adsorption energy fitting prediction effects of each algorithm under the attributes included in the major attribute category of "the system after carbon dioxide adsorption" between families.

[0135] The design of the diffusion model in S7 includes at least the following steps:

[0136] Based on the analysis results of S4, key attributes with relatively high contributions to the adsorption energy are screened out. The key attributes include but are not limited to the number of electrons in each orbital of the doping atom and the change in charge distribution before and after carbon dioxide adsorption;

[0137] These attributes are combined with the corresponding structural data (based on one-hot encoding) to form the input of the diffusion model. The diffusion model is introduced to learn these data. The diffusion model generates new structures through a process of gradually adding noise and denoising, and ensures that the newly generated data is consistent with the real data in terms of physical and chemical properties;

[0138] To adapt to the magnesium oxide structure, a voxel grid - based diffusion model is introduced. As Figure 2 shown, in traditional molecular structure diffusion models, most tasks are designed based on small molecule structures. For small molecule structures composed of benzene ring hydroxyl groups, there is no multi - layer structure like magnesium oxide. Therefore, the traditional molecular structure diffusion model can achieve the goal by simply introducing noise to atoms and then denoising. In order to retain the characteristics of the multi - layer structure, the voxel grid - based diffusion model cuts the entire structure into multiple grids through grids and introduces noise into the grids. As Figure 2 shown by the red frame in, each grid is 0.25 angstroms, which ensures that it will maintain the hierarchical structure regardless of how noise is added or removed, thus being more adaptable to the magnesium oxide molecular structure to expand the dataset;

[0139] The process of the diffusion model is divided into adding noise and denoising. First, a denoising neural network is trained during the noise - adding process. This denoising neural network learns the mapping from the noise molecular distribution to the real molecular distribution;

[0140] The noise - adding process includes at least the following steps:

[0141] Noise addition: Gaussian noise is added to each voxel grid to generate the noise sample at the next moment;

[0142] Denoising neural network: The goal of this neural network is to recover the original clean molecule from the noise sample during the denoising process. Therefore, the denoising neural network starts to be trained during the noise - adding process. The training goal is to minimize the mean square error (MSE) between the voxel grid at time t + 1, that is, the voxel grid after adding noise, and the voxel grid at time t after denoising.

[0143] The denoising process is different from traditional diffusion models. After adding noise, in fact, the entire voxel grid has become smoother. Here, the denoising method of the VoxMol voxel grid diffusion model is adopted. The core steps of the denoising method of the VoxMol voxel grid diffusion model include walking and jumping.

[0144] The goal of the walking step is to sample noise molecules yk from the smooth noise distribution p(y) through the MCMC (Markov Chain Monte Carlo) method. By gradually reversing through MCMC in the smooth noise distribution, noise molecules that conform to the distribution are generated. Since the noise distribution p(y) is the result of the convolution of the original molecular distribution p(x) and Gaussian noise (smoother and easier to sample), and by iteratively updating the noise molecules yk, the walking step can better approximate the distribution characteristics of p(y) and provide high-quality initial noise samples for subsequent denoising. The formula is as follows:

[0145]

[0146] where, is the score function of the smooth distribution (indirectly learned through the denoising network), δ is the step size, and ∈k is Gaussian noise;

[0147] The goal of the walking step is to map the noise molecule yk to the true molecular distribution p(x) through single-step denoising. Through the denoising network that has been trained during the previous noise addition process, a clean molecule xk can be directly generated from the noise molecule yk in one step. This step skips the multi-step iterative denoising process of traditional diffusion models. The formula is as follows:

[0148] xk = yk + σ2gθ(yk)

[0149] where, gθ(yk) is the score function learned by the denoising network, and σ2 is the fixed noise level;

[0150] In summary, it can be considered that the walking step is responsible for efficiently sampling in the smooth noise distribution p(y) to provide "rough" noise samples (because p(y) is easier to sample), while the jumping step is responsible for mapping the noise samples to the true molecular distribution p(x) in one step and refining the true molecular distribution p(x) through the neural network;

[0151] This enables the VoxMol voxel grid diffusion model to avoid the multi-step iteration (e.g., 1000 steps) of other diffusion models and significantly accelerate the generation.

[0152] At the same time, in addition, various strategies can also be adopted to optimize the generated data during the training process, including introducing physical constraints (such as stability and rationality tests) and property constraints (such as bias adjustment of high-contribution properties), thereby expanding the diversity of the original dataset.

[0153] S8 at least includes the following steps:

[0154] Use a neural network model to fit and predict the adsorption energy for the dataset expanded by S7;

[0155] Design a multi-layer graph neural network according to the scale of the dataset; the input layer includes all attributes of the expanded dataset; the number of hidden layers and the number of neurons in each layer are determined by cross-validation, and the activation function uses ReLU or other non-linear functions; the output layer is a single node for predicting the adsorption energy; use stratified cross-validation to divide the dataset into a training set, a validation set, and a test set to ensure that each stratification evenly covers different categories of structures and adsorption energy values;

[0156] Select a suitable optimizer and loss function for training, and adjust the learning rate to avoid overfitting or underfitting;

[0157] In addition, use an early stopping strategy to monitor the performance of the validation set and stop training before the model overfits;

[0158] Introduce a multi-layer perceptron (MLP) during the training process, and combine the results of feature selection to preferentially learn high-contribution attributes. Introduce residual connections to facilitate increasing the number of layers of the graph neural network to extract data features. After optimizing the model hyperparameters (such as learning rate, number of layers, etc.) and obtaining good results on the public dataset, apply it to the expanded dataset to ensure that the model can predict the adsorption energy with high accuracy.

[0159] The execution process of the neural network model is as follows:

[0160] First is the input layer, which assigns an initial feature vector to each node and edge, usually using one-hot encoding or embedding representation of atomic types;

[0161] Then, information is passed through graph convolutional layers. Usually, 2 to 3 layers of graph convolutional layers are sufficient to capture local structure information. In each layer, the feature vector of the node is aggregated and updated with the feature vectors of neighboring nodes to capture local structure information;

[0162] The aggregation methods at least include summation, averaging, and max pooling;

[0163] To ensure the equivariance of the model to graph isomorphism, symmetric aggregation functions such as summation or averaging can be used, which can avoid the sensitivity of the model to the node order and prevent isomorphism differences;

[0164] After multi-layer graph convolution, use a global pooling layer (such as global average pooling or global max pooling) to aggregate the node features into a fixed-length global feature vector representing the features of the entire crystal structure;

[0165] Finally, the global feature vector is input into the fully connected layer to perform regression prediction on the adsorption energy of magnesium oxide. At the same time, in addition to the adsorption energy, according to the task requirements, the prediction results of other properties are output. For example, by classifying other properties, the class probability distribution can be output to achieve the result of exploring the adsorption mechanism of magnesium oxide.

[0166] The specific process of dataset construction is proposed:

[0167] First, construct modified structures of various dopings of magnesium oxide using computer simulation software (Material Studio) and optimize them to ensure structural stability. Subsequently, place carbon dioxide molecules at multiple possible adsorption sites, calculate the adsorption energy of each structure and its corresponding electrical properties, such as the number of orbital electrons, average net charge, etc. By eliminating unstable structures and invalid data, a large-scale dataset containing 141 systems and 68 properties is finally constructed [the dataset contains the following properties: "adsorption energy"; "atomic number"; "element symbol"; "atomic mass amu"; "atomic radius pm"; "atomic covalent radius pm"; "number of s orbital electrons of the doped atom in the free state"; "number of p orbital electrons of the doped atom in the free state"; "number of d orbital electrons of the doped atom in the free state"; "number of f orbital electrons of the doped atom in the free state"; "number of valence electrons of the doped atom in the free state"; "first ionization energy of the doped atom in the free state"; "oxide formation enthalpy of the doped atom in the free state"; "electronegativity of the doped atom in the free state"; "adsorption site of carbon dioxide"; "number of doped atoms"; "average net charge of O atoms around the doped atom before adsorption"; "average net charge between the doped atom and surrounding O atoms before adsorption"; "number of s orbital electrons of the doped atom before adsorption"; "number of p orbital electrons of the doped atom before adsorption"; "number of d orbital electrons of the doped atom before adsorption"; "number of f orbital electrons of the doped atom before adsorption"; "number of valence electrons of the doped atom before adsorption"; "average Bondorder of the Mg atom bonds around the doped atom before adsorption"; "average bond length between the Mg atoms around the doped atom and the atom before adsorption"; "average Bondorder of the O atom bonds around the doped atom before adsorption"; "average bond length between the O atoms around the doped atom and the atom before adsorption"; "band gap of the doped substrate before adsorption"; "net charge of CO2_C atom under the system after adsorption"; "net charge of the doped atom in the substrate after adsorption"; "number of s orbital electrons of the doped atom after adsorption"; "number of p orbital electrons of the doped atom after adsorption"; "number of d orbital electrons of the doped atom after adsorption"; "number of f orbital electrons of the doped atom after adsorption"; "number of valence electrons of the doped atom after adsorption"; "average Bondorder of the Mg atom bonds around the doped atom after adsorption"; "average bond length between the Mg atoms around the doped atom and the atom after adsorption"; "average Bondorder of the O atom bonds around the doped atom after adsorption"; "average bond length between the O atoms around the doped atom and the atom after adsorption"; "Bondorder between the doped atom and C after adsorption"; "bond length between the doped atom and C after adsorption"; "band gap of the doped substrate after adsorption"; "HOMO"; "LUMO"; "electron gain and loss situation of carbon dioxide between carbon dioxide and the substrate in the charge difference density map calculated by sets"]"In the charge difference density map calculated by sets, the part where the electron gain and loss between carbon dioxide and the substrate is the strongest on the substrate surface facing carbon dioxide"; "In the charge difference density map calculated by single atoms, the electron gain and loss of carbon dioxide between carbon dioxide and the substrate"; "In the charge difference density map calculated by single atoms, the electron gain and loss of the doped atom between carbon dioxide and the substrate"; "In the charge difference density map calculated by single atoms, the electron gain and loss directly below carbon dioxide between carbon dioxide and the substrate"; "In the charge difference density map calculated by single atoms, the electron gain and loss of the doped atom in the entire substrate before adsorption"; "The difference between HOMO and LUMO"; "The difference in the band gap of the doped substrate before and after adsorption"; "The difference in electron gain and loss between carbon dioxide and the substrate in the charge difference density map calculated by single atoms"; "The difference in electron gain and loss between carbon dioxide and the substrate in the charge difference density map calculated by sets"; "The difference in the number of s orbital electrons of the doped atom before and after doping"; "The difference in the number of p orbital electrons of the doped atom before and after doping"; "The difference in the number of d orbital electrons of the doped atom before and after doping"; "The difference in the number of f orbital electrons of the doped atom before and after doping"; "The difference in the number of valence electrons of the doped atom before and after doping"; "The difference in the number of s orbital electrons of the doped atom before and after adsorption"; "The difference in the number of p orbital electrons of the doped atom before and after adsorption"; "The difference in the number of d orbital electrons of the doped atom before and after adsorption"; "The difference in the number of f orbital electrons of the doped atom before and after adsorption"; "The difference in the number of valence electrons of the doped atom before and after adsorption"; "The average difference in Bondorder of the surrounding Mg atom bonds"; "The average difference in bond length between the surrounding Mg atoms and atoms"; "The average difference in Bondorder of the surrounding O atom bonds"; "The average difference in bond length between the surrounding O atoms and atoms"; "Compound structure"; 】。;

[0168] In summary:

[0169] 1. The method for predicting adsorption energy based on a hybrid driving strategy avoids cumbersome and expensive experimental procedures, is completely based on theoretical calculations, reduces costs, and at the same time reduces the dependence on first-principles calculations and shortens the research cycle. By combining material properties with a diffusion model, not only the prediction accuracy is enhanced, but also the adaptability to unknown structures is improved.

[0170] 2. The method for generating diverse data by combining high-contribution attributes with a diffusion model. By screening and extracting high-contribution attributes of adsorption energy and inputting these attributes and the structural data of modified magnesium oxide into the diffusion model for learning, the model can generate material structure data with rich diversity and high reliability. This technical point significantly expands the diversity of the dataset and improves the generalization ability of the machine learning model, which is the core innovation point for improving the prediction performance of the present invention.

[0171] 3. Expand the dataset through the diffusion model and establish a high-precision and high-efficiency adsorption energy prediction model to accurately predict any adsorption site of carbon dioxide on the modified magnesium oxide material, providing theoretical guidance for the design of new adsorption materials.

[0172] 4. The model constructed by the present invention is simple and can be further widely applied in the industrial field.

[0173] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for predicting the adsorption of carbon dioxide by modified magnesium oxide based on hybrid drive, characterized in that: At least the following steps are included: S1: Optimize the structure of various modified magnesium oxides through computer simulation software to ensure their stability and reliability as materials; S2: Consider multiple possible adsorption sites, place carbon dioxide gas molecules on the adsorption sites, and calculate the adsorption energy and electrical properties of the corresponding structures to construct a data set for machine learning; S3: Considering the complexity of the dataset, clustering is used to divide the dataset into multiple clusters for research; S4: Construct a correlation analysis model for arbitrary machine learning, whose main function is to filter out attributes with high correlation to adsorption energy from the data set; S5: Determine whether the existing data set needs to be expanded; S6: Apply the attributes highly correlated with adsorption energy to the machine learning model to fit and predict adsorption energy, thereby constructing a prediction model for adsorption energy; S7: Learning from existing data sets through diffusion models can expand data sets and increase data diversity. Diffusion models can supplement more diverse training data for artificial intelligence systems by generating more random adsorption situations and carbon dioxide adsorption results under different conditions. S8: Use the neural network model to fit the expanded data set to predict the adsorption energy, so that the relationship between "adsorption energy-structure" can be mapped; S9: Collect the statistical data of all trained adsorption energy prediction models to evaluate the differences in the adsorption capacity of different modified magnesium oxides for carbon dioxide gas and construct an "adsorption energy-structure" mapping relationship.

2. The method for predicting carbon dioxide adsorption by modified magnesium oxide based on hybrid drive according to claim 1, characterized in that: The S1 at least comprises the following steps: The doping simulation of magnesium oxide was carried out using Material Studio computer simulation software; During the simulation, 17 elements from six different groups in the periodic table were selected; These 17 elements will be doped into magnesium oxide at doping concentrations of 1%, 2%, and 3% to obtain simulation results after doping; The six different groups include alkali metals, alkaline earth metals, transition metals, post-transition metals, metal compounds and nonmetals; The alkali metals include at least Li, Na and K, the transition metals include at least Fe, Co, Ir and Pt, the post-transition metals include at least Al and Ga, the metal compounds include at least Si and Ge, and the non-metals include at least N, P, S and Cl.

3. The method for predicting carbon dioxide adsorption by modified magnesium oxide based on hybrid drive according to claim 2, characterized in that: The S2 at least comprises the following steps: For each element-doped MgO system, there are three doping concentrations with a total of 12 adsorption sites; A total of 204 systems were established through 17 elements; After screening and eliminating the unstable structures, a total of 141 systems were obtained; The screening and exclusion method is to analyze properties, and the properties at least include energy band, state density, differential charge density and charge population; Each system is named ECA, i.e., Element-Concentration-Adsorption sites, where E represents the doping atom, C represents the doping concentration, and A represents the adsorption site; These 141 systems correspond to data sets of 68 types of properties, which can be divided into the following four categories: free state of doped atoms, system before carbon dioxide adsorption, system after carbon dioxide adsorption, and difference properties of the system before and after carbon dioxide adsorption.

4. The method for predicting carbon dioxide adsorption by modified magnesium oxide based on hybrid drive according to claim 3, characterized in that: The data set in S3 contains data from 17 different groups of elements at different doping concentrations and adsorption sites, which is relatively complicated; Although the selected elements represent the six main groups of elements in the periodic table, the order of the elements is not determined by adsorption energy, so cluster analysis is used to divide the data set into different clusters and conduct correlation analysis on each cluster, which is more reliable than studying directly by group; For complex attribute data, spectral clustering method is used to divide the data into multiple clusters, and the data in each cluster has similar adsorption behavior characteristics; Next, Spearman correlation analysis was used to screen out high-contribution attributes that have a significant impact on adsorption energy, which provides an important reference for the subsequent expansion of the diffusion model data set.

5. The method for predicting carbon dioxide adsorption by modified magnesium oxide based on hybrid drive according to claim 4, characterized in that: Since the adsorption energy has a certain probability of containing outliers, and the electrical properties of the compound do not conform to the normal distribution but are fixed properties determined by the structure and composition, the Spearman correlation analysis was selected instead of the Pearson correlation analysis in the correlation analysis in S4. The Spearman correlation does not rely on the distribution assumption of the data and is more suitable for processing non-normally distributed data.

6. The method for predicting the adsorption of carbon dioxide by modified magnesium oxide based on hybrid drive according to claim 5, characterized in that: The S6 at least comprises the following steps: Perform regression prediction on the data set using a variety of fitting algorithms, including but not limited to decision tree, gradient, Lasso regression and MLP; The object of algorithm fitting is not just the whole data set, but also includes labeling each data in the data set with doping element family, doping element period and cluster label after clustering; Further combined with the four attribute categories in S2; In this way, the performance of different algorithms under specific conditions can be studied separately.

7. The method for predicting the adsorption of carbon dioxide by modified magnesium oxide based on hybrid drive according to claim 6, characterized in that: The design of the diffusion model in S7 at least includes the following steps: Through the analysis results of S4, the key attributes that contribute more to the adsorption energy are screened out. The key attributes include but are not limited to the number of electrons in each orbit of the doping atom and the change in charge distribution before and after carbon dioxide adsorption; These properties are combined with the corresponding structural data (based on one-hot codes) to form the input of the diffusion model, which is introduced to learn these data. The diffusion model generates new structures through a gradual process of denoising and denoising, and ensures that the newly generated data is consistent with the real data in terms of physical and chemical properties; In order to adapt to the structure of magnesium oxide, a voxel grid diffusion model was introduced. The voxel grid diffusion model cuts the entire structure into multiple squares through a grid, and introduces noise into the grid. Each square is 0.25 angstroms, which ensures that the hierarchical structure will be maintained regardless of whether it is denoised or not. This is more suitable for the molecular structure of magnesium oxide to expand the data set. The process of the diffusion model is divided into noise addition and denoising. First, a denoising neural network is trained in the noise addition process, and the denoising neural network learns the mapping from the noise molecular distribution to the real molecular distribution; The noise adding process comprises at least the following steps: Noise addition: Add Gaussian noise to each voxel grid to generate noise samples for the next moment; Denoising neural network: The goal of this neural network is to recover the original clean molecules from the noisy samples during the denoising process. Therefore, the denoising neural network should be trained during the denoising process. The training goal is to minimize the mean square error between the voxel grid at time t+1, that is, the voxel grid after denoising, and the voxel grid at time t. The denoising process is different from the traditional diffusion model. After adding noise, the entire voxel grid is actually close to smoothness. The denoising method of the VoxMol voxel grid diffusion model is cited here. The core steps of the denoising method of the VoxMol voxel grid diffusion model include walking and jumping. The goal of the walking step is to sample noise molecules yk from the smooth noise distribution p(y) through the MCMC method, and to generate noise molecules that conform to the distribution by gradually back-propagating in the smooth noise distribution through MCMC. Because the noise distribution p(y) is the result of the convolution of the original molecule distribution p(x) and Gaussian noise, and by iteratively updating the noise molecule yk, the walking step can better approximate the distribution characteristics of p(y) and provide high-quality initial noise samples for subsequent denoising. The formula is as follows: in, is the score function of the smooth distribution, δ is the step size, and ∈k is Gaussian noise; The goal of the walking step is to map the noise molecule yk to the true molecular distribution p(x) through single-step denoising. Through the denoising network that has been trained in the previous denoising process, the clean molecule xk is directly generated from the noise molecule yk in one step. This step skips the multi-step iterative denoising process of the traditional diffusion model. The formula is as follows: xk=yk+σ2gθ(yk) Where gθ(yk) is the score function learned by the denoising network and σ2 is the fixed noise level; In summary, it can be considered that the walking step is responsible for efficiently sampling and providing "rough" noise samples in the smooth noise distribution p(y), while the jumping step is responsible for mapping the noise samples to the true molecular distribution p(x) in one step and refining the true molecular distribution p(x) through the neural network; This allows the VoxMol voxel grid diffusion model to avoid the multiple iterations of other diffusion models and significantly speed up generation. At the same time, in addition to this, a variety of strategies can be used to optimize the generated data during the training process, including the introduction of physical constraints and attribute constraints, thereby expanding the diversity of the original data set.

8. The method for predicting the adsorption of carbon dioxide by modified magnesium oxide based on hybrid drive according to claim 6, characterized in that: The S8 at least comprises the following steps: The neural network model was used to fit the expanded data set of S7 to predict the adsorption energy; According to the scale of the data set, a multi-layer graph neural network is designed; the input layer includes all the attributes of the extended data set; the number of hidden layers and the number of neurons in each layer are determined by cross-validation, and the activation function uses ReLU or other nonlinear functions; the output layer is a single node for predicting adsorption energy; the data set is divided into training set, validation set and test set by stratified cross-validation to ensure that each stratum covers different categories of structures and adsorption energy values ​​evenly; Choose the appropriate optimizer and loss function for training, and adjust the learning rate to avoid overfitting or underfitting; In addition, an early stopping strategy is used to monitor the performance of the validation set to prevent the training from being stopped before the model is overfitted; During the training process, a multi-layer perceptron was introduced, and high-contribution attributes were learned preferentially based on the results of feature selection. Residual connections were introduced to increase the number of graph neural network layers to extract data features. The model hyperparameters were optimized and good results were obtained on public datasets before being applied to the expanded dataset, ensuring that the model could predict adsorption energy with high accuracy.

9. The method for predicting the adsorption of carbon dioxide by modified magnesium oxide based on hybrid drive according to claim 8, characterized in that: The execution process of the neural network model is as follows: First is the input layer, which assigns initial feature vectors to each node and edge; Then, information is passed through the graph convolution layer. In each layer, the feature vector of the node is aggregated and updated with the feature vectors of neighboring nodes to capture local structural information. The aggregation method includes at least summation, averaging and maximum pooling; In order to ensure the equivariance of the model to graph isomorphism, a symmetric aggregation function can be used, which can avoid the model's sensitivity to node order and prevent isomorphism inconsistency; After multiple layers of graph convolution, a global pooling layer is used to aggregate node features into a global feature vector of fixed length, representing the features of the entire crystal structure; Finally, the global feature vector is input into the fully connected layer for regression prediction of the adsorption energy of magnesium oxide. At the same time, in addition to the adsorption energy, the prediction results of other attributes are output according to the task requirements to achieve the results of the research on the adsorption mechanism of magnesium oxide.

10. The method for predicting the adsorption of carbon dioxide by modified magnesium oxide based on hybrid drive according to claim 1, characterized in that: When determining whether the data set needs to be expanded in S5, a systematic evaluation is performed based on multi-dimensional standards; The systematic evaluation based on multi-dimensional criteria at least includes drawing a learning curve of the machine learning model, using statistical power analysis and feature space coverage of the data set; Draw the learning curve of the machine learning model. If the training error and validation error do not reach the convergence platform, it indicates that increasing the amount of data may improve performance. Statistical power analysis was used to calculate the confidence interval coverage ability of the current sample size for the target prediction accuracy; Feature space coverage of the dataset: t-SNE is used to visualize the high-dimensional feature distribution and identify sparse areas in the feature space to supplement the uncovered adsorption site configurations. However, considering that some characteristic compounds may not exist in reality, or stable structures are likely to be distributed in the same threshold space, this solution is only an alternative. The main judgment method is still to rely on whether the training error and verification error mentioned above have reached convergence, as well as the current sample size’s confidence interval coverage of the target prediction accuracy to decide whether to supplement the dataset.