Method for evaluating influence law of entropy increase on OER performance of catalytic material
By combining machine learning algorithms and data analysis, the impact of entropy increase on the OER performance of catalytic materials was evaluated. The negative correlation between configurational entropy and OER overpotential was revealed, which solved the research limitations in the existing technology and achieved efficient OER performance evaluation and catalyst development guidance.
Patent Information
- Application Number
- CN202310025082.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-01-04
AI Technical Summary
Existing technologies do not consider the impact of entropy increase on the OER performance of catalytic materials, which limits research to traditional chemical reaction mechanisms and makes it impossible to effectively evaluate the performance improvement of high-entropy OER catalysts.
By combining machine learning algorithms and data analysis, we constructed a regression model and performed SHAP analysis to evaluate the influence of entropy increase on the OER performance of catalytic materials. Using the information entropy of the physicochemical properties of elements as features, we conducted unsupervised learning and feature engineering to reveal the influence of configuration entropy on OER overpotential.
It boasts high accuracy, small error, and resource conservation, providing new research directions, guiding the development of high-efficiency, high-entropy OER catalysts, reducing experimental costs, and improving work efficiency.
Smart Images

Figure CN116092609B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to catalysis technology, and in particular to a method for evaluating the influence of entropy increase on the OER performance of catalytic materials. Background Technology
[0002] Hydrogen energy is one of the most important and efficient clean energy sources in the 21st century, and electrochemical water splitting for hydrogen production is a key research area. The oxygen evolution reaction (OER) at the anolyte is the bottleneck in electrochemical water splitting for hydrogen production. Researchers are increasingly focusing on electrocatalysts (including metal oxides, hydroxides / hydroxyoxides, chalcogenides, selenides, nitrides, phosphides / phosphates, borides, carbides, metal-organic frameworks, and non-metallic compounds), and have achieved considerable results.
[0003] High-entropy materials, due to their high configurational entropy, induce a large number of coordinated unsaturated reactive sites; the entropy stabilization effect can play a role in stabilizing the structure during electrochemical reactions, and the unique "cocktail effect" and tunable performance endow them with great potential for structural performance regulation. These advantages make high-entropy materials a very promising class of oxygen evolution reaction catalysts.
[0004] Currently, most high-entropy OER catalysts are prepared based on fourth-period transition metals, thus inherently possessing excellent OER performance. Researchers typically focus on improving the performance of catalyst materials under different preparation conditions or the impact of different element combinations on catalytic performance, but generally have not considered the potential influence of entropy increase within the material on OER performance. Therefore, there are currently no related research results or progress reports.
[0005] This invention aims to propose an evaluation method for the influence of entropy increase on the OER performance of catalytic materials, so as to provide guidance for further exploration of efficient and high-entropy OER catalysts. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a method for evaluating the influence of entropy increase on the OER performance of catalytic materials.
[0007] To solve the technical problem, the solution of the present invention is:
[0008] A method for evaluating the effect of entropy increase on the OER performance of catalytic materials is provided, comprising the following steps:
[0009] (1) Establish the elemental composition-overpotential dataset D;
[0010] (2) Select several elements with physical and chemical properties related to OER performance to form set A;
[0011] (3) For each physicochemical property A in set A i Use the Pymatgen program to obtain the corresponding values of the elements other than the noble gases in the first 6 periods of the periodic table, forming a set.
[0012] (4) For each set The data in the dataset was classified using the K-Means clustering method from Scikit-learn, with the number of clusters denoted as k.
[0013] (5) For each sample in dataset D, obtain the A in that sample based on its element composition. i Category distribution of attributes when the number of clusters is set to k pass A was calculated i The corresponding entropy value of the attribute when the number of clusters is set to k and under different k values Merge them into a set S, and use set S as the feature of the sample; construct a new dataset X in this way, and divide it into training sets X. train and test set X test ;
[0014] (6) Perform mean and variance normalization on the training set and test set obtained in step (5) so that all data are normalized to a distribution with a mean of 0 and a variance of 1.
[0015] (7) Use machine learning algorithms to perform regression fitting on the training set processed in step (6) to obtain regression model M; perform hyperparameter tuning on the model and use the test set processed in step (6) to test generalization performance.
[0016] (8) Use SHAP to analyze model M on the test set to obtain the importance of each feature in the model and the influence of each feature on the prediction of OER overpotential by model M.
[0017] As a preferred embodiment of the present invention, in step (2), the physicochemical properties of the element specifically include at least one of the following: atomic number, atomic radius, common valence state, Pauling electronegativity, average radius of common ions, outermost atomic orbital energy, thermal conductivity, and electrical conductivity.
[0018] As a preferred embodiment of the present invention, in step (4), when A i When setting the atomic number and common valence state as described in step (2), k is set to The number of distinct values in the range; otherwise, k is set to 3, 4, 5, 6, 7 respectively, i.e., k∈{3, 4, 5, 6, 7} for multiple classifications.
[0019] As a preferred embodiment of the present invention, in step (5), the entropy calculation adopts the following formula:
[0020]
[0021] In the above formula, For A with the number of clusters set to k i The entropy value of the attribute, For A with the number of clusters set to k i The attribute is category x j The probability of.
[0022] As a preferred embodiment of the present invention, in step (5), the dataset X is randomly divided into training set X in a ratio of 7:3. train and test set X test .
[0023] As a preferred embodiment of the present invention, in step (7), the machine learning algorithm refers to the algorithm provided by the Autogluon framework, which is either Random Forest or XGBoost.
[0024] Description of the invention principle:
[0025] 1. This invention creatively addresses the technical problem of "how entropy increase affects the OER performance of catalytic materials" and proposes a reasonable and effective solution.
[0026] In the process of developing high-entropy OER catalysts with superior performance, the impact of entropy increase is a crucial and unavoidable issue. However, due to limitations in previous research directions, the common approach has remained at the level of traditional chemical reaction mechanisms. Therefore, research on the OER performance of catalytic materials cannot exclude various factors affecting the reaction mechanism and cannot simply consider the influence of configurational entropy on OER performance. Furthermore, existing published literature lacks specific reports on how the pure entropy increase effect specifically affects the OER performance of catalytic materials.
[0027] During their research on catalytic materials, the applicant's team of inventors, through in-depth analysis and research on the physicochemical microscopic mechanism of high-entropy OER catalysts, creatively proposed a method for evaluating the impact of entropy increase on the performance of catalytic materials. This solution breaks through the traditional practices and conventional thinking in the field of catalyst research.
[0028] 2. High entropy is a scientific definition, generally referring to compounds containing five or more elements at the same chemical structural site. High-entropy materials exhibit better structural stability. Furthermore, the heterogeneous elements and structural distortions introduced by high entropy enhance the material's designability; desired material properties can be achieved by adjusting the types of elements present. Current research commonly employs a method of preparing materials with novel combinations through numerous repetitive experiments, followed by screening through material performance testing. Therefore, conventional methods for understanding the impact of compositional adjustments on material properties require substantial human and financial resources and are highly inefficient.
[0029] This invention creatively proposes a data-driven approach that combines machine learning algorithms and data analysis to analyze the influence of entropy increase on the performance of catalytic materials. By utilizing existing publicly available data, a machine learning model is trained, and SHAP is used to study the effect of configurational entropy on overpotential within this model. This not only calculates the importance of features to the overall model but also reveals how each feature affects the model's predicted values.
[0030] 3. This invention employs feature engineering based on prior knowledge of OER catalytic materials. In feature engineering, Pymatgen is first used to obtain elemental physicochemical features. Then, through unsupervised learning, the continuous values of these elemental physicochemical features are discretized, and the entropy of the elemental physicochemical properties in the compound is further calculated. During model training, this invention uses only a series of entropy features obtained in the feature engineering as input vectors, and various machine learning algorithms are used to train the dataset to obtain a regression model with high generalization ability. To study the influence of configurational entropy on OER overpotential, this invention also uses the SHAP method to solve and rank the importance of each entropy feature in the model. Simultaneously, the influence of configurational entropy on OER overpotential is observed based on the honeycomb diagram drawn by SHAP. These are innovative approaches not previously documented in any published literature in this technical field.
[0031] Compared with the prior art, the beneficial effects of the present invention are:
[0032] 1. This invention is the first to propose a correlation between the information entropy of elemental physicochemical properties and the OER performance of catalytic materials. Using the SHAP method, this invention analyzes the impact of features on model prediction results, demonstrating that configuration entropy has a significant influence on overpotential. Furthermore, through honeycomb diagrams, it reveals a negative correlation between configuration entropy and OER overpotential. OER overpotential is an indicator of OER performance. This invention obtains specific influence patterns by evaluating the relationship between entropy increase and the OER performance of catalytic materials, providing a new research direction for OER catalytic materials.
[0033] 2. This invention is the first to propose combining machine learning algorithms and data analysis to evaluate the impact of entropy increase on the OER performance of catalytic materials, thereby providing OER performance guidance for the development of new catalytic materials. This method eliminates the need for extensive laboratory operations, saving significant resources and improving work efficiency.
[0034] 3. The method proposed in this invention has high accuracy and small error. By using information entropy based on the physicochemical properties of elements as descriptors to characterize the OER performance of materials, and utilizing easily obtainable elemental physicochemical properties as initial features, this invention generates information entropy features through unsupervised learning to train the model. The resulting model achieves a final generalization RMSE error as low as 26.3mV, which is significantly lower than the 38mV RMSE error of similar machine learning models.
[0035] 4. This invention reveals for the first time the impact of pure entropy increase on OER performance. By using the SHAP method to interpret the model and calculating the marginal contribution of entropy to OER performance, this invention obtains the proportion of entropy's influence on OER performance among numerous influencing factors, ultimately deriving the law governing the impact of pure entropy increase on OER performance. Attached Figure Description
[0036] Figure 1 This illustrates the impact of each sample feature on the importance of the model when using SHAP analysis in Example 1.
[0037] Figure 2 This is a honeycomb diagram of the model prediction values for each sample feature in Example 1;
[0038] Figure 3 This illustrates the impact of each sample feature on the importance of the model when using SHAP analysis in Example 2.
[0039] Figure 4 This is a honeycomb diagram of the model prediction values for each sample feature in Example 2;
[0040] Figure 5 This illustrates the impact of each sample feature on the importance of the model when using SHAP analysis in Example 3.
[0041] Figure 6 This is a honeycomb diagram of the model prediction values for each sample feature in Example 3. Detailed Implementation
[0042] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0043] Example 1:
[0044] A method for evaluating the effect of entropy increase on the OER performance of catalytic materials includes the following steps:
[0045] (1) Based on the OER overpotential data provided in the published literature (Rohr, B., et al. (2020). "Benchmarking the acceleration of materials discovery by sequential learning." Chemical Science 11(10): 2696-2706), the component dataset D was constructed with a size of 8424.
[0046] Each sample in dataset D represents a catalytic material, including its elemental composition and overpotential.
[0047] (2) Select several physicochemical properties of elements related to OER performance to form a set A; A = {atomic number, atomic radius, common valence state, Pauling electronegativity, average radius of common ions, outermost atomic orbital energy, thermal conductivity, resistivity}.
[0048] (3) For each physicochemical property A in set A i Use the Pymatgen program to obtain the corresponding values of the elements other than the noble gases in the first 6 periods of the periodic table, forming a set.
[0049] This step is preparatory work for calculating the entropy of an element's physicochemical properties. For example, to calculate the entropy of thermal conductivity, the value of thermal conductivity needs to be classified before the calculation can proceed. Classification requires statistical analysis of all thermal conductivity values in the periodic table, and then, based on the statistical results, defining a classification standard for thermal conductivity values: how many are considered high, how many are considered medium, and how many are considered low. Catalytic materials do not contain rare gases and should be excluded when statistically analyzing the specific values of the element's physicochemical properties.
[0050] (4) For each set The data in the dataset was classified using Scikit-learn's K-Means clustering method, with the number of clusters denoted as k; when A i When the atomic number and common valence state are as described in step (2), k is set to The number of distinct values is determined; otherwise, k is set to 3, 4, 5, 6, 7 respectively, i.e., k∈{3, 4, 5, 6, 7} for multiple classifications.
[0051] (5) For each sample in dataset D, obtain the A in that sample based on its element composition. i Category distribution of attributes when the number of clusters is set to k
[0052] pass A was calculated i The corresponding entropy value of the attribute when the number of clusters is set to k and under different k values Merge them into a set S; S serves as the feature of this sample.
[0053]
[0054] In the above formula, For A with the number of clusters set to k i The entropy value of the attribute, For A with the number of clusters set to k i The attribute is category x j The probability of;
[0055] Let S be the set of features of this sample;
[0056] Construct a new dataset X in this manner, and then partition it into a training set X in a 7:3 ratio using a random partitioning method. train and test set X test The training set accounts for 70% and the test set accounts for 30%.
[0057] (6) For X train and X test Perform mean and variance normalization, which means normalizing all data to a distribution with a mean of 0 and a variance of 1.
[0058] (7) Use the random forest algorithm on the training set X train Regression fitting was performed to obtain a regression model. The model's hyperparameters were then tuned. The final hyperparameters were determined to be n_estimators = 220, max_depth = 10, min_samples_leaf = 5, and min_samples_split = 7. The resulting model M had an RMSE error of 24.3 mV on the training set. Further evaluation of the model's generalization performance using the test set showed an RMSE error of 27.9 mV on the test set.
[0059] (8) Use SHAP to analyze model M on the test set to obtain the importance of each feature in the model and the influence of each feature on the prediction of OER overpotential by model M.
[0060] The SHAP analysis refers to the Shapley Additive explanations method proposed by Lundberg in his paper "A Unified Approach to Interpreting Model Predictions".
[0061] The abbreviations for the various features are as follows:
[0062] en_COS: Common valence state entropy, en_AN: Atomic number entropy (configuration entropy), TC i B 热导率 Thermal conductivity entropy obtained by classifying into i categories, ER i B 电阻率 The resistivity entropy obtained by classifying it into class i, ACR i B 常见离子平均半径 The average radius entropy of common ions obtained by classifying them into class i, PE i B 泡林电负性 Pauling electronegativity entropy obtained by classifying it into class i, AR i B 原子半径 The atomic radius entropy obtained by classifying into i, AO i B 最外层原子轨道能量 The outermost atomic orbital energy entropy is obtained by classifying it into type i.
[0063] The importance of each feature to the model is as follows: Figure 1 As shown, the cellular structure of all sample features affects the model's predicted values. Figure 2 As shown.
[0064] Based on the data relationships in the figure, it can be concluded that configuration entropy and OER overpotential are negatively correlated.
[0065] Example 2:
[0066] Compared with Example 1: In step (2), the physicochemical property set A removes the term "resistivity"; in step (4), the number of clusters is set to 3, 4, 5, 6, k∈{3, 4, 5, 6}; in step (7), the XGBoost algorithm is used to cluster X train Regression fitting was performed, and hyperparameters were tuned. The final hyperparameters were determined to be: learning_rate = 0.1, n_estimators = 280, max_depth = 3, gama = 0, subsample = 0.7, and colsample_bytree = 1. The resulting model M had an RMSE error of 28.7 mV on the training set and an RMSE error of 30.1 mV on the test set. All other operational steps remained consistent with those in Example 1.
[0067] The importance of each feature to the model is as follows: Figure 3 As shown, the cellular structure of all sample features affects the model's predicted values. Figure 4 As shown.
[0068] Based on the data relationships in the figure, it can be concluded that configuration entropy and OER overpotential are negatively correlated.
[0069] Example 3:
[0070] Compared to Example 1: In step (2), the physicochemical property set A removes the term "resistivity"; in step (4), the cluster numbers are set to 3, 4, 5, 6, 7, k∈{3, 4, 5, 6, 7}; Autogluon is used to analyze X. train Regression fitting was performed. The final model M had an RMSE error of 25.3 mV on the training set and an RMSE error of 26.6 mV on the test set. The remaining steps were consistent with those in Example 1.
[0071] The importance of each feature to the model is as follows: Figure 5 As shown, the cellular structure of all sample features affects the model's predicted values. Figure 6 As shown.
[0072] Based on the data relationships in the figure, it can be concluded that configuration entropy and OER overpotential are negatively correlated.
[0073] Further explanation:
[0074] In the final step of this invention, the object of analysis is the model M, and the analysis tool used is the SHAP method; that is, SHAP is used to analyze the model's performance on the test set, and then the analysis results are output. If the model is regarded as a function f, then the analysis of the model using SHAP in this invention is similar to using a certain method to analyze the monotonicity, maxima, minima, and boundedness of the function f.
[0075] Machine learning can be used to fit a model for prediction, or to interpret the fitted model to discover patterns in the data. Predicting involves fitting a function f to the training data and using f to predict the output value f(x) for an input x. Interpreting involves studying why an input x yields an output value f(x) for the established function f. In this invention, a function f is first fitted to predict the OER performance of a material, and then the function f is interpreted. In other words, by interpreting the fitted function f, patterns in the existing data are discovered, revealing a direct negative correlation between configuration entropy and overpotential.
[0076] This invention innovatively uses the entropy of an element's physicochemical properties as a feature, rather than the element's physicochemical properties themselves. The entropy in high-entropy materials refers to configurational entropy. The conclusions of the three different embodiments provided in this invention also verify that configurational entropy is a relatively important feature, and it was found that the higher the configurational entropy, the lower the OER overpotential, and a lower overpotential means higher OER performance. Therefore, the high OER performance of high-entropy OER catalytic materials is not due to the introduction of multiple elements with inherently good performance, but rather the increase in entropy itself leads to improved OER performance. Therefore, the influence of configurational entropy on overpotential is something that those skilled in the art need to pay attention to during the research of catalytic materials.
[0077] Furthermore, the model used in this invention has smaller errors compared to similar technologies. Smaller errors indicate that the discovered patterns are closer to the patterns truly reflected in the data, meaning the discovered patterns are more reliable.
Claims
1. A method for evaluating the influence of entropy increase on the OER performance of catalytic materials, characterized in that, Includes the following steps: (1) Establish the elemental composition-overpotential dataset D; (2) Select several elements with physical and chemical properties related to OER performance to form set A; (3) For each physicochemical property in set A Use the Pymatgen program to obtain the corresponding values of the elements other than the noble gases in the first 6 periods of the periodic table, forming a set. ; (4) For each set The data in the dataset was classified using the K-Means clustering method from Scikit-learn, with the number of clusters denoted as k. (5) For each sample in dataset D, obtain the composition of the sample based on its elements. Category distribution of attributes when the number of clusters is set to k ; pass Calculated The corresponding entropy value of the attribute when the number of clusters is set to k and under different k values Merge into a set , in sets As a feature of this sample; Build a new dataset in this way And divided into training sets and test set ; (6) Perform mean and variance normalization on the training set and test set obtained in step (5) so that all data are normalized to a distribution with a mean of 0 and a variance of 1. (7) Use machine learning algorithms to perform regression fitting on the training set processed in step (6) to obtain regression model M; perform hyperparameter tuning on the model and use the test set processed in step (6) to perform generalization performance testing; (8) Use SHAP to analyze model M on the test set to obtain the importance of each feature in the model and the influence of each feature on the prediction of OER overpotential by model M.
2. The method according to claim 1, characterized in that, In step (2), the physicochemical properties of the element specifically include at least one of the following: atomic number, atomic radius, common valence state, Pauling electronegativity, average radius of common ions, outermost atomic orbital energy, thermal conductivity, and electrical conductivity.
3. The method according to claim 2, characterized in that, In step (4), when When the atomic number and common valence state are as described in step (2), k is set to The number of distinct values; otherwise, set k to 3, 4, 5, 6, 7 respectively, i.e. Perform multiple classifications.
4. The method according to claim 1, characterized in that, In step (5), the entropy is calculated using the following formula: (1) In the above formula, For clusters with the number of clusters set to k The entropy value of the attribute, For cluster number set to k The attribute is category The probability of.
5. The method according to claim 1, characterized in that, In step (5), the dataset is... The data were randomly divided into training sets at a ratio of 7:
3. and test set .
6. The method according to claim 1, characterized in that, In step (7), the machine learning algorithm refers to the algorithm provided by the Autogluon framework, which is either the Random Forest algorithm or XGBoost.
Citation Information
Patent Citations
First principle prediction method for regulating and controlling Co3S4 hydrolysis catalytic performance by P doping
CN111415714A
Intelligent design method and system for transition metal hydroxide oxygen evolution electrocatalyst
CN115331747A