A machine learning based high capacity hydrogen storage alloy design method
By optimizing the design of hydrogen storage alloy composition through machine learning and genetic algorithms, the problems of low efficiency and low accuracy in existing technologies have been solved, enabling rapid and accurate design of high-capacity hydrogen storage materials and improving the development efficiency and precision of hydrogen storage materials.
Patent Information
- Application Number
- CN202211673962.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-12-26
AI Technical Summary
Existing hydrogen storage alloy design methods are inefficient and inaccurate, making it difficult to effectively improve the capacity of hydrogen storage materials and limiting the application and promotion of solid-state hydrogen storage technology.
By combining machine learning methods with genetic algorithms, and through data preprocessing, modeling and optimization using multiple machine learning algorithms, a high-capacity hydrogen storage alloy composition was designed. The Xgboost algorithm was used as the fitness function to achieve rapid and accurate design of the alloy composition.
It has achieved high-precision prediction of hydrogen storage capacity of hydrogen storage alloys with a relative error as low as 0.54%, which significantly improves the efficiency and accuracy of hydrogen storage material development.
Smart Images

Figure CN116092606B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of solid-state hydrogen storage technology, and in particular to a high-capacity hydrogen storage alloy design method based on machine learning. BACKGROUND
[0002] Solid-state hydrogen storage technology has the characteristics of low hydrogen storage pressure and high hydrogen storage density, and is considered a safe and efficient hydrogen storage technology, which is a hot spot in the field of hydrogen energy at home and abroad. Solid-state hydrogen storage materials mainly include LaNi5 series, TiFe series, TiMn2 series, titanium-vanadium solid solution system, magnesium-based and other types. At present, LaNi5, TiFe, TiMn2 and other series of alloys that have realized commercial application can quickly absorb and release hydrogen near room temperature, and have excellent application prospects in the fields of fuel cell special vehicles and distributed power generation, but such materials have the disadvantage of low hydrogen storage mass density (<1.8wt%), which restricts the application and promotion of solid-state hydrogen storage technology. Therefore, it is crucial to further improve the hydrogen storage capacity of LaNi5 series, TiFe series, TiMn2 series and other alloys, and to develop higher-capacity titanium-vanadium solid solutions (theoretical maximum hydrogen storage capacity 3.7wt%) and magnesium-based hydrogen storage materials (theoretical maximum hydrogen storage capacity 7.6wt%) that can reversibly absorb and release hydrogen near room temperature.
[0003] Existing hydrogen storage alloy design methods are mainly based on empirical criteria, such as large unit cell parameters and large metal-hydrogen electronegativity differences, but empirical design methods have low accuracy and low development efficiency, and still need to be continuously tested and improved through experimental methods to obtain good alloy compositions. There are also linear fitting methods for alloy design, but the performance prediction accuracy is low. Therefore, it is urgent to develop a high-efficiency and accurate hydrogen storage material performance prediction model to improve the development efficiency of high-capacity hydrogen storage materials. SUMMARY
[0004] In view of the problems of low efficiency and high cost of the empirical criterion method currently used for hydrogen storage alloy design, and the low prediction accuracy of the linear fitting model, the present application uses a machine learning method to accurately and quickly predict the maximum hydrogen storage capacity of hydrogen storage alloys, and further uses a genetic algorithm to optimize and experimentally verify alloy compositions with high maximum hydrogen storage capacity, significantly improving the development efficiency of high-capacity hydrogen storage materials and reducing the development cost.
[0005] The high-capacity hydrogen storage alloy design method provided by the present application specifically includes the following steps:
[0006] 1) Obtain data: obtain the hydrogen storage capacity of m kinds of hydrogen storage alloys and the data of their alloy compositions. The alloy composition and hydrogen storage capacity of each hydrogen storage alloy are a group of original data, and the m groups of original data of the m kinds of hydrogen storage alloys constitute an original data set.
[0007] 2) data preprocessing: according to the hydrogen storage amount and the alloy composition of the hydrogen storage alloy obtained in step 1), the atomic radius, the electronegativity, the average value and the variance of the number of valence electrons, and the average value of the bulk modulus of the hydrogen storage alloy are calculated; d mean , d var , X mean , X var , e mean , e var , G mean , B mean After standardization, the hydrogen storage amount of the corresponding alloy is composed of standard data set;
[0008] The calculation formula is as follows:
[0009]
[0010]
[0011]
[0012]
[0013]
[0014]
[0015]
[0016]
[0017] Wherein, c i and w i represent the atomic percentage and the mass percentage of the i-th component in the alloy, n is the number of alloy components, d i , X i , e i , G i respectively represent the atomic radius, the electronegativity, the number of valence electrons and the bulk modulus of the i-th component, d mean , d var , X mean , X var , e mean , e var , G mean , B mean respectively represent the average value of the atomic radius, the variance of the atomic radius, the average value of the electronegativity, the variance of the electronegativity, the average value of the number of valence electrons, the variance of the number of valence electrons, the weighted average of the bulk modulus with respect to the atomic percentage, and the weighted average of the bulk modulus with respect to the mass percentage.
[0018] 3) Data set division: using leave-one-out method to divide the data set obtained in step 2) into training set and test set, in turn, one sample in the data set is taken as test set, and the remaining samples are taken as training set, and so on until each sample is traversed.
[0019] 4) Establishing prediction model: using multiple machine learning algorithms, inputting the parameters of the machine learning algorithm, training and testing the training set and the test set obtained in step 3) respectively using grid search, and accordingly establishing multiple prediction models.
[0020] 5) Determining the optimal prediction model: comparing the predicted values and actual values of the multiple prediction models obtained in step 4), calculating the relative error of the two, and the prediction model with the smallest relative error is determined as the optimal prediction model.
[0021] 6) Alloy composition design: setting the optimal prediction model determined in step 5) as the fitness function, and using genetic algorithm to optimize the alloy composition with the hydrogen storage capacity of the hydrogen storage alloy in step 1) to achieve fast and accurate design of the hydrogen storage alloy composition.
[0022] In step 1), the data sources of the hydrogen storage capacity of the hydrogen storage alloy and its alloy composition include existing technical literature, experimental results or tool manuals.
[0023] In step 1), when collecting data, the test temperature of hydrogen storage capacity should be near room temperature to exclude the influence of temperature on hydrogen storage capacity, and the collected data is used as the original data set for subsequent modeling.
[0024] In step 1), the hydrogen storage alloy includes rare earth AB5 type, titanium AB type, AB2 type, vanadium solid solution type, and magnesium-based hydrogen storage alloy.
[0025] In step 4), the machine learning algorithm includes linear regression, support vector machine, random forest, and Xgboost.
[0026] In step 4), the parameters of the linear regression algorithm include the penalty term alpha; the parameters of the support vector machine include the regularization term C, the kernel function name kernel, the interval width epsilon, and the kernel function parameter degree; the parameters of the random forest include the maximum depth max_depth, the number of decision trees n_estimators; and the parameters of the Xgboost algorithm include the maximum depth max_depth, the minimum loss reduction required for node branching gamma, the regularization coefficient, and the number of decision trees n_estimators.
[0027] In step 5), the optimal prediction model is Xgboost, and the optimal parameter combination is {max-depth = 6, gamma = 0.5, alpha = 1, n_estimators = 200}.
[0028] The step further comprises:
[0029] 7) According to the hydrogen storage alloy composition obtained in step 6), a hydrogen storage alloy sample is prepared, and the actual hydrogen storage capacity is tested and compared with the predicted hydrogen storage capacity of the hydrogen storage alloy, and if the preset error requirement is not met, the model parameters in step 4) are adjusted, and steps 4)-7) are repeated until the preset error requirement is met.
[0030] Compared with the traditional hydrogen storage alloy design method based on empirical criteria, the high-capacity hydrogen storage alloy design method based on machine learning used in the present application has the following beneficial effects:
[0031] (1) The present application can realize high-precision prediction of the hydrogen storage capacity of the hydrogen storage alloy by establishing a suitable prediction model, which is more efficient than the traditional empirical criteria method and the linear fitting method.
[0032] (2) The present application trains and tests the training set and test set respectively by using multiple machine learning algorithms, establishes multiple prediction models accordingly, and selects the prediction model with the smallest error as the optimal prediction model, so that the relative error between the predicted value and the actual value is as low as within 15%, which can further improve the high-precision prediction of the hydrogen storage capacity of the hydrogen storage alloy compared with the alloy composition design method using a single prediction model.
[0033] (3) The present application uses Xgboost algorithm as the fitness function, and the average and maximum fitness function values of the population are converged, and the relative error between the actual hydrogen storage capacity and the predicted hydrogen storage capacity is as low as 0.54% after the preparation of the alloy, i.e. high-precision prediction is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to make the content of the present application more easily understood, the present application will be further described in detail below according to specific embodiments of the present application and in conjunction with the accompanying drawings, in which,
[0035] Figure 1 Figure is the predicted value and actual value graph obtained by using the linear model leave-one-out method in the present application embodiment 1;
[0036] Figure 2 Figure is the predicted value and actual value graph obtained by using the Xgboost model leave-one-out method in the present application embodiment 1;
[0037] Figure 3 Figure is the evolution curve of the genetic algorithm in the present application embodiment 1;
[0038] Figure 4 For the V 85 Ti 10 Cr1Fe4 room temperature hydrogen absorption and desorption PCT curve. DETAILED DESCRIPTION
[0039] For the purpose, technical solutions and advantages of the present application, the following will be combined with the drawings of the embodiments of the present application to describe the technical solutions of the embodiments of the present application in more detail. The described embodiments are part of the embodiments of the present application, but not all the embodiments. Similar improvements and adjustments made by those skilled in the art according to the content of the present application, and other embodiments obtained without creative labor, are all considered to be within the scope of protection of the present application.
[0040] The high-capacity hydrogen storage alloy design method provided by the present application specifically comprises the following steps:
[0041] 1) Obtain data: obtain the hydrogen storage capacity of m kinds of hydrogen storage alloys and the data of their alloy components, the alloy components and hydrogen storage capacity of each kind of hydrogen storage alloy are a group of original data, and the m groups of original data of the m kinds of hydrogen storage alloys constitute an original data set.
[0042] 2) Data preprocessing: according to the hydrogen storage capacity and alloy components of the hydrogen storage alloy obtained in step 1), the average value and variance of the atomic radius, electronegativity, number of valence electrons, and the average value of the bulk modulus of the hydrogen storage alloy are calculated; d mean , d var , X mean , X var , e mean , e var , G mean , B mean After standardization, the hydrogen storage capacity of the corresponding alloy forms a standard data set;
[0043] Among them, the calculation formula is as follows:
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] wherein c i and w i represent the atomic percentage and the mass percentage of the i-th component in the alloy, n is the number of components of the alloy, d i , X i , e i , G i represent the atomic radius, the electronegativity, the number of valence electrons and the bulk modulus of the i-th component, d mean , d var , X mean , X var , e mean , e var , G mean , B mean represent the average value of the atomic radius, the variance of the atomic radius, the average value of the electronegativity, the variance of the electronegativity, the average value of the number of valence electrons, the variance of the number of valence electrons, the weighted average of the bulk modulus with respect to the atomic percentage, the weighted average of the bulk modulus with respect to the mass percentage.
[0053] 3) Data set division: using leave-one-out method to divide the data set obtained in step 2) into training set and test set, in turn taking one sample in the data set as test set and taking the remaining samples as training set, repeating until each sample is traversed.
[0054] 4) Establishing prediction model: using multiple machine learning algorithms, inputting the parameters of the machine learning algorithms, training and testing the training set and the test set obtained in step 3) respectively using grid search, and accordingly establishing multiple prediction models.
[0055] 5) Determining the optimal prediction model: comparing the predicted values and the actual values of the multiple prediction models obtained in step 4), calculating the relative error of the two, and the prediction model with the smallest relative error is determined as the optimal prediction model.
[0056] 6) Alloy composition design: setting the optimal prediction model determined in step 5) as the fitness function, and using genetic algorithm to optimize the alloy composition with the hydrogen storage capacity of the hydrogen storage alloy in step 1), to realize the rapid and accurate design of the hydrogen storage alloy composition.
[0057] 7) According to the hydrogen storage alloy composition obtained in step 6), preparing a hydrogen storage alloy sample and testing its actual hydrogen storage capacity, and comparing it with the predicted hydrogen storage capacity of the hydrogen storage alloy, if it does not meet the preset error requirement, adjusting the model parameters in step 4), repeating steps 4)-7), until the preset error requirement is met.
[0058] In step 1), the data sources of the hydrogen storage amount of the hydrogen storage alloy and its alloy composition include existing technical literature, experimental results or tool manuals.
[0059] In step 1), when collecting data, the test temperature of the hydrogen storage amount should be near room temperature to exclude the influence of temperature on the hydrogen storage amount, and the collected data is used as the original data set for subsequent modeling.
[0060] In step 1), the hydrogen storage alloy includes rare earth AB5 type, titanium AB type, AB2 type, vanadium solid solution type, and magnesium-based hydrogen storage alloy.
[0061] In step 4), the machine learning algorithm includes linear regression, support vector machine, random forest, and Xgboost.
[0062] In step 4), the parameters of the linear regression algorithm include the penalty term alpha; the parameters of the support vector machine include the regularization term C, the kernel function name kernel, the interval width epsilon, and the kernel function parameter degree; the parameters of the random forest include the maximum depth max_depth, the number of decision trees n_estimators; and the parameters of the Xgboost algorithm include the maximum depth max_depth, the minimum loss reduction required for node branching gamma, the regularization coefficient, and the number of decision trees n_estimators.
[0063] In step 5), the optimal prediction model is Xgboost, and the optimal parameter combination is {max_depth=6, gamma=0.5, alpha=1, n_estimators=200}.
[0064] The specific implementation is as follows:
[0065] Example 1 :
[0066] The embodiment includes a high-capacity hydrogen storage alloy design method, which specifically includes the following steps:
[0067] 1) Obtain the maximum hydrogen storage amount of 85 kinds of V-Ti-Fe hydrogen storage alloys at 20-30℃, alloy composition and corresponding cell parameter data from literature or experimental results;
[0068] 2) Data preprocessing: calculate the structural parameters such as the average value and variance of the electron concentration, electronegativity, and bulk modulus of the alloy as a whole according to the alloy composition and the calculation formula as shown below:
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077] wherein c i and w i represent the atomic percentage and the mass percentage of the i-th component in the alloy, n is the number of components of the alloy, d i , X i , e i , G i represent the atomic radius, the electronegativity, the number of valence electrons and the bulk modulus of the i-th component, d mean , d var , X mean , X var , e mean , e var , G mean , B mean represent the average value of the atomic radius, the variance of the atomic radius, the average value of the electronegativity, the variance of the electronegativity, the average value of the number of valence electrons, the variance of the number of valence electrons, the weighted average of the bulk modulus with respect to the atomic percentage, the weighted average of the bulk modulus with respect to the mass percentage. Each feature is calculated according to the formula and standardized to establish a standard data set.
[0078] The basic parameters of the components of the V-Ti-Fe hydrogen storage alloy are as follows in the table:
[0079] Table 1 Basic parameters of the components of the V-Ti-Fe hydrogen storage alloy
[0080]
[0081] 3) The data set is divided into training set and test set using the leave-one-out method, one sample in the data set is sequentially taken as the test set, and the remaining samples are taken as the training set, and this is repeated until each sample is traversed.
[0082] 4) Using multiple machine learning algorithms, using grid search, input machine learning algorithm training and testing on training set and test set, respectively, to establish multiple prediction models. Among them, the above features are input into the Xgboost model together with the alloy element composition percentage, and the optimal parameter combination of {max_depth=6, gamma=0.5, alpha=1, n_estimators=200} is obtained by grid search, and the training and test of the data set are carried out.
[0083] 5) Compare the predicted values and actual values of multiple prediction models, and calculate the relative error of the two. The above feature data is input into the three base learners of random forest, support vector machine and Xgboost strengthened by Adaboost algorithm, and the output of the three base learners is input into the secondary learner, and the secondary learner gives the final output result. Some mathematical indexes are introduced to predict the accuracy, and the following formula is used to calculate the error and correlation between the measured value and the predicted value.
[0084] (1) Mean Square Error (MSE)
[0085]
[0086] (2) Mean Absolute Error (MAE)
[0087]
[0088] (3) goodness of fit (R 2 )
[0089]
[0090] (4) Pearson Correlation (PC)
[0091]
[0092] Where y i is the actual value, is the predicted value. Generally, the smaller the MAE and MSE and the larger the R 2 and PC represent the higher the accuracy of the model prediction.
[0093] The errors (MAE, MSE) and correlations (R 2 , PC) of different prediction models are listed in the following table 1, among which the error of the Xgboost model is the lowest.
[0094]
[0095] Table 1
[0096] Figure 2 The graph shows the predicted and actual values obtained from the leave-one-out test of the XgBoost model. The relative error of the sample is within 30%, while more than half of the samples have a prediction relative error within 15%, with an overall mean squared error of 0.180. Compared to... Figure 1 The overall mean squared error of the predicted and actual values obtained by the leave-one-out test of the linear model is 0.323, and the prediction error of the Xgboost model has been significantly reduced.
[0097] 6) The genetic algorithm used the Xgboost algorithm as the fitness function, with a population size of 100, 25 generations, a genetic probability of 0.3, and a mutation probability of 0.01. After the algorithm converged, an alloy composition range with a hydrogen storage capacity as high as 3.65 wt% was obtained. (85-86) Ti (10-11) Cr (1-2) Fe (1-4) . Figure 3 The evolution curve of the genetic algorithm shows that the average and maximum fitness function values of the population converged within 15 generations.
[0098] 7) With V 85 Ti 10 The alloy was prepared by mixing Cr1Fe4 components and then using a vacuum arc melting method. The hydrogen absorption and desorption PCT curves of the samples were measured using Sievert's constant volume method. (See figure). Figure 4 The actual hydrogen storage capacity obtained was 3.67 wt%, with a relative error of approximately 0.54%. This demonstrates that the present invention achieves accurate prediction of the composition of high-capacity hydrogen storage alloys through machine algorithm learning.
[0099] In summary, the present invention achieves the following:
[0100] (1) By establishing a suitable prediction model, this invention can achieve high-precision prediction of hydrogen storage capacity of hydrogen storage alloys, which is more efficient than the traditional empirical criterion method and linear fitting method.
[0101] (2) This invention uses a variety of machine learning algorithms to train and test the training set and test set respectively, and establishes a number of prediction models accordingly. The prediction model with the smallest error is selected as the optimal prediction model. The relative error between the predicted value and the actual value is as low as 15%. Compared with the alloy composition design method of a single prediction model, it can further improve the high-precision prediction of hydrogen storage capacity of hydrogen storage alloy.
[0102] (3) By using the Xgboost algorithm as the fitness function, the average and maximum fitness function values of the population are converged. After the preparation of the alloy, the relative error between the actual hydrogen storage capacity and the predicted hydrogen storage capacity is as low as 0.54%, which achieves high-precision prediction.
Claims
1. A design method for a high-capacity hydrogen storage alloy, characterized in that, The method includes the following steps: 1) Data acquisition: Acquire data on the hydrogen storage capacity and alloy composition of m hydrogen storage alloys. The alloy composition and hydrogen storage capacity of each hydrogen storage alloy constitute a set of raw data. The m sets of raw data of m hydrogen storage alloys constitute the raw dataset. 2) Data Preprocessing: Based on the hydrogen storage capacity and alloy composition of the hydrogen storage alloy obtained in step 1), calculate the average and variance of the atomic radius, electronegativity, and number of valence electrons of the hydrogen storage alloy, as well as the average value of the bulk modulus; d mean d var X mean X var e mean e var G mean B mean After standardization, it forms a standard dataset with the hydrogen storage capacity of the corresponding alloy; The calculation formula is as follows: Among them, c i and w i Represents the atomic percentage and mass percentage of the i-th component in the alloy, where n is the number of alloy components, and d i X i e i G i d represents the atomic radius, electronegativity, number of valence electrons, and bulk modulus of the i-th component, respectively. mean d var X mean X var e mean e var G mean B mean These represent the average atomic radius, the variance of atomic radius, the average electronegativity, the variance of electronegativity, the average number of valence electrons, the variance of valence electrons, the weighted average of bulk modulus with respect to atomic percentage, and the weighted average of bulk modulus with respect to mass percentage, respectively. 3) Dataset partitioning: Using the leave-one-out method, the dataset obtained in step 2) is partitioned into training and test sets. One sample from the dataset is used as the test set, and the remaining samples are used as the training set. Repeat this process until every sample has been traversed; 4) Establish prediction models: Using multiple machine learning algorithms, input the parameters of the machine learning algorithms, and use grid filtering to train and test the training set and the test set obtained in step 3) respectively, thereby establishing multiple prediction models accordingly; 5) Determine the optimal prediction model: Compare the predicted values and actual values of the multiple prediction models obtained in step 4), calculate the relative error between the two, and determine the prediction model with the smallest relative error as the optimal prediction model. 6) Alloy composition design: The optimal prediction model determined in step 5) is set as the fitness function, and the genetic algorithm is used to optimize the alloy composition with the hydrogen storage capacity of the hydrogen storage alloy described in step 1) so as to achieve rapid and accurate design of the hydrogen storage alloy composition.
2. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, In step 1), the data sources for the hydrogen storage capacity and alloy composition of the hydrogen storage alloy include existing technical literature, experimental results, or tool manuals.
3. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, In step 1), when acquiring data, the test temperature for hydrogen storage should be kept near room temperature to eliminate the influence of temperature on hydrogen storage, and the collected data should be used as the original dataset for subsequent modeling.
4. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, In step 1), the hydrogen storage alloy includes rare earth AB5 type, titanium AB type, AB2 type, vanadium solid solution type, and magnesium-based hydrogen storage alloy.
5. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, In step 4), the machine learning algorithms include linear regression, support vector machine, random forest, and XGboost.
6. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, In step 4), the parameters of the linear regression algorithm include the penalty term alpha; the parameters of the support vector machine include the regularization term C, the kernel function name kernel, the margin width epsilon, and the kernel function parameter degree; the parameters of the random forest include the maximum depth max_depth and the number of decision trees n_estimators; the parameters of the Xgboost algorithm are the maximum depth max_depth, the minimum loss reduction required for a node to branch gamma, the regularization coefficient, and the number of decision trees n_estimators.
7. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, In step 5), the optimal prediction model is Xgboost, and its optimal parameter combination is {max_depth=6,gamma=0.5,alpha=1,n_estimators=200}.
8. The design method for a high-capacity hydrogen storage alloy as described in claim 1, characterized in that, The method further includes: 7) Based on the hydrogen storage alloy composition obtained in step 6), prepare a hydrogen storage alloy sample and test its actual hydrogen storage capacity. Compare it with the predicted hydrogen storage capacity of the hydrogen storage alloy. If it does not meet the preset error requirements, adjust the model parameters in step 4) and repeat steps 4)-7) until it meets the preset error requirements.
Citation Information
Patent Citations
Hydrogen storage alloy performance prediction method, prediction model thereof and model establishment method
CN114724655A
Method, device and system for determining optimal alloy performance and preparation strategy
CN115274020A
Cited By
Hydrogen storage alloy performance prediction method based on feature engineering and multi-model screening
CN122290835A