A high-entropy alloy machine learning phase prediction method and device for extracting features by combining an empirical parameter and a convolutional neural network with the periodic table of elements
By combining empirical parameters and convolutional neural networks, using the periodic table of elements to extract features, and through feature screening and genetic algorithm optimization, the traditional machine learning model was finally used to predict the phase composition of high-entropy alloys, which solved the problems of poor prediction performance and feature redundancy in the existing technology, and achieved higher prediction accuracy.
Patent Information
- Application Number
- CN202310317884.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-03-28
AI Technical Summary
The prior art has poor prediction performance when predicting the composition of high-entropy alloys, and the automatically extracted features lack chemical and physical significance, too many feature dimensions, and redundant information.
Combining empirical parameters and convolutional neural networks, the features are extracted using the periodic table of elements, the optimal feature combination is screened through five-fold cross-validation and feature engineering methods, and feature screening is performed with genetic algorithms, and finally the traditional machine learning model is used for prediction.
The accuracy of prediction of phase composition of high-entropy alloys is improved, and the information extracted by domain knowledge and machine learning is effectively combined, redundant information is removed, and complementary information is mined to obtain the optimal feature combination.
Smart Images

Figure CN116364211B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-entropy alloy phase prediction. Specifically, it relates to a machine learning phase prediction method and device for high-entropy alloys that extracts features by combining empirical parameters and a convolutional neural network with the periodic table of elements. Background Art
[0002] High-entropy alloys (HEAs) are multi-principal element alloys that contain five or more elements mixed in equal or nearly equal atomic percentages. Since the design concept of high-entropy alloys was introduced, it has become increasingly popular and is a promising alloy family. High-entropy alloys have a huge composition space, which makes it possible for high-entropy alloys to obtain unique properties. Machine learning is very suitable for designing high-entropy alloys and has achieved good results. Generally, empirical parameters are used as the input of machine learning. The empirical parameters depend on element properties such as valence electrons, atomic size, electronegativity, melting point, mixing entropy, and mixing enthalpy. These empirical parameters are manually constructed according to formulas. The setting of such empirical parameters requires reliance on relevant domain knowledge and is too broad, making it difficult to comprehensively reflect the various different properties of elements in the alloy. In addition, there is another input method that can avoid manually constructing features. Only the element composition and element concentration need to be known. According to the element composition and element concentration, the alloy sample is mapped into a unique two-dimensional matrix in the form of a periodic table of elements, and then a convolutional neural network is constructed for automatic feature extraction and prediction of the target attribute. This method can directly learn the properties of elements from the periodic table of elements and generate features pointing to the prediction of the target attribute. However, the automatically extracted features currently cannot be given chemical and physical meanings and cannot be used to analyze the internal relationship between them and the target attribute. Moreover, the feature dimension is too large and there is information redundancy, which should be reduced when combined with other machine learning models for prediction. And the manually generated features contain human understanding of the phase composition of high-entropy alloys. The empirical parameters should be organically combined with the automatically extracted features. Summary of the Invention
[0003] The technical problem to be solved by the present invention is:
[0004] To solve the problem of poor prediction performance in predicting the phase composition of high-entropy alloys in the prior art.
[0005] The technical solution adopted by the present invention to solve the above technical problem:
[0006] On the one hand, the present invention proposes a machine learning phase prediction method for high-entropy alloys that extracts features by combining empirical parameters and a convolutional neural network with the periodic table of elements, including the following steps:
[0007] Step 1: Collect high-entropy alloy phase classification data. The data includes the composition information of each high-entropy alloy, which consists of elemental composition and elemental concentration, as well as the phase label of each high-entropy alloy. Then, calculate the empirical parameters according to the formula based on the composition information, and use these data as the original dataset. Adopt five-fold cross-validation to divide the dataset into five groups of 80% training sets and 20% test sets;
[0008] Step 2: Use the empirical parameters as input to construct five traditional machine learning models. Adopt three feature engineering methods to screen features, predict the test set, and obtain the optimal feature combination and optimal model with the highest accuracy according to the prediction results;
[0009] Step 3: Map each high-entropy alloy data into a unique two-dimensional pseudo-image in the form of the periodic table according to the high-entropy alloy composition information;
[0010] First, assign a 9*18 matrix to each high-entropy alloy, corresponding to the periodic table. The positions in the matrix where there is no element distribution in the periodic table are set to 0; then, fill in the element concentrations at the corresponding positions in the matrix according to the positions of the elements in the periodic table in each high-entropy alloy; finally, set the positions corresponding to other elements in the matrix to 0;
[0011] Step 4: Build a convolutional neural network model, train it in the training set, automatically extract features of the input high-entropy alloy, and predict the high-entropy alloy phase composition;
[0012] Step 5: Combine all the empirical parameters and the features automatically extracted in Step 4, adopt the optimal model described in Step 2, combine with the genetic algorithm for feature screening, select the optimal feature combination, and finally use the optimal feature combination as the input feature, and retrain and predict the high-entropy alloy phase composition using the optimal model in Step 2.
[0013] Furthermore, the high-entropy alloy described in Step 1 contains fifty-three elements, namely Al, Cr, Fe, Co, Ni, Cu, Mn, Ti, Zr, Nb, Mo, Hf, V, Ta, C, Sn, W, Re, Si, Pd, N, B, Ce, Ag, Pt, Au, Zn, Ge, Mg, Rh, Ir, Be, Li, Y, Nd, Cd, Ca, In, Sb, Bi, Na, Ru, Gd, Tb, Dy, Er, Sr, Yb, P, La, Ho, Pr, Sc elements, which is an alloy dataset without limited systems.
[0014] Furthermore, the high-entropy alloy described in Step 1 contains four categories of phase labels, namely solid solution phase (SS phase), intermetallic compound phase (IM phase), amorphous phase (AM phase), and solid solution and intermetallic compound mixed phase (SS+IM phase).
[0015] Further, the empirical parameters in Step 1 include: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), valence electron concentration variance (δVEC), mixing enthalpy (ΔH mix ), four mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ), numbered as Feature 0, 1, 2, 3, ……, 17 respectively.
[0016] Further, the five traditional machine learning models in Step 2 are Support Vector Machine (SVC), Random Forest (RF), K-Nearest Neighbor (KNN), Logistic Regression (LR), and Decision Tree (DT).
[0017] Further, the three feature engineering methods in Step 2 include: recursive feature elimination method based on random forest feature importance ranking, Pearson correlation coefficient method, and principal component analysis method.
[0018] Further, the optimal model selected in Step 2 is Random Forest (RF); the optimal feature combination includes: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), mixing enthalpy (ΔH mix ), three mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ).
[0019] Further, the convolutional neural network model structure described in step four includes: 3 convolutional layers, 3 pooling layers, a flattening layer, and an output layer; the convolutional kernel size of the convolutional layer is 3*3, the number of convolutional kernels is 8, 16, and 32 respectively, the stride is 1, and each convolutional layer is followed by a non-linear activation layer ReLU layer; the pooling layer is a max pooling layer, the pooling window is 2*2, and the stride is 2; the flattening layer contains 192 neurons, the output layer contains 4 neurons, and the Softmax activation function is used; in backpropagation, the sparse categorical cross-entropy loss function (sparse_categorical_crossentropy) is used, and the AdamOptimizer optimizer is used to adjust the weights; the automatic feature extraction is to input the two-dimensional pseudo-image corresponding to the high-entropy alloy data into the trained convolutional neural network model, and extract the 192-dimensional vector output in the flattening layer, which is numbered as feature 18, 19, 20, 21, ……, 208, 209 respectively.
[0020] Further, the selected optimal feature combination numbers in step five are feature 0, 1, 4, 8, 9, 10, 17, 18, 21, 29, 47, 48, 54, 84, 85, 88, 90, 96, 101, 102, 107, 128, 129, 139, 142, 161, 171, 174, 182, 184, 200, 202, 204.
[0021] A high-entropy alloy machine learning phase prediction device for extracting features by combining empirical parameters and a convolutional neural network with the periodic table of elements, the device includes:
[0022] An acquisition unit, used to collect high-entropy alloy phase classification data, the data contains the composition information of each high-entropy alloy composed of elements and element concentrations and the phase label of each high-entropy alloy, and then calculate the empirical parameters according to the formula from the composition information, use these data as the original data set, and use five-fold cross-validation to divide the data set into five groups of 80% training set and 20% test set;
[0023] A machine learning model selection unit, used to select the optimal machine learning model;
[0024] Using the empirical parameters as input, construct five traditional machine learning models, use three feature engineering methods to screen features, predict the test set, and obtain the optimal feature combination and the optimal model according to the prediction results;
[0025] A feature extraction unit, used to automatically extract features from the two-dimensional matrix in the form of the periodic table of elements mapped from the composition information;
[0026] First, map each high-entropy alloy data into a unique two-dimensional pseudo-image in the form of the periodic table according to the high-entropy alloy composition information; then build a convolutional neural network model and train it. Then, input the two-dimensional pseudo-image corresponding to the high-entropy alloy data into the trained convolutional neural network model, and extract the 192-dimensional vector output in the flatten layer as the features extracted by the convolutional neural network;
[0027] A feature screening unit for combining the empirical parameters and the features extracted by the convolutional neural network and screening out the optimal feature combination;
[0028] A machine learning model training unit for inputting the optimal feature combination obtained in the feature screening unit into the optimal machine learning model obtained in the machine learning model selection unit to retrain the model and obtain a high-entropy alloy phase composition prediction model;
[0029] A phase composition detection unit for inputting the sample to be detected into the high-entropy alloy phase composition prediction model obtained in the machine learning model training unit for phase composition prediction.
[0030] Furthermore, the high-entropy alloy obtained by the acquisition unit contains fifty-three elements, namely Al, Cr, Fe, Co, Ni, Cu, Mn, Ti, Zr, Nb, Mo, Hf, V, Ta, C, Sn, W, Re, Si, Pd, N, B, Ce, Ag, Pt, Au, Zn, Ge, Mg, Rh, Ir, Be, Li, Y, Nd, Cd, Ca, In, Sb, Bi, Na, Ru, Gd, Tb, Dy, Er, Sr, Yb, P, La, Ho, Pr, Sc elements, which is an alloy data set without limited systems; the phase labels of the acquisition unit include solid solution phase (SS phase), intermetallic compound phase (IM phase), amorphous phase (AM phase) and solid solution and intermetallic compound mixed phase (SS+IM phase); the empirical parameters of the acquisition unit include: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), valence electron concentration variance (δVEC), mixing enthalpy (ΔH mix ), four mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), the combined effect parameter (Ω) of mixing enthalpy and mixing entropy, average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ) and local atomic distortion (a 2 ), numbered as feature 0, 1, 2, 3, ……, 17 respectively.
[0031] Further, the five traditional machine learning models in the machine learning model selection unit are Support Vector Machine (SVC), Random Forest (RF), K-Nearest Neighbor (KNN), Logistic Regression (LR), and Decision Tree (DT); the three feature engineering methods in the machine learning model selection unit include: Recursive Feature Elimination method based on the importance ranking of random forest features, Pearson correlation coefficient method, and Principal Component Analysis method; the optimal feature combination in the machine learning model selection unit includes: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), mixing enthalpy (ΔH mix ), three mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ); the optimal model in the machine learning model selection unit is Random Forest (RF).
[0032] Further, the convolutional neural network model structure in the feature extraction unit includes: 3 convolutional layers, 3 pooling layers, a flattening layer, and an output layer; the convolutional kernel size of the convolutional layer is 3*3, the number of convolutional kernels is 8, 16, and 32 respectively, the stride is 1, and each convolutional layer is followed by a non-linear activation layer ReLU layer; the pooling layer is max pooling, the pooling window is 2*2, and the stride is 2; the flattening layer contains 192 neurons, the output layer contains 4 neurons, and the Softmax activation function is used; in backpropagation, the sparse categorical cross-entropy loss function (sparse_categorical_crossentropy) is used, and the AdamOptimizer optimizer is used to adjust the weights; the number of features automatically extracted by the convolutional neural network is 192, and they are numbered as feature 18, 19, 20, 21, ……, 208, 209 respectively.
[0033] Further, the optimal feature combination numbers in the feature screening unit are feature 0, 1, 4, 8, 9, 10, 17, 18, 21, 29, 47, 48, 54, 84, 85, 88, 90, 96, 101, 102, 107, 128, 129, 139, 142, 161, 171, 174, 182, 184, 200, 202, 204.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] The present invention relates to a method and device for predicting the phases of high-entropy alloys by extracting features based on empirical parameters and a convolutional neural network combined with the periodic table of elements. First, empirical parameters related to the phase classification of high-entropy alloys are constructed based on empirical knowledge, and the optimal features and the optimal machine learning model are selected based on three feature engineering methods. Then, features are automatically extracted from two-dimensional images using a combination of the periodic table of elements and a convolutional neural network. The generated features are information generated with respect to the target attributes, which contain the knowledge in the periodic table of elements and the distribution information of the types and concentrations of alloying elements. Then, for the feature pool composed of all the empirical parameters and the extracted features, genetic algorithm feature screening is performed based on the selected optimal model, so that the domain knowledge and the information extracted by machine learning can be organically combined, the redundant parts in the information of both are removed, and the complementary information of both is mined, thereby obtaining the optimal feature combination. Based on the finally obtained feature combination, the accuracy of predicting the phase composition of high-entropy alloys is improved using traditional machine learning algorithms. This method is superior to the method of only using empirical parameters or only using the representation method of the periodic table of elements mapped by compositional information.
[0036] By predicting the phase composition of high-entropy alloys using the method of the present invention, the performance of the machine learning model in predicting the phase composition of high-entropy alloys can be effectively improved, which has good application prospects and actively promotes the research on predicting the phase composition of high-entropy alloys.
[0037] The method of the present invention has universality. For the feature extraction and input problems in the research of other properties in the field of high-entropy alloys or even other material fields, the work can be carried out through the idea of the present invention to improve the performance of the machine learning model in predicting the properties of materials. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of a two-dimensional matrix in the form of a periodic table of elements mapped by compositional information in an embodiment of the present invention;
[0039] Figure 2 It is a schematic diagram for comparing the prediction performance of a machine learning model for the phase composition of high-entropy alloys under three input features in an embodiment of the present invention;
[0040] Figure 3 It is a schematic structural diagram of a device for predicting the phases of high-entropy alloys by extracting features based on empirical parameters and a convolutional neural network combined with the periodic table of elements in an embodiment of the present invention. Detailed Embodiments
[0041] In the description of the present invention, it should be noted that in the embodiments of the present invention, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.
[0042] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following detailed description of the specific embodiments of the present invention will be provided in conjunction with the accompanying drawings.
[0043] Specific Embodiment 1: This embodiment provides a method for predicting the phases of high-entropy alloys in machine learning by combining empirical parameters and a convolutional neural network to extract features from the periodic table of elements, including the following steps:
[0044] Step 1: Collect high-entropy alloy phase classification data, which includes the composition information of each high-entropy alloy consisting of element composition and element concentration, as well as the phase labels of each high-entropy alloy. Then, calculate the empirical parameters according to the formula from the composition information, and use this data as the original data set. Adopt five-fold cross-validation to divide the data set into five groups of 80% training sets and 20% test sets;
[0045] Among them, the high-entropy alloy contains fifty-three elements, namely Al, Cr, Fe, Co, Ni, Cu, Mn, Ti, Zr, Nb, Mo, Hf, V, Ta, C, Sn, W, Re, Si, Pd, N, B, Ce, Ag, Pt, Au, Zn, Ge, Mg, Rh, Ir, Be, Li, Y, Nd, Cd, Ca, In, Sb, Bi, Na, Ru, Gd, Tb, Dy, Er, Sr, Yb, P, La, Ho, Pr, Sc elements, which is an alloy data set without system limitation;
[0046] Among them, the high-entropy alloy contains four types of phase labels, namely solid solution phase (SS phase), intermetallic compound phase (IM phase), amorphous phase (AM phase), and mixed phase of solid solution and intermetallic compound (SS+IM phase);
[0047] Among them, the empirical parameters include: average atomic radius (r), variance of atomic radius (δr), valence electron concentration (VEC), variance of valence electron concentration (δVEC), mixing enthalpy (ΔH mix ), four variances of mixing enthalpy (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), variance of electronegativity (δχ), average melting point (T), variance of melting point (δT), two geometric parameters (λ, γ), and local atomic distortion (a2 ) are numbered as feature 0, 1, 2, 3, ……, 17 respectively;
[0048] Step 2: Use empirical parameters as input, construct five traditional machine learning models, use three feature engineering methods to screen features, predict the test set, and obtain the optimal feature combination and optimal model with the highest accuracy according to the prediction results;
[0049] Among them, the five traditional machine learning models are Support Vector Machine (SVC), Random Forest (RF), K-Nearest Neighbor (KNN), Logistic Regression (LR) and Decision Tree (DT);
[0050] Among them, the three feature engineering methods include: recursive feature elimination method based on random forest feature importance ranking, Pearson correlation coefficient method and principal component analysis method;
[0051] Among them, the optimal model selected is Random Forest (RF). The optimal feature combination includes: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), mixing enthalpy (ΔH mix ), three mixing enthalpy variances (δH mix , ) mixing entropy (ΔS mix ), the combined effect parameter (Ω) of mixing enthalpy and mixing entropy, average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ) and local atomic distortion (a 2 );
[0052] Step 3: Map each high-entropy alloy data into a unique two-dimensional pseudo-image in the form of the periodic table according to the high-entropy alloy composition information;
[0053] First, assign a 9*18 matrix to each high-entropy alloy, corresponding to the periodic table, and set the positions without element distribution in the periodic table in the matrix to 0; then, fill the element concentrations into the corresponding positions in the matrix according to the positions of the elements in the periodic table in each high-entropy alloy; finally, set the corresponding positions of other elements in the matrix to 0;
[0054] Step 4: Build a convolutional neural network model, train it in the training set, automatically extract features of the input high-entropy alloy and predict the high-entropy alloy phase composition;
[0055] Among them, the convolutional neural network model structure includes: 3 convolutional layers, 3 pooling layers, a flattening layer, and an output layer; the convolutional kernel size of the convolutional layer is 3*3, the number of convolutional kernels is 8, 16, and 32 respectively, the stride is 1, and each convolutional layer is followed by a non-linear activation layer ReLU layer; the pooling layer is a max pooling layer, the pooling window is 2*2, and the stride is 2; the flattening layer contains 192 neurons, the output layer contains 4 neurons, and the Softmax activation function is used. In backpropagation, the sparse categorical cross-entropy loss function (sparse_categorical_crossentropy) is used, and the AdamOptimizer optimizer is used to adjust the weights;
[0056] Among them, the automatic feature extraction is to input the two-dimensional pseudo-image corresponding to the high-entropy alloy data into the trained convolutional neural network model, and extract the 192-dimensional vector output in the flattening layer, and number them as feature 18, 19, 20, 21,..., 208, 209 respectively;
[0057] Step 5: Combine all the empirical parameters and the features automatically extracted in Step 4, adopt the optimal model described in Step 2, combine with the genetic algorithm for feature screening, select the optimal feature combination, and finally use the optimal feature combination as the input feature, and adopt the optimal model in Step 2 to retrain and predict the high-entropy alloy phase composition;
[0058] Among them, the selected optimal feature combination numbers are feature 0, 1, 4, 8, 9, 10, 17, 18, 21, 29, 47, 48, 54, 84, 85, 88, 90, 96, 101, 102, 107, 128, 129, 139, 142, 161, 171, 174, 182, 184, 200, 202, 204.
[0059] Example 1
[0060] By comparing the prediction performance of the machine learning model for the high-entropy alloy phase composition under the same test set when only using empirical parameters, only using the features automatically extracted by the convolutional neural network, and using the combined features of the features extracted by the convolutional neural network and empirical parameters, the effectiveness of the method of the present invention is further verified.
[0061] Step 1: Using the 1130 high-entropy alloy composition and phase composition data in Reference [1], 18 empirical parameters are calculated from the high-entropy alloy composition information. Using five-fold cross-validation, it is divided into five groups of 80% training set and 20% test set; Step 2: Construct five traditional machine learning models, using the empirical parameters as input, and adopting three feature engineering methods to select the optimal feature combination and the optimal model; Step 3: Map each high-entropy alloy data into a unique two-dimensional pseudo-image in the form of the periodic table according to the high-entropy alloy composition information, construct a convolutional neural network to automatically extract features, train on the training set, and test on the test set; Step 4: Combine the empirical parameters and the features automatically extracted in Step 3, and screen out the optimal feature combination through a genetic algorithm. Use this optimal feature combination as the input of the optimal model in Step 2, retrain and make predictions; Step 5: Compare the prediction effects of the optimal model in Step 2, the convolutional neural network in Step 3, and the optimal model in Step 4.
[0062] All three methods are trained on the same training set and evaluated on the same test set to ensure fairness, and five-fold cross-validation is performed to prevent overfitting. The average accuracy on five test sets is used as the evaluation index.
[0063] From Figure 2 It can be seen that the model of the method of the present invention has a higher accuracy rate than the models that only use empirical parameters and the features extracted by only using a convolutional neural network, indicating that the model obtained by using the method of the present invention has better performance, and the results are significantly better than the model performances obtained by the existing two feature input methods.
[0064] Example 2
[0065] This embodiment provides a high-entropy alloy machine learning phase prediction device for extracting features by combining empirical parameters and a convolutional neural network with the periodic table, as Figure 3 shown. The device includes:
[0066] An acquisition unit, used to collect high-entropy alloy phase classification data, where the data includes the composition information of each high-entropy alloy composed of elements and element concentrations and the phase label of each high-entropy alloy. Then, empirical parameters are calculated according to the formula from the composition information, and these data are used as the original data set. Using five-fold cross-validation, the data set is divided into five groups of 80% training set and 20% test set;
[0067] A machine learning model selection unit, used to select the optimal machine learning model;
[0068] Using the empirical parameters as input, construct five traditional machine learning models, use three feature engineering methods to screen features, predict the phase classification, and obtain the optimal feature combination and the optimal model according to the prediction results;
[0069] A feature extraction unit for automatically extracting features from a two-dimensional matrix in the form of a periodic table of elements mapped from composition information.
[0070] First, map each high-entropy alloy data into a unique two-dimensional pseudo-image in the form of a periodic table of elements according to the high-entropy alloy composition information; then build and train a convolutional neural network model, and then input the two-dimensional pseudo-image corresponding to the high-entropy alloy data into the trained convolutional neural network model, and extract the 192-dimensional vector output in the flatten layer as the features extracted by the convolutional neural network.
[0071] A feature screening unit for combining the empirical parameters and the features extracted by the convolutional neural network and screening out the optimal feature combination.
[0072] A machine learning model training unit for inputting the optimal feature combination obtained in the feature screening unit into the optimal machine learning model obtained in the machine learning model selection unit to retrain the model and obtain the final high-entropy alloy phase composition prediction model.
[0073] A phase composition detection unit for inputting the sample to be detected into the high-entropy alloy phase composition prediction model obtained in the machine learning model training unit for phase composition prediction.
[0074] Furthermore, the high-entropy alloy obtained by the acquisition unit contains fifty-three elements, namely Al, Cr, Fe, Co, Ni, Cu, Mn, Ti, Zr, Nb, Mo, Hf, V, Ta, C, Sn, W, Re, Si, Pd, N, B, Ce, Ag, Pt, Au, Zn, Ge, Mg, Rh, Ir, Be, Li, Y, Nd, Cd, Ca, In, Sb, Bi, Na, Ru, Gd, Tb, Dy, Er, Sr, Yb, P, La, Ho, Pr, Sc elements, which is an alloy data set without limited system; the phase labels obtained by the acquisition unit include solid solution phase (SS phase), intermetallic compound phase (IM phase), amorphous phase (AM phase) and solid solution and intermetallic compound mixed phase (SS + IM phase); the empirical parameters obtained by the acquisition unit include: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), valence electron concentration variance (δVEC), mixing enthalpy (ΔH mix ), four mixing enthalpy variances (δH mix ), ), mixing entropy (ΔS mix ), the combined effect parameter (Ω) of mixing enthalpy and mixing entropy, average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ) and local atomic distortion (a 2 ), numbered as feature 0, 1, 2, 3, ……, 17 respectively.
[0075] 2. Further, the five traditional machine learning models in the machine learning model selection unit are Support Vector Machine (SVC), Random Forest (RF), K-Nearest Neighbor (KNN), Logistic Regression (LR), and Decision Tree (DT); the three feature engineering methods in the machine learning model selection unit include: Recursive Feature Elimination method based on the importance ranking of random forest features, Pearson correlation coefficient method, and Principal Component Analysis method; the optimal feature combination in the machine learning model selection unit includes: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), mixing enthalpy (ΔH mix ), three mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ); the optimal model in the machine learning model selection unit is Random Forest (RF).
[0076] Further, the convolutional neural network model structure in the feature extraction unit includes: 3 convolutional layers, 3 pooling layers, a flattening layer, and an output layer; the convolutional kernel size of the convolutional layer is 3*3, the number of convolutional kernels are 8, 16, and 32 respectively, the stride is 1, and each convolutional layer is followed by a non-linear activation layer ReLU layer; the pooling layer is max pooling, the pooling window is 2*2, and the stride is 2; the flattening layer contains 192 neurons, the output layer contains 4 neurons, and the Softmax activation function is used; in backpropagation, the sparse categorical cross-entropy loss function (sparse_categorical_crossentropy) is used, and the AdamOptimizer optimizer is used for weight adjustment; the number of features automatically extracted by the convolutional neural network is 192, and they are numbered as feature 18, 19, 20, 21, ……, 208, 209 respectively.
[0077] Further, the optimal feature combination numbers in the feature screening unit are feature 0, 1, 4, 8, 9, 10, 17, 18, 21, 29, 47, 48, 54, 84, 85, 88, 90, 96, 101, 102, 107, 128, 129, 139, 142, 161, 171, 174, 182, 184, 200, 202, 204.
[0078] The function of the high-entropy alloy machine learning phase prediction device based on the combination of empirical parameters and convolutional neural network to extract features described in this embodiment can be illustrated by the aforementioned high-entropy alloy machine learning phase prediction method based on the combination of empirical parameters and convolutional neural network to extract features. Therefore, for the parts not detailed in this embodiment, reference can be made to the above method embodiments and will not be elaborated here.
[0079] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.
[0080] The literature cited in the present invention is as follows: [1] Han, Q.A.; Lu, Z.L.; Cui, H.T. Data-driven based phase constitution prediction in high entropy alloys. Comput. Mater. Sci. 2022, 215, 111774.
Claims
1. A method for predicting high-entropy alloy machine learning phases by extracting features based on the combination of empirical parameters and convolutional neural networks with the periodic table of elements, characterized in that it includes the following steps: Step 1: Collect high-entropy alloy phase classification data, which contains the composition information of each high-entropy alloy composed of elements and element concentrations, as well as the phase labels of each high-entropy alloy. Then, calculate the empirical parameters according to the formula from the composition information. Take these data as the original data set and use five-fold cross-validation to divide the data set into five groups of 80% training sets and 20% test sets; Step 2: Use the empirical parameters as input, construct five traditional machine learning models, use three feature engineering methods to screen features, predict the test set, and obtain the optimal feature combination and the optimal model with the highest accuracy according to the prediction results; Step 3: Map each high-entropy alloy data into a unique two-dimensional pseudo-image in the form of the periodic table of elements according to the high-entropy alloy composition information; First, assign a 9*18 matrix to each high-entropy alloy, corresponding to the periodic table of elements. The positions in the matrix where there is no element distribution in the periodic table of elements are set to 0; then, fill in the element concentrations at the corresponding positions in the matrix according to the positions of the elements in each high-entropy alloy in the periodic table of elements; finally, set the positions corresponding to other elements in the matrix to 0; Step 4: Build a convolutional neural network model, train it in the training set, automatically extract features of the input high-entropy alloy, and predict the phase composition of the high-entropy alloy; Step 5: Combine all the empirical parameters and the features automatically extracted in Step 4, use the optimal model described in Step 2, combine with the genetic algorithm for feature screening, select the optimal feature combination, and finally use the optimal feature combination as the input feature, and retrain and predict the phase composition of the high-entropy alloy using the optimal model in Step 2.
2. The method according to claim 1, characterized in that: The high-entropy alloy described in Step 1 contains fifty-three elements, namely Al, Cr, Fe, Co, Ni, Cu, Mn, Ti, Zr, Nb, Mo, Hf, V, Ta, C, Sn, W, Re, Si, Pd, N, B, Ce, Ag, Pt, Au, Zn, Ge, Mg, Rh, Ir, Be, Li, Y, Nd, Cd, Ca, In, Sb, Bi, Na, Ru, Gd, Tb, Dy, Er, Sr, Yb, P, La, Ho, Pr, Sc elements, which is an alloy data set without system limitations; the high-entropy alloy described in Step 1 contains four types of phase tags, namely solid solution phase (SS phase), intermetallic compound phase (IM phase), amorphous phase (AM phase), and solid solution and intermetallic compound mixed phase (SS+IM phase); the empirical parameters described in Step 1 include: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), valence electron concentration variance (δVEC), mixing enthalpy (ΔH mix ), four mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), the combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ), numbered as feature 0, 1, 2, 3, ……, 17 respectively.
3. The method according to claim 1, characterized in that: The five traditional machine learning models described in Step 2 are Support Vector Machine (SVC), Random Forest (RF), K-Nearest Neighbor (KNN), Logistic Regression (LR), and Decision Tree (DT); the three feature engineering methods described in Step 2 include: Recursive Feature Elimination method based on the importance ranking of Random Forest features, Pearson correlation coefficient method, and Principal Component Analysis method; the optimal model obtained in Step 2 is Random Forest (RF); the optimal feature combination described in Step 2 includes: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), mixing enthalpy (ΔH mix ), three mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ).
4. The method according to claim 1, characterized in that: The convolutional neural network model structure described in Step 4 includes: 3 convolutional layers, 3 pooling layers, a flattening layer, and an output layer; the convolutional kernel size of the convolutional layer is 3*3, the number of convolutional kernels is 8, 16, 32 respectively, the stride is 1, and each convolutional layer is followed by a non-linear activation layer ReLU layer; the pooling layer is max pooling, the pooling window is 2*2, and the stride is 2; the flattening layer contains 192 neurons, the output layer contains 4 neurons, and the Softmax activation function is used; in backpropagation, the sparse categorical cross-entropy loss function (sparse_categorical_crossentropy) is used, and the AdamOptimizer optimizer is used to adjust the weights; the automatic feature extraction described in Step 4 is to input the two-dimensional pseudo-image corresponding to the high-entropy alloy data into the trained convolutional neural network model, and extract the 192-dimensional vector output in the flattening layer, and number them as feature 18, 19, 20, 21,..., 208, 209 respectively.
5. The method according to claim 1, characterized in that: The optimal feature combination numbers selected in step five are feature 0, 1, 4, 8, 9, 10, 17, 18, 21, 29, 47, 48, 54, 84, 85, 88, 90, 96, 101, 102, 107, 128, 129, 139, 142, 161, 171, 174, 182, 184, 200, 202, 204.
6. A high-entropy alloy machine learning phase prediction device for extracting features by combining empirical parameters and a convolutional neural network with the periodic table of elements, characterized in that the device includes: An acquisition unit, configured to collect high-entropy alloy phase classification data, where the data includes the composition information of each high-entropy alloy composed of elements and element concentrations, and the phase label of each high-entropy alloy. Then, according to the composition information, empirical parameters are calculated according to the formula, and these data are used as the original data set. Using five-fold cross-validation, the data set is divided into five groups of 80% training sets and 20% test sets; A machine learning model selection unit, configured to select the optimal machine learning model; Using the empirical parameters as input, five traditional machine learning models are constructed, three feature engineering methods are used to screen features, the test set is predicted, and according to the prediction results, the optimal feature combination and the optimal model are obtained; A feature extraction unit, configured to automatically extract features from the two-dimensional matrix in the form of the periodic table of elements mapped from the composition information; First, according to the high-entropy alloy composition information, each high-entropy alloy data is mapped into a unique two-dimensional pseudo-image in the form of the periodic table of elements; then a convolutional neural network model is built and trained, and then the two-dimensional pseudo-image corresponding to the high-entropy alloy data is input into the trained convolutional neural network model, and the 192-dimensional vector output in the flatten layer is extracted as the features extracted by the convolutional neural network; A feature screening unit, configured to combine the features extracted by the empirical parameters and the convolutional neural network and screen out the optimal feature combination; A machine learning model training unit, configured to input the optimal feature combination obtained in the feature screening unit into the optimal machine learning model obtained in the machine learning model selection unit, and retrain the model to obtain the final high-entropy alloy phase composition prediction model; A phase composition detection unit, configured to input the sample to be detected into the high-entropy alloy phase composition prediction model obtained in the machine learning model training unit for phase composition prediction.
7. The device according to claim 6, characterized in that: The high-entropy alloy described in the acquisition unit contains fifty-three elements, namely Al, Cr, Fe, Co, Ni, Cu, Mn, Ti, Zr, Nb, Mo, Hf, V, Ta, C, Sn, W, Re, Si, Pd, N, B, Ce, Ag, Pt, Au, Zn, Ge, Mg, Rh, Ir, Be, Li, Y, Nd, Cd, Ca, In, Sb, Bi, Na, Ru, Gd, Tb, Dy, Er, Sr, Yb, P, La, Ho, Pr, Sc, which is an alloy dataset without restricted systems; the phase tags described in the acquisition unit include solid solution phase (SS phase), intermetallic compound phase (IM phase), amorphous phase (AM phase), and solid solution and intermetallic compound mixed phase (SS+IM phase); The empirical parameters of the acquisition unit include: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), valence electron concentration variance (δVEC), mixing enthalpy (ΔH mix ), four mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ) and local atomic distortion (a 2 ), numbered as feature 0, 1, 2, 3, ……, 17 respectively.
8. The device according to claim 6, wherein: the five traditional machine learning models in the machine learning model selection unit are Support Vector Machine (SVC), Random Forest (RF), K-Nearest Neighbor (KNN), Logistic Regression (LR), and Decision Tree (DT); The three feature engineering methods in the machine learning model selection unit include: recursive feature elimination method based on the importance ranking of random forest features, Pearson correlation coefficient method, and principal component analysis method; the optimal feature combination in the machine learning model selection unit includes: average atomic radius (r), atomic radius variance (δr), valence electron concentration (VEC), mixing enthalpy (ΔH mix ), three mixing enthalpy variances (δH mix , ), mixing entropy (ΔS mix ), combined effect parameter of mixing enthalpy and mixing entropy (Ω), average electronegativity (χ), electronegativity variance (δχ), average melting point (T), melting point variance (δT), two geometric parameters (λ, γ), and local atomic distortion (a 2 ); the optimal model in the machine learning model selection unit is random forest (RF).
9. The device according to claim 6, wherein: the convolutional neural network model structure in the feature extraction unit includes: 3 convolutional layers, 3 pooling layers, a flattening layer, and an output layer; the convolutional kernel size of the convolutional layers is 3*3, the number of convolutional kernels is 8, 16, 32 respectively, the stride is 1, and each convolutional layer is followed by a non-linear activation layer ReLU layer; the pooling layers are max pooling, the pooling window is 2*2, and the stride is 2; the flattening layer contains 192 neurons, the output layer contains 4 neurons, and the Softmax activation function is used; in backpropagation, the sparse categorical cross-entropy loss function (sparse_categorical_crossentropy) is used, and the AdamOptimizer optimizer is used for weight adjustment; the number of features automatically extracted by the convolutional neural network in the feature extraction unit is 192, numbered as feature 18, 19, 20, 21, ……, 208, 209 respectively.
10. The device according to claim 6, wherein: the numbers of the optimal feature combinations in the feature screening unit are feature 0, 1, 4, 8, 9, 10, 17, 18, 21, 29, 47, 48, 54, 84, 85, 88, 90, 96, 101, 102, 107, 128, 129, 139, 142, 161, 171, 174, 182, 184, 200, 202, 204.
Citation Information
Patent Citations
Laser radar three-dimensional target rapid detection method based on pseudo-image technology
CN111242041A
Method for realizing a multi-channel convolutional recurrent neural network EEG emotion recognition model using transfer learning
US20230039900A1