Method for predicting phase forming ability based on machine learning of unbalanced data of high-entropy carbide ceramics
By oversampling and feature engineering of the high-entropy carbide ceramic dataset and optimizing the prediction model with multiple machine learning models, the data imbalance problem is solved, the accuracy and efficiency of high-entropy carbide ceramic phase formation prediction is improved, and the optimal components and concentration combinations are achieved quickly screened.
Patent Information
- Application Number
- CN202510477194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has data imbalance in the phase formation prediction of high-entropy carbide ceramics, which leads to low prediction accuracy and poor generalization capabilities of machine learning models, making it difficult to efficiently screen out the optimal components and concentration combinations of single-phase high-entropy ceramic materials that meet specific needs.
The high-entropy carbide ceramic dataset is oversampled by unbalanced learning strategy, the best feature subset is obtained through feature engineering processing, and a variety of machine learning models are used to train and experimentally verify the optimized prediction model to build a phase diagram of the phase formation capability.
It significantly improves the prediction and generalization capabilities of machine learning algorithms, reduces experimental costs, and quickly obtains the single-phase formation probability of high-entropy carbide components, solving the problem of difficulty in efficiently screening the optimal components and concentration combinations in traditional methods.
Smart Images

Figure CN120388660A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of ceramic material design, and relates to a method for predicting the phase formation ability of machine learning based on unbalanced data of high-entropy carbide ceramics. Background Art
[0002] With the increasingly complex service environment in the aerospace field, traditional single-component ceramic materials are difficult to meet the development needs of new technologies, and the need to research and develop high-performance new materials is becoming more urgent. High-entropy carbide ceramics (HECCs) developed on the basis of traditional ultra-high temperature ceramics have advantages such as higher hardness, higher melting point, better high-temperature oxidation resistance and ablation resistance than corresponding single-component carbides, making them one of the most promising material systems in ultra-high temperature extreme environments. In addition, studies have shown that the introduction of rare earth elements (Sc, Y, La) is beneficial to improving the oxidation resistance of ultra-high temperature ceramics and plays a beneficial role in stabilizing the oxide layer. Therefore, the development of new rare earth element-containing high-entropy carbides has become a research hotspot in the field of ultra-high temperature ceramic materials.
[0003] Although the field of high-entropy ceramics has developed vigorously in recent years, its huge composition space and complex atomic structure have hindered people's comprehensive exploration. Traditional new material synthesis routes often rely on a large number of "trial and error experiments", with a long development cycle and heavy workload. It is difficult to effectively screen out the best component and concentration combinations of single-phase high-entropy ceramic materials that meet specific requirements through large-scale experiments. With the rapid development of simulation calculations in materials science, high-throughput density functional theory, molecular dynamics, and computational phase diagrams have also been applied to the phase formation prediction of high-entropy ceramics. Although descriptors such as entropy formation ability (EFA) and mixing enthalpy (ΔH mix ) can predict the phase formation ability of HECCs under most conditions, they require thousands of costly first-principles calculations for each component to achieve the required accuracy. The high computational cost greatly limits the possibility of simulation calculations to explore and predict new high-entropy ceramics in a wide composition space.
[0004] In recent years, machine learning (ML) has not only accelerated material discovery and design strategies but also predicted the properties of unknown materials, thereby accelerating the development and application of new materials. Data-driven machine learning has made some remarkable progress in accelerating the rapid design and high-throughput prediction of novel high-entropy ceramics. The team of Shijun Zhao from the City University of Hong Kong collected 34 single-phase and 19 multi-phase HECCs from the literature as a dataset and successfully predicted the single-phase formation probability of novel HECCs using artificial neural network (ANN) and support vector machine (SVM) models (npj Computational Materials, 2022, 8(1): 5). The team of Yanhui Chu from South China University of Technology synthesized 91 high-entropy carbides as the original dataset through high-throughput methods and screened out the descriptor combinations for HECCs phase formation ability by combining high-throughput simulation calculations and machine learning methods (Cell Reports Physical Science, 2023, 4(8): 101512). However, the phase formation prediction of HECCs is currently limited by data imbalance, which has an adverse impact on the generalization ability and prediction performance of ML models. Specifically, most of the HECCs reported in the current literature focus on the B (Zr, Ti, Hf), V B (V, Nb, Ta), VI B (Cr, Mo, W) subgroups of the periodic table, while there is relatively little research on HECCs containing the B III B th B 、IV B th B 、V Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings of the prior art and provide a method for predicting the phase formation ability of high-entropy carbide ceramics based on imbalanced data in machine learning, so as to solve the technical problems that it is difficult to efficiently screen out the best component and concentration combinations of single-phase high-entropy ceramic materials that meet specific requirements using traditional experimental methods in the prior art, and the technical problems of low prediction accuracy and poor model generalization ability caused by data imbalance when applying machine learning algorithms to predict the phase formation ability.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A method for predicting the phase formation ability of high-entropy carbide ceramics based on machine learning for unbalanced data, comprising the following steps:
[0008] Step 1, according to the elemental composition ratio and elemental inherent properties of high-entropy carbide ceramics, calculate and obtain the original input characteristic parameters affecting the phase formation ability of high-entropy carbide ceramics, and obtain the initial data set;
[0009] Step 2, process the initial data set through an imbalance learning strategy to obtain a balanced data set, and after normalizing the balanced data set, obtain the normalized data set;
[0010] Step 3, perform feature engineering on the normalized data set to obtain the best feature subset;
[0011] Step 4, based on the best feature subset, compare multiple machine learning models to obtain the best machine learning model;
[0012] Step 5, predict the phase formation ability of the components of high-entropy carbide ceramics through the best machine learning model, verify through experiments, add the results of the experimental verification to the normalized data set in Step 2, re-obtain the best feature subset, and retrain and correct the best machine learning model through the best feature subset;
[0013] Step 6, predict and construct a phase formation ability phase diagram through the best machine learning model obtained in Step 5, and obtain the phase formation ability of high-entropy carbide ceramics corresponding to each component of high-entropy carbide ceramics through the phase formation ability phase diagram;
[0014] The endpoints of the phase formation ability phase diagram are elements, and the color of each point in the phase formation ability phase diagram represents the probability value of forming a phase at the corresponding elemental content.
[0015] A further improvement of the present invention lies in:
[0016] Preferably, in Step 1, the input of the initial data set is the original input characteristic parameters, and the output is the phase formation ability; the original input features are the weighted average and standard deviation of elemental inherent properties, as well as configurational entropy and geometric parameters; the phase formation ability is divided into 0 and 1.
[0017] Preferably, in Step 2, the data in the initial data set are divided into majority class samples and minority class samples according to the occurrence frequency of elements, and the minority class samples are oversampled through an imbalance learning strategy to obtain a balanced data set;
[0018] The imbalance learning strategy is the Borderline-SMOTE algorithm.
[0019] Preferably, in step 3, the feature engineering process is as follows: the normalized dataset is successively processed by Pearson correlation coefficient, recursive feature elimination, exhaustive search of all feature combinations, and hyperparameter optimization to obtain the optimal feature subset.
[0020] Preferably, in step 3, the specific process of the feature engineering process is as follows: redundant feature parameters in the normalized dataset are removed by Pearson correlation coefficient, the feature parameters in the dataset after removing the redundant feature parameters are sorted by the recursive feature elimination method, and unimportant feature parameters are removed; the feature combinations corresponding to the top 10 important feature parameters are exhaustively enumerated in the dataset to obtain the optimal feature subset; the hyperparameters of the machine learning model are optimized by random search to obtain the preliminary prediction model corresponding to the optimal feature subset.
[0021] Preferably, in step 4, the machine learning models include random forest, extreme gradient boosting, K-nearest neighbor, logistic regression, adaptive boosting, support vector machines with different kernel functions, decision tree, and naive Bayes.
[0022] Preferably, in step 4, during the comparison of multiple machine learning models, the area under the receiver operating characteristic curve, F1-score, Recall, and G-mean values of each machine learning model are compared.
[0023] Preferably, in step 5, prediction samples with the prediction probability of the optimal machine learning model between 0.3 and 0.7 are obtained, the components corresponding to the prediction samples are experimented, and the optimal machine learning model is corrected through the experimental results.
[0024] Preferably, in step 6, the high-entropy carbide ceramic can be a non-equimolar high-entropy carbide ceramic.
[0025] Preferably, when the high-entropy carbide ceramic is a non-equimolar high-entropy carbide ceramic, with a total of M elements and capable of generating N systems of non-equimolar high-entropy carbide ceramics, the process of obtaining the phase formation ability phase diagram of the high-entropy carbide ceramic is as follows: the components in all possible non-equimolar multi-component high-entropy carbide systems in the N systems are obtained, the characteristic parameters are calculated, and the characteristic parameters are input into the prediction model to obtain the phase formation ability phase diagram.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] The present invention discloses a method for predicting the phase formation ability of high-entropy carbide ceramics based on machine learning with imbalanced data. By collecting the literature data and experimental data on the phase formation ability of high-entropy carbide ceramics in existing research, an initial data set is constructed; the data distribution in the initial data set is statistically analyzed, and an imbalanced learning strategy is adopted to oversample the minority-class samples in the data set to balance the data set; the normalized data set is subjected to feature engineering processing, gradually deleting redundant features and unimportant features to obtain the best feature subset; the data set is divided into a training set and a test set, and several machine learning models are trained using different machine learning algorithms, and the performance of several trained models on the test set is calculated and compared; based on the optimized prediction model, the phase formation ability of high-entropy carbide ceramics in the unknown phase composition space is predicted; through experimental verification, whether the actual results and the prediction results are consistent is verified, and the experimental verification data are successively added to the data set for multiple model iterations to continuously optimize the machine learning prediction model; the iteratively optimized machine learning prediction model is extended to predict the phase formation ability of non-equimolar high-entropy carbide ceramics to obtain the corresponding prediction results. The present invention predicts the phase formation ability of high-entropy carbide ceramics through the combination of imbalanced learning and machine learning algorithms, significantly improving the prediction ability and generalization ability of machine learning algorithms.
[0028] The method of the present invention first uses machine learning algorithms to predict the phase formation ability of high-entropy carbide ceramics, which can quickly obtain the single-phase formation probability of specific high-entropy carbide components in a huge phase composition space, reducing the economic and time costs consumed by a large number of experiments, increasing scientificity and effectiveness, and solving the technical problem that the existing technology faces a huge phase composition space of high-entropy carbide ceramics, and it is difficult to efficiently screen out the best components and concentration combinations of single-phase high-entropy ceramic materials that meet specific requirements by using traditional experimental methods. On the other hand, based on imbalanced data, the Borderline-SMOTE algorithm of the imbalanced learning strategy is used to oversample the minority-class samples in the original data set. When the sample distribution in the original data set is unbalanced and easily leads to misclassification of minority-class data by the classifier, the data distribution in the data set is made more balanced by synthesizing a certain number of minority-class samples, thereby significantly improving the prediction ability and generalization ability of the machine learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flowchart of a machine learning prediction phase formation design method based on high-entropy carbide ceramic imbalanced data provided by the present invention;
[0030] Figure 2 It is the data set structure before and after Borderline-SMOTE processing of the present invention;
[0031] Figure 3 Comparison of the recursive feature elimination results of three different ML models of the present invention: (a) XGB, (b) RF, (c) SVM.rbf;
[0032] Figure 4 Comparison of the results of exhaustive all feature combinations of three different ML models of the present invention: (a) XGB, (b) RF, (c) SVM.rbf;
[0033] Figure 5 Performance evaluation results of three ML models of the present invention on the training set and the test set: (a) without Borderline-SMOTE processing, (b) after Borderline-SMOTE processing;
[0034] Figure 6 Predicted phase diagram of the optimized RF model of the present invention in the (HfZrTaYLa)C system. Detailed implementation manners
[0035] The present invention will be further described in detail below with reference to the accompanying drawings:
[0036] To enable those skilled in the art to understand the features and effects of the present invention, the following provides a general description and definition of the terms and phrases mentioned in the specification and claims. Unless otherwise specified, all technical and scientific terms used herein have the ordinary meaning understood by those skilled in the art regarding the present invention. In case of conflict, the definition in this specification shall prevail.
[0037] In this article, unless otherwise specified, "comprising", "including", "containing", "having" or similar terms cover the meanings of "consisting of" and "consisting essentially of". For example, "A comprises a" covers the meanings of "A comprises a and others" and "A only comprises a".
[0038] The present invention will be further described below with reference to specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0039] Conventional instrument equipment in the art is used in the following embodiments. For the experimental methods without specific conditions noted in the following embodiments, they are generally carried out under conventional conditions or according to the conditions recommended by the manufacturer. Various raw materials are used in the following embodiments. Unless otherwise stated, commercially available products are used, and their specifications are the conventional specifications in the art. In the specification of the present invention and the following embodiments, unless otherwise specified, "%" represents weight percentage, "parts" represents weight parts, and the ratio represents weight ratio.
[0040] The present invention discloses a machine learning prediction phase formation design method based on unbalanced data of high-entropy carbide ceramics, comprising the following steps:
[0041] Step 1, constructing a data set: collecting literature data and experimental data on the phase formation ability of high-entropy carbide ceramics in existing research, calculating the original input characteristic parameters affecting the phase formation ability of high-entropy carbide ceramics by using empirical formulas according to the element group allocation ratio and element inherent properties of high-entropy carbide ceramics, and constructing an initial data set. The input parameters in the initial data set are the original input characteristic parameters, and the output is the phase formation ability. The phase formation ability is 0 or 1, where 0 means not a high-entropy carbide ceramic phase and 1 means a high-entropy carbide ceramic phase.
[0042] Step 2, balancing the data set: statistically analyzing the data distribution in the initial data set, adopting an unbalanced learning strategy to oversample the minority-class samples in the data set, making the distribution of the minority-class samples and the majority-class samples in the data set more balanced by generating a certain number of minority-class samples, and then normalizing all the input feature data in the balanced data set to unify all the input features to the same order of magnitude;
[0043] Step 3, feature engineering: performing feature engineering processing on the normalized data set, gradually deleting redundant features and unimportant features to obtain the best feature subset.
[0044] Step 4, model establishment and evaluation: dividing the data set into a training set and a test set according to a certain ratio, inputting the training set data into different machine learning models for training respectively, establishing several machine learning models after training, calculating and comparing the performance of several trained models on the test set, and screening out the machine learning model with the best prediction performance.
[0045] Step 5, model prediction: based on the optimized prediction model for the phase formation ability of high-entropy carbide ceramics, predicting the phase formation ability of high-entropy carbide ceramics in the unknown phase composition space by randomly setting the element types and group allocation ratios.
[0046] Step 6, experimental verification and model iteration: selecting a certain number of representative high-entropy carbide ceramic compositions for experimental preparation, verifying whether the actual results are consistent with the predicted results, and based on the experimental verification results, sequentially adding the experimental verification data to the normalized data set in Step 2 for multiple model iterations, continuously optimizing the machine learning prediction model until the prediction performance remains stable and meets the expected requirements.
[0047] Step 7, Non-equimolar ratio high-entropy prediction: Generalize the iteratively optimized machine learning prediction model to predict the phase formation ability of non-equimolar ratio high-entropy carbide ceramics. Obtain the best machine learning model to predict and construct the phase formation ability phase diagram. Intuitively obtain the phase formation ability of high-entropy carbide ceramics corresponding to each component through the phase formation ability phase diagram.
[0048] The endpoints of the phase formation ability phase diagram are elements. The color of each point in the phase formation ability phase diagram represents the probability value of forming a single phase at the corresponding element content.
[0049] In some embodiments of the present invention, in step 1, the elements include twelve elements: Sc, Y, La, Ti, Zr, Hf, V, Nb, Ta, Cr, Mo, and W. The high-entropy carbide ceramics are a multi-component high-entropy ceramic system composed of any four to nine of the above elements.
[0050] In some embodiments of the present invention, in step 1, the inherent properties of the elements include valence electron concentration (VEC), Pauling electronegativity (χ p ), Mulliken electronegativity (χ m ), density (ρ), mass (m), lattice size (l), metal ion radius (r Me ), first ionization energy (I1), effective nuclear charge number (Z * ), configurational entropy (ΔSconf), and geometric parameter (Λ). The input features are the weighted average of relevant parameters such as VEC, χ p , χ m , ρ, m, l, r Me , I1, and Z * and the standard deviation (σ prop ), as well as configurational entropy (ΔSconf) and geometric parameter (Λ). During the initial calculation, for the selected inherent properties and input feature parameters, the inherent properties and input feature parameters can be inversely deduced according to the known phase situation. The empirical formula is the weighted average of the inherent properties prop the standard deviation (σ prop ) and the relevant expressions of ΔSconf and Λ:
[0051]
[0052] where X i is the molar fraction of each element i in the high-entropy carbide ceramics, and prop i is the value of a certain characteristic property of the carbide corresponding to the i-th element in the high-entropy carbide ceramics.
[0053] In some embodiments of the present invention, in step 2, the imbalanced learning strategy is specifically the Borderline-SMOTE algorithm.
[0054] Specifically, the data distribution in the dataset is statistically analyzed according to the occurrence frequencies of different subgroup elements in the periodic table of chemical elements during data collection. For some elements, there are more relevant studies, so there are more corresponding phase formation ability data, which are majority class samples. For some elements, there are fewer relevant studies, so there is a lack of corresponding phase formation ability data, which are minority class samples. When determining whether a specific element is a majority class sample or a minority class sample, it can be set according to the actual situation, or a judgment threshold can be set for judgment. The minority class samples (including subgroup III B subgroup element HECCs) in the dataset are oversampled using the imbalanced learning strategy Borderline-SMOTE algorithm. By generating a certain number of minority class samples, the distribution of minority class samples and majority class samples in the dataset becomes more balanced, and all input features after balancing are normalized.
[0055] Furthermore, for some elements, there are more high-entropy carbide ceramic phase data, and for some elements, there are fewer. The minority class samples in the present invention refer to high-entropy carbide ceramics containing Sc, Y, or La elements.
[0056] In some embodiments of the present invention, in step 2, the formula for the normalization process is as follows:
[0057]
[0058] In some embodiments of the present invention, where X i and respectively represent the values of the i-th feature of the input feature X before and after normalization, X max and X min are respectively the maximum and minimum values of the input feature X.
[0059] In step 3, the feature engineering is a four-step feature screening process including Pearson correlation coefficient, recursive feature elimination, exhaustive search of all feature combinations, and hyperparameter optimization. Among them, the screening criterion for the Pearson correlation coefficient is to retain one of the features with the absolute value of the Pearson correlation coefficient between two features greater than 0.9.
[0060] Specifically, the specific process of feature engineering is as follows: First, calculate the correlation between pairwise features in the dataset based on the Pearson correlation coefficient and determine the system threshold of the Pearson correlation coefficient. Select and retain the features with a value less than the system threshold and one feature from each pair of features with a value greater than or equal to the system threshold to obtain a feature set after initially removing redundant features. Then, train an initial model using the recursive feature elimination method and rank the feature importance to gradually remove unimportant features from the feature set. On this basis, exhaust all feature combinations corresponding to the top 10 important features and select the best feature subset. Finally, optimize the hyperparameters of the machine learning model with the best feature subset using the RandomizedSearchCV method to obtain a preliminary prediction model corresponding to the best feature subset.
[0061] In step 4, the process of model establishment and model evaluation is as follows: Randomly divide the training set and the test set in the ratio of training set:test set = 3:1. Use multiple algorithms such as Random Forest (RF), Extreme Gradient Boosting (XGB), K-Nearest Neighbor (KNN), Logistic Regression (LOG), Adaptive Boosting (Adaboost), Support Vector Machines with different kernel functions (SVM.rbf, SVM.poly, and SVM.linear), Decision Tree (DT), and Naive Bayes (NBayes) to construct machine learning models. Calculate through five-fold repeated cross-validation and comprehensively compare the area under the Receiver Operating Characteristic curve (AUC), F1-score, Recall, and G-mean values of different machine learning models to select the model with the best performance and strong generalization ability.
[0062] For AUC, F1-score, Recall, and G-mean, the corresponding calculation formulas are as follows:
[0063]
[0064] Among them, TP, FN, FP, and TN represent true positive rate, false negative rate, false positive rate, and true negative rate respectively.
[0065] In step 5, the process of model prediction is as follows: Randomly generate five-element high-entropy carbides in the phase composition space composed of nine elements Sc, Y, La, Ti, Zr, Hf, V, Nb, and Ta. Calculate the characteristic parameters required by the screened phase formation ability prediction model using an empirical formula and input them into the trained phase formation ability prediction model for phase formation ability prediction to obtain the prediction results of the phase formation ability corresponding to all high-entropy carbide components in the input dataset.
[0066] In step 6, the representative high-entropy carbide ceramic is a prediction sample with a relatively large classification uncertainty and a prediction probability between 0.3 and 0.7, and the number of model iterations is 5 times.
[0067] In step 6, the processes of experimental verification and model iteration are as follows: a certain number of prediction samples with a relatively large classification uncertainty and a prediction probability between 0.3 and 0.7 are selected according to the prediction results, and are experimentally prepared by the high-temperature solid-state reaction method. The prediction accuracy of the preliminary model is obtained by comparing the experimental verification results and the prediction results. Then, the experimental verification data are successively added to the balanced dataset for multiple model iterations to continuously optimize the machine learning prediction model until the prediction performance remains stable and meets the expected requirements.
[0068] In step 7, the non-equimolar high-entropy carbide ceramics are (HfZrTaScY)C, (HfZrTaScLa)C, and (HfZrTaYLa)C systems.
[0069] As Figure 6 shown, in step 7, the process of predicting non-equimolar high-entropy carbide ceramics is as follows: it is assumed that there are M elements in total, and M elements can generate N systems of non-equimolar high-entropy carbide ceramics, where both M and N are integers. The process of obtaining the phase formation ability phase diagram of high-entropy carbide ceramics by the trained model is as follows: obtain the components in all possible non-equimolar multi-component high-entropy carbide systems in N systems, calculate the characteristic parameters, and input them into the prediction model for phase formation ability prediction. The prediction phase diagram of the phase formation ability of the corresponding system is obtained by using the prediction results of all non-equimolar high-entropy carbide components.
[0070] For high-entropy carbide ceramics, due to the presence of multiple elements, there are also multiple elements in the formed phase diagram. Each coordinate line represents the change in the content of one element, and the formed phase diagram is a multi-dimensional graph. The color of each point in the multi-dimensional phase diagram represents the probability value of the ability to form a single phase at the corresponding content of each element.
[0071] Exemplarily, all possible non-equimolar five-component high-entropy carbide components are randomly generated in the phase composition spaces of (HfZrTaScY)C, (HfZrTaScLa)C, and (HfZrTaYLa)C systems, and the characteristic parameters required for the phase formation ability prediction model after multiple iterations are calculated by means of empirical formulas, and are input into the prediction model for phase formation ability prediction. The prediction phase diagram of the phase formation ability of the corresponding system is obtained by using the prediction results of all non-equimolar high-entropy carbide components.
[0072] A machine learning design method for predicting the formation ability of high-entropy carbide ceramics based on imbalanced data has the following general idea. First, collect the literature data and experimental data on the formation ability of high-entropy carbide ceramics in existing research, and calculate the original input feature set affecting the formation ability of high-entropy carbide ceramics through empirical data. Statistically analyze the data distribution in the initial dataset, and use the imbalanced learning strategy to oversample the minority-class samples in the dataset to make the sample distribution in the dataset more balanced, and normalize all the input feature data in the balanced dataset. Perform feature dimensionality reduction through a four-step feature selection method of Pearson correlation coefficient, recursive feature elimination, exhaustive feature combination, and hyperparameter optimization. Then use multiple algorithms such as random forest (RF), extreme gradient boosting (XGB), K-nearest neighbor (KNN), logistic regression (LOG), adaptive boosting (Adaboost), support vector machines with different kernel functions (SVM.rbf, SVM.poly, and SVM.linear), decision tree (DT), and naive Bayes (NBayes) to construct and train a machine learning model, calculate and compare the AUC, F1-score, Recall, and G-mean values of different machine learning models, and select the model with the best performance and strong generalization ability. Use the trained model to predict the formation ability of high-entropy carbide ceramics in the unknown phase composition space, select some predicted components for experimental verification based on the prediction results, and compare the verification results with the prediction results to obtain the prediction accuracy of the preliminary model.
[0073] The above technical solutions will be further described below in conjunction with specific embodiments. The preferred embodiments of the present invention are described in detail as follows:
[0074] Example 1
[0075] Step 1: Based on the literature data and experimental data in existing research, 171 pieces of data were collected, and a dataset of quaternary to nonary high-entropy carbide ceramics composed of twelve elements, namely Sc, Y, La, Ti, Zr, Hf, V, Nb, Ta, Cr, Mo, and W, was constructed. Twenty original input feature parameters affecting the formation ability of high-entropy carbide ceramics were calculated using empirical formulas according to the element group ratios and inherent properties of high-entropy carbide ceramics, specifically including σ VEC , σ ρ , σ m , σ l , σ Z * , ΔSconf, Λ.
[0076] Step 2: According to the occurrence frequency of different subgroups of elements in the periodic table of chemical elements, the data distribution in the dataset is statistically analyzed. For the minority class samples in the dataset (including the subgroup III B subgroup element HECCs), the imbalanced learning strategy Borderline-SMOTE algorithm is used for oversampling, generating 80 minority class samples to increase the sample size of the dataset to 251, and normalizing all the input features after balancing.
[0077] From Figure 2 It can be seen that after being processed by the Borderline-SMOTE technique, the data volume of high-entropy carbides containing Sc / Y / La elements in the dataset increases from 18 to 98, and the total number of data in the dataset increases from 171 before processing to 251, effectively alleviating the data imbalance problem between the majority class samples and the minority class samples in the original dataset.
[0078] Step 3: First, calculate the correlation between pairwise features in the dataset based on the Pearson correlation coefficient and take the absolute value equal to 0.9 as the system threshold of the Pearson correlation coefficient. Select and retain the features with a value less than the system threshold and one of the pairwise features with a value greater than or equal to the system threshold, reducing the number of features from the original 20 to 16; then, train six different initial machine learning models through the recursive feature elimination method and rank the feature importance, so as to gradually remove the unimportant features in the feature set, thus roughly screening out the optimal number of features corresponding to different machine learning models; on this basis, enumerate all the feature combinations corresponding to the top 10 important features, and screen out the optimal number of features and the optimal feature subset corresponding to different machine learning models; then, optimize the hyperparameters of the machine learning model through the RandomizedSearchCV method to obtain the preliminary prediction model corresponding to the optimal feature subset.
[0079] From Figure 3 It can be seen that the XGB, RF, and SVM.rbf models trained by the recursive feature elimination method show a similar pattern, that is, as the number of features gradually decreases, the average AUC value shows a trend of first increasing and then decreasing. This is mainly because in the initial stage of feature elimination, as the number of features decreases, the unimportant features are removed one by one, thus improving the model performance. However, as the number of features further decreases, some important features may be gradually eliminated, resulting in the model being unable to fully capture the key information in the dataset, which will lead to a significant decline in the model performance.
[0080] From Figure 4It can be seen that after exhausting all the feature combinations corresponding to the top 10 important features, the model presents a similar pattern to the recursive feature elimination method, and the optimal feature subsets and the optimal number of features corresponding to the three models are further optimized. Specifically, the optimal number of features corresponding to the XGB, RF, and SVM.rbf models are 6, 5, and 6 respectively.
[0081] Step 4: Randomly divide the training set and the test set at a ratio of training set: test set = 3:1. Use multiple algorithms such as Random Forest (RF), Extreme Gradient Boosting (XGB), K-Nearest Neighbor (KNN), Logistic Regression (LOG), Adaptive Boosting (Adaboost), Support Vector Machines with different kernel functions (SVM.rbf, SVM.poly, and SVM.linear), Decision Tree (DT), and Naive Bayes (NBayes) to construct machine learning models. Calculate and comprehensively compare the AUC, F1-score, Recall, and G-mean values of different machine learning models through five-fold repeated cross-validation, so as to select the RF model with the best performance and strong generalization ability as the final prediction model.
[0082] From Figure 5 it can be seen that the three machine learning models without Borderline-SMOTE processing all have good performance on the training set ( Figure 5 (a)), especially the Recall value of the XGB model on the training set reaches 100%, indicating that the model can correctly identify the phase formation ability of all samples in the training set. However, the performance of the three models on the test set is poor, which shows that data imbalance exacerbates the overfitting of the model. From Figure 5 (b), it can be seen that after Borderline-SMOTE, the three models show good prediction performance on both the training set and the test set, indicating that the Borderline-SMOTE-assisted machine learning method helps to solve the overfitting problem and effectively improves the generalization ability of the model.
[0083] Step 5: Randomly generate quinary high-entropy carbides in the phase composition space composed of nine elements Sc, Y, La, Ti, Zr, Hf, V, Nb, and Ta. Calculate the characteristic parameters required by the selected phase formation ability prediction model with the help of empirical formulas, and input them into the trained phase formation ability prediction model for phase formation ability prediction to obtain the prediction results of the phase formation ability corresponding to all high-entropy carbide components in the input dataset.
[0084] Step 6: Select a certain number of predicted samples with relatively large classification uncertainty and predicted probabilities between 0.3 and 0.7 through the high-temperature solid-state reaction method for experimental preparation. Compare the experimental verification results with the predicted results to obtain the prediction accuracy of the preliminary model. Then, sequentially add the experimental verification data to the dataset after normalization in Step 2 for multiple model iterations to continuously optimize the machine learning prediction model until the prediction performance remains stable and meets the expected requirements.
[0085] Step 7: Randomly generate all possible non-equimolar ratio quinary high-entropy carbide components in the phase composition spaces of the (HfZrTaScY)C, (HfZrTaScLa)C, and (HfZrTaYLa)C systems. Calculate the characteristic parameters required for the phase formation ability prediction model after multiple iterations using empirical formulas, and input them into the prediction model for phase formation ability prediction. Use the prediction results of all non-equimolar ratio high-entropy carbide components to obtain the predicted phase diagrams of the corresponding system's phase formation ability.
[0086] Figure 6 is the predicted phase diagram of (HfZrTaYLa)C. Figure 6 In the predicted phase diagram, in order to visualize high-dimensional data, the quinary high-entropy system is usually simplified to a pseudo-ternary system, that is, usually three of the elements are combined and regarded as a "synthetic end point" to adapt to the dimension of the ternary phase diagram. Each coordinate represents the content of each element in a specific non-equimolar ratio high-entropy carbide, and the percentage content ranges from 0 to 1; when the total content of Hf, Zr, and Ta elements is less than 0.25, that is, when the rare earth element concentration is relatively high, the (HfZrTaLaY) system tends to form a single-phase high-entropy carbide within a relatively large concentration range; when the total content of Hf, Zr, and Ta elements exceeds 0.85, that is, corresponding to a lower rare earth element concentration, there is a higher possibility of forming a single-phase high-entropy carbide in the lower right corner of the predicted phase diagram. The establishment of the predicted phase diagram provides a valuable reference for the experimental work of screening single-phase non-equimolar ratio rare earth high-entropy carbides in the (HfZrTaYLa)C system, and provides an intuitive guidance for optimizing the element ratio and exploring potential single-phase high-entropy carbide materials.
[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for predicting the phase formation ability of machine learning based on unbalanced data of high-entropy carbide ceramics, characterized in that It includes the following steps: Step 1: According to the elemental composition ratio and elemental inherent properties of the high-entropy carbide ceramic, calculate and obtain the original input characteristic parameters affecting the phase formation ability of the high-entropy carbide ceramic, and obtain the initial dataset; Step 2: Process the initial dataset through an imbalance learning strategy to obtain a balanced dataset, and after normalizing the balanced dataset, obtain the normalized dataset; Step 3: Process the normalized dataset through feature engineering to obtain the best feature subset; Step 4: Based on the best feature subset, compare multiple machine learning models to obtain the best machine learning model; Step 5: Predict the phase formation ability of the high-entropy carbide ceramic components through the best machine learning model. Through experimental verification, add the results of the experimental verification to the normalized dataset in Step 2, re-obtain the best feature subset, and retrain and correct the best machine learning model through the best feature subset; Step 6: Predict and construct a phase formation ability phase diagram through the best machine learning model obtained in Step 5, and obtain the high-entropy carbide ceramic phase formation ability corresponding to each component of the high-entropy carbide ceramic through the phase formation ability phase diagram; The endpoints of the phase formation ability phase diagram are elements, and the color of each point in the phase formation ability phase diagram represents the probability value of forming a phase at the corresponding element content.
2. A method for machine learning to predict the phase formation ability based on unbalanced data of high-entropy carbide ceramics according to claim 1, characterized in that, In Step 1, the input of the initial dataset is the original input characteristic parameters, and the output is the phase formation ability; the original input features are the weighted average and standard deviation of the elemental inherent properties, as well as the configurational entropy and geometric parameters; the phase formation ability is divided into 0 and 1.
3. A method for predicting the phase formation ability of machine learning based on high-entropy carbide ceramics with unbalanced data according to claim 1, characterized in that In Step 2, the data in the initial dataset is divided into majority class samples and minority class samples according to the occurrence frequency of elements, and the minority class samples are oversampled through an imbalance learning strategy to obtain a balanced dataset; The imbalance learning strategy is the Borderline-SMOTE algorithm.
4. A method for predicting the phase formation ability of machine learning based on high-entropy carbide ceramics with unbalanced data according to claim 1, characterized in that, In Step 3, the feature engineering processing process is: sequentially pass the normalized dataset through the Pearson correlation coefficient, recursive feature elimination, exhaustive search of all feature combinations, and hyperparameter optimization processing to obtain the best feature subset.
5. A method for machine learning to predict the phase formation ability based on unbalanced data of high-entropy carbide ceramics according to claim 1, characterized in that In Step 3, the specific process of the feature engineering processing is: remove the redundant feature parameters in the normalized dataset through the Pearson correlation coefficient, sort the feature parameters in the dataset after removing the redundant feature parameters through the recursive feature elimination method, and remove the unimportant feature parameters; exhaustively search the feature combinations corresponding to the top 10 important feature parameters in the dataset to obtain the best feature subset; optimize the hyperparameters of the machine learning model through random search to obtain the preliminary prediction model corresponding to the best feature subset.
6. A method for machine learning to predict the phase formation ability based on unbalanced data of high-entropy carbide ceramics according to claim 1, characterized in that, In Step 4, the machine learning models include random forest, extreme gradient boosting, K-nearest neighbor, logistic regression, adaptive boosting, support vector machines with different kernel functions, decision tree, and naive Bayes.
7. A method for machine learning to predict the phase formation ability based on unbalanced data of high-entropy carbide ceramics according to claim 1, characterized in that In Step 4, during the process of comparing multiple machine learning models, compare the area under the receiver operating characteristic curve, F1-score, Recall, and G-mean values of each machine learning model.
8. A method for machine learning to predict the phase formation ability based on unbalanced data of high-entropy carbide ceramics according to claim 1, characterized in that In step 5, obtain the prediction samples with the predicted probability of the best machine learning model between 0.3 and 0.7, conduct experiments on the components corresponding to the prediction samples, and correct the best machine learning model based on the experimental results.
9. A method for machine learning to predict the phase formation ability based on unbalanced data of high-entropy carbide ceramics according to claim 1, characterized in that, In step 6, the high-entropy carbide ceramic can be a non-equimolar high-entropy carbide ceramic.
10. A method for predicting the phase formation ability of machine learning based on high-entropy carbide ceramic unbalanced data according to claim 1, characterized in that, When the high-entropy carbide ceramic is a non-equimolar high-entropy carbide ceramic and there are M elements in total, and N systems of non-equimolar high-entropy carbide ceramics can be generated, the process of obtaining the phase formation ability phase diagram of the high-entropy carbide ceramic is as follows: obtain the components in all possible non-equimolar multi-component high-entropy carbide systems in the N systems, calculate the characteristic parameters, input the characteristic parameters into the prediction model, and obtain the phase formation ability phase diagram.
Citation Information
Cited By
Fine-grained soil saturated permeability coefficient prediction method based on machine learning
CN121051717A