Iron-based high-temperature alloy design method based on multi-model machine learning and alloy thereof

By combining multi-model machine learning with thermodynamic simulation, the microcrack problem in the manufacturing of iron-based superalloys in LPBF was solved, achieving efficient and low-cost crack-free forming and improving the mechanical properties and design efficiency of the alloy.

CN121789858APending Publication Date: 2026-04-03SHANDONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing iron-based superalloy composition design methods suffer from microcrack problems in laser powder bed melting manufacturing, making it difficult to achieve efficient and low-cost crack-free forming. Furthermore, existing design methods are inefficient, costly, and have large prediction biases, making it difficult to meet multi-objective optimization requirements in LPBF environments.

Method used

By combining a multi-model machine learning framework with thermodynamic simulation, an end-to-end design closed loop is constructed. Through feature selection, model optimization, and small-scale sample verification, a crack-free iron-based superalloy is designed. By combining multi-objective optimization and laser powder bed melting technology, the manufacturing of iron-based superalloys with high-temperature strength and low cost is achieved.

Benefits of technology

It significantly reduces crack sensitivity, improves the mechanical properties of the alloy, shortens the R&D cycle, reduces costs, and achieves efficient crack-free forming, meeting the multi-objective optimization requirements of the LPBF process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789858A_ABST
    Figure CN121789858A_ABST
Patent Text Reader

Abstract

The invention relates to an iron-based high-temperature alloy design method based on multi-model machine learning and an alloy thereof, and the method comprises the following steps: 1, obtaining iron-based high-temperature alloy components, carrying out high-throughput calculation, and constructing a data set; 2, correlation analysis is carried out on the data set, feature importance sorting is carried out based on a random forest model, and key components are screened out; 3, constructing a machine learning algorithm model, training by using the key components, and adjusting and optimizing hyper-parameters of the machine learning algorithm model through an optimization algorithm to obtain a trained prediction model; 4, constraint conditions are constructed and optimized, an optimal solution set is obtained, the optimal solution set is input into the trained prediction model, and optimal alloy components are obtained; and 5, laser powder bed melting forming is conducted according to the optimal alloy components, and the iron-based high-temperature alloy is obtained. A data-driven machine learning method is adopted to replace a traditional trial and error method, the alloy research and development period is greatly shortened, and the research and development cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a design method for iron-based high-temperature alloys based on multi-model machine learning and the alloy thereof, belonging to the field of metal additive manufacturing technology. Background Technology

[0002] Iron-based superalloys are precipitation-strengthened superalloys. Due to their high volume fraction of γ′ strengthening phase, these alloys exhibit excellent yield strength, good fatigue resistance, and outstanding creep resistance at high temperatures. They are widely used in the manufacture of key hot-end components in aero-engines, such as turbine disks, compressor disks, rotor blades, and various fasteners. While traditional manufacturing processes (such as casting and forging) are technically mature, they often suffer from long production cycles, low material utilization, and high manufacturing costs when dealing with complex structures and high precision requirements, limited by mold design accuracy and processing capabilities. This is particularly true for efficiently achieving integrated forming of complex parts with internal flow channels and honeycomb structures. Metal additive manufacturing technology, especially Laser Powder Bed Fusion (LPBF), selectively melts pre-spread metal powder layers with a high-energy laser beam and accumulates them layer by layer. This allows for the direct fabrication of high-precision, geometrically complex metal parts from digital models, effectively overcoming the limitations of traditional manufacturing methods in forming complex components. This provides an important technological path for manufacturing lightweight, functionally integrated advanced components. However, most alloys face significant technical challenges in LPBF rapid prototyping: extremely high cooling rates and drastic temperature gradients lead to the accumulation of internal residual stress, which in turn triggers the formation and propagation of microcracks. These microscopic defects severely impair the mechanical properties and reliability of parts, limiting the application of this alloy in additive manufacturing. Currently, simply adjusting process parameters such as laser power, scanning speed, and scanning strategy is insufficient to completely eliminate microcracks, indicating that the root cause involves a deep contradiction between metallurgical mechanisms and thermodynamic processes, necessitating a more fundamental material solution.

[0003] Current methods for designing the composition of iron-based superalloys mainly rely on three types of approaches: First, the traditional design approach based on experience and analogy. This involves using existing mature alloy systems (such as the Inconel and Haynes series) as a foundation, with experts obtaining new compositional combinations through fine-tuning and iterative experiments. However, this method is highly dependent on experience, has limited design space exploration, high experimental costs, and long cycles, making it difficult to achieve global optimization under multiple constraints such as strength, cost, and printability. Furthermore, its ability to predict microstructure evolution under the rapid solidification environment of LPBF is limited. Second, the quantitative thermodynamic design method based on CALPHAD calculates phase diagrams, phase stability, and solidification paths using a thermodynamic database. While widely used in traditional casting and forging processes, under the non-equilibrium solidification conditions of LPBF, it is insufficient in predicting metastable phase formation, phase transformations caused by intense thermal cycling, and solid-liquid interface dynamics. Its computational efficiency is also limited, making it difficult to support rapid screening of large-scale, high-dimensional design spaces and to achieve multi-objective synergistic optimization of key indicators such as crack sensitivity (FR), solid-liquid interface (SCI), and solid-liquid interface (SAC). Third, while single-model-based machine learning design methods, such as XGBoost, GBDT, and ANN, have emerged in recent years and can be used for performance prediction and composition optimization, these methods often suffer from uneven fitting capabilities to different performance indices (FR / SCI / SAC) due to model bias. They also exhibit insufficient generalization, limited predictive stability, and lack the ability to explore the global design space, resulting in deficiencies in interpretability. In summary, existing design methods still suffer from low efficiency, high cost, large prediction bias, and difficulty in systematic optimization when developing compositions for iron-based superalloys in LPBF environments. These limitations hinder the rapid development of novel iron-based superalloys that offer "low cost, high strength, and low crack sensitivity." Therefore, there is an urgent need for an intelligent composition design method that integrates multi-model prediction capabilities, is suitable for large-scale design space searches, and better aligns with the non-equilibrium solidification characteristics of LPBF. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a design method for iron-based superalloys based on multi-model machine learning and the resulting alloy. It deeply couples a multi-model machine learning framework (model selection, feature engineering, integration strategy, and uncertainty-based active learning closed loop) with customized objective functions and constraints for LPBF non-equilibrium solidification / crack sensitivity indices (FR / SCI / SAC, etc.). This constructs an end-to-end design closed loop encompassing computational prediction, physical / thermodynamic simulation (including CALPHAD assistance), and small-scale LPBF experimental verification. This achieves, for the first time, near-crack-free forming of the iron-based superalloy AM-SD during the LPBF process, while simultaneously maintaining high-temperature strength and low cost. This technology is not simply a "scraping together of multiple existing algorithms," but represents a substantial breakthrough in integration architecture, objective definition, verification methods (process-microstructure-performance closed loop), and uncertainty handling, resulting in quantifiable engineering effects (such as significantly reduced crack sensitivity and mechanical properties of printed parts that meet or exceed those of existing alloys).

[0005] The technical solution of the present invention is as follows: A design method for iron-based superalloys based on multi-model machine learning, comprising: Step 1: Obtain the composition of iron-based high-temperature alloys and perform high-throughput calculations to build a dataset; Step 2: Perform correlation analysis on the dataset and rank the features by importance based on the random forest model to select key components; Step 3: Build a machine learning algorithm model, train it using key components, and fine-tune the hyperparameters of the machine learning algorithm model through optimization algorithms to obtain a trained prediction model; Step 4: Construct and optimize constraints to obtain the optimal solution set. Input the optimal solution set into the trained prediction model to obtain the optimal alloy composition. Step 5: Perform laser powder bed melting and forming according to the optimal alloy composition to obtain an iron-based high-temperature alloy.

[0006] According to a preferred embodiment of the present invention, step 1 specifically comprises: Randomly generate several combinations of gold components; High-throughput thermodynamic calculations were performed on several composite gold compositions using thermodynamic calculation software to obtain non-equilibrium solidification paths and equilibrium phase transition diagrams. Non-equilibrium solidification pathways include liquidus temperature, solidus temperature, and solid fraction. The equilibrium phase transformation diagram includes the volume fraction of precipitated phases and the volume fraction of γ matrix during alloy solidification; The cracking index of the alloy composition is calculated by using the non-equilibrium solidification path and equilibrium phase transformation diagram, including the solidification temperature range (FR), the solidification cracking factor (SCI), and the strain aging crack index (SAC). The alloy composition and the corresponding cracking indices FR, SCI and SAC were constructed into a dataset.

[0007] According to a preferred embodiment of the present invention, correlation analysis is performed on the dataset; including: Using several gold components in the dataset as component variables, and FR, SCI and SAC as target variables, correlation analysis was performed to calculate the Pearson correlation coefficient and the Spearman correlation coefficient. The Pearson correlation coefficient r is shown below: ; in, and These represent the first and second components of the variable and the target variable, respectively. One observation value; and These are the sample means of the two variables, respectively. It is the sample size; Spearman correlation coefficient As shown below: ; in, It is each pair of observations The difference in grade.

[0008] According to a preferred embodiment of the present invention, feature importance is ranked based on a random forest model, and key components are screened using correlation analysis; including: The random forest model was used to calculate feature importance. The alloy composition in the dataset was used as the feature, and FR, SCI and SAC were used as the target variables. The random forest model was trained and the MDI based on node purity decrease and the importance PI based on permutation were calculated. Mean Impurity Reduction (MDI) includes: In a random forest model, for a given set of features... The importance of node m, which is to be split, is calculated as follows: ; in, It is the proportion of samples reaching node m, i.e., the node weight. It is the impurity of the parent node. and It refers to the impurity of the left and right child nodes after the split; characteristic The MDI in the entire random forest model is shown below: ; in, Indicates the calculated MDI, This represents the total number of decision trees in the random forest model. This represents a single decision tree in a forest. This indicates a specific node in the decision tree. The features selected for splitting; Indicates at node Features used The amount of impurity reduction resulting from the splitting process; The importance of features (PI) includes calculating the performance score of the random forest model, as shown below: ; in, This represents the performance score of the random forest model. Let M represent the evaluation function, M represent the random forest model, and D represent the input to the random forest model. The features in the random forest model are randomly shuffled; the shuffled features and the target variable form a new input. Use new input Recalculate the performance score of the random forest model: ; in, The new performance score is represented; the importance of the ranked features is the difference between the baseline score and the score after the disruption, as shown below: ; in, Indicates the importance of arrangement features; Repeat the above steps of randomly shuffling features, calculating performance scores, and calculating the importance of ranked features to obtain the average value of the ranked feature importance, which is used as the final ranked feature importance PI. In the correlation analysis, the top k features with the largest absolute values ​​were selected, and the top k features with the largest MDI and PI were selected as key components.

[0009] According to a preferred embodiment of the present invention, step 3 specifically comprises: The machine learning algorithm models are XGBoost and RF models; The key components were divided into training and testing sets, and trained using XGBoost and RF models respectively. During the training of XGBoost and RF models, Bayesian optimization is used to automatically search in a preset hyperparameter space; the optimal model hyperparameters corresponding to the three prediction objectives FR, SCI and SAC are determined respectively, and the models corresponding to the three prediction objectives are obtained, which are denoted as FR prediction model, SCI prediction model and SAC prediction model respectively.

[0010] According to a preferred embodiment of the present invention, step 4 specifically includes, for example... Figure 5 As shown, it includes: Target constraints for FR, SCI, and SAC are constructed: FR ≥ 200 MPa, SCI ≥ 1800 MPa, and SAC ≥ 0.15. Boundary constraints for alloy composition are also constructed: Ni: 30–40%, Cr: 5–15%, Al: 2–4%, Ti: 3–8%, W: 1–3%, Mo: 1–3%, Mn: 0–1%, Si: 0–1%, and B: 0–0.5%. The multi-objective optimization algorithm NSGA-II is used to optimize the target constraints and boundary constraints to obtain the optimal solution set. The optimal solution set is input into the trained FR, SCI, and SAC prediction models to predict the optimal solution set. At the same time, through the non-dominated ranking mechanism and the crowding distance selection mechanism, iterative updates are continuously generated to generate new components with better overall performance. After the prediction model converges, a Pareto front solution set is obtained, which is a set of optimal candidate components that cannot be improved simultaneously on the three objectives and are mutually balanced. Based on actual needs, the Pareto front solution set is further screened to obtain the final optimal alloy composition, namely the iron-based high-temperature alloy AM-SD.

[0011] According to a preferred embodiment of the present invention, AM-SD alloy powder is subjected to LBPF forming to obtain an iron-based superalloy; comprising: AM-SD alloy powder was obtained by gas atomization, with a particle size distribution of 15~53μm and an average particle size of 30-40μm. AM-SD alloy powder is formed by laser powder bed melting (LPBF) using metal additive manufacturing equipment. The laser power is 180-200W, the scanning speed is 800-1000mm / s, the powder bed thickness is 30-50μm, and the scanning spacing is 100-120μm.

[0012] Further preferred features include a laser power of 190W, a scanning speed of 900mm / s, a layer thickness of 40μm, and a scanning spacing of 110μm.

[0013] A multi-model machine learning-based iron-based superalloy is obtained through the aforementioned multi-model machine learning-based iron-based superalloy design method. The alloy powder composition by mass percentage includes: C<0.08%, Cr 9.00-15.00%, Ni 35.00-40.00%, Mo 1.50-2.50%, W 1.50-2.50%, Al 1.50-3.00%, Ti 3.00-10.00%, B 0.15-0.30%, Si 0.30-0.50%, Mn 0.20-0.40%, S<0.02%, P<0.02%, O<0.02%, N<0.02%, with the balance being Fe and unavoidable impurities.

[0014] According to a preferred embodiment of the present invention, the particle size of the iron-based high-temperature alloy, namely AM-SD alloy powder, is 15~53μm.

[0015] The beneficial effects of this invention are as follows: 1. The adoption of data-driven machine learning methods has replaced the traditional trial-and-error approach, greatly shortening the alloy R&D cycle and reducing R&D costs.

[0016] 2. Through rigorous feature selection and model optimization, the constructed performance prediction model has high accuracy (R²>0.85) and strong generalization ability, laying a solid foundation for reliable design.

[0017] 3. The constructed composition design system can handle single-objective or multi-objective optimization tasks, and can flexibly set performance targets and constraints according to specific application requirements to achieve "design on demand".

[0018] 4. The AM-SD alloy designed by the method of this invention effectively solves the microcrack problem caused by high residual stress and sensitive solidification range during LPBF process. It can be formed without cracks through LPBF, and its room temperature and high temperature mechanical properties are significantly improved.

[0019] 5. Through SHAP interpretability analysis, the impact mechanism of each key element on performance was clarified, guiding the reasonable setting of component ranges and making the design process no longer a "black box". Attached Figure Description

[0020] Figure 1 This is a flowchart of the iron-based superalloy design method based on multi-model machine learning of the present invention; Figure 2 This is a schematic diagram showing the analysis results of various features of the FR prediction model of the present invention on the MDI and PI indices; Figure 3 This is a scatter plot of the predicted and calculated values ​​of the three cracking index prediction models after hyperparameter optimization according to the present invention. Figure 3(a) is a scatter plot of the predicted and calculated values ​​from the FR prediction model; Figure 3 (b) is a scatter plot of the predicted and calculated values ​​of the SCI prediction model; Figure 3 (c) is a scatter plot of the predicted and calculated values ​​of the SAC prediction model; Figure 4 This is a schematic diagram of SHAP analysis using the FR prediction model as an example in this invention; Figure 5 This is a schematic diagram of the multi-objective optimization process of the present invention; Figure 6 This is a schematic diagram of the microstructure of the AM-SD alloy prepared according to the present invention; Figure 7 This is a schematic diagram showing the results of the tensile stress-strain curves of the AM-SD alloy compared with GH4169 at room temperature and high temperature. Detailed Implementation

[0021] The present invention will be further described below with reference to the embodiments and accompanying drawings, but is not limited thereto.

[0022] Example 1: Terminology Explanation: 1. Laser Powder Bed Fusion (LPBF) technology: By selectively melting a pre-spread layer of metal powder with a high-energy laser beam and accumulating it layer by layer, it can directly manufacture high-precision metal parts with complex geometries from digital models. It effectively overcomes the limitations of traditional manufacturing methods in forming complex components and provides an important technical path for manufacturing advanced components that are lightweight and functionally integrated.

[0023] 2. Pearson correlation coefficient: measures the strength and direction of the linear relationship between two continuous variables; : Perfectly positive linear correlation; : Perfect negative linear correlation; Wireless correlation; absolute value The closer to 1, the stronger the linear relationship; 2. Spearman correlation coefficient: The Spearman correlation coefficient measures the strength of the monotonic relationship between two variables; that is, whether the other variable tends to increase (or decrease) when one variable increases, regardless of whether the rate of increase is constant. : Perfectly monotonically positively correlated; : Perfectly monotonically negatively correlated; No monotonic relationship; 4. MDI: Mean Impurity Reduction, also known as Gini Importance, is calculated based on the sum of node impurity reductions during the decision tree construction process. It measures the average contribution of a feature during model training to reducing node impurity in the decision tree (i.e., making predictions more accurate). Node impurity: For classification problems, Gini impurity is typically used; for regression problems (such as cracking index prediction), variance is typically used.

[0024] 5. PI: Ranking Feature Importance. It is calculated based on the degree of performance degradation after a feature is altered. It measures how much the model's predictive performance decreases when a feature becomes randomized. The greater the decrease, the more important the feature is to the model's predictions.

[0025] 6. XGBoost: An efficient machine learning algorithm based on the gradient boosting framework. It iteratively constructs a series of decision trees and combines regularization terms to prevent overfitting, demonstrating excellent performance and computational efficiency in various prediction tasks.

[0026] 7. Random Forest (RF): Random forest is an ensemble learning method that improves model accuracy and stability by constructing multiple decision trees and combining their prediction results (such as voting in classification tasks or averaging in regression tasks), while reducing the risk of overfitting.

[0027] 8. Bayesian optimization: A sequence model optimization method for optimizing black-box functions. It constructs a probabilistic surrogate model of the objective function (such as a Gaussian process) and intelligently selects the next evaluation point by combining the acquisition function (such as the desired improvement in EI). This allows it to find the global optimum with fewer evaluations and is often used for hyperparameter tuning.

[0028] 9. SHAP Analysis: Short for Shapley Additive exPlanations, it's a model-independent machine learning interpretability method based on Shapley values ​​in cooperative game theory. Its core objective is to quantify the contribution of each feature (input variable) to a single prediction, thus decomposing the output of a complex "black box" model into a linear sum of the contributions of each feature. The core idea is to treat the model's prediction as the result of cooperation among all feature "players," and to allocate the feature's contribution value by fairly calculating the average of the marginal contributions of each feature to the prediction across all possible feature combinations ("alliances").

[0029] 10. Polynomial Feature Engineering: This is a data preprocessing method that expands the feature space by creating higher-order terms (such as squares and cubes) and interaction terms (such as the product of two features) of the original features. Its core idea is that the inherent laws of many systems are not simple linear relationships, but contain nonlinear effects and interactions between features. By explicitly creating these new features, a linear model can be transformed into a nonlinear model, or a more direct and learnable relational expression can be provided for nonlinear models (such as decision trees). Its working principle and formula are as follows: Assume there are two original features: and Original feature set: If we perform a second polynomial feature expansion, we will generate the following new features: constant term: (Usually processed by the model's intercept term, can be ignored); Original features (first-order terms): , Interactive items: * Higher-order terms (square terms): Expanded feature set: , ;

[0030] 11. Non-dominated ranking mechanism: A hierarchical ranking strategy in multi-objective optimization. It ranks solutions according to their "dominance" relationship: if one solution is not inferior to another solution in all objectives, and is strictly superior in at least one objective, then the former is said to dominate the latter. Non-dominated solutions are assigned to the first level (Pareto front), removed, and then the remaining solutions are ranked in the next round, forming multiple front levels in sequence.

[0031] 12. Crowding Distance Selection Mechanism: Used in multi-objective optimization to measure the distribution density of solutions within the same non-dominated layer. It calculates the sum of distances between each solution and its neighboring solutions in each objective direction; a larger crowding distance indicates a sparser distribution of solutions around that solution. This mechanism often prioritizes solutions with large crowding distances during selection operations to maintain population diversity and a uniform distribution of the Pareto front.

[0032] 13. Multi-objective optimization algorithm (NSGA-II): also known as "Non-dominated sorting genetic algorithm II", it is a classic multi-objective evolutionary algorithm. By combining non-dominated sorting and crowding distance comparison, it simultaneously optimizes multiple conflicting objectives during the evolutionary process, and finally obtains a set of uniformly distributed and well-converged Pareto optimal solutions.

[0033] A design method for iron-based superalloys based on multi-model machine learning, such as Figure 1 As shown, it includes: Step 1: Obtain the composition of iron-based high-temperature alloys and perform high-throughput calculations to build a dataset; Step 2: Perform correlation analysis on the dataset and rank the features by importance based on the random forest model to select key components; Step 3: Build a machine learning algorithm model, train it using key components, and fine-tune the hyperparameters of the machine learning algorithm model through optimization algorithms to obtain a trained prediction model; Step 4: Construct and optimize constraints to obtain the optimal solution set. Input the optimal solution set into the trained prediction model to obtain the optimal alloy composition. Step 5: Perform laser powder bed melting and forming according to the optimal alloy composition to obtain an iron-based high-temperature alloy; Example 2: The difference between the iron-based superalloy design method based on multi-model machine learning described in Example 1 and the following is: Step 1 is as follows: Within the common composition range of iron-based superalloys (e.g., Ni: 40–60%, Cr: 10–20%, Co: 10–20%, Al: 1–4%, Ti: 1–4%, Ta: 1–6%, W: 0–5%, Mo: 0–3%), several groups (3000 groups) of alloy compositions were randomly generated in increments of 0.1 wt.%. High-throughput thermodynamic calculations were performed on several composite gold compositions using thermodynamic calculation software (such as Thermo-Calc) to obtain non-equilibrium solidification paths and equilibrium phase transition diagrams. Non-equilibrium solidification pathways include liquidus temperature, solidus temperature, and solid fraction. The equilibrium phase transformation diagram includes the volume fraction of precipitated phases and the volume fraction of γ matrix during alloy solidification; The cracking index of the alloy composition is calculated by using the non-equilibrium solidification path and equilibrium phase transformation diagram, including the solidification temperature range (FR), the solidification cracking factor (SCI), and the strain aging crack index (SAC). FR (freezing range) assesses the hot cracking tendency of an alloy by measuring the temperature difference between the liquidus and solidus. The larger the solidification temperature range, the greater the hot cracking tendency of the alloy. The calculation is shown below: FR=T l -T S ; Among them, T l T represents the liquidus temperature. S Indicates the solidus temperature; The SCI (solidification cracking index) is used to evaluate the hot cracking tendency of an alloy by the derivative of the square root of the solidification end temperature with respect to the solid content. The larger the value, the greater the hot cracking tendency of the alloy. ; Where T represents temperature, and fs represents the volume fraction of precipitated phases during alloy solidification; SAC (strain age cracking index) indirectly obtains the precipitation rate of the γ phase by measuring the rate of change of the γ phase, thereby judging the stress aging cracking tendency of the alloy. The larger the value, the greater the cracking tendency of the alloy. ; in, This represents the volume fraction of the γ matrix. This indicates the critical temperature at which the volume fraction of the γ phase reaches 0.7. The γ phase composition is Ni3Ti, which is an important intermetallic compound strengthening phase in precipitation-strengthened iron-based superalloys and can precipitate during the LPBF forming process. A large-scale, high-quality dataset was constructed by combining alloy composition and corresponding cracking indices FR, SCI, and SAC.

[0034] Perform correlation analysis on the dataset; including: Several combinations of gold components in the dataset (i.e., all component variables mentioned in the dataset, including the content of elements such as Ni, Cr, Co, Al, Ti, Ta, W, and Mo) are used as component variables, and FR, SCI, and SAC are used as target variables, respectively. Correlation analysis is performed to calculate the Pearson correlation coefficient and the Spearman correlation coefficient. The Pearson correlation coefficient r is shown below: ; in, and Represent the component variable and the target variable (FR, SCI, or SAC), respectively. One observation (actual value); and These are the sample means of the two variables, respectively. It refers to the number of samples (the number of alloy components). Spearman correlation coefficient As shown below: ; in, It is each pair of observations The rank difference (sorting the data of two variables to determine the rank of each data point within its respective variable); this formula is used when there are no data points with equal ranks; its range is also [missing information]. .

[0035] Feature importance was ranked using a random forest model, and key components were selected based on correlation analysis; including: A random forest model was used to calculate feature importance. Alloy composition in the dataset was used as a feature, and FR, SCI, and SAC were used as target variables. These were input into the random forest model for training. Mean Decrease in Impurity (MDI) based on node purity decrease and Permutation Importance (PI) based on permutation were calculated. Figure 2 As shown; Mean Impurity Reduction (MDI) includes: In a random forest model, for a given set of features... The importance of node m (the decision point for data partitioning in any decision tree during the growth process of a random forest) is calculated as follows: (Alloy composition) ; in, It is the proportion of samples reaching node m, i.e., the node weight. It is the impurity (such as variance) of the parent node. and It refers to the impurity of the left and right child nodes after the split; characteristic The MDI in the entire random forest model is shown below (features used in all trees). (The sum of the importance of the nodes that are splitting, then averaged). ; in, Indicates the calculated MDI, This represents the total number of decision trees in the random forest model. This represents a single decision tree in a forest. This indicates a specific node in the decision tree. The features selected for splitting; Indicates at node Features used The amount of impurity reduction resulting from the splitting process; The importance of features (PI) includes calculating the performance score of the random forest model, as shown below: ; in, This represents the performance score of the random forest model. Let M represent the evaluation function (e.g., R², 1-MSE, etc.), M represent the random forest model, and D represent the input of the random forest model. The features in the random forest model are permuted; this breaks the feature set. Any relationship between the shuffled features and the true label is preserved, while keeping other features and labels unchanged; the shuffled features and the target variable form a new input. Use new input Recalculate the performance score of the random forest model: ; in, The new performance score is represented; the importance of the ranked features is the difference between the baseline score and the score after the disruption, as shown below: ; in, Indicates the importance of arrangement features; Repeat the above steps of randomly shuffling features, calculating performance scores, and calculating the importance of ranked features (e.g., repeat 5 times) to obtain the average value of the importance of ranked features, so as to obtain a more stable estimate as the final importance PI of ranked features. In the correlation analysis, the top k features with the largest absolute values ​​are selected, and the top k features with the largest MDI and PI are selected respectively as key components (alloy components that meet the selection criteria and their corresponding cracking indices). The final key component features will be used as input variables for subsequent multi-model integrated modeling (such as XGBoost, GBDT, RF, etc.) to improve the accuracy, stability and generalization ability of model prediction, while avoiding interference from invalid features and improving computational efficiency.

[0036] Step 3 specifically involves: The machine learning algorithm models are XGBoost and RF (Random Forest) models; The key components are divided into training and test sets, and XGBoost and RF models are used for training respectively. This invention trains XGBoost and RF models on the same training set respectively and compares their fitting effects in the validation stage, thereby selecting the algorithm with better prediction performance as the final model. During the training of XGBoost and RF models, Bayesian optimization is used to automatically search within a predefined hyperparameter space. Bayesian optimization, as an existing global hyperparameter optimization technique, can efficiently find the optimal parameter combination with fewer trials. The optimal model hyperparameters for the three prediction objectives (FR, SCI, and SAC) are determined, resulting in models corresponding to the three prediction objectives, denoted as the FR prediction model, SCI prediction model, and SAC prediction model, respectively. Figure 3 As shown.

[0037] The model's performance on the test set is measured by the coefficient of determination (R²), and the model's performance on the test set is as follows: FR prediction model: R² = 0.88 on the test set; SCI prediction model: R² = 0.85 on the test set; SAC prediction model: R² = 0.89 on the test set; The following conclusions can be drawn from Figure 3: (1) All three target models have high coefficients of determination (R²>0.85), indicating that the models can effectively capture the nonlinear mapping relationship between components and target values; (2) The point cloud is basically distributed near the diagonal, indicating that the predicted value is highly consistent with the true value, and the model has good generalization ability; (3) SAC and FR have relatively higher prediction accuracy, indicating that the contribution patterns of components to these two indicators are more easily learned by the model; (4) Overall, the established multi-model system has stability and reliability and can be used as a basic prediction tool for subsequent component design and optimization; The model was subjected to SHAP analysis, as shown below:

[0038] in, Representation of features For the sample The SHAP value, i.e., the feature Contribution to this forecast Represents the set of all features; Indicates that it does not contain features A subset of features; Indicates using only subsets When considering features in a sample, The predicted values ​​of (samples in the dataset), Indicates when features Join a subset Then, for the sample The predicted value; Representation of features Join the alliance That is, subset The marginal contribution it brings; Represents a weight term used for all possible subsets. By using a weighted average, it is ensured that all possible orderings of features are considered fairly. After completing the SHAP analysis, to further explore the interactions between features, multinomial feature engineering was performed on the original dataset. That is, for a dataset with... Data set with features, The polynomial expansion generates new features of degree n, and the total number of new features is n; ; Where n represents the number of samples in the dataset; After polynomial feature engineering, we obtain an expanded new feature set and its total number. The total number of features is used to quantify the expansion of the feature space, while the new feature set will be used to retrain or optimize the model to capture potential nonlinear and interaction effects between the original features, thereby potentially improving model performance. After polynomial feature engineering, the model can effectively identify nonlinear and interaction effects between components; among them, the interaction term (Al*C) has the most significant impact on the model output, and its higher eigenvalue corresponds to a higher SHAP value, indicating that the synergistic effect of Al and C is significantly positively correlated with FR; in addition, features such as Ti–Nb and Cr² also show high importance, indicating that these elements play a key role in alloy strengthening and microstructure stability.

[0039] Step 4 is as follows: Target constraints are established for FR, SCI, and SAC: FR ≥ 200 MPa, SCI ≥ 1800 MPa, and SAC ≥ 0.15. (Referring to existing data on iron-based and iron-nickel-based additive manufacturing high-temperature alloys, FR≈200 provides an acceptable process window for suppressing hot cracking; SCI≈1800 aims to ensure the alloy has low hot cracking susceptibility; SAC≤0.15 is considered a low threshold for aging-induced cracking susceptibility, therefore 0.15 is used as a safety upper limit.) Boundary constraints are also established for the alloy composition: Ni: 30–40%, Cr: 5–15%, Al: 2–4%, Ti: 3–8%, W: 1–3%, Mo: 1–3%, Mn: 0–1%, Si: 0–1%, B: 0–0.5%; (Based on the traditional GH2135 alloy composition, and considering cost and performance settings) The multi-objective optimization algorithm NSGA-II is used to optimize the objective constraints and value boundary constraints to obtain a set of uniformly distributed and well-converged Pareto optimal solutions; The optimal solution set is input into the trained FR, SCI, and SAC prediction models to predict the optimal solution set. At the same time, through the non-dominated sorting mechanism and the crowding distance selection mechanism, iterative updates are made to continuously generate new components with better overall performance (which can be combined with genetic operations such as crossover and mutation). After the prediction model converges, a Pareto front solution set is obtained, which is a set of optimal candidate components that cannot be improved simultaneously on the three objectives and are mutually balanced. Based on actual needs (such as adaptability to laser powder bed melting (LPBF) process, material cost, and microstructure stability), the Pareto front solution set is further screened to obtain the final optimal alloy composition, namely the iron-based superalloy AM-SD.

[0040] LBPF forming of AM-SD alloy powder yields an iron-based superalloy; including: AM-SD alloy powder was obtained by gas atomization, with a particle size distribution of 15~53μm and an average particle size of 30-40μm. The AM-SD alloy powder was formed by laser powder bed melting (LPBF) using a metal additive manufacturing equipment (AVIMETAL MT170). The laser power was 180-200W, the scanning speed was 800-1000mm / s, the powder bed thickness (the thickness of each AM-SD layer in the LPBF forming process) was 30-50μm, and the scanning spacing was 100-120μm. The laser power was 190W, the scanning speed was 900mm / s, the layer thickness was 40μm, and the scanning spacing was 110μm. Compared with the GH2135 alloy formed by LPBF, the AM-SD alloy showed no cracks. Figure 6 As shown;

[0041] The crack-free AM-SD alloy, GH4169 alloy, and AM-SD alloy obtained tensile curves through room temperature and high temperature tensile tests as described in Example 1 are shown below. Figure 7 As shown, the AM-SD alloy has significant advantages in both strength and ductility compared to the GH4169 alloy.

[0042] Example 3: A multi-model machine learning-based iron-based superalloy is obtained through the aforementioned multi-model machine learning-based iron-based superalloy design method. The powder composition by mass percentage includes: C<0.08%, Cr 9.00-15.00%, Ni 35.00-40.00%, Mo 1.50-2.50%, W 1.50-2.50%, Al 1.50-3.00%, Ti 3.00-10.00%, B 0.15-0.30%, Si 0.30-0.50%, Mn 0.20-0.40%, S<0.02%, P<0.02%, O<0.02%, N<0.02%, with the balance being Fe and unavoidable impurities.

[0043] The particle size of the iron-based superalloy, namely AM-SD alloy powder, is 15~53μm.

Claims

1. A design method for iron-based superalloys based on multi-model machine learning, characterized in that, include: Step 1: Obtain the composition of iron-based high-temperature alloys and perform high-throughput calculations to build a dataset; Step 2: Perform correlation analysis on the dataset and rank the features by importance based on the random forest model to select key components; Step 3: Build a machine learning algorithm model, train it using key components, and fine-tune the hyperparameters of the machine learning algorithm model through optimization algorithms to obtain a trained prediction model; Step 4: Construct and optimize constraints to obtain the optimal solution set. Input the optimal solution set into the trained prediction model to obtain the optimal alloy composition. Step 5: Perform laser powder bed melting and forming according to the optimal alloy composition to obtain an iron-based high-temperature alloy.

2. The iron-based superalloy design method based on multi-model machine learning as described in claim 1, characterized in that, Step 1 is as follows: Several composite gold components are generated; High-throughput thermodynamic calculations were performed on several composite gold compositions using thermodynamic calculation software to obtain non-equilibrium solidification paths and equilibrium phase transition diagrams. Non-equilibrium solidification pathways include liquidus temperature, solidus temperature, and solid fraction. The equilibrium phase transformation diagram includes the volume fraction of precipitated phases and the volume fraction of γ matrix during alloy solidification; The cracking index of the alloy composition is calculated by using the non-equilibrium solidification path and equilibrium phase transformation diagram, including the solidification temperature range (FR), the solidification cracking factor (SCI), and the strain aging crack index (SAC). The alloy composition and the corresponding cracking indices FR, SCI and SAC were constructed into a dataset.

3. The iron-based superalloy design method based on multi-model machine learning as described in claim 2, characterized in that, Perform correlation analysis on the dataset; including: Using several gold components in the dataset as component variables, and FR, SCI and SAC as target variables, correlation analysis was performed to calculate the Pearson correlation coefficient and the Spearman correlation coefficient. The Pearson correlation coefficient r is shown below: ; in, and These represent the first and second components of the variable and the target variable, respectively. One observation value; and These are the sample means of the two variables, respectively. It is the sample size; Spearman correlation coefficient As shown below: ; in, It is each pair of observations The difference in grade.

4. The iron-based superalloy design method based on multi-model machine learning as described in claim 3, characterized in that, Feature importance was ranked using a random forest model, and key components were selected based on correlation analysis; including: The random forest model was used to calculate feature importance. The alloy composition in the dataset was used as the feature, and FR, SCI and SAC were used as the target variables. The random forest model was trained and the MDI based on node purity decrease and the importance PI based on permutation were calculated. Mean Impurity Reduction (MDI) includes: In a random forest model, for a given set of features... The importance of node m, which is to be split, is calculated as follows: ; in, It is the proportion of samples reaching node m, i.e., the node weight. It is the impurity of the parent node. and It refers to the impurity of the left and right child nodes after the split; characteristic The MDI in the entire random forest model is shown below: ; in, Indicates the calculated MDI, This represents the total number of decision trees in the random forest model. This represents a single decision tree in a forest. This indicates a specific node in the decision tree. The features selected for splitting; Indicates at node Features used The amount of impurity reduction resulting from the splitting process; The importance of features (PI) is calculated by assigning a performance score to the random forest model, as shown below: ; in, This represents the performance score of the random forest model. Let M represent the evaluation function, M represent the random forest model, and D represent the input to the random forest model. The features in the random forest model are randomly shuffled; the shuffled features and the target variable form a new input. Use new input Recalculate the performance score of the random forest model: ; in, The new performance score is represented; the importance of the ranked features is the difference between the baseline score and the score after the disruption, as shown below: ; in, Indicates the importance of arrangement features; Repeat the above steps of randomly shuffling features, calculating performance scores, and calculating the importance of ranked features to obtain the average value of the ranked feature importance, which is used as the final ranked feature importance PI. In the correlation analysis, the top k features with the largest absolute values ​​were selected, and the top k features with the largest MDI and PI were selected as key components.

5. The iron-based superalloy design method based on multi-model machine learning as described in claim 4, characterized in that, Step 3 specifically involves: The machine learning algorithm models are XGBoost and RF models; The key components were divided into training and testing sets, and trained using XGBoost and RF models respectively. During the training of XGBoost and RF models, Bayesian optimization is used to automatically search in a preset hyperparameter space; the optimal model hyperparameters corresponding to the three prediction objectives FR, SCI and SAC are determined respectively, and the models corresponding to the three prediction objectives are obtained, which are denoted as FR prediction model, SCI prediction model and SAC prediction model respectively.

6. The iron-based superalloy design method based on multi-model machine learning as described in claim 5, characterized in that, Step 4 specifically involves: Target constraints for FR, SCI, and SAC are constructed: FR ≥ 200 MPa, SCI ≥ 1800 MPa, and SAC ≥ 0.

15. Boundary constraints for alloy composition are also constructed: Ni: 30–40%, Cr: 5–15%, Al: 2–4%, Ti: 3–8%, W: 1–3%, Mo: 1–3%, Mn: 0–1%, Si: 0–1%, and B: 0–0.5%. The multi-objective optimization algorithm NSGA-II is used to optimize the target constraints and boundary constraints to obtain the optimal solution set. The optimal solution set is input into the trained FR, SCI, and SAC prediction models to predict the optimal solution set. At the same time, through the non-dominated ranking mechanism and the crowding distance selection mechanism, iterative updates are continuously generated to generate new components with better overall performance. After the prediction model converges, a Pareto front solution set is obtained, which is a set of optimal candidate components that cannot be improved simultaneously on the three objectives and are mutually balanced. Based on actual needs, the Pareto front solution set is further screened to obtain the final optimal alloy composition, namely the iron-based high-temperature alloy AM-SD.

7. The iron-based superalloy design method based on multi-model machine learning as described in claim 6, characterized in that, LBPF forming of AM-SD alloy powder yields an iron-based superalloy; including: AM-SD alloy powder was obtained by gas atomization, with a particle size distribution of 15~53μm and an average particle size of 30-40μm. AM-SD alloy powder is formed by laser powder bed melting (LPBF) using metal additive manufacturing equipment. The laser power is 180-200W, the scanning speed is 800-1000mm / s, the powder bed thickness is 30-50μm, and the scanning spacing is 100-120μm.

8. The iron-based superalloy design method based on multi-model machine learning as described in claim 7, characterized in that, Laser power 190W, scanning speed 900mm / s, layer thickness 40μm, scanning spacing 110μm.

9. A multi-model machine learning-based iron-based superalloy, obtained through the above-mentioned multi-model machine learning-based iron-based superalloy design method, characterized in that, The alloy powder composition by mass percentage includes: C <0.08%, Cr 9.00-15.00%, Ni 35.00-40.00%, Mo 1.50-2.50%, W 1.50-2.50%, Al 1.50-3.00%, Ti 3.00-10.00%, B 0.15-0.30%, Si 0.30-0.50%, Mn 0.20-0.40%, S <0.02%, P <0.02%, O <0.02%, N <0.02%, with the balance being Fe and unavoidable impurities.

10. The design method for iron-based superalloys based on multi-model machine learning as described in claim 1, characterized in that, The particle size of the iron-based superalloy, namely AM-SD alloy powder, is 15~53μm.

Citation Information

Cited By

  • Method for regulating and controlling type of iron-containing phase in secondary aluminum, medium and computer equipment

    CN122024897A