A method for designing the composition and process of ductile iron based on an active learning strategy

The construction of ductile iron model through active learning strategies and small sample machine learning algorithms solves the problems of low data utilization and multi-objective design in traditional design methods, and achieves efficient and accurate ductile iron composition and process design, meeting a variety of performance requirements.

CN119811526BActive Publication Date: 2025-07-04HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411880238.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-07-04
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

The existing ductile iron design methods rely on experience, resulting in low data utilization, incomplete design content, inability to take into account multiple performance goals at the same time, and lack of feedback optimization mechanisms. The design process is time-consuming and labor-intensive and the design results are not ideal.

Method used

Based on active learning strategy, a ductile iron component-process-microstructure characteristics-performance model is constructed through a small sample machine learning algorithm. Combined with optimization algorithm and experimental feedback, the model performance is optimized, covering a comprehensive component-process-organization-performance relationship, and solving the problem of uneven data distribution.

Benefits of technology

It improves the prediction accuracy and efficiency of ductile iron design, reduces the number of experiments and costs, and can optimize the composition ratio and process parameters according to a variety of design goals to meet practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811526B_ABST
    Figure CN119811526B_ABST
Patent Text Reader

Abstract

A method for designing the composition and process of ductile iron based on an active learning strategy, belonging to the technical field of material process design. To improve the prediction accuracy and design efficiency of material process design, the present invention includes collecting experimental data including the composition, process, microstructure characteristics, and properties of ductile iron in the literature, constructing a ductile iron composition-process-microstructure characteristic model and a ductile iron composition-process-microstructure characteristic-property model based on a small-sample machine learning algorithm, setting different weights for different target properties based on an optimization algorithm, using a weight-based objective function to obtain optimized composition and process parameters of ductile iron with different performance requirements; through an active learning strategy, preferentially selecting data with a larger uncertainty for experiments and supplementing them to the data set, then updating the model to optimize the model prediction accuracy, and finally designing the composition and process parameters of ductile iron that meet the expected performance requirements. The present invention reduces the number of experiments and costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of material process design, and particularly relates to a method for designing the composition and process of ductile iron based on an active learning strategy. Background Art

[0002] In recent years, driven by domestic energy demand and industrial needs, the large-scale of cast iron structures has received particular attention. As an important casting material, ductile iron material has good strength, toughness and fatigue resistance, and is widely used in fields such as automobiles, buildings, pipelines, etc., and is often used as a key component of various cast iron castings, so it is highly favored. Many devices usually operate in very harsh environments. Therefore, designing ductile iron that meets multiple properties has become a key issue in related industries.

[0003] However, the traditional method for designing ductile iron almost completely relies on experience and requires repeated experimental adjustments. This method not only consumes time and effort, but also is difficult to quickly optimize the material properties because the interaction of each parameter cannot be accurately grasped. Therefore, the existing ductile iron design mainly faces the following problems:

[0004] 1. Low data utilization rate: In the design process based on data modeling, a large amount of experimental data is usually relied on, but these data are often unevenly distributed. Under the existing technical methods, some combinations of key parameters are not fully covered, and the design results are not ideal.

[0005] 2. Incomplete design content: The traditional design method hardly involves the comprehensive relationship between composition - process - microstructure - properties, but only involves several of them, so it is impossible to effectively predict the microstructure and properties of cast iron under different compositions and processes, resulting in a slow design process and insufficient design content.

[0006] 3. Unable to meet multi-objective requirements: The existing design method is difficult to simultaneously consider multiple performance objectives, such as the balanced requirements of tensile strength, toughness and hardness in different application scenarios, and often requires a large number of experiments for verification, increasing the complexity and cost of design.

[0007] 4. Lack of feedback optimization mechanism: The existing material design method cannot perform dynamic optimization through the feedback of the experiments that have been carried out, resulting in a deviation between the model prediction and the actual result, and unable to effectively reduce the design error. Summary of the Invention

[0008] The problem to be solved by the present invention is to improve the prediction accuracy and design efficiency of material process design, and propose a method for designing the composition and process of ductile iron based on an active learning strategy.

[0009] To achieve the above object, the present invention is realized through the following technical solutions:

[0010] A method for designing the composition and process of ductile iron based on an active learning strategy, comprising the following steps:

[0011] S1. Collect experimental data including the composition, process, microstructure characteristics, and properties of ductile iron in the literature. Add the associated data with the composition, process, microstructure characteristics, and properties of ductile iron in the experimental data to dataset 1, and add the associated data with only the composition, process, and properties of ductile iron in the experimental data to dataset 2;

[0012] S2. Visualize the distribution of the composition or process data of ductile iron in dataset 1, and then conduct an orthogonal experiment based on the blank area of the visualized composition or process data of ductile iron, and add the obtained orthogonal experiment data to dataset 1;

[0013] S3. Preprocess the data in dataset 1 and dataset 2 to obtain preprocessed dataset 1 and preprocessed dataset 2;

[0014] S4. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microstructure characteristics model, set the composition and process of ductile iron as the input, and the microstructure characteristics as the output. Use the preprocessed dataset 1 to train the ductile iron composition-process-microstructure characteristics model, and use the trained ductile iron composition-process-microstructure characteristics model to predict the preprocessed dataset 2 to obtain the microstructure characteristics corresponding to each group of data in the preprocessed dataset 2, and supplement the obtained microstructure characteristic data to the preprocessed dataset 2 after normalization;

[0015] S5. Merge the preprocessed dataset 1 and the preprocessed dataset 2 obtained in step S4 into dataset 3;

[0016] S6. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microstructure characteristics-performance model, set the composition, process, and microstructure characteristics of ductile iron as the input, and the performance as the output. Use dataset 3 to train the ductile iron composition-process-microstructure characteristics-performance model to obtain a trained ductile iron composition-process-microstructure characteristics-performance model;

[0017] S7. Based on the optimization algorithm, set different weights for different target performances, and calculate each target performance using the trained ductile iron composition-process-microstructure characteristics-performance model in step S6. Adopt a weight-based objective function to continuously screen the composition closest to the target performance to obtain optimized ductile iron compositions and process parameters for different performance requirements;

[0018] S8. For the optimized ductile iron compositions and process parameters with different performance requirements obtained in step S7, through an active learning strategy, preferentially select data with a larger uncertainty for experiments and supplement them to the dataset, then return to step S4 for model update to optimize the model prediction accuracy, and finally design ductile iron compositions and process parameters that meet the expected performance requirements.

[0019] Further, the data of ductile iron compositions in step S1 mainly include carbon, silicon, manganese, rare earths, sulfur, phosphorus, copper, nickel, magnesium, and chromium, the process data mainly include cooling rate, inoculant type, inoculant content, spheroidizing agent type, spheroidizing agent content, inoculation temperature, inoculation time, and the data of microstructural characteristics mainly include graphite ball diameter, graphite ball number, graphite spheroidization rate, pearlite ratio in the matrix, ferrite ratio in the matrix, and spheroidization grade, and the performance data mainly include tensile strength, yield strength, ductility, and impact toughness.

[0020] Further, in step S2, select the top two with the most microstructural characteristic data in dataset 1 as the output features of the model or customize the microstructural characteristics in dataset 1 as the output features of the model, visualize the distribution of ductile iron composition or process data associated with the selected or customized microstructural characteristic data, and extract the blank areas.

[0021] Further, the preprocessing in step S3 includes data cleaning, normalization processing, and encoding processing. The formula for normalization processing is:

[0022]

[0023] where X is the original data, X min and X max are the minimum and maximum values in the dataset respectively, and X' is the normalized data;

[0024] The encoding processing is to perform encoding processing on the inoculant type and spheroidizing agent type.

[0025] Further, the specific implementation method of step S4 includes the following steps:

[0026] S4.1. Based on small-sample machine learning algorithms, construct a ductile iron composition-process-microstructural characteristic model, set the ductile iron composition and process as inputs, and the microstructural characteristics as outputs. The small-sample machine learning algorithms include the extremely randomized trees algorithm, the CatBoost algorithm, the random forest algorithm, and the adaptive boosting regression algorithm;

[0027] S4.2. Set methods such as k-fold cross-validation to evaluate the performance of the model, where k = 5 - 10, and select the best-performing nodular cast iron composition-process-microstructure feature model. When predicting multiple microstructure features, adopt the strategy of establishing separate models;

[0028] S4.3. Set the evaluation criteria to include the mean absolute error MAE, root mean square error RMSE, coefficient of determination R 2 and mean absolute percentage error MAPE, and the calculation formulas are as follows:

[0029]

[0030]

[0031]

[0032]

[0033] where, represents the predicted value of the machine learning model, and y i represents the true value;

[0034] Then select the algorithm with the best evaluation criteria as the final trained nodular cast iron composition-process-microstructure feature model.

[0035] Furthermore, in the extremely randomized trees algorithm in step S4.1, the number of trees is 100, the maximum depth is 10, and the minimum number of samples required for splitting each node is 2; in the CatBoost algorithm, the number of iterations for training is 300, the learning rate is 0.02, and the L2 regularization coefficient is 3; in the random forest algorithm, the number of trees in the forest is 100, the minimum number of samples for splitting is 2, and the minimum number of samples in the leaf nodes is 2; in the adaptive boosting regression algorithm, the number of base models is 80, the learning rate is 0.02, and the loss function type is linear.

[0036] Furthermore, the specific implementation method of step S6 includes the following steps:

[0037] S6.1. Based on the small-sample machine learning algorithm, construct a nodular cast iron composition-process-microstructure feature-performance model, set the nodular cast iron composition, process, and microstructure features as inputs, and the performance as the output. The small-sample machine learning algorithm includes the extremely randomized trees algorithm, CatBoost algorithm, random forest algorithm, and adaptive boosting regression algorithm;

[0038] S6.2. Set methods such as k-fold cross-validation to evaluate the performance of the model, where k = 5 - 10, and select the best-performing ductile iron composition-process-microstructure feature-performance model. When predicting multiple performance features, adopt the strategy of establishing models separately;

[0039] S6.3. Set the evaluation criteria to include the mean absolute error MAE, root mean square error RMSE, coefficient of determination R 2 and mean absolute percentage error MAPE.

[0040] Furthermore, in step S7, intelligent optimization of composition and process is carried out through the particle swarm optimization algorithm. The objective function f of the particle swarm optimization algorithm is defined as:

[0041]

[0042] where K i is the weight of the i-th objective, T i is the target value, and T' i is the model prediction value.

[0043] Furthermore, in step S8, the calculation method of uncertainty is as follows: Under the condition that other things remain unchanged, change the hyperparameters of the composition-process-microstructure feature-performance model, and calculate the performance prediction values of the composition and process optimized by the particle swarm optimization algorithm in the previous step respectively. The uncertainty U is the variance of the prediction values of different models, and the expression is:

[0044]

[0045] where is the average prediction value of all models, is the prediction value of the m-th model, and M is the total number of models.

[0046] Advantages of the present invention:

[0047] A method for designing the composition and process of ductile iron based on an active learning strategy according to the present invention realizes dynamic optimization of the model performance through experimental feedback by introducing the active learning strategy; through a step-by-step method, it covers the comprehensive composition-process-structure-performance relationship; through supplementary sampling, it solves the problem of uneven data distribution, effectively improving the prediction accuracy and design efficiency of the model. At the same time, this method can intelligently optimize the composition ratio and process parameters of ductile iron according to various different design goals (such as tensile strength, impact toughness, etc.), effectively serving the actual application requirements.

[0048] A method for designing the composition and process of ductile iron based on an active learning strategy according to the present invention improves the design efficiency of ductile iron, reduces the number of experiments and costs, solves many drawbacks of the existing data modeling design methods, and has good promotion value. Brief Description of the Drawings

[0049] Figure 1 It is a flowchart of a method for designing the composition and process of ductile iron based on an active learning strategy according to the present invention;

[0050] Figure 2 It is a visualization diagram of the data distribution of C, Si, and Mn in Embodiment 2 of the present invention;

[0051] Figure 3 It is a visualization diagram of the data distribution of Cu, Ni, and Re in Embodiment 2 of the present invention. Detailed Embodiments

[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention usually described and shown in the drawings here can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0053] Therefore, the detailed description of the specific embodiments of the present invention provided in the drawings below is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0054] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and are accompanied by the attached Figure 1 - Attached Figure 3 The details are as follows:

[0055] Embodiment 1:

[0056] A method for designing the composition and process of ductile iron based on an active learning strategy includes the following steps:

[0057] S1. Collect experimental data including the composition, process, microstructure characteristics and properties of ductile iron in the literature, add the associated data with the composition, process, microstructure characteristics and properties of ductile iron in the experimental data to Dataset 1, and add the associated data with only the composition, process and properties of ductile iron in the experimental data to Dataset 2;

[0058] Further, the data of the ductile iron composition in step S1 mainly includes carbon, silicon, manganese, rare earth, sulfur, phosphorus, copper, nickel, magnesium and chromium, and the data of the process mainly includes cooling rate, inoculant type, inoculant content, spheroidizing agent type, spheroidizing agent content, inoculation temperature, inoculation time. The data of the microscopic tissue characteristics mainly includes graphite ball diameter, graphite ball number, graphite spheroidization rate, pearlite ratio in the matrix, ferrite ratio in the matrix, spheroidization grade. The data of the performance mainly includes tensile strength, yield strength, ductility and impact toughness;

[0059] S2. Visualize the distribution of the ductile iron composition or process data in dataset 1, and then conduct an orthogonal experiment based on the blank area of the visualized ductile iron composition or process data, and add the obtained orthogonal experiment data to dataset 1;

[0060] Further, in step S2, select the top two with the most microscopic tissue characteristic data in dataset 1 as the output features of the model or customize the microscopic tissue characteristics in dataset 1 as the output features of the model, visualize the ductile iron composition or process associated with the selected or customized microscopic tissue characteristic data, and extract the blank area;

[0061] S3. Preprocess the data in dataset 1 and dataset 2 to obtain the preprocessed dataset 1 and the preprocessed dataset 2;

[0062] Further, the preprocessing in step S3 includes data cleaning, normalization processing and encoding processing. The formula for the normalization processing is:

[0063]

[0064] where X is the original data, X min and X max are the minimum value and the maximum value in the dataset respectively, and X' is the normalized data;

[0065] The encoding processing is to encode the inoculant type and the spheroidizing agent type;

[0066] S4. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microscopic tissue characteristic model, set the ductile iron composition and process as the input, and the microscopic tissue characteristics as the output. Use the preprocessed dataset 1 to train the ductile iron composition-process-microscopic tissue characteristic model, and use the trained ductile iron composition-process-microscopic tissue characteristic model to predict the preprocessed dataset 2 to obtain the microscopic tissue characteristics corresponding to each group of data in the preprocessed dataset 2, and supplement the obtained microscopic tissue characteristic data to the preprocessed dataset 2 after normalization;

[0067] Further, the specific implementation method of step S4 includes the following steps:

[0068] S4.1. Based on small-sample machine learning algorithms, construct a nodular cast iron composition-process-microstructure feature model, set the nodular cast iron composition and process as inputs, and the microstructure features as outputs. The small-sample machine learning algorithms include the extremely randomized trees algorithm, the CatBoost algorithm, the random forest algorithm, and the adaptive boosting regression algorithm;

[0069] Further, in the extremely randomized trees algorithm in step S4.1, the number of trees is 100, the maximum depth is 10, and the minimum number of samples required for splitting each node is 2; in the CatBoost algorithm, the number of iterations for training is 300, the learning rate is 0.02, and the L2 regularization coefficient is 3; in the random forest algorithm, the number of trees in the forest is 100, the minimum number of samples for splitting is 2, and the minimum number of samples in the leaf nodes is 2; in the adaptive boosting regression algorithm, the number of base models is 80, the learning rate is 0.02, and the loss function type is linear;

[0070] Further, the random forest is composed of multiple decision trees, and the final output is the mean of the predictions of each tree. The output of the random forest can be expressed as:

[0071]

[0072] where Y i (x) represents the prediction result of the i-th tree, and N is the number of decision trees.

[0073] AdaBoost trains the model by updating the weights of the learners, and its model output is:

[0074]

[0075] where α t is the weight of the t-th base learner, and h t (x) is the prediction result of the t-th weak learner.

[0076] CatBoost uses decision trees based on gradient boosting for regression to minimize the loss function, and the loss function is expressed as:

[0077]

[0078] S4.2. Set methods such as k-fold cross-validation to evaluate the performance of the model, where k = 5 - 10, and select the best-performing nodular cast iron composition-process-microstructure feature model. When predicting multiple microstructure features, adopt the strategy of establishing models separately;

[0079] S4.3. Set the evaluation criteria to include the Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Coefficient of Determination (R 2 and Mean Average Percentage Error (MAPE). The calculation formulas are as follows:

[0080]

[0081] where, represents the predicted value of the machine learning model, and y i represents the true value;

[0082] Then select the algorithm with the best evaluation criteria as the final trained nodular cast iron composition - process - microstructure feature model

[0083] S5. Merge the pre - processed dataset 1 and the pre - processed dataset 2 obtained in step S4 into dataset 3;

[0084] S6. Based on the small - sample machine learning algorithm, construct a nodular cast iron composition - process - microstructure feature - property model. Set the nodular cast iron composition, process, and microstructure features as inputs and the property as the output. Use dataset 3 to train the nodular cast iron composition - process - microstructure feature - property model to obtain the trained nodular cast iron composition - process - microstructure feature - property model;

[0085] Furthermore, the specific implementation method of step S6 includes the following steps:

[0086] S6.1. Based on the small - sample machine learning algorithm, construct a nodular cast iron composition - process - microstructure feature - property model. Set the nodular cast iron composition, process, and microstructure features as inputs and the property as the output. The small - sample machine learning algorithms include the Extremely Randomized Trees algorithm, CatBoost algorithm, Random Forest algorithm, and Adaptive Boosting Regression algorithm;

[0087] Furthermore, in the Extremely Randomized Trees algorithm, n_estimators (the number of trees) is 100, the maximum depth is 10, and the minimum number of samples required for each node split is 2; in the CatBoost algorithm, the number of iterations for training is 300, the learning_rate is 0.02, and the L2 regularization coefficient is 3; in the Random Forest algorithm, n_estimators: the number of trees in the forest is,00, the minimum number of samples for splitting is 2, and the minimum number of samples for leaf nodes is 2; in the Adaptive Boosting Regression algorithm, the number of n_estimators (the number of base models) is 80, the learning_rate is 0.02, and the loss function type is 'linear';

[0088] S6.2. Set methods such as k-fold cross-validation to evaluate the performance of the model, where k = 5 - 10, and screen out the best-performing nodular cast iron composition-process-microstructure feature-performance model. When predicting multiple performance features, adopt the strategy of establishing separate models;

[0089] S6.3. Set the evaluation criteria to include the mean absolute error MAE, root mean square error RMSE, coefficient of determination R 2 and mean absolute percentage error MAPE;

[0090] S7. Based on the optimization algorithm, set different weights for different target performances, and use the nodular cast iron composition-process-microstructure feature-performance model trained in step S6 for each target performance calculation. Adopt a weighted objective function to continuously screen out the composition closest to the target performance, and obtain the optimized nodular cast iron compositions and process parameters for different performance requirements;

[0091] Further, in step S7, intelligent optimization of the composition and process is carried out through the particle swarm optimization algorithm. The objective function f of the particle swarm optimization algorithm is defined as:

[0092]

[0093] where, K i is the weight of the i-th target, T i is the target value, and T' i is the model prediction value;

[0094] S8. For the optimized nodular cast iron compositions and process parameters with different performance requirements obtained in step S7, through the active learning strategy, preferentially select data with larger uncertainties for experiments and supplement them to the dataset, then return to step S4 for model update to optimize the model prediction accuracy, and finally design nodular cast iron compositions and process parameters that meet the expected performance requirements.

[0095] Further, the calculation method of uncertainty in step S8 is: with other conditions unchanged, change the hyperparameters of the composition-process-microstructure feature-performance model, and calculate the performance prediction values of the composition and process optimized by the previous particle swarm optimization algorithm respectively. The uncertainty U is the variance of the prediction values of different models, and the expression is:

[0096]

[0097] where, is the average prediction value of all models, is the prediction value of the m-th model, and M is the total number of models.

[0098] Example 2:

[0099] Based on Example 1, the target properties of ductile iron are set as follows: the tensile strength is greater than 550 MPa, and the V-notch impact energy at room temperature is greater than 12 J. A detailed example is given as follows:

[0100] This embodiment solves the problems of low data utilization rate, incomplete design content, inability to meet multi-objective design requirements, and lack of feedback optimization mechanism in the existing data modeling design methods.

[0101] S1. Collect experimental data on the composition, process, microstructure characteristics, and properties of ductile iron in the literature. Add the associated data with the composition, process, microstructure characteristics, and properties of ductile iron in the experimental data to Dataset 1, and add the associated data with only the composition, process, and properties of ductile iron in the experimental data to Dataset 2;

[0102] Furthermore, the target properties of ductile iron are set as follows: the tensile strength is greater than 550 MPa, and the V-notch impact energy at room temperature is greater than 12 J. First, using "ductile iron" as the keyword, collect 202 groups of ductile iron data containing tissue information as Dataset 1 and 67 groups of data without tissue information as Dataset 2 from databases such as CNKI. The data cover the chemical composition, melting process, tissue appearance, and various mechanical properties of different ductile irons under the same environmental conditions. The detailed data types are shown in Table 1:

[0103] Table 1 Ductile iron datasets to be collected

[0104]

[0105] S2. Visualize the distribution of the ductile iron composition or process data in Dataset 1, and then conduct an orthogonal experiment based on the blank areas of the visualized ductile iron composition or process data, and add the obtained orthogonal experiment data to Dataset 1;

[0106] Furthermore, orthogonal experiment design: Based on Dataset 1, analyze to find blank or sparse areas in certain chemical compositions or process parameters. Visualize the distribution of carbon-manganese-silicon composition data and copper-nickel-rare earth composition data in the dataset as Figure 2 and Figure 3As shown. It can be found that the data distribution is uneven for carbon elements in the range of 2.2-2.4wt%, silicon elements in the range of 3.0-4.3wt%, manganese elements in the range of 0.5-0.8wt%, and rare earth elements in the range of 0.02-0.05wt%. In order to make up for these blank areas and improve the data quality, this embodiment uses an orthogonal experiment for supplementary sampling. The gradients of carbon, silicon, manganese, and rare earth components are set to 0.1, 0.2, 0.1, and 0.01, respectively, and an orthogonal experiment is designed. The smelting process adopts the impact spheroidization method, the spheroidization temperature is 1450°C, the spheroidization treatment time is 60s, and rare earth magnesium alloy spheroidizer and FeSi75 inoculant are used. The data obtained from the experiment are supplemented to the data set.

[0107] S3. Preprocess the data in data set 1 and data set 2 to obtain preprocessed data set 1 and preprocessed data set 2;

[0108] Furthermore, necessary preprocessing is performed on the collected experimental data to ensure data quality and consistency. Normalization: variables of different scales (such as graphite spheroidization rate and tensile strength) are converted into a unified [0,1] range.

[0109] For some classification features (such as inoculant type and spheroidizer type), coding processing is performed. After preliminary processing to delete invalid, abnormal and duplicate data, the ductile iron data set has a total of 218 data sets, of which 37 data sets do not contain organizational information.

[0110] S4. Based on a small sample machine learning algorithm, a ductile iron composition-process-microstructure feature model is constructed, the ductile iron composition and process are set as input, and the microstructure feature is used as output, and the ductile iron composition-process-microstructure feature model is trained using the preprocessed data set 1, and the preprocessed data set 2 is predicted using the trained ductile iron composition-process-microstructure feature model to obtain the microstructure feature corresponding to each group of data in the preprocessed data set 2, and the obtained microstructure feature data is normalized and added to the preprocessed data set 2;

[0111] Furthermore, the performance of different algorithms is shown in Table 2 and Table 3:

[0112] Table 2 Composition-process-spheroidization grade model indicators

[0113] Model MAE RMSE R2 Percentage % ExtraTrees 0.3332 0.5071 0.7666 14.77 CatBoost 0.3448 0.5172 0.7498 15.52 RandomForest 0.374 0.5471 0.6839 16.81 AdaBoostRegressor 0.4406 0.5453 0.6739 19.64

[0114] Table 3 Composition-processing-pearlite content model index

[0115] Model MAE RMSE <![CDATA[R 2 > Percentage % ExtraTrees 4.6619 6.2455 0.8217 21.63 CatBoost 5.4326 7.387 0.7024 26.08 AdaBoostRegressor 6.1398 7.8395 0.6631 38.82 RandomForest 5.849 7.8791 0.6341 29.73

[0116] By comparing different algorithms, the Extra Trees algorithm with the best performance in each index is selected as the subsequent algorithm. For the 37 groups of data without tissue information in the dataset, the trained composition-process-tissue model is used to predict the missing tissue information respectively. The predicted tissue data is filled into the dataset, and together with the original 218 groups of data, it is used as the input data for the next step.

[0117] S5. Merge the preprocessed dataset 1 and the preprocessed dataset 2 obtained in step S4 into dataset 3;

[0118] S6. Based on the small-sample machine learning algorithm, construct a nodular cast iron composition-process-microstructure feature-performance model. Set the nodular cast iron composition, process, and microstructure features as inputs, and the performance as the output. Use dataset 3 to train the nodular cast iron composition-process-microstructure feature-performance model to obtain the trained nodular cast iron composition-process-microstructure feature-performance model;

[0119] Furthermore, the optimal algorithm is the Extra Trees algorithm, and the performances of the two performance models are shown in Table 4:

[0120] Table 4 Indexes of the optimal composition-process-tissue-performance model

[0121] Target MAE RMSE <![CDATA[R 2 > Percentage % Impact energy / J 1.6027 1.9027 0.8400 15.99 Tensile strength / MPa 22.0486 33.1265 0.9356 5.90

[0122] S7. Based on the optimization algorithm, set different weights for different target performances. For each target performance, use the nodular cast iron composition-process-microstructure feature-performance model trained in step S6 for calculation. Adopt the weight-based objective function, and continuously screen the composition closest to the target performance to obtain the optimized nodular cast iron compositions and process parameters with different performance requirements;

[0123] Furthermore, use the selected optimal machine learning model and the particle swarm optimization algorithm to optimize the composition-process that meets the performance target. Set the targets as the tensile strength of 550 MPa and the V-notch impact energy of 12 J at room temperature, and the performance weights are 0.4 and 0.6 respectively. Set the objective function in the optimization algorithm as the weighted average of different targets, that is:

[0124] f = 0.4 * |T′1 - 550| + 0.6 * |T′2 - 12|

[0125] where T'1 and T'2 are the tensile strength and the V-notch impact energy at room temperature predicted by the model respectively. By continuously iteratively adjusting the chemical composition and process parameters, find 8 groups of composition and process combinations with the minimum objective function.

[0126] S8. For the optimized nodular cast iron compositions and process parameters with different performance requirements obtained in step S7, through the active learning strategy, preferentially select the data with larger uncertainty for experiments and supplement them to the dataset, then return to step S4 for model update to optimize the model prediction accuracy, and finally design nodular cast iron compositions and process parameters that meet the expected performance requirements.

[0127] Further, for the 8 compositions and processes optimized by the previous particle swarm optimization algorithm, calculate the uncertainty measures of their predicted values, and preferentially select the data with larger uncertainty for experiments. The calculation method of uncertainty is as follows: Under the condition that other factors remain unchanged, change the hyperparameters of the composition-process-structure-performance model, that is: keep the maximum depth as 10, the minimum number of samples required for each node split as 2. In the extremely randomized tree algorithm, n_estimators are 100, 140, 180, 220, 260 respectively, calculate the performance predicted values of the 8 compositions and processes respectively, and calculate the uncertainty U (the variance of the predicted values of different models):

[0128]

[0129] Among them, is the average predicted value of all models, is the predicted value of the m-th model.

[0130] Respectively screen out the 3 compositions-processes with the largest uncertainty, and the results are shown in Table 5:

[0131] Table 5 Compositions-processes with the largest predicted value uncertainty screened out

[0132] C Si Mn … Mg RE Process Uncertainty Type 3.92 2.48 0.17 … 0.056 0.043 … 9.4 Impact toughness 3.84 2.46 0.17 … 0.058 0.039 … 8.8 Impact toughness 3.78 2.21 0.25 … 0.062 0.021 … 7.1 Impact toughness 2.23 3.76 0.33 … 0.007 0.009 … 836.44 Tensile strength 2.29 3.09 0.64 … 0.033 0.004 … 817.5 Tensile strength 3.68 3.21 0.62 … 0.013 0.043 … 786.60 Tensile strength

[0133] Conduct experimental verification on the screened compositions-processes, and record the obtained composition, process, structure and performance data. Add the composition-process-structure-performance data obtained after the experiment to the dataset, and re-establish the two-step method model, correct and update the model, so as to improve the prediction accuracy and form a closed-loop optimization process of data-model-experiment.

[0134] Further, use the final model with high accuracy optimized by the active learning strategy to optimize the compositions-processes that meet the performance goals again. Set the goals as tensile strength of 550 MPa and V-notch impact energy of 12 J at room temperature, and the performance weights are 0.4 and 0.6 respectively. Finally, design nodular cast iron compositions and melting process parameters that meet the target performance, and conduct experimental verification. The experiment shows that the designed as-cast nodular cast iron has a tensile strength of 550 MPa and a V-notch impact energy of 12 J at room temperature.

[0135] Example 3:

[0136] The difference between this embodiment and Embodiment 2 is that in the steps of establishing the two-step machine learning model, the method of 5-fold cross-validation is not adopted. Instead, the data set is divided into two parts, where 85% is used as the training set and the other 15% is used as the test set. The algorithms used include the support vector machine algorithm, decision tree algorithm, and logistic regression. To evaluate the performance of different models, the root mean square error (RMSE), coefficient of determination (R 2 ) and mean absolute percentage error (MAPE) are introduced as evaluation criteria.

[0137] Embodiment 4:

[0138] The difference between this embodiment and Embodiment 2 is that the target ductile iron properties are a tensile strength of 500 MPa and an elongation at break greater than 18%. In the data collection step, the collected data covers the chemical composition, melting process, microstructure, and various mechanical properties of different ductile irons under the same environmental conditions. Specifically, the composition (wt.%) includes C, Si, Mn, P, S, Cu, RE, Ni, Cr, Mg; the melting process includes the type of spheroidizing agent, the type of inoculant, and the inoculation time s; the tissue characteristics include the pearlite content %, spheroidization grade, and graphite sphere diameter μm; the mechanical properties include the tensile strength MPa and the elongation at break %. The finally designed ductile iron was tested with a tensile strength of 501 MPa, a yield strength of 331 MPa, and an elongation at break of 21%.

[0139] Embodiment 5:

[0140] The difference between this embodiment and Embodiment 2 is that the target ductile iron properties are a tensile strength of 460 MPa and a V-notch impact value of 7.0 J at -40°C. In the data collection step, the mechanical properties include the tensile strength MPa and the V-notch impact value at -40°C. The finally designed ductile iron was tested with a tensile strength of 458 MPa and a V-notch impact value of 6.8 J at -40°C.

[0141] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0142] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of the situations of these combinations is not given in this specification only for the consideration of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for designing the composition and process of ductile iron based on an active learning strategy, characterized in that, It includes the following steps: S1. Collect the experimental data including the composition, process, microstructure characteristics and properties of ductile iron in the literature. Add the associated data with the composition, process, microstructure characteristics and properties of ductile iron in the experimental data to Dataset 1, and add the associated data with only the composition, process and properties of ductile iron in the experimental data to Dataset 2; S2. Visualize the distribution of the composition or process data of ductile iron in Dataset 1, and then conduct orthogonal experiments based on the blank areas of the visualized composition or process data of ductile iron, and add the obtained orthogonal experiment data to Dataset 1; S3. Preprocess the data in Dataset 1 and Dataset 2 to obtain the preprocessed Dataset 1 and the preprocessed Dataset 2; S4. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microstructure characteristics model. Set the composition and process of ductile iron as the input and the microstructure characteristics as the output. Use the preprocessed Dataset 1 to train the ductile iron composition-process-microstructure characteristics model, and use the trained ductile iron composition-process-microstructure characteristics model to predict the preprocessed Dataset 2 to obtain the microstructure characteristics corresponding to each group of data in the preprocessed Dataset 2, and supplement the obtained microstructure characteristic data to the preprocessed Dataset 2 after normalization; S5. Merge the preprocessed Dataset 1 and the preprocessed Dataset 2 obtained in step S4 into Dataset 3; S6. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microstructure characteristics-performance model. Set the composition, process and microstructure characteristics of ductile iron as the input and the performance as the output. Use Dataset 3 to train the ductile iron composition-process-microstructure characteristics-performance model to obtain the trained ductile iron composition-process-microstructure characteristics-performance model; S7. Based on the optimization algorithm, set different weights for different target performances. Each target performance is calculated using the trained ductile iron composition-process-microstructure characteristics-performance model in step S6. Adopt the weight-based objective function to continuously screen the composition closest to the target performance to obtain the optimized ductile iron composition and process parameters with different performance requirements; S8. For the optimized ductile iron composition and process parameters with different performance requirements obtained in step S7, through the active learning strategy, select the data with large uncertainty for experiments and supplement them to Dataset 1 and Dataset 2, and then return to step S4 for model update to optimize the model prediction accuracy, and finally design the ductile iron composition and process parameters that meet the expected performance requirements; The calculation method of uncertainty in step S8 is: Under other unchanged conditions, change the hyperparameters of the composition-process-microstructure characteristics-performance model, and calculate the performance prediction values of the composition and process optimized by the previous particle swarm optimization algorithm respectively. The uncertainty U is the variance of the prediction values of different models, and the expression is: Among them, is the average predicted value of all models, is the predicted value of the m-th model, and M is the total number of models.

2. The method for designing the composition and process of ductile iron based on an active learning strategy according to claim 1, wherein, The data of the ductile iron composition in step S1 include carbon, silicon, manganese, rare earth, sulfur, phosphorus, copper, nickel, magnesium, and chromium. The data of the process include cooling rate, inoculant type, inoculant content, spheroidizing agent type, spheroidizing agent content, inoculation temperature, and inoculation time. The data of the microstructural characteristics include graphite ball diameter, graphite ball number, graphite spheroidization rate, pearlite ratio in the matrix, ferrite ratio in the matrix, and spheroidization grade. The data of the properties include tensile strength, yield strength, ductility, and impact toughness.

3. The method for designing the composition and process of ductile iron based on an active learning strategy according to claim 2, wherein In step S2, select the top two with the most microstructural characteristic data in dataset 1 as the output features of the model, or customize the microstructural characteristics in dataset 1 as the output features of the model, and visualize the distribution of the ductile iron composition or process data associated with the selected or customized microstructural characteristic data, and extract the blank areas.

4. A method for designing the composition and process of ductile iron based on an active learning strategy according to claim 3, characterized in that, The preprocessing in step S3 includes data cleaning, normalization processing, and encoding processing. The formula for the normalization processing is: where X is the original data, X min and X max are the minimum and maximum values in the dataset respectively, and X′ is the normalized data; The encoding processing is to encode the inoculant type and the spheroidizing agent type.

5. A method for designing the composition and process of ductile iron based on an active learning strategy according to claim 4, characterized in that The specific implementation method of step S4 includes the following steps: S4.

1. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microstructural characteristic model, set the ductile iron composition and process as the input, and the microstructural characteristics as the output. The small-sample machine learning algorithm includes the extremely randomized tree algorithm, the CatBoost algorithm, the random forest algorithm, or the adaptive boosting regression algorithm; S4.

2. Set to use the k-fold cross-validation method to evaluate the performance of the model, where k = 5 - 10, and select the best-performing ductile iron composition-process-microstructural characteristic model. When predicting multiple microstructural characteristics, adopt the strategy of establishing models separately; S4.

3. Set the evaluation criteria including Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Coefficient of Determination (R 2 and Mean Average Percentage Error (MAPE), and the calculation formulas are as follows: Among them, represents the predicted value of the machine learning model, y i represents the true value; Then select the algorithm with the best evaluation criteria as the finally trained ductile iron composition-process-microstructural characteristic model.

6. The method for designing the composition and process of ductile iron based on an active learning strategy according to claim 5, wherein In the extremely randomized tree algorithm in step S4.1, the number of trees is 100, the maximum depth is 10, and the minimum number of samples required for each node split is 2; in the CatBoost algorithm, the number of iterations for training is 300, the learning rate is 0.02, and the L2 regularization coefficient is 3; in the random forest algorithm, the number of trees in the forest is 100, the minimum number of split samples is 2, and the minimum number of samples in the leaf nodes is 2; In the adaptive boosting regression algorithm, the number of base models is 80, the learning rate is 0.02, and the loss function type is linear.

7. A method for designing the composition and process of ductile iron based on an active learning strategy according to claim 6, characterized in that, The specific implementation method of step S6 includes the following steps: S6.

1. Based on the small-sample machine learning algorithm, construct a ductile iron composition-process-microstructural characteristic-property model, set the ductile iron composition, process, and microstructural characteristics as the input, and the property as the output. The small-sample machine learning algorithm includes the extremely randomized tree algorithm, the CatBoost algorithm, the random forest algorithm, or the adaptive boosting regression algorithm; S6.

2. Set the k-fold cross-validation method to evaluate the performance of the model, where k = 5-10, and screen out the best-performing nodular cast iron composition-process-microstructure feature-performance model. When predicting multiple performance features, adopt the strategy of establishing separate models; S6.

3. Set the evaluation criteria to include the mean absolute error (MAE), root mean square error (RMSE), coefficient of determination (R 2 ), and mean average percentage error (MAPE).

8. A method for designing the composition and process of ductile iron based on an active learning strategy according to claim 7, characterized in that, In step S7, intelligent optimization of composition and process is carried out through the particle swarm optimization algorithm. The objective function f of the particle swarm optimization algorithm is defined as: Among them, K i is the weight of the i-th target, T i is the target value, and T' i is the model prediction value.

Citation Information

Patent Citations

  • High-entropy alloy hardness prediction method and device based on machine learning and two-step method data expansion

    CN115394381A

  • Structural material high-throughput design method based on microstructure image visual analysis

    CN115472243A