Magnesium alloy mechanical property prediction method based on MOPSO data enhancement and hardness assistance
By using MOPSO data augmentation and hardness-assisted methods, an extended dataset was generated and combined with SHAP analysis. This solved the problems of universality and accuracy of the prediction model for the mechanical properties of magnesium alloys, and achieved efficient prediction and interpretable guidance without the need for complex microscopic parameters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing models for predicting the mechanical properties of magnesium alloys lack universality, rely on complex and costly microscopic parameters, and suffer from overfitting due to data scarcity, resulting in poor prediction performance.
Virtual samples were generated using the MOPSO data augmentation algorithm and combined with hardness-assisted input to construct an extended dataset for multi-objective optimization. SHAP analysis was used to achieve model interpretability and key threshold output, thus constructing a method for predicting the mechanical properties of magnesium alloys based on MOPSO data augmentation and hardness assistance.
It significantly improves the accuracy and stability of predicting the mechanical properties of magnesium alloys, simplifies input features, reduces dependence on complex parameters, provides key threshold guidance for process-composition-heat treatment, and solves the problems of data scarcity and insufficient model generalization ability.
Smart Images

Figure CN121885045A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of metallic materials and machine learning, and relates to a method for predicting the mechanical properties of magnesium alloys based on MOPSO data augmentation and hardness assistance. In particular, it relates to an interpretable machine learning method for predicting the mechanical properties of magnesium alloys based on multi-objective particle swarm optimization (MOPSO) data augmentation and hardness assistance, which is particularly suitable for performance prediction and process optimization of extruded magnesium alloy parts in aerospace, new energy vehicles and other fields. Background Technology
[0002] Magnesium alloys, with their excellent properties of low density, high specific strength, and high specific stiffness, are promising lightweight structural materials in the engineering field, showing significant application potential in lightweight applications such as aerospace and new energy vehicles. However, their relatively low plasticity and strength have become core bottlenecks restricting large-scale engineering applications. Although the academic community has explored ways to improve performance through alloy composition design and processing technology optimization, the mechanism by which the complex interactions of multiple elements within the magnesium alloy system affect strength and elongation remains unclear. Furthermore, different application scenarios require customized processing routes, which not only increases the design complexity of composition-process matching but also significantly prolongs the development cycle of new high-performance magnesium alloys, resulting in high R&D costs.
[0003] Machine learning, as a data-driven tool, can accurately construct nonlinear relationships between multiple variables and target properties, effectively handle large datasets in the field of materials, and significantly reduce the time and cost of traditional "trial and error" experiments. It has been widely applied to the prediction of electromagnetic properties, phase transformation behavior, corrosion properties, and mechanical properties of materials. However, in the field of predicting the mechanical properties of magnesium alloys, existing research has obvious limitations: most studies only target a single mechanical property or a specific series of magnesium alloys, lacking universality; and the model inputs often contain complex microstructure parameters or physicochemical parameters, which are costly to detect and require cumbersome preprocessing, severely limiting the practical engineering applicability of the models.
[0004] Meanwhile, magnesium alloy experimental data relies on costly smelting-extrusion-heat treatment and performance testing processes, resulting in generally small dataset sizes. This data scarcity not only affects the sufficiency of machine learning model training but also easily leads to overfitting, significantly weakening the model's generalization ability. To address this issue, data augmentation algorithms expand the dataset by applying "distribution-preserving" transformations to existing data. This avoids the high cost and long cycle of acquiring new data while ensuring that the augmented data conforms to the inherent characteristics of the original data, providing diverse training samples for the model. Among existing data augmentation methods, although there have been attempts to increase the data size through distribution-preserving transformations and single-objective particle swarm optimization (PSO), the application of MOPSO remains relatively scarce. As a multi-objective algorithm with strong global search capabilities, MOPSO can better balance "maintaining the consistency of the original data distribution" and "improving sample diversity" in data augmentation, better ensuring the physical rationality and feature coverage of the augmented data. Its application potential has not yet been fully explored, resulting in room for improvement in the quality of magnesium alloy data augmentation. Furthermore, incorporating prior knowledge directly related to mechanical properties into machine learning models can improve prediction accuracy and reduce parameter complexity. Hardness, as an easily measurable macroscopic mechanical parameter, is significantly correlated with the strength of magnesium alloys. It has been proven to be an effective auxiliary feature for strength prediction, which can greatly reduce the model's dependence on complex parameters and improve its practicality.
[0005] Therefore, there is an urgent need for a method to predict the mechanical properties of magnesium alloys that does not require complex microscopic parameters, can efficiently expand data, and has both model interpretability and key threshold output, in order to solve the pain points of traditional design and existing technology. Summary of the Invention
[0006] In view of this, in order to solve the problems of insufficient model universality, reliance on complex and costly microscopic parameters, and overfitting due to data scarcity in the prediction of mechanical properties of magnesium alloys, the present invention provides a method for predicting the mechanical properties of magnesium alloys based on MOPSO data augmentation and hardness assistance. The method achieves data augmentation without dependence on microscopic parameters through the MOPSO algorithm, improves model accuracy by combining hardness assistance input, and visualizes the model output through SHAP analysis to determine key thresholds for process, composition and heat treatment. Finally, it achieves the integration of "accurate prediction, interpretability and process guidance" for the mechanical properties of extruded magnesium alloys.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] A method for predicting the mechanical properties of magnesium alloys based on MOPSO data enhancement and hardness assistance includes the following steps:
[0009] S1. Collect data; construct a raw dataset containing magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and mechanical properties;
[0010] S2, Data Augmentation; Data augmentation is performed based on the MOPSO algorithm to construct an expanded dataset consisting of "original dataset + virtual samples"; Data preprocessing is performed on the expanded dataset, including missing value handling, outlier removal, standardization, and feature selection;
[0011] S3. Model Construction: Using the magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and hardness (HV) from the database in step S1 as input parameters, and yield strength (YS), ultimate tensile strength (UTS), and elongation (EL) as output performance parameters, construct multiple machine learning models and select the optimal model. Optimize the hyperparameters of the optimal model, and divide the extended dataset from step S2 into training and testing sets according to different proportions. Use 10-fold cross-validation to evaluate performance. Use R² and root mean square error (RMSE) as performance evaluation metrics to compare the model performance and generalization ability before and after data augmentation and hardness assistance. When selecting the machine learning model, R² is chosen as the optimal model. 2 The machine learning model with the largest value and the smallest RMSE value is used as the prediction model.
[0012] S4. Interpretability Analysis: Calculate the mean SHAP of the input features in step S3, determine the importance ranking of features affecting yield strength YS, ultimate tensile strength UTS, and elongation EL, draw a feature-SHAP value dependency graph to visualize the influence trend, and quantify the critical threshold of key features to understand the degree of influence of features on model prediction.
[0013] S5. Experimental verification: Verify the model's prediction accuracy using real experimental samples and samples not present in the original dataset.
[0014] Furthermore, in step S1, the magnesium alloy composition includes the contents of Al, Zn, Mn and other elements; the extrusion process parameters include extrusion temperature ET, extrusion ratio ER, and extrusion speed ES; the heat treatment parameters include solution temperature ST, solution time St, aging temperature AT, and aging time At; and the mechanical properties include hardness HV, yield strength YS, ultimate tensile strength UTS, and elongation EL.
[0015] Furthermore, in step S2, the MOPSO algorithm simulates the collective intelligent behavior of flocks of birds or schools of fish to search for the optimal solution in the solution space to generate virtual samples. Its core functionality is achieved through iterative updates of particle velocity and position: the velocity update formula is... Inertial weight To maintain the current motion trend of the particles and ensure search stability, the learning factor... , Each particle is driven toward its own historical best solution. (Individual optimal solution), Group global optimal solution near, , Use random numbers in the range [0,1] to enhance the breadth of solution space exploration; the position update formula is: The particles move to new positions in the solution space based on the updated velocity. These new positions correspond to a set of parameter combinations (including alloy composition, process parameters, and performance indicators), which are the candidate virtual samples. After multiple iterations, the effective virtual samples are finally obtained by combining the physical property boundaries of the magnesium alloy (such as the range of hardness and strength values) and the distribution consistency with the original data.
[0016] Furthermore, step S2 data augmentation focuses on optimizing KL divergence <0.05 and sample coverage ≥95% as multiple objectives, expanding the original dataset by 2 to 4 times to improve dataset distribution consistency and sample diversity.
[0017] Furthermore, in step S2, the boundary values of the physical properties of the virtual samples are set as follows: hardness HV 30-140, yield strength YS 50-500MPa, ultimate tensile strength UTS 150-600MPa, and elongation EL 1-40%, to avoid generating invalid samples that exceed the actual properties.
[0018] Furthermore, the data standardization preprocessing formula for step S2 is as follows: Where X is the original feature value, μ is the feature mean, and σ is the feature standard deviation, ensuring that the input feature dimensions are consistent; key features are selected based on the Pearson correlation coefficient.
[0019] Furthermore, in step S3, multiple machine learning models are included, such as AdaBoost, Gradient Boosting GBDT, Extreme Gradient Boosting XGBoost, Gaussian Regression Process (GPR), Random Forest (RF), Support Vector Machine (SVM), and Multilayer Perceptron (MLP).
[0020] Furthermore, the hyperparameter optimization method in step S3 is one of random search, grid search, or Bayesian optimization.
[0021] Furthermore, in step S3, the test set accounts for 10% to 30% of the expanded dataset.
[0022] Furthermore, in step S3, the various machine learning models are extreme gradient boosting XGBoost, the hyperparameter optimization method is grid search, and the extended dataset is divided into training and test sets in an 8:2 ratio.
[0023] Furthermore, the performance evaluation metrics for the 10-fold cross-validation in step S3 include the coefficient of determination R² (the closer to 1, the better) and the root mean square error RMSE (the closer to 0, the better), with the following formulas: , ,in For predicted values, For the true value, is the mean of the true values, and n is the number of samples.
[0024] The beneficial effects of this invention are as follows:
[0025] 1. The method for predicting the mechanical properties of magnesium alloys based on MOPSO data augmentation and hardness assistance disclosed in this invention achieves data augmentation without dependence on complex microscopic parameters through the MOPSO algorithm: it does not require input of difficult-to-measure parameters such as grain size and second phase distribution, but only based on the original process and composition data, it iteratively generates 800 qualified virtual samples with "KL divergence <0.05 (distribution consistency)" and "sample coverage ≥95% (diversity)" as multi-objective optimization directions, and merges them with 271 original data to form 1071 extended datasets, which increases the size of the dataset by 3.9 times, effectively solving the problem of insufficient generalization ability of ML models caused by the scarcity of experimental data of extruded magnesium alloys.
[0026] 2. The magnesium alloy mechanical property prediction method based on MOPSO data enhancement and hardness assistance disclosed in this invention uses easily measurable hardness parameters as model inputs. Pearson correlation verification (hardness correlation coefficients with YS PCC=0.88, with UTS PCC=0.92, and with EL PCC=-0.79) clarifies its predictive value. After being incorporated into the model, the XGBoost model's YS prediction R² increased from 0.24 to 0.98, and RMSE decreased from 88.44 MPa to 17.65 MPa; UTS prediction R² increased from 0.18 to 0.97, and RMSE decreased from 76.21 MPa to 15.19 MPa; EL prediction R² increased from 0.29 to 0.91, and RMSE decreased from 5.81 MPa to 1.83 MPa. This successfully replaces costly micro-parameter measurements, simplifies input features, and significantly improves prediction accuracy.
[0027] 3. The method for predicting the mechanical properties of magnesium alloys based on MOPSO data enhancement and hardness assistance disclosed in this invention determines the key thresholds of process-composition-heat treatment based on SHAP analysis: In response to the "black box" problem of XGBoost, the SHAP mean of each input feature is calculated, the feature-performance dependency graph is drawn to visualize the correlation trend, and the contribution of features to mechanical properties is quantified. For the first time, it provides a theoretical basis for determining the key thresholds of the correlation between "composition-process-performance" of magnesium alloys, and overcomes the limitation of traditional ML models that cannot explain the prediction logic.
[0028] 4. The method for predicting the mechanical properties of magnesium alloys based on MOPSO data enhancement and hardness assistance disclosed in this invention can ensure prediction accuracy and stability: the XGBoost model has a prediction R² of ≥0.91 for yield strength (YS), tensile strength (UTS) and elongation (EL) on the extended dataset; in the verification of 2 sets of real experimental samples and 4 sets of samples that did not appear in the original dataset, the prediction error is <10%.
[0029] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0030] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0031] Figure 1 This is an overall flowchart of the method for predicting the mechanical properties of magnesium alloys based on MOPSO data enhancement and hardness assistance according to the present invention.
[0032] Figure 2 This is a comparison chart showing the distribution of the original data and the enhanced virtual sample data of this invention across four different material performance indicators; wherein... Figure 2 (a) is a comparison chart of the fitting of the original data and the extended data under the yield strength (YS) index. Figure 2 (b) is a comparison chart of the fitting of the original data and the extended data under the tensile strength (UTS) index. Figure 2 (c) is a comparison chart of the fitting of the original data and the extended data under the elongation (EL) index. Figure 2 (d) is a comparison chart of the fitting of the original data and the extended data under the Vickers hardness (HV) index;
[0033] Figure 3 This is a comparison chart of the predictive performance of the XGBoost model for yield strength (YS), tensile strength (UTS), and elongation (EL) under different data scenarios according to embodiments of the present invention; wherein Figure 3 (a) is the coefficient of determination R for the raw and reinforced data containing hardness. 2 Histogram comparison of prediction performance of root mean square error (RMSE). Figure 3 (b) To enhance the coefficient of determination R in scenarios including and excluding Vickers hardness (HV). 2 Histogram comparing the prediction performance of Root Mean Square Error (RMSE).
[0034] Figure 4The image shows a scatter plot of the XGBoost model's predictions of mechanical properties on a test set (214 items) according to an embodiment of the present invention. From left to right, the images show the prediction results for three scenarios, including yield strength (YS), tensile strength (UTS), and elongation (EL).
[0035] Figure 5 This invention presents an input feature importance ranking graph based on SHAP and a dependency graph between key features and SHAP values, as described in this embodiment. Figure 5 (a) is a summary chart of SHAP values. Figure 5 (b) is a scatter plot of the Zn content SHAP values. Figure 5 (c) is a scatter plot of the SHAP values of Mn content. Figure 5 (d) is a scatter plot of the Zr content SHAP values. Detailed Implementation
[0036] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0037] like Figure 1 As shown, the method for predicting the mechanical properties of magnesium alloys based on MOPSO data enhancement and hardness assistance includes the following steps:
[0038] S1. Collect data; construct a raw dataset containing magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and mechanical properties; where magnesium alloy composition includes the content of Al, Zn, Mn, and other elements; extrusion process parameters include extrusion temperature ET, extrusion ratio ER, and extrusion speed ES; heat treatment parameters include solution temperature ST, solution time St, aging temperature AT, and aging time At; mechanical properties include hardness HV, yield strength YS, ultimate tensile strength UTS, and elongation EL.
[0039] S2, Data Augmentation; Data augmentation is performed based on the MOPSO algorithm to construct an expanded dataset consisting of "original dataset + virtual samples"; Data preprocessing is performed on the expanded dataset, including missing value handling, outlier removal, standardization, and feature selection;
[0040] Since magnesium alloy experimental data rely on a high-cost smelting-extrusion-heat treatment and performance testing process, the dataset size is generally small. To solve the problem of data scarcity, data augmentation can be used to apply a "distribution-preserving" transformation to the existing data to expand the dataset, thereby achieving consistency in the distribution of the expanded dataset and sample diversity.
[0041] The particle position update formula for the MOPSO algorithm is as follows: The speed update formula is: ,in This is the optimal solution for the individual. This is the globally optimal solution. , It is a random number in the range [0,1].
[0042] With "KL divergence < 0.05 (distribution consistency)" and "sample coverage ≥ 95% (diversity)" as the multi-objective optimization directions, 800 qualified virtual samples were generated iteratively and merged with 271 original data to form 1071 extended datasets, increasing the dataset size by 3.9 times.
[0043] The data standardization preprocessing formula is as follows: , where X is the original feature value, μ is the feature mean, and σ is the feature standard deviation, ensuring that the input feature dimensions are consistent.
[0044] S3. Model Construction: Using the magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and hardness (HV) from the database in step S1 as input parameters, and yield strength (YS), ultimate tensile strength (UTS), and elongation (EL) as output performance parameters, construct various machine learning models and select the best model. Optimize the hyperparameters of the best model, divide the extended dataset from step S2 into training and testing sets in an 8:2 ratio, and use 10-fold cross-validation to evaluate performance. Use R² and root mean square error (RMSE) as performance evaluation metrics.
[0045] Multiple machine learning models are available, including one of the following: Adaptive Boosting (AdaBoost), Gradient Boosting (GBDT), Extreme Gradient Boosting (XGBoost), Gaussian Regression Process (GPR), Random Forest (RF), Support Vector Machine (SVM), and Multilayer Perceptron (MLP).
[0046] Hyperparameter optimization methods include random search, grid search, and Bayesian optimization.
[0047] The performance evaluation metrics for 10-fold cross-validation include the coefficient of determination R² (the closer to 1, the better) and the root mean square error RMSE (the closer to 0, the better), with the following formulas: , ,in For predicted values, For the true value, is the mean of the true values, and n is the number of samples.
[0048] S4. Interpretability Analysis: Calculate the mean SHAP of the input features in step S3, determine the importance ranking of features affecting yield strength YS, ultimate tensile strength UTS, and elongation EL, draw a feature-SHAP value dependency graph to visualize the influence trend, and quantify the critical threshold of key features to understand the degree of influence of features on model prediction.
[0049] S5. Experimental verification: Verify the model's prediction accuracy using real experimental samples and samples not present in the original dataset.
[0050] Example
[0051] This example focuses on four core performance parameters: hardness (HV), yield strength (YS), ultimate tensile strength (UTS), and elongation (EL). Through quantitative indicators and visualization analysis, it verifies the consistency and diversity of the performance dimensions between the virtual samples generated by MOPSO and the original data, ensuring the physical rationality of the augmented data. Furthermore, using the XGBoost model as a benchmark, it sets up two scenarios: "no hardness input" and "with hardness input," evaluating the performance differences in the original and augmented datasets respectively, verifying the predictive value of hardness. The steps include:
[0052] S1. Collect data by category; Combination 1 (no hardness input scenario): extrusion process parameters (ET, ER, ES) + heat treatment parameters (ST, St, AT, At) + alloy composition + load, a total of 32 features; Combination 2 (with hardness input scenario): add hardness HV to Combination 1, a total of 33 features;
[0053] S2. Dataset expansion specifically includes the following steps:
[0054] S21. Using four mechanical performance parameters as the target dimensions, the MOPSO particle swarm size is set to 120, the maximum number of iterations is 50, the inertia weight decreases linearly from 0.9 to 0.5, and the learning factor is c1=c2=2.0; the multi-objective optimization objectives are set as "KL divergence < 0.05 (distribution consistency)" and "sample coverage ≥ 95% (diversity)".
[0055] S22. Based on the physical property boundaries of magnesium alloys, the range of virtual sample values is limited (HV 30-140, YS 50-500MPa, UTS 150-600MPa, EL 1-40%) to avoid generating invalid samples that exceed the actual performance.
[0056] S23. Iteratively generate 830 virtual samples, and remove 30 outliers through a second screening using the 3σ criterion. Finally, retain 800 qualified virtual samples, which are then merged with the original 271 sets of data to form an extended dataset of 1071 sets.
[0057] S24. Using KL divergence (to measure distribution similarity) and KS test (to verify distribution consistency), the performance parameter distributions of the original data and the virtual sample were compared. The results are shown in the table below:
[0058]
[0059] S25. Draw histograms of the four main performance parameters (see attached diagram). Figure 2 ),in Figure 2 (a) is a comparison chart of the fitting of the original data and the extended data under the yield strength (YS) index. Figure 2 (b) is a comparison chart of the fitting of the original data and the extended data under the tensile strength (UTS) index. Figure 2 (c) is a comparison chart of the fitting of the original data and the extended data under the elongation (EL) index. Figure 2 (d) is a comparison chart of the fitting of the original data and the extended data under the Vickers hardness (HV) index;
[0060] In the yield strength (YS) distribution comparison chart, the mean (μ) of the original data is 226.55 MPa, and the standard deviation (σ) is 90.61 MPa. The mean (μ) of the reinforced data is 253.07 MPa, and the standard deviation (σ) is 97.26 MPa. The reinforced data distribution is more biased towards higher yield strength. In the tensile strength (UTS) distribution comparison chart, the mean (μ) of the original data is 304.79 MPa, and the standard deviation (σ) is 77.39 MPa. The mean (μ) of the reinforced data is 322.55 MPa, and the standard deviation (σ) is 83.79 MPa. The reinforced data distribution is also more biased towards higher tensile strength. In the elongation (EL) distribution comparison chart, the mean (μ) of the original data is 15.70%, and the standard deviation (σ) is 7.82%. The mean (μ) of the reinforced data is 13.85%, and the standard deviation (σ) is 6.50%. The augmented data exhibits a more concentrated distribution of elongation, with a slightly lower mean. In the Vickers hardness (HV) distribution comparison chart, the mean (μ) of the original data is 78.55, and the standard deviation (σ) is 22.54. The mean (μ) of the augmented data is 78.55, and the standard deviation (σ) is 29.90. The augmented data is more dispersed, but the mean remains the same as the original data. The histogram curves of the original data and the virtual sample overlap by ≥90%, and the mean and standard deviation deviations of the box plots are ≤5%, visually demonstrating the consistency in distribution between the original and augmented data. The augmented data achieves a balance between diversity and consistency.
[0061] S3. Model Construction: The magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and hardness (HV) in the mechanical properties from the database in step S1 are used as input parameters. The yield strength (YS), ultimate tensile strength (UTS), and elongation (EL) are used as output performance parameters. The XGBoost model is selected as the machine learning model and hyperparameter optimization is performed. Mesh search is selected for hyperparameter optimization.
[0062] S31. Divide the original data (271 records) and augmented data (271 original records + 800 virtual records = 1071 records) into training and test sets in an 8:2 ratio, and use 10-fold cross-validation to eliminate the splitting bias.
[0063] S32. Using the predicted indices (R², RMSE) of core mechanical properties (YS, UTS, EL) as performance evaluation indicators, the results are shown in the table below:
[0064]
[0065] The feature combinations of the original data in the table above are assumed to include hardness, which is equivalent to the attached data. Figure 3 (a) The original data.
[0066] S4. Plot the original data and the enhanced data, and the coefficient of determination R between the enhanced data without hardness and the enhanced data with hardness. 2 Root Mean Square Error (RMSE) Histogram (with attached) Figure 3 (and a comparison chart of generalization capabilities.)
[0067] in Figure 3 (a) is the coefficient of determination R for the raw and reinforced data containing hardness. 2 The root mean square error (RMSE) histogram compares the predictive performance of the original and augmented data models: YS model: Rm for the original data. 2 The R value for augmented data is 0.82. 2 The RMSE is 0.98; the RMSE for the original data is 37.05, and the RMSE for the augmented data is 17.65. UTS model: R0 for the original data... 2 The R value for augmented data is 0.82. 2 The RMSE is 0.97; the RMSE for the original data is 29.99, and the RMSE for the augmented data is 15.19. EL model: R0 for the original data 2 The R value for augmented data is 0.73. 2 The RMSE for the original data is 4.09, and the RMSE for the augmented data is 1.83. It can be seen that the model using the augmented data has a higher R-value than the model using the original data across all three performance metrics. 2 The higher values and lower RMSE values indicate that augmented data can significantly improve the model's predictive performance.
[0068] Figure 3 (b) The determination coefficient R is given for scenarios including and excluding Vickers hardness (HV). 2 The root mean square error (RMSE) prediction performance histogram compares the performance of models that include and do not include Vickers hardness (HV) as a feature: YS model: RSE without HV. 2 The value is 0.24, including HV's R. 2 The RMSE is 0.98; the RMSE without HV is 88.44, and the RMSE with HV is 17.65. UTS model: R without HV 2 The value is 0.18, including the R value of HV. 2 The RMSE is 0.97; the RMSE without HV is 76.21, and the RMSE with HV is 15.19. EL model: R without HV 2 The value is 0.29, including HV's R. 2 The RMSE is 0.91; the RMSE without HV is 5.81, and the RMSE with HV is 1.83. It can be seen that for both the YS and UTS models, including HV as a feature significantly improves the model's R-value. 2 The increased value and reduced RMSE value indicate that HV is an important feature. However, for the EL model, including HV did not significantly change the model's performance, suggesting that HV has a relatively small impact on the EL model. In summary, these two figures demonstrate that data augmentation and the selection of appropriate features (such as HV) can significantly improve the model's predictive performance. The combination of virtual samples and hardness further optimizes the feature space and enhances the model's generalization ability.
[0069] Figure 4 This is a scatter plot showing the mechanical property predictions of the XGBoost model for the test set (214 data points) in this embodiment of the invention. From left to right, it compares the predicted results with the actual values in three scenarios: yield strength (YS), tensile strength (UTS), and elongation (EL). Each subplot shows scatter plots of the training and test sets. In each subplot, black dots represent training set data, and colored dots represent test set data (blue in the left and middle plots, and purple in the right plot). The red dashed line represents the ideal case where the predicted value is completely consistent with the actual value (i.e., the y=x line). The figure also provides the coefficient of determination (R²) for each model on the training and test sets. 2 ) and root mean square error (RMSE): Left figure (YS scenario): Training set: R 2 =0.98, RMSE =17.01 MPa. Test set: R 2 =0.99, RMSE = 14.95 MPa. (Chinese image, UTS scenario): Training set: R 2=0.98, RMSE= 13.15 MPa. Test set: R 2 =0.99, RMSE = 11.19 MPa. Right figure (EL scenario): Training set: R 2 =0.91, RMSE= 1.94%. Test set: R 2 =0.98, RMSE = 0.87%. These figures show that the yield strength (YS), tensile strength (UTS), and elongation (EL) all exhibit high R-values on both the training and test sets. 2 The high RMSE value indicates a high degree of model fit to the data. A relatively low RMSE value indicates a small error between the predicted and actual values. The fact that most points in the scatter plot are concentrated near the red dashed line further demonstrates the accuracy of the model's predictions. Overall, these figures show that the developed model has good performance in predicting yield strength, tensile strength, and elongation.
[0070] Figure 5 (a) is a summary plot of SHAP values, showing the impact of each feature on the model output. The color represents the magnitude of the feature value (red represents a high value, and blue represents a low value), and the size of the dot represents the absolute value of the feature value. As can be seen from the plot, hardness (HV) has the greatest impact on the model output, followed by the weight percentage of elements such as magnesium (Mg) and zinc (Zn). Figure 5 (b) is a SHAP dependency plot of Zn content, showing the impact of Zn mass percentage on the model output. Again, the colors represent Zn values (red for high, blue for low). The plot shows that increasing Zn content has a positive impact on the model output. Figure 5 (c) is a SHAP dependency plot of Mn content, showing the impact of the percentage of Mn mass on the model output. Each point in the plot represents a sample, and the color indicates the value of Mn (red for high, blue for low). It can be seen that as the Mn content increases, the positive impact on the model output also increases. Figure 5 (d) is a scatter plot of Zr content SHAP values, showing the impact of Zr mass percentage on the model output. Colors represent Zr values (red for high, blue for low). The plot shows that increasing Zr content primarily has a negative impact on the model output. These plots demonstrate that SHAP values provide a visual understanding of how the model utilizes input features to make predictions. In particular, they identify which features have the greatest impact on the model's predictions and how changes in these feature values affect the model's output. This is crucial for model interpretability and transparency, especially in materials science where model predictions require validation and interpretation.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for predicting the mechanical properties of magnesium alloys based on MOPSO data enhancement and hardness assistance, characterized in that, Includes the following steps: S1. Collect data; construct a raw dataset containing magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and mechanical properties; S2. Data Augmentation: Data augmentation is performed based on the MOPSO algorithm to construct an expanded dataset consisting of "original dataset + virtual samples". The expanded dataset is then preprocessed, including missing value handling, outlier removal, standardization, and feature selection. S3. Model Construction: Using the magnesium alloy composition, extrusion process parameters, heat treatment parameters, load, and hardness (HV) from the database in step S1 as input parameters, and yield strength (YS), ultimate tensile strength (UTS), and elongation (EL) as output performance parameters, construct multiple machine learning models and select the optimal model. Optimize the hyperparameters of the optimal model, and divide the extended dataset from step S2 into training and testing sets according to different proportions. Use 10-fold cross-validation to evaluate performance. Use R² and root mean square error (RMSE) as performance evaluation metrics to compare the model performance and generalization ability before and after data augmentation and hardness assistance. When selecting the machine learning model, R² is chosen as the optimal model. 2 The machine learning model with the largest value and the smallest RMSE value is used as the prediction model. S4. Interpretability Analysis: Calculate the mean SHAP of the input features in step S3, determine the importance ranking of features affecting yield strength YS, ultimate tensile strength UTS, and elongation EL, draw a feature-SHAP value dependency graph to visualize the influence trend, and quantify the critical threshold of key features to understand the degree of influence of features on model prediction. S5. Experimental verification: Verify the model's prediction accuracy using real experimental samples and samples not present in the original dataset.
2. The method for predicting the mechanical properties of magnesium alloys as described in claim 1, characterized in that, In step S1, the magnesium alloy composition includes the contents of Al, Zn, Mn and other elements; the extrusion process parameters include extrusion temperature ET, extrusion ratio ER, and extrusion speed ES; the heat treatment parameters include solution temperature ST, solution time St, aging temperature AT, and aging time At; and the mechanical properties include hardness HV, yield strength YS, ultimate tensile strength UTS, and elongation EL.
3. The method for predicting the mechanical properties of magnesium alloys as described in claim 2, characterized in that, In step S2, the MOPSO algorithm simulates the collective intelligent behavior of flocks of birds or schools of fish to search for the optimal solution in the solution space to generate virtual samples; its core is achieved through iterative updates of particle velocity and position: the velocity update formula is... Inertial weight To maintain the current motion trend of the particles and ensure search stability, the learning factor... , Each particle is driven toward its own historical best solution. swarm global optimal solution near, , Use random numbers in the range [0,1] to enhance the breadth of the solution space exploration; The position update formula is The particles move to a new position in the solution space based on the updated velocity. The combination of alloy composition, process parameters, and mechanical property indicators corresponding to the new position is the candidate virtual sample. After multiple iterations, the effective virtual samples are finally obtained by combining the physical property boundaries of the hardness and strength range of magnesium alloy and the distribution consistency with the original data.
4. The method for predicting the mechanical properties of magnesium alloys as described in claim 3, characterized in that, Step S2, data augmentation, focuses on optimizing KL divergence <0.05 and sample coverage ≥95% as multiple objectives, expanding the original dataset by 2 to 4 times to improve dataset distribution consistency and sample diversity.
5. The method for predicting the mechanical properties of magnesium alloys as described in claim 4, characterized in that, In step S2, the boundary values of the physical properties of the virtual samples are set as follows: hardness HV 30-140, yield strength YS 50-500MPa, ultimate tensile strength UTS 150-600MPa, and elongation EL 1-40%, in order to avoid generating invalid samples that exceed the actual properties.
6. The method for predicting the mechanical properties of magnesium alloys as described in claim 1, characterized in that, The formula for data standardization preprocessing in step S2 is as follows: Where X is the original feature value, μ is the feature mean, and σ is the feature standard deviation, ensuring that the input feature dimensions are consistent; key features are selected based on the Pearson correlation coefficient.
7. The method for predicting the mechanical properties of magnesium alloys as described in claim 1, characterized in that, Step S3 involves multiple machine learning models, including one of the following: Adaptive Boosting (AdaBoost), Gradient Boosting (GBDT), Extreme Gradient Boosting (XGBoost), Gaussian Regression (GPR), Random Forest (RF), Support Vector Machine (SVM), and Multilayer Perceptron (MLP); and the hyperparameter optimization method is one of the following: random search, grid search, and Bayesian optimization.
8. The method for predicting the mechanical properties of magnesium alloys as described in claim 7, characterized in that, In step S3, the test set accounts for 10% to 30% of the expanded dataset.
9. The method for predicting the mechanical properties of magnesium alloys as described in claim 8, characterized in that, In step S3, multiple machine learning models are extreme gradient boosting XGBoost, and the hyperparameter optimization method is grid search. The extended dataset is divided into training and test sets in an 8:2 ratio.
10. The method for predicting the mechanical properties of magnesium alloys as described in claim 1, characterized in that, The performance evaluation metrics for 10-fold cross-validation in step S3 include the coefficient of determination R² and the root mean square error RMSE, respectively, with the following formulas: , ,in For predicted values, For the true value, is the mean of the true values, and n is the number of samples.