Wheel material reverse design method based on machine learning
By constructing a positive prediction model and Bayesian optimization through machine learning, combined with SHAP analysis, reverse design of wheel materials is achieved, solving the problems of long design cycle and high cost in traditional design, improving design efficiency and service performance, and is applicable to the field of intelligent manufacturing of rail transit equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional wheel material design has a long design cycle and high cost, and fails to effectively combine service performance, resulting in a disconnect between design and actual needs. This is especially true in the field of intelligent manufacturing of rail transit equipment, where there is a lack of fast and efficient material design methods.
A forward prediction model is constructed using machine learning methods. Combined with Bayesian optimization and SHAP analysis, a reverse design method for chemical composition and process parameters is constructed based on the service performance optimization objective, forming a closed-loop optimization design process. The model is continuously updated by combining experimental verification.
It significantly shortens the design cycle, reduces R&D costs, extends wheel service life, improves design accuracy and adaptability, and meets the design needs of different operating scenarios.
Smart Images

Figure CN121789850A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent manufacturing of computers and rail transit equipment, specifically to a reverse design method for wheel materials based on machine learning. Background Technology
[0002] High-speed train wheels are core components ensuring safe train operation. With the continuous development of railway technology, the requirements for wheel materials in terms of performance and cost are increasing, with performance demands shifting towards a synergistic optimization of high strength, low cost, and long service life. Therefore, developing new high-performance materials is a crucial approach to meeting these demands. However, the influence of elemental composition, heat treatment, and deformation processes on the mechanical properties and microstructure of wheel materials is complex, and the mechanisms underlying these influences remain unclear. Furthermore, as the composition of wheel materials increases and processing performance is continuously optimized, the combination of various material parameters becomes increasingly complex, and there is a lack of models that quantitatively describe the relationship between alloy composition, structure, processing, and performance. Quickly clarifying the complex relationship between the mechanical properties of wheel materials and their composition and processing, and further improving the mechanical properties of wheels, is one of the keys to developing new types of wheels.
[0003] Traditional wheel material design employs a "forward trial-and-error" approach, which involves adjusting composition and process parameters to create prototypes and test their performance. This method is characterized by long lead times (typically 6-10 months), high costs, and heavy reliance on expert experience. Due to the vast potential parameter space, finding the optimal combination of parameters in wheel design and manufacturing through traditional trial-and-error methods often consumes significant time and resources, potentially hindering the development of high-performance wheels.
[0004] Furthermore, traditional wheel designs often focus on mechanical properties (such as strength and hardness) but neglect key performance characteristics such as wear and contact fatigue in actual service, leading to a disconnect between design and application scenarios. For example, current methods do not consider wear rate as a core indicator, often resulting in insufficient wheel service life.
[0005] To address these issues, recent research has attempted to combine traditional materials design with machine learning. Machine learning (ML) methods are computationally inexpensive and have short development cycles. They do not require manual solution of complex equations; instead, they learn the relationships between data and identify implicit relationships between features and labels. Therefore, they are one of the most effective methods to replace repetitive laboratory experiments. However, research on the application of machine learning methods in the field of intelligent manufacturing for rail transit equipment is limited. Similarly, their application in alloy material design is more common. Invention patent CN116844673B discloses a reverse design method for high-performance magnesium alloys based on machine learning. By collecting data on magnesium alloy composition, processing technology, and mechanical properties, machine learning algorithms analyze the implicit structure-property relationships between magnesium alloy composition, processing technology, and mechanical properties, thereby enabling the efficient design of new high-performance magnesium alloys based on mechanical performance requirements. This invention combines optimization algorithms with forward models to optimize parameters such as alloy composition, effectively improving prediction accuracy (calculation precision up to 99%) and reducing bias (error only 0.5%). It can achieve good fitting results even with small alloy data volumes and complex process combinations. The calculation method is simple and easy to implement. However, this method has not yet been applied to the field of intelligent manufacturing of rail transit equipment, especially in the design of wheel materials. This is because wheel material design involves complex thermal process parameters, deformation processes, different service conditions, and various service performance indicators. Integrating service performance with wheel design is one of the keys to the rapid and efficient development of new types of wheels. Therefore, a new method that can solve the above problems is urgently needed for wheel material design. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a machine learning-based reverse design method for wheel materials. For the first time, it incorporates service performance into the reverse design of wheels, using wear rate and contact stress as core optimization targets. This breaks through the limitations of traditional design that "emphasizes mechanical performance and neglects service performance." Through experimental verification data, the training set is continuously updated to achieve model self-optimization. This solves the problems of existing machine learning methods, such as computational complexity, low efficiency, low accuracy, weak generalization ability, and poor model fitting effect for small amounts of data.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a machine learning-based reverse design method for wheel materials, comprising the following steps: S1. Data Acquisition and Preprocessing: Acquire the chemical composition, process parameters, and corresponding mechanical and service performance data of the wheels. After cleaning and normalization, obtain the modeling dataset. S2. Forward Prediction Model Training: Train a LightGBM regression prediction model to construct a forward prediction model with "chemical composition + process parameters" as input and "mechanical properties" and "service performance" as output; optimize the model hyperparameters through Bayesian optimization algorithm, introduce the SHAP method to calculate the feature contribution, and then calculate the MASH value of the feature to quantify the global importance of the feature, and screen out the features that have a decisive impact on the design goal; S3. Reverse design model construction and solution: With preset service performance targets (such as wear rate) as constraints, and with the equipment safety range and production cost of process parameters as boundary conditions, the chemical composition and process parameter combination that meets the performance requirements is solved by Bayesian optimization in reverse. S4. Parameter Optimization and Verification: Taking into account performance satisfaction, process feasibility and economy, select the best 1 to 3 candidate schemes for experimental verification; add the new data obtained from the experimental verification to the dataset in step S1, and retrain the positive prediction model in step S2 to achieve continuous iteration and performance improvement of the model, forming a self-evolving material closed-loop design method.
[0008] Preferably, step S1 specifically includes: S11. Data Acquisition: Collect historical data of the target wheel system through multiple channels such as literature review, laboratory experiments, industrial production records and line service monitoring, including the chemical composition, process parameters and corresponding mechanical and service performance data of the wheel (such as wear rate and contact stress). S12. Missing value handling: Missing data are filled using random forest imputation (implemented using the Python sklearn library, with 100 trees), with an imputation error ≤5%; S13. Outlier handling: Outliers are identified and removed based on the 3σ principle (i.e., data exceeding the range of "mean ± 3 × standard deviation"). The proportion of outliers is controlled within 8% of the total data volume. S14. Feature Normalization: The input features are standardized using the Min-Max normalization method. The normalization formula is as follows:
[0009] in, x These are the original eigenvalues. , These are the minimum and maximum values of the feature in the dataset, respectively. These are the normalized eigenvalues (mapped to the [0,1] interval).
[0010] Preferably, in step S11, the wheel data specifically includes: Industrial production data: Wheel material range includes ER7, ER8, ER9, and CL65 steel; Laboratory-based experimental data: The experimental equipment consisted of a double-disc rolling tester and a Vickers hardness tester; the experimental environment was at room temperature of 25±2℃. Literature data: Publications were published within the last 10 years; data validity verification standards were followed. Preferably, the chemical composition data includes carbon, silicon, manganese, chromium, molybdenum, phosphorus, sulfur, chromium, nickel, vanadium, copper, aluminum, titanium, niobium and the corresponding mass percentages of these elements; the process parameters include quenching temperature (800℃-1000℃) and tempering time (1-4 hours); the mechanical properties include tensile strength (≥800MPa), yield strength (≥500MPa), Vickers hardness (HV230-406), and impact toughness (≥20J); and the service performance data includes wear rate (≤0.1mm / 10,000km) and contact stress (≤400MPa).
[0011] Preferably, the positive prediction model is LightGBM, and its parameter selection range is: n_estimators: 100-300; learning_rate: 0.05-0.2; max_depth: 3-8; min_child_samples:10-30.
[0012] Preferably, in step S2, the tuning of model hyperparameters using the Bayesian optimization algorithm specifically involves: searching for the optimal combination of hyperparameters using the Bayesian optimization algorithm, with the objective of minimizing the mean squared error (MSE); the hyperparameters include the learning rate, tree depth, number of leaf nodes, and number of iterations. Performance evaluation: The model must satisfy the coefficient of determination R. 2 Mean absolute error (MAE) and mean square error (MSE) are used to ensure that the prediction accuracy meets the actual situation and engineering design requirements.
[0013] Preferably, in step S2, the SHAP (SHapley Additive exPlanations) method is introduced to analyze the contribution of model features. Features are ranked by importance and redundant features are removed based on the Shapley value. The top 80% of key features by contribution are retained. The formula for calculating the Shapley value is defined as follows:
[0014] in, F is the contribution of a feature to the model output, i.e., the Shapley value of the feature; F is the set of all features; S is the subset of input that does not contain features. The total number of characteristics; The performance predictions of the forward model when only a subset S is input; The marginal contribution of adding a feature to the input subset S to the performance (a positive contribution indicates that the feature improves performance, and a negative contribution indicates the opposite).
[0015] The MASH value is the "mean absolute contribution" of a feature to all samples, quantifying the global importance of the feature. The formula for calculating the MASH value (Mean Absolute SHAP) of a feature is as follows:
[0016] in The mean absolute Shapley value of the i-th feature; N is the total number of samples in the modeling dataset; Let be the Shapley value of the i-th feature in the j-th sample.
[0017] Feature filtering rules: Retain MASH i ≥1.2×MASH avg And MASH i valid >MASH i total The features of MASH are used to eliminate redundant features; avg MASH is the average of all feature MASH values. i valid MASH is the MASH value of the i-th feature in the service performance compliance sample. i total is the MASH value of the i-th feature in the global sample.
[0018] Preferably, in step S3, the objective function of the Bayesian optimization is defined as:
[0019] Where X is a vector of chemical composition and process parameters; n is the number of service performance indicators; For performance weighting coefficients; Let i be the model prediction value of the i-th service performance item. Let i be the target value for the service performance of the i-th item; This represents the standard deviation of process parameters (reflecting process stability; for example, the larger the fluctuation range of quenching temperature, the higher the value). This represents the process stability coefficient.
[0020] The process stability coefficient λ is determined based on the equipment accuracy, and its calculation formula is as follows: λ
[0021] in, The standard deviation of the process parameters; The maximum allowable standard deviation for the equipment (e.g., maximum temperature fluctuation of the quenching furnace ±8℃). =8; λ∈[0.1,0.5] is required to ensure a balance between process stability and performance objectives.
[0022] Preferably, in step S3, the search space constraints for Bayesian optimization include: Chemical composition constraints: The total chemical composition shall be 100±0.5%, and trace impurity elements are allowed, with a total content of ≤0.1%; Process parameter constraints: meet the safety range of the equipment and the feasibility of industrial production.
[0023] Preferably, in step S4, the experimental verification involves: inputting the parameter combinations of candidate schemes into the forward prediction model, calculating the relative error between the prediction performance and the target performance, with the verification standard being a relative error ≤ 3%; conducting small-batch trial production on the verified parameter combinations, and using the measured performance data to update the training dataset; supplementing the modeling dataset with the measured data, retraining the forward prediction model, and achieving iterative improvement in model accuracy.
[0024] The beneficial effects of this invention are: 1) This invention constructs a full-chain nonlinear mapping model encompassing "chemical composition—manufacturing process—mechanical properties—service performance," and combines SHAP analysis and Bayesian optimization to achieve efficient reverse design starting from service performance targets. The design cycle is significantly shortened, and R&D costs are substantially reduced. This invention uses LightGBM to construct a high-precision forward prediction model and leverages Bayesian optimization to efficiently search for the optimal solution in a high-dimensional parameter space, avoiding the repeated trials and tests of traditional "trial and error methods." Experimental results show that the design cycle can be shortened from the traditional 6 to 12 months to 2 to 3 weeks, improving efficiency by over 90%; R&D costs are reduced by over 80%, making it particularly suitable for scenarios with high R&D efficiency requirements, such as high-speed rail and heavy-haul railways.
[0025] 2) Service performance-oriented approach to improve the actual performance of wheels. This invention is the first to directly incorporate service performance indicators such as wear rate, contact stress, and fatigue life into the objective function of reverse engineering, realizing a direct correlation between "performance requirements and parameter design". In practical applications, this can increase the service life of wheels by more than 10%, reducing the frequency of track maintenance and operating costs.
[0026] 3) Closed-loop optimization to continuously improve design accuracy. This invention establishes a closed-loop mechanism of "design-verification-update," where measured data after each experimental verification can be fed back into the model training stage to continuously optimize the prediction model. As the number of applications increases, the model accuracy continuously improves (e.g., R² after verification). 2 (From 0.88 to 0.93), resulting in long-term technological accumulation.
[0027] 4) Strong cross-scenario adaptability and customized design: By adjusting the weight of service performance indicators, this invention can flexibly adapt to the wheel design requirements of different operating scenarios such as high-speed rail, heavy-haul railway, and subway, without the need for repeated model development, which significantly improves the versatility and scalability of the technology. Attached Figure Description
[0028] Figure 1 This is a system flowchart of the reverse design method for wheel materials based on machine learning in an embodiment of the present invention; Figure 2 This is a feature importance ranking chart for SHAP analysis in this embodiment of the invention; Figure 3 The figure shows the Bayesian optimization iteration curve in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] The present invention aims to develop a product with the following target properties: hardness ≥ 295 HV and wear rate ≤ 2.0 × 10⁻⁶. -6 mm 3 Taking a high-speed train wheel with a capacity of / N·m as an example, the provided technical solution is adopted: a reverse design method for wheel materials based on machine learning, such as... Figure 1 As shown, it includes the following steps: S1. Data Acquisition and Preprocessing: Acquire the chemical composition, process parameters, and corresponding mechanical and service performance data of the wheels (such as wear rate and contact stress), and obtain the modeling dataset after cleaning and normalization. Furthermore, based on the wheel factory's recent production data and laboratory test results (a total of 3400 samples): Industrial production data (3000 sets): Production records of a wheel manufacturing company from 2018 to 2023, including the composition, heat treatment process and mechanical properties (rim tensile strength, hardness, rim yield strength) of wheel steel (such as ER7, ER8, ER9, CL65 materials). Laboratory-developed experimental data (400 sets): Data from a laboratory's double-disc rolling test, including variables such as the compositional combinations of C (0.48-0.72%), Si (0.20-0.60%), Mn (0.60-1.20%), Cr (0.10-0.50%), and Mo (0.05-0.20%), and the process combinations of quenching temperature (800-1000℃) and tempering time (1-4h), corresponding to the measured wear rate and contact stress; including composition (e.g., carbon content C: 0.48-0.72%), process parameters (e.g., normalizing temperature: 860-920℃), and wear rate ≤ 2.0 × 10⁻⁶. -6 mm 3 / N·m; Literature data (80 sets): Retrieved relevant research data on high-speed rail wheels published in journals such as "China Railway Science" and "Materials Engineering", and supplemented the sample of the influence of special components (such as V and Nb); the publication time of the literature is within the last 10 years, and the data validity verification standard (must include a complete correspondence between "component-process-performance", and the experimental method is consistent with the present invention).
[0031] The collected data were grouped according to the correspondence between alloy composition, processing technology and performance to obtain data groups.
[0032] Missing value handling (for 60 sets of missing data) Missing features: mainly Al content (32 groups), tempering time (18 groups), and wear rate (10 groups). Solution: Implement random forest imputation based on RandomForestRegressor from the Python sklearn library, with parameters set to n_estimators=100 (number of trees) and max_depth=8 (tree depth). Error verification: 20 sets of known complete data were randomly selected, and the missing data were simulated and then imputed using this method. The results showed that the average imputation error was 3.2% (e.g., if the true Mn content of a sample is 0.90%, the imputed value is 0.87%, and the error is 3.3%), which meets the requirement of "imputation error ≤ 5%".
[0033] After preprocessing to remove 10 sets of outlier data, the remaining data was normalized using a Min-Max formula. (Mapped to the [0,1] interval), taking the "quenching temperature" feature as an example: Original data range: 820℃ (x min -980℃ (x)max If the quenching temperature of a sample is 920℃, then the normalized value is: =0.625; other features are processed using the same logic to ensure consistent input dimensions for the model.
[0034] S2. Forward Prediction Model Training: The data sets are randomly divided, with 80% of the data sets used as the training set and the remaining 20% as the test set. A LightGBM regression prediction model is trained to construct a forward prediction model with "chemical composition + process parameters" as input and "mechanical properties" and "service performance" as outputs. The model hyperparameters are tuned using a Bayesian optimization algorithm, and the SHAP method is introduced to calculate the feature contribution. Then, the MASH value of the features is calculated to quantify the global importance of the features, and features that have a decisive impact on the design objectives are selected.
[0035] The process of tuning model hyperparameters using Bayesian optimization involves searching for the optimal combination of hyperparameters with the goal of minimizing the mean squared error (MSE). Hyperparameters include the learning rate, tree depth, number of leaf nodes, and number of iterations.
[0036] The SHAP method is introduced to analyze the contribution of model features. Features are ranked by importance and redundant features are removed based on their Shapley values. Figure 2 As shown; the formula for calculating the Shapley value is defined as follows:
[0037] in, F is the contribution of a feature to the model output, i.e., the Shapley value of the feature; F is the set of all features; S is the subset of input that does not contain features. The total number of characteristics; The performance predictions of the forward model when only a subset S is input; The marginal contribution to performance after adding the input subset S to the features.
[0038] The importance of each feature was analyzed using SHAP, as shown in Table 1.
[0039] Table 1. Importance of each feature in SHAP analysis ; The MASH value is the "mean absolute contribution" of a feature to all samples, quantifying the global importance of the feature. The formula for calculating the MASH value (Mean Absolute SHAP) of a feature is as follows:
[0040] in The mean absolute Shapley value of the i-th feature; N is the total number of samples in the modeling dataset; Let be the Shapley value of the i-th feature in the j-th sample.
[0041] Feature filtering rules: Retain MASH i ≥1.2×MASH avg And MASH i valid >MASH i total The features of MASH are used to eliminate redundant features; avg MASH is the average of all feature MASH values. i valid MASH is the MASH value of the i-th feature in the service performance compliance sample. i total is the MASH value of the i-th feature in the global sample.
[0042] In this embodiment, the data with input features from the training set are used as input, and wear rate and hardness are used as output. The model selected is LightGBM, and the parameter optimization range is: Learning rate: 0.01-0.3 Tree depth (max_depth): 3-10 Number of iterations (n_estimators): 100-1000 The final prediction error of the LightGBM model is ≤3.5%, R0 2 =0.8.
[0043] S3. Reverse Design Model Construction and Solution: With preset service performance targets as constraints and the equipment safety range and production cost of process parameters as boundary conditions, the chemical composition and process parameter combination that meets the performance requirements is solved through Bayesian optimization.
[0044] The objective function of Bayesian optimization is defined as:
[0045] Where X is a vector of chemical composition and process parameters; n is the number of service performance indicators; For performance weighting coefficients; Let i be the model prediction value of the i-th service performance item. Let i be the target value for the service performance of the i-th item; The standard deviation of the process parameters; This represents the process stability coefficient.
[0046] The process stability coefficient λ is determined based on the equipment accuracy, and its calculation formula is as follows: λ
[0047] in, The standard deviation of the process parameters; The maximum allowable standard deviation for the equipment (e.g., maximum temperature fluctuation of the quenching furnace ±8℃). =8; λ∈[0.1,0.5] is required to ensure a balance between process stability and performance targets (in the examples). =5, λ=0.375, which meets the requirements.
[0048] The search space constraints for the Bayesian optimization include: Chemical composition constraints: The total chemical composition shall be 100±0.5%, and trace impurity elements are allowed, with a total content of ≤0.1%; Process parameter constraints: meet the safety range of the equipment and the feasibility of industrial production.
[0049] In this embodiment of the invention, the optimization targets are: hardness ≥ 295 HV, wear rate ≤ 2.0 × 10⁻⁶. -6 mm 3 / N·m; In the constraints: Service performance: Contact stress ≤ 300 MPa; Process feasibility: Quenching temperature 820-980℃ (equipment safety range); Economic efficiency: Molybdenum (Mo) content ≤ 0.1% (reduces the cost of precious metals); The optimization algorithm was run, using a Gaussian process as the surrogate model and the expected improvement (EI) as the acquisition function, for 50 iterations: Iterative convergence: With an initial objective function value of f(X) = 1.2, the function converges to f(X) = 0.28 after 35 iterations. The fluctuation in the subsequent 15 iterations is ≤0.02, thus convergence is determined. The final Bayesian optimization iteration curve is shown below. Figure 3 As shown, the convergence trend of the objective function value with the number of iterations is compared.
[0050] S4. Parameter Optimization and Verification: Taking into account performance satisfaction, process feasibility and economy, select the best 1 to 3 candidate schemes for experimental verification; add the new data obtained from the experimental verification to the dataset in step S1, and retrain the positive prediction model in step S2 to achieve continuous iteration and performance improvement of the model, forming a self-evolving material closed-loop design method.
[0051] The experimental verification was conducted by inputting the parameter combinations of candidate schemes into the forward prediction model and calculating the relative error between the prediction performance and the target performance. The verification standard was that the relative error was ≤3%.
[0052] In this embodiment, from the 10 feasible solutions obtained through optimization, the top 3 optimal solutions are selected according to "performance satisfaction + process cost", as shown in Table 2: Table 2 Optimal Solutions for the First 3 Groups ; The selected schemes were prepared and tested in the laboratory. The results showed that the actual performance of one of the schemes was: hardness 298 HV, wear rate 1.8 × 10⁻⁶. -6 mm 3 / N·m, which fully meets the preset target.
[0053] This new set of "composition-process-performance" data is added to the original database. The expanded database is then used to incrementally update the LightGBM model in S2, further enhancing the model's prediction accuracy and generalization ability in future designs. Model R 2 The mean square error (MSE) has been improved to 0.88, meeting engineering accuracy requirements. The updated model can be directly used for the next high-speed rail wheel design, shortening the design cycle from the traditional 6-10 months to within 2 weeks and reducing the number of trial productions by more than 90%.
[0054] This invention not only solves the problems of long design cycles, high costs, and disconnect from actual service requirements in traditional wheel design, but also significantly improves the design accuracy and efficiency in small sample scenarios through a closed-loop optimization mechanism, which has important engineering application value and promotion prospects.
[0055] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0056] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0057] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0058] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0059] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A reverse design method for wheel materials based on machine learning, characterized in that, Includes the following steps: S1. Data Acquisition and Preprocessing: Acquire the chemical composition, process parameters, and corresponding mechanical and service performance data of the wheels. After cleaning and normalization, obtain the modeling dataset. S2. Forward Prediction Model Training: Train a LightGBM regression prediction model to construct a forward prediction model with "chemical composition + process parameters" as input and "mechanical properties" and "service performance" as output; optimize the model hyperparameters through Bayesian optimization algorithm, introduce the SHAP method to calculate the feature contribution, and then calculate the MASH value of the feature to quantify the global importance of the feature, and screen out the features that have a decisive impact on the design goal; S3. Reverse design model construction and solution: With the preset service performance target as a constraint and the equipment safety range and production cost of the process parameters as boundary conditions, the chemical composition and process parameter combination that meets the performance requirements is solved by Bayesian optimization. S4. Parameter optimization and verification: Taking into account performance satisfaction, process feasibility and economy, select the best 1 to 3 candidate solutions for experimental verification. The new data obtained from the experimental verification is added to the dataset in step S1, and the positive prediction model in step S2 is retrained to achieve continuous iteration and performance improvement of the model, forming a self-evolving material closed-loop design method.
2. The reverse design method for wheel materials based on machine learning according to claim 1, characterized in that: Step S1 specifically includes: S11. Data Acquisition: Collect historical data of the target wheel system through multiple channels such as literature review, laboratory experiments, industrial production records and line service monitoring, including the chemical composition, process parameters and corresponding mechanical and service performance data of the wheel. S12. Missing value handling: Missing data are filled using random forest imputation, with an imputation error ≤5%; S13. Outlier handling: Outliers are identified and removed based on the 3σ principle, and the proportion of outliers is controlled within 8% of the total data volume. S14. Feature Normalization: The input features are standardized using the Min-Max normalization method. The normalization formula is as follows: ; in, x These are the original eigenvalues. , These are the minimum and maximum values of the feature in the dataset, respectively. These are the normalized eigenvalues.
3. The reverse design method for wheel materials based on machine learning according to claim 2, characterized in that: In step S11, the wheel data specifically includes: Industrial production data: Wheel material range includes ER7, ER8, ER9, and CL65 steel; Laboratory-based experimental data: The experimental equipment consisted of a double-disc rolling tester and a Vickers hardness tester; the experimental environment was at room temperature of 25±2℃. Literature data: The publication time of the literature is within the last 10 years, and the data validity verification standard is used.
4. The machine learning-based reverse engineering method for wheel materials according to claim 1 or 2, characterized in that: The chemical composition data includes carbon, silicon, manganese, chromium, molybdenum, phosphorus, sulfur, chromium, nickel, vanadium, copper, aluminum, titanium, niobium and the corresponding mass percentages of these elements; the process parameters include quenching temperature and tempering time; the mechanical properties include tensile strength, yield strength, Vickers hardness and impact toughness; and the service performance data includes wear rate and contact stress.
5. The reverse design method for wheel materials based on machine learning according to claim 1, characterized in that: The positive prediction model is LightGBM, and its parameter selection range is as follows: n_estimators: 100-300; learning_rate: 0.05-0.2; max_depth: 3-8; min_child_samples:10-30.
6. The reverse design method for wheel materials based on machine learning according to claim 1, characterized in that: In step S2, the process of tuning the model hyperparameters using the Bayesian optimization algorithm specifically involves: searching for the optimal combination of hyperparameters using the Bayesian optimization algorithm, with the goal of minimizing the mean squared error (MSE) of the model; the hyperparameters include the learning rate, tree depth, number of leaf nodes, and number of iterations.
7. The reverse design method for wheel materials based on machine learning according to claim 1, characterized in that: In step S2, the SHAP method is introduced to analyze the contribution of model features, and the features are ranked by importance and redundant features are removed based on the Shapley value. The formula for calculating the Shapley value is defined as follows: ; in, F is the contribution of a feature to the model output, i.e., the Shapley value of the feature; F is the set of all features; S is the subset of input that does not contain features. The total number of characteristics; The performance predictions of the forward model when only a subset S is input; The marginal contribution of adding the input subset S to the features to the performance; The MASH value is the "average absolute contribution" of a feature to all samples, which quantifies the global importance of a feature. The MASH value of a feature is calculated using the following formula: ; in The mean absolute Shapley value of the i-th feature; N is the total number of samples in the modeling dataset; The Shapley value of the i-th feature in the j-th sample; Feature filtering rules: Retain MASH i ≥1.2×MASH avg And MASH i valid >MASH i total The features of MASH are used to eliminate redundant features; avg MASH is the average of all feature MASH values. i valid MASH is the MASH value of the i-th feature in the service performance compliance sample. i total is the MASH value of the i-th feature in the global sample.
8. The reverse design method for wheel materials based on machine learning according to claim 1, characterized in that: In step S3, the objective function of the Bayesian optimization is defined as: ; Where X is a vector of chemical composition and process parameters; n is the number of service performance indicators; For performance weighting coefficients; Let i be the model prediction value of the i-th service performance item. Let i be the target value for the service performance of the i-th item; The standard deviation of the process parameters; This represents the process stability coefficient; the process stability coefficient λ is determined based on the equipment accuracy, and its calculation formula is: l ; in, The standard deviation of the process parameters; λ represents the maximum allowable standard deviation of the equipment; λ∈[0.1,0.5], ensuring a balance between process stability and performance objectives.
9. The machine learning-based reverse design method for wheel materials according to claim 1, characterized in that: In step S3, the search space constraints for Bayesian optimization include: Chemical composition constraints: The total chemical composition shall be 100±0.5%, and trace impurity elements are allowed, with a total content of ≤0.1%; Process parameter constraints: meet the safety range of the equipment and the feasibility of industrial production.
10. The machine learning-based reverse design method for wheel materials according to claim 1, characterized in that: In step S4, the experimental verification is as follows: input the parameter combination of the candidate scheme into the positive prediction model, calculate the relative error between the prediction performance and the target performance, and the verification standard is that the relative error is ≤3%.
Citation Information
Patent Citations
A reverse design method for high-performance magnesium alloys based on machine learning
CN116844673B