Heat treatment simulation database-oriented material parameter extraction method

By training linear regression, decision tree, random forest and neural network models in the heat treatment simulation database, and using a combined hyperparameter optimization algorithm, the problems of insufficient data integrity and single retrieval function of the heat treatment database are solved, efficient data utilization and extensive retrieval are achieved, and database dynamic expansion is adapted to.

CN120278027APending Publication Date: 2025-07-08SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510412131.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing heat treatment database has problems such as insufficient data integrity and single retrieval function, which limits the application and development of heat treatment simulation calculation.

Method used

The linear regression model, decision tree model, random forest model and neural network model are trained using material heat treatment simulation database, and the model hyperparameters are optimized through a combined hyperparameter self-optimization algorithm to improve data utilization and retrieval capabilities.

Benefits of technology

It realizes efficient data utilization and extensive retrieval of material heat treatment simulation database, solves the problem of data loss, reduces the blindness and inconvenience of manually adjusting hyperparameters, and adapts to the dynamic expansion of database scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278027A_ABST
    Figure CN120278027A_ABST
Patent Text Reader

Abstract

A heat treatment simulation database-oriented material parameter extraction method comprises the following steps: generating a training set according to a material heat treatment simulation database, and respectively training and constructing an obtained linear regression model, a decision tree model, a random forest model and a neural network model; and in the online stage, the trained model is adopted to calculate data required by material heat treatment simulation. According to the method, the linear regression model, the decision tree model, the random forest model and the neural network model are trained by using the original data of the database, so that a user can conveniently and efficiently obtain the required material heat treatment parameters by using the models, and the utilization rate and retrieval capability of the material heat treatment simulation database are improved; inconvenience caused by data integrity deficiency of the heat treatment simulation database to material research work is avoided to a certain extent, and data mining can be performed by utilizing existing data in the database so as to provide missing material data in the database for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of material heat treatment, and specifically to a method for extracting material parameters for a heat treatment simulation database. Background Art

[0002] Heat treatment is an important material processing technology. By means of heating, heat preservation and cooling, etc., solid materials can obtain the expected microstructure and properties. With the development of computer science and technology, using simulation software to simulate the heat treatment process has become an important means in material research. Heat treatment simulation involves numerous material parameters, and these parameters are often complex functions related to elements such as microstructure, temperature, and pressure. Therefore, the problems of collecting, storing, and retrieving them need to be solved by establishing a heat treatment material database. Limited by the existing experimental technologies and calculation means, the current heat treatment databases generally have the problem of insufficient data integrity. At the same time, many databases can only provide a single data retrieval and query function, and there are great limitations in the types and ranges of elements available for retrieval. These deficiencies greatly restrict the application and development of heat treatment simulation calculations. Summary of the Invention

[0003] Aiming at the above deficiencies existing in the prior art, the present invention proposes a method for extracting material parameters for a heat treatment simulation database, which uses the original data in the database to train linear regression models, decision tree models, random forest models, and neural network models. Users can use these models to conveniently and efficiently obtain the required material heat treatment parameters, improve the utilization rate and retrieval ability of the material heat treatment simulation database, and to a certain extent avoid the inconvenience caused by the lack of data integrity in the heat treatment simulation database to material research work.

[0004] The present invention is realized by the following technical solutions:

[0005] The present invention relates to a method for extracting material parameters for a heat treatment simulation database, which generates a training set according to the material heat treatment simulation database and is used to train and construct the obtained linear regression model, decision tree model, random forest model, and neural network model respectively; in the online stage, the trained models are used to calculate the data required for material heat treatment simulation. Technical Effects

[0006] The present invention uses the original data in the material heat treatment simulation database to train a machine learning model and designs a combined hyperparameter self-optimization algorithm to achieve hyperparameter self-optimization. Users can use these models to conveniently and efficiently obtain the required material heat treatment parameters. Compared with the traditional material heat treatment simulation database that can only provide the existing data in the database to users, this technology improves the data utilization rate and retrieval scope of the material heat treatment simulation database, and can better cope with the problem of the variety and easy loss of material heat treatment data in practical applications. In addition, as the database scale dynamically expands in practical applications, machine learning models will also be continuously generated and updated. For this scenario, the machine learning hyperparameter optimization method provided by the present invention can play a better role, and to a great extent solve the problems of blindness, randomness, and inconvenient maintenance faced by manual adjustment of hyperparameters during the continuous update of the database machine learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 is a flowchart of the present invention;

[0008] Figure 2 is a flowchart of combined hyperparameter optimization;

[0009] Figure 3 is a schematic diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] As Figure 3 shown, a material parameter extraction system for a heat treatment simulation database according to this embodiment includes: a material data storage unit, a model training unit, a model storage unit, and a user access unit, where: the material data storage unit stores material heat treatment data, the model training unit periodically extracts the data in the material data storage unit, trains a machine learning model and realizes hyperparameter self-optimization, the model storage unit stores the machine learning model output by the model training unit, and the user access unit extracts the model in the model storage unit and calculates the corresponding material heat treatment data according to user requirements.

[0011] As Figure 1 shown, a material parameter extraction method for a heat treatment simulation database based on the above system according to this embodiment includes:

[0012] S1. Circularly extract material parameter data from the material heat treatment simulation database to form a classified data set, check whether each data set needs to be retrained for the machine learning model. When retraining is required, the data in the database is divided into several data sets according to different material parameters and different single-phase structures, and each data set independently trains and tests the machine learning model.

[0013] The sample features of the dataset include: material composition, temperature, ambient pressure, and equivalent plastic strain; the sample labels of the dataset include yield strength.

[0014] The basis for whether the dataset is to be used to train a machine learning model is as follows: when specified by the administrator, train; if the current data volume meets the conditions, then train. Specifically: (n > N ∧ n p = None) ∨ n > m · n p where: n is the current data volume of this dataset, n p is the data volume at the last training of this type of dataset, and m is a manually set magnification factor.

[0015] The condition being met means: when the dataset has never been used to train a machine learning model and the data volume reaches a certain scale, then train; when the data volume of the dataset exceeds a certain degree of the data volume used in the last training, then retrain to update the machine learning model.

[0016] S2. Divide the dataset into a training set and a test set. The training set is used to train the machine learning model, and the test set is used to test the model.

[0017] S3. Use the training set obtained in step S2 to train the constructed linear regression model, decision tree model, random forest model, and neural network model respectively;

[0018] For the linear regression model mentioned above, no hyperparameter optimization is required. The test result is the R2 score.

[0019] The test result, the R2 score, specifically is: where: R 2 is the R2 score of the test set, n is the number of samples in the test set, y i is the sample label of the test set, is the predicted value of the model, is the average value of the sample labels of the test set.

[0020] For the decision tree model mentioned above, it is trained in the way shown by Figure 2 , that is: use the combined hyperparameter optimization algorithm to optimize the hyperparameters of the decision tree model, compare the optimization results of each optimization algorithm, and save the optimal model and test results. Specifically include:

[0021] i) Set the hyperparameter search space: For the decision tree model, the minimum number of samples required to split internal nodes (non - leaf nodes), the minimum number of samples required for leaf nodes, and the number of features considered when looking for the best split point need to be optimized. Take a roughly reasonable range for each of these parameters to form the hyperparameter search space.

[0022] ii) TPE-based Bayesian optimization, covariance matrix adaptive evolution strategy, random search and grid search algorithms are used in parallel to optimize the hyperparameters of the decision tree model.

[0023] The TPE-based Bayesian optimization specifically includes:

[0024] A1) Set the maximum optimization round n1 and initialize the optimization round i=1.

[0025] A2) Use the TPE-based Bayesian optimization algorithm to calculate a set of hyperparameter combinations to be observed through the maximization formula, specifically: Where: x is the hyperparameter combination vector selected in the current round, D is the hyperparameter search space, g(x) = P(x|y≥y * ) is the distribution of the hyperparameter combination vector x when the R2 score of the target test set is not lower than any threshold, g(x) is constructed using kernel density estimation based on the historical data of the previous round, l(x) = P(x|y <y * ) is the distribution of the hyperparameter combination vector x when the R2 score of the target test set is lower than any threshold, and g(x) and l(x) are constructed using kernel density estimation based on historical data from previous rounds.

[0026] A3) Use hyperparameter combinations to construct a decision tree model.

[0027] A4) Use the training set obtained in step 2 to train the decision tree model, and use the test set obtained in step S2 to test the decision tree model. The test result is the R2 score.

[0028] A5) Judgment when It indicates that a machine learning model that has reached the predetermined R2 score target has been trained, and hyperparameter optimization is stopped in advance. The current model and test results are saved, and other parallel optimization algorithm processes are terminated. Otherwise, step A2 is continued, where: i is the current optimization round of the Bayesian optimization algorithm based on TPE, is the R2 score of the neural network model trained in the current round, and C is the predetermined R2 score target.

[0029] A6) Determine i=n1: When i≤n1, go to step A2. Otherwise, the optimization rounds reach the predetermined total optimization rounds, and record the optimal model and test results that appear in the optimization process, where: n1 is the maximum optimization round of the Bayesian optimization algorithm based on TPE set manually.

[0030] The covariance matrix adaptive evolution strategy, the hyperparameter combination to be observed is calculated by the covariance matrix adaptive evolution strategy algorithm, specifically: in: The hyperparameter combination vector selected for the current round, assumed to be the i-th candidate solution in the t-th generation, m t is the mean vector of the t-th generation, σ t is the step size, N i (0, C t ) is a random vector generated from a zero-mean, multivariate normal distribution, C t is the covariance.

[0031] The mean vector m t+1 , covariance matrix C t+1 , step size σ t+1 are updated respectively in the following ways: where: μ is the number of candidate solutions for calculating the new mean (usually the top μ candidate solutions with the highest fitness), represents the candidate solution with the i-th best fitness in the t-th generation, ω i is the weight of the i-th candidate solution, c1 and c μ are the learning rates, p c,t+1 is the covariance update path, p σ,t+1 is the step size control path, c σ is the path decay coefficient, d σ is the damping coefficient, E‖N(0, I)‖ is the expected length of the standard normal distribution.

[0032] The maximum optimization rounds n2 and n1 of the artificially set covariance matrix adaptive evolution strategy algorithm do not have to be the same.

[0033] For the random search algorithm described, the hyperparameter combination to be observed is calculated by the random search algorithm, specifically: x ∼ U D , where: x is the hyperparameter combination vector selected for the current round, U D is the uniform distribution sampling within the hyperparameter search space D.

[0034] The maximum optimization rounds n3 and n1 of the artificially set random search algorithm do not have to be the same.

[0035] For the grid search algorithm described, the hyperparameter combination to be observed is obtained by the grid search algorithm: First, a set of candidate values is artificially set for each hyperparameter, and then the Cartesian product of these candidate values is generated to form a combination grid of hyperparameters. The hyperparameter combinations are sequentially selected in each optimization round. The number of hyperparameter combinations n4 and n1 do not have to be the same.

[0036] The comparison of the optimization results of each optimization algorithm refers to: when all parallel hyperparameter optimization algorithms complete the optimization of the predetermined number of rounds without early stopping, compare the optimal models obtained by each optimization algorithm, select the model with the highest R2 score as the global optimal model and save it in the database, and save its R2 score. The saving of the decision tree model and the test results is realized through the database.

[0037] The described random forest model is trained in the following way: use the training set obtained in step S2 to train the random forest model, use the test set obtained in step S2 to test the random forest model, use the combined hyperparameter optimization algorithm to optimize the hyperparameters of the random forest model, and save the optimized model and test results.

[0038] The optimization of the hyperparameters of the random forest model includes the number of decision trees in the model to be optimized, the minimum number of samples required to split internal nodes (non-leaf nodes), the minimum number of samples required for leaf nodes, and the number of features considered when searching for the best split point.

[0039] The described neural network model is trained in the following way: use the training set obtained in step S2 to train the neural network model, use the test set obtained in step S2 to test the neural network model, use the combined hyperparameter optimization algorithm to optimize the hyperparameters of the neural network model, and save the optimized model and test results.

[0040] The optimization of the hyperparameters of the neural network model includes the learning rate to be optimized, the L2 regularization coefficient, and the number of neurons in each layer of the hidden layer.

[0041] S4. In the online stage, use the trained model to calculate the data required for material heat treatment simulation and wait for the data set to meet the conditions for model update. The model can be further updated: the call of the machine learning model is realized through the database, and the user can arbitrarily select from the four obtained machine learning models. Input the target sample features (variables such as material composition, temperature, and pressure) into the model, and the required material heat treatment simulation data can be output.

[0042] Through specific actual experiments, take the martensite yield strength data set as an example to illustrate the application effect of the present invention. The sample features of this data set are: material composition, temperature, ambient pressure, equivalent plastic strain; the sample label of this data set is yield strength, with a total of 2311 pieces of data. Apply the machine learning application solution and the hyperparameter automatic optimization algorithm of the related model proposed by the present invention to this data set, and finally one linear regression model, one decision tree model, one random forest model, and one neural network model and their test results will be saved. Table 2 is a summary of the test results.

[0043] Table 1 Partial data example of the martensite yield strength data set of metal materials Material Temperature / °C Ambient pressure / atm Equivalent plastic strain Yield strength / Pa 34CrNi3Mo 300.0 1.0 0.01 1.49602E9 20CrMo 600.0 1.0 1.7 3.506176E8 20CrMnMo 700.0 1.0 0.0075 2.18177E8 20CrMo 1100.0 1.0 2.5 2.01524E7 20Cr2Ni4 200.0 1.0 0.095 1.31146E9 20CrMnTi 400.0 1.0 0.005 9.313361E8 20CrNi2Mo 1100.0 1.0 1.3 2.3006E7

[0044] Table 2 Summary of Test Results of Martensite Yield Strength Dataset R2 score of linear regression model 0.8571 R2 score of decision tree model 0.9952 R2 score of random forest model 0.9951 R2 score of neural network model 0.9928

[0045] The closer the R2 score is to 1, the better the prediction performance of the model. It can be seen from the results that the present invention obtains decision tree, random forest and neural network models with excellent prediction performance within a limited number of optimization times (the significance of linear regression in the present invention is to provide reference for linear material data, and in this embodiment, a non-linear material data is used). Compared with the traditional material heat treatment simulation database without machine learning function, the present invention improves the data utilization rate and retrieval range of the material heat treatment simulation database, can better cope with the problem of various and easily missing material heat treatment data in practical applications, and to a certain extent avoids the inconvenience caused by the lack of data integrity in the heat treatment simulation database to material research work. In addition, as the database scale dynamically expands in practical applications, machine learning models will also be continuously generated and updated. If hyperparameters rely on manual adjustment, it will face the problems of time-consuming, laborious and difficult to maintain. For this scenario, the combined hyperparameter optimization method obtained by the present invention can play a better role, and to a great extent solve the problems of blindness, randomness and inconvenient maintenance faced by manual adjustment of hyperparameters during the continuous update of the database machine learning model.

[0046] The above specific implementation can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation, and all implementation solutions within its scope are subject to the present invention.

Claims

1. A method for extracting material parameters for a heat treatment simulation database, characterized in that Generate a training set based on the material heat treatment simulation database, and train the constructed linear regression model, decision tree model, random forest model, and neural network model respectively; in the online stage, use the trained models to calculate the data required for material heat treatment simulation.

2. The method for extracting material parameters for a heat treatment simulation database according to claim 1, characterized in that specifically Including: S1. Repeatedly extract material parameter data from the material heat treatment simulation database to form a classified data set, and check whether each data set needs to be retrained for the machine learning model. When retraining is required, divide the data in the database into several data sets according to different material parameters and different single-phase microstructures, and each data set independently trains and tests the machine learning model; S2. Divide the data set into a training set and a test set. The training set trains the machine learning model, and the test set tests the model; S3. Use the training sets obtained in step S2 to train the constructed linear regression model, decision tree model, random forest model, and neural network model respectively; S4. In the online stage, use the trained models to calculate the data required for material heat treatment simulation and wait for the data set to meet the conditions for model update.

3. The method for extracting material parameters for a heat treatment simulation database according to claim 1 or 2, characterized in that, The training method of the decision tree model is as follows: Use the combined hyperparameter optimization algorithm to optimize the hyperparameters of the decision tree model, compare the optimization results of each optimization algorithm, and save the optimal model and test results. Specifically, it includes: i) Set the hyperparameter search space: For the decision tree model, the internal nodes to be optimized for splitting, that is, the minimum number of samples required for non-leaf nodes, the minimum number of samples required for leaf nodes, and the number of features considered when finding the best splitting point, are each taken within a roughly reasonable range to form the hyperparameter search space; ii) Parallelly use the Bayesian optimization based on TPE, covariance matrix adaptation evolution strategy, random search, and grid search algorithms to optimize the hyperparameters of the decision tree model.

4. The method for extracting material parameters for a heat treatment simulation database according to claim 3, wherein The Bayesian optimization based on TPE specifically includes: A1) Set the maximum number of optimization rounds n1, and initialize the optimization round i = 1; A2) Use the TPE-based Bayesian optimization algorithm to calculate a set of hyperparameter combinations to be observed by maximizing the formula, specifically: where: x is the hyperparameter combination vector selected in the current round, D is the hyperparameter search space, and g(x) = P(x|y≥y * ) is the distribution of the hyperparameter combination vector x when the R2 score of the target test set is not lower than any threshold. g(x) is constructed using kernel density estimation based on the historical data of previous rounds. l(x) = P(x|y<y * ) is the distribution of the hyperparameter combination vector x when the R2 score of the target test set is lower than any threshold. g(x) and l(x) are constructed using kernel density estimation based on the historical data of previous rounds; A3) Use the hyperparameter combination to construct the decision tree model; A4) Use the training set obtained in step S2 to train the decision tree model, and use the test set obtained in step S2 to test the decision tree model. The test result is the R2 score; A5) Judgment when It indicates that a machine learning model that has reached the predetermined R2 score target has been trained, and hyperparameter optimization is stopped in advance. The current model and test results are saved, and other parallel optimization algorithm processes are terminated. Otherwise, go to step A2, where: i is the current optimization round of the Bayesian optimization algorithm based on TPE, is the R2 score of the neural network model trained in the current round, and C is the predetermined R2 score target; A6) Judge i = n1: When i ≤ n1, go to step A2. Otherwise, when the optimization round reaches the predetermined total number of optimization rounds, record the optimal model and test results that appear during the optimization process, where: n1 is the maximum number of optimization rounds of the Bayesian optimization algorithm based on TPE set by humans; The R2 score mentioned above is specifically as follows: Where: R 2 is the R2 score of the test set, n is the number of samples in the test set, y i is the label of the test set sample, is the predicted value of the model, is the average value of the labels of the test set samples.

5. The material parameter extraction method for a heat treatment simulation database according to claim 3, characterized in that For the covariance matrix adaptation evolution strategy, the hyperparameter combination to be observed is calculated by the covariance matrix adaptation evolution strategy algorithm, specifically as follows: Where: is the hyperparameter combination vector selected in the current round, assumed to be the i-th candidate solution in the t-th generation, m t is the mean vector of the t-th generation, σ t is the step size, N i (0, C t ) is a random vector generated from a zero-mean, multivariate normal distribution, and C t is the covariance; The mean vector m t+1 , the covariance matrix C t+1 , and the step size σ t+1 are updated respectively in the following ways: where: μ is the number of candidate solutions for calculating the new mean (usually the top μ candidate solutions with the highest fitness), represents the candidate solution with the i-th best fitness in the t-th generation, ω i is the weight of the i-th candidate solution, c1 and c μ are the learning rates, p c,t+1 is the covariance update path, p σ,t+1 is the step size control path, c σ is the path decay coefficient, d σ is the damping coefficient, and E‖N(0,I)‖ is the expected length of the standard normal distribution.

6. The method for extracting material parameters for a heat treatment simulation database according to claim 3, characterized in that For the described random search algorithm, the hyperparameter combination to be observed is calculated by the random search algorithm, specifically: x ∼ U D , where: x is the hyperparameter combination vector selected in the current round, and U D is a uniform distribution sampling within the hyperparameter search space D; For the grid search algorithm, the hyperparameter combinations to be observed are obtained by the grid search algorithm: First, a set of candidate values is set for each hyperparameter by humans, and then the Cartesian product of these candidate values is generated to form a combination grid of hyperparameters. Each optimization round sequentially selects the hyperparameter combinations in the grid.

7. The method for extracting material parameters for a heat treatment simulation database according to claim 3, characterized in that The comparison of the optimization results of each optimization algorithm means that when all parallel hyperparameter optimization algorithms complete the optimization of a predetermined number of rounds without early stopping, the optimal models obtained by each optimization algorithm are compared, and the model with the highest R2 score is selected as the global optimal model and saved in the database, and its R2 score is saved. The saving of the decision tree model and the test results is realized through the database.

8. The method for extracting material parameters for a heat treatment simulation database according to claim 1 or 2, characterized in that, The described random forest model is trained in the following way: The training set obtained in step S2 is used to train the random forest model, the test set obtained in step S2 is used to test the random forest model, the combined hyperparameter optimization algorithm is used to optimize the hyperparameters of the random forest model, and the optimized model and test results are saved; The optimization of the hyperparameters of the random forest model includes the number of decision trees in the model to be optimized, the minimum number of samples required to split internal nodes, i.e., non-leaf nodes, the minimum number of samples required for leaf nodes, and the number of features considered when finding the best split point.

9. The method for extracting material parameters for a heat treatment simulation database according to claim 1 or 2, characterized in that, The described neural network model is trained in the following way: The training set obtained in step S2 is used to train the neural network model, the test set obtained in step S2 is used to test the neural network model, the combined hyperparameter optimization algorithm is used to optimize the hyperparameters of the neural network model, and the optimized model and test results are saved; The optimization of the hyperparameters of the neural network model includes the learning rate to be optimized, the L2 regularization coefficient, and the number of neurons in each layer of the hidden layer.

10. A material parameter extraction system for a heat treatment simulation database implementing any one of the methods described in claims 1-9, characterized in that, It includes: A material data storage unit, a model training unit, a model storage unit, and a user access unit. Among them: The material data storage unit stores material heat treatment data. The model training unit periodically extracts the data from the material data storage unit, trains a machine learning model and realizes hyperparameter self-optimization. The model storage unit stores the machine learning model output by the model training unit. The user access unit extracts the model from the model storage unit and calculates the corresponding material heat treatment data according to the user's needs.

Citation Information

Cited By

  • Intelligent research and development system for new material research and development and material research and development method

    CN121687336A

  • Intelligent research and development system for new material research and development and material research and development method

    CN121687336B