A method and system for predicting the performance of zinc alloys and optimizing their composition
By combining machine learning and non-dominated sorting genetic algorithm, the problems of large zinc alloy composition design space and insufficient data were solved, and efficient zinc alloy performance prediction and composition optimization were achieved, which reduced development costs and improved prediction accuracy.
Patent Information
- Application Number
- CN202410728592.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-06-06
AI Technical Summary
Existing technologies in zinc alloy development have problems such as large composition design space, long development cycle, and high cost. In addition, the application of machine learning is limited by insufficient data sample size, making it difficult to achieve high-precision comprehensive mechanical property prediction and composition optimization.
A machine learning strategy combined with a non-dominated sorting genetic algorithm is used to perform data expansion and feature screening through a data processing module, and a multi-objective optimization model is constructed to achieve zinc alloy performance prediction and composition optimization.
High-precision zinc alloy performance prediction and composition optimization are achieved, which reduces development cycle and cost, improves the generalization ability of the model, and ensures the accuracy of alloy composition.
Smart Images

Figure CN118675666B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of metal material science and technology, and in particular to a method and system for predicting zinc alloy properties and optimizing composition. Background Art
[0002] Zinc, a key industrial metal, is in enormous demand, second only to steel, aluminum, and copper. Zinc mineral resources are abundant and inexpensive, but pure zinc has very low strength and hardness, and the performance of pure zinc parts often fails to meet the strength requirements of engineering components. Therefore, in practical applications, various elements are often added to form zinc alloys to meet these requirements. For example, Al, a key alloying element in zinc alloys, enhances strength, but different performance indicators correspond to different peak concentrations. Mg provides solid solution strengthening, but excessive addition can easily lead to hot cracking. The addition of Cu increases the number of strengthening phases and enhances the strengthening effect, but excessive addition can make the alloy brittle. In addition to the primary elements Al, Cu, and Mg, rare earth elements, salt modifiers, and elements such as Si, Mn, Ti, and Ni also affect the alloy.
[0003] However, faced with so many elements and their contents to choose from, traditional trial-and-error experiments require a long development cycle and high development costs for the vast unexplored composition space, posing a huge challenge to the development and growth of zinc alloys. Therefore, it is necessary to find an efficient alloy composition design method to accelerate the development of high-performance zinc alloys, thereby opening up a broader application space. If the various performance indicators of zinc alloys can truly replace copper alloys and aluminum alloys, it will be of great significance for saving energy and reducing raw material costs.
[0004] In recent years, machine learning technology has been proven to have great potential in the research and application of materials. In particular, supervised learning and deep learning methods can explore the intrinsic connection between influencing factors such as composition, processing technology, chemical properties and performance indicators such as tensile strength, elongation, and hardness from historical data. By constructing predictive models, they can quickly and accurately evaluate the performance of influencing variables such as new alloy composition ratios and different processes, thereby reducing the number of experiments and R&D costs.
[0005] Regarding the application of machine learning in the selection of alloying elements and their contents, patent application number 202410044916.0 discloses a method for predicting the plastic hardening behavior of metal materials using machine learning, which improves the accuracy of prediction of metal material properties while reducing the cost of research and engineering design to a certain extent. Patent application number 202311600473.0 discloses a method for predicting the compressive strength of magnesium phosphate cement-based composites based on machine learning, which provides a new method for determining the optimal mix ratio and compressive strength through a large number of pouring and curing tests. However, they only provide a more accurate prediction model and do not propose further optimization solutions for specific problems.
[0006] In the patent application number 202210139501.2, a design method for simultaneously optimizing the conductivity and hardness of a multi-component electrical contact alloy is disclosed, so as to simultaneously optimize the conductivity and hardness performance of the multi-component electrical contact alloy material. However, this technical solution only comprehensively optimizes the two properties of the material, and the contradiction between the performance of the material is not limited to the two properties, and only dual-target optimization is achieved.
[0007] In summary, the above-mentioned patents all describe model data. For the application of machine learning in the field of materials, the data sample size is a key issue. Most materials do not have large data sets, especially metal materials such as zinc alloys. Therefore, these design methods have great limitations. A suitable data expansion method is needed to better realize the widespread application of machine learning. Summary of the Invention
[0008] The technical problem to be solved by the present invention is how to stably predict the comprehensive mechanical properties and elemental composition optimization of zinc alloy materials with high precision. In order to overcome the defects of the above-mentioned prior art, the present invention provides a zinc alloy performance prediction and composition optimization method and system, including a zinc alloy performance prediction and composition optimization method and a zinc alloy performance prediction and composition optimization system.
[0009] The present invention provides a method for predicting zinc alloy properties and optimizing composition, comprising the following steps:
[0010] S1: Collecting data on the composition, forming process, and mechanical properties of zinc alloys based on public databases. The mechanical properties include tensile strength, elongation, and hardness, and synthesizing a total data set from the data.
[0011] S2: decomposing the total data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data by a data processing module, and performing data expansion using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain an expanded data set;
[0012] S3: Screening key features corresponding to mechanical properties from the expanded data set through the Pearson correlation coefficient and random forest feature importance evaluation in the expanded data set to obtain overall key features;
[0013] S4: Based on the overall key features, a ten-fold cross validation method is used to screen the optimal prediction model from multiple mechanical property prediction models using the average determination coefficient as an evaluation index to obtain the current performance prediction model;
[0014] S5: Optimizing hyperparameters of the current performance prediction model using grid search, and evaluating the regression model performance using the coefficient of determination and root mean square error to obtain an optimized model of the current performance prediction model;
[0015] S6: verifying the error rate of the optimization model of the current performance prediction model through a zinc alloy performance prediction experiment, and executing the next step when the error rate is not greater than a predetermined value; and when the error rate is greater than the predetermined value, collecting the experimental results into a total data set and returning to executing step S1;
[0016] S7: Perform multi-objective optimization on the overall key features through a non-dominated sorting genetic algorithm to obtain the optimal composition ratio of the mechanical properties of the zinc alloy.
[0017] Compared with the existing technology, the zinc alloy performance prediction and composition optimization method of the present application has the following advantages:
[0018] First, the present invention provides a zinc alloy performance prediction and composition optimization method, which adopts a combination of machine learning strategy and non-dominated sorting genetic algorithm, can achieve multi-objective optimization, quickly design zinc alloys with high comprehensive mechanical properties, and provide design ideas and references for the development of new zinc alloys.
[0019] Secondly, the present invention provides a zinc alloy performance prediction and composition optimization method, which performs data expansion based on the idea of mean and standard deviation, providing a solution for the lack of sample data for some alloys when constructing machine learning prediction models, improving the generalization ability of the model and reducing the risk of overfitting.
[0020] Finally, the present invention provides a method for zinc alloy performance prediction and composition optimization. Based on a machine learning algorithm, a target performance prediction model is established for alloy development. To a certain extent, it solves the problems of long cycle and high cost caused by the large composition design space of traditional design, and realizes accurate and effective performance prediction. It has important reference value and practical significance for the composition design of high-strength and toughness zinc alloys.
[0021] In a possible implementation, step S2 includes the following steps:
[0022] S21: removing all data whose forming process is not a casting process from the total data set by the data processing module to obtain a casting process data set;
[0023] S22: Decomposing the casting process data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data by the data processing module, to obtain three data sets;
[0024] S23: performing data expansion using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain multiple sub-data sets, which are used as the expanded data sets;
[0025] By removing all data where the forming process is not a casting process from the total data set, process consistency can be ensured, avoiding the impact of subsequent model quality due to different processing techniques. Compared with existing technologies, the above-mentioned technical solution based on the idea of mean and standard deviation for data expansion can provide a solution to the lack of sample data for some alloys when constructing machine learning prediction models, improve the generalization ability of the model, and reduce the risk of overfitting.
[0026] In a possible implementation, step S3 includes the following steps:
[0027] S31: deleting, from each sub-dataset of the expanded data set, the component items whose Pearson correlation coefficient with the mechanical property of the sub-dataset is less than a first predetermined value, to obtain respective initial component features;
[0028] S32: Deleting the component items whose random forest feature importance evaluation values are less than a second predetermined value from the initial component features of each sub-dataset of the expanded data set by using a random forest feature importance ranking method, to obtain respective final component features;
[0029] S33: combining all the final component features obtained in step S32 to obtain the overall key feature, and using all component items corresponding to the overall key feature as the final component features of each sub-dataset of the expanded data set;
[0030] Compared with the existing technology, the above technical solution can realize the screening of key features corresponding to mechanical properties in the expanded data set and obtain the overall key features, so as to provide clear and reasonable training sets and test sets for subsequent algorithm screening.
[0031] In a possible implementation, step S4 includes the following steps:
[0032] S41: constructing a plurality of mechanical property prediction models;
[0033] S42: using the final component features of each sub-dataset as input and the mechanical properties of the sub-dataset as the ideal output, dividing the training set and the test set into two sets in proportion;
[0034] S43: using the test set and the training set, adopting the ten-fold cross validation method, and taking the average determination coefficient as an evaluation index to screen the optimal prediction model from the multiple mechanical property prediction models to obtain the current performance prediction model;
[0035] Compared with the existing technology, the above technical solution can screen the optimal prediction model from multiple mechanical property prediction models, thereby obtaining the current performance prediction model.
[0036] In a possible implementation, in step S6, the process of verifying the error rate of the optimization model of the current performance prediction model through the zinc alloy performance prediction experiment includes the following steps:
[0037] S61: Using zinc alloy samples other than the total data set as experimental objects, testing their tensile strength and elongation using a universal mechanical testing machine, and testing their hardness using a hardness tester to obtain experimental results;
[0038] S62: Predicting the tensile strength, elongation, and hardness of the experimental object using the optimized model of the current performance prediction model to obtain a model prediction result;
[0039] S63: Compare the model prediction result with the experimental result to obtain the error rate of the optimized model of the current performance prediction model.
[0040] Compared with the existing technology, the above technical solution can obtain the error rate of the optimization model of the current performance prediction model, and decide whether to execute the next step based on the error rate, thereby ensuring accurate prediction of the alloy composition.
[0041] Another technical solution of the present invention is to provide a zinc alloy performance prediction and composition optimization system, based on the zinc alloy performance prediction and composition optimization method described in the present invention, comprising:
[0042] A data acquisition module collects data on the composition, forming process, and mechanical properties of the zinc alloy, including tensile strength, elongation, and hardness, and synthesizes the data into a total data set;
[0043] a data processing module, decomposing the total data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data, and performing data expansion using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain an expanded data set;
[0044] A feature module, which screens key features corresponding to mechanical properties from the expanded data set through the Pearson correlation coefficient and random forest feature importance evaluation in the expanded data set to obtain overall key features;
[0045] A model generation module, based on the overall key features, adopts a ten-fold cross validation method and uses the average determination coefficient as an evaluation index to screen the optimal prediction model from multiple mechanical property prediction models to obtain a current performance prediction model;
[0046] An optimization module, which uses grid search to optimize the hyperparameters of the current performance prediction model, and evaluates the regression model performance using the coefficient of determination and root mean square error to obtain an optimized model of the current performance prediction model;
[0047] A judgment module verifies the error rate of the optimization model of the current performance prediction model through a zinc alloy performance prediction experiment, and does not output if the error rate is greater than the predetermined value, and feeds the experimental results back to the data acquisition module to notify the data acquisition module to re-collect data; otherwise, output based on the error rate.
[0048] A prediction module performs multi-objective optimization on the overall key features by using a non-dominated sorting genetic algorithm to obtain the optimal composition ratio of the zinc alloy mechanical properties;
[0049] in,
[0050] The data processing module is electrically connected to the data acquisition module, the feature module is electrically connected to the data processing module, the model generation module is electrically connected to the feature module, the optimization module is electrically connected to the model generation module, and the judgment module is electrically connected to the optimization module, the prediction module, and the data acquisition module.
[0051] Compared with the existing technology, the zinc alloy performance prediction and composition optimization system of this application has the following advantages:
[0052] First, the present invention provides a zinc alloy performance prediction and composition optimization system. Through the orderly coordination of a data processing module, a feature module, a model generation module, an optimization module, a judgment module, and a prediction module, and by adopting a combination of a machine learning strategy and a non-dominated sorting genetic algorithm, the system can achieve multi-objective optimization and rapidly design zinc alloys with high comprehensive mechanical properties, providing design ideas and references for the development of new zinc alloys.
[0053] Secondly, the present invention provides a zinc alloy performance prediction and composition optimization system, which uses a data processing module to expand data based on the idea of mean value and standard deviation, providing a solution for the lack of sample data for some alloys when constructing machine learning prediction models, improving the generalization ability of the model and reducing the risk of overfitting.
[0054] Finally, the present invention provides a zinc alloy performance prediction and composition optimization system, which establishes a target performance prediction model based on a machine learning algorithm for alloy development. To a certain extent, it solves the problems of long cycle and high cost caused by the large composition design space of traditional design, and realizes accurate and effective performance prediction. It has important reference value and practical significance for the composition design of high-strength and toughness zinc alloys.
[0055] In one possible implementation, the data processing module includes:
[0056] a screening unit, removing all data whose forming process is a non-casting process from the total data set to obtain a casting process data set;
[0057] a decomposition unit, decomposing the casting process data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data, to obtain three data sets;
[0058] An expansion unit, which expands the data using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain multiple sub-data sets as the expanded data sets;
[0059] in,
[0060] The screening unit is electrically connected to the data acquisition module, and the expansion unit is electrically connected to the feature module;
[0061] Using screening units to remove all data whose forming processes are not casting processes from the total data set can ensure process consistency and avoid the impact of subsequent model quality on different processing processes; using disassembly units and expansion units, and a technical solution based on the idea of mean and standard deviation to expand data can provide a solution for the lack of sample data for some alloys when constructing machine learning prediction models, improve the generalization ability of the model, and reduce the risk of overfitting.
[0062] In one possible implementation, the feature module includes:
[0063] A Pearson correlation unit is configured to delete, from each sub-dataset of the expanded data set, component terms having a Pearson correlation coefficient with the mechanical property of the sub-dataset that is less than a first predetermined value, to obtain respective initial component features;
[0064] The random forest unit deletes the component items whose random forest feature importance evaluation value is less than a second predetermined value from the initial component features of each sub-dataset of the expanded data set through a random forest feature importance ranking method to obtain respective final component features;
[0065] A feature generation unit combines all the final component features to obtain the overall key feature, and uses all component items corresponding to the overall key feature as the final component features of each sub-dataset of the expanded data set;
[0066] in,
[0067] The Pearson correlation unit is electrically connected to the expansion unit, and the feature generation unit is electrically connected to the model generation module;
[0068] Compared with the existing technology, the above technical solution can realize the screening of key features corresponding to mechanical properties in the expanded data set and obtain the overall key features, so as to provide clear and reasonable training sets and test sets for subsequent algorithm screening.
[0069] In a possible implementation, the model generation module includes:
[0070] A storage unit, collecting and storing a plurality of the mechanical property prediction models;
[0071] a partitioning unit, taking the final component characteristics of each sub-dataset as input and the mechanical properties of the sub-dataset as ideal output, and partitioning the training set and the test set in proportion;
[0072] a model selection unit, using the test set and the training set, adopting the ten-fold cross validation method, and using the average determination coefficient as an evaluation index to screen the optimal prediction model from the plurality of mechanical property prediction models to obtain the current performance prediction model;
[0073] in,
[0074] The storage unit is electrically connected to the model selection unit, the division unit is electrically connected to the feature generation unit, the division unit is electrically connected to the model selection unit, and the model selection unit is electrically connected to the optimization module;
[0075] This enables the optimal prediction model to be screened from a plurality of mechanical property prediction models to obtain a current performance prediction model.
[0076] In a possible implementation, the judgment module includes:
[0077] an experimental data unit for obtaining the tensile strength, elongation, and hardness of zinc alloy samples other than the total data set to obtain experimental results;
[0078] a prediction unit, which predicts the tensile strength, elongation and hardness of the experimental object by using an optimized model of the current performance prediction model to obtain a model prediction result;
[0079] A comparison unit, which compares the model prediction result with the experimental result to obtain an error rate of the optimized model of the current performance prediction model;
[0080] a judgment unit, based on the error rate, not outputting when the error rate is greater than the predetermined value, and feeding back the experimental result to the data acquisition module to notify the data acquisition module to re-collect data; otherwise, outputting;
[0081] in,
[0082] The experimental data unit is electrically connected to the comparison unit, the prediction unit is electrically connected to the comparison unit and the optimization module, and the judgment unit is electrically connected to the comparison unit, the data acquisition module, and the prediction module.
[0083] Compared with the existing technology, the above technical solution can obtain the error rate of the optimization model of the current performance prediction model, and decide whether to output based on the error rate, thereby ensuring accurate prediction of the alloy composition. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 This is a flow chart of a zinc alloy performance prediction and composition optimization method disclosed in Example 1 of the present application;
[0085] Figure 2 This is a flow chart of step S2 in Example 1 of this application;
[0086] Figure 3 This is a schematic diagram of the data expansion method in Example 1 of the present application;
[0087] Figure 4 This is the flow chart of step S3 in Example 1 of this application;
[0088] Figure 5 This is the flow chart of step S4 in Example 1 of this application;
[0089] Figure 6 This is an evaluation diagram of the cross-validation of each algorithm in Example 1 of the present application;
[0090] Figure 7 Schematic diagram of the performance of each performance prediction model in Example 1 of the present application;
[0091] Figure 8 This is a flow chart of the error rate of the optimization model of the current performance prediction model verified by the zinc alloy performance prediction experiment in Example 1 of the present application;
[0092] Figure 9 This is a schematic diagram of the structure of a zinc alloy performance prediction and composition optimization system in Example 2 of the present application. DETAILED DESCRIPTION
[0093] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.
[0094] In the description of the embodiments of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "electrical connection" and "electrical connection relationship" should be understood in a broad sense, that is, referring to a connection method with an electrical relationship, for example, it can be a circuit connection through a wire, or it can be an electrical connection through a radio signal channel (channel), or a combination of the two. In addition, "electrical connection" and "electrical connection relationship" can be based on a mechanical connection (such as a wire set in a connecting key); it can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.
[0095] Two embodiments will be used below to further illustrate the present application in detail in conjunction with the accompanying drawings and specific embodiments.
[0096] See also Figure 1~Figure 8 As shown, the embodiment of the present application discloses a method for predicting the performance and optimizing the composition of zinc alloys, see Figure 1 As shown, the method includes the following steps:
[0097] S1: Collect data on the composition, forming process, and mechanical properties of zinc alloys based on public databases. The mechanical properties include tensile strength, elongation, and hardness. The data are then synthesized into a comprehensive dataset.
[0098] As a specific example, data related to zinc alloy composition, forming process and mechanical properties are collected as a total data set through literature and public material databases. The alloy composition values are expressed in mass percentage, and the mechanical properties include ultimate tensile strength (UTS), elongation (EL), and hardness. The processes are divided into casting, rolling, and extrusion. The data of some zinc alloy samples are shown in Table 1. In Table 1, the elements Al, Cu, Mg, Ca, Ti, Mn, Si, B, Re, La, Ce, and Zn in the first row represent all the elemental components contained in the collected zinc alloy data, UTS, EL, and Hardness represent tensile strength, elongation, and hardness, respectively. Process represents zinc alloys under different forming processes, and the unit of alloy element composition is mass percentage (wt%). For example, the Al content "11±0.5" in the second row of Table 1 indicates that the Al content of the zinc alloy is in the range of 10.5~11.5 (wt%), bal. represents the balance, the tensile strength unit is MPa, the elongation unit is (%), and the hardness unit is (HB). For example, the UTS value in the first row is "327.5±17.5", which indicates that the tensile strength of the zinc alloy is 310~345 MPa. The processes are divided into casting, rolling, and extrusion. For the alloy element composition in the table, the blank part indicates that the content is 0. For the three mechanical properties, the blank part indicates that the value was not published or tested at the time of collection.
[0099] Table 1: Total dataset table
[0100]
[0101]
[0102] S2: The data processing module decomposes the total data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of composition data and elongation data, and a composition-hardness data set consisting of composition data and hardness data. The data are then expanded using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain an expanded data set.
[0103] See also Figure 2 As shown, in this embodiment, step S2 includes the following steps:
[0104] S21: removing all data whose forming process is not a casting process from the total data set by a data processing module to obtain a casting process data set;
[0105] S22: Decomposing the casting process data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of composition data and elongation data, and a composition-hardness data set consisting of composition data and hardness data by a data processing module, to obtain three data sets;
[0106] S23: Data expansion is performed using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain multiple sub-data sets, which are used as expanded data sets.
[0107] Specifically, in the specific example of this embodiment, by executing step S21, all the casting process data in the total data are screened to ensure process consistency and avoid the subsequent impact of different processing processes on model quality. Then, by step S22, the total data set is divided into three sub-data sets of "Composition-Tensile Strength (UTS)", "Composition-Elongation (EL)", and "Composition-Hardness (Hardness)" as shown in Tables 2, 3, and 4.
[0108] Table 2: Composition-Ultrasonic Tensile Strength (UTS) Dataset
[0109]
[0110]
[0111] Table 3: Composition-Elongation (EL) Dataset
[0112] Al With Mg That IT min And B Re to What Zn HE / % 11±0.5 0.85±0.35 0.0225±0.0075 ball. 2±1 3.9±0.4 2.85±0.35 0.045±0.015 ball. 1 5.8±0.2 1.4±0.2 ball. 1.5 4.1 2.97 0.02 0.3 ball. 14.3 3.99 2.98 0.01 0.1 0.21 ball. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 ball. 10.5 4.12 0.049 0.05 0.02 ball. 8.5 … … … … … … … … … … … … …
[0113] Table 4: Composition-Hardness dataset
[0114] Al With Mg That IT min And B Re to What Zn Handness / HB 11±0.5 0.85±0.35 0.0225±0.0075 ball. 89 3.9±0.4 2.85±0.35 0.045±0.015 ball. 100 5.8±0.2 1.4±0.2 ball. 105 4.1 2.97 0.02 0.3 ball. 114 3.99 2.98 0.01 0.1 0.21 ball. 118 4.12 0.049 0.04 0.1 0.1 0.05 0.03 ball. 83 4.12 0.049 0.05 0.02 ball. 76 … … … … … … … … … … … … …
[0115] See also Figure 3 As shown, in step S23, the fluctuation range of each component content in these three sub-datasets, as well as the median and standard deviation of each mechanical property are used to expand the data. The specific method will be combined with Figure 3Taking the zinc alloy data of the first item in Table 1 or 2 as an example, the Al content of the alloy is in the range of 10.5-11.5 (wt%), the Cu content is in the range of 0.5-1.2 (wt%), the Mg content is in the range of 0.015-0.03 (wt%), the median tensile strength is 327.5 MPa, the standard deviation is 17.5 MPa, the median elongation is 2%, the standard deviation is 1%, and the median hardness is 89 HB. By taking the maximum, median and minimum values of the composition data and the performance data and combining them, we will obtain three component distribution ratios and three groups of target performance combinations. The three component distribution ratios are as follows: the maximum ratio Al content is 11.5wt%, Cu content is 1.2wt%, and Mg content is 0.03wt%; the median composition ratio Al content is 11wt%, Cu content is 0.85wt%, and Mg content is 0.0225wt%; the minimum composition ratio Al content is 10.5wt%, Cu content is 0.5wt%, and Mg content is 0.015wt%; the three groups of target performance are as follows: the maximum performance combination tensile strength is 345Mpa, elongation is 3%, and hardness is 89HB; the median performance combination tensile strength is 327.5Mpa, elongation is 2%, and hardness is 89HB; the minimum performance combination tensile strength is 310Mpa, elongation is 1%, and hardness is 89HB.
[0116] The three component ratios are further combined with the three groups of target properties, that is, the maximum component ratio and the three groups of target properties are combined into three sample data, the median component ratio and the three groups of target properties are combined into three sample data, and the minimum component ratio and the three groups of properties are combined into three sample data. So far, 9 sample data have been obtained from the first zinc alloy data. The missing values of the components in each sub-dataset are supplemented with 0, and the element Zn content is obtained by subtracting the sum of the contents of other elements from the total content 100wt%.
[0117] Therefore, in the specific example of this embodiment, Table 2 is expanded into sub-datasets 1 to 9, which are listed as follows:
[0118] Subdataset 1
[0119] Al With Mg That IT min And B Re to What Zn STU / Mpa 11.5 1.2 0.03 87.27 345 4.3 3.2 0.06 92.44 240 6 1.6 92.4 315 4.1 2.97 0.02 0.3 92.61 287 3.99 2.98 0.01 0.1 0.21 92.71 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0120] Subdataset 2
[0121] Al With Mg That IT min And B Re to What Zn STU / Mpa 11 0.85 0.0225 88.1275 345 3.9 2.85 0.045 93.205 240 5.8 1.4 92.8 315 4.1 2.97 0.02 0.3 92.61. 287 3.99 2.98 0.01 0.1 0.21 92.71. 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0122] Subdataset 3
[0123] Al With Mg That IT min And B Re to What Zn STU / Mpa 10.5 0.5 0.015 88.985 345 3.5 2.5 0.03 93.97 240 5.6 1.2 93.2 315 4.1 2.97 0.02 0.3 92.61. 287 3.99 2.98 0.01 0.1 0.21 92.71. 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0124] Subdataset 4
[0125] Al With Mg That IT min And B Re to What Zn STU / Mpa 11.5 1.2 0.03 87.27 327.5 4.3 3.2 0.06 92.44 240 6 1.6 92.4 315 4.1 2.97 0.02 0.3 92.61 287 3.99 2.98 0.01 0.1 0.21 92.71 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0126] Subdataset 5
[0127] Al With Mg That IT min And B Re to What Zn STU / Mpa 11 0.85 0.0225 88.1275 327.5 3.9 2.85 0.045 93.205 240 5.8 1.4 92.8 315 4.1 2.97 0.02 0.3 92.61. 287 3.99 2.98 0.01 0.1 0.21 92.71. 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0128] Subdataset 6
[0129]
[0130]
[0131] Subdataset 7
[0132] Al With Mg That IT min And B Re to What Zn STU / Mpa 11.5 1.2 0.03 87.27 310 4.3 3.2 0.06 92.44 240 6 1.6 92.4 315 4.1 2.97 0.02 0.3 92.61 287 3.99 2.98 0.01 0.1 0.21 92.71 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0133] Subdataset 8
[0134] Al With Mg That IT min And B Re to What Zn STU / Mpa 11 0.85 0.0225 88.1275 310 3.9 2.85 0.045 93.205 240 5.8 1.4 92.8 315 4.1 2.97 0.02 0.3 92.61. 287 3.99 2.98 0.01 0.1 0.21 92.71. 295 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 287 4.12 0.049 0.05 0.02 95.761 234 … … … … … … … … … … … … …
[0135] Subdataset 9
[0136]
[0137]
[0138] Table 3 is expanded into sub-datasets 10 to 18, which are listed as follows: Sub-dataset 10
[0139] Al With Mg That IT min And B Re to What Zn HE / % 11.5 1.2 0.03 87.27 3 4.3 3.2 0.06 92.44 1 6 1.6 92.4 1.5 4.1 2.97 0.02 0.3 92.61 14.3 3.99 2.98 0.01 0.1 0.21 92.71 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0140] Subdataset 11
[0141] Al With Mg That IT min And B Re to What Zn HE / % 11.5 1.2 0.03 87.27 2 4.3 3.2 0.06 92.44 1 6 1.6 92.4 1.5 4.1 2.97 0.02 0.3 92.61 14.3 3.99 2.98 0.01 0.1 0.21 92.71 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0142] Subdataset 12
[0143]
[0144]
[0145] Subdataset 13
[0146] Al With Mg That IT min And B Re to What Zn HE / % 11 0.85 0.0225 88.1275 3 3.9 2.85 0.045 93.205 1 5.8 1.4 92.8 1.5 4.1 2.97 0.02 0.3 92.61. 14.3 3.99 2.98 0.01 0.1 0.21 92.71. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0147] Subdataset 14
[0148] Al With Mg That IT min And B Re to What Zn HE / % 11 0.85 0.0225 88.1275 2 3.9 2.85 0.045 93.205 1 5.8 1.4 92.8 1.5 4.1 2.97 0.02 0.3 92.61. 14.3 3.99 2.98 0.01 0.1 0.21 92.71. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0149] Subdataset 15
[0150] Al With Mg That IT min And B Re to What Zn HE / % 11 0.85 0.0225 88.1275 1 3.9 2.85 0.045 93.205 1 5.8 1.4 92.8 1.5 4.1 2.97 0.02 0.3 92.61. 14.3 3.99 2.98 0.01 0.1 0.21 92.71. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0151] Subdataset 16
[0152] Al With Mg That IT min And B Re to What Zn HE / % 10.5 0.5 0.015 88.985 3 3.5 2.5 0.03 93.97 1 5.6 1.2 93.2 1.5 4.1 2.97 0.02 0.3 92.61. 14.3 3.99 2.98 0.01 0.1 0.21 92.71. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0153] Subdataset 17
[0154] Al With Mg That IT min And B Re to What Zn HE / % 10.5 0.5 0.015 88.985 2 3.5 2.5 0.03 93.97 1 5.6 1.2 93.2 1.5 4.1 2.97 0.02 0.3 92.61. 14.3 3.99 2.98 0.01 0.1 0.21 92.71. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0155] Subdataset 18
[0156] Al With Mg That IT min And B Re to What Zn HE / % 10.5 0.5 0.015 88.985 1 3.5 2.5 0.03 93.97 1 5.6 1.2 93.2 1.5 4.1 2.97 0.02 0.3 92.61. 14.3 3.99 2.98 0.01 0.1 0.21 92.71. 15.3 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 10.5 4.12 0.049 0.05 0.02 95.761 8.5 … … … … … … … … … … … … …
[0157] Table 4 is expanded into sub-datasets 19 to 21, which are listed as follows: Sub-dataset 19
[0158] Al With Mg That IT min And B Re to What Zn Handness / HB 11.5 1.2 0.03 87.27 89 4.3 3.2 0.06 92.44 100 6 1.6 92.4 105 4.1 2.97 0.02 0.3 92.61 114 3.99 2.98 0.01 0.1 0.21 92.71 118 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 83 4.12 0.049 0.05 0.02 95.761 76 … … … … … … … … … … … … …
[0159] Subdataset 20
[0160] Al With Mg That IT min And B Re to What Zn Handness / HB 11 0.85 0.0225 88.1275 89 3.9 2.85 0.045 93.205 100 5.8 1.4 92.8 105 4.1 2.97 0.02 0.3 92.61. 114 3.99 2.98 0.01 0.1 0.21 92.71. 118 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 83 4.12 0.049 0.05 0.02 95.761 76 … … … … … … … … … … … … …
[0161] Subdataset 21
[0162] Al With Mg That IT min And B Re to What Zn Handness / HB 10.5 0.5 0.015 88.985 89 3.5 2.5 0.03 93.97 100 5.6 1.2 93.2 105 4.1 2.97 0.02 0.3 92.61. 114 3.99 2.98 0.01 0.1 0.21 92.71. 118 4.12 0.049 0.04 0.1 0.1 0.05 0.03 95.511 83 4.12 0.049 0.05 0.02 95.761 76 … … … … … … … … … … … … …
[0163] S3: Screening key features corresponding to mechanical properties from the expanded data set through the Pearson correlation coefficient and random forest feature importance evaluation in the expanded data set to obtain overall key features;
[0164] See also Figure 4 As shown, in this embodiment, step S3 includes the following steps:
[0165] S31: deleting, from each sub-dataset of the expanded data set, the component items whose Pearson correlation coefficient with the mechanical properties of the sub-dataset is less than a first predetermined value, to obtain respective initial component characteristics;
[0166] S32: Deleting the component items whose random forest feature importance evaluation values are less than a second predetermined value from the initial component features of each sub-dataset of the expanded data set by using a random forest feature importance ranking method, to obtain respective final component features;
[0167] S33: All the final component features obtained in step S32 are combined to obtain the overall key feature, and all component items corresponding to the overall key feature are used as the final component features of each sub-dataset of the expanded data set.
[0168] Specifically, in the specific example of this embodiment, features are preliminarily screened based on the Pearson correlation coefficient value. Each sub-dataset deletes component features with a correlation coefficient less than 0.1 with the target performance (tensile strength, elongation, hardness), and the features screened by each sub-dataset are combined to preliminarily obtain the overall features, namely the initial component features, which include the component elements Al, Cu, Mg, Ca, Ti, Mn, Si, B, Re, La, Ce, and Zn. Further, the initial component features are evaluated for importance according to the random forest feature importance ranking, and features with importance values close to 0 are deleted. The results of the screening of each sub-dataset are further combined to obtain the overall key features, which include the component elements Al, Cu, Mg, Ti, Mn, Si, La, Ce, and Zn, and then the final component features are obtained, which simplify the model and improve the generalization ability.
[0169] S4: Based on the overall key features, the ten-fold cross-validation method was used, and the average determination coefficient was used as the evaluation index to screen the optimal prediction model from multiple mechanical property prediction models to obtain the current performance prediction model;
[0170] See also Figure 5 As shown, in this embodiment, step S4 includes the following steps:
[0171] S41: Construct multiple mechanical property prediction models;
[0172] S42: Using the final component characteristics of each sub-dataset as input and the mechanical properties of this sub-dataset as the ideal output, the training and test sets are divided proportionally;
[0173] S43: Using the test set and training set, the ten-fold cross-validation method is adopted, and the average determination coefficient is used as the evaluation index to screen the optimal prediction model from multiple mechanical property prediction models to obtain the current performance prediction model.
[0174] See also Figure 6 As shown, in the specific example of this embodiment, the machine learning model algorithm is implemented in Python language, the overall features obtained by screening in step S33 are used as input, and the data set is divided into a training set and a test set at an 8:2 ratio. Three integrated algorithms, random forest (RF), extreme gradient boosting (XGBoost), and AdaBoost, are initially selected to construct the tensile strength, elongation, and hardness prediction model of zinc alloy. The ten-fold cross validation method is used and the average determination coefficient R 2 As an evaluation indicator to screen the appropriate algorithm model, according to Figure 6 It can be seen that the XGBoost algorithm performs best in the construction of various mechanical property prediction models.
[0175] S5: Use grid search to optimize the hyperparameters of the current performance prediction model, and evaluate the regression model performance using the coefficient of determination and root mean square error to obtain the optimized model of the current performance prediction model;
[0176] See also Figure 7 As shown, in the specific example of this embodiment, grid search is used to optimize the hyperparameters of the prediction model constructed by the XGBoost algorithm, and the coefficient of determination and root mean square error RMSE are used to evaluate the performance of the regression model. Figure 7 The figure shows the performance of the final performance prediction model. (a), (b) and (c) in the figure correspond to the mechanical properties respectively, among which R 2 _train, R 2 _test, RMSE_train, and RMSE_test represent the prediction accuracy of the training set, the prediction accuracy of the test set, the root mean square error of the training set, and the root mean square error of the test set of each model, respectively. It can be seen that the prediction accuracy of each model set has reached above 0.9.
[0177] S6: verifying the error rate of the optimization model of the current performance prediction model through zinc alloy performance prediction experiments, and executing the next step when the error rate is not greater than a predetermined value; when the error rate is greater than the predetermined value, collecting the experimental results into the total data set and returning to step S1;
[0178] See also Figure 8 As shown, in step S6 of this embodiment, the process of verifying the error rate of the optimization model of the current performance prediction model through the zinc alloy performance prediction experiment includes the following steps:
[0179] S61: Use zinc alloy samples other than the total data set as experimental objects, test their tensile strength and elongation using a universal mechanical testing machine, and test their hardness using a hardness tester to obtain experimental results;
[0180] S62: predicting the tensile strength, elongation, and hardness of the experimental object by using the optimized model of the current performance prediction model to obtain a model prediction result;
[0181] S63: Compare the model prediction results with the experimental results to obtain the error rate of the optimized model of the current performance prediction model.
[0182] Specifically, in a specific example of this embodiment, experiments were conducted on zinc alloy samples that were not included in the total data set. The tensile strength and elongation were tested using a universal mechanical testing machine, and the hardness was tested using a hardness tester. The results were compared with the model prediction values to further verify the reliability of the model. As shown in the following table, the composition ratios of two zinc alloys that were not included in the total data set, as well as the machine learning prediction performance and experimental values, are shown. The comparison between the predicted values and the experimental values shows that the errors are within an acceptable range, and the composition ratio can be optimized to obtain a zinc alloy with higher comprehensive performance.
[0183]
[0184]
[0185] S7: performing multi-objective optimization on the overall key features by a non-dominated sorting genetic algorithm to obtain the optimal composition ratio of the zinc alloy with respect to mechanical properties;
[0186] Specifically, in the specific example of this embodiment, the search space for algorithm optimization is set as Al: 3-8wt%, Cu: 0-5wt%, Mg: 0-1wt%, Ti; 0-0.5wt%, Mn: 0-1wt%, Si: 0-1wt%, La: 0-0.5wt%, Ce: 0-0.5wt%, and Zn participates in the prediction in the form of remainder. The NSGA-II algorithm is used for iterative optimization, and several component distribution ratios with better comprehensive prediction performance are selected from the Pareto optimal solution set, providing design ideas and references for the development of new zinc alloys.
[0187] The present embodiment discloses a zinc alloy performance prediction and composition optimization method, which establishes a target performance prediction model based on a machine learning algorithm for alloy development. This method solves the problems of long cycles and high costs caused by the large composition design space in traditional designs, and achieves accurate and effective performance prediction. It has important reference value and practical significance for the composition design of high-strength and toughness zinc alloys.
[0188] Example 2
[0189] This embodiment further discloses a zinc alloy performance prediction and composition optimization system, see Figure 9 As shown, the system includes a data acquisition module, a data processing module, a feature module, a model generation module, an optimization module, a judgment module, and a prediction module. The data processing module is electrically connected to the data acquisition module, the feature module is electrically connected to the data processing module, the model generation module is electrically connected to the feature module, the optimization module is electrically connected to the model generation module, and the judgment module is electrically connected to the optimization module, the prediction module, and the data acquisition module. The functions corresponding to each module are as follows:
[0190] The data acquisition module is configured to collect data on the composition, forming process, and mechanical properties of the zinc alloy, the mechanical properties including tensile strength, elongation, and hardness, and synthesize a total data set from the data;
[0191] The data processing module is configured to decompose the total data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of composition data and elongation data, and a composition-hardness data set consisting of composition data and hardness data, and perform data expansion using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain an expanded data set;
[0192] The feature module is configured to screen key features corresponding to mechanical properties from the expanded data set by using the Pearson correlation coefficient and random forest feature importance evaluation in the expanded data set to obtain overall key features;
[0193] The model generation module is set to be based on the overall key features, using the ten-fold cross-validation method and the average determination coefficient as the evaluation index to screen the optimal prediction model from multiple mechanical property prediction models to obtain the current performance prediction model;
[0194] The optimization module is configured to use grid search to optimize the hyperparameters of the current performance prediction model, and evaluate the regression model performance using the coefficient of determination and root mean square error to obtain the optimized model of the current performance prediction model;
[0195] The judgment module is configured to verify the error rate of the optimization model of the current performance prediction model through a zinc alloy performance prediction experiment, and if the error rate is greater than a predetermined value, no output is given, and the experimental results are fed back to the data acquisition module to notify the data acquisition module to re-collect data; otherwise, the output is based on the error rate.
[0196] The prediction module is set to perform multi-objective optimization of the overall key features through a non-dominated sorting genetic algorithm to obtain the optimal composition ratio of the mechanical properties of the zinc alloy.
[0197] Please continue to see Figure 9 In this embodiment, the data processing module includes a screening unit, a disassembly unit and an expansion unit that establish an electrical connection relationship. The screening unit is configured to remove all data whose molding process is not a casting process from the total data set to obtain a casting process data set; the disassembly unit is configured to disassemble the casting process data set into a component-tensile strength data set consisting of component data and tensile strength data, a component-elongation data set consisting of component data and elongation data, and a component-hardness data set consisting of component data and hardness data to obtain three data sets; the expansion unit is configured to use the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to expand the data to obtain multiple sub-data sets, which are used as expanded data sets; wherein the screening unit is electrically connected to the data acquisition module, and the expansion unit is electrically connected to the feature module.
[0198] Please continue to see Figure 9In this embodiment, the feature module includes a Pearson correlation unit, a random forest unit and a feature generation unit that establish an electrical connection relationship. The Pearson correlation unit is configured to delete the component items whose Pearson correlation coefficients with the mechanical properties of the sub-dataset are less than a first predetermined value from each sub-dataset of the expanded data set, and obtain their respective initial component features; the random forest unit is configured to delete the component items whose random forest feature importance evaluation values are less than a second predetermined value from the initial component features of each sub-dataset of the expanded data set through a random forest feature importance ranking method, and obtain their respective final component features; the feature generation unit is configured to merge all the final component features to obtain the overall key features, and use all the component items corresponding to the overall key features as the final component features of each sub-dataset of the expanded data set; wherein the Pearson correlation unit is electrically connected to the expansion unit, and the feature generation unit is electrically connected to the model generation module.
[0199] Please continue to see Figure 9 In this embodiment, the model generation module includes a storage unit, a partitioning unit and a model selection unit. The storage unit is configured to collect and store multiple mechanical property prediction models; the partitioning unit is configured to take the final component characteristics of each sub-data set as input, the mechanical properties of the sub-data set as the ideal output, and divide the training set and the test set in proportion; the model selection unit is configured to use the test set and the training set, adopt a ten-fold cross-validation method, and use the average determination coefficient as an evaluation index to screen the optimal prediction model from multiple mechanical property prediction models to obtain the current performance prediction model; wherein the storage unit is electrically connected to the model selection unit, the partitioning unit is electrically connected to the feature generation unit, the partitioning unit is electrically connected to the model selection unit, and the model selection unit is electrically connected to the optimization module.
[0200] Please continue to see Figure 9 In this embodiment, the judgment module includes an experimental data unit, a prediction unit, a comparison unit and a judgment unit. The experimental data unit is configured to obtain the tensile strength, elongation and hardness of the zinc alloy samples outside the total data set to obtain experimental results; the prediction unit is configured to predict the tensile strength, elongation and hardness of the experimental object through the optimization model of the current performance prediction model to obtain the model prediction results; the comparison unit is configured to compare the model prediction results with the experimental results to obtain the error rate of the optimization model of the current performance prediction model; the judgment unit is configured to be based on the error rate, and when the error rate is greater than a predetermined value, no output is given, and the experimental results are fed back to the data acquisition module to notify the data acquisition module to re-collect data; otherwise, based on the output; wherein, the experimental data unit is electrically connected to the comparison unit, the prediction unit is electrically connected to the comparison unit and the optimization module, and the judgment unit is electrically connected to the comparison unit, the data acquisition module and the prediction module.
[0201] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.
[0202] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0203] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for predicting zinc alloy properties and optimizing composition, characterized in that: The steps include: S1: Collecting data on the composition, forming process, and mechanical properties of zinc alloys based on public databases. The mechanical properties include tensile strength, elongation, and hardness, and synthesizing a total data set from the data. S2: decomposing the total data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data by a data processing module, and performing data expansion using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain an expanded data set; S3: Screening key features corresponding to mechanical properties from the expanded data set by using the Pearson correlation coefficient and random forest feature importance evaluation in the expanded data set to obtain overall key features; S4: Based on the overall key features, a ten-fold cross validation method is used to screen the optimal prediction model from multiple mechanical property prediction models using the average determination coefficient as an evaluation index to obtain the current performance prediction model; S5: Optimizing hyperparameters of the current performance prediction model using grid search, and evaluating the regression model performance using the coefficient of determination and root mean square error to obtain an optimized model of the current performance prediction model; S6: verifying the error rate of the optimization model of the current performance prediction model through a zinc alloy performance prediction experiment, and executing the next step when the error rate is not greater than a predetermined value; and when the error rate is greater than the predetermined value, collecting the experimental results into a total data set and returning to executing step S1; S7: Perform multi-objective optimization on the overall key features through a non-dominated sorting genetic algorithm to obtain the optimal composition ratio of the mechanical properties of the zinc alloy.
2. The zinc alloy performance prediction and composition optimization method according to claim 1, characterized in that: The step S2 comprises the following steps: S21: removing all data whose forming process is not a casting process from the total data set by the data processing module to obtain a casting process data set; S22: Decomposing the casting process data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data by the data processing module, to obtain three data sets; S23: Expand the data using the fluctuation range of the content of each component and the median and standard deviation of each mechanical property in each data set to obtain multiple sub-data sets, which are used as the expanded data sets.
3. The zinc alloy performance prediction and composition optimization method according to claim 2, characterized in that: The step S3 comprises the following steps: S31: deleting, from each sub-dataset of the expanded data set, the component items whose Pearson correlation coefficient with the mechanical property of the sub-dataset is less than a first predetermined value, to obtain respective initial component features; S32: Deleting the component items whose random forest feature importance evaluation values are less than a second predetermined value from the initial component features of each sub-dataset of the expanded data set by using a random forest feature importance ranking method, to obtain respective final component features; S33: All the final component features obtained in step S32 are combined to obtain the overall key feature, and all component items corresponding to the overall key feature are used as the final component features of each sub-dataset of the expanded data set.
4. The zinc alloy performance prediction and composition optimization method according to claim 3, characterized in that: The step S4 comprises the following steps: S41: constructing a plurality of mechanical property prediction models; S42: using the final component features of each sub-dataset as input and the mechanical properties of the sub-dataset as the ideal output, dividing the training set and the test set into two sets in proportion; S43: Using the test set and the training set, adopting the ten-fold cross validation method, and taking the average determination coefficient as the evaluation index, the optimal prediction model is screened from the multiple mechanical property prediction models to obtain the current performance prediction model.
5. The zinc alloy performance prediction and composition optimization method according to claim 4, characterized in that: In step S6, the process of verifying the error rate of the optimization model of the current performance prediction model through the zinc alloy performance prediction experiment includes the following steps: S61: Using zinc alloy samples other than the total data set as experimental objects, testing their tensile strength and elongation using a universal mechanical testing machine, and testing their hardness using a hardness tester to obtain experimental results; S62: Predicting the tensile strength, elongation, and hardness of the experimental object using the optimized model of the current performance prediction model to obtain a model prediction result; S63: Compare the model prediction result with the experimental result to obtain the error rate of the optimized model of the current performance prediction model.
6. A zinc alloy performance prediction and composition optimization system, characterized in that: The method for predicting zinc alloy properties and optimizing composition according to any one of claims 1 to 5 comprises: A data acquisition module collects data on the composition, forming process, and mechanical properties of the zinc alloy, including tensile strength, elongation, and hardness, and synthesizes the data into a total data set; a data processing module, decomposing the total data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data, and performing data expansion using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain an expanded data set; A feature module, which screens key features corresponding to mechanical properties from the expanded data set through the Pearson correlation coefficient and random forest feature importance evaluation in the expanded data set to obtain overall key features; A model generation module, based on the overall key features, adopts a ten-fold cross validation method and uses the average determination coefficient as an evaluation index to screen the optimal prediction model from multiple mechanical property prediction models to obtain a current performance prediction model; An optimization module, which uses grid search to optimize the hyperparameters of the current performance prediction model, and evaluates the regression model performance using the coefficient of determination and root mean square error to obtain an optimized model of the current performance prediction model; A judgment module verifies the error rate of the optimization model of the current performance prediction model through a zinc alloy performance prediction experiment, and does not output if the error rate is greater than the predetermined value, and feeds the experimental results back to the data acquisition module to notify the data acquisition module to re-collect data; otherwise, output based on the error rate. A prediction module performs multi-objective optimization on the overall key features by using a non-dominated sorting genetic algorithm to obtain the optimal composition ratio of the zinc alloy mechanical properties; in, The data processing module is electrically connected to the data acquisition module, the feature module is electrically connected to the data processing module, the model generation module is electrically connected to the feature module, the optimization module is electrically connected to the model generation module, and the judgment module is electrically connected to the optimization module, the prediction module, and the data acquisition module.
7. The zinc alloy performance prediction and composition optimization system according to claim 6, characterized in that: The data processing module includes: a screening unit, removing all data whose forming process is a non-casting process from the total data set to obtain a casting process data set; a decomposition unit, decomposing the casting process data set into a composition-tensile strength data set consisting of composition data and tensile strength data, a composition-elongation data set consisting of the composition data and elongation data, and a composition-hardness data set consisting of the composition data and hardness data, to obtain three data sets; An expansion unit, which expands the data using the fluctuation range of each component content and the median and standard deviation of each mechanical property in each data set to obtain multiple sub-data sets as the expanded data sets; in, The screening unit is electrically connected to the data acquisition module, and the expansion unit is electrically connected to the feature module.
8. The zinc alloy performance prediction and composition optimization system according to claim 7, characterized in that: The feature module includes: A Pearson correlation unit is configured to delete, from each sub-dataset of the expanded data set, component terms having a Pearson correlation coefficient with the mechanical property of the sub-dataset that is less than a first predetermined value, to obtain respective initial component features; The random forest unit deletes the component items whose random forest feature importance evaluation value is less than a second predetermined value from the initial component features of each sub-dataset of the expanded data set through a random forest feature importance ranking method to obtain respective final component features; A feature generation unit combines all the final component features to obtain the overall key feature, and uses all component items corresponding to the overall key feature as the final component features of each sub-dataset of the expanded data set; in, The Pearson correlation unit is electrically connected to the expansion unit, and the feature generation unit is electrically connected to the model generation module.
9. The zinc alloy performance prediction and composition optimization system according to claim 8, characterized in that: The model generation module includes: A storage unit, collecting and storing a plurality of the mechanical property prediction models; a partitioning unit, taking the final component characteristics of each sub-dataset as input and the mechanical properties of the sub-dataset as ideal output, and partitioning the training set and the test set in proportion; a model selection unit, using the test set and the training set, adopting the ten-fold cross validation method, and using the average determination coefficient as an evaluation index to screen the optimal prediction model from the plurality of mechanical property prediction models to obtain the current performance prediction model; in, The storage unit is electrically connected to the model selection unit, the division unit is electrically connected to the feature generation unit, the division unit is electrically connected to the model selection unit, and the model selection unit is electrically connected to the optimization module.
10. The zinc alloy performance prediction and composition optimization system according to claim 9, characterized in that: The judgment module includes: an experimental data unit for obtaining the tensile strength, elongation, and hardness of zinc alloy samples other than the total data set to obtain experimental results; A prediction unit, which predicts the tensile strength, elongation and hardness of the experimental object by using the optimization model of the current performance prediction model to obtain a model prediction result; A comparison unit, which compares the model prediction result with the experimental result to obtain an error rate of the optimized model of the current performance prediction model; a judgment unit, based on the error rate, not outputting when the error rate is greater than the predetermined value, and feeding back the experimental result to the data acquisition module to notify the data acquisition module to re-collect data; otherwise, outputting; in, The experimental data unit is electrically connected to the comparison unit, the prediction unit is electrically connected to the comparison unit and the optimization module, and the judgment unit is electrically connected to the comparison unit, the data acquisition module, and the prediction module.
Citation Information
Patent Citations
Design method for simultaneous optimization of conductivity and hardness of multicomponent electrical contact alloys
CN114580272B
Metal material temperature and strain rate related plastic hardening model calculation method
CN117558381A
Method for predicting compressive strength of magnesium phosphate cement-based composite material based on machine learning
CN117577241A
Method for screening high-hardness and high-entropy alloy through particle swarm optimization algorithm assisted machine learning
CN116312890A
Method and system for determining technological parameters of selective laser melting technology
CN116502455A