Method for designing multi-component alloy components based on component extrapolation of symbol regression
By screening key element features and constructing relational equations based on symbol regression, the prediction problem of the impact of new elements import in alloy design is solved, and efficient prediction and optimization of multi-component alloy performance is achieved.
Patent Information
- Application Number
- CN202510641269.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-19
AI Technical Summary
The existing technology cannot efficiently judge the impact of new element introduction on alloy performance in alloy design. The traditional method has limitations and cannot effectively conduct component extrapolation prediction.
By establishing a low-component alloy dataset, screening key element features, constructing relational equations using symbol regression methods based on multiple group evolution algorithms, combining Latin hypercube sampling for exploration and verification of multicomponent alloy components, the symbol regression model is optimized to achieve high-performance multicomponent alloy design.
Accurate prediction and optimization of the performance of multi-component alloys is achieved, and the accuracy of alloy performance prediction after the introduction of new elements is improved, and it is suitable for a variety of alloy systems.
Smart Images

Figure CN120510962A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of material calculation, and in particular relates to a method for designing the composition of a multi-component alloy by component extrapolation based on symbolic regression. Background Art
[0002] In recent years, machine learning (ML) methods have become increasingly popular in the field of materials design. However, in the field of alloy design, machine learning models are generally limited to accurate interpolation predictions. Exploring the high-dimensional space of new alloys still relies on trial-and-error methods, making it difficult to effectively determine the impact of new elements on alloy properties.
[0003] For predictions beyond the scope of components, solution geometry model extrapolation, transfer learning, or the combination of physical-based theoretical models (such as the CALPHAD method) and data-driven models are currently possible approaches. Solution geometry model extrapolation can extrapolate most melt properties, including viscosity, surface tension, and activity, based on low-component melt data. In 1989, Zhou Guozhi and YAChang published a comprehensive article on "numerical models" in the journal CALPHAD, systematically exploring the principles of numerical models. In 1995, Zhou Guozhi further proposed a new generation of solution geometry models, also known as the General Solution Model (GSM), which can describe the thermodynamic properties of various complex multi-component solutions. Transfer learning is typically applied to extrapolate different properties from similar datasets. For example, Zhu et al. proposed a transfer learning architecture based on a TCN-GRU to transfer knowledge from quasi-dynamic data to dynamic conditions, enabling dynamic monitoring of fuel cell health. Kim et al. used active transfer learning and data augmentation to gradually expand the reliable prediction domain of DNNs for forward design of composite materials. Hybrid modeling, combining physics-based theoretical models with data-driven models, is more suitable for property extrapolation in multi-component situations in materials and chemistry. Sansana et al. compared hybrid modeling approaches with physics-based and data-driven models and found that hybrid modeling consistently outperformed baseline models in both extrapolation and transfer learning tasks. Wang et al. collected 8,920 data points from a materials project and generated 4,262 features for predicting the shear modulus and bulk modulus of solid electrolytes. In the case of nickel element extrapolation, they significantly improved the poor prediction performance by adding a small number of samples to the training set. While the three component extrapolation methods are widely used in their respective fields, each has its limitations. Solution geometry models can only calculate solution properties and are unable to calculate macroscopic properties influenced by phases. Furthermore, the accuracy of geometric models is highly dependent on low-level metadata, which determines the correlation coefficient for extrapolating high-level metadata properties, requiring a large amount of accurate data for fitting. Transfer learning methods can effectively predict the same properties from similar datasets. A common approach is to use a broader range of features for modeling and fine-tune a model that predicts a narrow range of features. However, the introduction of component extrapolation requires transfer learning to readjust parameters to find patterns. In this case, it is generally necessary to transfer through models that predict different properties. However, the relationship between different properties and microscopic characteristics in the materials field varies greatly, so there are limitations to prediction. The most widely used method is a hybrid modeling method that combines theoretical models based on physical laws with data-driven models because it integrates domain knowledge unique to materials science, making extrapolation more reasonable. However, the current common practice is to collect all relevant features through calculations and experiments, and then construct an interpolation prediction model after layer-by-layer screening, which has not yet developed into extrapolation prediction.Sansana et al. applied a hybrid modeling approach based on physical laws and data-driven methods to extrapolation. This approach simply used a data-driven model to fit the residuals of first-principles calculations, aiming to improve the accuracy of the theoretical model. While this approach improved the accuracy of the extrapolation, it remained limited by traditional computational methods, preventing efficient calculations or extension to other systems and properties.
[0004] Therefore, it is necessary to design a method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression to solve or alleviate one or more of the above problems. Summary of the Invention
[0005] In view of this, the present invention provides a method for designing the composition of multi-component alloys by component extrapolation based on symbolic regression. By constructing a relationship equation based on the key element characteristics calculated based on the low-component alloy composition, and predicting the performance of the multi-component alloy, the symbolic regression model is continuously iterated and optimized to achieve high-performance multi-component alloy design.
[0006] The object of the present invention is to provide a method for designing multi-component alloy compositions based on component extrapolation of symbolic regression, characterized in that the steps include: S1. Establish an initial low-component element data set. According to the selected low-component alloy system, establish an initial low-component element data set including alloy composition, material element characteristics and alloy properties; S2. Screening meta-features, gradually screening features through Spearman correlation coefficient screening, feature importance analysis and recursive feature elimination, and selecting key material meta-feature combinations that affect the performance of low-component alloys; S3. Establishing a relationship equation, using a symbolic regression method based on a multi-population evolutionary algorithm to construct a relationship equation between the key element characteristic data screened in step S2 and the alloy properties; S4. Introduce new elements. Based on domain knowledge, introduce a new element (Cn) into the low-component alloy (C1, C2, C3...) to determine the multi-component alloy system. S5. Multi-component alloy sampling and evaluation: Latin hypercube sampling is used to uniformly sample the multi-component alloy composition space and select sampling points for experimental testing; S6. Evaluate the relationship equation, calculate the key element characteristics of the multi-component alloy sampling point, use the relationship equation obtained in step S3 to predict the performance of the multi-component alloy sampling point, and evaluate the validity of the relationship equation; S7. Optimize the symbolic regression model based on the evaluation results and experimental data, and iterate steps S3-S6 to improve the alloy design, ultimately producing a multi-component alloy with high performance.
[0007] Furthermore, the initial low-component dataset in step S1 includes three parts: S1-1: The alloy composition comprises Al, Fe, Cr, Ni, Mo, Ti and other main elements and other trace elements, which are combined in equimolar or nearly equimolar proportions (5%-49% atomic percentage), and the elements are of at least three types; S1-2: The material meta-characteristics include molecular orbital occupancy, atomic packing efficiency, band center, cation properties, atomic cohesive energy, atomic affinity, electronegativity difference, compound formation enthalpy, oxidation state, valence electron orbital, mixing heat and size mismatch term, element bond neighbor ratio, crystal structure information, Coulomb matrix, adjacent site electrostatic interaction, symmetry information, Voronoi polyhedron information, Brillouin band information, Fermi level information, local site chemical fingerprint, and at least two of the short-range order characteristics; S1-3: The alloy properties include at least one of alloy hardness, tensile strength and elongation data.
[0008] Furthermore, the meta-features in step S1 are directly associated with microscopic atomic scale information, and both low-component alloy data and multi-component alloy data can realize the calculation of meta-features through composition information.
[0009] Furthermore, the specific steps of the meta-feature screening in step S2 include: S2-1: Spearman correlation coefficient screening, by calculating the Spearman coefficient between any two material meta-features. If the correlation coefficient is greater than 0.95, one feature is eliminated from the two and the more important feature is retained, and finally n meta-features are obtained; S2-2: Feature importance analysis: take the n meta-features remaining after the Spearman correlation coefficient screening as input and the alloy properties as output, establish an RF machine learning model, calculate the importance of each feature, and select the top m material meta-features; S2-3: Recursive elimination, with the m material meta-features obtained from the feature importance analysis as input and the alloy properties as output, removing the corresponding extracted features when the model error is minimum, leaving m-1 features, and then repeating the above feature elimination until the minimum error changes from decreasing to stable, and stopping, and selecting the meta-feature finally retained as the final key meta-feature.
[0010] Furthermore, the multi-population evolutionary algorithm of step S3 can aim to automatically discover mathematical expressions from the data (X, y) , making Able to fit as accurately as possible , construct candidate expressions within the allowed parameter range and operator combination, optimize the objective function as formula (1), and continuously cross-mutate to reduce the loss function until it stabilizes and stops, to obtain the improved expression: (1); Where y is the real data label; f(X) is the model prediction value; It is the L3 norm loss function, which measures the gap between the model and the true value based on the cube of the error, reducing the impact of large outliers on the overall model; Indicates the model complexity, indicating the depth of expression and the number of operators; is the regularization coefficient, which is used to control complexity and avoid overfitting.
[0011] Furthermore, the evolutionary algorithm parameters of step S3 include a regularization coefficient, a population size, a maximum number of iterations, a crossover probability, a mutation probability, a weight of an L3 norm loss function, an expression depth, and a model complexity; the parameter ranges are 10-6~103, 5~1000, 100~10000, 0.6~0.9, 0.01~0.2, 0.1~10, 2~10, and 2~50, respectively; and the operator combinations include ("+", "-", "*", " / ", "^", "cos", "exp", "sin", "cube", "log10", "sqrt", and "1 / x").
[0012] Furthermore, the new element in step S4 needs to utilize an element substitution strategy, that is, selecting elements with similar atomic size, electronegativity and electronic configuration in the same crystal position or similar chemical environment to reduce lattice distortion and maintain structural stability.
[0013] Furthermore, the Latin hypercube sampling in step S5 sets the range of the new element content according to the original low-component alloy composition, determines the overall composition space, evenly distributes the sampling points in the composition space, selects 3 to 10 sampling points for experimental verification, and the sampling process is as shown in formula (2), and ensures that the sum of the components of each group of samples is 1: (2); Where, represents the mole fraction of element i in the jth group of alloy samples; represents the candidate value of element i in the jth group of samples, which is randomly selected by evenly dividing the component space based on Latin hypercube sampling; The random order of element i among N sampling points; d represents the total number of elements in the alloy.
[0014] Furthermore, the relationship equation evaluation in step S6 uses the determination coefficient R2 to evaluate the degree of fit of the relationship equation to the multi-component alloy data set. The R2 calculation process is as shown in formula (3). When the R2 between the predicted value and the experimental value of the sampling point is greater than 0.8, the symbolic regression model is retained: (3); Where, is the coefficient of determination, and its value range is [0,1]. The closer it is to 1, the higher the prediction accuracy. is the i-th true value; is the i-th predicted value; is the mean of the true value; n is the number of samples.
[0015] The beneficial effects of the above technical solution are: the present invention provides a method for designing multi-component alloy compositions by component extrapolation based on symbolic regression. This method combines Spearman correlation coefficient screening, feature importance analysis, and recursive feature elimination for element feature screening, constructs a relationship equation based on a symbolic regression method of a multi-population evolutionary algorithm, and uses Latin hypercube sampling to explore and verify the multi-component alloy composition space, continuously optimizing the relationship equation, and realizing modeling based on low-component alloy compositions to predict multi-component alloy properties. Therefore, it can be said that the method proposed by the present invention can be implemented to meet the engineering application needs of extrapolating the performance of multi-component alloys from low-component alloys in the materials field. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of the method for designing multi-component alloy compositions by component extrapolation based on symbolic regression provided by the present invention.
[0017] Figure 2 Schematic diagram of the key element characteristics of "composition-Vickers hardness" of the AlCoCrFeNi alloy provided in Example 1 of the present invention.
[0018] Figure 3 Schematic diagram comparing the predicted results and experimental results of the Vickers hardness of AlCoCrFeNi alloy and AlCoCrCuFeNi alloy using the symbolic regression equation provided in Example 1 of the present invention.
[0019] Figure 4 This is a schematic diagram comparing the predicted results and experimental results of the Vickers hardness of AlCoCrFeNi alloy and AlCoCrCuFeNi alloy using the extreme gradient boosting model XGBR provided in Comparative Example 1 of the present invention.
[0020] Figure 5 Schematic diagram comparing the predicted results of the total elongation of low-activated ferrite-martensitic steel and low-activated ferrite-martensitic steel (with Zr addition) using the symbolic regression equation provided in Example 2 of the present invention and the experimental results. DETAILED DESCRIPTION
[0021] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected. Example
[0022] The alloy system selected in this example is Al-Co-Cr-Fe-Ni high entropy alloy; like Figure 1 As shown in the figure, a method for designing multi-component alloy composition based on component extrapolation based on symbolic regression is presented. The specific process is as follows: Figure 1 As shown, the following steps are included: S1. Establish an initial low-component element data set. Based on the selected Al-Co-Cr-Fe-Ni alloy system, establish an initial low-component element data set including alloy composition, material element characteristics, and alloy properties. S2. Screening meta-features, gradually screening features through Spearman correlation coefficient screening, feature importance analysis and recursive feature elimination, and selecting key material meta-feature combinations that affect the performance of low-component alloys; S3. Establishing a relationship equation, using a symbolic regression method based on a multi-population evolutionary algorithm to construct a relationship equation between the key element characteristic data screened in step S2 and the alloy properties; S4. Introduce new elements. Based on domain knowledge, introduce a new element Cu into the Al-Co-Cr-Fe-Ni low-component alloy to determine the Al-Co-Cr-Cu-Fe-Ni multi-component alloy system. S5. Multi-component alloy sampling and evaluation: Latin hypercube sampling is used to uniformly sample the multi-component alloy composition space and select sampling points for experimental testing; S6. Evaluate the relationship equation, calculate the key element characteristics of the multi-component alloy sampling point, use the relationship equation obtained in step S3 to predict the performance of the multi-component alloy sampling point, and evaluate the validity of the relationship equation; S7. Optimize the symbolic regression model based on the evaluation results and experimental data, and iterate steps S3-S6 to improve the alloy design, ultimately producing a multi-component alloy with high performance.
[0023] The initial low-component dataset in step S1 above includes three parts: S1-1: The alloy composition includes Al, Co, Cr, Fe and Ni elements, which are combined in equimolar or nearly equimolar proportions (5%-49% atomic percentage), with a total of 98 data points; S1-2: The material meta-characteristics include molecular orbital occupancy, atomic packing efficiency, band center, cation properties, atomic cohesive energy, atomic affinity, electronegativity difference, compound formation enthalpy, oxidation state, valence electron orbital, mixing heat and size mismatch term, element bond neighbor ratio, crystal structure information, Coulomb matrix, adjacent site electrostatic interaction, symmetry information, Voronoi polyhedron information, Brillouin band information, Fermi level information, local site chemical fingerprint, and short-range order characteristics; S1-3: The alloy property is the alloy Vickers hardness.
[0024] The meta-features in the above step S1 are directly related to the microscopic atomic scale information. Both low-component alloy data and multi-component alloy data can realize the calculation of meta-features through composition information.
[0025] The specific steps of the meta-feature screening in step S2 above include: S2-1: Spearman correlation coefficient screening: By calculating the Spearman coefficient between any two material meta-features, if the correlation coefficient is greater than 0.95, one feature is eliminated from the two and the more important feature is retained, ultimately obtaining 28 meta-features; S2-2: Feature importance analysis: The 28 meta-features remaining after the Spearman correlation coefficient screening are used as input, and the alloy properties are used as output. The RF machine learning model is established to calculate the importance of each feature and select the top 15 material meta-features; S2-3: Recursive elimination, with the 15 material features analyzed by feature importance as input and the alloy properties as output, remove the features corresponding to the minimum model error, and leave m-1 features. Then repeat the above feature elimination until the minimum error changes from decreasing to stable, and select the 12 features retained as the final key features. The detailed key features are as follows: Figure 2 shown.
[0026] The key element features of the above step S2 specifically include: average number of unfilled electrons, minimum difference in solid solution formation enthalpy, Young's ω parameter, Young's Δ parameter, radius γ parameter, λ entropy, electronegativity difference, average cohesive energy, local mismatch of shear modulus, average heat capacity, average value of simulated absolute stacking efficiency, and average sound velocity of simulated material.
[0027] The multi-population evolutionary algorithm in step S3 can automatically discover mathematical expressions from data (X, y) , making Able to fit as accurately as possible , construct candidate expressions within the allowed parameter range and operator combination, optimize the objective function as formula (1), and continuously cross-mutate to reduce the loss function until it stabilizes and stops, to obtain the improved expression: (1); Where y is the real data label; f(X) is the model prediction value; It is the L3 norm loss function, which measures the gap between the model and the true value based on the cube of the error, reducing the impact of large outliers on the overall model; Indicates the model complexity, indicating the depth of expression and the number of operators; is the regularization coefficient, which is used to control complexity and avoid overfitting.
[0028] The evolutionary algorithm parameters in step S3 above include regularization coefficient, population size, maximum number of iterations, crossover probability, mutation probability, weight of L3 norm loss function, expression depth and model complexity; the parameter ranges are 10 -1 ~10 2 , 5~400, 100~6000, 0.6~0.9, 0.01~0.2, 0.1~10, 2~10, 2~50; operator combinations include ("+", "-", "*", " / ", "^", "cos", "exp", "sin", "cube", "log10", "sqrt" and "1 / x").
[0029] The new element in step S4 needs to utilize an element substitution strategy, that is, selecting elements with similar atomic size, electronegativity and electronic configuration in the same crystal position or similar chemical environment to reduce lattice distortion and maintain structural stability.
[0030] The Latin hypercube sampling in step S5 above sets the range of the new element content based on the original low-component alloy composition, determines the overall composition space, evenly distributes sampling points in the composition space, and selects 10 sampling points for experimental verification. The sampling process is as shown in formula (2), and ensures that the sum of the components of each group of samples is 1: (2); Where, represents the mole fraction of element i in the jth group of alloy samples; represents the candidate value of element i in the jth group of samples, which is randomly selected by evenly dividing the component space based on Latin hypercube sampling; The random order of element i among N sampling points; d represents the total number of elements in the alloy.
[0031] The relationship equation evaluation in step S6 above uses the determination coefficient R2 to evaluate the degree of fit of the relationship equation to the multi-component alloy data set.2 The calculation process is as shown in formula (3), the R between the predicted value and the experimental value at the sampling point is 2 Greater than 0.8, retain the sign regression model: (3); Where, is the coefficient of determination, and its value range is [0,1]. The closer it is to 1, the higher the prediction accuracy. is the i-th true value; is the i-th predicted value; is the mean of the true value; n is the number of samples.
[0032] like Figure 3 The comparison results of the predicted values and experimental values of all sampling points in the iterative process are shown. The symbolic regression relationship equation constructed based on the key element features generated and screened for the Al-Co-Cr-Fe-Ni low-component alloy predicts the Vickers hardness of the Al-Co-Cr-Fe-Ni low-component alloy. 2 The determination coefficient R for the prediction of Vickers hardness of Al-Co-Cr-Cu-Fe-Ni multi-component alloy is 0.9086. 2 It is 0.8829.
[0033] Therefore, according to Example 1, the following conclusion can be drawn: the method of designing the composition of multi-component alloys by component extrapolation based on symbolic regression, and the symbolic regression relationship equation constructed only by the key element characteristics generated and screened by the Al-Co-Cr-Fe-Ni low-component alloy, can effectively predict the Vickers hardness of the Al-Co-Cr-Cu-Fe-Ni multi-component alloy.
[0034] In order to demonstrate the effect of symbolic regression on the accuracy of extrapolation of new elements in multi-component alloys, comparative example 1 is provided, and the extreme gradient boosting regression XGBR model is used for modeling evaluation. Comparative Example 1
[0035] A design method for multi-component alloy composition extrapolation based on extreme gradient boosting regression and low-component data sets, the difference being that: in step S3, a symbolic regression method based on a multi-population evolutionary algorithm is not used to construct a relationship equation, but an extreme gradient boosting regression XGBR model is used to construct a "meta-feature-Vickers hardness" prediction model for Al-Co-Cr-Fe-Ni alloy and Al-Co-Cr-Cu-Fe-Ni alloy.
[0036] The comparison between the predicted value and the true value of the prediction model built based on extreme gradient boosting regression XGBR is as follows: Figure 4 The results show that the XGBR model has a high determination coefficient R for predicting the Vickers hardness of Al-Co-Cr-Fe-Ni alloy. 2The determination coefficient R for the prediction of Vickers hardness of Al-Co-Cr-Cu-Fe-Ni alloy is 0.9744. 2 The predicted determination coefficient R of the Vickers hardness of the Al-Co-Cr-Fe-Ni low-component alloy obtained by the symbolic regression method in Example 1 is 0.6748. 2 The determination coefficient R for the prediction of Vickers hardness of Al-Co-Cr-Cu-Fe-Ni alloy is 0.9086. 2 The result shows that XGBR is overfitting for low-component data, and the symbolic regression method is helpful to improve the Vickers hardness prediction of Al-Co-Cr-Fe-Ni alloy after adding Cu element.
[0037] To further demonstrate that the method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression provided by the present invention is also applicable to other alloy systems, Example 2 is provided to add new elements to low-activation ferrite-martensitic steel and predict its total elongation. Example 2
[0038] A method for designing multi-component alloy compositions by component extrapolation based on symbolic regression. Steps not specifically specified are the same as those in Example 1, except that the low-component alloy system in step S1 is a low-activation ferrite-martensitic steel, wherein the main elements include Cr, W, Mn, V, and Ta, and the trace elements include C, Si, N, Y, and Ti. The low-component data set comprises 232 data items, and the newly added element is Zr. The predicted alloy property is total elongation.
[0039] like Figure 5 The comparison results of the predicted and experimental values of all sampling points in the iterative process are shown. The symbolic regression equation constructed based on the key element features generated and screened for low-component low-activation ferrite-martensite steel predicts the total elongation of low-component low-activation ferrite-martensite steel with a determination coefficient R. 2 The determination coefficient R for the total elongation prediction of low activation ferrite-martensite steel after adding Zr is 0.9143. 2 The result proves that the method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression provided by the present invention can be used in a variety of alloy systems and has a wide range of applications.
[0040] As summarized, the present invention provides a method for designing multi-component alloy compositions by component extrapolation based on symbolic regression. This method combines Spearman correlation coefficient screening, feature importance analysis, and recursive feature elimination for feature screening. A symbolic regression method based on a multi-population evolutionary algorithm constructs a relationship equation. Latin hypercube sampling is used to explore and validate the multi-component alloy composition space, continuously optimizing the relationship equation. This enables modeling based on low-component alloy compositions and predicting multi-component alloy properties. Therefore, the method proposed in this invention can be applied to meet the engineering application needs of extrapolating the properties of multi-component alloys from low-component alloys in the materials field.
[0041] In this embodiment, specific examples are used to illustrate the principles and implementation plans of the present invention. The description of the above embodiments is only used to help understand the methods and core ideas of the present invention. Any changes, modifications, substitutions, combinations or simplifications made according to the spirit and principles of the technical solutions of the present invention should be equivalent replacement methods. As long as they comply with the purpose of the invention and do not deviate from the technical principles and inventive concepts of the present invention, they belong to the scope of protection of the present invention.
Claims
1. A method for designing multi-component alloy compositions based on component extrapolation of symbolic regression, characterized in that the steps include: S1. Establish an initial low-component element data set. According to the selected low-component alloy system, establish an initial low-component element data set including alloy composition, material element characteristics and alloy properties; S2. Screening meta-features, gradually screening features through Spearman correlation coefficient screening, feature importance analysis and recursive feature elimination, and selecting key material meta-feature combinations that affect the performance of low-component alloys; S3. Establishing a relationship equation, using a symbolic regression method based on a multi-population evolutionary algorithm to construct a relationship equation between the key element characteristic data screened in step S2 and the alloy properties; S4. Introduce new elements. Based on domain knowledge, introduce a new element (Cn) into the low-component alloy (C1, C2, C3...) to determine the multi-component alloy system. S5. Multi-component alloy sampling and evaluation: Latin hypercube sampling is used to uniformly sample the multi-component alloy composition space and select sampling points for experimental testing; S6. Evaluate the relationship equation, calculate the key element characteristics of the multi-component alloy sampling point, use the relationship equation obtained in step S3 to predict the performance of the multi-component alloy sampling point, and evaluate the validity of the relationship equation; S7. Optimize the symbolic regression model based on the evaluation results and experimental data, and iterate steps S3-S6 to improve the alloy design, ultimately producing a multi-component alloy with high performance.
2. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The data set in step S1 specifically includes the following: S1-1: The alloy composition comprises Al, Fe, Cr, Ni, Mo, Ti and other main elements and other trace elements, which are combined in equimolar or nearly equimolar proportions (5%-49% atomic percentage), and the elements are of at least three types; S1-2: The material meta-characteristics include molecular orbital occupancy, atomic packing efficiency, band center, cation properties, atomic cohesive energy, atomic affinity, electronegativity difference, compound formation enthalpy, oxidation state, valence electron orbital, mixing heat and size mismatch term, element bond neighbor ratio, crystal structure information, Coulomb matrix, adjacent site electrostatic interaction, symmetry information, Voronoi polyhedron information, Brillouin band information, Fermi level information, local site chemical fingerprint, and at least two of the short-range order characteristics; S1-3: The alloy properties include at least one of alloy hardness, tensile strength and elongation data.
3. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The meta-features in step S1 are directly related to microscopic atomic scale information, and both low-component alloy data and multi-component alloy data can realize the calculation of meta-features through composition information.
4. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The meta-feature screening process in step S2 specifically includes the following steps: S2-1: Spearman correlation coefficient screening, by calculating the Spearman coefficient between any two material meta-features. If the correlation coefficient is greater than 0.95, one feature is eliminated from the two and the more important feature is retained, and finally n meta-features are obtained; S2-2: Feature importance analysis: take the n meta-features remaining after the Spearman correlation coefficient screening as input and the alloy properties as output, establish an RF machine learning model, calculate the importance of each feature, and select the top m material meta-features; S2-3: Recursive elimination, with the m material meta-features obtained from the feature importance analysis as input and the alloy properties as output, removing the corresponding extracted features when the model error is minimum, leaving m-1 features, and then repeating the above feature elimination until the minimum error changes from decreasing to stable, and stopping, and selecting the meta-feature finally retained as the final key meta-feature.
5. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The multi-population evolutionary algorithm of step S3 can aim to automatically discover mathematical expressions from data (X, y) , making Able to fit as accurately as possible , construct candidate expressions within the allowed parameter range and operator combination, optimize the objective function as formula (1), and continuously cross-mutate to reduce the loss function until it stabilizes and stops, to obtain the improved expression: (1); Where y is the real data label; f(X) is the model prediction value; It is the L3 norm loss function, which measures the gap between the model and the true value based on the cube of the error, reducing the impact of large outliers on the overall model; Indicates the model complexity, indicating the depth of expression and the number of operators; is the regularization coefficient, which is used to control complexity and avoid overfitting.
6. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The evolutionary algorithm parameters of step S3 include regularization coefficient, population size, maximum number of iterations, crossover probability, mutation probability, weight of L3 norm loss function, expression depth and model complexity; the parameter ranges are 10-6~103, 5~1000, 100~10000, 0.6~0.9, 0.01~0.2, 0.1~10, 2~10, and 2~50 respectively; operator combinations include ("+", "-", "*", " / ", "^", "cos", "exp", "sin", "cube", "log10", "sqrt" and "1 / x").
7. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The new element in step S4 needs to utilize an element substitution strategy, that is, selecting elements with similar atomic size, electronegativity and electronic configuration in the same crystal position or similar chemical environment to reduce lattice distortion and maintain structural stability.
8. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The Latin hypercube sampling in step S5 sets the range of the new element content according to the original low-component alloy composition, determines the overall composition space, evenly distributes sampling points in the composition space, selects 3 to 10 sampling points for experimental verification, and the sampling process is as shown in formula (2), and ensures that the sum of the components of each group of samples is 1: (2); Where, represents the mole fraction of element i in the jth group of alloy samples; represents the candidate value of element i in the jth group of samples, which is randomly selected by evenly dividing the component space based on Latin hypercube sampling; The random order of element i among N sampling points; d represents the total number of elements in the alloy.
9. The method for designing multi-component alloy compositions based on component extrapolation based on symbolic regression according to claim 1, characterized in that: The relationship equation evaluation in S6 uses the determination coefficient R2 to evaluate the degree of fit of the relationship equation to the multi-component alloy data set. The R2 calculation process is as shown in formula (3). When the R2 between the predicted value and the experimental value of the sampling point is greater than 0.8, the symbolic regression model is retained: (3); Where, is the coefficient of determination, and its value range is [0,1]. The closer it is to 1, the higher the prediction accuracy. is the i-th true value; is the i-th predicted value; is the mean of the true value; n is the number of samples.
Citation Information
Cited By
Prediction method for hardenability of medium-low carbon alloy steel
CN121637256A