Method for clustering and screening adaptive machine learning-assisted design of high-reliability lead-free tin-based solder alloy

Through cluster screening and adaptive machine learning assisted in the design of lead-free tin-based solder alloys, the problems of long R&D cycle and high cost in traditional methods are solved, and fast and low-cost alloy performance optimization is achieved.

CN115700574BActive Publication Date: 2025-07-11SHANGHAI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210826948.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-07-11
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and at low cost to develop lead-free tin-based solder alloys with excellent performance, and the traditional methods have long cycles and high costs.

Method used

Adaptive machine learning assisted design method after clustering screening is adopted, models are established through k-means clustering, feature screening and multiple machine learning algorithms, and alloy components are screened and designed based on expert domain knowledge.

Benefits of technology

It shortens the alloy R&D cycle, reduces costs, improves the comprehensive performance of lead-free tin-based solder alloys, and achieves rapid conversion from R&D to application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115700574B_ABST
    Figure CN115700574B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for assisting in the design of lead-free tin-based solder alloys through adaptive machine learning after clustering and screening. First, material data of lead-free tin-based solder alloys are collected and obtained to establish a data set. Then, the k-means clustering method is used to cluster the mechanical properties, the clusters with poor performance are removed, and the samples are classified. The alloy compositions of different categories and the atomic features after their characteristics are screened are used as inputs, and their mechanical properties are used as outputs. A single-objective machine learning model is established. For each machine learning model, the leave-one-out cross-validation method and the Pearson index R are used as the accuracy indicators of the machine learning model. For each different mechanical property, the machine learning model with the largest Pearson index R is selected. For the collected lead-free alloy composition data, internal interpolation and orthogonal permutation and combination are performed to serve as virtual samples. Finally, the virtual samples are input into the machine learning model to obtain the predicted values of the mechanical properties, and the alloy compositions with excellent performance are preferably selected according to the predicted values to achieve the auxiliary design of the alloy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of material design, and particularly relates to a method for adaptively machine learning-assisted design of lead-free tin-based solder alloys. Background Art

[0002] Pb-Sn alloys are widely used for connecting two metal surfaces in the electronics industry due to their low melting point, high strength, good electrical conductivity, and good wettability to most substrate materials commonly used in engineering. So far, no solder alloy can completely replace Pb-Sn. However, lead pollutes the environment and endangers human health. With the promulgation and implementation of the EU "WEEE Directive" and "RoHS Directive", as well as relevant regulations and management measures in countries such as the United States, Japan, and China, Pb-Sn alloys have been prohibited from use in most countries. Developing lead-free solders with excellent performance is also a hot issue in the current research in the fields of electronic packaging and micro-connection.

[0003] Based on the concept of material R & D based on the materials genome, combined with machine learning, developing new lead-free solders, establishing a machine learning model applicable to the development of tin-based solders, realizing accurate prediction of the performance of tin-based solders with different compositions, and developing lead-free tin-based solders with excellent comprehensive performance are of great significance to the development of the current electronics industry. Traditional material R & D means are usually the "trial and error method", which has a long cycle and high cost. By combining technologies such as machine learning and big data, and combining expert domain knowledge, the speed from R & D, manufacturing to application of new materials can be accelerated, and the R & D cost can be reduced. Through this method of adaptively machine learning-assisted design of lead-free tin-based solder alloys after clustering and screening, an efficient path can be provided for developing lead-free tin-based solder alloys with excellent performance and other new materials. Summary of the Invention

[0004] The purpose of the present invention is to accelerate the development and design of lead-free tin-based solder alloys by using machine learning methods to reduce R & D costs. To achieve the above purpose, the present invention provides a method for adaptively machine learning-assisted design of lead-free tin-based solder alloys after clustering and screening, which can develop lead-free tin-based solder alloys with excellent performance.

[0005] A method for adaptively machine learning-assisted design of lead-free tin-based solder alloys after clustering and screening includes the following steps:

[0006] Step S1: Collect and obtain material data of lead-free tin-based solder alloys to establish a basic data set;

[0007] Step S2: Use the k-means clustering method to cluster the mechanical properties. Take the product of different mechanical property data as the comprehensive target performance data, and use the comprehensive performance parameter of the mean value of each cluster as the mechanical property index to evaluate the cluster. Select the clusters with excellent comprehensive performance and eliminate the clusters with poor comprehensive performance. Then use the k-means clustering method to cluster the alloy compositions and divide the samples into at least two categories;

[0008] Step S3: Use the alloy compositions to construct atomic features, perform high-correlation filtering and feature screening on the atomic features. Use the alloy compositions and the screened atomic features of different category samples as input data, and their mechanical property data as output data;

[0009] Step S4: Use 12 different algorithms to establish single-objective machine learning models respectively. For each machine learning model, use the leave-one-out cross-validation method and the Pearson index R as the accuracy index of the machine learning model. For each different mechanical property, select the machine learning model with the largest Pearson index R;

[0010] Step S5: Perform interpolation and orthogonal permutation combination on the collected lead-free tin-based alloy composition data to obtain virtual samples. Input the virtual samples into the machine learning model to obtain mechanical property prediction values, and screen the alloy compositions according to the prediction values, so as to realize the design of alloy parameters.

[0011] Preferably, a method for self-adaptive machine learning-assisted design of lead-free tin-based solder alloys after clustering and screening includes the following steps:

[0012] Step S1: Collect and obtain lead-free tin-based solder alloy material data;

[0013] Collect the mass ratios of alloy compositions of lead-free tin-based solders;

[0014] Collect the mechanical properties of lead-free tin-based solders, and the mechanical properties include at least tensile strength and elongation at break;

[0015] Step S2: Use the k-means clustering method to cluster the mechanical properties. Take the product of several mechanical properties as the comprehensive target performance, and use the comprehensive performance of the mean value of each cluster as the mechanical property index to evaluate the cluster. Select the clusters with excellent comprehensive performance and eliminate the clusters with poor comprehensive performance. Then use the k-means clustering method to cluster the alloy compositions and divide the samples into several categories;

[0016] Use the k-means clustering method to cluster the mechanical properties of tensile strength and elongation at break, and eliminate a cluster of data with low tensile strength and low elongation at break, and retain the remaining data;

[0017] Use the k-means clustering method to cluster the alloy compositions, and divide the samples into several categories according to the clustering results;

[0018] Preferably, after removing a cluster of data with low tensile strength and low elongation at break, the remaining several categories of data are used to train a machine learning model, and the model accuracy is greatly improved compared with the model trained without removing data.

[0019] Step S3: Convert the alloy composition ratio into an atomic ratio as a supplementary feature, and add the atomic radius, valence electron number, and Pauling electronegativity as supplementary features. Perform high-correlation filtering and shapvalue feature importance ranking screening on the supplementary features.

[0020] Step S4: Use the alloy compositions of different categories and the screened atomic features as inputs, and their mechanical properties as outputs. Use 12 different algorithms to establish single-objective machine learning models respectively. For each machine learning model, use the leave-one-out cross-validation method and the Pearson index R as the machine learning model accuracy index. For each different mechanical property, select the machine learning model with the largest Pearson index R.

[0021] Use 12 different algorithms to establish single-objective machine learning models respectively. The 12 models include: Linear, Ridge, LASSO, MLP, Decision Tree, Random Forest, Xgboost, Adaboost, GBDT, Bagging, SVM, KNN.

[0022] Each machine learning model uses the leave-one-out cross-validation method.

[0023] The Pearson index R is used as the machine learning model accuracy index. The mathematical expression of R is:

[0024]

[0025] For each different mechanical property, select the machine learning model with the largest Pearson index R. Among them, n represents the number of samples, X i represents the predicted value, Y i represents the measured value, represents the mean of the predicted values, represents the mean of the measured values.

[0026] Step S5: For the collected lead-free tin-based alloy composition data, combine expert domain knowledge to perform interpolation and orthogonal permutation and combination to generate virtual samples, and input the virtual samples into the machine learning model to obtain the mechanical property prediction values of different composition combinations.

[0027] Based on the collected lead-free tin-based alloy composition data, combine expert domain knowledge to perform interpolation from the maximum value to zero for each alloy composition, set the step size, and perform permutation and combination as virtual samples.

[0028] Input the virtual samples into a machine learning model to obtain predicted values of mechanical properties. Use the product of several mechanical properties as a comprehensive index to measure strength and toughness. Select samples with relatively large comprehensive performance from the Pareto boundary of tensile strength and elongation at break for experimental verification, so as to achieve the auxiliary design of alloys.

[0029] Preferably, the composition range of the alloy suitable for design by the method of the present invention is in weight percentage composition: 3.0 - 5.5% of Ag, 0.5 - 1.0% of Cu, 1.0 - 5.0% of Bi, 0.01 - 1.0% of Ti, 0.01 - 1.0% of Ni, and the balance is Sn. The present invention proposes a composition range of a high-reliability lead-free solder alloy through the above method.

[0030] A system for auxiliary design of lead-free tin-based solder alloys, including an operation analysis module, a storage module, and an input / output module, is characterized in that: a program of the method for auxiliary design of lead-free tin-based solder alloys by clustering screening and adaptive machine learning described in the present invention is executed by using an operation analysis device.

[0031] Compared with the prior art, the present invention has the following obvious outstanding substantial features and remarkable advantages:

[0032] The present invention utilizes existing experimental data for machine learning prediction, and guides the design of material composition according to the prediction results, thereby shortening the experimental period and reducing the R & D cost, and having the effect of accelerating the speed of lead-free solder alloys from R & D, manufacturing to application. Brief Description of the Drawings

[0033] Figure 1 is the general flowchart of the present invention.

[0034] Figure 2 is the clustering diagram of the mechanical properties of tensile strength and elongation at break of the present invention.

[0035] Figure 3 is the schematic diagram of high-correlation filtering of the present invention.

[0036] Figure 4 is the feature importance ranking diagram of the present invention.

[0037] Figure 5 is the schematic diagram of the leave-one-out cross-validation method of the present invention.

[0038] Figure 6 is the flowchart of adaptive machine learning of the present invention.

[0039] Figure 7 is the experimental verification diagram of the present invention. Detailed Description of the Invention

[0040] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings:

[0041] A method for adaptively machine learning-assisted design of lead-free tin-based solder alloys after clustering and screening, comprising the following steps:

[0042] Step S1: Collect and obtain material data of lead-free tin-based solder alloys, and establish a basic data set;

[0043] Step S2: Use the k-means clustering method to cluster the mechanical properties, use the product of different mechanical property data as the comprehensive target performance data, use the comprehensive performance parameters of the mean value of each cluster as the mechanical property index to evaluate the cluster, select the clusters with excellent comprehensive performance, and eliminate the clusters with poor comprehensive performance; then use the k-means clustering method to cluster the alloy components and divide the samples into at least two categories;

[0044] Step S3: Use the alloy components to construct atomic features, perform high-correlation filtering and feature screening on the atomic features, use the alloy components and the screened atomic features of different category samples as input data, and use their mechanical property data as output data;

[0045] Step S4: Use 12 different algorithms to establish single-objective machine learning models respectively. For each machine learning model, use the leave-one-out cross-validation method and the Pearson index R as the accuracy index of the machine learning model. For each different mechanical property, select the machine learning model with the largest Pearson index R;

[0046] Step S5: Perform interpolation and orthogonal permutation combination on the collected lead-free tin-based alloy component data as virtual samples, input the virtual samples into the machine learning model, obtain the predicted values of mechanical properties, and screen the alloy components according to the predicted values, so as to realize the design of alloy parameters.

[0047] In this embodiment, 90 pieces of material data of lead-free tin-based solder alloys are collected and obtained. It includes the mass ratio of alloy components of lead-free tin-based solder, and the alloy components are tin, silver, copper, bismuth, indium, antimony, nickel, zinc, titanium, and aluminum. The mechanical properties of lead-free tin-based solder include tensile strength and elongation at break.

[0048] Use the k-means clustering method to cluster the mechanical properties of tensile strength and elongation at break, as Figure 2 shown, divide them into 3 clusters according to the clustering results, eliminate the first cluster data with low tensile strength and low elongation at break, that is, the lowest comprehensive performance, and retain the remaining data. Input the remaining data into the machine learning model, and there is a significant improvement in accuracy compared with inputting the data without elimination into the machine learning model. The R of tensile strength is increased from 0.822 to 0.881, and the R of elongation at break is increased from 0.581 to 0.749.

[0049] The k-means clustering method is used to cluster the alloy compositions. The clustering results are two clusters with silver content greater than 3% and less than or equal to 3%. According to the clustering results, the samples are divided into two categories: high-silver and low-silver.

[0050] The alloy composition ratio is converted into atomic ratio as supplementary feature, and atomic radius, number of valence electrons, and Pauling electronegativity are added as supplementary features. High-correlation filtering ( Figure 3 ) and shapvalue feature importance ranking and screening ( Figure 4 ) are performed on the supplementary features.

[0051] Taking the atomic features after alloy composition screening as input and a mechanical property as output, single-objective machine learning models are established using 12 different algorithms respectively. These 12 machine learning algorithms include: Linear, Ridge, LASSO, MLP, Decision Tree, Random Forest, Xgboost, Adaboost, GBDT, Bagging, SVM, KNN. Each machine learning model adopts the leave-one-out cross-validation method. As Figure 5 shown, the Pearson index R is used as the accuracy index of the machine learning model. For each different mechanical property, the machine learning model with the largest Pearson index R is selected. The machine learning model for high-silver tensile strength is Xgboost, the machine learning model for high-silver elongation at break is SVM, the machine learning model for low-silver tensile strength is KNN, and the machine learning model for low-silver elongation at break is LASSO.

[0052] Based on the collected lead-free tin-based alloy composition data and combined with expert domain knowledge, the alloy element range is determined. From the maximum value to zero of each alloy composition, an internal difference is set with a step size of 0.1%, and 945536328 combinations are obtained through permutation and combination as virtual samples.

[0053] Input the virtual samples into the machine learning model to obtain the predicted values of mechanical properties. We use the product of tensile strength and elongation at break as a comprehensive index to measure strength and toughness, and use this comprehensive index to screen out the solder alloy composition combinations with excellent performance to assist alloy composition design. Through the above method, we have determined the composition range of lead-free solder alloys with excellent mechanical properties, and their weight percentage composition is: 3.0 - 5.5% Ag, 0.5 - 1.0% Cu, 1.0 - 5.0% Bi, 0.01 - 1.0% Ti, 0.01 - 1.0% Ni, and the balance is Sn. Two samples with the best comprehensive performance are selected through the comprehensive index for experiments, and the designed two alloys and their properties are shown in Table 1. The comprehensive performance of the two lead-free tin-based solder alloys designed with the assistance of machine learning is higher than that of the alloys in the training set, and it has increased by 33% compared with the average comprehensive performance of the training set, as Figure 7 shown.

[0054] Table 1 Two example lead-free tin-based solder alloys designed by adaptive machine learning assisted by clustering screening

[0055]

[0056] The method for designing lead-free tin-based solder alloys by adaptive machine learning assisted by clustering screening in the above embodiments of the present invention first collects and obtains the material data of lead-free tin-based solder alloys to establish a data set; then, uses the k-means clustering method to cluster the mechanical properties and eliminates the clusters with poor performance; then uses the k-means clustering method to cluster the alloy compositions and divides the samples into several categories; takes the alloy compositions of different categories and their atomic features after feature screening as inputs and their mechanical properties as outputs; uses 12 different algorithms to establish single-objective machine learning models respectively; for each machine learning model, uses the leave-one-out cross-validation method and the Pearson index R as the accuracy index of the machine learning model, and for each different mechanical property, selects the machine learning model with the largest Pearson index R; performs interpolation and orthogonal permutation combination on the collected lead-free tin-based alloy composition data as virtual samples; finally, inputs the virtual samples into the machine learning model to obtain the predicted values of mechanical properties, and optimizes the alloy compositions with excellent performance according to the predicted values, so as to realize the auxiliary design of alloys.

[0057] The above description only explains the embodiments of the present invention in combination with the drawings, but the present invention is not limited to the above embodiments, and various changes can be made according to the purpose of the invention of the present invention. Any changes, modifications, substitutions, combinations or simplifications made based on the spirit and principle of the technical solution of the present invention shall be equivalent replacement methods, as long as they meet the invention purpose of the present invention and do not deviate from the technical principle and invention concept of the present invention, they all belong to the protection scope of the present invention.

Claims

1. A method for designing lead-free tin-based solder alloys assisted by adaptive machine learning after clustering screening, characterized in that The method includes the following steps: Step S1: Collect and obtain the material data of the lead-free tin-based solder alloy, and establish a basic data set; Step S2: Use the k-means clustering method to cluster the mechanical properties. Take the product of different mechanical property data as the comprehensive target performance data, use the comprehensive performance parameter of the mean value of each cluster as the mechanical property index to evaluate the cluster, select the clusters with excellent comprehensive performance, and eliminate the clusters with poor comprehensive performance; Then use the k-means clustering method to cluster the alloy composition and divide the samples into at least two categories; Step S3: Use the alloy composition to construct atomic features, perform high-correlation filtering and feature screening on the atomic features. Use the alloy composition of different category samples and the screened atomic features as input data, and their mechanical property data as output data; Step S4: Use 12 different algorithms to establish single-objective machine learning models respectively. For each machine learning model, use the leave-one-out cross-validation method and the Pearson index R as the accuracy index of the machine learning model. For each different mechanical property, select the machine learning model with the largest Pearson index R; Step S5: Perform interpolation and orthogonal permutation combination on the collected lead-free tin-based alloy composition data to obtain virtual samples. Input the virtual samples into the machine learning model to obtain the predicted values of the mechanical properties, and screen the alloy composition according to the predicted values, so as to realize the design of alloy parameters.

2. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloy after clustering screening according to claim 1, wherein The specific content of step S1 includes: Collect the mass ratio of the alloy composition of the lead-free tin-based solder; Collect the mechanical property data of the lead-free tin-based solder, including at least the tensile strength and the elongation at break.

3. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloy after clustering screening according to claim 1, wherein, The specific content of step S2 includes: Use the k-means clustering method to cluster the mechanical properties of tensile strength and elongation at break, eliminate the data classes with relatively low mechanical properties, and retain the remaining data; Use the k-means clustering method to cluster the alloy composition, and classify the samples according to the clustering results.

4. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloy after clustering screening according to claim 3, characterized in that In step S2, use the k-means clustering method to cluster the mechanical properties of tensile strength and elongation at break, eliminate the data classes with relatively low mechanical properties, and retain the remaining data. The accuracy of the screened data set input into the machine learning model is greatly improved compared with that of the data without elimination input into the machine learning model.

5. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloy after clustering screening according to claim 1, wherein The specific content of step S3 includes: Convert the alloy composition ratio to the atomic ratio as a supplementary feature, and add the atomic radius, the number of valence electrons, and the Pauling electronegativity as supplementary features, and perform high-correlation filtering and shapvalue feature importance ranking screening on the supplementary features.

6. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloy after clustering screening according to claim 1, wherein, The specific content of step S4 includes: Use the alloy composition and the screened atomic features as input and a mechanical property as output to establish a single-objective machine learning model; Use 12 different algorithms to establish single-objective machine learning models respectively. The 12 models include: Linear, Ridge, LASSO, MLP, Decision Tree, Random Forest, Xgboost, Adaboost, GBDT, Bagging, SVM, KNN; Each machine learning model adopts the leave-one-out cross-validation method; The Pearson index R is used as an accuracy metric for machine learning models. The mathematical expression of R is as follows: For each different mechanical property, select the machine learning model with the largest Pearson index R; where n represents the number of samples, X i represents the predicted value, Y i represents the measured value, represents the mean of the predicted values, represents the mean of the measured values.

7. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloy after clustering screening according to claim 1, wherein The specific steps of step S5 include: Based on the collected lead-free tin-based alloy composition data, perform internal interpolation from the maximum value to zero for each alloy composition with a set step size, and then perform permutations and combinations to obtain virtual samples; Input the virtual samples into the machine learning model to obtain mechanical property prediction values. Use the product of different mechanical properties as a comprehensive index to measure strength and toughness, and select samples with better comprehensive performance from the Pareto boundary of mechanical properties for experimental preparation, thereby realizing the auxiliary alloy design process.

8. The method for adaptively machine learning-assisted design of lead-free tin-based solder alloys after clustering and screening according to claim 1, wherein The composition range of the alloy suitable for design is by weight percentage: 3.0 - 5.5% Ag, 0.5 - 1.0% Cu, 1.0 - 5.0% Bi, 0.01 - 1.0% Ti, 0.01 - 1.0% Ni, and the balance is Sn.

Citation Information

Patent Citations

  • High-entropy alloy hardness prediction method based on machine learning

    CN112216356A

  • Method for efficiently predicting stability of perovskite based on integrated machine learning

    CN113052367A