A bayesian network based tin-based material composition performance inference method

By using a Bayesian network-based approach, combined with the normalization of tin-based material data and the quantile method, a tin-based Bayesian network is constructed. This solves the problem that existing component ratio prediction models cannot consider intrinsic correlations, and achieves both accuracy and efficiency in inferring component ratios from given performance.

CN120951147BActive Publication Date: 2026-04-21YUNNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNNAN UNIV
Filing Date
2025-10-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing tin-based material composition prediction models fail to effectively consider the intrinsic relationship between composition ratios and material properties, and cannot infer the remaining composition ratios based on given expected material properties and some composition ratios.

Method used

A Bayesian network-based approach is adopted. By normalizing the data of tin-based materials and dividing the intervals using the quantile method, combined with Gini impurity assessment and decision tree algorithm, a tin-based Bayesian network is constructed. GFlowNet is used to learn the posterior distribution of directed edges, and Monte Carlo sampling and KL divergence calculation are used to determine the causal strength and importance, thereby realizing the inference of component ratio.

Benefits of technology

This method enables the deduction of the remaining component ratios from given desired material properties and partial component ratios, thus meeting material performance requirements and improving the accuracy and efficiency of tin-based material composition deduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951147B_ABST
    Figure CN120951147B_ABST
Patent Text Reader

Abstract

The application relates to the field of material design, and in particular to a tin-based material composition performance inference method based on a Bayesian network. The method comprises the following steps: taking material performance variables and composition proportion variables as nodes in a DAG, learning a posterior distribution of the DAG based on a reward function based on GFlowNet; based on a tin-based Bayesian network, performing a maximum likelihood estimation algorithm to learn CPT parameters of each node in the tin-based Bayesian network, and obtaining a target tin-based Bayesian network; performing a Monte Carlo sampling algorithm in the target tin-based Bayesian network to obtain a data set, and based on preset composition proportion variables or preset material performance variables, combining a preset candidate cause set to perform KL divergence calculation, and determining the causal strength and importance between the nodes of each target tin-based Bayesian network. The purpose of inferring the remaining composition proportion from the given expected material performance and part of the composition proportion, and then obtaining the composition proportion meeting the material performance requirement is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of materials design, and in particular to a method for inferring the composition and properties of tin-based materials based on Bayesian networks. Background Technology

[0002] Tin-based materials generally refer to alloys or composite materials with tin as the main component, combined with other metals (such as copper, silver, zinc, lead, antimony, etc.), and are widely used in the energy, electronics, and machinery fields. Inferring the performance of tin-based materials by giving different component ratios, or inferring the ratios of remaining components based on desired material properties and some component ratios, is of great significance for improving the reliability of tin-based material products and realizing the lean use of precious metals such as tin.

[0003] Typically, methods for predicting the properties of tin-based materials include using traditional machine learning models such as Support Vector Machines, Decision Trees, Random Forests, and Neural Networks. For example, by utilizing XGBoost and Random Forests, predictors can be built to forecast material properties based on composition or process, and genetic algorithms can be used to search for the optimal combination of composition, process, and material properties. Alternatively, multi-objective genetic algorithms can be used to establish a complete data model of the entire process from wax patterns and shells to castings and optimize process parameters, thereby precisely controlling the dimensions and quality of castings to ensure they meet the expected performance requirements.

[0004] However, common models for predicting material properties from component ratios do not consider the inherent relationship between different component ratios and material properties. Therefore, they cannot infer the remaining component ratios based on the given expected material properties and some component ratios, and thus cannot build a knowledge model for component inference for all possible scenarios.

[0005] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main objective of this invention is to provide a method for inferring the composition and properties of tin-based materials based on Bayesian networks, aiming to solve the problem of how to infer the proportions of the remaining components based on given desired material properties and partial component proportions.

[0007] To achieve the above objectives, this invention provides a method for inferring the composition and properties of tin-based materials based on Bayesian networks. The method includes the following steps:

[0008] Optionally, the material performance variables in the tin-based material data are normalized to obtain the material performance variable normalization result, and the performance value range is determined in the material performance variable normalization result based on the quantile method to obtain the material performance variable discrete result;

[0009] The normalized results of the material performance variables are summed and averaged to obtain a comprehensive performance index. Based on the comprehensive performance index, the Gini impurity assessment and decision tree algorithm are executed to determine the division interval of the component ratio variables in the tin-based material data. Based on the division interval of the component ratio variables, the discrete results of the component ratio variables are obtained.

[0010] The material property variables and composition ratio variables are used as nodes in the DAG, and the posterior distribution of the DAG based on the reward function is learned based on GFlowNet. The posterior distribution of the DAG based on the reward function is learned based on GFlowNet, and sampling is performed with the BIC scoring function as the benchmark. High-frequency edges are selected from the sampling results to construct a tin-based Bayesian network.

[0011] Based on the tin-based Bayesian network, the maximum likelihood estimation algorithm is executed to learn the CPT parameters of each node in the tin-based Bayesian network, and the target tin-based Bayesian network is obtained.

[0012] The Monte Carlo sampling algorithm is executed to obtain a dataset from the target tin-based Bayesian network, and KL divergence is calculated based on preset component ratio variables or preset material property variables and a preset candidate cause set to determine the causal strength and importance between the nodes of each target tin-based Bayesian network.

[0013] Optionally, based on a set of discrete results of the material property variables or a set of discrete results of the component ratio variables, and for a preset set of candidate causes, the KL divergence calculation is performed to obtain the result information content, wherein the result information content includes the information content of each cause pair in the preset set of candidate causes that produces a given result in the KL divergence calculation;

[0014] The importance of the resulting information is ranked to determine the causal strength and importance among the nodes of each target tin-based Bayesian network.

[0015] Optionally, each of the aforementioned material property variables is normalized to the [0, 1] interval to obtain the material property variable normalization result;

[0016] The performance category cutoff point is determined based on the quantile method in the normalized results of the material performance variables;

[0017] The performance value range is determined based on the performance category segmentation point;

[0018] Based on the performance value range, the discrete results of the material performance variables are obtained.

[0019] Optionally, the number of composition ratio categories can be determined from the comprehensive performance index based on the quantile method;

[0020] Based on the number of component ratio categories and the component ratio variable, a decision tree algorithm is executed to determine the partition interval of the component ratio variable;

[0021] The execution decision tree algorithm includes:

[0022] A. Evaluate the Gini impurity of the component proportion variables based on the number of component proportion categories and the component proportion variables;

[0023] B. Calculate the Gini gain corresponding to each candidate segmentation point of the component ratio variable based on the Gini impurity.

[0024] C. Based on the candidate segmentation point with the largest Gini gain as the target segmentation point, divide the component ratio variable into intervals, and return to execute steps A and B until the Gini impurity of the component ratio variable is lower than the preset Gini impurity threshold, thereby obtaining the division interval of the component ratio variable.

[0025] Optionally, the material property variables and composition ratio variables are used as nodes in the DAG, and a preset forbidden edge list is used as the structural constraint of the DAG. The degree of fit between the DAG and the tin-based material data is measured using the BIC scoring function as a benchmark.

[0026] The degree of fit is used as the reward function to guide the generation of directed edges in the DAG.

[0027] Optionally, a policy network is trained based on the reward function, and the policy network is iterated by executing the GFlowNet loss function to obtain a DAG posterior sample set;

[0028] The policy network is executed to learn the posterior distribution of the DAG based on the reward function, and the directed edges of the DAG are generated.

[0029] Calculate the frequency of the directed edges of the DAG in the posterior sample set of the DAG to determine the confidence level of the association relationship;

[0030] Directed edges of the DAG with a confidence level greater than a preset confidence threshold are selected to construct the tin-based Bayesian network.

[0031] Optionally, a Monte Carlo sampling algorithm is performed to obtain a dataset in the target tin-based Bayesian network, wherein the dataset includes at least two weighted samples, and each weighted sample consists of a component allocation variable and a material property variable, the weight of the weighted sample representing the probability that the values ​​of its component allocation variable and material property variable are consistent with a given result;

[0032] Perform KL divergence calculation to measure the impact of changes in the distribution of the preset candidate cause set on whether or not it is based on the preset component ratio variable or the preset material performance variable, and obtain the amount of information of the preset candidate cause set on the generation of the preset component ratio variable or the preset material performance variable based on the dataset.

[0033] Furthermore, to achieve the above objectives, the present invention also provides a device for inferring the composition and performance of tin-based materials based on Bayesian networks. The device includes a memory, a processor, and a program for inferring the composition and performance of tin-based materials based on Bayesian networks, which is stored in the memory and can run on the processor. When the program for inferring the composition and performance of tin-based materials based on Bayesian networks is executed by the processor, it implements the steps of the method for inferring the composition and performance of tin-based materials based on Bayesian networks as described above.

[0034] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a Bayesian network-based method for inferring the composition and performance of tin-based materials. When executed by a processor, the Bayesian network-based method for inferring the composition and performance of tin-based materials implements the steps of the Bayesian network-based method for inferring the composition and performance of tin-based materials as described above.

[0035] This invention provides a method for inferring the composition and performance of tin-based materials based on Bayesian networks. By comprehensively considering the correlation between different component ratios and material properties in tin-based material data, the method infers the remaining component ratios from given desired material properties and some component ratios, thereby obtaining the component ratios that meet the material performance requirements. Attached Figure Description

[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the invention. To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0037] Figure 1This is a schematic diagram of the hardware operating environment of the device for inferring the composition and properties of tin-based materials based on Bayesian networks, as described in an embodiment of the present invention.

[0038] Figure 2 This is a flowchart illustrating an embodiment of the method for inferring the composition and properties of tin-based materials based on Bayesian networks according to the present invention.

[0039] Figure 3 This is a GFlowNet generation strategy diagram of an embodiment of the method for inferring the composition and properties of tin-based materials based on Bayesian networks according to the present invention;

[0040] Figure 4 This is a decision tree-based Ag value segmentation diagram of an embodiment of the method for inferring the composition and properties of tin-based materials based on Bayesian networks according to the present invention.

[0041] Figure 5 This is a SnBN structure and CPT diagram corresponding to the nodes in an embodiment of the method for inferring the composition and properties of tin-based materials based on Bayesian networks according to the present invention.

[0042] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0043] This application discloses a method for inferring the composition and properties of tin-based materials based on Bayesian networks. The method normalizes the material property variables in tin-based material data to obtain normalized results. Then, based on the quantile method, it determines the performance value intervals within these normalized results, obtaining discrete results for the material property variables. Finally, it sums and averages these normalized results to obtain a comprehensive performance index. Based on this comprehensive performance index, it performs Gini impurity assessment and a decision tree algorithm to determine the partitioning intervals of the component ratio variables in the tin-based material data. Finally, based on these partitioning intervals, it obtains the composition and properties of the tin-based materials. The process involves discretizing the allocation ratio variables; using material property variables and component allocation ratio variables as nodes in a Directed Acyclic Graph (DAG), and learning the posterior distribution of the DAG based on the reward function using GFlowNet; employing a tin-based Bayesian network, performing maximum likelihood estimation to learn the CPT parameters of each node in the tin-based Bayesian network to obtain the target tin-based Bayesian network; using a Monte Carlo sampling algorithm to obtain a dataset from the target tin-based Bayesian network, and calculating KL divergence based on preset component allocation ratio variables or preset material property variables, combined with a preset candidate cause set, to determine the causal strength and importance between nodes in each target tin-based Bayesian network. This achieves the goal of inferring the remaining component allocation ratios from given desired material properties and some component allocation ratios, thereby obtaining the component allocation ratios that meet the material performance requirements.

[0044] To better understand the above technical solutions, exemplary embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While exemplary embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art.

[0045] As one implementation scheme, Figure 1 This is a schematic diagram of the hardware operating environment of the device for inferring the composition and performance of tin-based materials based on Bayesian networks, which is involved in the embodiments of the present invention.

[0046] like Figure 1 As shown, the Bayesian network-based device for inferring the composition and properties of tin-based materials may include: a processor 101, such as a central processing unit (CPU), a memory 102, and a communication bus 103. The memory 102 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. Optionally, the memory 102 may also be a storage device independent of the aforementioned processor 101. The communication bus 103 is used to enable communication between these components.

[0047] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the device for inferring the composition and properties of tin-based materials based on Bayesian networks, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0048] like Figure 1 As shown, the memory 102, which is a computer-readable storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a program for inferring the composition and properties of tin-based materials based on Bayesian networks.

[0049] exist Figure 1 In the Bayesian network-based tin-based material composition and performance inference device shown, the processor 101 and memory 102 can be installed in the device. The device uses the processor 101 to call the Bayesian network-based tin-based material composition and performance inference program stored in the memory 102 and performs the following operations:

[0050] The material property variables in the tin-based material data are normalized to obtain the normalized results of the material property variables. Based on the quantile method, the performance value range is determined in the normalized results of the material property variables to obtain the discrete results of the material property variables.

[0051] The normalized results of the material performance variables are summed and averaged to obtain a comprehensive performance index. Based on the comprehensive performance index, the Gini impurity assessment and decision tree algorithm are executed to determine the division interval of the component ratio variables in the tin-based material data. Based on the division interval of the component ratio variables, the discrete results of the component ratio variables are obtained.

[0052] The material property variables and composition ratio variables are used as nodes in the DAG, and the posterior distribution of the DAG based on the reward function is learned based on GFlowNet. The posterior distribution of the DAG based on the reward function is learned based on GFlowNet, and sampling is performed with the BIC scoring function as the benchmark. High-frequency edges are selected from the sampling results to construct a tin-based Bayesian network.

[0053] Based on the tin-based Bayesian network, the maximum likelihood estimation algorithm is executed to learn the CPT parameters of each node in the tin-based Bayesian network, and the target tin-based Bayesian network is obtained.

[0054] The Monte Carlo sampling algorithm is executed to obtain a dataset from the target tin-based Bayesian network, and KL divergence is calculated based on preset component ratio variables or preset material property variables and a preset candidate cause set to determine the causal strength and importance between the nodes of each target tin-based Bayesian network.

[0055] In one embodiment, the processor 101 can be used to call a Bayesian network-based tin-based material composition and performance inference program stored in the memory 102, and perform the following operations:

[0056] Based on a set of discrete results of the material property variables or a set of discrete results of the component ratio variables, and for a preset set of candidate causes, the KL divergence calculation is performed to obtain the result information content, wherein the result information content includes the information content of each cause pair in the preset set of candidate causes that produces a given result in the KL divergence calculation;

[0057] The importance of the resulting information is ranked to determine the causal strength and importance among the nodes of each target tin-based Bayesian network.

[0058] In one embodiment, the processor 101 can be used to call a Bayesian network-based tin-based material composition and performance inference program stored in the memory 102, and perform the following operations:

[0059] Each of the aforementioned material property variables is normalized to the [0, 1] interval to obtain the material property variable normalization result;

[0060] The performance category cutoff point is determined based on the quantile method in the normalized results of the material performance variables;

[0061] The performance value range is determined based on the performance category segmentation point;

[0062] Based on the performance value range, the discrete results of the material performance variables are obtained.

[0063] In one embodiment, the processor 101 can be used to call a Bayesian network-based tin-based material composition and performance inference program stored in the memory 102, and perform the following operations:

[0064] The number of component ratio categories is determined based on the quantile method in the comprehensive performance index;

[0065] Based on the number of component ratio categories and the component ratio variable, a decision tree algorithm is executed to determine the partition interval of the component ratio variable;

[0066] The execution decision tree algorithm includes:

[0067] A. Evaluate the Gini impurity of the component proportion variables based on the number of component proportion categories and the component proportion variables;

[0068] B. Calculate the Gini gain corresponding to each candidate segmentation point of the component ratio variable based on the Gini impurity.

[0069] C. Based on the candidate segmentation point with the largest Gini gain as the target segmentation point, divide the component ratio variable into intervals, and return to execute steps A and B until the Gini impurity of the component ratio variable is lower than the preset Gini impurity threshold, thereby obtaining the division interval of the component ratio variable.

[0070] In one embodiment, the processor 101 can be used to call a Bayesian network-based tin-based material composition and performance inference program stored in the memory 102, and perform the following operations:

[0071] The material property variables and composition ratio variables are used as nodes in the DAG, and a preset forbidden edge list is used as the structural constraint of the DAG. The degree of fit between the DAG and the tin-based material data is measured using the BIC scoring function as a benchmark.

[0072] The degree of fit is used as the reward function to guide the generation of directed edges in the DAG.

[0073] In one embodiment, the processor 101 can be used to call a Bayesian network-based tin-based material composition and performance inference program stored in the memory 102, and perform the following operations:

[0074] The policy network is trained based on the reward function, and the policy network is iterated by executing the GFlowNet loss function to obtain the DAG posterior sample set;

[0075] The policy network is executed to learn the posterior distribution of the DAG based on the reward function, and the directed edges of the DAG are generated.

[0076] Calculate the frequency of the directed edges of the DAG in the posterior sample set of the DAG to determine the confidence level of the association relationship;

[0077] Directed edges of the DAG with a confidence level greater than a preset confidence threshold are selected to construct the tin-based Bayesian network.

[0078] In one embodiment, the processor 101 can be used to call a Bayesian network-based tin-based material composition and performance inference program stored in the memory 102, and perform the following operations:

[0079] The Monte Carlo sampling algorithm is performed to obtain a dataset in the target tin-based Bayesian network, wherein the dataset includes at least two weighted samples, and each weighted sample consists of a component allocation variable and a material property variable, the weight of the weighted sample representing the probability that the values ​​of its component allocation variable and material property variable are consistent with a given result;

[0080] Perform KL divergence calculation to measure the impact of changes in the distribution of the preset candidate cause set on whether or not it is based on the preset component ratio variable or the preset material performance variable, and obtain the amount of information of the preset candidate cause set on the generation of the preset component ratio variable or the preset material performance variable based on the dataset.

[0081] Based on the hardware architecture of the Bayesian network-based tin-based material composition and performance inference device described above, an embodiment of the Bayesian network-based tin-based material composition and performance inference method of the present invention is proposed.

[0082] Reference Figure 2 In the first embodiment, the method for inferring the composition and properties of tin-based materials based on Bayesian networks includes the following steps:

[0083] Step S100: Normalize the material property variables in the tin-based material data to obtain the normalized results of the material property variables, and determine the performance value range in the normalized results of the material property variables based on the quantile method to obtain the discrete results of the material property variables.

[0084] It should be noted that in the original tin-based material data, the value ranges of the component ratio variables and material property variables are different, and learning the correlation between variables with large value ranges requires a large amount of time and space overhead. Therefore, it is necessary to divide the value ranges of the component ratio variables and material property variables, and improve the learning efficiency of SnBN (Sn-based Bayesian Network) by discretizing the value ranges.

[0085] In this embodiment, each material performance variable is normalized to the [0, 1] interval to obtain the material performance variable normalization result; the performance category segmentation point is determined in the material performance variable normalization result based on the quantile method; the performance value interval is determined according to the performance category segmentation point; and then, the material performance variable discrete result is obtained based on the performance value interval.

[0086] Optionally, given tin-based material data containing d types of component ratio variables and k types of material property variables, the value spaces for the component ratio variables and material property variables are respectively represented as follows: and Let N be the data samples of tin-based materials. , For the i-th tin-based material composition sample, its j-th dimension express The sample value of the j-th component ratio, This represents the value of the material properties of the k-th type in the i-th tin-based material sample.

[0087] To discretize the values ​​of the component ratio variable based on the values ​​of the material property variable, based on The maximum value in D and minimum value Take the value of each sample Normalize to the interval [0, 1]. This can be expressed by the following formula:

[0088]

[0089] To ensure comparability among different material performance variables and to achieve uniformity in classification standards, a quantile method is used to find performance category dividing points for each material performance variable in D. The values ​​are divided into multiple categories. For example, when there are two categories, this method finds the median of the data distribution as the split point, thus dividing the material properties into two value ranges: "low" and "high". Ultimately, the resulting set of material property variable values ​​is obtained. That is, the normalized result of material property variables.

[0090] Step S200: Perform a summation and averaging operation on the normalized results of the material performance variables to obtain a comprehensive performance index. Based on the comprehensive performance index, execute the Gini impurity assessment and decision tree algorithm to determine the division interval of the component ratio variables in the tin-based material data, and obtain the discrete results of the component ratio variables based on the division interval of the component ratio variables.

[0091] It should be noted that different values ​​of the material property variables determine the composition ratio of tin-based materials. For example, to construct a "highly conductive" tin-based material, it is necessary to increase the proportion of "silver." Therefore, based on the aforementioned values ​​of the material property variables, the range of values ​​for the composition ratio variables is discretized.

[0092] In this embodiment, the number of component ratio categories is determined in the comprehensive performance index based on the quantile method; then, according to the number of component ratio categories and the component ratio variable, a decision tree algorithm is executed to determine the division interval of the component ratio variable. The aforementioned decision tree algorithm includes: evaluating the Gini impurity of the component ratio variable based on the number of component ratio categories and the component ratio variable; calculating the Gini gain corresponding to each candidate segmentation point of the component ratio variable based on the Gini impurity; then, using the candidate segmentation point with the largest Gini gain as the target segmentation point, dividing the component ratio variable into intervals, and executing the steps of evaluating the Gini impurity of the component ratio variable based on the number of component ratio categories and the component ratio variable, calculating the Gini gain corresponding to each candidate segmentation point of the component ratio variable based on the Gini impurity, and dividing the component ratio variable into intervals based on the candidate segmentation point with the largest Gini gain as the target segmentation point, until the Gini impurity of the component ratio variable is lower than a preset Gini impurity threshold, thereby obtaining the division interval of the component ratio variable.

[0093] In other words, in the decision tree algorithm, the merits of each candidate split point for the component ratio variable are evaluated by calculating the Gini gain; then, the candidate split point that maximizes the Gini gain is selected, and the partition interval of the component ratio variable is determined based on this split point. This partitioning process is performed recursively until the Gini impurity of the component ratio variable used as a node falls below a preset Gini impurity threshold, at which point the partitioning stops.

[0094] Optionally, the comprehensive performance index for each data point is first calculated based on the normalized material property variables. The formula is expressed as follows:

[0095]

[0096] Then, in order to base on comprehensive performance indicators Discretize the component ratio variable, use the quantile method to determine the cut-off point, divide the continuous value comprehensive performance index into several discrete intervals, and use these intervals as the category labels of the component ratio variable.

[0097] When a component ratio variable is divided into any discrete value interval of the component ratio variable, the higher the probability that the samples' labels belong to the same category, the purer the samples within the discrete value interval, and the better the interval division effect. Therefore, Gini impurity, a widely used indicator to measure the purity of dataset segmentation, is used to find the optimal split point for the component ratio variable's value interval, ensuring that the comprehensive performance index values ​​of the subsets after the component ratio variable is divided belong to the same category as much as possible. The formula for calculating Gini impurity is as follows:

[0098]

[0099] in, For the number of categories, For the data sample set of the component ratio variable, Indicates that the category in S belongs to The sample proportion.

[0100] A lower Gini impurity indicates a purer sample within a discrete value range. Therefore, the following Gini gain metric data sample set is defined. Divided into and Subsequently, the degree to which the impurity of the ginnig decreased:

[0101]

[0102] Furthermore, to find the optimal segmentation threshold for the component ratio variables, with the objective of maximizing the Gini gain, a decision tree is used to discretize the value interval of each component ratio variable. Specifically:

[0103] First, given the set of all possible values ​​for a given distribution ratio. As the root node, sort all its values, take the midpoint of every two adjacent values, and construct an initial set of candidate split points.

[0104] Then, iterate through all candidate split points and perform the following operations:

[0105] Operation 1: Move the current node Divided into two subsets, left and right. and The impurity of the Gini is calculated according to the formula for calculating Gini impurity. , and .

[0106] Step 2: Calculate the segmentation based on the formula for calculating the degree of decrease in Gini impurity. , and Gini gain .

[0107] Finally, select the split point that maximizes the Gini gain as the candidate split point, and repeat the above operations 1 and 2 until the Gini impurity of the split set is lower than the preset threshold.

[0108] The purpose of this is to ensure that the decision tree finds the optimal split point for the component ratio variables, dividing the value space of the d component ratio variables into... Each discrete interval corresponds to a category, thus mapping continuous material property values ​​to... A discrete range of values. For example, At that time, the median of the data distribution of the "Ag" component ratio variable was found as the dividing point, thus dividing the component ratio into two value intervals: "low" and "high". Finally, the set of values ​​for the segmented material property variable was obtained. Thus, the dataset after discretizing the range of variable values ​​is obtained. .

[0109] Step S300: The material performance variables and composition ratio variables are used as nodes in the DAG, and the posterior distribution of the DAG based on the reward function is learned based on GFlowNet. The posterior distribution of the DAG based on the reward function is learned based on GFlowNet by sampling with the BIC (Bayesian Information Criterion) scoring function as the benchmark, and high-frequency edges are selected from the sampling results to construct a tin-based Bayesian network.

[0110] It should be noted that, in order to avoid introducing noise by constructing a single DAG (Directed Acyclic Graph) as a SnBN structure from sparse tin-based material data, which could lead to erroneous directed edges between component ratio variables and material property variables, GFlowNet (Generative Flow Network) is used to learn the posterior distribution of the DAG in the tin-based material data. By sampling multiple DAGs to obtain edges with high confidence, a more robust correlation is learned on the sparse data, laying the foundation for effective inference of subsequent component properties.

[0111] In this embodiment, SnBN uses a Directed Acyclic Graph (DAG) to represent the relationship between the component ratio variables and material property variables in the tin-based material. The component ratio variables are... and material property variables Represented as nodes in the diagram, material property nodes. The value is the L categories into which the component ratio variable is divided, and the node corresponding to the component ratio variable. The value of is found by the decision tree. Each node represents a discrete interval. Directed edges represent the direct relationships between variables, and the degree of mutual influence between the component ratio variables and material property variables in tin-based materials is quantitatively described through the CPT (Conditional Probability Table) of each node.

[0112] The structure learning of SnBN aims to learn from... Learn an optimal graph structure The distribution ratio variable will be formed. and material property variables Represented as a node in SnBN, a material property node. The value is divided into L categories, and the component ratio node is... The value of is found by the decision tree. A discrete interval.

[0113] Optionally, the material performance variables and composition ratio variables are used as nodes in the DAG, and a preset forbidden edge list is used as the structural constraint of the DAG. The degree of fit between the DAG and the tin-based material data is measured using the BIC scoring function as a benchmark. Then, the degree of fit is used as the reward function to guide the generation of directed edges in the DAG.

[0114] As an alternative implementation, firstly, all component ratio variables and material property variables are treated as nodes in a DAG as the initial structure, and a preset forbidden edge list is given based on expert knowledge. As a structural constraint, it is used to avoid including unreasonable component-performance relationships in the model.

[0115] Then, based on the generative model GFlowNet, the posterior distribution of the DAG between the composition ratio variables and material property variables in the tin-based material data is learned. The DAG learning process is transformed into a directed edge generation process based on a policy network, thereby efficiently sampling DAG structures that conform to the posterior distribution and avoiding the overfitting problem caused by learning a single DAG under sparse tin-based material data.

[0116] To ensure that the DAG structure learned by the policy network accurately reflects the relationship between the allocation ratio variable and the material property variable, a reward function is designed. This guides the generation of directed edges. The formula is as follows:

[0117]

[0118] According to Bayes' theorem, we know that , The structure of all SnBNs is assumed to be uniformly distributed as a priori, which can be regarded as a constant term. .

[0119] The BIC scoring function, widely used for model selection, is employed to measure the goodness of fit between the DAG structure constructed from composition-performance variables and the tin-based material data, thereby approximating the calculation. This can be expressed by the following formula:

[0120]

[0121] in, for exist Log-likelihood of the optimal parameters These are independent parameters in the structure. This represents the number of samples.

[0122] Based on the above derivation, the reward function It can be approximated by the BIC score. The formula is as follows:

[0123]

[0124] In tin-based materials, any component ratio variable may be correlated with material property variables. To address this, it is necessary to... Training Policy Network This allows it to learn the next directed edge based on the current structure G, using a Graph Neural Network (GNN) architecture and a self-attention mechanism. The structural characterization is used to fully capture the global and local correlations among variables in tin-based materials. To this end, The adjacency matrix representation is used as input to obtain its global representation. ,as well as Each component distribution variable and Characterization of individual material property variable nodes The formula is as follows:

[0125]

[0126] Let G be the transition through adding edges. The state is , Indicates the construction is terminated; otherwise... Construct the state transition probability distribution for termination. From global representation calculate:

[0127]

[0128] If the construction is not terminated A new correlation between tin-based material variables is obtained through the "adding edge" action, and the relationship is transferred to a new state. The probability distribution of this transition Based on node representation Calculation. Expressed as a formula:

[0129]

[0130] in, A mask matrix introduced to prevent DAG from forming cycles.

[0131] In construction When, if a new directed edge is introduced An unreasonable component-performance relationship, or the generation of a directed cycle in G, will affect the probability distribution of the transition. Based on node representation The calculated output probability is forcibly set to 0. Finally, from state... Transferred to The process is as follows Figure 3 As shown, in Figure 3 In this context, tin-based materials consist of tin (Sn), copper (Cu), and silver (Ag). Their total probability is expressed by the formula:

[0132]

[0133] Furthermore, the steps described above, including learning the posterior distribution of the DAG based on the reward function using GFlowNet, sampling using the BIC scoring function as a benchmark, and selecting high-frequency edges from the sampling results to construct a tin-based Bayesian network, include: training a policy network based on the reward function; iterating the policy network using the GFlowNet loss function to obtain a DAG posterior sample set; executing the policy network to learn the posterior distribution of the DAG based on the reward function to generate directed edges of the DAG; calculating the frequency of the DAG directed edges appearing in the DAG posterior sample set to determine the association confidence; and then selecting DAG directed edges whose association confidence is greater than a preset confidence threshold to construct the tin-based Bayesian network.

[0134] Optionally, to address the sparsity of tin-based material data and the uncertainty of composition-performance variable relationships, the GFlowNet loss function is used to optimize the policy network, making... Directed edges are stably generated by learning the entire posterior distribution. The GFlowNet loss function can be expressed by the following formula:

[0135]

[0136] in, Indicates in Removing an edge with moderate probability yields The probability of.

[0137] Implemented using an experience playback mechanism The core of the training is to optimize the network parameters of the strategy through alternating generative and learning loops.

[0138] Specifically, it includes four stages: trajectory generation, experience storage, network training, and iterative optimization.

[0139] Trajectory generation. From an unbounded graph with component ratio variables and material property variables as nodes. start, Through a series of "adding edges" actions Build G until G is in a certain state. The construction process ends and the trajectory of the relationship construction is obtained. .

[0140] Experience storage. Transition all states in the trajectory ( ) and its corresponding structure and Stored in an experience replay pool, where .

[0141] Network training. Mini-batch data is randomly sampled from the experience replay pool, the loss is calculated using the GFlowNet loss function described above, and the network parameters are updated via backpropagation. .

[0142] Iterative optimization. Repeat the above trajectory generation, experience storage, and network training steps until the model converges.

[0143] Based on the trained GFlowNet, a posterior sample set containing B DAGs representing the correlation between the composition and properties of tin-based materials is efficiently generated. Then, the directed edges in the posterior sample set are computed. The frequency of occurrence of a term is used to assess the confidence level of the association; a higher confidence level indicates that the association exists. The greater the probability, the higher the likelihood. Therefore, a confidence threshold is preset. Only associations with a confidence level exceeding a preset confidence threshold are considered reliable. The final structure of SnBN. edge set All confidence levels exceeding the threshold The edges constitute the boundary. This can be expressed by the following formula:

[0144]

[0145] Among them, confidence level Defined as:

[0146]

[0147] in, To generate structure edge set, It is an indicator function. In the generation Finally, a final loop test is required to ensure... Its directed acyclic property.

[0148] Step S400: Based on the tin-based Bayesian network, perform the maximum likelihood estimation algorithm to learn the CPT parameters of each node in the tin-based Bayesian network, and obtain the target tin-based Bayesian network.

[0149] In this embodiment, based on tin-based material data and the aforementioned SnBN structure... Parameter learning is performed. The conditional probability table for each node is calculated by statistically analyzing the number of instances in the training data using the maximum likelihood estimation algorithm, and this table serves as the result of parameter learning. For a variable V with parent node set U, the influence of the frequency of entities in U in D on the frequency of entities in V is used as the conditional probability. The value of quantifies the dependency between V and U in BN, and is calculated as follows:

[0150]

[0151] in, This indicates that the variable V takes the value of The number of instances when U takes the value u. This represents the number of instances when U takes the value u; both can be obtained by counting from D.

[0152] Then, The results are entered into the corresponding positions to obtain the CPT parameters of SnBN; Each node will generate a CPT, ultimately resulting in the SnBN model. .

[0153] Step S500: Execute the Monte Carlo sampling algorithm to obtain a dataset in the target tin-based Bayesian network, and perform KL divergence calculation based on preset component ratio variables or preset material property variables and a preset candidate cause set to determine the causal strength and importance between the nodes of each target tin-based Bayesian network.

[0154] In this embodiment, based on a set of discrete results of the material property variables or a set of discrete results of the component ratio variables, and for a preset set of candidate causes, the KL divergence calculation is performed to obtain the result information content, wherein the result information content includes the information content of each cause pair in the preset set of candidate causes that produces a given result in the KL divergence calculation; then, the result information content is sorted by importance to determine the causal strength and importance between the nodes of each target tin-based Bayesian network.

[0155] It should be noted that SnBN-based inference of the composition and properties of tin-based materials aims to measure the strength of the correlation between component ratios and material properties, in order to obtain the set of component rankings that give tin-based materials their properties. To this end, given a material property variable or component ratio variable as the result, and the component ratio variable or material property variable leading to that result as the cause, the problem of inferring the composition and properties of tin-based materials is transformed into a SnBN-based probabilistic reasoning task.

[0156] Further, the steps of obtaining a dataset from the target tin-based Bayesian network by performing the Monte Carlo sampling algorithm, and calculating KL divergence based on a preset component ratio variable or preset material performance variable and a preset candidate cause set to determine the causal strength and importance between the nodes of the target tin-based Bayesian network include: obtaining a dataset from the target tin-based Bayesian network by performing the Monte Carlo sampling algorithm, wherein the dataset includes at least two weighted samples, and each weighted sample consists of a component ratio variable and a material performance variable, the weight of the weighted sample representing the probability that the values ​​of its component ratio variable and material performance variable are consistent with a given result; then, performing KL divergence calculation to measure the impact of changes in the distribution of the preset candidate cause set under conditions whether or not based on the preset component ratio variable or preset material performance variable, to obtain, based on the dataset, the amount of information of the preset candidate cause set on the generation of the preset component ratio variable or preset material performance variable.

[0157] Optionally, firstly, to efficiently calculate the strength of the causal relationship between different component ratios and their combinations and material properties, a set of component ratio variables or material property variables with given values ​​is used. As a result, the Monte Carlo method was used to analyze the constructed SnBN model. Generated in Weighted sample set Each sample consists of a set of tin-based material composition ratio variables and material property variables with given values, and its weight represents the probability that the value of the sample variable is consistent with the given result.

[0158] Specifically, first initialize the i-th weighted sample set. The weight is Then, according to Topological sorting traverses all nodes For nodes set of parent nodes The set obtained after sampling is If node For evidence variables, search from CPT probability ,calculate Update the sample weight ;like For non-evidence variables, the conditional probability distribution in CPT A value is randomly sampled from the middle as The value. Finally, repeat according to... Topological sorting traverses all nodes For nodes set of parent nodes The set obtained after sampling is The action, to obtain Each sample and its corresponding weight Finally generated A weighted sample set used to compute the posterior distribution. .

[0159] In order to base on weighted sample sets Measuring the candidate cause set For the results The amount of information is measured using the KL divergence metric, regardless of whether it is given. Under the conditions The effect of distribution variation. The calculation formula is as follows:

[0160]

[0161] in, Directly from discretized datasets This was obtained from statistics. It can be expressed as a formula:

[0162]

[0163] in, It is in the dataset Status is The number of samples, That is the total number of samples. Through the generated weighted posterior sample set Make an estimate:

[0164]

[0165] Further based on the KL divergence metric, whether or not... Under the conditions The effect of distribution changes, calculate all The information content was analyzed, and all candidate causal variables were ranked in descending order based on this. The posterior probability distributions of the top-ranked and most important variables were further analyzed. And find the state with the highest posterior probability. This is considered its optimal recommended value. The formula is expressed as follows:

[0166]

[0167] In the technical solution provided in this embodiment, the value space of material property variables is divided using the quantile method. For the component ratio variables, a decision tree algorithm with the objective of maximizing Gini gain is used to search for and determine the optimal split point for each component ratio variable. This allows the discretized variables to better reflect the intrinsic relationship between composition and performance, effectively ensuring the model training convergence efficiency, stability, and generalization ability. This solves the problems of high time and space overhead, complex model structure, and poor generalization ability associated with strategies that construct correlations for every possible precise value of a variable.

[0168] By employing GFlowNet to learn the posterior distribution of a directed acyclic graph, structures are sampled from the posterior distribution, and reliable edges are selected from the sampled structures by setting a confidence threshold. This ultimately constructs a tin-based Bayesian network, providing a theoretically sound and interpretable knowledge model for accurate composition inference. This addresses the problems of noise overfitting, spurious associations, and inaccurate composition inference results caused by simply constructing a single Bayesian network due to the sparse data samples in tin-based materials.

[0169] By transforming the inference of material composition performance into a probabilistic inference task based on tin-based Bayesian networks, and ranking the information content of causal variables by calculating KL divergence, the state with the highest posterior probability is identified as the recommended solution. This enables the screening and optimization of material composition ratios for specific material performance requirements. It solves the problems of numerous and uncertain combinations of correlations between different composition ratios of tin-based materials and material performance indicators, as well as the difficulty of inference.

[0170] In other words, the present invention provides a method for inferring the composition and performance of tin-based materials based on Bayesian networks. This method comprehensively considers the correlation between different component ratios and material properties in tin-based material data, and infers the remaining component ratios from given desired material properties and some component ratios, thereby obtaining the component ratios that meet the material performance requirements.

[0171] For example, suppose 400 data points out of 500 tin-based material composition / property data are selected as samples to train the model, and the composition / property inference is performed on the remaining 100 data points. Statistical information for some tin-based solders is shown in Table 1.

[0172] Table 1. Examples of Composition Ratios and Material Property Data for Tin-Based Solder

[0173]

[0174] Selecting component ratio variables For Sn, Ag, and Cu, the material property variables are... For tensile strength, liquidus temperature, and Vickers hardness.

[0175] The value ranges of the three selected material performance indicators are discretized. For each material performance indicator, the number of discretization levels is set. These correspond to the "low" and "high" states, respectively. For the "tensile strength" material property variable in the selected 6 samples... First find minimum value and maximum value ,pass Obtain normalized data Use the quantile method to... Segmentation, in Find the corresponding The quartile 591, using quartiles to divide Divided into A range of values, i.e. and The intervals correspond to "low" and "high" respectively. Finally, based on the segmentation threshold, a discrete material performance index consisting of 6 "low" and "high" labels is generated. The specific segmentation points and value intervals of all calculated material performance variables are shown in Table 2.

[0176] Table 2. Cutoff points and value ranges for material property variables.

[0177]

[0178] Then, based on the obtained normalized material property vector Calculate comprehensive performance indicators For the first data point in the dataset, first find its normalized performance vector. Through the formula Calculate the comprehensive performance index In obtaining each piece of data Subsequently, based on the quantile method... Segmentation, find The quartile in the data is 0.4805. Use the quartile to divide... Divided into A range of values, i.e. and These correspond to two levels: "Excellent" and "Poor". Finally, based on the segmentation threshold, a discrete comprehensive performance index consisting of six "Excellent" and "Poor" labels is generated. The normalized material performance variables and the obtained comprehensive performance index are shown in Table 3.

[0179] Table 3 Normalized performance indicators

[0180]

[0181] The Ag content, a component ratio variable, is discretized using a decision tree. The discrete comprehensive performance index has two states: {excellent, poor}. The sample data is shown in Table 4.

[0182] Table 4 Sample data for Ag

[0183]

[0184] Ag component ratio variable dataset The dataset contains 6 samples; 3 are rated "excellent" and 3 are rated "poor" in terms of overall performance. The formula used is... calculate Impurity of the gin:

[0185] The midpoint of the ratio of adjacent components in Ag is used as the candidate split point. For example, the midpoint of the ratios of adjacent samples 1 and 2 is (3.0+3.5) / 2=3.25. Taking the candidate split point of 3.25 as an example, the Gini gain calculation steps for candidate splits based on decision trees include:

[0186] Partition the dataset. Left subset. The right subset contains samples 1 and 2, for a total of 2 samples. It includes samples 3, 4, 5, and 6, for a total of 4 samples.

[0187] Calculate the Gini impurity of the subset. The overall performance index distribution is {Excellent: 0, Poor: 2}, and its Gini impurity is: . The overall performance index distribution is {Excellent: 3, Poor: 1}, and its Gini impurity is:

[0188] According to the formula Calculate the Gini gain of this segment.

[0189] By calculating and comparing the remaining candidate segmentation points, the candidate segmentation point 3.25 yielded the largest Gini gain of 0.25 so far. Therefore, 3.25 was selected as the first segmentation point. The schematic diagram of segmenting Ag is shown below. Figure 4 As shown.

[0190] After evaluating different candidate cut-off points for each component ratio variable, the Gini gain results were calculated. 85.75, 3.25, and 1.75 were the candidate cut-off points with the largest Gini gains for Sn, Ag, and Cu, respectively. Taking Ag as an example, its optimal cut-off point is 3.25. Based on this, the value space of Ag can be divided into two new categories: when the original Ag content is ≤3.25%, its category is recorded as "low," and otherwise as "high." Table 5 shows the results of discretizing some variable values ​​after the value range is discretized.

[0191] Table 5 Discretized Datasets

[0192]

[0193] Then, taking “Sn”, “Ag” and “Cu” as examples as component ratio variables, “tensile strength (TL)”, “liquidothermal temperature (TS)” and “Vickers hardness (HV)” are used as material property variables.

[0194] Transferring datasets Train the GFlowNet model to generate a graph structure. Taking the process as an example, from the process of using component ratio variables and material property variables as nodes, the boundaryless graph is used. Begin. Using the formula. The constructed neural network is obtained global representation ,as well as Each component distribution variable and Characterization of individual material property variable nodes According to the formula Calculate the state transition probability distribution for construction termination This indicates that GFlowNet has a 10% probability of stopping the construction of this structure and directly generating... If the construction is not terminated, A new correlation between tin-based material variables is obtained through the "adding edge" action, and the relationship is transferred to a new state. The probability distribution of this transition. Depend on Through Calculate and obtain the added edge The probability is 0.23, adding an edge The probability is 0.17, add an edge The probability is 0.02, equal to the probability of all addable edges. Based on the obtained probability distribution, choose [the edge] with a probability of 17%. Add edge get And continue calculating And continue computation while the construction is not terminated. Finally, at step t, with The probability of terminating the construction of the structure ultimately generates the structure. .

[0195] The above from arrive The complete construction process is a trajectory To make policy networks Learning the entire posterior distribution to stably generate directed edges requires training it with a large number of generated trajectories. The threshold for loss convergence is set to 0.01, according to the formula... Calculate the loss, repeat the buffer sample and model parameter update 100,000 times, and consider the loss to be less than 0.01. The iterative optimization converges and ends. Then, using... Generate a posterior sample set containing 1,000 DAGs. Furthermore, by statistically analyzing each possible directed edge... The confidence level of an edge is assessed by evaluating the frequency of its occurrence in the given context. Table 6 shows the confidence assessment results for some edges.

[0196] Table 6. Confidence assessment results for some candidate edges

[0197]

[0198] Set the confidence threshold to τ=0.8, and the edge ( The confidence level of ) is 0.985, because Therefore, add this edge. ;and the side ( The confidence level of ) is only 0.045, because If an edge is not found, discard it. Perform the above filtering process on all possible edges, and then perform a final loop check to ensure... The directed acyclic property yields... ,like Figure 5 As shown.

[0199] Then, using and Perform parameter learning. For example, for the CPT of node TL, in a dataset of 400 discretized data points, It appeared a total of 200 times, that is .in The data appeared a total of 164 times, that is... Then it can be expressed by formula Calculate Table 7 shows the CPT of TL calculated through the above steps. The columns corresponding to different values ​​of the child nodes represent different combinations of values ​​for their parent node Sn. The values ​​in the table are conditional probability distribution values. The final SnBN is as follows: Figure 5 As shown.

[0200] Table 7 Conditional Probability Table for Node TL

[0201]

[0202] Given the result set The candidate reason set is Further in upper combination Sampling is performed, starting with nodes Sn, Ag, and Cu for a single sample, and weights are initialized. .because Therefore, sampling is performed according to its corresponding CPT to obtain... When sampled At that time, the value of TL is fixed to "high", and updated according to the CPT of TL. Similarly, update the weights of TS and HV nodes using CPT. After sampling all nodes and updating their weights, the current sample can be obtained. and their corresponding weights Using the above method, 10,000 weighted samples were obtained.

[0203] Then, according to the formula:

[0204]

[0205] Calculate query variables For the results The amount of information, according to the formula:

[0206] and the formula:

[0207]

[0208] Calculated prior probabilities and posterior probability As shown in Table 8.

[0209] Table 8 Comparison of Prior and Posterior Probabilities for Component Proportion Variables

[0210]

[0211] Using Table 8 and The information content of each target variable was calculated, and the results are shown in Table 9.

[0212] Table 9 Information content rating and ranking of each variable

[0213]

[0214] Based on the above analysis, the following recommended solutions can be provided to achieve the comprehensive performance indicators of "high tensile strength, high liquidus temperature, and high Vickers hardness":

[0215] Recommendation 1: Ag has the highest score (0.263), making it the most critical factor in achieving the goal. Based on the posterior probabilities in Tables 8 and 9, the Ag content should be selected in the "high" range. The confidence level of this recommendation is 0.92.

[0216] Recommendation 2: Cu's score is the second highest (0.176), making it the second most important factor. Its content should also be selected in the "high" range. The confidence level is 0.88.

[0217] Recommendation 3: Sn has the lowest score (0.092), but its posterior probability of reaching the "high" state is 0.95. Its content should be selected within the "high" range. The confidence level is 0.95.

[0218] This indicates that while a high Sn content is necessary, its prevalence in the original data means that the emergence of new material performance indicators has not significantly altered the understanding of its proportions. In contrast, Ag and Cu scored higher, implying that controlling the content of Ag and Cu is more sensitive and critical than controlling Sn content during formulation adjustments. To obtain high-strength, high-melting-point, and high-hardness tin-based solder, the component proportions should follow... , , The principle is that the contents of Ag and Cu are the two most critical and sensitive factors affecting the final material properties.

[0219] It should be noted that the system embodiments described above are merely illustrative. The devices such as the Bayesian network-based tin-based material composition and performance inference device may or may not be physically separate. The Bayesian network-based tin-based material composition and performance inference device may or may not be a physical unit; that is, it may be located in one place or mapped to a terminal backend via a network. Some or all of the devices can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the system embodiment drawings provided by this invention, the connection relationships between devices indicate that they have communication connections, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without creative effort.

[0220] Therefore, the present invention also provides a computer-readable storage medium storing a Bayesian network-based method for inferring the composition and performance of tin-based materials. When executed by a processor, the Bayesian network-based method for inferring the composition and performance of tin-based materials implements the various steps of the Bayesian network-based method for inferring the composition and performance of tin-based materials as described in the above embodiments.

[0221] The computer-readable storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0222] It should be noted that, since the storage medium provided in the embodiments of this application is the storage medium used to implement the methods of the embodiments of this application, those skilled in the art can understand the specific structure and variations of the storage medium based on the methods described in the embodiments of this application, and therefore will not be repeated here. All storage media used in the methods of the embodiments of this application fall within the scope of protection of this application.

[0223] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0224] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0225] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0226] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0227] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, third, etc., does not indicate any order. These words can be interpreted as names.

[0228] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0229] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for inferring the composition and properties of tin-based materials based on Bayesian networks, characterized in that, The tin-based material composition performance inference method based on the Bayesian network comprises the following steps: The material performance variables in the tin-based material data are normalized to obtain material performance variable normalization results, and the performance value intervals are determined in the material performance variable normalization results based on the quantile method to obtain material performance variable discretization results; The material performance variable normalization results are summed and averaged to obtain a comprehensive performance index, and based on the comprehensive performance index, Gini impurity evaluation and decision tree algorithm are performed to determine the division intervals of the composition ratio variables in the tin-based material data, and based on the division intervals of the composition ratio variables, composition ratio variable discretization results are obtained; The material performance variables and the composition ratio variables are taken as nodes in a DAG, and the posterior distribution of the DAG based on a reward function is learned based on GFlowNet, wherein the posterior distribution of the DAG based on the reward function is learned based on GFlowNet, sampled based on a BIC score function, and high-frequency edges are selected from the sampling results to construct a tin-based Bayesian network; Based on the tin-based Bayesian network, a maximum likelihood estimation algorithm is used to learn the CPT parameters of each node in the tin-based Bayesian network to obtain a target tin-based Bayesian network; A Monte Carlo sampling algorithm is performed in the target tin-based Bayesian network to obtain a data set, and based on a preset composition ratio variable or a preset material performance variable, a KL divergence calculation is performed in combination with a preset candidate cause set to determine the causal strength and importance between the nodes of each target tin-based Bayesian network; The step of performing the KL divergence calculation in combination with the preset candidate cause set to determine the causal strength and importance between the nodes of each target tin-based Bayesian network comprises: Based on a group of material performance variable discretization results or a group of composition ratio variable discretization results, and for a preset candidate cause set, the KL divergence calculation is performed to obtain result information, wherein the result information includes the information amount of each reason in the preset candidate cause set for a given result in the KL divergence calculation; The result information is sorted by importance to determine the causal strength and importance between the nodes of each target tin-based Bayesian network.

2. The Bayesian network based tin-based material composition performance inference method of claim 1, wherein, The step of normalizing the material performance variables in the tin-based material data to obtain material performance variable normalization results, and determining performance value intervals in the material performance variable normalization results based on the quantile method to obtain material performance variable discretization results comprises: Each material performance variable is normalized to the [0, 1] interval to obtain material performance variable normalization results; The performance category segmentation points are determined in the material performance variable normalization results based on the quantile method; The performance value intervals are determined according to the performance category segmentation points; Based on the performance value intervals, the material performance variable discretization results are obtained.

3. The Bayesian network based tin-based material composition performance inference method of claim 1, wherein, The step of performing Gini impurity evaluation and decision tree algorithm based on the comprehensive performance index to determine the division intervals of the composition ratio variables in the tin-based material data comprises: The number of composition ratio categories is determined in the comprehensive performance index based on the quantile method; According to the component proportion category number and the component proportion variable, a decision tree algorithm is executed to determine the division interval of the component proportion variable; The execution of the decision tree algorithm includes: A. According to the component proportion category number and the component proportion variable, the Gini impurity of the component proportion variable is evaluated; B. According to the Gini impurity, the Gini gain corresponding to each candidate split point of the component proportion variable is calculated; C. Based on the candidate split point with the maximum Gini gain as the target split point, the interval of the component proportion variable is divided, and steps A and B are returned until the Gini impurity of the component proportion variable is lower than the preset Gini impurity threshold, and the division interval of the component proportion variable is obtained.

4. The Bayesian network based tin-based material composition performance inference method of claim 1, wherein, The step of taking the material performance variable and the component proportion variable as nodes in the DAG and learning the posterior distribution of the DAG based on the GFlowNet based on the reward function includes: Taking the material performance variable and the component proportion variable as nodes in the DAG, and taking the preset forbidden edge list as the structural constraint of the DAG, and taking the BIC score function as the benchmark, the fitting degree of the DAG and the tin-based material data is measured; Taking the fitting degree as the reward function, the generation of the DAG directed edge is guided.

5. The Bayesian network-based tin-based material composition performance inference method of claim 4, wherein, The step of learning the posterior distribution of the DAG based on the reward function based on the GFlowNet, sampling based on the BIC score function, and selecting high-frequency edges from the sampling results to construct the tin-based Bayesian network includes: Training a policy network based on the reward function, and iterating the policy network based on the GFlowNet loss function to obtain a DAG posterior sample set; The policy network learns the posterior distribution of the DAG based on the reward function to generate the DAG directed edge; The frequency of the DAG directed edge appearing in the DAG posterior sample set is calculated to determine the association relationship confidence; The DAG directed edge with the association relationship confidence greater than the preset confidence threshold is screened to construct the tin-based Bayesian network.

6. The Bayesian network based tin-based material composition performance inference method of claim 1, wherein, The step of executing the Monte Carlo sampling algorithm in the target tin-based Bayesian network to obtain a data set, and based on a preset component proportion variable or a preset material performance variable, combining a preset candidate cause set to perform KL divergence calculation to determine the causal strength and importance between the nodes of each target tin-based Bayesian network includes: The Monte Carlo sampling algorithm is executed in the target tin-based Bayesian network to obtain a data set, wherein the data set includes at least two weighted samples, and each weighted sample is composed of a group of component proportion variables and material performance variables, and the weight of the weighted sample represents the possibility that the values of the component proportion variables and the material performance variables are consistent with the given result; The KL divergence calculation is performed to measure the influence of the distribution change of the preset candidate cause set based on whether the preset component proportion variable or the preset material performance variable, and to obtain the information amount of the preset candidate cause set to generate the preset component proportion variable or the preset material performance variable based on the data set.

7. A Bayesian network-based tin-based material composition performance inference device characterized by comprising: The tin-based material composition performance inference device based on the Bayesian network comprises a memory, a processor, and a tin-based material composition performance inference program based on the Bayesian network stored on the memory and executable on the processor, and the tin-based material composition performance inference program based on the Bayesian network is configured to implement the steps of the tin-based material composition performance inference method based on the Bayesian network as claimed in any one of claims 1 to 6.

8. A readable storage medium, characterized by, The tin-based material composition performance inference program based on the Bayesian network is stored on the readable storage medium, and when executed by the processor, implements the steps of the tin-based material composition performance inference method based on the Bayesian network as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-field data analysis method based on Bayesian information enhanced neural network

    CN120277515A