Quantification and grading method of material performance sensitivity based on multi-dimensional statistical analysis
By using multi-dimensional statistical analysis methods to quantify the sensitivity of material properties, the problems of strong subjectivity and poor versatility in existing technologies are solved. This enables efficient and automated material property assessment and knowledge base generation, supporting process optimization and quality control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-10
AI Technical Summary
In current materials research and development, performance sensitivity assessment relies on expert experience, is highly subjective, lacks unified standards, is difficult to quantify systematically, fails to fully utilize historical data, has a low degree of automation, poor versatility, and is difficult to adapt to the differences in characteristics of different material systems.
A multi-dimensional statistical analysis method is adopted to quantify the sensitivity of material performance by fusing multiple statistical characteristic values such as compositional variability coefficient, coefficient of variation, quantile range ratio, absolute value of skewness, absolute value of kurtosis, and proportion of outliers. The compositional variability coefficient is then used for correction, and finally a structured knowledge base is generated.
It enables calculable, reproducible, and quantitative assessment of material performance sensitivity, eliminates subjective bias, improves assessment efficiency, adapts to different material systems, and generates a reliable knowledge base to support process optimization and quality control.
Smart Images

Figure CN121565343B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of material informatics and data science, in particular to a method for quantifying and grading material performance sensitivity based on multi-dimensional statistical analysis. BACKGROUND
[0002] Materials are the core foundation of industrial manufacturing, new energy, electronic information and other fields, and their performance directly determines the function realization, reliability and service life of products. Performance sensitivity, as a key indicator of measuring the change of material performance with composition or process parameters, is the core basis for resource allocation, process optimization and quality control in material research and development. In the process of material research and development, accurate evaluation of the sensitivity of each performance indicator is of great significance to material design, process optimization and quality control. Performance sensitivity reflects the sensitivity of material performance indicators to changes in composition, process and other factors, and its evaluation results directly guide the allocation of research resources and the development of optimization strategies.
[0003] However, the existing evaluation of performance sensitivity in material research and development has the following significant problems:
[0004] Performance sensitivity evaluation is highly dependent on expert experience and is highly subjective, and the results of different experts often differ, lacking a unified standard;
[0005] There is a lack of systematic and quantitative evaluation methods, and existing technologies rely on a single statistical quantity (such as range, standard deviation) for judgment, making it difficult to fully characterize the volatility characteristics of performance data;
[0006] The accumulated historical experimental data is not fully utilized, and a large amount of valuable performance fluctuation information is wasted, resulting in a lack of representativeness of the evaluation results;
[0007] The evaluation process is low in automation and efficiency, and is difficult to meet the needs of high-throughput experiments and big data era of material research and development;
[0008] Existing methods are highly targeted and difficult to adapt to the differences in characteristics of different material systems, with poor universality. SUMMARY
[0009] The purpose of the present application is to provide a universal, efficient and quantitative method for evaluating material performance sensitivity, which converts qualitative experience judgment into calculable and reproducible quantitative results, meeting the urgent needs of the material research and development field.
[0010] To achieve the above purpose, the present application provides the following technical solution: a method for quantifying and grading material performance sensitivity based on multi-dimensional statistical analysis, comprising the following steps:
[0011] S1, data input and performance index identification: input a historical data set containing material chemical formula and at least one column of performance index data, automatically identify and map each performance index to a standardized identifier through a pre-defined pattern matching rule library;
[0012] S2, composition variability coefficient calculation: calculate the composition diversity index and element complexity index based on the material chemical formula, and weight and fuse them into a composition variability coefficient;
[0013] S3, multi-dimensional statistical feature extraction: for each identified performance index data column, calculate five statistical feature values of coefficient of variation, quantile range ratio, absolute value of skewness, absolute value of kurtosis and proportion of outliers;
[0014] S4, sensitivity score synthesis and correction: after standardizing the five statistical feature values, weighted sum is performed according to a pre-defined weight vector to obtain a basic statistical sensitivity score; the basic statistical sensitivity score is corrected by the composition variability coefficient to obtain a final sensitivity score;
[0015] S5, grading determination and knowledge output: according to the pre-set threshold interval where the final sensitivity score is located, the sensitivity level of the corresponding performance index is determined, and a structured knowledge base containing the performance index name, typical value range and sensitivity level is automatically generated.
[0016] Preferably, the historical data set of performance index data refers to a past accumulated data set containing material chemical formula and at least one column of quantifiable performance index, and the performance index is a quantifiable parameter representing the core function or characteristics of the material, specifically including thermoelectric figure of merit, Seebeck coefficient, power factor and thermal conductivity.
[0017] Preferably, the composition variability coefficient calculation step in step S2 is:
[0018] Calculate the composition diversity index DI:
[0019] Parse each non-repeated chemical formula in the data set into an element presence vector v, which is a 118-dimensional binary vector (corresponding to the first 118 elements in the periodic table), and the value of the i-th dimension is 1 if the chemical formula contains the i-th element, otherwise 0;
[0020] For all N non-repeated chemical formulas in the data set, calculate the Jaccard distance between the element presence vectors of any two chemical formulas and The definition of the Jaccard distance is: where is the number of dimensions that are 1 in both vectors (i.e., the number of common elements). It is the number of dimensions (i.e., the total number of unique element types) of two vectors, where at least one of them is 1.
[0021] Calculate all The arithmetic mean of the Jaccard distances between non-repeating chemical formulas , the average value It is directly used as the diversity index (DI).
[0022] Calculate the element complexity index CI:
[0023] For each chemical formula in the dataset, calculate its weighted element-wise complexity. Define the complexity weight of each chemical element as one-tenth of its atomic number Z. If a chemical formula contains m unique elements, the set of elements is {E1, E2, ..., E...}. m The corresponding atomic numbers are {Z1, Z2, ..., Z}. m The complexity of this chemical formula is... ;
[0024] The complexity of calculating the total number of chemical formulas (or all unique chemical formulas) in the dataset. arithmetic mean ;
[0025] Set the upper limit of complexity C max The C max Based on the statistical characteristics of the dataset, the complexity of the chemical formula is determined. Adding a certain number of standard deviations to the higher quantiles or mean of the distribution will increase the average complexity. Normalizing to the [0,1] interval yields the element complexity index. ;
[0026] The composition diversity index DI and the element complexity index CI are linearly combined according to preset weights α and β to obtain the composition variability coefficient CVC, that is, CVC=α×DI+β×CI, and α+β=1.
[0027] Preferably, in step S3:
[0028] coefficient of variation ,in Standard deviation, The average value measures the normalized dispersion of the performance value relative to its average level.
[0029] Quantile range ratio ,in It is the 25th percentile. It is the 75th percentile. The maximum value of the sample. For the sample minimum, evaluate the concentration of the data body, effectively resist extreme value interference;
[0030] Skewness absolute value Wherein The absolute value of skewness is the absolute value of the skewness of the data distribution, which reflects the asymmetry of the data distribution. The absolute value of skewness is the absolute value of the skewness of the data distribution, which reflects the asymmetry of the data distribution. The absolute value of skewness is the absolute value of the skewness of the data distribution, which reflects the asymmetry of the data distribution. The absolute value of skewness is the absolute value of the skewness of the data distribution, which reflects the asymmetry of the data distribution. The absolute value of skewness is the absolute value of the skewness of the data distribution, which reflects the asymmetry of the data distribution.
[0031] The above reflects the symmetry of the data distribution, and the absolute value is because of the concern about the degree of deviation from the normal distribution rather than the direction.
[0032] Kurtosis absolute value Excess kurtosis, reflects the sharpness of the data distribution, and the absolute value is concerned about the degree of deviation from the normal distribution.
[0033] The proportion of outliers ;
[0034] Wherein , , The proportion of outliers is calculated as the proportion of outliers in the total data, and a high proportion of outliers is a strong indicator that the performance indicator is highly sensitive to uncontrollable factors.
[0035] Preferably, the function of the characteristic value standardized in step S4 is specifically:
[0036] Coefficient of variation CV norm =min(CV×s cv ,1.0), wherein s cv is a scaling factor;
[0037] Interquartile range ratio IQR norm =1-IQRRatio;
[0038] Outlier ratio OR norm =min(OR×s or ,1.0), wherein s or is a scaling factor;
[0039] Skewness absolute value Skew norm =min(|Skewness| / t skew ,1.0), wherein t skew is a threshold value;
[0040] Kurtosis absolute value Kurt norm= min(|Kurtosis-3| / t kurt , 1.0), where t kurt is a threshold value.
[0041] Preferably, the weight vector W in the step S4 is:
[0042] W = [w cv , w iqr , w or , w skew , w kurt ] = [0.30, 0.25, 0.20, 0.15, 0.10].
[0043] Preferably, the correction step in the step S4 is specifically:
[0044] FS = BS x (k1 + k2 x CVC);
[0045] where FS represents the final sensitivity score, BS represents the basic statistical sensitivity score, CVC represents the coefficient of variation of composition, k1 and k2 are preset constants, and k1 + k2 = 1.
[0046] Preferably, the sensitivity level in the step S5 includes: high sensitivity, medium sensitivity and low sensitivity.
[0047] The preset threshold is: high sensitivity >= 0.7, medium sensitivity 0.4~0.7, and low sensitivity < 0.4.
[0048] Compared with the prior art, the beneficial effects of the present application are:
[0049] The present application realizes significant beneficial effects through a complete technical system of performance index adaptive identification, multi-dimensional statistical feature fusion quantization, coefficient of variation of composition correction and automatic grading and knowledge base construction: qualitative experience judgment is converted into calculable and reproducible quantitative evaluation, eliminating subjective bias; a comprehensive analysis of five complementary statistical dimensions is used to replace a single statistical quantity, forming a systematic quantitative method; the value of historical data is fully tapped to avoid data idling, and the representativeness of the results is improved through data-driven analysis; the evaluation cycle is shortened to minutes through an automatic process, greatly improving the efficiency of material design optimization; with the help of the coefficient of variation of composition correction mechanism, the composition complexity difference of different material systems is adapted to ensure the fairness and accuracy of cross-system evaluation, and the generated structured knowledge base can directly support downstream automatic design and provide reliable data support for process optimization and quality control. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 is a schematic diagram of the overall architecture of the material performance sensitivity quantification and grading method based on multi-dimensional statistical analysis of the present application;
[0051] Figure 2 The sensitivity analysis process schematic diagram of the material performance sensitivity quantification and grading method based on multi-dimensional statistical analysis of the application;
[0052] Figure 3 The multi-dimensional sensitivity evaluation model schematic diagram of the material performance sensitivity quantification and grading method based on multi-dimensional statistical analysis of the application;
[0053] Figure 4 The technical scheme application scene schematic diagram of the material performance sensitivity quantification and grading method based on multi-dimensional statistical analysis of the application. DETAILED DESCRIPTION
[0054] The technical scheme in the embodiments of the application will be apparently and completely described in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the application.
[0055] Please refer to Figures 1-4 The application provides a technical scheme: a material performance sensitivity quantification and grading method based on multi-dimensional statistical analysis, comprising the following steps:
[0056] S1, data input and performance index identification: input a historical data set containing a material chemical formula and at least one column of performance index data, the historical data set of performance index data refers to a past accumulated data set containing a material chemical formula and at least one column of quantifiable performance index, the performance index is a quantifiable parameter representing the core function or characteristic of the material, specifically including thermoelectric merit value, Seebeck coefficient, power factor and thermal conductivity, after the historical data set is input, data preprocessing is first performed through a preset abnormal value filtering rule (such as deleting data exceeding 3 times of the standard deviation, filling random missing values with an arithmetic mean, and filling continuous missing values with a linear interpolation method) to ensure the validity of the performance index data, through a pre-defined pattern matching rule library, each performance index is automatically identified and mapped to a standardized identifier, the rule library includes the standard name, common abbreviation, English alias and regular expression template of the performance index, the matching algorithm adopts fuzzy matching with a string similarity threshold ≥ 85%, specifically, string fuzzy matching and regular expression scanning technology are adopted, the column name of the input data set is automatically traversed, and a non-standard column name is accurately mapped to the standardized performance index identifier defined in the system. This process does not require manual intervention, and realizes a unified entrance for cross-data set analysis;
[0057] S2, Composition Variability Coefficient Calculation: Based on the chemical formula, calculate the composition diversity index and element complexity index, and fuse them into the composition variability coefficient. The calculation steps of the composition variability coefficient (CVC) are as follows:
[0058] Calculate the composition diversity index DI:
[0059] Vector representation: Analyze each non-repeated chemical formula in the data set into a mathematical representation of chemical composition. A preferred embodiment is to convert it into an element presence vector v. Assuming that the first 118 elements in the periodic table are considered, v is a 118-dimensional binary vector, where the value of the i-th dimension is 1 if the chemical formula contains the i-th element, otherwise 0.
[0060] Difference calculation: For all N non-repeated chemical formulas in the data set, calculate the composition difference between each pair. An effective method is to calculate the Jaccard distance. For two chemical formula pairs corresponding to element presence vectors and , the Jaccard distance is defined as:
[0061] ;
[0062] where is the number of dimensions that are both 1 in the two vectors (i.e., the number of common elements), is the number of dimensions that are at least one in the two vectors (i.e., the total number of non-repeated elements). The value range of is [0, 1], and the larger the value, the greater the difference in element composition between the two chemical formulas.
[0063] Average distance: Calculate the arithmetic mean of the Jaccard distances between all pairs of non-repeated chemical formulas, denoted as .
[0064] Index generation: Since is already in the interval [0, 1], it can be directly used as the composition diversity index DI, that is:
[0065] DI ;
[0066] DI approaches 1 indicates that the material composition explores widely and the difference between chemical formulas is significant. The index abandons the simple "proportion of non-repeating chemical formulas" calculation and instead uses a difference measurement based on the element composition vector. Specifically, each chemical formula is parsed into a normalized element presence vector, and the average difference between all pairs of non-repeating chemical formulas in the dataset is calculated, i.e., as the Jaccard distance, to accurately measure the breadth and heterogeneity of the composition space exploration.
[0067] Calculate the element complexity index CI:
[0068] Single chemical formula complexity: For each valid chemical formula in the dataset, calculate its weighted element complexity. One specific implementation is to define the complexity weight of each chemical element as one-tenth of its atomic number Z. For a chemical formula, let its contain m non-repeating elements, the element set is {E1, E2,..., Em}, and the corresponding atomic numbers are {Z1, Z2,..., Zm}, then the complexity of the chemical formula is: m} is: m
[0069] ;
[0070] Parameter explanation: The above uses one-tenth of the atomic number as the weight, which is a preferred and easy-to-implement linear weighting scheme. The element complexity weight function is not limited to this. For example, the atomic number itself, its logarithmic value, or other monotonic increasing functions based on the periodic table of elements (such as electronegativity, atomic radius, common oxidation states, etc.) can be used to define the weight, the core is to reflect the relative contribution trend of element species to the intrinsic complexity of the material system. The scaling factor of 10 is to make the calculation result in a convenient numerical range, and the subsequent normalized complexity upper limit C max Co-design to ensure that the final element complexity index CI can fully utilize the interval range of [0, 1], maintaining good sensitivity and discrimination for different complexity material systems. It can be adjusted according to the data size of specific applications and the typical element composition of different material fields, but its core role is in numerical scale optimization, which does not affect the relative contribution ranking of each element to complexity.
[0071] Average complexity: Calculate the arithmetic mean of the complexity of all chemical formulas (or all non-repeating chemical formulas) in the dataset, denoted as .
[0072] Index generation: To normalize the average complexity to the interval [0, 1], set a reasonable "complexity upper limit" C max The upper limit can be set based on domain knowledge, for example, for a complex compound system containing multiple heavy elements, C max = 40. Then the element complexity index CI is calculated as follows:
[0073] ;
[0074] The index not only considers the number of element types, but also introduces a weighted complexity based on the intrinsic properties of elements (such as atomic number). By calculating the weighted complexity of each element in each chemical formula and taking the average of the dataset, the inherent complexity of the element composition of the material system is quantified.
[0075] Parameter description:
[0076] The upper limit of the normalized complexity C max is used to map the average complexity to the interval [0, 1] so that the element complexity index CI can effectively distinguish materials of different complexity.
[0077] Setting principle: The value of C max should be determined based on the statistical characteristics of the target dataset in a data-driven manner to ensure that the CI index has good discrimination.
[0078] Preferred method: By analyzing the complexity distribution of all chemical formulas in the dataset, select a statistical quantity representing the "typical complexity upper limit" as the benchmark, for example:
[0079] 1) The high quantile of the distribution (such as the 95th, 99th quantile);
[0080] 2) or the average value plus several times the standard deviation.
[0081] Example value: In this example, the 99th quantile of the dataset , is about 36.2. To balance discrimination and calculation margin, set C max = 40, and CI = 0.55 is calculated. It has been verified that this value can effectively distinguish CI values of simple to complex compounds.
[0082] Adjustability statement: C max can be adjusted according to the data characteristics of specific datasets or material fields. For most inorganic functional materials, values in the range of 30 to 50 can achieve good results. Any method that determines the normalization benchmark based on data distribution to quantify complexity is within the scope of the present application.
[0083] Synthetic composition variability coefficient CVC:
[0084] The calculated DI and CI are fused by linear weighting to generate a composition variability coefficient CVC:
[0085] CVC = a x DI + b x CI;
[0086] wherein a and b are preset weights, a e [0.5, 0.7], b e [0.3, 0.5], and a + b = 1; wherein a = 0.6, b = 0.4 for a metal material system, a = 0.55, b = 0.45 for a ceramic material system, and a = 0.65, b = 0.35 for a polymer material system. The weights reflect the judgment of the relative importance of diversity and complexity in constituting the overall variability, and the composition variability coefficient linearly weights the above more representative diversity index DI and complexity index CI according to the preset ratio (for example, 6:4) to generate a composition variability coefficient CVC between 0 and 1. The coefficient more scientifically quantifies the inherent fluctuation background of the material database in the composition space.
[0087] Weight source and optimization logic: based on prior knowledge in the field, in material research and development, exploring a variety of different chemical formulas (high diversity) usually introduces more extensive performance fluctuation sources than studying a small number of complex chemical formulas. Therefore, a slightly higher weight can be given to the diversity index, and in one metal material embodiment, a = 0.6, b = 0.4, and the weight setting basis is: in material research and development, exploring a variety of different chemical formulas (diversity) usually introduces more extensive performance fluctuations than studying a small number of complex chemical formulas, and this weight can be verified and fine-tuned through historical data review analysis.
[0088] S3, multi-dimensional statistical feature extraction: for each identified performance index data column, five statistical feature values of coefficient of variation, quantile range ratio, absolute value of skewness, absolute value of kurtosis, and proportion of outliers are calculated, the quantile is calculated by linear interpolation method, and the skewness and kurtosis are sample statistics, wherein:
[0089] coefficient of variation , wherein is the standard deviation, is the average value, which measures the normalized dispersion degree of the performance value relative to its average level;
[0090] quantile range ratio , wherein is the 25th percentile, is the 75th percentile, is the maximum value of the sample, is the minimum value of the sample, which evaluates the concentration of the data subject and effectively resists extreme value interference;
[0091] absolute value of skewness , wherein For skewness, For kurtosis, For skewness, For kurtosis of 3 / 2 power, used to eliminate dimension effect, make skewness a dimensionless index, facilitate comparison of different performance indicators;
[0092] The above reflects the symmetry of data distribution, and the absolute value is because the degree of deviation from the normal distribution is concerned, not the direction;
[0093] The absolute value of kurtosis (excess kurtosis) reflects the sharpness of data distribution, and the absolute value is concerned about the degree of deviation from the normal distribution;
[0094] The proportion of outliers ;
[0095] Among them , , , the proportion of outliers in the total data is calculated, and a high outlier ratio is a strong indicator that the performance indicator is highly sensitive to uncontrollable factors.
[0096] S4, sensitivity score synthesis and correction: map the five statistical characteristic values to the [0, 1] interval through the pre-set standardization function, and perform weighted summation according to the pre-set weight vector W to obtain the basic statistical sensitivity score BS; The variation coefficient is used to modify the basic statistical sensitivity score to obtain the final sensitivity score;
[0097] The function of standardizing the characteristic values is as follows:
[0098] The coefficient of variation CV norm =min(CV×s cv ,1.0), wherein s cv is a scaling factor, and in an embodiment, s cv =2.0. Meaning: amplification effect, assuming CV>0.5 represents extremely high dispersion, belongs to extremely high volatility, and is mapped to 1.0;
[0099] The interquartile range ratio IQR norm =1-IQRRatio, because IQRRatio is small, indicating that the data set has small volatility; after taking the inverse, the larger the value, the more dispersed, the greater the volatility;
[0100] The proportion of outliers OR norm =min(OR×s or ,1.0), wherein s or is a scaling factor, and in an embodiment, sor =5.0. Meaning: amplify the impact of outliers, because even a small number of outliers often means a high sensitivity interval;
[0101] Skewness norm = min(|Skewness| / t skew ,1.0), where t skew is a threshold, in one embodiment, t skew = 2.0, skewness beyond this threshold is considered significant;
[0102] Kurtosis norm = min(|Kurtosis-3| / t kurt ,1.0), where t kurt is a threshold, in one embodiment, t kurt = 5.0, the absolute deviation from the normal distribution (kurtosis = 3) is calculated;
[0103] The above parameter values are verified by historical data backtracking to ensure that the standardized results within the [0, 1] interval effectively reflect the performance fluctuation characteristics;
[0104] The weight vector W is:
[0105] W = [w cv , w iqr , w or , w skew , w kurt ] = [0.30, 0.25, 0.20, 0.15, 0.10];
[0106] Parameter source and design logic:
[0107] 1) Coefficient of variation (w cv = 0.3): the highest weight, because it is the most classic standard for measuring relative dispersion, and most directly reflects the fluctuation amplitude of performance values.
[0108] 2) Quantile range ratio (w iqr = 0.25): the second highest weight, because it is not sensitive to outliers and can stably reflect the concentration of the data body, which is an important complement to CV.
[0109] 3) Outlier ratio (w or = 0.2): a high outlier ratio is a strong signal of sensitivity, and is given a higher weight.
[0110] 4) Skewness and kurtosis (w skew = 0.15, w kurtt=0.10): describes the shape of the distribution, can reveal non-linear response or sensitive interval, but the interpretation is relatively indirect, and is given a lower weight.
[0111] 5) Normalization: the sum of all weights is 1, which ensures that the range of the basic sensitivity score (BS) is controllable;
[0112] The above weight system is optimized based on domain knowledge, the preliminary weights are determined through expert investigation, and the fine tuning is carried out through historical case backtracking, reflecting the priority of the contribution of different statistical characteristics to sensitivity, and the weight value is determined based on the priority of the contribution of statistical characteristics to sensitivity evaluation, among which the coefficient of variation weight is the highest, because it most directly reflects the performance fluctuation amplitude, and the skewness and kurtosis weights are lower, because their interpretation is relatively indirect.
[0113] The correction step is specifically:
[0114] FS=BS×(k1+k2×CVC);
[0115] Where FS represents the final sensitivity score, BS represents the basic statistical sensitivity score, CVC represents the coefficient of variation, k1 and k2 are preset constants, and k1+k2=1, if it is a multi-component material system, the correction factor k1=0.7, k2=0.3; if it is a binary compound material system, k1=0.8, k2=0.2, in a multi-component material system embodiment, k1=0.7, k2=0.3, k1 and k2 are set by empirical values to ensure that the correction amplitude is moderate. The consistency between the scores before and after correction and expert judgment can be analyzed to optimize;
[0116] The design logic of the correction model is:
[0117] When CVC is low (simple composition), the correction factor (0.7+0.3*CVC) is less than 1, the score is adjusted lower to avoid overestimating the fluctuation. Reason: The fluctuation observed in a simple system may be overestimated (because there are fewer sources of change), and it needs to be carefully adjusted downward;
[0118] When CVC is high (complex composition), the correction factor (0.7+0.3*CVC) is close to 1, the original statistical fluctuation score is accepted, ensuring the fairness of the evaluation results between different material systems, and the reason is: In a complex system, high fluctuation is likely to be caused by its inherent sensitivity, and should be accepted.
[0119] S5, grading determination and knowledge output: according to the preset threshold interval of the final sensitivity score, the sensitivity level of the corresponding performance indicator is determined, the analysis results of all performance indicators are summarized, and a structured knowledge base containing the performance indicator name, typical value range and sensitivity level is automatically generated, and the sensitivity level includes: high sensitivity, medium sensitivity and low sensitivity;
[0120] The preset threshold is: high sensitivity ≥ 0.7, medium sensitivity 0.4-0.7, and low sensitivity < 0.4.
[0121] A material performance sensitivity quantification and grading system based on multi-dimensional statistical analysis, characterized by comprising:
[0122] A performance index identification module for receiving a data set containing material chemical formula and performance index data, and automatically identifying and mapping performance indexes based on a predefined rule base;
[0123] A composition variability calculation module for calculating composition variability coefficients based on material chemical formula. To address the industry pain points of inconsistent and diverse column names in material databases, this module has an extensible performance index name pattern rule base. This base not only contains standard names (such as "ZT" and "Seebeck"), but also covers common abbreviations, aliases, and English and Chinese variants (such as "thermoelectric figure of merit" and "figure of merit"). This module uses string fuzzy matching and regular expression scanning technology to automatically traverse the column names of the input data set, accurately mapping non-standard column names to the standardized performance index identifiers defined in the system. This process does not require human intervention, achieving a unified entry for cross-data set analysis;
[0124] A multi-dimensional statistical feature extraction module that creatively constructs and integrates five core statistical dimensions to comprehensively evaluate the volatility of performance data:
[0125] Coefficient of variation: measures the normalized dispersion of performance values relative to their average level.
[0126] Quartile range ratio: by calculating the ratio of the inner quartile range to the full range, it evaluates the concentration of the data body, effectively resisting the interference of extreme values.
[0127] Skewness and kurtosis: joint analysis of the symmetry and sharpness of data distribution. Specific skewness or kurtosis may reveal the sensitive interval of performance to process parameters or the existence of nonlinear response.
[0128] Proportion of outliers: based on statistical criteria (such as 1.5 times the interquartile range method) to identify the proportion of outliers in the data. A high proportion of outliers is a strong indicator that the performance index is highly sensitive to certain uncontrolled factors;
[0129] Assign fixed weights based on domain knowledge optimization to the above five dimensions (for example: coefficient of variation: 0.30, quartile range ratio: 0.25, outlier proportion: 0.20, skewness: 0.15, and kurtosis: 0.10), and combine them into a unified "basic statistical sensitivity score" through weighted summation algorithm;
[0130] The above modules solve the problem that the traditional method for evaluating performance sensitivity usually relies on a single statistical quantity (such as a range or a standard deviation), and cannot comprehensively and robustly characterize the fluctuation characteristics of performance data.
[0131] The sensitivity score synthesis and correction module proposes a quantitative index of "component variability coefficient" for normalizing and weighted summing of the five statistical characteristic values to obtain a basic statistical sensitivity score, and uses the component variability coefficient to correct the basic statistical sensitivity score to obtain a final sensitivity score. The calculation method includes:
[0132] Component diversity analysis: Calculate the ratio of the number of different chemical formulas in the historical data set to the total number of chemical formulas to reflect the breadth of component exploration.
[0133] Element complexity analysis: Analyze each chemical formula, count the number of chemical elements contained in each chemical formula, calculate the average value and normalize it to reflect the complexity of the composition.
[0134] Coefficient synthesis: Linearly weight and fuse the diversity index and the complexity index according to a predetermined ratio (for example, 6:4) to generate a component variability coefficient between 0 and 1. This coefficient quantifies the inherent fluctuation background of the material database in the composition space;
[0135] The above modules solve the problem that the existing technology ignores the complexity of the material system itself, which can cause deviation when comparing sensitivity between data sets of different composition ranges;
[0136] The grading and knowledge base generation module is used to determine the sensitivity level according to the final sensitivity score and output a structured knowledge base containing the performance index name, typical value range and sensitivity level. Specifically:
[0137] Sensitivity score correction model: The "basic statistical sensitivity score" and the "component variability coefficient" are coupled through a specific mathematical relationship (for example: final sensitivity score = basic statistical sensitivity score x (0.7 + 0.3 x component variability coefficient)) to generate a final sensitivity score. This model ensures that high fluctuation performance is not underestimated in complex composition systems, and low fluctuation performance is not overestimated in simple composition systems.
[0138] Dynamic threshold grading: According to the final sensitivity score calculated by coupling, set a clear numerical threshold range (for example: high sensitivity ≥ 0.7, medium sensitivity 0.4-0.7, low sensitivity < 0.4) to achieve objective and consistent grading determination.
[0139] Automated knowledge base construction: the system traverses all performance metrics and automatically generates a structured record for each metric containing "typical value range" (based on historical data statistics) and "sensitivity level". This knowledge base can be used as a standardized output and directly embedded into the materials informatics pipeline;
[0140] The above modules provide quantitative basis for qualitative classification (high / medium / low), and the analysis results can be directly used in the downstream automated material design process.
[0141] Example 1
[0142] This example analyzes the "Thermoelectric Materials Database (Thermoelectric_Materials_DB.xlsx)".
[0143] Data preparation: the database contains 2000 material records, and the main fields include: Formula (chemical formula), ZT_avg (average thermoelectric figure of merit), Seebeck_coeff (Seebeck coefficient, μV / K), Power_Factor (power factor, mW / m·K²), Thermal_Conductivity (thermal conductivity, W / m·K).
[0144] Step S1: data input and feature recognition
[0145] a. The system loads the Excel file.
[0146] b. The built-in rule base successfully identifies: ZT_avg -> "ZT", Seebeck_coeff -> "Seebeck", Power_Factor -> "PF", Thermal_Conductivity fails to completely match the typical pattern and is temporarily excluded from the core analysis.
[0147] Step S2: Calculate the composition variability coefficient (CVC)
[0148] a. Calculate the composition diversity index (DI): parse the 43761 non-repeating chemical formulas (Formula) in the data set and convert them into 118-dimensional element presence (0 / 1) vectors. Calculate the Jaccard distance between all pairs of non-repeating chemical formulas (about 43761x43760 / 2 pairs) and find their average . After calculation, =0.88, so DI=0.88. This indicates that in the thermoelectric material database, any two chemical formulas have an average of only 12% of the element species in common, and the composition difference is extremely large and highly diverse.
[0149] b. Calculate the element complexity index (CI): For each chemical formula in the dataset, calculate one-tenth of the sum of the atomic numbers of its non-repeating elements, and use this as the complexity Cf of that chemical formula. For example, the complexity of the chemical formula Bi₂Te₃ is (83+52) / 10 = 13.5; The complexity is (47+82+51+52) / 10=23.2. Calculate the average complexity of all chemical formulas, and obtain... =22.0. Set the upper limit of normalization C. max =40 (referencing an extremely complex multi-element thermoelectric system), then CI = 22.0 / 40 = 0.55. This reflects that the compounds in this material library have an average intrinsic elemental complexity of moderate, which is consistent with the reality that most are ternary, quaternary solid solutions or doped systems.
[0150] c. Coefficient of Variation of Composite Composition (CVC): Taking weights α=0.6 and β=0.4, calculate CVC:
[0151] CVC=α×DI+β×CI=0.6×0.88+0.4×0.55=0.528+0.220=0.748
[0152] The value CVC=0.748 precisely indicates that this thermoelectric material database has a very rich compositional variability.
[0153] Steps S3 and S4: Multidimensional statistics and score calculation (taking ZT as an example)
[0154] a. The ZT column contains 1950 valid data entries, from which the following calculations are performed:
[0155] CV=0.82, IQRRatio=0.25, |Skewness|=1.5, |Kurtosis|=5.0, OutlierRatio=0.18.
[0156] b. Standardization (e.g., CV) norm =min(0.82*2,1)=1.0;IQR norm =1-0.25=0.75; Outlier norm =min(0.18*5,1)=0.9; Skew norm =min(1.5 / 2,1)=0.75; Kurt norm =min(|5.0-3| / 5,1)=0.4).
[0157] c. Weighted summation yields BS_ZT = 1.0 * 0.30 + 0.75 * 0.25 + 0.9 * 0.20 + 0.75 * 0.15 + 0.4 * 0.10 = 0.7975.
[0158] d. Correction: FS_ZT = 0.7975 * (0.7 + 0.3 * 0.748) ≈ 0.7975 * 0.9244 ≈ 0.737.
[0159] Step S5: Grading and output
[0160] a. FS_ZT = 0.737 ≥ 0.7, determine ZT as a high sensitivity index.
[0161] b. Similarly, assume that the calculated FS_S of the Seebeck coefficient is 0.65, which is determined as medium sensitivity; and the FS_PF of the PF is 0.82, which is determined as high sensitivity.
[0162] c. The system finally outputs the performance mode configuration dictionary as follows:
[0163] performance_patterns={
[0164] 'ZT':{'typical_range':(0.01,2.5),'sensitivity':'high','score':0.737},
[0165] 'Seebeck':{'typical_range':(-200,800),'sensitivity':'medium','score':0.650},
[0166] 'PF':{'typical_range':(1e-5,0.01),'sensitivity':'high','score':0.820},
[0167] }
[0168] d. Engineering significance: The results clearly inform material researchers that when optimizing thermoelectric materials, ZT and PF are highly sensitive performance indicators, and their values can easily change greatly with small adjustments in composition or process, so they need to be given the highest priority and more stringent monitoring in experimental design and process control. The Seebeck coefficient is relatively stable, and the optimization strategy can be different. This quantitative conclusion provides direct and reliable data support for resource allocation and research and development strategies.
[0169] The innovation of the present application lies in proposing a set of general data models for automatically quantifying the inherent volatility (i.e. sensitivity) of performance indicators from historical material data, the core of which is the creative integration of multi-dimensional statistical characteristics and the collaborative correction of material intrinsic complexity. The specific innovations are as follows:
[0170] Five statistically complementary descriptive features—coefficient of variation (reflecting relative dispersion), quantile range ratio (reflecting data concentration, anti-outlier), outlier ratio (directly indicating sensitive interval), skewness and kurtosis (reflecting distribution shape asymmetry and sharpness)—are systematically combined for the first time in material performance sensitivity analysis. By designing specific, non-obvious standardized mapping functions for the above five features (such as taking the opposite of IQRRatio, and amplifying and scaling the outlier ratio), and weighting them with a fixed weight vector optimized based on domain knowledge (for example: [0.30, 0.25, 0.20, 0.15, 0.10]), a unified "basic statistical sensitivity score (BS)" ranging from [0, 1] is generated. The weight system reflects the priority judgment of the contribution of different statistical features to "sensitivity", and is one of the key empirical parameters of the invention.
[0171] The above technical means overcomes the one-sidedness and subjectivity of traditional methods that rely on a single statistical quantity (such as range) or expert experience to evaluate performance fluctuations.
[0172] The invention proposes the quantitative indicator of "composition coefficient of variation (CVC)", which eliminates the bias caused by the difference in system composition complexity when directly comparing performance sensitivity scores between different material systems (such as simple binary compound library vs. complex high-entropy alloy library). This indicator is composed of two more refined sub-indicators: a composition diversity index (DI) based on the average Jaccard distance between chemical formulas, and an element complexity index CI based on the average atomic number weighted complexity. The two are fused in a preset ratio (such as 6:4) to obtain CVC, which more scientifically quantifies the inherent fluctuation background of the data set in the composition space. CVC is coupled to the basic sensitivity score BS through a linear correction factor (k1+k2xCVC) (such as 0.7+0.3xCVC) to obtain the final sensitivity score FS. This correction model has a clear physical meaning: when the material composition is single (CVC is low), the score is corrected downward to avoid underestimating the observed fluctuations within a limited range; when the composition is complex (CVC is high), the original statistical fluctuation score is preferred. This makes the sensitivity evaluation results more fair and comparable between different material databases;
[0173] The above technical means eliminates the bias caused by the difference in system composition complexity when directly comparing performance sensitivity scores between different material systems (such as simple binary compound library vs. complex high-entropy alloy library).
[0174] The present application is based on the final sensitivity score FS, and sets a clear numerical threshold (such as high >=0.7, medium [0.4,0.7), low <0.4) for automatic classification. By analyzing all performance indicators in the data set, a structured performance sensitivity knowledge base is automatically output. Each record contains "performance indicator name", "typical value range based on historical statistics" and "quantitative sensitivity level". The knowledge base can be used as a standard output, seamlessly embedded in materials informatics workflow or high-throughput experimental design system, realizing the closed loop from historical data mining to design rule generation.
[0175] The quantitative results are converted into structured knowledge that can be directly used to guide material design through the above technical means.
[0176] It should be noted that, in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.
[0177] Although the embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for quantifying and ranking material performance sensitivity based on multi-dimensional statistical analysis, characterized in that: The method comprises the following steps: S1, data input and performance index identification: input a historical data set containing a material chemical formula and at least one column of performance index data, automatically identify and map each performance index to a standardized identifier through a pre-defined pattern matching rule library; S2, composition variability coefficient calculation: calculate a composition diversity index and an element complexity index based on the material chemical formula, and weight and fuse them into a composition variability coefficient, the calculation steps of the composition variability coefficient being: calculate the composition diversity index DI: parse each non-repeated chemical formula in the data set into an element presence vector v, the vector v being a 118-dimensional binary vector; For all N non-redundant chemical formulas in the dataset, calculate the Jaccard distance between the element presence vectors of any two chemical formulas and The Jaccard distance between two vectors is defined as: where is the number of dimensions that are 1 in both vectors, is the number of dimensions that are at least 1 in either vector. Calculate all The arithmetic mean of the Jaccard distances between non-repeating chemical formulas The average value Is directly taken as the compositional diversity index DI; calculate the element complexity index CI: For each chemical formula in the dataset, calculate its weighted element complexity , define the complexity weight of each chemical element as one-tenth of its atomic number Z, if a chemical formula contains m non-repeating elements, the element set is {E1, E2,..., Em}, and the corresponding atomic numbers are {Z1, Z2,..., Zm}, then the complexity of the chemical formula is m . m ; Complexity of all chemical formulae in the data set is calculated the arithmetic mean of ; Setting a complexity upper limit C max , the C max is determined based on statistical features of the dataset, for chemical formula complexity distribution, the high quantile or the average value plus several times of the standard deviation, the average complexity is normalized to the interval [0, 1] to obtain the element complexity index ; linearly combine the composition diversity index DI and the element complexity index CI according to preset weights α and β to obtain the composition variability coefficient CVC, that is, CVC = α × DI + β × CI, and α + β = 1; S3, multi-dimensional statistical feature extraction: for each identified performance index data column, calculate five statistical feature values of a coefficient of variation, a quantile range ratio, an absolute value of skewness, an absolute value of kurtosis and an abnormal value proportion; S4, sensitivity score synthesis and correction: after standardizing the five statistical feature values respectively, weight and sum them according to a preset weight vector to obtain a basic statistical sensitivity score; correct the basic statistical sensitivity score by using the composition variability coefficient to obtain a final sensitivity score; S5, grading determination and knowledge output: according to a preset threshold interval in which the final sensitivity score is located, determine the sensitivity level of the corresponding performance index, and automatically generate a structured knowledge base containing a performance index name, a typical value range and a sensitivity level.
2. The method for quantification and ranking of material performance sensitivity based on multi-dimension statistical analysis according to claim 1, wherein: The historical data set of the performance index data refers to a past accumulated data set containing a material chemical formula and at least one column of quantifiable performance indexes, the performance index being a quantifiable parameter representing the core function or characteristics of the material, specifically including a thermoelectric merit value, a Seebeck coefficient, a power factor and a thermal conductivity.
3. The method of quantifying and ranking material performance sensitivity based on multi-dimension statistical analysis according to claim 2, wherein: In the step S3: coefficient of variation wherein is the standard deviation, is the mean value; quantile range ratio wherein Q25 is the 25th percentile, Q75 is the 75th percentile, Qmax is the sample maximum, Qmin is the sample minimum; skewness absolute value wherein is the mean deviation, is the expectation operator; kurtosis absolute value ; Outlier proportion , wherein , , .
4. The method for quantifying and ranking material performance sensitivity based on multi-dimension statistical analysis according to claim 3, wherein: The function of the feature value standardization in the step S4 is specifically: Coefficient of variation CV norm = min(CV x s cv , 1.0), where s cv is a scaling factor; Interquartile Range Ratio IQR norm = 1 - IQR Ratio; Outlier ratio OR norm = min(OR x s or , 1.0), where s or is a scaling factor; Skewness norm = min(|Skewness| / t skew ,1.0), where t skew is a threshold value; Kurtosis absolute value Kurt norm = min(|Kurtosis - 3| / t kurt , 1.0), where t kurt is a threshold value.
5. The method for quantification and ranking of material performance sensitivity based on multi-dimension statistical analysis according to claim 4, wherein: The weight vector W in the step S4 is: W=[w cv ,w iqr ,w or ,w skew ,w kurt ], where w cv w iqr w or w skew w kurt =1.
6. The method for quantification and ranking of material performance sensitivity based on multi-dimension statistical analysis according to claim 5, wherein: The correction step in the step S4 is specifically: FS = BS × (k1 + k2 × CVC); wherein FS represents the final sensitivity score, BS represents the basic statistical sensitivity score, CVC represents the composition variability coefficient, k1 and k2 are preset constants, and k1 + k2 = 1.
7. The method for quantification and ranking of material performance sensitivity based on multi-dimension statistical analysis according to claim 6, wherein: The sensitivity level in the step S5 includes high sensitivity, medium sensitivity and low sensitivity; The preset threshold is: high sensitivity ≥ 0.7, medium sensitivity 0.4-0.7, and low sensitivity < 0.4.
Citation Information
Patent Citations
Automatic targeted proteomics qualitative and quantitative analysis method
CN116153392A
Pollutant toxicity assessment method and device based on RAG enhancement and knowledge graph reasoning, equipment and medium
CN121278146A