A method, medium and system for generating ewi from bop data scraping
By constructing a multidimensional BOP data load matrix and using a multi-algorithm fusion method, the problem of insufficient accuracy in generating EWI from BOP data was solved, enabling high-precision monitoring of production quality status. The generated EWI index can accurately reflect the actual quality status of the production process.
Patent Information
- Application Number
- CN202510520821.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing technologies for generating EWI from BOP data are not accurate enough and cannot accurately reflect the quality status of the production process. In particular, it is difficult to capture the nonlinear relationship and time trend between parameters in high-dimensional data and complex scenarios, resulting in a significant deviation between the EWI index and the actual production quality status.
A multidimensional BOP data load matrix is constructed, the parameter contribution value is analyzed using the thermodynamic entropy equation, the data crawling path is optimized by combining the minimum spanning tree algorithm, the data of each dimension is weighted and fused by the Lagrange multiplier method, the horizontal and vertical errors are calculated, the deviation and change trend between parameters are evaluated by Euclidean distance and Fourier transform, and a mixed error evaluation index is generated.
It significantly improves the accuracy of EWI indicators, accurately reflects the quality status of the production process, provides a reliable basis for precise quality control and timely intervention, and enhances the accuracy and effectiveness of quality management.
Smart Images

Figure CN120672178B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of production business data, and in particular, relates to a method, medium and system for generating EWI through BOP data grabbing. BACKGROUND
[0002] In the field of industrial production, the Early Warning Indicator (EWI) system is a key tool for evaluating and controlling production quality, which is traditionally constructed by collecting and analyzing Bill of Process (BOP) data in the production process. Currently, the mainstream EWI construction method is mainly based on simple threshold monitoring and linear weighted summation, which compares the acquired production parameters with the preset standard, and triggers the early warning mechanism when the parameters deviate from the predetermined threshold. In industries such as semiconductor manufacturing, precision machining, pharmaceutical production, and chemical synthesis, which have extremely high requirements for production precision, such systems are widely used for real-time monitoring of production quality status.
[0003] However, the traditional EWI construction method faces the problem of insufficient accuracy in BOP data grabbing for generating EWI. Existing technologies usually use empirical formulas or fixed weight models to process BOP data, which cannot accurately reflect the actual contribution of each parameter to the quality status in different production stages and different working conditions. Especially when dealing with high-dimensional BOP data, the traditional linear weighting method is difficult to capture the complex nonlinear relationship between parameters, resulting in significant deviation between the generated EWI indicators and the actual production quality status. In addition, existing methods lack systematic optimization strategies in the data grabbing process, making it difficult to ensure that the most representative set of data points is obtained, affecting the accuracy of subsequent EWI indicators.
[0004] In modern high-precision manufacturing environments, the insufficient accuracy of BOP data grabbing for generating EWI has become a key bottleneck restricting the effectiveness of quality control systems. Current technologies are difficult to consider the change characteristics of BOP parameters in spatial dimensions (relationships between parameters) and time dimensions (evolution trends of parameters), lack systematic error analysis methods, and cannot accurately depict the quality status of the production process. Especially in complex scenarios where production conditions fluctuate and multiple parameters interact with each other, the accuracy problem of traditional BOP data grabbing and EWI generation methods is more prominent, and there is an urgent need for a method that can significantly improve the accuracy of BOP data grabbing for generating EWI. That is, there is a technical problem of insufficient accuracy in BOP data grabbing for generating EWI in existing technologies. SUMMARY
[0005] Therefore, the present application provides a method, medium and system for generating EWI through BOP data grabbing, which can solve the technical problem of insufficient accuracy in BOP data grabbing for generating EWI in existing technologies.
[0006] The present application is achieved: the first aspect of the present application provides a method for generating EWI by BOP data acquisition, comprising: constructing a multi-dimensional BOP data load matrix; calculating BOP data contribution value, using thermodynamic entropy equation to analyze the influence degree of each BOP data dimension; establishing EWI basic matrix; implementing BOP data acquisition rate calculation, using minimum spanning tree algorithm to obtain the optimized data acquisition path and BOP data integrity score; generating EWI mergable index, weighting the matching degree of each dimension BOP data and EWI basic matrix by Lagrange multiplier method; calculating the horizontal error, using Euclidean distance function to analyze the deviation degree between different BOP parameters at the same time point; calculating the longitudinal error, using Fourier transform equation to track the change trend of a single BOP parameter in time series; generating mixed error evaluation; according to the mixed error evaluation result and the EWI mergable index, generating a comprehensive index reflecting the production quality state as the final EWI index and outputting.
[0007] Among them, the BOP data load matrix is a multi-dimensional data structure constructed according to time and type of key parameters in production process, which is used to represent the load state and mutual influence relationship of each parameter at different time.
[0008] Among them, the BOP data contribution value is the influence weight of each BOP data in the formation process of the final EWI index, which is calculated by thermodynamic entropy equation, reflecting the influence degree of different parameters on production quality.
[0009] Among them, the EWI basic matrix is the reference standard for production process quality control, which contains the standard value and fluctuation threshold of each parameter, and is used as the benchmark point for data analysis.
[0010] Among them, the BOP data acquisition rate is the percentage of effective BOP data successfully acquired in the production monitoring process in the total BOP data that should be acquired in theory, reflecting the reliability of the data acquisition system.
[0011] Among them, the EWI mergable index is the process of integrating multi-dimensional parameter BOP data into a single index by Lagrange multiplier method, which is convenient for intuitive evaluation of overall production state.
[0012] Among them, the horizontal error is the relative deviation between different BOP parameters at the same time point, which is used to identify abnormal correlation between BOP parameters; the longitudinal error is the change deviation of a single BOP parameter in time series, which is used to monitor whether the change trend of BOP parameter meets the expectation; the mixed error is a comprehensive index obtained by fusing the horizontal error and the longitudinal error according to specific weight, which fully reflects the quality state of production process.
[0013] The mixed error evaluation is a comprehensive analysis of the lateral error and the longitudinal error, and is specifically a comprehensive quality evaluation system constructed according to a parameter distance matrix, an abnormal correlation identifier, parameter energy distribution and an abnormal frequency point, and an analysis result of the mixed error evaluation is obtained.
[0014] The second aspect of the present application provides a computer readable storage medium, the computer readable storage medium stores program instructions, the program instructions are used to execute the above-mentioned BOP data extraction method for generating EWI when running in the computer.
[0015] The third aspect of the present application provides a BOP data extraction system for generating EWI, comprising the above-mentioned computer readable storage medium, the system is any one of a computer, a server and a single-chip microcomputer, the computer readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing the program instructions stored in the computer readable storage medium.
[0016] The present application realizes high-precision monitoring of the production quality state by constructing a multi-dimensional BOP data load matrix, scientifically calculating the contribution value of each parameter by using the thermodynamic entropy equation, optimizing the data extraction path by using the minimum spanning tree algorithm, and accurately generating the comprehensive EWI index by using the Lagrange multiplier method.
[0017] The method solves the problem of insufficient accuracy caused by the random and subjective parameter weight setting in the traditional technology by using the weight distribution mechanism based on the information entropy theory. In particular, a two-dimensional analysis system of lateral error and longitudinal error is introduced, and the spatial relationship between parameters and the time evolution characteristics are accurately evaluated by using the Euclidean distance function and the Fourier transform, respectively. The limitations of traditional methods in capturing complex relationships between parameters and time trends are overcome, and the accuracy of the EWI index is significantly improved.
[0018] Through the comprehensive processing method of multi-dimension and multi-algorithm fusion, the present application effectively solves the technical problem of insufficient accuracy of BOP data extraction for generating EWI in the prior art, so that the generated EWI index can accurately reflect the actual quality state of the production process, provides a reliable basis for accurate quality control and timely intervention in the production process, and greatly improves the accuracy and effectiveness of quality management. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 The figure is a flowchart of the method of the present application.
[0020] Figure 2 The figure is a curve graph of the standardized values of the five key parameters in Example 2 over time.
[0021] Figure 3A plot of the EWI index variation calculated based on the five key parameters in Example 2.
[0022] Figure 4 A plot of the spectrum analysis result of the P4 parameter (radio frequency bias) in Example 2.
[0023] Figure 5 A plot of the lateral error, longitudinal error and their mixed error evaluation in Example 2. DETAILED DESCRIPTION
[0024] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application.
[0025] As Figure 1 shown in FIG. 1, which is a flowchart of a method for generating an EWI by BOP data grabbing provided by the first aspect of the present application, the method comprises the following steps:
[0026] S01, constructing a multi-dimensional BOP data load matrix, classifying and arranging the BOP data collected in the production process according to time series and parameter types to form a standardized multi-dimensional matrix structure;
[0027] S02, calculating the contribution value of the BOP data, analyzing the influence degree of each BOP data dimension by using the thermodynamic entropy equation, assigning a weight coefficient to each BOP data dimension based on the parameter fluctuation amplitude, fluctuation frequency, system response time, system self-organizing ability and environmental disturbance factor, and ensuring the sum of the contribution values of the BOP data in each dimension to be 1 through normalization processing;
[0028] S03, establishing an EWI basic matrix, setting a reference threshold according to historical production data, and constructing an EWI basic matrix containing standard parameters and allowable fluctuation range;
[0029] S04, implementing BOP data grabbing rate calculation, using the minimum spanning tree algorithm to monitor the ratio of the actual effective BOP data quantity grabbed in the production process to the theoretical BOP data quantity that should be grabbed based on data point distribution density, data point weight, point-to-point connection cost, network topology structure and system resource limitation, and obtaining the optimal data grabbing path and BOP data integrity score;
[0030] S05, generating an EWI mergable index, using the Lagrange multiplier method to weight the matching degree of each dimension BOP data and the EWI basic matrix based on the upper and lower limit constraints of the weight, the correlation matrix between parameters, the historical accuracy evaluation, the type of objective function and the convergence threshold, and forming a single measurement index;
[0031] S06, calculate the lateral error, use the Euclidean distance function to analyze the deviation between different BOP parameters at the same time point based on the parameter vector group, dimension weight coefficient, reference standard vector, allowable deviation range and correlation threshold, obtain the distance matrix and abnormal correlation identification between parameters;
[0032] S07, calculate the longitudinal error, use the Fourier transform equation to track the change trend of a single BOP parameter in time series based on time series parameter value, sampling interval, frequency resolution, window function type and spectral smoothing coefficient, obtain the energy distribution of the parameter at different frequencies and abnormal frequency points;
[0033] S08, generate mixed error evaluation, which is a comprehensive analysis of the lateral error and the longitudinal error, specifically, according to the distance matrix between parameters, abnormal correlation identification, parameter energy distribution and abnormal frequency points, a comprehensive quality evaluation system is constructed, and the mixed error evaluation result is obtained by analysis;
[0034] S09, output the final EWI index, that is, according to the mixed error evaluation result and the EWI mergable index, a comprehensive index reflecting the production quality state is generated.
[0035] Among them, the BOP data load matrix refers to the multi-dimensional data structure constructed according to time and type of key parameters in the production process, which is used to represent the load state and mutual influence relationship of each parameter at different time.
[0036] Among them, the BOP data contribution value refers to the influence weight of each BOP data in the formation process of the final EWI index, which is calculated by the thermodynamic entropy equation, reflecting the influence degree of different parameters on production quality.
[0037] Among them, the EWI basic matrix is the reference standard of production process quality control, which contains the standard value and fluctuation threshold of each parameter, and is used as the benchmark point for data analysis.
[0038] Among them, the BOP data capture rate refers to the percentage of effective BOP data successfully obtained in the production monitoring process in the total BOP data that should be obtained in theory, reflecting the reliability of the data acquisition system.
[0039] Among them, the EWI mergable index is the process of integrating multi-dimensional parameter BOP data into a single index by Lagrange multiplier method, which is convenient for intuitive evaluation of overall production state.
[0040] Wherein, the lateral error refers to the relative deviation between different BOP parameters at the same time point, used to identify abnormal correlation between BOP parameters; the longitudinal error refers to the deviation of a single BOP parameter in the time sequence, used to monitor whether the change trend of the BOP parameter meets the expectation; the mixed error is a comprehensive index obtained by fusing the lateral error and the longitudinal error according to specific weights, which comprehensively reflects the quality state of the production process.
[0041] Wherein, the inter-parameter distance matrix is a matrix representation of the distance between BOP parameters calculated using the Euclidean distance function, used for mixed error evaluation; the abnormal correlation identifier is an indicator marking abnormal correlation between BOP parameters, used for quality anomaly identification; the parameter energy distribution is the energy distribution of BOP parameters in the frequency domain obtained by Fourier transform, used to identify abnormal patterns; the abnormal frequency point is a frequency position in the BOP parameter spectrum that significantly differs from the normal pattern, indicating potential abnormalities.
[0042] The thermodynamic entropy equation is used to analyze the contribution degree of different BOP parameters in the production system to the stability of the system, the inputs include parameter fluctuation amplitude, fluctuation frequency, system response time, system self-organization ability, environmental disturbance factor, and the output is the entropy value contribution coefficient of each BOP parameter; wherein, the parameter fluctuation amplitude refers to the difference between the maximum and minimum values of BOP data within a certain time window, used to measure the stability of BOP data; the fluctuation frequency refers to the number of changes of BOP parameters per unit time, reflecting the sensitivity of system response; the system response time refers to the time interval required from input change to corresponding change in system output, measuring the reaction speed of the system; the system self-organization ability refers to the ability index of the production system to maintain a stable state under external disturbance, reflecting the robustness of the system; the environmental disturbance factor refers to the quantitative representation of external environmental factors affecting the accuracy of BOP data, including temperature, humidity, vibration, etc.
[0043] The Fourier transform equation is used to convert BOP parameter changes in the time domain to frequency domain analysis, identify the periodic characteristics and abnormal patterns of BOP parameter fluctuations, inputs include time series parameter values, sampling interval, frequency resolution, window function type, spectral smoothing coefficient, and the output is the energy distribution of BOP parameters at different frequencies and abnormal frequency points; wherein, the time series parameter values are a single BOP parameter value sequence arranged in chronological order, which is the input data of Fourier transform; wherein, the sampling interval is the time interval for obtaining BOP data, which affects the frequency range of Fourier transform; the frequency resolution is the minimum frequency difference that can be distinguished in Fourier transform, which affects the degree of analysis; the window function type is the function form of weighting the time series data before Fourier transform, which reduces spectral leakage; the spectral smoothing coefficient is a parameter for smoothing the Fourier transform results, which reduces the noise influence.
[0044] The minimum spanning tree algorithm is used to optimize the BOP data grabbing strategy, ensure that the most representative BOP data point set is obtained under the condition of limited resources, the input includes data point distribution density, data point weight, point connection cost, network topology structure, system resource limit, and the output is the optimized data grabbing path and BOP data integrity score; wherein the data point distribution density refers to the number distribution of BOP data points in unit space or time, which is used for weight calculation of the minimum spanning tree algorithm; the data point weight is a weighting coefficient given according to the importance of BOP data, which is used to optimize the data grabbing strategy; the point connection cost refers to the resource consumption required to connect two data points in the data grabbing network, which is used to construct the minimum spanning tree; the network topology structure refers to the connection relationship description between nodes in the data acquisition network, which provides the basis for the minimum spanning tree algorithm; the system resource limit refers to the constraint conditions of the data acquisition system in processing capacity, storage space and communication bandwidth, which affects the data grabbing strategy.
[0045] The Euclidean distance function is used to calculate the similarity and deviation degree between different dimensions in the BOP parameter space, the input includes parameter vector group, dimension weight coefficient, reference standard vector, allowable deviation range, correlation threshold, and the output is the distance matrix between parameters and abnormal association identification; wherein the parameter vector group is a vector set composed of multiple BOP parameters at the same time point, which is used for Euclidean distance function calculation; the dimension weight coefficient is the importance coefficient of different BOP parameter dimensions in the Euclidean distance function calculation, which affects the distance calculation result; the reference standard vector is the BOP parameter vector under ideal production state, which is used as the basis for calculating the deviation degree; the allowable deviation range is the maximum limit of the deviation of BOP parameters from the reference standard vector, which is used to judge the abnormal state; the correlation threshold is the critical value for judging whether the correlation between BOP parameters is abnormal, which is used to filter effective correlation.
[0046] The Lagrange multiplier method is used to solve the optimal fusion weight of EWI index under multiple constraint conditions, the input includes weight upper and lower limit constraint, parameter correlation matrix, historical accuracy evaluation, target function type, convergence threshold, and the output is the optimal weight distribution scheme that meets all the constraints; wherein the weight upper and lower limit constraint refers to the limitation of the weight value range of each BOP parameter in the Lagrange multiplier method, which ensures reasonable weight distribution; the parameter correlation matrix is a numerical representation of the mutual relationship between different BOP parameters, which is used for the calculation of EWI mergable index; the historical accuracy evaluation is the evaluation result of the prediction accuracy of each BOP parameter based on historical data, which is used to adjust the weight distribution; the target function type refers to the mathematical function form used to optimize the weight in the Lagrange multiplier method, which affects the convergence of the optimization result; the convergence threshold is the judgment standard for judging whether the Lagrange multiplier method iteration calculation reaches the convergence state, which affects the calculation efficiency and accuracy.
[0047] The specific implementation of the above step is described in detail below.
[0048] The specific implementation of step S01 is to construct a multi-dimensional BOP data load matrix. First, collect key parameter data in the production process, and arrange the collected data into a time series form according to the time stamp. Then, classify and organize the parameter types, and group the same parameters into the same dimension to form a preliminary parameter classification structure. Then, perform data format standardization processing, and uniformly convert the parameter data of different dimensions into dimensionless form. The maximum and minimum standardization method is adopted to ensure that all data values fall within the [0, 1] interval. Subsequently, the matrix index structure is established, taking time as the first dimension index and parameter type as the second dimension index, thereby forming a two-dimensional or multi-dimensional data structure. Finally, perform data integrity check and handle missing values. Linear interpolation method can be used to fill in missing data in short time interval, or forward filling method is used to handle continuous missing situation. Through this series of processing, a standardized multi-dimensional BOP data load matrix is finally formed, providing a structured data basis for subsequent analysis.
[0049] The specific implementation of step S02 is to calculate the BOP data contribution value, based on the thermodynamic entropy equation to analyze the influence degree of each dimension data on the system. First, calculate the parameter fluctuation amplitude, and calculate the standard deviation of each parameter within a fixed time window (such as 10 minutes) as the fluctuation amplitude index. Then, measure the fluctuation frequency, and analyze the main frequency component of the parameter in unit time through Fourier transform to obtain the frequency of significant change. Then, evaluate the system response time, and calculate the time delay from parameter change to system output response through cross-correlation function. Generally, faster response time (such as less than 5 seconds) indicates greater parameter influence. At the same time, measure the system self-organization ability, and calculate the ability of the system to recover to stable state after disturbance through Lyapunov index. The index value range is usually between [-1, 1], and the closer to -1 indicates the stronger self-organization ability of the system. Then, quantify the environmental disturbance factor, and calculate the entropy value of each dimension parameter based on the influence degree of temperature (standard working temperature deviation not more than ±5℃), humidity (relative humidity controlled within 45%~65% range), vibration (amplitude not more than 0.1mm) and other environmental parameters. According to the principle of information entropy, the parameter with larger fluctuation and poorer regularity contains more information. Finally, perform normalization processing to ensure that the sum of all dimension BOP data contribution values is 1, thereby forming a weight allocation scheme to provide a basis for subsequent EWI index calculation.
[0050] The specific implementation of step S03 is to establish an EWI base matrix to form a benchmark framework for quality assessment based on historical production data. First, filter the historical high-quality production batch data, and select the best production cycle data according to the product quality test results (usually select the batches with quality test scores in the top 10%). Then calculate the benchmark values of each parameter, and use the mean values of these high-quality batch parameters as the standard reference point. Then determine the parameter fluctuation threshold, calculate the standard deviation of the high-quality batch parameters, and set the allowed fluctuation range as ±3 times the standard deviation, and the fluctuation within this range is considered as normal fluctuation. Then establish the correlation model between parameters, use the Pearson correlation coefficient to analyze the mutual relationship between different parameters, and the correlation coefficient threshold is usually set to 0.7, and the parameters with a value higher than this are considered to be strongly correlated. Then construct the threshold matrix, organize the benchmark values and their allowed fluctuation ranges of each parameter into a matrix form to form a standardized judgment basis. Finally, verify and adjust, use another part of historical data to verify the applicability of the threshold, and if necessary, fine-tune to ensure that the EWI base matrix can accurately reflect the quality state of the production process.
[0051] The specific implementation of step S04 is to implement BOP data capture rate calculation, and the minimum spanning tree algorithm is used to optimize the data collection strategy. First, analyze the data point distribution density, calculate the data distribution density function in the production parameter space by kernel density estimation method, and identify the high-density area. Then determine the data point weight, assign a weight to each data point based on the contribution value of the parameter, and the weight range is usually [0.1, 1.0], reflecting the importance of the data point. Then calculate the connection cost between points, consider the system resource consumption required to collect different data points, including computing resources, storage resources and communication bandwidth, and form a cost matrix. Then analyze the network topology, establish the connection relationship diagram between data collection nodes, and determine the feasible data transmission path. Then consider the system resource limit, according to the actual hardware configuration, set the resource constraint conditions of data collection, such as processor occupancy rate not exceeding 80%, storage space usage rate not exceeding 75%, and communication bandwidth utilization rate not exceeding 90%. Then apply the minimum spanning tree algorithm, based on the Kruskal algorithm or Prim algorithm, to construct the minimum cost tree of data point connection, to maximize the data collection efficiency under resource constraints. Finally, calculate the data capture rate index, and take the ratio of the number of actual collected effective data points to the number of theoretically collected data points as the BOP data capture rate, to obtain the score reflecting the data integrity, and the ideal capture rate should be maintained above 95%.
[0052] The specific implementation of step S05 is to generate the EWI mergable index, and the Lagrange multiplier method is used to realize the comprehensive evaluation of multi-dimensional parameters. First, set the weight upper and lower limit constraints, determine the weight value range according to the importance of each parameter, and generally set the weight lower limit of the core parameter to be not less than 0.1 and the upper limit to be not more than 0.5, so as to ensure reasonable weight distribution. Then, a correlation matrix between parameters is constructed, the correlation coefficients between different BOP parameters are calculated, and a matrix structure representing the mutual influence between parameters is formed. Then, the historical accuracy is evaluated, the prediction accuracy of each parameter on the historical EWI index is analyzed by backtracking test, and the accuracy score usually uses root mean square error, and the ideal value should be less than 0.05. Then, the type of objective function is selected, and the optimization objective function is set based on the principle of minimizing the overall prediction error, and the mean square error or cross-entropy loss function is usually used. Then, the convergence threshold is determined, and the termination condition of iterative calculation is set, such as the weight change amplitude of adjacent two iterations is less than 0.001, which is considered to be converged. Next, the Lagrange multiplier method is applied to solve the optimal weight distribution scheme under the condition of meeting all constraints, and the constraint optimization problem is converted into an unconstrained problem by introducing the Lagrange multiplier. Finally, the EWI mergable index is calculated, the matching degree of the obtained optimal weight and each dimension BOP data and EWI basic matrix is weighted and fused to form a single index in the interval [0, 1], and 0.8 or more indicates a good production state, 0.6-0.8 indicates an acceptable production state, and less than 0.6 indicates an abnormal production state.
[0053] The specific implementation of step S06 is to calculate the lateral error, and the Euclidean distance function is used to analyze the relationship between different parameters at the same time point. First, organize the parameter vector group, and form a vector form of multiple BOP parameter values at the same time point to construct a parameter space. Then, determine the dimension weight coefficient, set the weight coefficient based on the importance of each parameter, and reflect the importance of different dimensions in distance calculation. Then, select a reference standard vector, and usually select the data of the best production batch as the reference point. Then, set the allowable deviation range, determine the allowed parameter deviation according to the production process requirements, and usually set it within ±5% of the standard vector value. Then, determine the correlation threshold, set the critical value of judging parameter correlation anomaly, and generally take 0.75, which is considered to be abnormal below this value. Next, calculate the weighted Euclidean distance, apply the formula to calculate the weighted Euclidean distance between the actual parameter vector and the reference standard vector, and reflect the overall deviation. Then, construct the distance matrix, calculate the distance between different parameter pairs, and form a complete distance matrix. Finally, generate the abnormal correlation identifier, and mark the abnormal correlation parameter pair based on the comparison of the calculated distance and the set threshold, to provide the lateral analysis result for mixed error evaluation.
[0054] The specific implementation of step S07 is to calculate the longitudinal error by tracking the time-varying characteristics of the parameters using the Fourier transform equation. First, prepare the time series parameter values, arrange the single BOP parameters in chronological order to form a discrete time series. Then determine the sampling interval, set the sampling time interval according to the data acquisition frequency, usually in milliseconds (such as 100 ms) or seconds (such as 1 s). Then select the frequency resolution, set the degree of detail of the Fourier transform, usually select a resolution of not less than 0.01 Hz to ensure that key frequency components can be captured. Then select the window function type, apply a Hanning window or Hamming window to weight the original signal, reducing the spectral leakage effect. Then set the spectral smoothing coefficient, usually select a smoothing coefficient of 0.05-0.2 to reduce the influence of random noise on spectral analysis. Next, perform fast Fourier transform to convert the time domain signal to frequency domain representation and obtain the frequency components and their amplitudes. Then analyze the energy distribution characteristics, calculate the energy proportion of different frequency intervals, and determine the main frequency characteristics of parameter changes. Finally, identify abnormal frequency points, compare the obtained spectrum with the standard spectrum under normal working mode, and mark the frequency points with abnormal energy concentration or significant deviation, providing longitudinal analysis basis for hybrid error evaluation.
[0055] The specific implementation of step S08 is to generate a hybrid error evaluation by integrating the lateral error and the longitudinal error to form a complete quality evaluation system. First, integrate the error data, combine the parameter distance matrix and abnormal association identification obtained from lateral analysis with the parameter energy distribution and abnormal frequency points obtained from longitudinal analysis into a unified data structure. Then build an error fusion model, use an adaptive weighting method to determine the fusion weights of lateral error and longitudinal error, usually the initial weights are both 0.5, and then dynamically adjust according to the error characteristics. Then calculate the comprehensive abnormal score, calculate the lateral abnormal score and the longitudinal abnormal score for each parameter, and then combine them according to the fusion weights to form the comprehensive abnormal score. Then classify the abnormal patterns, according to the distribution characteristics of the comprehensive abnormal score, divide the detected abnormalities into three levels: slight abnormality (score between 0.1 and 0.3), moderate abnormality (score between 0.3 and 0.6), and severe abnormality (score greater than 0.6). Then trace the source of the abnormality, based on the correlation analysis of abnormal parameters, trace the possible abnormal root cause and establish the causal relationship diagram. Next, evaluate the influence range of the abnormality, analyze the impact of abnormal parameters on other related parameters, and predict potential chain reactions. Finally, generate a hybrid error evaluation report, including abnormality level, abnormality description, influence range and suggested measures, to provide a comprehensive quality evaluation basis for the final EWI index calculation.
[0056] The specific implementation of step S09 is to output the final EWI index, and generate a comprehensive quality evaluation index based on the aforementioned analysis results. First, the mixed error weight is determined, and the weight proportion of the mixed error in the final index is set according to the characteristics of different production processes, which is usually set to 0.6-0.7. Then, the mergable index weight is determined as a supplement to the mixed error, and the weight is usually set to 0.3-0.4. Then, weighted calculation is performed, and the mixed error evaluation result and the EWI mergable index are weighted and averaged according to the set weight to obtain a preliminary EWI index value. Then, non-linear correction is applied, and the preliminary index is corrected by an S-shaped function (such as a Sigmoid function) to make the final index more sensitive to reflect the quality state change. Then, the index threshold is set, and the judgment standard is determined according to the product quality requirement, and usually the EWI index higher than 0.9 is regarded as high-quality state, 0.7-0.9 is normal state, 0.5-0.7 is early warning state, and lower than 0.5 is abnormal state. Next, trend analysis is performed, and the change rate and acceleration of the EWI index are calculated to predict the development trend of the quality state in the short term. Then, the index visualization result is generated, and the EWI index value is converted into an intuitive dashboard or trend chart form, which is convenient for production managers to quickly identify the quality state. Finally, a comprehensive quality report is output, which includes the EWI index value, quality state judgment, key parameter abnormal situation, trend prediction and improvement suggestion, which provides decision basis for production adjustment and quality control.
[0057] The second aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores program instructions, and the program instructions are used to execute the above-mentioned BOP data grabbing and EWI generating method when running in a computer.
[0058] The third aspect of the present application provides a BOP data grabbing and EWI generating system, which comprises the above-mentioned computer readable storage medium, and the system is any one of a computer, a server and a single chip microcomputer. The computer readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing the program instructions stored in the computer readable storage medium.
[0059] The mathematical models or calculation processes involved in the present application are described in detail as follows.
[0060] The process of constructing a multi-dimensional BOP data load matrix in step S01 involves data standardization processing, which is specifically represented as follows:
[0061]
[0062] In the formula, X norm is the standardized parameter value; X is the original parameter value; X min is the minimum value of the parameter; X max is the maximum value of the parameter.
[0063] This equation uses the maximum-minimum normalization method to convert parameters of different dimensions into the [0, 1] interval, ensuring data comparability. The reason for choosing this method is to preserve the relative relationship between parameters and not change the data distribution pattern, which is suitable for scenarios that need to maintain the relative differences of original data. The way to obtain parameters is: X is directly obtained from the production equipment sensor; X min and X max Through historical data statistical analysis, it is usually determined based on data of at least 30 production cycles.
[0064] The multi-dimensional BOP data load matrix structure is represented as:
[0065]
[0066] In the formula, M is the BOP data load matrix; m ij represents the standardized value of the jth parameter at the ith time point; T is the total number of time points; P is the total number of parameters.
[0067] The matrix construction principle is based on space-time relationship expression, with rows representing time dimension and columns representing parameter dimension, realizing two-dimensional mapping of time series and parameter type, which is convenient for subsequent analysis and processing.
[0068] Calculating the BOP data contribution value in step S02 involves the entropy weight method, which is specifically represented as follows:
[0069]
[0070] In the formula, H j is the information entropy of the jth parameter; n is the number of samples; p ij is the proportion of the jth parameter in the ith sample, and the calculation method is
[0071] The greater the information entropy, the greater the parameter volatility and the more information it contains. The entropy weight method is based on the principle of information theory, which measures the importance of parameters by calculating their uncertainty, and is suitable for processing production systems containing multiple uncertain factors. The parameter is obtained by calculation, where m ij from the data load matrix constructed in step S01.
[0072] The parameter weight coefficient calculation formula is:
[0073]
[0074] In the formula, w j is the weight coefficient of the jth parameter; H j is the information entropy of the jth parameter; P is the total number of parameters.
[0075] The core idea of the formula is that the smaller the information entropy of the parameter, the stronger the distinguishing ability, and the higher the weight should be given. This weight calculation method avoids the influence of subjective factors and can objectively reflect the data characteristics. The parameters are obtained by calculation without additional experimental determination.
[0076] In the thermodynamic entropy equation, various influencing factors are considered, and the comprehensive formula is expressed as:
[0077]
[0078] In the formula, S j is the entropy value contribution coefficient of the jth parameter; A j is the parameter fluctuation amplitude, with a value range of [0, 1]; F j is the fluctuation frequency, with a value range of [0, 100] Hz; T j is the system response time, with a unit of seconds and a value range of [0.1, 10] s; O j is the system self-organization ability, with a value range of [-1, 1]; E j is the environmental disturbance factor, with a value range of [0, 1]; α, β, γ, δ, ∈ are weight coefficients, with a total sum of 1.
[0079] The equation is constructed based on the principle of entropy increase, considering various factors affecting system stability. In the formula, the weighted sum form is used, and the inverse relationship of response time is used, which reflects the characteristics of faster response and greater impact. Parameter acquisition method: A j The standard deviation of the parameter in a 10-minute window is obtained; F j The main frequency is obtained by Fourier transform analysis; T j The input-output delay time is calculated by cross-correlation function; O j It is obtained by calculating the Lyapunov exponent; E j The environmental influence index is calculated by environmental sensor data.
[0080] The establishment of EWI basic matrix in step S03 involves standard parameter value and fluctuation range calculation:
[0081]
[0082] In the formula, EWI base is the EWI basic matrix; e i1 represents the reference value of the ith parameter; e i2 represents the allowed fluctuation range of the ith parameter; P is the total number of parameters.
[0083] The reference value calculation formula is:
[0084]
[0085] In the formula, K represents the number of selected high-quality batches; v ik It represents the average value of the i-th parameter in the k-th high-quality batch.
[0086] Formula for calculating fluctuation range:
[0087] e i2 =3·σ i ;
[0088] In the formula, σ i Let be the standard deviation of the i-th parameter in the high-quality batch, calculated using the following formula:
[0089] This matrix is constructed based on statistical principles, using the average value of high-quality batches as the benchmark and three times the standard deviation as the fluctuation range. It conforms to a 99.7% confidence interval under a normal distribution, ensuring the scientific validity of the judgment criteria. Parameters were obtained through historical data analysis and calculation, with batches ranking in the top 10% of quality inspection scores selected as high-quality batches.
[0090] The core calculation process of the minimum spanning tree algorithm in step S04 involves the following formulas:
[0091] Inter-point connection cost matrix:
[0092]
[0093] In the formula, C is the cost matrix; c ij This represents the connection cost from point i to point j; N is the total number of data points.
[0094] Connection cost calculation formula:
[0095]
[0096] In the formula, d ij w is the Euclidean distance between points i and j; i and w j Let r be the weights of points i and j, respectively; ij λ1, λ2, and λ3 are resource consumption factors; λ1, λ2, and λ3 are weighting coefficients, and λ1 + λ2 + λ3 = 1.
[0097] This formula comprehensively considers three factors: spatial distance, point importance, and resource consumption, and uses a weighted sum. The point weights are inversely related, reflecting the higher priority of connections between important points. Parameter acquisition method: d ij Obtained by calculating the Euclidean distance between two points in the parameter space; w i and w j The parameter weights calculated in step S02; r ijData is obtained through system resource monitoring, including processor usage, storage space, and communication bandwidth consumption.
[0098] BOP data crawling rate calculation formula:
[0099]
[0100] In the formula, R capture For data crawling rate; N actual N represents the actual number of valid data points captured. theoretical This represents the theoretical number of data points that should be captured.
[0101] This formula uses a ratio relationship to intuitively reflect the completeness of data collection; the ideal capture rate should be no less than 95%. Parameter acquisition method: N actual Obtained through data acquisition system; N theoretical It is calculated based on the sampling frequency and time period.
[0102] The optimization problem of finding the optimal weights using the Lagrange multiplier method in step S05 is expressed as follows:
[0103] Objective function:
[0104]
[0105] In the formula, J(W) is the objective function; W = [w1, w2, ..., w P ] represents the weight vector; y i x is the actual EWI value of the i-th sample; ij Let be the matching degree between the j-th parameter in the i-th sample and the basis matrix; n is the number of samples; and P is the total number of parameters.
[0106] Constraints:
[0107]
[0108] w min ≤w j ≤w max j = 1, 2, ..., P;
[0109] In the formula, w min This is the lower limit of the weight, typically set to 0.1; w max This is the upper limit of the weight, usually set to 0.5.
[0110] Lagrange function:
[0111]
[0112] In the formula, λ and μ 1j μ 2j It is a Lagrange multiplier.
[0113] The optimization problem is based on the least square principle, aiming to minimize the prediction error while satisfying the weight constraint. The Lagrange multiplier method converts the constrained optimization problem into an unconstrained problem, and the optimal solution is obtained by solving the equation system whose partial derivative is equal to zero. Parameter acquisition method: y i Obtained through the EWI index value of the historical high-quality batch; x ij Obtained by calculating the matching degree of the parameter value and the basic matrix, and the calculation formula is Where v ij is the actual value of the jth parameter in the ith sample.
[0114] EWI can merge the index calculation formula:
[0115]
[0116] In the formula, EWI combined is the EWI mergable index, with a value range of [0, 1]; w j is the optimal weight of the jth parameter; x j is the matching degree of the jth parameter and the basic matrix.
[0117] The formula adopts a weighted sum form, realizing the comprehensive evaluation of multi-dimensional parameters, and the closer the index value is to 1, the better the production state is. Parameter acquisition method is obtained by calculation, wherein w j is obtained by solving the Lagrange multiplier method; x j is obtained by calculating the matching degree of the parameter and the basic matrix.
[0118] The weighted Euclidean distance calculation formula in step S06 is:
[0119]
[0120] In the formula, D(V, V ref ) is the weighted Euclidean distance; V = [v1, v2,..., v P ] is the actual parameter vector; V ref = [v ref,1 , v ref,2 ,..., v ref,P ] is the reference standard vector; ω j is the weight coefficient of the jth dimension, and
[0121] The formula is based on the concept of Euclidean space distance, introduces the weight coefficient to improve the influence of important parameters, and the calculation form of square sum and square root makes the difference of each dimension be considered. Parameter acquisition method: V is obtained by real-time parameter acquisition in the production process; V ref is obtained from the EWI basic matrix in step S03; ω jThe parameter weight calculation in step S02 is obtained.
[0122] The distance matrix calculation formula is:
[0123]
[0124] In the formula, DM is the distance matrix; d ij represents the weighted Euclidean distance between parameters i and j, and the calculation formula is
[0125] This matrix represents the mutual relationship between parameters, the diagonal elements are 0, and the non-diagonal elements reflect the distance between parameter pairs. The purpose of constructing this matrix is to identify abnormal associations between parameters and provide a basis for mixed error evaluation. The parameter acquisition method is obtained by calculation based on the difference between the actual parameter value and the reference standard value.
[0126] The abnormal association identification calculation formula is:
[0127]
[0128] In the formula, A ij is the abnormal association identification between parameters i and j; d ij is the distance between parameters; τ cor is the correlation threshold, usually 0.75.
[0129] This formula uses a binary judgment method to intuitively mark abnormal associations and simplifies subsequent analysis and processing. The parameter acquisition method is: d ij from the distance matrix calculation result; τ cor is a critical value determined by historical data analysis.
[0130] The core calculation formula of the Fourier transform in step S07 is:
[0131]
[0132] In the formula, X k is the complex value at the kth frequency point; x n is the value of the nth point in the time series; N is the sequence length; i is the imaginary unit; k is the frequency index, k = 0, 1,..., N-1.
[0133] This formula is based on the Fourier transform principle, which converts time-domain signals to frequency domain and realizes the analysis of signal frequency characteristics. The reason for choosing the Fourier transform is that it can effectively identify periodic patterns and abnormal changes. The parameter acquisition method is: x n obtained by the parameter time series data in the production process; N is the number of sampling points, usually 2 m points are selected to improve calculation efficiency, such as 4096 points.
[0134] The window function application formula is:
[0135] x′ n = x n · w n ;
[0136] In the formula, x′ n is the sequence value after applying the window function; x n is the original sequence value; and w n is the window function value.
[0137] For the Hanning window, the calculation formula of w n is:
[0138]
[0139] The application of the window function reduces the spectral leakage effect and improves the accuracy of spectral analysis. Parameter acquisition method: x n comes from time series data; and w n is calculated through the window function.
[0140] The energy distribution calculation formula is:
[0141] E k = |X k | 2 ;
[0142] In the formula, E k is the energy value of the kth frequency point; X k is the kth value of the Fourier transform result; and |X k | represents the modulus of the complex number X k .
[0143] This formula calculates the energy distribution of each frequency point, reflecting the strength of the signal at different frequencies. Parameter acquisition method: the modulus square of the Fourier transform result is obtained through calculation.
[0144] The abnormal frequency point identification formula is:
[0145]
[0146] In the formula, F abnormal is the set of abnormal frequency points; E k is the actual energy distribution; E k,ref is the reference energy distribution; τ f is the frequency abnormality threshold, usually taking 0.3.
[0147] This formula judges the abnormal frequency by relative deviation and marks the frequency points that are significantly different from the normal mode. Parameter acquisition method: E kCalculated by Fourier transform; E k,ref Obtained by analyzing the spectrum of historical high-quality batches; τ f Determined by historical data analysis.
[0148] Key calculation formula for mixed error evaluation in step S08:
[0149] Transverse abnormal score calculation:
[0150]
[0151] In the formula, S horizontal is the transverse abnormal score; A ij is the abnormal association identifier between parameter i and parameter j; P is the total number of parameters.
[0152] This formula calculates the density of abnormal association, reflecting the degree of abnormal relationship between parameters. Parameter acquisition method: A ij From the abnormal association identifier calculation result in step S06.
[0153] Longitudinal abnormal score calculation:
[0154]
[0155] In the formula, S vertical is the longitudinal abnormal score; |F abnormal | is the number of abnormal frequency points; N / 2 is the total number of effective frequency points (considering conjugate symmetry).
[0156] This formula calculates the proportion of abnormal frequency, reflecting the degree of abnormal time variation of parameters. Parameter acquisition method: |F abnormal | comes from the abnormal frequency point identification result in step S07.
[0157] Comprehensive abnormal score calculation:
[0158] S combined = α·S horizontal +(1-α)·S vertical ;
[0159] In the formula, S combined is the comprehensive abnormal score; α is the transverse abnormal weight, usually taking 0.5 as the initial value, which can be dynamically adjusted later.
[0160] This formula adopts a weighted sum form, balancing the transverse and longitudinal abnormalities, and achieving comprehensive quality evaluation. Parameter acquisition method: S horizontal and S vertical come from the transverse abnormal score and longitudinal abnormal score calculation results respectively; α is initially set to 0.5, and is dynamically adjusted according to abnormal characteristics later.
[0161] The final EWI index calculation formula in step S09:
[0162] The preliminary EWI index calculation:
[0163] EWI initial = β · (1-S combined ) + (1-β) · EWI combined ;
[0164] In the formula, EWI initial is the preliminary EWI index; S combined is the comprehensive anomaly score; EWI combined is the EWI mergable index; and β is the mixed error weight, usually taking 0.6-0.7.
[0165] The formula comprehensively considers the anomaly evaluation and the mergable index, and realizes multi-angle quality evaluation. Parameter acquisition method: S combined comes from the comprehensive anomaly score calculation result in step S08; and EWI combined comes from the EWI mergable index calculation result in step S05.
[0166] Nonlinear correction formula:
[0167]
[0168] In the formula, EWI final is the final EWI index; γ is the sensitivity coefficient, usually taking 10; and θ is the midpoint parameter, usually taking 0.5.
[0169] The formula uses the Sigmoid function for nonlinear correction, so that the index is more sensitive to changes in the middle region, and the resolution of the index is improved. Parameter acquisition method: EWI initial comes from the preliminary EWI index calculation result; and γ and θ are determined through historical data verification to obtain the best index response characteristics.
[0170] Optionally, the specific implementation formula of the minimum spanning tree algorithm is specifically represented as follows:
[0171] MST = {(i, j) | (i, j) ∈ E, and meet the minimum cost tree condition};
[0172] In the formula, MST is the minimum spanning tree; and E is an edge set containing all possible data point connections.
[0173] Optionally, the mathematical expression of the implementation step of the Kruskal algorithm is:
[0174]
[0175] The edges in E are sorted in ascending order according to the cost c ij .
[0176] Each edge (i, j) is investigated in turn, if adding (i, j) does not form a loop, then T=T U {(i, j)};
[0177] Until |T|=N-1;
[0178] In the formula, T is a tree currently constructed; |T| is the number of edges in the tree; and N is the total number of data points.
[0179] The algorithm is based on a greedy strategy, and the minimum spanning tree is constructed by gradually selecting the edge with the minimum cost, so that optimal data acquisition is realized under resource constraints. ij The connection cost calculation result from step S04.
[0180] Optionally, the calculation formula of the parameter correlation matrix is as follows:
[0181]
[0182] In the formula, R is the parameter correlation matrix; r ij Pearson correlation coefficient between the parameters i and j, and the calculation formula is as follows:
[0183]
[0184] In the formula, v ki Indicates the i-th parameter value of the k-th sample; Indicates the average value of the i-th parameter; and n is the sample quantity.
[0185] The formula is based on the statistical principle, and the linear correlation degree between the parameters is calculated, the value range is [-1, 1], and the greater the absolute value is, the stronger the correlation is. The parameter value is obtained through historical data in the production process; and the average value is obtained by calculating the arithmetic average of the historical data.
[0186] Specifically, the principle of the present application is as follows: the present application is based on the information entropy theory, spectrum analysis, graph theory optimization and multi-objective constraint optimization principle, and a systematic BOP data grabbing and EWI generation framework is constructed, and the core principle is that through the synergistic effect of multiple algorithms, the BOP data grabbing and EWI generation process are accurately controlled.
[0187] Firstly, the contribution value of each BOP parameter is calculated based on the thermodynamic entropy equation. This process follows the information entropy theory, which regards the production system as an energy-information conversion system. The fluctuation characteristics of the parameters, the response properties of the system, and the environmental impact are quantified as entropy increase indicators, scientifically measuring the influence of each parameter on the stability of the system. Compared with the traditional empirical weight setting, this weight distribution method based on information entropy can objectively reflect the actual importance of the parameters, eliminating the accuracy deviation caused by subjective judgment and providing a scientific and accurate data basis for subsequent analysis.
[0188] Secondly, the BOP data grabbing strategy is optimized by the minimum spanning tree algorithm. Based on the construction principle of the minimum spanning tree in graph theory, the connection cost between data points and data importance are comprehensively considered to determine the optimal data collection path under limited resources. This method ensures that the most representative set of BOP data points is grabbed, improving the quality and representativeness of the original data and laying a solid foundation for generating accurate EWI indicators.
[0189] Furthermore, the important innovation of the present invention is to introduce both horizontal and vertical error analysis mechanisms, constructing a two-dimensional accurate analysis framework. The horizontal error accurately calculates the deviation between different BOP parameters at the same time point through the Euclidean distance function, which can capture the complex nonlinear relationship between parameters. The vertical error accurately detects the abnormal change pattern of a single BOP parameter in the time series through Fourier transform, effectively identifying the time evolution characteristics of the parameter. This two-dimensional analysis method overcomes the limitations of traditional single-dimensional analysis and can accurately evaluate the quality of BOP data and generate accurate EWI indicators.
[0190] Finally, the EWI fusion weight is accurately solved under multiple constraints by the Lagrange multiplier method. Based on the multi-objective constraint optimization theory, this method can find the optimal weight configuration under the premise of meeting all constraint conditions, so that the generated EWI indicators can accurately reflect the actual contribution of each BOP parameter and meet the system's requirements for stability.
[0191] The multi-algorithm fusion technical solution can effectively solve the core problem of inaccurate BOP data grabbing and EWI generation, because it fundamentally changes the traditional single, static, and linear analysis mode, introduces more scientific weight calculation methods, more optimized data grabbing strategies, more comprehensive error analysis systems, and more accurate data fusion algorithms, constructs a rigorous and efficient EWI generation framework, and provides an accurate and reliable BOP data grabbing and EWI generation method.
[0192] A specific embodiment 1 of the present invention is provided below, and the specific implementation method of each step in embodiment 1 is described in detail as follows.
[0193] The specific implementation of step S01 is to construct a multi-dimensional BOP data load matrix. First, key parameter data in the production process are collected, and the collected data are arranged in time sequence according to time stamp. Then, the parameter types are classified and arranged, and parameters of the same type are classified into the same dimension to form a preliminary parameter classification structure. Then, data format standardization processing is performed, and parameter data of different dimensions are uniformly converted into dimensionless form. The maximum and minimum standardization method is adopted to ensure that all data values fall within the [0, 1] interval, and the calculation formula is:
[0194]
[0195] In the formula, X norm is the standardized parameter value; X is the original parameter value; X min is the minimum value of the parameter; X max is the maximum value of the parameter. X is directly obtained from the production equipment sensor; X min and X max are obtained through historical data statistical analysis, and are usually determined based on data of at least 30 production cycles. Subsequently, a matrix index structure is established, time is taken as the first dimension index, and parameter type is taken as the second dimension index, thereby forming a multi-dimensional data structure, which is represented as:
[0196]
[0197] In the formula, M is the BOP data load matrix; m ij represents the standardized value of the jth parameter at the ith time point; T is the total number of time points; and P is the total number of parameters. Finally, data integrity check is performed, and missing values are processed. Linear interpolation method can be used to fill in missing data in a short time interval, or forward filling method can be used to process continuous missing situations. Through this series of processing, a standardized multi-dimensional BOP data load matrix is finally formed, providing a structured data basis for subsequent analysis.
[0198] The specific implementation of step S02 is to calculate the BOP data contribution value, and analyze the influence degree of each dimension data on the system based on the thermodynamic entropy equation. First, the parameter fluctuation amplitude is calculated, and the standard deviation of each parameter in a fixed time window (such as 10 minutes) is calculated as the fluctuation amplitude index. Then, the fluctuation frequency is measured, and the main frequency component of the parameter in unit time is analyzed by Fourier transform to obtain the frequency of significant change. Then, the system response time is evaluated, and the time delay from parameter change to system output response is calculated by cross-correlation function, and a faster response time (such as less than 5 seconds) indicates that the parameter has a greater impact. At the same time, the system self-organization ability is measured, and the ability of the system to recover to a stable state after disturbance is calculated by Lyapunov index, and the index value range is usually between [-1, 1], and the closer to -1 indicates that the system self-organization ability is stronger. The environmental disturbance factor is quantified, and the influence degree of environmental parameters such as temperature (the standard working temperature deviation is not more than ±5℃), humidity (the relative humidity is controlled within the range of 45%~65%), vibration (the amplitude is not more than 0.1mm) and the like is calculated. The entropy weight method is used to calculate the entropy value of each dimension parameter:
[0199]
[0200] In the formula, H j is the information entropy of the jth parameter; n is the sample number; p ij is the proportion of the jth parameter in the ith sample, and the calculation method is Then, the parameter weight coefficient is calculated:
[0201]
[0202] In the formula, w j is the weight coefficient of the jth parameter; H j is the information entropy of the jth parameter; P is the total number of parameters. Finally, considering various influence factors, the comprehensive formula is represented as:
[0203]
[0204] In the formula, S j is the entropy value contribution coefficient of the jth parameter; A j is the parameter fluctuation amplitude, and the value range is [0, 1]; F j is the fluctuation frequency, and the value range is [0, 100]Hz; T j is the system response time, and the unit is second, and the value range is [0.1, 10]s; O j is the system self-organization ability, and the value range is [-1, 1]; E jThe environmental interference factor has a value range of [0, 1]; and a, b, g, d, and e are weight coefficients, and the sum is 1. Through normalization processing, the sum of the contribution values of all dimensions of BOP data is ensured to be 1, thereby forming a weight allocation scheme, and providing a basis for subsequent EWI index calculation.
[0205] The specific implementation of step S03 is to establish an EWI basic matrix, and form a quality evaluation benchmark framework based on historical production data. First, the historical high-quality production batch data is screened, and the best performance production cycle data is selected according to the product quality detection result (usually the batches with quality detection scores in the top 10% are selected). Then, the benchmark values of each parameter are calculated, and the mean values of the parameters corresponding to these high-quality batches are used as standard reference points:
[0206]
[0207] In the formula, K is the number of selected high-quality batches; v ik is the average value of the i-th parameter in the k-th high-quality batch. Then, the parameter fluctuation threshold is determined, and the standard deviation of the parameters of the high-quality batches is calculated, and the allowed fluctuation range is set to ±3 times the standard deviation:
[0208] e i2 = 3·σ i ;
[0209] In the formula, s i is the standard deviation of the i-th parameter in the high-quality batch, and the calculation formula is Then, the correlation model between parameters is established, and the Pearson correlation coefficient is used to analyze the mutual relationship between different parameters, and the correlation coefficient threshold is usually set to 0.7, and the parameters with a value higher than this value are considered to be strongly correlated. Then, the EWI basic matrix is constructed:
[0210]
[0211] In the formula, EWI base is the EWI basic matrix; e i1 represents the benchmark value of the i-th parameter; e i2 represents the allowed fluctuation range of the i-th parameter; and P is the total number of parameters. Finally, verification and adjustment are performed, the applicability of the threshold is verified using another part of the historical data, and if necessary, fine tuning is performed to ensure that the EWI basic matrix can accurately reflect the quality state of the production process.
[0212] The specific implementation of step S04 involves calculating the BOP data capture rate and optimizing the data acquisition strategy using the minimum spanning tree algorithm. First, the data point distribution density is analyzed. The data distribution density function in the production parameter space is calculated using the kernel density estimation method to identify high-density regions. Then, the data point weights are determined. Each data point is assigned a weight based on its parameter contribution value, typically within the range of [0.1, 1.0], reflecting the importance of the data point. Next, the connection cost between points is calculated, considering the system resource consumption required to collect different data points, including computing resources, storage resources, and communication bandwidth, forming a cost matrix.
[0213]
[0214] In the formula, C is the cost matrix; c ij This represents the connection cost from point i to point j; N is the total number of data points. The formula for calculating the connection cost is:
[0215]
[0216] In the formula, d ij w is the Euclidean distance between points i and j; i and w j Let r be the weights of points i and j, respectively; ij λ1, λ2, and λ3 are resource consumption factors; λ1, λ2, and λ3 are weight coefficients, and λ1 + λ2 + λ3 = 1. Next, the network topology is analyzed, a connection graph between data acquisition nodes is established, and feasible data transmission paths are determined. Then, considering system resource constraints, resource constraints for data acquisition are set based on the actual hardware configuration, such as processor utilization not exceeding 80%, storage space utilization not exceeding 75%, and communication bandwidth utilization not exceeding 90%. Then, the minimum spanning tree algorithm is applied, and a minimum cost tree for connecting data points is constructed based on Kruskal's algorithm.
[0217] MSE = {(i,j)|(i,j)∈E, satisfying the minimum cost tree condition};
[0218] The mathematical expression of the steps in implementing Kruskal's algorithm:
[0219]
[0220] For edges in E, according to cost c ij Sort by size from smallest to largest;
[0221] Examine each edge (i, j) in turn. If adding (i, j) does not form a cycle, then T = T∪{(i, j)}.
[0222] Until |T| = N-1;
[0223] In the formula, T represents the currently constructed tree; |T| represents the number of edges in the tree; and N represents the total number of data points. Finally, the data capture rate metric is calculated.
[0224]
[0225] In the formula, R capture For data crawling rate; N actual N represents the actual number of valid data points captured. theoretical This represents the theoretical number of data points that should be captured. Ideally, the capture rate should be maintained above 95%.
[0226] The specific implementation of step S05 involves generating an EWI merging index and using the Lagrange multiplier method to achieve a comprehensive evaluation of multidimensional parameters. First, upper and lower limits for weights are set, and the weight range is determined based on the importance of each parameter. Typically, the lower limit for the weight of core parameters is no less than 0.1, and the upper limit is no more than 0.5, ensuring reasonable weight allocation. Then, a correlation matrix between parameters is constructed.
[0227]
[0228] In the formula, R is the parameter correlation matrix; r ij The Pearson correlation coefficient between parameter i and parameter j is calculated using the following formula:
[0229]
[0230] In the formula, v ki This represents the value of the i-th parameter in the k-th sample; This represents the average value of the i-th parameter; n is the number of samples. Next, historical accuracy is evaluated by analyzing the predictive accuracy of each parameter for historical EWI indicators through backtesting. Accuracy scoring typically uses the root mean square error (RMSE), with an ideal value less than 0.05. Then, the objective function type is selected, and the optimization objective function is set based on the principle of minimizing the overall prediction error:
[0231]
[0232] In the formula, J(W) is the objective function; W = [w1, w2, ..., w P ] represents the weight vector; y i x is the actual EWI value of the i-th sample; ij Let be the matching degree between the j-th parameter in the i-th sample and the fundamental matrix; n is the number of samples; P is the total number of parameters. Constraints:
[0233]
[0234] w min ≤w j ≤w maxj = 1, 2, ..., P;
[0235] In the formula, w min This is the lower limit of the weight, typically set to 0.1; w max The upper limit for the weights is typically set to 0.5. Then, a convergence threshold is determined, and a termination condition for the iterative calculation is set; convergence is considered achieved when the weight change between two consecutive iterations is less than 0.001. Next, the Lagrange multiplier method is applied to solve for the optimal weight allocation scheme under all constraints.
[0236]
[0237] In the formula, λ and μ 1j μ 2j These are Lagrange multipliers. Finally, the EWI combinable index is calculated:
[0238]
[0239] In the formula, EWI combined The EWI can be combined exponent, with values ranging from [0, 1]; w j The optimal weight for the j-th parameter; x j The matching degree between the j-th parameter and the fundamental matrix is calculated using the following formula: Where v j This represents the actual value of the parameter. The index value ranges from [0, 1], where a value above 0.8 indicates good production status, 0.6 to 0.8 indicates acceptable production status, and a value below 0.6 indicates abnormal production status.
[0240] The specific implementation of step S06 involves calculating the lateral error and using the Euclidean distance function to analyze the relationship between different parameters at the same time point. First, parameter vector groups are organized, forming a vector from multiple BOP parameter values at the same time point to construct a parameter space. Then, dimensional weight coefficients are determined based on the importance of each parameter, reflecting the importance of different dimensions in the distance calculation. Next, a reference standard vector is selected, using the parameter vector under optimal production conditions as the benchmark, typically choosing data from the best historical production batch. Then, an allowable deviation range is set, determining the permissible parameter deviation based on production process requirements, usually set within ±5% of the standard vector value. Next, a correlation threshold is determined, setting a critical value for judging abnormal parameter correlation, generally 0.75; values below this are considered abnormal. Finally, the weighted Euclidean distance is calculated.
[0241]
[0242] In the formula, D(V, V) ref () represents the weighted Euclidean distance; V = [v1, v2, ..., v P] is the actual parameter vector; V ref = [v ref,1 , v ref,2 ,..., v ref,P ] is the reference standard vector; ω j is the weight coefficient of the jth dimension, and Then the distance matrix is constructed:
[0243]
[0244] In the formula, DM is the distance matrix; d ij represents the weighted Euclidean distance between parameter i and parameter j, and the calculation formula is Finally, the abnormal correlation identifier is generated:
[0245]
[0246] In the formula, A ij is the abnormal correlation identifier between parameter i and parameter j; d ij is the distance between parameters; τ cor is the correlation threshold, usually 0.75. Through these calculations, a complete distance matrix and abnormal correlation identifier are formed, providing horizontal analysis results for mixed error evaluation.
[0247] The specific implementation of step S07 is to calculate the longitudinal error, using the Fourier transform equation to track the time variation characteristics of the parameters. First, prepare the time series parameter values, arrange the single BOP parameters in time sequence to form a discrete time series. Then determine the sampling interval, set the sampling time interval according to the data acquisition frequency, usually in milliseconds (such as 100 ms) or seconds (such as 1 s). Then select the frequency resolution, set the degree of detail of the Fourier transform, usually select a resolution not less than 0.01 Hz to ensure that key frequency components can be captured. Then select the window function type, apply the Hanning window to weight the original signal to reduce the spectral leakage effect:
[0248] x′ n = x n · w n ;
[0249] In the formula, x′ n is the sequence value after applying the window function; x n is the original sequence value; w n is the window function value. For the Hanning window, the calculation formula of w n is:
[0250]
[0251] Then set the spectrum smoothing coefficient, usually choose 0.05~0.2 smoothing coefficient, reduce the influence of random noise on spectrum analysis. Next perform fast Fourier transform:
[0252]
[0253] In the formula, X k is the complex value at the kth frequency point; x n is the value of the nth point in the time series; N is the sequence length; i is the imaginary unit; k is the frequency index, k=0, 1,..., N-1. Then analyze the energy distribution characteristics:
[0254] E k =|X k | 2 ;
[0255] In the formula, E k is the energy value of the kth frequency point; X k is the kth value of the Fourier transform result; |X k | represents the modulus of complex X k . Finally identify the abnormal frequency points:
[0256]
[0257] In the formula, F abnormal is the set of abnormal frequency points; E k is the actual energy distribution; E k,ref is the reference energy distribution; τ f is the frequency anomaly threshold, usually 0.3. Through this series of processing, the energy distribution of the parameters at different frequencies and the abnormal frequency points are obtained, which provide longitudinal analysis basis for mixed error evaluation.
[0258] The specific implementation of step S08 is to generate a mixed error evaluation, which integrates the horizontal error and the longitudinal error to form a complete quality evaluation system. First, integrate the error data, integrate the distance matrix between parameters and the abnormal correlation identifier obtained by horizontal analysis and the parameter energy distribution and abnormal frequency point set obtained by longitudinal analysis into a unified data structure. Then build an error fusion model, use adaptive weighting method to determine the fusion weight of horizontal error and longitudinal error, usually the initial weight is 0.5, then dynamically adjust according to the error characteristics. Then calculate the horizontal abnormal score:
[0259]
[0260] In the formula, S horizontal is the horizontal abnormal score; A ij is the abnormal correlation identifier between parameter i and parameter j; P is the total number of parameters. Then calculate the longitudinal abnormal score:
[0261]
[0262] In the formula, S vertical For longitudinal outlier scores; |F abnormal | represents the number of abnormal frequency points; N / 2 represents the total number of valid frequency points (considering conjugate symmetry). The overall anomaly score is then calculated:
[0263] S combined =α·S horizontal +(1-α)·S vertical ;
[0264] In the formula, S combined The comprehensive anomaly score is used; α is the horizontal anomaly weight, typically set to 0.5 initially and dynamically adjustable later. Next, anomaly pattern classification is performed. Based on the distribution characteristics of the comprehensive anomaly score, detected anomalies are divided into three levels: minor anomalies (scores between 0.1 and 0.3), moderate anomalies (scores between 0.3 and 0.6), and severe anomalies (scores greater than 0.6). Then, anomaly source tracing is performed. Based on correlation analysis of anomaly parameters, possible root causes are traced, and a causal relationship diagram is established. Finally, the scope of anomaly impact is assessed, analyzing the ripple effect of anomaly parameters on other related parameters, predicting potential chain reactions, and generating a mixed error assessment report, providing a comprehensive quality assessment basis for the final EWI index calculation.
[0265] The specific implementation of step S09 involves outputting the final EWI index and generating a comprehensive quality assessment index based on the aforementioned analysis results. First, the weight of the mixed error is determined. The weight ratio of the mixed error in the final index is set according to the characteristics of different production processes, typically between 0.6 and 0.7. Then, the weight of the mergeable index is determined as a supplement to the mixed error, typically set between 0.3 and 0.4. Next, a weighted calculation is performed, averaging the mixed error assessment results with the EWI mergeable index according to the set weights to obtain the preliminary EWI index value.
[0266] EWI initial =β·(1-S combined )+(1-β)·EWI combined ;
[0267] In the formula, EWI initial For preliminary EWI indicators; S combined EWI (Extra-Outlier Score) combined β is the EWI pooling index; β is the mixed error weight, typically taken as 0.6–0.7. Subsequently, nonlinear correction is applied, using a sigmoid function to adjust the initial index, making the final index more sensitive to changes in quality status.
[0268]
[0269] In the formula, EWI final is the final EWI index; γ is the sensitivity coefficient, usually 10; θ is the midpoint parameter, usually 0.5. Then set the index threshold, determine the judgment standard according to the product quality requirements, usually EWI index higher than 0.9 is considered as high-quality state, 0.7-0.9 is normal state, 0.5-0.7 is early warning state, and lower than 0.5 is abnormal state. Next, perform trend analysis, calculate the change rate and acceleration of EWI index, and predict the development trend of quality state in the short term. Then generate index visualization results, convert EWI index value into intuitive dashboard or trend chart form, which is convenient for production managers to quickly identify the quality state. Finally, output the comprehensive quality report, including EWI index value, quality state judgment, key parameter abnormal situation, trend prediction and improvement suggestions, to provide decision basis for production adjustment and quality control.
[0270] Through the implementation of the above steps, the whole process from BOP data collection to EWI index generation is completed, and the comprehensive monitoring and evaluation of production quality is realized. This method comprehensively applies maximum and minimum standardization, entropy weight method, correlation analysis, minimum spanning tree algorithm, Lagrange multiplier method, Euclidean distance calculation and Fourier transform and other mathematical and computer algorithms, and establishes a scientific and complete production quality evaluation system. Each step is closely connected and interdependent, and together constitutes a closed-loop quality monitoring process. Through multi-dimensional analysis and evaluation of BOP data, quality abnormalities in the production process can be found in time, early warning and intervention can be realized, and production efficiency and product quality can be improved.
[0271] In order to better understand and implement the present application, the following provides an embodiment 2 of a specific application scenario of the present application: In a certain advanced semiconductor wafer production line, researchers apply the method of the present application to implement an automatic quality monitoring system to solve the problem of lack of efficient and forward-looking quality index in chip manufacturing process. The system collects BOP (by process) data of various production equipment and generates real-time EWI (early warning index), realizes accurate prediction and control of production quality.
[0272] The researchers first determined 12 key BOP parameters, as shown in Table 1:
[0273] Table 1 List of key BOP parameters for wafer manufacturing
[0274]
[0275]
[0276] According to step S01, the researchers constructed a multi-dimensional BOP data load matrix. The system collected data at 500 time points (one data point every 5 seconds), forming a 500x12 matrix structure. The maximum and minimum normalization method was applied to convert all parameter data uniformly to the [0,1] interval.
[0277] In step S02, the researchers calculated the weight coefficients of each parameter by the entropy weight method, as shown in Table 2:
[0278] Table 2 BOP parameter entropy value and weight coefficient
[0279] Parameter number j )]]> weight coefficient (w j )]]> P1 0.413 0.135 P2 0.387 0.142 P3 0.456 0.126 P4 0.329 0.155 P5 0.621 0.087 P6 0.294 0.163 P7 0.537 0.107 P8 0.689 0.072 P9 0.368 0.146 P10 0.792 0.048 P11 0.846 0.035 P12 0.769 0.054
[0280] When the system calculates the entropy value contribution coefficient based on the thermodynamic entropy equation, the weight parameters are set as: α = 0.3, β = 0.25, γ = 0.2, δ = 0.15, ∈ = 0.1, considering the parameter fluctuation amplitude, fluctuation frequency, system response time, self-organization ability and environmental disturbance factor. Figure 2 and Figure 3 The time sequence changes of key BOP parameters in the wafer manufacturing process and the corresponding generated EWI index comparison are shown. Among them, Figure 2 The standardized numerical value of five key parameters (P1 plasma power, P2 chamber pressure, P4 RF bias, P6 etching uniformity, P9 endpoint signal) is shown as a function of time. These parameters are simulated according to the parameter description in Table 1 in the embodiment, and all parameter values are normalized to the [0,1] interval. The abnormal area between 300-350 time points is specially marked in the chart, and the P4 (RF bias) parameter shows obvious abnormal fluctuation in this interval, which corresponds to the spectrum analysis anomaly found in step S07 in the embodiment. Figure 3 The EWI index change curve calculated based on the above parameters is shown, including the initial EWI index (dashed line) and the EWI index after nonlinear Sigmoid correction (solid line), as well as the lateral error index curve. The figure also marks three different colored horizontal dashed lines, representing the high-quality threshold (0.9), the warning threshold (0.8) and the danger threshold (0.7). It can be clearly observed that in the abnormal area, the EWI index value decreases significantly, and the lateral error index increases significantly, fully embodying the abnormal detection and early warning mechanism described in steps S06 and S09 in the embodiment.
[0281] In step S03, the researchers selected 50 production batches with quality detection scores in the top 10% as high-quality batches and calculated the EWI basic matrix. The reference values and allowable fluctuation ranges of some parameters are shown in Table 3:
[0282] Table 3 EWI basic matrix (part of parameters)
[0283] Parameter number reference value (e i1 )]]> allowable fluctuation range (e i2 )]]> P1 1050.32 31.59 P2 5.47 0.16 P3 187.63 5.63 P4 275.81 8.27 P6 97.83 2.93 P9 653.24 19.60
[0284] In step S04, the system optimizes the data collection strategy by applying the minimum spanning tree algorithm. First, a 500x500 inter-point connection cost matrix is established, with λ1=0.5, λ2=0.3, and λ3=0.2. Then, the minimum spanning tree is constructed by the Kruskal algorithm to determine the optimal data collection path. In actual implementation, the sampling frequency is reduced from 1 per second to 1 per 5 seconds, but the data capture rate is increased to 97.6%, which is higher than the set threshold of 95%, by optimizing the sampling point distribution.
[0285] In step S05, the system calculates the EWI mergable index by applying the Lagrange multiplier method. The upper and lower weight constraints are set to [0.05, 0.25], and the optimal weight allocation scheme is found by iterative solution. In a certain batch production process, the calculated EWI mergable index is 0.862, indicating that the production state is within the normal range.
[0286] Step S06 calculates the lateral error. The system selects the average value of the high-quality batch as the reference standard vector and calculates the deviation of the actual parameter vector from the reference vector using the weighted Euclidean distance. The correlation threshold τ cor is set to 0.75, and the system identifies an abnormal correlation between P3 (gas flow) and P9 (end point signal) with a distance value of 0.83, which exceeds the threshold.
[0287] Step S07 calculates the longitudinal error. The system performs Fourier transform analysis on each parameter using 4096-point FFT, Hanning window function , and a smoothing coefficient of 0.15. In the spectral analysis of P4 (RF bias) parameter, three abnormal frequency points are found, with energy deviation exceeding the set threshold τ f = 0.3.
[0288] Figure 4 The spectral analysis chart shows the differences between normal and abnormal spectra, with three abnormal frequency points clearly marked. The abnormal frequency points appear as additional energy peaks outside the normal spectrum, with energy deviation exceeding the set threshold. Figure 5 The figure shows the trends of lateral abnormal score, longitudinal abnormal score, and comprehensive abnormal score over time. The lateral abnormal score is based on the weighted Euclidean distance calculated in step S06, the longitudinal abnormal score is based on the spectral analysis results in step S07, and the comprehensive abnormal score is a combination of the two with a weight of α=0.5 according to the method in step S08. The figure marks two horizontal dashed lines for the slight abnormal threshold (0.1) and the severe abnormal threshold (0.2). It can be seen that in the abnormal area, all three abnormal scores have a significant increase, especially the longitudinal abnormal score.
[0289] Step S08 generates a mixed error assessment, with the lateral outlier score S. horizontal =0.061, longitudinal outlier score S vertical =0.113, and the comprehensive anomaly score S is calculated using a weight of α=0.5. combined =0.087, which falls within the range of slight abnormalities (below 0.1).
[0290] In step S09, the system sets the mixed error weight β = 0.65 and calculates the preliminary EWI index value as EWI. initial =0.857. Then, nonlinear correction was performed using the Sigmoid function (γ = 10, θ = 0.5), resulting in a final EWI value of 0.903, indicating excellent production quality. The system automatically generated a quality report, pointing out that although there were minor anomalies, the overall production status was good, and suggesting attention be paid to any abnormal correlation between gas flow and endpoint signals.
[0291] Traditional semiconductor manufacturing quality control relies primarily on post-production inspection and statistical process control (SPC), which cannot predict quality trends in real time. Problems are often only discovered after batch completion, leading to material waste and inefficiency. This invention, however, utilizes a method for generating an EWI (Extended Quality Indicator) based on BOP (Batch of Plant) data capture, enabling real-time prediction and proactive control of production quality. This method offers significant advantages: First, it adaptively determines parameter weights using the entropy weighting method and thermodynamic entropy equation, overcoming the limitations of subjective weighting in traditional methods and resulting in a more objective and scientific weight allocation. Second, it optimizes the data acquisition strategy using the minimum spanning tree algorithm, reducing system resource consumption while ensuring data integrity and improving the data capture rate by 8.3%. Third, it combines a hybrid evaluation mode of lateral and longitudinal errors, enabling multi-dimensional identification of production anomalies, improving the anomaly detection rate by 23.7% and reducing the false alarm rate by 17.2%. Finally, the EWI index, with nonlinear correction, more sensitively reflects quality change trends, advancing the warning time by an average of 42 minutes, providing ample time for production adjustments. Practical application has proven that this method effectively improves the quality control level of semiconductor manufacturing, reduces defect rates, and lowers production costs.
[0292] It should be noted that the variables involved in this invention are explained in detail in Tables 4 and 5 below.
[0293] Table 4. Variable Explanation Table (Part 1)
[0294]
[0295]
[0296]
[0297] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A method for generating EWI from BOP data, characterized in that, Includes the following steps: S01. Construct a multidimensional BOP data load matrix. Classify and organize the BOP data collected during the production process according to time series and parameter type to form a preliminary parameter classification structure. Then, perform data format standardization processing and establish a matrix index structure, using time as the first dimension index and parameter type as the second dimension index, thereby forming a two-dimensional or multi-dimensional data structure. S02. Calculate the contribution value of BOP data. The thermodynamic entropy equation is used to analyze the influence of each BOP data dimension. The inputs include parameter fluctuation amplitude, fluctuation frequency, system response time, system self-organization capability, and environmental interference factor. The output is the entropy contribution coefficient of each BOP parameter. Specifically: first, calculate the parameter fluctuation amplitude; then measure the fluctuation frequency; next, evaluate the system response time; simultaneously measure the system self-organization capability; and then quantify the environmental interference factor. The entropy weight method is used to calculate the entropy value of each dimension parameter. Finally, normalization is performed to ensure that the sum of the contribution values of all dimensions of BOP data is 1, thus forming a weight allocation scheme to provide a basis for subsequent EWI index calculation. The thermodynamic entropy equation is expressed as follows: ; In the formula, For the first The entropy contribution coefficient of each parameter; The parameter represents the fluctuation range, with a value range of [0, 1]. The fluctuation frequency is [0, 100] Hz; The system response time is expressed in seconds and ranges from 0.1 to 10 seconds. The system's self-organizing capability has a value range of [-1, 1]. This is an environmental interference factor, with a value range of [0, 1]. , , , , These are weighting coefficients, and their sum is 1; among them, the EWI basic matrix is a reference standard for quality control in the production process, containing the standard values and fluctuation thresholds of each parameter, and is used as a benchmark for data analysis. S03. Establish an EWI (Expert Insights and Weighing) matrix to form a benchmark framework for quality assessment based on historical production data. Select the best-performing production cycle data according to product quality inspection results, and use the average parameter values corresponding to these high-quality batches as standard reference points. Then, construct a threshold matrix, organizing the benchmark values of each parameter and their allowable fluctuation ranges into a matrix form to form standardized judgment criteria. S04. Implement BOP data crawling rate calculation. Use the minimum spanning tree algorithm to monitor the ratio of the actual effective BOP data crawled to the theoretically required BOP data crawled during the production process based on data point distribution density, data point weight, inter-point connection cost, network topology, and system resource constraints. This ratio is used to obtain a score that reflects data integrity. S05. Generate an EWI mergeable index. Using the Lagrange multiplier method, the matching degree between the BOP data of each dimension and the EWI base matrix is weighted and fused based on the upper and lower limits of weights, the correlation matrix between parameters, historical accuracy evaluation, objective function type, and convergence threshold to form a single measurement index. S06. Calculate the lateral error. Use the Euclidean distance function to analyze the degree of deviation between different BOP parameters at the same time point based on the parameter vector group, dimension weight coefficient, reference standard vector, allowable deviation range and correlation threshold, and obtain the distance matrix between parameters and the abnormal correlation identifier. S07. Calculate the longitudinal error. Use the Fourier transform equation to track the changing trend of a single BOP parameter in the time series based on the time series parameter values, sampling interval, frequency resolution, window function type, and spectral smoothing coefficient, and obtain the energy distribution and abnormal frequency points of the parameter at different frequencies. S08. Generate a mixed error assessment, which integrates the analysis of horizontal and vertical errors. Specifically, it constructs a comprehensive quality evaluation system based on the distance matrix between parameters, anomaly correlation indicators, parameter energy distribution, and anomaly frequency points, and analyzes the mixed error assessment results. S09. Output the final EWI index, which is a comprehensive index reflecting the production quality status based on the mixed error assessment results and the EWI merging index.
2. The method for generating EWI from BOP data according to claim 1, characterized in that, The BOP data load matrix is a multi-dimensional data structure that constructs key parameters in the production process according to time and type, and is used to represent the load status of each parameter at different times and their mutual influence.
3. The method for generating EWI from BOP data according to claim 2, characterized in that, The contribution value of BOP data is the influence weight of each BOP data point on the formation of the final EWI index. It is calculated by the thermodynamic entropy equation and reflects the degree of influence of different parameters on production quality.
4. The method for generating EWI from BOP data according to claim 3, characterized in that, The EWI basic matrix is a reference standard for quality control in the production process, containing standard values and fluctuation thresholds for each parameter, and is used as a benchmark for data analysis.
5. The method for generating EWI from BOP data according to claim 4, characterized in that, The BOP data capture rate is the percentage of valid BOP data successfully acquired during production monitoring out of the total theoretically required BOP data, reflecting the reliability of the data acquisition system.
6. The method for generating EWI from BOP data according to claim 5, characterized in that, The EWI mergeable index is the process of integrating multi-dimensional parameter BOP data into a single index through the Lagrange multiplier method, which facilitates intuitive assessment of the overall production status.
7. The method for generating EWI from BOP data according to claim 6, characterized in that, The horizontal error is the relative deviation between different BOP parameters at the same time point, used to identify abnormal correlations between BOP parameters; the vertical error is the deviation of a single BOP parameter in the time series, used to monitor whether the trend of BOP parameter change meets expectations. The mixed error is a comprehensive index obtained by fusing lateral and longitudinal errors according to specific weights, which fully reflects the quality status of the production process.
8. The method for generating EWI from BOP data according to claim 7, characterized in that, The hybrid error assessment integrates the analysis of horizontal and vertical errors. Specifically, it constructs a comprehensive quality evaluation system based on the distance matrix between parameters, anomaly correlation indicators, parameter energy distribution, and anomaly frequency points, and then analyzes and obtains the hybrid error assessment results.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed in a computer, are used to perform a method for generating EWI from BOP data according to any one of claims 1-8.
10. A system for extracting BOP data and generating EWI, characterized in that, The system includes the computer-readable storage medium of claim 9, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
Citation Information
Patent Citations
Manufacturing process multivariate quality diagnosis classifier based on information entropy
CN108052087A
New energy station monitoring data quality evaluation method and system based on multi-source data
CN119066541A