Iterative Parameter Exclusion for Chemical Process Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis methods for chemical and biological processes face challenges in efficiently filtering out non-significant process parameters, leading to complex and cumbersome interpretation of results, especially when dealing with large datasets from batch processes.
Innovation Solution
A computer-implemented method that iteratively calculates correlation values and confidence intervals for process parameters, excludes non-significant parameters based on average ratios, and uses partial least squares (PLS) or orthogonal PLS (OPLS) analysis to construct models representing relationships between process parameters and output values, thereby simplifying the data set and improving interpretability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all process parameters are included in the data analysis, then the completeness of the data set is improved, but the complexity of interpretation and model construction increases
Solution Approach 1:
The patent extracts and removes non-significant process parameters from the data set by calculating correlation values and confidence intervals for each parameter, then iteratively excluding parameters with the smallest average ratio of correlation to confidence. This extraction process eliminates irrelevant information while preserving significant parameters, thereby reducing interpretation complexity without losing important data.
Solution Approach 2:
The patent transforms the data set by changing the number of parameters through iterative exclusion. By dynamically adjusting the parameter set based on statistical criteria (correlation values and confidence intervals), the method adapts the data complexity to an optimal level for interpretation while maintaining the essential information needed for accurate process modeling.
2Loss of information
If non-significant parameters are not excluded, then the completeness of the analysis is maintained, but the identification of causal relationships becomes difficult
Solution Approach 1:
The patent extracts significant parameters from the complete data set by calculating correlation values between each parameter and process output, then determining confidence intervals. Parameters with low correlation or low confidence are systematically removed, leaving only those that contribute meaningfully to causal relationship identification. This extraction makes causal detection feasible by eliminating noise.
Solution Approach 2:
The patent applies different quality criteria to different parameters based on their individual statistical properties. Each parameter is evaluated locally using its specific correlation value and confidence interval, allowing the method to retain high-quality parameters while excluding low-quality ones. This local quality assessment enables precise identification of causal parameters without being overwhelmed by irrelevant data.
3Loss of information
If the data set includes many process parameters, then the comprehensiveness of process monitoring is improved, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary statistical analysis by calculating correlation values and confidence intervals for all parameters before the main modeling process. This preliminary action identifies and excludes non-significant parameters in advance, reducing the data set size before computationally intensive model construction. This pre-processing step significantly reduces processing time while maintaining comprehensive monitoring of significant parameters.
Solution Approach 2:
The patent applies partial action by selectively analyzing and retaining only the most significant parameters rather than processing all parameters equally. By focusing computational resources on parameters that meet statistical significance criteria, the method achieves effective process monitoring with reduced computational effort and shorter processing time.
4Measurement precision
If all parameters are retained in the model, then the accuracy of the model may be improved, but the interpretability and understanding of causal relationships deteriorates
Solution Approach 1:
The patent extracts parameters that have both statistically significant correlation with process output and sufficient confidence from the complete parameter set. By removing parameters with low correlation or low confidence, the method creates a streamlined model that maintains accuracy through retention of significant parameters while improving interpretability through elimination of irrelevant ones.
Solution Approach 2:
The patent optimizes the parameter set by changing the number and composition of parameters based on statistical criteria. The iterative process of calculating correlation values, determining confidence intervals, and excluding parameters with the smallest average ratios dynamically adjusts the model to achieve the optimal balance between accuracy and interpretability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A computer-implemented method is provided for analyzing data obtained with respect to a chemical and/or biological process. The method comprises: obtaining (S10) a result of statistical data analysis on a data set including the data obtained with respect to the chemical and/or biological process; calculating (S20), for values of each process parameter obtained at each group of corresponding time points during a plurality of batch processes of the chemical and/or biological process, a ratio of a correlation value to a confidence value of the correlation value, the correlation value indicating a correlation between the values of the process parameter and at least one process output value; calculating (S30), for each process parameter, an average of absolute values of the ratios calculated for the values of the process parameter obtained at different groups of the corresponding time points during the plurality of batch processes; excluding (S40), from the data set, the values of one of the process parameters having a smallest one of the averages calculated for the process parameters; and iterating, until at least one specified condition is met (S50), the steps of obtaining (S10) the result of the statistical data analysis, calculating (S20) the ratio, calculating (S30) the average of the absolute values of the ratios and excluding (S40) the values of the one of the process parameters.