Correlation Tolerance Limit Setting via Repetitive Cross-Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for setting tolerance limits in correlation fitting are prone to intentional or unintentional distortion of data characteristics, leading to unquantified risks and increased costs due to over-fitting and differences in design characteristics between training and validation datasets.
Innovation Solution
A system utilizing repetitive cross-validation with a variable extraction unit, normality test unit, and DNBR limit unit to iteratively extract variables, perform correlation fitting, and determine tolerance limits, thereby preventing and quantifying distortion risks through iterative data partitioning and statistical validation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional data partitioning methods are used for correlation fitting, then the process is simple and quick, but intentional or unintentional distortion of data characteristics occurs and cannot be prevented
Solution Approach 1:
The patent segments the correlation fitting process into multiple iterative rounds, each with distinct training and validation sets. This segmentation prevents distortion by ensuring that different data subsets are used in each iteration, thereby maintaining data characteristic integrity while achieving reliable correlation fitting results.
Solution Approach 2:
The patent performs preliminary actions by pre-defining multiple training and validation sets before the correlation fitting process begins. This preliminary preparation ensures that data characteristics are preserved from the outset, preventing both intentional and unintentional distortion during the fitting process.
2Loss of time
If limited number of cases are used for correlation fitting, then the process is faster and less resource-intensive, but over-fitting risk increases and cannot be properly validated
Solution Approach 1:
The patent implements continuous validation across multiple iterative rounds, where each round contributes to the overall validation of the correlation model. This continuous action ensures thorough over-fitting prevention without requiring excessive time, as each iterative round builds upon previous results and progressively validates the model.
Solution Approach 2:
The patent incorporates feedback mechanisms where validation results from each iterative round inform subsequent fitting processes. This feedback loop allows the system to adjust and improve the correlation model continuously, preventing over-fitting while maintaining efficient processing time through intelligent iteration.
3Measurement precision
If separately independent testing datasets with same design characteristics are used, then validation is more rigorous, but cost for additional production of testing data increases
Solution Approach 1:
The patent creates testing datasets that serve multiple functions: they validate the correlation model, assess over-fitting risk, and evaluate data characteristic integrity. This multi-functionality allows rigorous validation without proportionally increasing data production costs, as the same datasets are used for multiple validation purposes across different iterative rounds.
4Device complexity
If simple statistics analysis is used for setting tolerance limit, then the process is simpler and faster, but the influence of data characteristic distortion cannot be quantified
Solution Approach 1:
The patent applies a moderate level of statistical analysis that is sufficient to quantify distortion influence without excessive complexity. By performing statistical tests at appropriate stages of the iterative process, the system quantifies the influence of data characteristic distortion while maintaining reasonable process complexity and avoiding unnecessary computational overhead.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to a correlation tolerance limit setting system using repetitive cross-validation and a method therefor and, to a correlation tolerance limit setting system using repetitive cross-validation and a method therefor, the system and the method being for preventing data characteristic distortion due to accidental or human interference in correlation optimization and tolerance limit setting, and for preventing a risk created thereby, or for quantifying the influence thereof. According to the present invention, the correlation tolerance limit setting system using repetitive cross-validation comprises: a variable extraction unit for classifying training sets and validation sets, and optimizing a correlation coefficient so as to extract variables; a normality verification unit for verifying normality in accordance with the variable extraction result; a DNBR limit unit for verifying whether the same population is present according to the normality, and determining a tolerance limit of a departure from a nucleate boiling ratio by using a tolerance limit distribution for a departure from nucleate boiling; and a control unit.