Dispersion Parameter Estimation Using Transformed Samples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for estimating dispersion parameters in statistical models with outliers often result in underestimation of variance, leading to inaccurate outlier detection and control of false discovery rates.
Innovation Solution
An information processing apparatus and method that transforms responses into samples depending on covariates and an unbiased parameter, maximizing the distribution to estimate dispersion parameters and calculate p-values for accurate outlier identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional parameter estimation methods are used on data containing outliers, then the estimation process is simple, but the accuracy of parameter estimates deteriorates significantly
Solution Approach 1:
The patent segments the data analysis process into distinct stages: first estimating location parameters (mean, median) using robust methods, then using these estimates to transform the data, and finally estimating the dispersion parameter from the transformed data. This segmentation allows each stage to focus on specific parameters, improving overall accuracy while managing complexity.
Solution Approach 2:
The patent introduces transformed samples as an intermediary step between the original contaminated data and the final dispersion parameter estimate. By transforming the data using robust location estimates first, the method creates a cleaner intermediate dataset that facilitates more accurate dispersion estimation without directly confronting the outliers in the original data.
2Measurement precision
If robust estimation methods are used to handle outliers, then the accuracy of parameter estimates improves, but the complexity of the estimation procedure increases
Solution Approach 1:
The patent performs preliminary estimation of location parameters (mean, median, or trimmed mean) before attempting to estimate the dispersion parameter. This preliminary action prepares the data by removing or reducing the influence of outliers, making the subsequent dispersion estimation more accurate and less complex than trying to estimate all parameters simultaneously from contaminated data.
Solution Approach 2:
The patent replaces direct mechanical optimization of complex robust estimation procedures with a two-step process: first compute simple robust location estimates using straightforward formulas, then use these to transform the data before dispersion estimation. This substitution reduces the mechanical complexity of the overall system while maintaining robustness.
3Measurement precision
If the dispersion parameter is underestimated due to outliers, then the computational efficiency is maintained, but the accuracy of outlier detection deteriorates
Solution Approach 1:
The patent incorporates feedback by using the estimated location parameters to transform the data, then using the transformed data to estimate the dispersion parameter, which in turn can be used to identify outliers. This feedback loop ensures that the dispersion estimate accounts for the presence of outliers through the transformation step, improving both accuracy and reliability.
Solution Approach 2:
The patent changes the parameter representation by working with transformed samples rather than original samples for dispersion estimation. By applying a transformation based on robust location estimates, the method changes the parameter space to one where the dispersion parameter can be estimated more reliably even in the presence of outliers, improving both accuracy and reliability.
Data Source
AI summary
An information processing apparatus (100) is disclosed. The information processing apparatus (100) includes an input means (102), a statistic calculation means (104) and an optimization means (106). The input means (102) receives input samples including responses and covariates. The statistic calculation means (104) transforms the responses into transformed samples using a function depending on the covariates and an unbiased parameter. A distribution of the transformed samples only depends on a dispersion parameter. The optimization means (106) maximizes a distribution of the transformed samples to determine an estimate of the dispersion parameter.


