Variable Processing for Multicollinearity in Statistical Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In constructing statistical models, multicollinearity often leads to unstable models, and while stratification can improve accuracy, it reduces the number of included samples, making it difficult to build a reliable model.
Innovation Solution
An information processing apparatus that constructs models using a concept network, determines the correlation between variables, and performs creation, integration, or stratification processing based on these correlations to address multicollinearity, thereby improving model stability and sample inclusion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If stratification is minutely performed to improve accuracy and solve multicollinearity, then model accuracy is improved, but the number of included samples decreases
Solution Approach 1:
The patent applies segmentation by dividing variables into multiple categories based on their correlation relationships. Instead of minutely stratifying the entire dataset which reduces sample size, the method segments variables into groups (e.g., high correlation groups) and processes them differently. This allows solving multicollinearity by selecting representative variables from each segment while maintaining adequate sample sizes for model construction.
Solution Approach 2:
The patent implements partial action by selectively applying stratification only to variable selection rather than to the entire data sampling process. The method performs correlation analysis on all samples, identifies multicollinear variable groups, and then applies stratification logic only when selecting representative variables from these groups. This partial application of stratification solves multicollinearity without excessively reducing the overall sample size available for model construction.
2Reliability
If variables are highly correlated in a regression model, then model construction becomes difficult, but reducing variable correlation may lose important information
Solution Approach 1:
The patent introduces an intermediary mechanism - correlation analysis and variable classification - between the raw variables and the final regression model. Instead of directly including all variables or arbitrarily selecting them, the method uses correlation coefficients as an intermediary to identify multicollinear groups, then selects representative variables from each group. This intermediary process maintains model stability by reducing multicollinearity while preserving important information through systematic variable selection based on correlation relationships.
Solution Approach 2:
The patent applies parameter changes by transforming the variable selection problem into a correlation-based classification problem. The method changes the parameter perspective from individual variable importance to pairwise correlation relationships, using correlation coefficients to reorganize variables into segments. This parameter transformation allows the system to identify and handle multicollinearity systematically, selecting representative variables that maintain information content while reducing correlation-induced instability in the regression model.
Data Source
AI summary
An information processing apparatus has a model construction unit that constructs a model represented using a plurality of variables corresponding to a plurality of classes, an evaluated value calculation unit that calculates an evaluated value of the model constructed by the model construction unit, a correlation specification unit that specifies a correlation between some variables among the plurality of variables based on the calculated evaluated value, a variable processing determination unit that determines whether to perform at least one of creation, integration, and stratification of at least some variables among the plurality of variables based on the correlation specified by the correlation specification unit, and a variable processing unit that performs at least one of creation processing, integration processing, and stratification processing of the variables when the at least one of the creation, the integration, and the stratification of the variables is determined to be performed.


