Variable Processing for Multicollinearity in Statistical Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In constructing statistical models, multicollinearity often leads to unstable models, and while stratification can improve accuracy, it reduces the number of included samples, making it difficult to build a reliable model.

Innovation Solution

An information processing apparatus that constructs models using a concept network, determines the correlation between variables, and performs creation, integration, or stratification processing based on these correlations to address multicollinearity, thereby improving model stability and sample inclusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If stratification is minutely performed to improve accuracy and solve multicollinearity, then model accuracy is improved, but the number of included samples decreases

Engineering Contradiction:
Improvemodel accuracyVSAvoidnumber of samples
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies segmentation by dividing variables into multiple categories based on their correlation relationships. Instead of minutely stratifying the entire dataset which reduces sample size, the method segments variables into groups (e.g., high correlation groups) and processes them differently. This allows solving multicollinearity by selecting representative variables from each segment while maintaining adequate sample sizes for model construction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by selectively applying stratification only to variable selection rather than to the entire data sampling process. The method performs correlation analysis on all samples, identifies multicollinear variable groups, and then applies stratification logic only when selecting representative variables from these groups. This partial application of stratification solves multicollinearity without excessively reducing the overall sample size available for model construction.

Inventive Principle:
Principle #16Partial or excessive action

2Reliability

If variables are highly correlated in a regression model, then model construction becomes difficult, but reducing variable correlation may lose important information

Engineering Contradiction:
Improvemodel stabilityVSAvoidinformation loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary mechanism - correlation analysis and variable classification - between the raw variables and the final regression model. Instead of directly including all variables or arbitrarily selecting them, the method uses correlation coefficients as an intermediary to identify multicollinear groups, then selects representative variables from each group. This intermediary process maintains model stability by reducing multicollinearity while preserving important information through systematic variable selection based on correlation relationships.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by transforming the variable selection problem into a correlation-based classification problem. The method changes the parameter perspective from individual variable importance to pairwise correlation relationships, using correlation coefficients to reorganize variables into segments. This parameter transformation allows the system to identify and handle multicollinearity systematically, selecting representative variables that maintain information content while reducing correlation-induced instability in the regression model.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10592584B2Information processing apparatus, information processing method, and program
Publication Date: 2020.03.17 KK TOSHIBA
  • US10592584B2 patent drawing
  • US10592584B2 patent drawing
  • US10592584B2 patent drawing

AI summary

An information processing apparatus has a model construction unit that constructs a model represented using a plurality of variables corresponding to a plurality of classes, an evaluated value calculation unit that calculates an evaluated value of the model constructed by the model construction unit, a correlation specification unit that specifies a correlation between some variables among the plurality of variables based on the calculated evaluated value, a variable processing determination unit that determines whether to perform at least one of creation, integration, and stratification of at least some variables among the plurality of variables based on the correlation specified by the correlation specification unit, and a variable processing unit that performs at least one of creation processing, integration processing, and stratification processing of the variables when the at least one of the creation, the integration, and the stratification of the variables is determined to be performed.