Data Analysis Using Parent-Child Factors for Causal Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis methods struggle to accurately identify the factors causing changes in data, particularly when dealing with large volumes of data and categorical variables with many levels, leading to models with good fitness being misinterpreted and important factors being overlooked.
Innovation Solution
A data analysis apparatus that calculates an index value indicative of conditional dispersion between data elements, extracts parent elements with high association, and outputs an analysis result to identify parent-child relationships, thereby preventing important factors from being overlooked.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If regression analysis is executed with many data elements as explanatory variables to search for factors of data change, then the coverage of potential factors is improved, but the number of models with good fitness increases making it difficult to determine the correct answer
Solution Approach 1:
The patent introduces an intermediary mechanism (parent-child relationship structure) between the data elements and the regression analysis. By organizing data elements into hierarchical parent-child relationships, the system mediates the selection process, allowing only parent elements to be used as explanatory variables. This intermediary structure filters out redundant child elements that would otherwise create multiple equivalent models, thus resolving the complexity of model selection while maintaining comprehensive factor coverage.
2Reliability
If data elements with a large number of levels are used as explanatory variables, then the model fitness tends to be better, but important factors with a small number of levels may be overlooked
Solution Approach 1:
The patent segments the data elements into hierarchical levels (parent elements and child elements). By dividing the explanatory variables into these segments, the system can evaluate child elements (which may have many levels and high fitness) against their parent elements. This segmentation allows the identification of parent elements that represent broader categories, ensuring that important factors with fewer levels are not overlooked even when child elements with many levels show better model fitness.
3Ease of operation
If the score of data with a large number of levels is held down based on regularization or model selection, then the ranking balance is improved, but data elements that actually have an effect may still rank higher than they should
Solution Approach 1:
Instead of holding down the scores of data elements with many levels (conventional approach), the patent inverts the approach by using parent-child relationships to adjust rankings. The system calculates the difference in fitness between child elements and their parent elements, and uses this inverted logic to determine final rankings. This ensures that even if child elements naturally have better fitness due to having more levels, their parent elements can be identified as the true factors, correcting the ranking to reflect actual importance rather than just model fitness.
Data Source
AI summary
According to one embodiment, a data analysis apparatus includes processing circuitry. The processing circuitry acquires a data element group composed of a plurality of data elements, calculates, in regard to a first data element in the data element group, an index value indicative of a conditional dispersion from another data element in the data element group, extracts, based on the index value, a parent element having a high association with the first data element from the data element group, and outputs an analysis result including first information relating to the first data element and the parent element.


