Composite Relationship Discovery Framework for Data Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analytics systems face challenges in efficiently identifying and interpreting interactions between attributes in large datasets, making it difficult to derive actionable insights for decision-making.
Innovation Solution
The system employs a constrained composite interaction data mining approach to discover relationships between continuous and discrete features, using a feature selection component to identify key features and a composite relationship discovery analysis component to generate scores that rank features based on their relationships, enabling the identification of strong interactions and visualization of insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data mining methods are used to identify patterns in large datasets, then comprehensive pattern discovery is achieved, but the process becomes time-consuming and computationally infeasible
Solution Approach 1:
The patent segments the complex data mining process into distinct components: a feature selection component that identifies important features, and a composite relationship discovery analysis component that analyzes relationships between selected features. This segmentation reduces the computational complexity by focusing analysis on a subset of relevant features rather than all possible feature combinations, thereby reducing analysis time while maintaining pattern discovery accuracy.
Solution Approach 2:
The patent applies preliminary action by performing feature selection before conducting composite relationship discovery. The feature selection component pre-identifies and ranks important features based on their relevance to the target variable, so that subsequent analysis focuses only on these pre-selected features. This preliminary filtering step significantly reduces the search space and computational requirements for the subsequent relationship discovery phase.
2Loss of information
If comprehensive data analysis is performed to identify all attribute interactions, then complete knowledge discovery is achieved, but system complexity and computational requirements increase
Solution Approach 1:
The system segments the analysis into two distinct phases: feature selection and composite relationship discovery. Each phase has a specific function and operates independently, reducing overall system complexity. The feature selection component handles feature identification, while the composite relationship discovery component handles relationship analysis, allowing each component to be optimized separately and making the overall system more manageable.
Solution Approach 2:
The patent applies partial action by focusing analysis on the most relevant features identified through feature selection, rather than analyzing all possible feature interactions. By concentrating computational resources on a subset of high-priority features and their relationships, the system achieves sufficient knowledge discovery without the prohibitive complexity of exhaustive analysis of all attribute interactions.
3Loss of information
If detailed relationship analysis is performed between all features, then comprehensive insights are obtained, but computational resources and processing time increase
Solution Approach 1:
The feature selection component performs preliminary action by identifying and ranking features based on their importance to the target variable before the composite relationship discovery phase. This pre-ranking allows the subsequent analysis to focus computational resources on relationships involving high-priority features, thereby maintaining insight quality while significantly improving processing efficiency by avoiding analysis of relationships involving low-priority features.
Solution Approach 2:
The patent applies local quality by applying different levels of analysis depth to different features based on their importance. High-priority features identified through feature selection receive more thorough relationship analysis, while lower-priority features receive less analysis. This differentiated approach ensures that computational resources are allocated efficiently, with detailed analysis focused on features that contribute most to insight quality.
Data Source
AI summary
Systems and methods include reception of a set of data including continuous features and a discrete feature, each continuous feature associated with a plurality of values and the discrete feature associated with a plurality of discrete values, determine, for each continuous feature, a relationship factor representing a relationship between the discrete feature and the continuous feature based on the plurality of values associated with the continuous feature and the plurality of discrete values, identify one of the continuous features associated with a largest one of the determined relationship factors, generate, for each of the other features, a correlation factor representing a correlation between the continuous feature and the identified continuous feature, determine, for each of the continuous features other than the identified continuous feature, a composite relationship score based on the relationship factor and the correlation factor associated with the feature, and present a visualization associated with the discrete feature, the identified continuous feature, and a continuous feature associated with a largest composite relationship score.


