Interactive Data Modeling for Denormalized Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The analysis of denormalized data is hindered by duplicated or redundant data, making it difficult to identify and manage relationships effectively, which complicates data modeling and reasoning processes.
Innovation Solution
A system that uses processors to determine relationships between data fields in denormalized data sources, employing evaluators to identify redundant fields, associate confidence scores, and rank relationships, allowing for the display and modification of data models to remove redundant fields and indicate conflicts, thereby facilitating the creation of normalized data models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If denormalized data sources are used to store data, then storage efficiency and data accessibility are improved, but data redundancy and analysis complexity increase
Solution Approach 1:
The system segments denormalized data into normalized data models by identifying and separating redundant fields. Evaluators divide the data source into distinct data objects and relationships, organizing them into a structured normalized format that eliminates duplication while maintaining accessibility.
Solution Approach 2:
The system extracts redundant data fields from denormalized sources by using evaluators to identify and remove duplicate information. This extraction process isolates meaningful relationships and creates a cleaned normalized data model that preserves essential data while eliminating redundancy.
2Productivity
If automated relationship identification is implemented, then data modeling efficiency is improved, but system complexity increases
Solution Approach 1:
The system employs dynamic evaluators that can be selectively activated based on specific data characteristics and modeling requirements. Different evaluators implement various relationship identification strategies, allowing the system to adapt its complexity level to match the task at hand rather than maintaining fixed high complexity.
Solution Approach 2:
The evaluators automatically analyze data sources and identify relationships without requiring manual intervention. The system self-configures by selecting appropriate evaluators based on data characteristics, performing relationship identification autonomously while managing its own complexity through adaptive evaluator selection.
3Measurement precision
If multiple evaluators are used to identify relationships, then relationship accuracy is improved, but processing time increases
Solution Approach 1:
The system applies evaluators selectively rather than exhaustively to all data scenarios. Based on data characteristics and confidence thresholds, the system determines when partial evaluation is sufficient, avoiding unnecessary application of all evaluators and thus reducing processing time while maintaining adequate relationship identification accuracy.
Data Source
AI summary
Embodiments are directed to managing data models. A data source that includes records may be provided. Source fields may be determined based on the records and the source fields may be displayed in a source panel. A data model that includes a source data object may be displayed. Relationships between the source fields may be determined based on values in the records. In response to providing a relationship between the source fields, a data object that includes a key field and one or more data fields that correspond to the relationship may be generated. The data model may be modified to include the data object and to remove the source fields that correspond to the data fields.


