Interactive Data Modeling for Denormalized Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The analysis of denormalized data is hindered by duplicated or redundant data, making it difficult to identify and manage relationships effectively, which complicates data modeling and reasoning processes.

Innovation Solution

A system that uses processors to determine relationships between data fields in denormalized data sources, employing evaluators to identify redundant fields, associate confidence scores, and rank relationships, allowing for the display and modification of data models to remove redundant fields and indicate conflicts, thereby facilitating the creation of normalized data models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If denormalized data sources are used to store data, then storage efficiency and data accessibility are improved, but data redundancy and analysis complexity increase

Engineering Contradiction:
Improvedata accessibilityVSAvoiddata redundancy
Core Design Contradiction:
ProductivityVSLoss of substance

Solution Approach 1:

The system segments denormalized data into normalized data models by identifying and separating redundant fields. Evaluators divide the data source into distinct data objects and relationships, organizing them into a structured normalized format that eliminates duplication while maintaining accessibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts redundant data fields from denormalized sources by using evaluators to identify and remove duplicate information. This extraction process isolates meaningful relationships and creates a cleaned normalized data model that preserves essential data while eliminating redundancy.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If automated relationship identification is implemented, then data modeling efficiency is improved, but system complexity increases

Engineering Contradiction:
Improvedata modeling efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system employs dynamic evaluators that can be selectively activated based on specific data characteristics and modeling requirements. Different evaluators implement various relationship identification strategies, allowing the system to adapt its complexity level to match the task at hand rather than maintaining fixed high complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The evaluators automatically analyze data sources and identify relationships without requiring manual intervention. The system self-configures by selecting appropriate evaluators based on data characteristics, performing relationship identification autonomously while managing its own complexity through adaptive evaluator selection.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If multiple evaluators are used to identify relationships, then relationship accuracy is improved, but processing time increases

Engineering Contradiction:
Improverelationship accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies evaluators selectively rather than exhaustively to all data scenarios. Based on data characteristics and confidence thresholds, the system determines when partial evaluation is sufficient, avoiding unnecessary application of all evaluators and thus reducing processing time while maintaining adequate relationship identification accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11422985B2Interactive data modeling
Publication Date: 2022.08.23 TABLEAU SOFTWARE INC
  • US11422985B2 patent drawing
  • US11422985B2 patent drawing
  • US11422985B2 patent drawing

AI summary

Embodiments are directed to managing data models. A data source that includes records may be provided. Source fields may be determined based on the records and the source fields may be displayed in a source panel. A data model that includes a source data object may be displayed. Relationships between the source fields may be determined based on values in the records. In response to providing a relationship between the source fields, a data object that includes a key field and one or more data fields that correspond to the relationship may be generated. The data model may be modified to include the data object and to remove the source fields that correspond to the data fields.