Data Warehousing Standardized Format Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data warehouse systems face limitations in data quality, integrity, and performance due to format discrepancies, incomplete data, record merging issues, inadequate relationship identification, insufficient real-time alert capabilities, and inability to maintain persistent queries, leading to inaccurate and unreliable results.

Innovation Solution

A method and system that processes data into a database by converting it to a standardized format, analyzing and enhancing identifiers, creating hash keys, and utilizing algorithms to match and store records while maintaining attribution, enabling real-time relationship identification and alert issuance based on user-defined rules, and supporting persistent queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If records are merged to eliminate duplicates, then data quality improves, but data integrity deteriorates because original records cannot be separated later

Engineering Contradiction:
Improvedata qualityVSAvoiddata integrity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent divides the merged record into multiple segments, each containing original record information with unique identifiers. This allows the system to maintain separate original records while presenting a unified view during matching operations, enabling both data quality improvement through deduplication and data integrity preservation through traceability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary matching record structure that mediates between original records and query results. This intermediary layer preserves the relationship information and original record data while enabling efficient matching and deduplication operations, allowing separation of concerns between data storage and data presentation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If transformation and enhancement processes are applied to received data, then data accuracy improves, but query consistency deteriorates when different algorithms are used

Engineering Contradiction:
Improvedata accuracyVSAvoidquery consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements a universal transformation and enhancement algorithm that serves multiple functions: processing received data, processing queries, and ensuring consistency across both operations. This single algorithm handles all data transformation needs, eliminating the problem of inconsistent results when different algorithms are used for different operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If real-time relationship identification is implemented, then alert responsiveness improves, but system complexity increases

Engineering Contradiction:
Improvealert responsivenessVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-establishing alert rules, thresholds, and relationship identification criteria before actual monitoring occurs. This preparation enables real-time relationship identification and alerting without requiring complex real-time computation during operation, as the framework and parameters are already in place.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8452787B2Real time data warehousing
Publication Date: 2013.05.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8452787B2 patent drawing
  • US8452787B2 patent drawing
  • US8452787B2 patent drawing

AI summary

A method and system for processing data into and in a database and for retrieving the processed data is disclosed. The data comprises identifiers of a plurality of entities. The method and system comprises: (a) processing data into and in a database, (b) enhancing received data prior to storage in a database, (c) determining and matching records based upon relationships between the records in the received data and existing data without any loss of data, (d) enabling alerts based upon user-defined alert rules and relationships, (e) automatically stopping additional matches and separating previously matched records when identifiers used to match records are later determined to be common across entities and not generally distinctive of an entity, (f) receiving data queries for retrieving the processed data stored in the database, (g) utilizing the same algorithm to process the queries and (h) transferring the processed data to another database that uses the same algorithm.