Condenser Framework for Data Lake Redundancy Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data lakes lack the functionality to provide comprehensive, user-friendly views and data manipulation capabilities, leading to inefficiencies and resource wastage due to duplicate records across disparate source systems within an organization.

Innovation Solution

A system comprising a view creation framework, snapshot load framework, and condenser framework is introduced to migrate data from source systems to a data lake, enabling 360-degree data views, automating view generation, and condensing data to reduce redundancy, thereby enhancing data accessibility and performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If data is stored in multiple source systems (systems of record), then data accessibility for different departments is improved, but data duplicity and resource wastage occur

Engineering Contradiction:
Improvedata accessibilityVSAvoidresource wastage
Core Design Contradiction:
Ease of operationVSLoss of substance

Solution Approach 1:

The patent combines data from multiple source systems into a single data lake, merging previously分散 data storage and management operations. This consolidation eliminates duplicate records across departments while maintaining centralized accessibility, directly resolving the contradiction between data accessibility and resource wastage

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If data is stored in multiple source systems, then departmental data autonomy is maintained, but data discrepancies increase

Engineering Contradiction:
Improvedepartmental data autonomyVSAvoiddata consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces a data lake as an intermediary layer between source systems and end users. This mediator consolidates data from multiple autonomous source systems, applying standardization and validation processes that ensure data consistency while preserving the autonomy of original source systems

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If comprehensive data views are provided, then user-friendly data access is improved, but system complexity increases

Engineering Contradiction:
Improveuser-friendly data accessVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the data access system into distinct layers: source systems, data lake, and view generation framework. This segmentation allows complex data consolidation and view generation processes to be isolated from end users, providing user-friendly access while managing system complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If data manipulation capabilities are added to data lake, then functionality is improved, but processing overhead increases

Engineering Contradiction:
Improvedata manipulation capabilityVSAvoidprocessing overhead
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent performs data manipulation operations such as consolidation, standardization, and view generation in advance during the data loading process. This preliminary action prepares data for future queries without requiring intensive processing during actual data access, reducing processing overhead while maintaining enhanced functionality

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11308065B2Condenser framework
Publication Date: 2022.04.19 BANK OF AMERICA CORP
  • US11308065B2 patent drawing
  • US11308065B2 patent drawing
  • US11308065B2 patent drawing

AI summary

In entity transition from legacy systems to a big data distributed data platform, numerous system-based architectural gaps have surfaced. There exists a need for a bridge component for each of the architectural gaps in order to support the entity transition to the big data distributed data platform. These bridge components include a variety of frameworks that are configured to automate certain processes that are needed for the transition. These processes have only become necessary as a result of the Hadoop platform. The automated processes include a snapshot load platform. The snapshot load platform enables the addition of a new view to the historical tables. The platform includes replacing the entire table in a truncated scenario. The platform includes replacing cases in a refresh or update scenario.