Condenser Framework for Data Lake Redundancy Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data lakes lack the functionality to provide comprehensive, user-friendly views and data manipulation capabilities, leading to inefficiencies and resource wastage due to duplicate records across disparate source systems within an organization.
Innovation Solution
A system comprising a view creation framework, snapshot load framework, and condenser framework is introduced to migrate data from source systems to a data lake, enabling 360-degree data views, automating view generation, and condensing data to reduce redundancy, thereby enhancing data accessibility and performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data is stored in multiple source systems (systems of record), then data accessibility for different departments is improved, but data duplicity and resource wastage occur
Solution Approach 1:
The patent combines data from multiple source systems into a single data lake, merging previously分散 data storage and management operations. This consolidation eliminates duplicate records across departments while maintaining centralized accessibility, directly resolving the contradiction between data accessibility and resource wastage
2Adaptability or versatility
If data is stored in multiple source systems, then departmental data autonomy is maintained, but data discrepancies increase
Solution Approach 1:
The patent introduces a data lake as an intermediary layer between source systems and end users. This mediator consolidates data from multiple autonomous source systems, applying standardization and validation processes that ensure data consistency while preserving the autonomy of original source systems
3Ease of operation
If comprehensive data views are provided, then user-friendly data access is improved, but system complexity increases
Solution Approach 1:
The patent segments the data access system into distinct layers: source systems, data lake, and view generation framework. This segmentation allows complex data consolidation and view generation processes to be isolated from end users, providing user-friendly access while managing system complexity through modular architecture
4Adaptability or versatility
If data manipulation capabilities are added to data lake, then functionality is improved, but processing overhead increases
Solution Approach 1:
The patent performs data manipulation operations such as consolidation, standardization, and view generation in advance during the data loading process. This preliminary action prepares data for future queries without requiring intensive processing during actual data access, reducing processing overhead while maintaining enhanced functionality
Data Source
AI summary
In entity transition from legacy systems to a big data distributed data platform, numerous system-based architectural gaps have surfaced. There exists a need for a bridge component for each of the architectural gaps in order to support the entity transition to the big data distributed data platform. These bridge components include a variety of frameworks that are configured to automate certain processes that are needed for the transition. These processes have only become necessary as a result of the Hadoop platform. The automated processes include a snapshot load platform. The snapshot load platform enables the addition of a new view to the historical tables. The platform includes replacing the entire table in a truncated scenario. The platform includes replacing cases in a refresh or update scenario.


