Data Lake Loading Method for Distributed Messaging Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems face challenges in efficiently processing and loading large amounts of data from distributed messaging systems like Apache KAFKA, particularly in accessing and managing unstructured data, which hinders data analysis and increases computational complexity.

Innovation Solution

An electronic apparatus and method that allows users to select specific data fields from a distributed messaging system, process them according to a predetermined format, and load the data into a data lake, enhancing data access convenience and analysis efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If all data fields from distributed messaging system are loaded into data lake, then data completeness is improved, but data processing time and computational complexity increase

Engineering Contradiction:
Improvedata completenessVSAvoiddata processing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent extracts only the necessary data fields from messages in the distributed messaging system based on user selection, rather than loading all fields. This selective extraction reduces the volume of data transferred and processed, directly addressing the contradiction between data completeness and processing time by allowing users to specify which fields are actually needed for their analysis purposes

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If user interface provides detailed data field selection options, then data loading flexibility is improved, but system complexity increases

Engineering Contradiction:
Improvedata loading flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the data field selection process into distinct, manageable components presented through a user interface. Users can selectively choose which data fields to load by interacting with organized field options, breaking down the complex task of data field management into simple selection actions. This segmentation maintains high flexibility while keeping the interface and system manageable

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11734284B2Method for loading data and electronic apparatus therefor
Publication Date: 2023.08.22 COUPANG CORP
  • US11734284B2 patent drawing
  • US11734284B2 patent drawing
  • US11734284B2 patent drawing

AI summary

The present disclosure relates to a method of loading data. The method includes checking a topic corresponding to a search word among a plurality of topics in response to acquiring a search word for a topic of a distributed messaging system from a user, checking a data format including one or more fields of a message loaded into a topic, and then loading data generated based on the checked data format and the read message into a data lake.