Automated Data Transfer Tool for Simplified Big Data Operations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data transfers in 'big data' systems remain complex and require significant capital investment and specialized IT training, limiting accessibility to personnel with lower levels of expertise, even with open-source solutions like Hadoop and Sqoop.
Innovation Solution
A system with coded logical rules that allows lesser-trained users to initiate data transfer operations by inputting fewer parameter values, with the tool automatically retrieving missing parameters from operational or transactional metadata storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional data transfer tools (Sqoop, Hadoop) are used, then data transfer capability is achieved, but complexity of operation increases and requires specialized IT training
Solution Approach 1:
The patent introduces an intermediary layer between the user and the complex Hadoop/Sqoop infrastructure. This intermediary automatically generates Sqoop commands based on simple user inputs (source table, target table, where clause), translating high-level data transfer requirements into low-level technical commands without requiring users to understand the underlying complexity.
Solution Approach 2:
The system performs self-service by automatically retrieving metadata (table schemas, column definitions, data types) from the database and using it to construct complete data transfer commands. The tool serves itself by generating the necessary technical parameters (delimited-input-fields, delimited-output-fields, target-table-columns) without human intervention, reducing the operational burden on users.
2Adaptability or versatility
If proprietary big data solutions (Oracle, Teradata) are implemented, then comprehensive data management capability is achieved, but cost increases significantly
Solution Approach 1:
The patent creates a universal interface that works across different data sources and destinations (relational databases, flat files, Hadoop) through a single tool. This multi-functional capability replaces the need for multiple proprietary solutions, providing comprehensive data management versatility while using open-source infrastructure to reduce costs.
Solution Approach 2:
The system uses open-source, freely available tools (Hadoop, Sqoop, MySQL) instead of expensive proprietary software. These open-source components can be freely deployed and discarded without licensing costs, providing comparable functionality to proprietary solutions at a fraction of the cost.
3Manufacturing precision
If complete parameter specification is required for data transfer, then accuracy of data transfer is ensured, but time required for operation increases
Solution Approach 1:
The system performs preliminary actions by automatically retrieving and storing metadata (table structures, column definitions, data types) before the actual data transfer operation. This pre-fetching of information eliminates the need for users to manually specify these parameters during the transfer operation, ensuring precision while reducing setup time.
Solution Approach 2:
The system uses feedback from the database metadata to automatically complete parameter specification. By querying the database schema and using the returned information to populate command parameters, the system ensures accurate data transfer configuration without requiring users to spend time manually specifying each parameter.
Data Source
AI summary
Systems, methods, and articles of manufacture provide for simplified and partially-automated data operation services, such as data transfer, storage, management, and analysis operations. Non-IT data consumers may, for example, initiate such data operations by providing only a subset of the required parameters for the operation, with the specially-coded system automatically fetching any missing parameters or values from one or more metadata stores and initiating the requested operation.


