Data Flow Orchestrator for Analytics Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data scientists face difficulties in deploying data analytics applications due to the lack of standard components and the need to select from multiple types of data stores, requiring expertise in database selection and schema design, which is often beyond their knowledge and necessitating multiple engineer terminals and software packages for deployment.
Innovation Solution
A data flow orchestrator creates a data integration plan and application logical topology based on user-defined analytics plans, allowing data scientists to select databases and design schemas directly, generating executable deployment plans for cloud deployment without requiring data engineer or application engineer terminals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data scientists use multiple engineer terminals and specialized software packages for deployment, then deployment completeness is improved, but system complexity and deployment difficulty increase
Solution Approach 1:
The patent combines multiple engineer terminals (data engineer terminal, application engineer terminal) and their specialized software packages into a single integrated terminal. This terminal provides unified access to all deployment functions including data store selection, schema design, and application deployment, eliminating the need for multiple separate systems while maintaining complete deployment capabilities
Solution Approach 2:
The integrated terminal is designed to perform multiple functions that previously required separate specialized tools. It can select data stores, design schemas, create application logical topology, and execute deployment plans all within one interface, making the system more versatile and easier to use while reducing complexity
2Adaptability or versatility
If data scientists select from N*M data store combinations, then data store selection flexibility is improved, but selection complexity and time increase
Solution Approach 1:
The patent introduces an intermediary component that automatically matches data types with appropriate data store configurations. Instead of requiring data scientists to manually select from N*M combinations, the system provides intelligent recommendations and automated selection based on the specific data analytics requirements, significantly reducing selection time while maintaining flexibility
Solution Approach 2:
The system enables self-service deployment by providing automated data store selection and schema generation based on the data analytics plan. The integrated terminal guides data scientists through the deployment process with automated assistance, reducing the need for expert knowledge and minimizing manual configuration time
3Adaptability or versatility
If data scientists manually export, transform and load data, then data deployment flexibility is improved, but deployment complexity increases
Solution Approach 1:
The patent implements preliminary action by automatically generating data integration plans and ETL (Extract, Transfer, Load) configurations before the actual deployment process. The system prepares data transformation rules and integration pathways in advance based on the data analytics plan, eliminating the need for manual data handling and simplifying the deployment process
Data Source
AI summary
Example implementations are directed to a system and method to reduce deployment cost of data analytics application by designing both an application deployment plan and data integration plan, implementing the plans into an application template automatically and deploying application components and data in accordance with the desired implementation. Through example implementations, the need for separate terminals for a data engineer and an application engineer can be eliminated.


