In-Database Analytic Flow Decoupling for Data Mining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining systems are proprietary and tightly coupled to specific databases, making models non-portable and reusable across different data sources and database engines, and require unnecessary data movement, which increases latency and processing overhead, while also posing security risks.
Innovation Solution
A data mining system that uses a massively parallel processing architecture to create a data miner integrated with databases, allowing analytic flows to be executed within the database environment, decoupling them from specific data sources and enabling model portability and efficient data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is moved out of the database environment to perform analytic flow operations, then data mining operations can be performed, but processing overhead and latency increase significantly
Solution Approach 1:
The patent merges the data storage environment with the analytic flow execution environment by allowing the database management system to directly execute analytic flows on stored data without requiring data to be moved to external processing systems. This integration eliminates data movement overhead and latency while maintaining full data mining functionality.
2Productivity
If data is moved out of the database environment for analytic processing, then analysis can be performed, but data security risks increase due to increased data fetch and store cycles
Solution Approach 1:
The patent combines the analytic flow execution environment with the database management system, allowing data processing to occur in-place within the secure database environment. This eliminates the need for data to be fetched and stored externally, thereby maintaining data security while enabling full analytic capabilities.
3Productivity
If present data miners are used that are specific to a particular database, then data mining can be performed, but models are not portable or reusable across different enterprises
Solution Approach 1:
The patent implements a universal analytic flow execution environment within the database management system that can execute standardized analytic flows across different database instances and enterprises. The analytic flows are designed to be database-agnostic, allowing models to be ported and reused across different organizations while maintaining full data mining functionality.
4Productivity
If data is moved wholesale for data mining operations, then complete data sets can be analyzed, but memory and processor requirements increase significantly
Solution Approach 1:
The patent merges the analytic processing capabilities with the database storage environment, enabling the database management system to execute analytic flows directly on the stored data using its existing memory and processor resources. This eliminates the need for separate high-resource processing systems while maintaining the ability to analyze complete data sets.
Data Source
AI summary
Embodiments are described for a system and method of providing a data miner that decouples the analytic flow solution components from the data source. An analytic-flow solution then couples with the target data source through a simple set of data source connector, table and transformation objects, to perform the requisite analytic flow function. As a result, the analytic-flow solution needs to be designed only once and can be re-used across multiple target data sources. The analytic flow can be modified and updated at one place and then deployed for use on various different target data sources.


