Ad Hoc Data Exploration Tool for Multi-Source Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current enterprise databases are inadequate for simultaneous in-place analysis by data scientists and business analysts, requiring significant resources and user time to gather, clean, and organize data from multiple sources, limiting flexibility and convenience in data exploration and analysis.
Innovation Solution
A data exploration tool with a query pane, results pane, and analysis interface that allows users to connect to multiple data sources, store query results in a containerized temporary storage space, and apply dimensions as columns or filters within a scorecard, enabling flexible data manipulation and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If enterprise databases are used to gather data from multiple sources, then data can be obtained from various sources, but significant computing resources and user time are required to clean and organize the data
Solution Approach 1:
The patent introduces an intermediary data preparation layer between the enterprise database and analysis tools. This layer automatically performs data cleaning, organization, and validation tasks, serving as a mediator that transforms raw multi-source data into analysis-ready formats without requiring significant user intervention or computing resources.
Solution Approach 2:
The system performs preliminary data preparation actions automatically when data is first retrieved from enterprise databases. Data cleaning, validation, and organization are executed in advance before the user needs the data for analysis, eliminating the need for manual data preparation later.
2Ease of operation
If data scientists use existing analysis tools to execute analytic code, then data science models can be applied, but flexibility in data manipulation is limited
Solution Approach 1:
The patent creates a universal data exploration environment that combines multiple functions: it allows execution of analytic code (Python/R scripts), provides interactive data manipulation capabilities, enables visualization, and supports collaborative work. This multi-functional platform serves both data scientists needing code execution and those requiring flexible data exploration.
Solution Approach 2:
The system provides dynamic data exploration capabilities where users can interactively modify queries, adjust parameters, and explore data relationships in real-time. The environment adapts to user needs by allowing flexible data manipulation while maintaining the ability to execute predefined analytic code.
3Ease of operation
If business analysts use analysis tools to generate dashboards and reports, then graphical data analysis is available, but the tools are constrained by statically defined data connections
Solution Approach 1:
The patent enables dynamic data connections that can be modified on-the-fly without requiring static pre-definition. Business analysts can explore data relationships interactively, create ad-hoc connections between data sources, and adjust their analysis approach as insights emerge, rather than being constrained by predetermined data models.
4Quantity of substance
If distributed databases are used to store enterprise data, then data can be stored across multiple nodes, but significant resources are required to gather data from different nodes
Solution Approach 1:
The patent introduces an intermediary data retrieval and preparation layer that automatically manages data gathering from distributed database nodes. This mediator handles the complexity of querying multiple nodes, aggregating results, and preparing data for analysis, shielding users from the underlying distributed system complexity.
Data Source
AI summary
The disclosed application relates to a tool by which a user may create a cloud workspace that includes a data memory space, as well as a tool for automatically identifying ad-hoc analyses on that data. The solution allows a user to connect to data sources using SQL or GUI tools, combine data from different data sources, prepare and clean the data, mine the data for insights, and move that data into downstream reporting tools for visualization. The system is linked to a code repository to allow data scientists to execute code from the code repository in trial data spaces, investigate that data, and prepare more in-depth analytics for downstream reporting tools.


