Dynamic Schema Analytics for Changing Data Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current analytics tools, such as OLAP and ROLAP engines, are inadequate for handling dynamically-changing data models and fail to provide efficient multidimensional data analytics across multiple data sources, leading to inefficiencies in analyzing advertising campaign performance and resulting in revenue losses for businesses.
Innovation Solution
A system and method for providing big data analytics that parses user queries into sub-queries based on a logical data schema, sends them to multiple data stores selected by a physical data schema, and combines sub-result datasets using aggregation and join operations, enabling efficient analysis of dynamically-changing data models across multiple sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If OLAP or ROLAP engines are used for data analysis, then multidimensional data analytics can be provided, but they cannot efficiently handle dynamically-changing data models from multiple data sources
Solution Approach 1:
The system employs a dynamic data model that automatically adapts to changing data structures from multiple sources. The schema evolution mechanism allows the system to accommodate new data formats and structures without requiring manual intervention or query rewriting, enabling both adaptability and maintained analysis efficiency through automated schema detection and transformation.
Solution Approach 2:
The patent introduces an intermediary layer between multiple data sources and the analysis engine. This layer includes a unified schema definition and automatic data transformation mechanisms that mediate between diverse data formats and the analytical processing, allowing the system to handle dynamic data models while maintaining efficient query processing through standardized intermediate representations.
2Loss of information
If data is gathered from multiple ad-serving companies and social media websites, then comprehensive campaign analytics can be obtained, but the volume and format complexity make analysis almost impossible using existing tools
Solution Approach 1:
The system implements a universal data processing framework that can handle multiple data sources with different formats through a single unified interface. The platform provides multi-functional capabilities including automated schema detection, format normalization, and adaptive query processing that work across diverse data sources, reducing system complexity while maintaining comprehensive analytics coverage.
Solution Approach 2:
The patent utilizes parameter changes in the data model to accommodate different data sources. By implementing a flexible schema with configurable parameters that can adapt to various data formats and structures, the system maintains a consistent processing framework while handling the diversity of input data, thereby reducing complexity through parameterized adaptability rather than multiple specialized systems.
3Measurement precision
If ROLAP engines require precise queries to be written based on the data model, then accurate performance analysis can be retrieved, but database administrators must write and maintain queries which reduces interactivity and insight gathering
Solution Approach 1:
The system implements self-service capabilities where the data analysis platform automatically adapts to user needs without requiring manual query writing or database administration. The automated schema detection and dynamic query generation allow users to interact with the system through natural interfaces, with the system automatically handling the complexity of precise data retrieval, thereby maintaining accuracy while dramatically improving ease of operation and interactivity.
Data Source
AI summary
A system and method for providing big data analytics responsive to dynamically-changing data models are provided. The method includes parsing, based on a logical data schema, a user query into a plurality of sub-queries; sending the plurality of sub-queries to a plurality of data stores, wherein each data store is selected based on a physical data schema of a dynamic data schema; receiving a plurality of sub-result datasets, wherein each sub-result dataset corresponds to a sub-query; and combining the plurality of sub-result datasets into a single resulting data set based on a logical schema of the dynamic data schema, wherein the combining includes at least any of: an aggregation operation and a join operation.


