Machine Learning Data Structuring for Automated Dimensional Data Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Dimensional data modeling in data repositories is a manual, time-consuming process that requires deep understanding of business context and extensive data analysis, and any modifications to datasets necessitate manual adjustments, making it inefficient for automating data structuring.
Innovation Solution
A machine learning-based data structuring system that automates the dimensional data modeling process by obtaining datasets, labeling data based on importance, classifying attributes and measures, determining associations using primary and foreign keys, SQL logs, and fuzzy string matching, assigning weightages, validating associations, clustering datasets, and generating actionable insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual dimensional data modeling is used, then data analysis depth and breadth are improved, but time consumption and labor effort increase significantly
Solution Approach 1:
The system enables self-service dimensional data modeling by automatically detecting dimensions, attributes, and hierarchies from source datasets without requiring manual intervention. The machine learning algorithms autonomously analyze data patterns, identify relationships, and construct dimensional models, allowing the system to serve itself rather than requiring continuous human expertise for each modeling task.
Solution Approach 2:
The patent replaces the mechanical manual process of dimensional data modeling with an automated machine learning-based system. Instead of manually analyzing data, crafting hierarchies, and establishing relationships, the system uses algorithms to automatically perform these tasks, substituting human cognitive and manual operations with computational processes that achieve similar or superior results more efficiently.
2Manufacturing precision
If manual adjustments are made for dataset modifications, then data structure accuracy is maintained, but process efficiency deteriorates
Solution Approach 1:
The system incorporates feedback mechanisms where machine learning models continuously monitor and learn from changes in datasets. When new data is introduced or existing data is modified, the system automatically detects these changes, updates dimensional models accordingly, and maintains accuracy without requiring manual review or adjustment. The feedback loop ensures the model adapts to data evolution while maintaining structural integrity.
Solution Approach 2:
The dimensional data modeling system transitions from a static manual process to a dynamic automated system that continuously adapts to data changes. The machine learning algorithms dynamically adjust model parameters, detect schema changes, and reconfigure relationships in real-time as new datasets are introduced or existing ones are modified, maintaining both accuracy and efficiency throughout the process.
3Loss of time
If machine learning automation is implemented, then time consumption is reduced, but system complexity increases
Solution Approach 1:
The system segments the complex dimensional data modeling process into distinct functional modules: data ingestion, dimension detection, attribute identification, hierarchy construction, and model validation. Each module is handled by specialized machine learning algorithms that work independently but coordinate through standardized interfaces, making the overall complex system more manageable and easier to implement while achieving automation benefits.
Data Source
AI summary
A machine learning based data structuring method for automating dimensional data modelling process in data repositories is disclosed. The ML-based data structuring method includes obtaining datasets from databases; classifying the data comprising attributes and measures based on historical data using a ML-based classifier model; determining associations between the datasets using primary and foreign keys, SQL logs, usage of the datasets in creating transformations, and performance of fuzzy string match; assigning weightages to the determined associations between the datasets based on utilization of the determined associations using weighted network graphs; validating the determined associations between the datasets based on recurrent utilization of the determined associations; clustering the datasets based on the validated associations between the datasets by detecting dimensional models in the weighted network graphs; and generating actionable insights on each of the clustered datasets by performing exploratory data analysis, influencer analytics, and forecasting of the data.


