Machine Learning Dataset Consolidation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database management systems face challenges in efficiently managing datasets due to duplicate data storage, which increases infrastructure costs and can lead to inaccurate data analysis.
Innovation Solution
A computing system that utilizes machine learning processes to detect data redundancies across multiple datasets, involving entity data processing, validation, and training of model architectures to identify data similarities and facilitate dataset consolidation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplicate datasets are stored to various databases and servers, then data availability for multiple teams is improved, but infrastructure costs increase
Solution Approach 1:
The patent consolidates multiple duplicate datasets into a single centralized dataset, merging redundant storage locations into one unified data repository that serves multiple teams, thereby reducing infrastructure costs while maintaining data availability
Solution Approach 2:
The centralized dataset serves as a universal data source for multiple teams and projects, allowing one dataset to fulfill the data needs of various teams simultaneously, eliminating the need for separate duplicate datasets for each team
2Ease of operation
If duplicate datasets are stored to various databases and servers, then data access for different projects is improved, but data accuracy deteriorates
Solution Approach 1:
By merging duplicate datasets into a single centralized dataset, the patent eliminates inconsistencies between duplicates, ensuring that all teams access the same accurate data while maintaining ease of access through the unified repository
Data Source
AI summary
Systems and methods receive input(s) facilitating dataset management, the input(s) initiating a machine learning process configured to detect data redundancies of two or more datasets. Entity data stored to entity data storage location(s) are accessed and processed to conform with formatting requirements for the machine learning process. Validation is performed on the processed entity data, the validation ensuring the processed entity data satisfy the formatting requirements, where the validation produces training data that is inserted into an iterative training and testing loop. A model architecture is trained, based on weights and calculations, using the training data in the iterative training and testing loop to detect data redundancies, the training including predicting a target variable and iteratively adjusting the weights and calculations during each subsequent iteration to improve predictability of the target variable, where the model architecture is trained to identify data similarities among the two or more datasets.


