Proactive Database Management With Predictive Duplicate Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face inefficiencies in database management due to duplicate datasets being stored across multiple systems, leading to increased infrastructure costs and skewed data analysis.
Innovation Solution
A computing system that utilizes a predictive model to compare datasets and identify similarities, transmitting notifications when a threshold similarity is exceeded, prompting users to avoid duplicative data storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplicate datasets are stored to various databases and servers, then data availability for multiple teams is improved, but infrastructure costs increase
Solution Approach 1:
The patent implements a centralized data registry that consolidates dataset metadata and location information into a single system. Multiple teams can query this registry to find existing datasets, eliminating the need for separate storage infrastructures and reducing overall infrastructure costs while maintaining data availability.
Solution Approach 2:
The system introduces a backend comparison system as an intermediary between data storage and retrieval operations. This intermediary uses machine learning models to detect duplicate datasets across different storage locations, preventing redundant storage and reducing infrastructure requirements while ensuring data accessibility.
2Ease of operation
If duplicate datasets are stored across multiple systems, then team-specific data access is improved, but data analysis accuracy deteriorates
Solution Approach 1:
The system implements feedback mechanisms where the backend comparison system continuously monitors new dataset uploads, compares them against existing datasets using machine learning models, and provides feedback to users about potential duplicates. This feedback loop prevents duplicate data from entering the system, maintaining analysis accuracy while preserving easy data access.
Solution Approach 2:
The patent applies preliminary comparison actions before datasets are fully stored or made accessible. The machine learning model performs similarity checks on incoming datasets against the existing registry, identifying duplicates in advance and preventing them from being added to the system, thus ensuring data integrity before access operations occur.
3Device complexity
If traditional data storage methods are used, then implementation simplicity is maintained, but resource efficiency deteriorates
Solution Approach 1:
The system segments the data management functionality into distinct modular components: a data registry for metadata storage, a comparison system for duplicate detection, and machine learning models for similarity analysis. This segmentation allows each component to be independently optimized and implemented, maintaining simplicity while improving resource efficiency through targeted functionality.
Data Source
AI summary
Systems and methods receive, by a backend system and from a user device, input(s) indicating dataset(s) should be stored to storage location(s) of the backend system, and dynamically compare data of the dataset(s) to stored data of stored dataset(s), the comparing including applying the dataset(s) to a deployed predictive model that is trained to quantify a percentage of similarity between the dataset(s) and the stored dataset(s). Based on the applying, it is dynamically determined that the percentage of similarity of one dataset of the dataset(s) surpasses a predefined threshold percentage, and electronic notification(s) indicating the percentage of similarity between the one dataset and a stored dataset of the dataset(s) is transmitted to the user device based on the percentage of similarity surpassing the predefined threshold percentage.


