Proactive Database Management With Predictive Duplicate Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large organizations face inefficiencies in database management due to duplicate datasets being stored across multiple systems, leading to increased infrastructure costs and skewed data analysis.

Innovation Solution

A computing system that utilizes a predictive model to compare datasets and identify similarities, transmitting notifications when a threshold similarity is exceeded, prompting users to avoid duplicative data storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If duplicate datasets are stored to various databases and servers, then data availability for multiple teams is improved, but infrastructure costs increase

Engineering Contradiction:
Improvedata availabilityVSAvoidinfrastructure costs
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent implements a centralized data registry that consolidates dataset metadata and location information into a single system. Multiple teams can query this registry to find existing datasets, eliminating the need for separate storage infrastructures and reducing overall infrastructure costs while maintaining data availability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces a backend comparison system as an intermediary between data storage and retrieval operations. This intermediary uses machine learning models to detect duplicate datasets across different storage locations, preventing redundant storage and reducing infrastructure requirements while ensuring data accessibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If duplicate datasets are stored across multiple systems, then team-specific data access is improved, but data analysis accuracy deteriorates

Engineering Contradiction:
Improveteam data accessVSAvoiddata analysis accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where the backend comparison system continuously monitors new dataset uploads, compares them against existing datasets using machine learning models, and provides feedback to users about potential duplicates. This feedback loop prevents duplicate data from entering the system, maintaining analysis accuracy while preserving easy data access.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary comparison actions before datasets are fully stored or made accessible. The machine learning model performs similarity checks on incoming datasets against the existing registry, identifying duplicates in advance and preventing them from being added to the system, thus ensuring data integrity before access operations occur.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If traditional data storage methods are used, then implementation simplicity is maintained, but resource efficiency deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidresource efficiency
Core Design Contradiction:
Device complexityVSLoss of energy

Solution Approach 1:

The system segments the data management functionality into distinct modular components: a data registry for metadata storage, a comparison system for duplicate detection, and machine learning models for similarity analysis. This segmentation allows each component to be independently optimized and implemented, maintaining simplicity while improving resource efficiency through targeted functionality.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12430319B2Proactive database management systems
Publication Date: 2025.09.30 TRUIST BANK
  • US12430319B2 patent drawing
  • US12430319B2 patent drawing
  • US12430319B2 patent drawing

AI summary

Systems and methods receive, by a backend system and from a user device, input(s) indicating dataset(s) should be stored to storage location(s) of the backend system, and dynamically compare data of the dataset(s) to stored data of stored dataset(s), the comparing including applying the dataset(s) to a deployed predictive model that is trained to quantify a percentage of similarity between the dataset(s) and the stored dataset(s). Based on the applying, it is dynamically determined that the percentage of similarity of one dataset of the dataset(s) surpasses a predefined threshold percentage, and electronic notification(s) indicating the percentage of similarity between the one dataset and a stored dataset of the dataset(s) is transmitted to the user device based on the percentage of similarity surpassing the predefined threshold percentage.