Database Replica Generation for Data Obfuscation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Confidential information is often compromised due to easy identification, as it may be labeled with user IDs or incorporated into databases with minimal data, making it easily identifiable by nefarious actors.

Innovation Solution

A method that generates and stores a large number of replica data entries, substantially similar to a new data entry, to overwhelm and obscure the valid data, making it difficult for actors to determine which records are valid or invalid, by randomly selecting and inserting these entries into a database with a unique identifier.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If confidential information is stored with user IDs or minimal data, then ease of identification and access is improved, but vulnerability to compromise and easy identification by nefarious actors worsens

Engineering Contradiction:
Improveease of identificationVSAvoidvulnerability to compromise
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent creates multiple replica data entries that are substantially similar to the original confidential information. These replicas are stored alongside the original data, making it difficult for attackers to identify which entry is the genuine confidential information. The system generates replica entries with similar structure and content characteristics, thereby diluting the value of compromised data while maintaining operational ease.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If a single data entry is stored, then storage efficiency and simplicity are improved, but the ability to obscure valid data from attackers worsens

Engineering Contradiction:
Improvedata storage quantityVSAvoiddifficulty of identifying valid records
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent changes the parameter of data quantity by generating multiple replica entries (e.g., 100 replicas) for each original confidential entry. This parameter change transforms the storage structure from single-entry to multi-entry, thereby increasing the difficulty for attackers to identify valid records while the system automatically manages the proliferation through defined thresholds and replication rules.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If replica data entries are generated to obscure valid data, then security against compromise is improved, but system complexity and processing overhead worsen

Engineering Contradiction:
Improvesecurity against compromiseVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by establishing replica generation thresholds and replication rules before compromise occurs. The system proactively creates replica entries when certain conditions are met (e.g., when an entry meets replica threshold criteria), rather than reacting after compromise. This preliminary structuring automates the complexity management and reduces real-time processing burden.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically managing the replication process according to predefined thresholds and rules. The patent enables the system to autonomously determine when to create replicas, how many replicas to generate, and how to distribute them, thereby reducing manual intervention and simplifying operational complexity despite the increased security measures.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12189826B2Automatic devaluation of compromised data
Publication Date: 2025.01.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12189826B2 patent drawing
  • US12189826B2 patent drawing
  • US12189826B2 patent drawing

AI summary

A processor may identify that a new data entry is being generated. The processor may identify that the new data entry is associated with a replica data entry threshold. The replica data entry threshold may indicate a minimum amount of replica data entries to generate. The replica data entries may be substantially similar to the new data entry. The processor may generate an amount of replica data entries. The processor may store the new data entry and the amount of replica data entries in a repository.