Hash-Based Duplicate Message Detection in Database Arrays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Enterprises face inefficiencies in storing and processing large numbers of messages in databases, including duplicate messages, which consume significant memory and computing resources, and slow down data processing due to the use of slower memory devices like HDDs.

Innovation Solution

A system that converts data to a hash using a hash function, determines selected character positions of the hash, and identifies array elements based on these positions to efficiently store and identify messages, reducing the need for extensive memory and computing resources by using faster semiconductor memory for in-memory databases and slower HDDs for long-term storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If duplicate messages are stored in the database, then the database contains complete message history, but memory consumption and computing resources increase significantly

Engineering Contradiction:
Improvemessage history completenessVSAvoidmemory consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates a hash copy of each message and stores it in an array instead of storing the full message content. This hash copy serves as a representative identifier that consumes minimal memory while still enabling duplicate detection. When a new message arrives, its hash is computed and checked against existing hashes in the array, allowing duplicate identification without storing redundant full message copies.

Inventive Principle:
Principle #26Copying

2Reliability

If traditional sorting methods are used to locate messages in the database, then all messages can be searched, but the time and computing resources required increase

Engineering Contradiction:
Improvemessage location accuracyVSAvoidmessage location time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent pre-computes and stores hash values for all messages in the database array before actual message retrieval operations. This preliminary hashing action transforms the database into a hash-indexed structure, enabling O(1) average time complexity for message location operations. When a message needs to be located, only its hash needs to be computed and compared against the pre-stored hashes, eliminating the need for time-consuming full message comparisons or sorting operations.

Inventive Principle:
Principle #10Preliminary action

3Speed

If full message data is stored in fast semiconductor memory, then access speed is high, but the cost and resource consumption increase

Engineering Contradiction:
Improvedata access speedVSAvoidsemiconductor memory usage
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent segments the message data storage into two parts: (1) full message content stored in slower, cheaper storage media, and (2) hash values of messages stored in fast semiconductor memory array. This segmentation allows the system to maintain high access speeds for duplicate detection and message retrieval operations by keeping only the essential hash identifiers in fast memory, while the bulk message data resides in cost-effective storage. The hash array acts as a high-speed index that guides access to the actual message data.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8805795B2Identifying duplicate messages in a database
Publication Date: 2014.08.12 BANK OF AMERICA CORP
  • US8805795B2 patent drawing
  • US8805795B2 patent drawing
  • US8805795B2 patent drawing

AI summary

A system for storing data in a memory comprises a memory operable to store a database, wherein the database comprises an array, and the array comprises a number of elements uniquely identifiable by their location in relation to an origin point of the array, an interface operable to receive first data to be stored in the array; and a processor communicatively coupled to the memory and the interface, the processor operable to convert the first data to a hash using a hash function, determine a selected number of character positions of the hash, and identify an array element according to the character values of the selected character positions of the hash.