Hash-Based Duplicate Message Detection in Database Arrays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprises face inefficiencies in storing and processing large numbers of messages in databases, including duplicate messages, which consume significant memory and computing resources, and slow down data processing due to the use of slower memory devices like HDDs.
Innovation Solution
A system that converts data to a hash using a hash function, determines selected character positions of the hash, and identifies array elements based on these positions to efficiently store and identify messages, reducing the need for extensive memory and computing resources by using faster semiconductor memory for in-memory databases and slower HDDs for long-term storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If duplicate messages are stored in the database, then the database contains complete message history, but memory consumption and computing resources increase significantly
Solution Approach 1:
The patent creates a hash copy of each message and stores it in an array instead of storing the full message content. This hash copy serves as a representative identifier that consumes minimal memory while still enabling duplicate detection. When a new message arrives, its hash is computed and checked against existing hashes in the array, allowing duplicate identification without storing redundant full message copies.
2Reliability
If traditional sorting methods are used to locate messages in the database, then all messages can be searched, but the time and computing resources required increase
Solution Approach 1:
The patent pre-computes and stores hash values for all messages in the database array before actual message retrieval operations. This preliminary hashing action transforms the database into a hash-indexed structure, enabling O(1) average time complexity for message location operations. When a message needs to be located, only its hash needs to be computed and compared against the pre-stored hashes, eliminating the need for time-consuming full message comparisons or sorting operations.
3Speed
If full message data is stored in fast semiconductor memory, then access speed is high, but the cost and resource consumption increase
Solution Approach 1:
The patent segments the message data storage into two parts: (1) full message content stored in slower, cheaper storage media, and (2) hash values of messages stored in fast semiconductor memory array. This segmentation allows the system to maintain high access speeds for duplicate detection and message retrieval operations by keeping only the essential hash identifiers in fast memory, while the bulk message data resides in cost-effective storage. The hash array acts as a high-speed index that guides access to the actual message data.
Data Source
AI summary
A system for storing data in a memory comprises a memory operable to store a database, wherein the database comprises an array, and the array comprises a number of elements uniquely identifiable by their location in relation to an origin point of the array, an interface operable to receive first data to be stored in the array; and a processor communicatively coupled to the memory and the interface, the processor operable to convert the first data to a hash using a hash function, determine a selected number of character positions of the hash, and identify an array element according to the character values of the selected character positions of the hash.


