Edge Data Deduplication via Semantic Pattern Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In network environments, especially in IoT settings, similar data samples are redundantly stored and transmitted, consuming additional storage space and bandwidth, and requiring unnecessary processing, as existing deduplication techniques fail to detect semantically similar data streams effectively due to dynamic fields like timestamps.
Innovation Solution
The implementation of semantic pattern detection at the edge device to categorize data streams as semantically duplicate or unique, allowing for efficient storage and transmission, where semantically duplicate data is either discarded or transmitted in a compressed form, and unique data is selectively stored and transmitted, optimizing storage and bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing deduplication techniques are used to store and transmit data, then data storage and transmission can be performed, but semantically similar data streams are not detected, causing redundant storage and transmission of similar data
Solution Approach 1:
The patent transforms data by removing dynamic fields (timestamps, sequence numbers) and retaining only static semantic fields, thereby changing the data parameters to enable effective deduplication detection while reducing storage of redundant semantic information
Solution Approach 2:
The patent extracts and removes dynamic fields that cause false uniqueness from data streams, separating the semantic content from temporal metadata, thereby enabling accurate detection of semantically similar data without being misled by changing timestamps
2Loss of information
If all data streams are transmitted over the network, then complete data availability is maintained, but bandwidth consumption increases due to redundant transmission of semantically duplicate data
Solution Approach 1:
The patent changes the transmission parameter from raw data streams to processed data with removed dynamic fields, enabling identification and elimination of semantically duplicate transmissions while preserving unique semantic information
Solution Approach 2:
The patent transmits only the necessary semantic content without redundant dynamic fields, performing partial transmission that suffices for analytical purposes while reducing overall bandwidth consumption
3Reliability
If data with dynamic fields is processed for analysis, then temporal information is preserved, but processing time increases due to unnecessary handling of redundant data
Solution Approach 1:
The patent performs preliminary processing at the edge device to remove dynamic fields and identify semantically similar data before transmission, thereby reducing the processing burden on remote devices and overall system processing time while maintaining analysis accuracy
Data Source
AI summary
Example techniques of data management in a network environment are described. In an example, a semantic pattern in a data stream transmitted from a source device to an edge device in the network environment is determined. The semantic pattern indicates relevance of data samples in the data stream for analysis of the data stream. The data stream is processed based on the semantic pattern, for storage and transmission in the network environment.


