Edge Data Pattern Clustering for Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data management systems face inefficiencies in handling duplicate data from collection devices, leading to wastage of transmission bandwidth and storage space, particularly due to the limitations of edge computing devices in processing and storing large numbers of data patterns.
Innovation Solution
The method involves acquiring data patterns from multiple collection devices, clustering them to form groups, and determining shared data patterns within each group, which are then used to reduce duplicate data transmission and storage by replacing duplicate data segments with identifiers at the edge computing device before transmission to the core network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data patterns from multiple collection devices are stored and processed at edge computing devices, then data transmission efficiency is improved, but the burden on edge computing devices increases
Solution Approach 1:
The patent segments collection devices into multiple groups based on clustered data patterns, and processes data patterns at the group level rather than individually at each device. This segmentation allows the system to manage data patterns more efficiently by grouping similar data together, reducing the computational burden on individual edge devices while maintaining transmission efficiency.
Solution Approach 2:
The patent merges data patterns from multiple collection devices into shared data pattern groups. By identifying and merging common data patterns across devices, the system reduces redundancy and the overall burden on edge computing devices, as patterns are processed and managed collectively rather than separately at each device.
2Reliability
If duplicate data is transmitted to core network, then data completeness is maintained, but transmission bandwidth is wasted
Solution Approach 1:
The patent performs preliminary data pattern clustering and deduplication at the edge computing device before data transmission to the core network. By identifying and removing duplicate data patterns in advance at the edge, the system maintains data completeness through proper pattern representation while significantly reducing transmission bandwidth requirements.
Solution Approach 2:
The patent uses data pattern identifiers as copies or representations of actual data patterns. Instead of transmitting redundant duplicate data, the system transmits references to shared data patterns, maintaining data completeness through these pattern identifiers while reducing transmission bandwidth.
3Loss of information
If all data patterns are stored for later analysis, then data availability is improved, but storage space is wasted
Solution Approach 1:
The patent extracts and identifies duplicate data patterns from the total data set, separating them from unique patterns. By taking out duplicates and representing them through shared pattern identifiers, the system maintains data availability for analysis while significantly reducing the storage space required to store all data patterns.
Solution Approach 2:
The patent changes the storage parameter from storing complete data patterns to storing data pattern identifiers and their occurrence information. This parameter change allows the system to maintain data availability for analysis while reducing storage space requirements by storing only essential pattern information rather than all duplicate data.
Data Source
AI summary
Techniques for managing data patterns involve: acquiring multiple sets of data patterns respectively associated with multiple collection devices, wherein a set of data patterns in the multiple sets of data patterns represent patterns of duplicate data in data from one of the multiple collection devices; dividing the multiple collection devices into multiple groups based on clusters of the multiple sets of data patterns; and determining, based on sets of data patterns associated with collection devices in a group in the multiple groups, a set of shared data patterns for sharing among the collection devices in the group. Accordingly, data patterns that can be shared among multiple collection devices can be determined in a more accurate and effective manner, thereby facilitating the removal of duplicate data from the multiple collection devices.


