Data Storage System Relevance Optimization via Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High volumes of data from interconnected devices and instrumentation sources lead to inefficient storage and processing costs due to the inclusion of irrelevant data, which traditional storage systems cannot effectively distinguish from relevant data, resulting in unnecessary data transfers and processing costs that affect the timeliness of actionable insights.
Innovation Solution
A data storage system comprising a data summarization module, a clustering module, and a representative content selection module that associates each data object with a derived data object, clusters similar data objects, and selects representative data objects to optimize relevance for analytics workloads, providing only relevant data in response to access requests, thereby reducing redundant information and enhancing insight discovery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If traditional storage systems store all data objects, then data availability is maintained, but storage costs and processing overhead increase due to irrelevant data
Solution Approach 1:
The patent segments data objects into clusters based on similarity metrics, creating distinct groups of related data. This segmentation allows the system to identify and retain only representative data objects from each cluster, eliminating redundant information while preserving data availability for analytics workloads.
Solution Approach 2:
The patent extracts representative data objects from clusters of similar data objects based on fidelity metrics. By taking out only the most representative objects that capture the essential characteristics of each cluster, the system reduces storage requirements and processing overhead while maintaining data availability for analytics queries.
2Productivity
If all data objects are transferred to analytics applications, then complete data access is achieved, but data transfer costs and time increase due to redundant information
Solution Approach 1:
The patent performs preliminary clustering and representative object selection before analytics queries are executed. By pre-organizing data into clusters and identifying representative objects in advance, the system prepares optimized data subsets that can be quickly transferred to analytics applications, reducing data transfer time while maintaining productivity in insight discovery.
Solution Approach 2:
The patent extracts only the representative data objects from each cluster that are most relevant to analytics workloads. This extraction eliminates redundant data transfer of similar or duplicate information while ensuring that the transferred data maintains high fidelity for insight discovery, thereby improving productivity without increasing data transfer time.
3Measurement precision
If high fidelity data is provided to analytics applications, then analysis accuracy is improved, but storage and transfer costs increase
Solution Approach 1:
The patent creates representative copies of data objects that capture the essential characteristics and fidelity needed for accurate analytics. Instead of storing and transferring all original high-fidelity data, the system generates representative copies from clustered data objects, maintaining analysis accuracy while significantly reducing storage and transfer costs through these optimized representations.
Data Source
AI summary
Relevance optimized representative content associated with a data storage system is disclosed. One example is a system including a data summarization module, a clustering module, and a representative content selection module. The data summarization module associates, via a processor, each data object in a storage system with a derived data object. The clustering module determines clusters of similar data objects based on a similarity between associated derived data objects, and selects a representative data object for each determined cluster. The representative content selection module selects representative content associated with the storage system, where the representative content is based on the data objects, the derived data objects, and the representative data objects, and relevance optimizes of the selected representative content to an analytics application.


