Augmented Metadata Collection System for Storage Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for gathering metadata in data storage platforms using network protocols like NFS and CIFS are limited, as they restrict the types of metadata that can be collected and can negatively impact system performance and increase network traffic.
Innovation Solution
Implementing an augmented metadata collection system within the storage platform to identify and schedule metadata collection efficiently, allowing for the gathering of additional metadata not limited by standard protocols, by distinguishing between scheduled and unknown metadata and optimizing scanning and filtering processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If network protocols like NFS and CIFS are used to gather metadata from storage platforms, then metadata can be collected using standard client access protocols, but the types of metadata that can be gathered are limited to what can be expressed in the standard protocols
Solution Approach 1:
The patent introduces an intermediary component within the storage platform that translates between standard network protocols and internal metadata structures. This intermediary enables the system to respond to standard protocol requests while providing access to extended metadata types that go beyond what the standard protocols natively support, thus resolving the contradiction between protocol standardization and metadata versatility.
Solution Approach 2:
The storage platform is designed with multi-functional metadata collection capabilities that can handle both standard protocol requests and extended metadata queries. The system universally supports multiple metadata types through a unified architecture that can adapt to different protocol requirements while providing access to additional metadata fields, thereby achieving both protocol compatibility and enhanced adaptability.
2Loss of information
If network protocols are used to request metadata from storage platforms, then metadata can be gathered externally, but system performance is negatively impacted and network traffic increases
Solution Approach 1:
The system performs preliminary metadata collection and caching within the storage platform before external requests arrive. By pre-gathering and storing metadata information in an optimized format, the system can quickly respond to external requests without performing expensive scanning operations at the time of request, thus maintaining both complete metadata collection and high system performance.
Solution Approach 2:
The patent creates internal copies of metadata from the actual storage data. Instead of repeatedly scanning the underlying storage devices to answer external metadata requests, the system maintains copied metadata representations that can be quickly queried and returned, significantly reducing the performance impact on the primary storage system while maintaining complete metadata availability.
3Ease of operation
If external clients use network protocols to gather metadata, then metadata can be accessed remotely, but network traffic increases due to tree walking all network attached storage devices
Solution Approach 1:
The patent extracts metadata collection operations from the external network request path and performs them internally within the storage platform. By taking out the metadata gathering function from the external client-initiated process and implementing it as an internal optimization, the system can serve remote clients with metadata information without requiring them to perform expensive tree-walking operations across the network, thus maintaining ease of remote access while dramatically reducing network traffic.
Data Source
AI summary
Implementations are provided herein relating to augmenting metadata collection within a storage platform. The storage platform can be audited to determine the types of metadata currently being gathered within the storage platform, and the schedule for when that information is gathered. The storage platform can receive a request to generate metadata, compare the requested information with the previously generated and/or scheduled generation of metadata. Rather than redundantly gathering the same metadata via multiple requests, known metadata or scheduled retrieval of known metadata can be used to process portions of the metadata request, and any metadata that was not previously generated can then be separately generated. In this sense, the metadata collection within a storage platform can be augmented to gather additional metadata requested outside the storage platform in an efficient matter that does not unnecessarily increase scanning activity within the storage platform.


