Primary Storage Array Data Placement Using Backup Block Statistics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data storage systems, the primary storage array lacks the information to efficiently utilize the storage capabilities of storage virtualization arrays, leading to inefficiencies in data placement due to mismatched capabilities, resulting in suboptimal selection of storage resources for host application data.
Innovation Solution
A data analysis program running on a data backup storage array quantifies the suitability of data extents for processing by various storage capabilities, generating block backup statistics that guide the primary storage array to select the most suitable storage resources, such as virtualized managed drives with deduplication, compression, power conservation, or performance tiering, for optimal data placement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the primary storage array uses conventional data placement methods, then the storage system can operate with existing infrastructure, but the storage capabilities of virtualization arrays are not efficiently utilized due to mismatched capabilities
Solution Approach 1:
The system implements feedback by having the backup storage array analyze backup copies and generate block backup statistics that are fed back to the primary storage array. This feedback loop enables the primary array to make informed data placement decisions based on actual data characteristics and the capabilities of available storage resources, resolving the mismatch between storage capabilities and data placement efficiency
Solution Approach 2:
The system performs preliminary analysis of data characteristics by analyzing backup copies before final data placement decisions are made. The block backup statistics are generated in advance from backup data, allowing the primary storage array to pre-determine optimal placement locations that match data characteristics with appropriate storage capabilities, thereby improving both adaptability and productivity
2Productivity
If the primary storage array lacks information about data suitability, then the system architecture remains simple, but data placement becomes suboptimal due to inability to match data with appropriate storage capabilities
Solution Approach 1:
The block backup statistics serve as an intermediary that bridges the information gap between the primary storage array and the characteristics of stored data. Rather than requiring the primary array to directly analyze all data characteristics, the statistics act as a mediator that summarizes data suitability information, enabling optimized data placement without excessive complexity in the primary array's architecture
Solution Approach 2:
The system uses backup copies of data as proxies for analyzing data characteristics. Instead of analyzing the actual production data directly, the system analyzes copies from the backup storage array, which provides the necessary information about data suitability for different storage capabilities without requiring direct access to or complex processing of the primary data
3Adaptability or versatility
If data is placed without analyzing suitability for different storage capabilities, then the placement process is fast and simple, but storage resources are not optimally utilized leading to increased costs
Solution Approach 1:
The system performs data suitability analysis in advance by processing backup copies and generating block backup statistics before actual data placement occurs. This preliminary action separates the analysis phase from the execution phase, allowing optimal data placement decisions to be made quickly during actual data operations without sacrificing storage resource utilization
Solution Approach 2:
The system analyzes backup copies rather than actual data to determine data characteristics and suitability for different storage capabilities. This copying approach enables comprehensive analysis to be performed without impacting the speed of actual data placement operations, as the analysis is performed on duplicate data in parallel or beforehand
Data Source
AI summary
A backup copy of a production device is used to quantify suitability of host application data for placement on individual managed drives and virtualized managed drives based on storage capabilities associated with those drives. A data analysis program on a data backup storage array may generate block backup statistics to indicate that a production device or certain chunks, blocks or volumes of host application data are highly compressible or reducible via deduplication. The block backup statistics are sent from the data backup storage array to the primary storage array. The primary storage array uses the block backup statistics to select a particular storage resource with suitable storage capabilities for the data. Highly compressible data may be stored on a storage virtualization storage array with data compression capability, and data that is neither highly compressible nor reducible with deduplication may be stored on local resources.


