Primary Storage Array Data Placement Using Backup Block Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In data storage systems, the primary storage array lacks the information to efficiently utilize the storage capabilities of storage virtualization arrays, leading to inefficiencies in data placement due to mismatched capabilities, resulting in suboptimal selection of storage resources for host application data.

Innovation Solution

A data analysis program running on a data backup storage array quantifies the suitability of data extents for processing by various storage capabilities, generating block backup statistics that guide the primary storage array to select the most suitable storage resources, such as virtualized managed drives with deduplication, compression, power conservation, or performance tiering, for optimal data placement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the primary storage array uses conventional data placement methods, then the storage system can operate with existing infrastructure, but the storage capabilities of virtualization arrays are not efficiently utilized due to mismatched capabilities

Engineering Contradiction:
Improveutilization of storage capabilitiesVSAvoiddata placement efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system implements feedback by having the backup storage array analyze backup copies and generate block backup statistics that are fed back to the primary storage array. This feedback loop enables the primary array to make informed data placement decisions based on actual data characteristics and the capabilities of available storage resources, resolving the mismatch between storage capabilities and data placement efficiency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary analysis of data characteristics by analyzing backup copies before final data placement decisions are made. The block backup statistics are generated in advance from backup data, allowing the primary storage array to pre-determine optimal placement locations that match data characteristics with appropriate storage capabilities, thereby improving both adaptability and productivity

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the primary storage array lacks information about data suitability, then the system architecture remains simple, but data placement becomes suboptimal due to inability to match data with appropriate storage capabilities

Engineering Contradiction:
Improvedata placement optimizationVSAvoidinformation processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The block backup statistics serve as an intermediary that bridges the information gap between the primary storage array and the characteristics of stored data. Rather than requiring the primary array to directly analyze all data characteristics, the statistics act as a mediator that summarizes data suitability information, enabling optimized data placement without excessive complexity in the primary array's architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses backup copies of data as proxies for analyzing data characteristics. Instead of analyzing the actual production data directly, the system analyzes copies from the backup storage array, which provides the necessary information about data suitability for different storage capabilities without requiring direct access to or complex processing of the primary data

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If data is placed without analyzing suitability for different storage capabilities, then the placement process is fast and simple, but storage resources are not optimally utilized leading to increased costs

Engineering Contradiction:
Improvestorage resource utilizationVSAvoiddata placement time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs data suitability analysis in advance by processing backup copies and generating block backup statistics before actual data placement occurs. This preliminary action separates the analysis phase from the execution phase, allowing optimal data placement decisions to be made quickly during actual data operations without sacrificing storage resource utilization

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system analyzes backup copies rather than actual data to determine data characteristics and suitability for different storage capabilities. This copying approach enables comprehensive analysis to be performed without impacting the speed of actual data placement operations, as the analysis is performed on duplicate data in parallel or beforehand

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10565068B1Primary array data dedup/compression using block backup statistics
Publication Date: 2020.02.18 EMC IP HLDG CO LLC
  • US10565068B1 patent drawing
  • US10565068B1 patent drawing
  • US10565068B1 patent drawing

AI summary

A backup copy of a production device is used to quantify suitability of host application data for placement on individual managed drives and virtualized managed drives based on storage capabilities associated with those drives. A data analysis program on a data backup storage array may generate block backup statistics to indicate that a production device or certain chunks, blocks or volumes of host application data are highly compressible or reducible via deduplication. The block backup statistics are sent from the data backup storage array to the primary storage array. The primary storage array uses the block backup statistics to select a particular storage resource with suitable storage capabilities for the data. Highly compressible data may be stored on a storage virtualization storage array with data compression capability, and data that is neither highly compressible nor reducible with deduplication may be stored on local resources.