Memory health tracking for differentiated data recovery configuration

By receiving storage health data from storage devices and dynamically adjusting the data recovery configuration of the distributed storage system, the problem of unreasonable resource allocation in existing technologies is solved, achieving more efficient storage resource utilization and data protection.

CN114746843BActive Publication Date: 2026-01-20WESTERN DIGITAL TECHNOLOGIES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080081578.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-14
Filing Date
2020-05-29
Publication Date
2026-01-20
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

Existing distributed storage systems cannot dynamically adjust resource allocation based on the actual health status of storage devices in their backup configurations for edge storage devices, leading to resource waste or data loss risks, and failing to effectively cope with different usage scenarios and error rates of storage devices.

Method used

By receiving memory health data from storage devices, the data recovery configuration in the distributed storage system is dynamically adjusted. Based on changes in the memory health status of storage devices, storage resources for redundant datasets are reallocated, and different parity levels are adopted to adapt to different error rates.

Benefits of technology

It enables flexible allocation of storage resources based on the actual health status of storage devices, improving the reliability and efficiency of the storage system and reducing the risk of resource waste and data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114746843B_ABST
    Figure CN114746843B_ABST
Patent Text Reader

Abstract

Example systems and methods provide differentiated data recovery configurations based on memory health data. A distributed storage system, such as a cloud-based storage system, stores backup data from a remote storage device using a first data recovery configuration. Based on memory health data collected from the remote storage device, a change in a memory health state of the remote storage device can be determined. In response to the change in the memory health state, a different data recovery configuration can be used to store incoming backup data in the distributed storage system and to reallocate previously stored backup data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to data storage, and in more particular examples, to backup data storage data recovery configuration. BACKGROUND

[0002] Edge storage devices, such as computer hard drives, external hard drives, removable solid state storage devices (e.g., flash drives), and the like, can benefit from remote backup solutions to prevent data loss in the event of degradation or loss of the edge storage device. For example, such storage devices can be configured for periodic backup to a cloud-based storage system. The cloud-based storage system can provide a storage application for interacting with the edge storage device, implementing a backup configuration, and receiving data from the storage device to be backed up. In some configurations, the storage application can provide interface functionality for a distributed storage system supporting backup or other data storage applications for a plurality of end users.

[0003] Generally, distributed storage systems are used to store large amounts (e.g., terabytes, petabytes, exabytes, etc.) of data, such as objects or files, in a distributed and fault-tolerant manner with a predetermined level of redundancy. Such distributed storage systems can be particularly advantageous as active backup destinations for edge storage devices.

[0004] These large-scale storage systems can support storage of erasure coded and distributed data across many storage devices. Data, such as files or objects, can be divided into messages or similar data units with an upper bound on size. These data units are then divided into a plurality of symbols. The symbols are then used as input for erasure coding. For example, when using a systematic erasure coding algorithm, the output of the erasure coding process produces the original symbols and a fixed number of additional parity symbols. The total of these symbols is distributed across a selection of storage devices.

[0005] While erasure coding can enable a greater level of redundancy and increased error rate tolerance for recovering data, the processing, storage space, and other storage resources required still represent a large cost for storage providers. More specifically, selection of the level of parity and other aspects of data recovery configuration can increase or decrease the processing, storage space, network bandwidth, and other resources required for storage and recovery. Storage providers can need to balance the storage resources allocated to backups for any given edge storage device relative to the desired level of data backup and recovery service.

[0006] A general recovery configuration for backups of edge storage devices can be determined. However, individual storage devices have different actual usage, error rates, and degradation. The general configuration can be based on average or worst-case performance of the storage devices, and thus, more or less storage resources are allocated to backups than are actually guaranteed by the current conditions of the storage devices.

[0007] When using a distributed storage system to provide backup of edge storage devices, allocating storage resources for a storage device population against a fixed standard based on worst-case scenarios can result in wasted storage resources and / or unnecessary risk of data loss. Adaptive differential backup is needed based on the actual risk to data at the current point in time for a particular storage device. For example, a storage system can need to receive memory health data from individual storage devices and use the memory health data to allocate storage resources through a differentiated data recovery configuration. SUMMARY

[0008] Aspects are described for storing redundant backup data from storage devices to a distributed storage system, particularly using memory health tracking to differentiate data recovery configurations for individual storage devices.

[0009] One general aspect includes a computer-implemented method. The computer- implemented method includes storing a redundant data set from a remote storage device in a distributed storage system using a first data recovery configuration; receiving memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory medium of the remote storage device; determining a change in the memory health state of the non-transitory medium of the remote storage device based on the memory health data; and reallocating the redundant data set in the distributed storage system using a second data recovery configuration in response to the change in the memory health state.

[0010] Implementations can include one or more of the following features. The remote storage device can be a non-volatile memory device located at a site remote from the distributed storage system, and reallocating the redundant data set in the distributed storage system can include periodically backing up differences between a current data set stored on the remote storage device and a full copy of data stored on the remote storage device at an earlier time. The computer-implemented method can further include determining a periodic backup configuration of the remote storage device, determining at least one initial memory health value of the remote storage device, and determining the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration. The computer-implemented method can further include determining a service level of at least one system resource of the distributed storage system, determining an allocation of the at least one system resource for storing the redundant data set in the distributed storage system based on the service level, wherein determining the first data recovery configuration is further based on the allocation of the at least one system resource, and determining the second data recovery configuration based on the allocation of the at least one system resource and the change in memory health. Storing the redundant data set in the distributed storage system using the first data recovery configuration can include encoding the redundant data set in a first plurality of encoded data symbols according to a first parity level. Reallocating the redundant data set in the distributed storage system using the second data recovery configuration can include encoding at least a portion of the redundant data set in a second plurality of encoded data symbols according to a second parity level. The second parity level can be adapted for a different error rate for recovering the portion of the redundant data set compared to the first parity level. The memory health data can include at least one memory health value selected from a bit error rate value, a write / erase cycle value, a program cycle counter value, an erase cycle counter value, a leak detection measurement value, an unstable program disturbance value, a bad block value, or a voltage margin value. The computer-implemented method can further include receiving the redundant data set from the remote storage device according to a periodic backup schedule, wherein receiving memory health data from the remote storage device is performed in conjunction with receiving the redundant data set according to the periodic backup schedule. The computer-implemented method can further include determining a plurality of physical storage units in the remote storage device, and storing a reference value associating the redundant data set stored in the distributed storage system with the plurality of physical storage units in the remote storage device that store corresponding data. Receiving memory health data from the remote storage device can include receiving at least one memory health value for each physical storage unit of the plurality of physical storage units. Determining the change in the memory health state can include determining that at least one memory health value of a first physical storage unit of the plurality of physical storage units satisfies a reduced reliability condition, and determining that at least one memory health value of a second physical storage unit of the plurality of physical storage units does not satisfy the reduced reliability condition.Reallocating the set of redundant data in the distributed storage system using the second data recovery configuration can include storing data associated with the first physical storage unit using the second data recovery configuration in response to determining the reduced reliability condition. Data associated with the second physical storage unit can remain stored using the first data recovery configuration. Determining the change in the memory health status can include determining at least one reduced reliability threshold and evaluating the memory health data against the at least one reduced reliability threshold. The computer-implemented method can further include collecting historical memory health data of a population of remote storage devices of a remote storage device type associated with the remote storage device, determining a data reliability model for the remote storage device type based on the collected historical memory health data, and determining at least one reduced reliability threshold based on the data reliability model, wherein determining the change in the memory health includes evaluating the memory health data against the at least one reduced reliability threshold.

[0011] Another general aspect includes a system. The system includes a storage system configured to store a set of redundant data from a remote storage device using a first data recovery configuration, a memory health monitor configured to receive memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health status of a non-transitory media of the remote storage device, a reliability manager configured to determine a change in the memory health status of the remote storage device based on the memory health data and initiate a second data recovery configuration in response to the change in the memory health status, wherein the storage system is further configured to store the set of redundant data from the remote storage device using the second data recovery configuration.

[0012] Implementations can include one or more of the following features. The remote storage device can be a non-volatile memory device located at a site remote from the storage system, and the storage system can be further configured to periodically store differences between a current data set stored on the remote storage device and a full copy of data stored on the remote storage device at an earlier time. The system can also include a backup interface configured to determine a periodic backup configuration of the remote storage device. The reliability manager can be further configured to: determine at least one initial memory health value of the remote storage device; and determine the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration. The reliability manager can be further configured to: determine a service level of at least one system resource of the storage system; determine an allocation of the at least one system resource for storing redundant data in the storage system based on the service level, wherein the first data recovery configuration is further based on the allocation of the at least one system resource; and determine the second data recovery configuration based on the allocation of the at least one system resource and a change in the memory health state. The storage system can be further configured to: encode the redundant data set in a first plurality of encoded data symbols according to a first parity level in response to the first data recovery configuration; and encode redundant data in a second plurality of encoded data symbols according to a second parity level in response to the second data recovery configuration. The second parity level can be adapted for a different error rate in recovering data compared to the first parity level. The system can also include a backup interface configured to receive backup data from the remote storage device according to a periodic backup schedule, wherein the memory health monitor can be further configured to receive memory health data from the remote storage device in conjunction with the backup interface receiving backup data according to the periodic backup schedule. The memory health monitor can be further configured to: determine a plurality of physical storage units in the remote storage device; store reference values associating data stored in the storage system with the plurality of physical storage units storing corresponding user data in the remote storage device; and receive at least one memory health value for each physical storage unit of the plurality of physical storage units. The reliability manager can be further configured to: determine that at least one memory health value of a first physical storage unit of the plurality of physical storage units satisfies a reduced reliability condition; and determine that at least one memory health value of a second physical storage unit of the plurality of physical storage units does not satisfy the reduced reliability condition. The storage system can be further configured to store redundant data associated with the first physical storage unit using the second data recovery configuration in response to determining the reduced reliability condition. The redundant data associated with the second physical storage unit can remain stored using the first data recovery configuration.The reliability manager can be further configured to determine at least one reduced reliability threshold and evaluate the memory health data against the at least one reduced reliability threshold. The reliability manager is further configured to access historical memory health data of a population of remote storage devices of a remote storage device type associated with the remote storage device, determine a data reliability model of the remote storage device type based on the historical memory health data, determine at least one reduced reliability threshold based on the data reliability model, and evaluate the memory health data against the at least one reduced reliability threshold.

[0013] Another general aspect includes a system. The system includes a storage system configured to store a redundant data set from a remote storage device using a first data recovery configuration, means for receiving memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device, means for determining a change in the memory health state of the remote storage device based on the memory health data, and means for initiating a second data recovery configuration in response to the change in the memory health state, wherein the storage system is further configured to store a redundant data from the remote storage device using the second data recovery configuration. Various embodiments advantageously apply the teachings of distributed storage networks and / or systems to improve the functionality of such computer systems. Various embodiments include operations that overcome or at least reduce the problems in the previously described storage networks and / or systems, and are therefore more reliable and / or effective than other computing networks. That is, various embodiments disclosed herein include hardware and / or software with functionality to improve the use of storage resources by differentiating data recovery configurations for backing up individual storage devices using memory health data. Accordingly, the embodiments disclosed herein provide various improvements to storage networks and / or storage systems, particularly for cloud-based storage.

[0014] It should be appreciated that the language used in the subject disclosure is chosen primarily for readability and instructional purposes and can not have a special or critical meaning. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 An example of a cloud-based system for backing up storage devices to a distributed storage system is schematically illustrated.

[0016] Figure 2 An example of an exemplary backup architecture operable in the system of Figure 1 is schematically illustrated.

[0017] Figure 3 Some exemplary elements of a storage system for the system of Figure 1 are schematically illustrated.

[0018] Figure 4 An example method of differentiating data recovery configurations based on memory health data is shown.

[0019] Figure 5 An example method of managing backup configurations based on memory health data is shown.

[0020] Figure 6 An example method of differentiating data recovery configurations within a storage unit of a storage device based on memory health data is shown.

[0021] Figure 7 An example method of changing data recovery configurations using reduced reliability conditions is shown. DETAILED DESCRIPTION

[0022] The following embodiments allow for flexible cloud redundancy allocation based on user locally stored memory health. For example, a purchaser of an edge storage device product, such as a flash drive or solid state drive (SSD), can have the option to allow the manufacturer to track their memory health in order to better manage data backups and / or proactively alert the purchaser of changes in memory health. The purchaser can receive incentives for selecting to collect memory health data from their storage device, such as discounted cloud services. The memory health information can be used by a cloud storage provider, providing the purchaser with a cloud backup application to optimize backup storage allocation based on the memory health data.

[0023] A cloud storage provider can provide a distributed storage system that provides two levels of protection for data stored on a particular storage device. A backup service can include full backups with super-slow recovery latency as a first level of backup - slow and cheap protection to store a snapshot of the entire contents of the storage device at a particular point in time. These full backups can be performed periodically according to a periodic backup schedule or other backup conditions. This first level of protection can be most useful for handling sudden catastrophic loss of the storage device (e.g., due to theft, loss, or mechanical failure). An example slow storage system can include a redundant array of independent disks (RAID) system and / or low-cost storage backups, such as magnetic tape, high-capacity hard disk drives (HDDs), X4 NAND flash memory, etc.

[0024] The second backup protection level can include flexible protection with fast data recovery response times, such as frequent fast backups to store the delta between the data currently stored on the storage device and the most recent snapshot stored in the first level of protection. This second protection level can be most useful for handling gradual degradation of reliability of the storage device (e.g., high bit error rate (BER) due to high write / erase cycle values). An example fast response storage system can include a server and a storage array configured with fast response storage, such as an all-flash storage array.

[0025] The edge storage devices can be of any type, such as consumer or enterprise, SSD, HDD, hybrid drive, secure digital (SD) card, universal serial bus (USB) stick, or other form factor flash drive. In the illustrated embodiment, the storage devices are backed up to a distributed storage system, which typically operates as a component of a cloud storage system. However, other storage system configurations are possible, and the described systems and methods can be implemented for backup to any type of local storage system, even possibly to another edge storage device of the same type as the storage device being backed up.

[0026] In some embodiments, the cloud backup service can differentiate between two types of reliability threats to the physical storage media of the storage device: gradual reliability degradation (e.g., high bit error rate due to high W / E cycle values) and “catastrophic” sudden loss (the following events: theft, loss, failure, etc.). Thus, the cloud backup service can implement two levels of cloud protection of the data.

[0027] The first protection level can include full backups with super-slow recovery latency. To protect against “catastrophic” sudden loss events, the cloud provider can maintain a full updated backup of the data, although it can be stored in very high recovery latency duration. For example, for such rare events of storage device failure, the response time until the entire storage dataset is provided can be in hours. A slow and cheap backup method for the entire device can be used for the proposed first protection level, such as RAID and / or low-cost storage backup (tape, low-cost HDD, X4 NAND, etc.) on multiple clients.

[0028] The second protection level can include flexible protection for gradual reliability degradation (reduced reliability condition based on memory health) with fast data recovery response times. The second protection level can allocate redundancy in a flexible manner based on the physical health of the storage device. Memory health parameters of the physical storage media in each storage device are tracked by the storage device, and the memory health data can be sent to the storage system. The storage system can then use this information to dynamically allocate redundancy in the data recovery configuration of a particular storage device or even those settings within the physical units (device, die, page / block, etc.).

[0029] To protect against gradual reliability degradation (i.e., high BER), cloud service providers can be enabled to access memory health parameters collected by the storage device. Memory manufacturers can include remote access to physical health parameters of the non-volatile memory device in the storage device (adjusted according to user approval). Collection of memory health data can allow memory manufacturers to provide memory health warnings to users, as well as provide additional capabilities related to the user's cloud backup service. Physical memory health parameters can include:

[0030] • BER (bit error rate) values

[0031] • W / E (write / erase) cycle values

[0032] • PLC (program cycle counter) values

[0033] • ELC (erase cycle counter) values

[0034] • leakage detection measurement values

[0035] • EPD (erroneous program disturb) error values

[0036] • bad block statistics

[0037] • voltage margin values.

[0038] This second level of protection can be stored in a fast response memory backup within the distributed storage system, which is configured to allow fast recovery of data. The storage system can utilize knowledge of the memory health data to adjust optimal redundancy allocation. In some embodiments, the second level of protection can be used to store incremental backups from the last time a full copy of the entire memory backup (first level of protection) was completed. An incremental backup can be changes in data between the full copy taken at an earlier time and the current data set in the storage device.

[0039] Figure 1A block diagram of an example cloud-based system 100 is shown in which tiered storage with memory health tracking of a differentiated recovery configuration can be implemented for a fast storage system (second level of protection). As shown, the system 100 includes client systems 102 (e.g., client systems 102.1 and 102.n), distributed storage systems 120.1, 102.2...120.n, object stores 140.1, 140.2...140.n associated with the distributed storage systems, and server systems 150 (e.g., server systems 150.1 and 150.n). The components 102, 120, 140, and / or 150 and / or subcomponents thereof can be interconnected directly or via a communications network 110. In some cases, for simplicity, depending on the context, the client systems 102.1...102.n can also be referred to herein individually or collectively as client systems 102 or clients 102, the distributed storage systems 120.1, 120.2...120.n can be referred to herein individually or collectively as distributed storage systems 120 or DSSs 120, the storage applications 124.1, 124.2...124.n can be referred to herein individually or collectively as storage applications 124, the metadata stores 130.1, 130.2...130.n can be referred to herein individually or collectively as metadata stores 130, the object stores 140.1, 140.2...104.n can be referred to herein individually or collectively as object stores 140, and the server systems 150.1 and 150.n can be referred to herein individually or collectively as server systems 150.

[0040] The communications network 110 can include any number of private and public computer networks. The communications network 110 can include networks having any of a variety of network types, including local area networks (LANs), wide area networks (WANs), wireless networks, virtual private networks, wired networks, the Internet, personal area networks (PANs), object buses, computer buses, and / or any suitable combination of communication media via which devices can communicate in a secure or unsecured manner.

[0041] Data can be transmitted via network 110 using any suitable protocol. Exemplary protocols include, without limitation, transmission control protocol / internet protocol (TCP / IP), user datagram protocol (UDP), transmission control protocol (TCP), hypertext transfer protocol (HTTP), secure hypertext transfer protocol (HTTPS), dynamic adaptive streaming over HTTP (DASH), real-time streaming protocol (RTSP), real-time transport protocol (RTP) and real-time transport control protocol (RTCP), voice over internet protocol (VOIP), file transfer protocol (FTP), WebSocket (WS), wireless access protocol (WAP), various messaging protocols (short message service (SMS), internet message access protocol (IMAP), etc.), or other suitable protocols.

[0042] Client system 102 can include an electronic computing device such as a personal computer (PC), a laptop computer, a smartphone, a tablet computer, a mobile phone, a wearable electronic device, a server, a server farm, or any other electronic device or computing system capable of communicating with communication network 110. Client system 102 can store one or more client applications in non-transitory memory, including internal memory (not shown) and / or storage devices 106.1-106.n, for use by users 104.1-104.n. Exemplary external or removable storage devices 106 can include SSDs, HDDs, hybrid drives, secure digital (SD) cards, universal serial bus (USB) sticks, or other forms of flash drives, including non-transitory storage media. The client applications can be executed by a computer processor of client system 102. In some exemplary embodiments, the client applications include one or more applications such as, but not limited to, a data storage application, a search application, a communication application, a productivity application, a gaming application, a word processing application, or any other application. In some cases, the client applications can include a web browser and / or code that can be executed thereby.

[0043] In some embodiments, client systems 102 can include applications for creating, modifying, and deleting objects that can be stored in object store 140. For example, an application can be specifically tailored for communicating with cloud application 152 and / or storage application 124, such as an application adapted to configure and / or utilize the programmatic interfaces of storage application 124. In some embodiments, cloud application 152 and / or storage application 124 can embody a backup application of client system 102 and / or storage device 106. For example, redundant copies of data units stored in non-transitory memory of client system 102 and / or storage device 106 can be stored as objects in object store 140 using storage application 124 and / or cloud application 152. In some embodiments, cloud application 152 hosted by server system 150.1 can embody a client of storage application 124, as it can use the various programmatic interfaces it faces to access the functionality of storage application 124 (e.g., to make creations, stores, retrievals, deletions, etc. of objects stored in object storage). Client systems 102 can be remote from distributed storage system 120 and server system 150, and connected only via communication network 110. For example, distributed storage system 120 and / or server system 150 can be located in secure sites, such as commercial data centers, and client systems 102 can be personal and business computer systems and storage devices operating at remote home, business, and mobile sites.

[0044] Client systems 102, distributed storage system 120, and / or server system 150 can send / receive requests and / or send / receive responses to / from each other, such as but not limited to HTTP(S) requests / responses. Client systems 102 can present information to user 104 via output devices, such as displays, audio reproduction devices, vibrating mechanisms, etc., based on information generated by client systems 102 and / or received from server system 128 and / or distributed storage system 120, such as visual, audio, tactile, and / or other information.

[0045] User 104 can interact with various client systems 102 to provide input and receive information. For example, as shown, user 104.1 and 104.n can interact with client systems 102.1 and 102.n by utilizing operating systems and / or various applications executing on client systems 102.1 and 102.n.

[0046] In some embodiments, client applications (e.g., client applications executing on client systems 102, cloud applications 152, etc.) can send requests (also referred to as object storage requests) to distributed storage system 120 or object store 140 over communication network 110 to store, update, delete, or retrieve particular files, data objects, or other units of data stored in distributed storage system 120 and / or object store 140. For example, but not by way of limitation, a user 104 can configure a backup application to store a backup data set (such as a full snapshot of a storage device or selected volumes or files therein and / or incremental updates of changes since a previous snapshot or incremental update) to distributed storage system 120 or object store 140 periodically, in which case the backup application transmits requests to distributed storage system 120 or object store 140 to store the updates. A full copy of a storage device or selected volumes, directories, buckets, etc. can include all data objects or files contained in the selected storage device or selected sub-portions thereof.

[0047] Object storage requests can include information describing the objects being created and / or updated (such as file names, data including updates, client identifiers, operation types, etc.), and storage application 124 can use this information to record the updates, as described herein. In another example, a client application (e.g., an application executing on a client system 102, cloud application 152, etc.) can request objects or portions thereof, lists of objects matching particular criteria, etc., in which case the requests can include corresponding information (e.g., object identifiers, search criteria (e.g., time / date, keywords, etc.)) and receive the lists of objects or the objects themselves from storage application 124. Many other use cases are also applicable and contemplated.

[0048] Storage application 124 can provide object storage services, manage data storage (e.g., store, retrieve, and / or otherwise manipulate data in metadata store 130 and object store 140, etc.) using metadata store 130 and object store 140, process requests received from various entities (e.g., client systems 102, server systems 150, local applications, etc.), provide concurrency, provide data redundancy and duplication, perform garbage collection, and perform other actions, as discussed further herein. Storage application 124 can include various interfaces, such as software and / or hardware interfaces (e.g., application programming interfaces (APIs)) that can be accessed (e.g., locally, remotely, etc.) by components of system 100, such as various client applications, cloud applications 152, memory health monitor 154, etc.

[0049] In some embodiments, the storage application 124 can be a distributed application implemented in two or more computing systems (e.g., distributed storage systems 120.1-120.n). For example, the storage application 124 can be configured to provide a first level of protection for backup data using a slow distributed storage system 120.1 and a second level of protection for backup data using a fast distributed storage system 120.2. In some embodiments, the object store 140 can include multiple storage devices, servers, software applications, and other components such as, but not limited to, any suitable enterprise data grade storage hardware and software. In some embodiments, the storage application 124 can be a local application that receives local and / or remote storage requests from other clients (e.g., local applications, remote applications, etc.).

[0050] In non-limiting examples, the distributed storage system 120 can provide an object storage service such as a storage service that provides enterprise scale object storage functionality. Further examples of such storage services can include Amazon Simple Storage Service (S3) object storage services, such as the Amazon S3 TM , other local and / or cloud-based S3 storage systems / services.

[0051] The distributed storage system 120 can be coupled to and / or include an object store 140. The object store 140 can include one or more data stores for storing data objects. The object store 140 can be implemented across multiple physical storage devices. In some example embodiments, the multiple physical storage devices can be located at different locations. Objects stored in the object store 140 can be referenced by metadata entries stored in the metadata store 130. In some example embodiments, multiple copies (e.g., erasure coded copies) of a given object or portions thereof can be stored at different physical storage devices for protection against data loss due to system failures or to enable fast access to objects from different geographic locations.

[0052] The metadata store 130 can include a database that stores a collection of ordered metadata entries. Entries can be stored in response to object storage requests (such as, but not limited to, put, get, delete, list, etc.) received by a storage service. The storage service provided by the storage application 124 can instruct a metadata controller of the metadata store 130 to record data manipulation operations. For example, the storage service provided by the storage application 124 can call a corresponding method of the metadata controller of the metadata store 130 that is configured to perform various storage functions and act as needed depending on the configuration.

[0053] In some embodiments, metadata store 130 can comprise a horizontally partitioned database with two or more shards, although other suitable database configurations are possible and contemplated. As horizontal partitioning is a database design principle whereby rows of a database table are kept separately rather than split into columns (as is done to varying degrees by normalization and vertical partitioning), each partition can form part of a shard, which in turn can be located on a separate database server or physical location. Depending on the configuration, in some implementations, database shards can be implemented on different physical storage devices, as virtual partitions on the same physical storage device, or as any combination thereof.

[0054] Metadata store 130 and / or object store 140 can be included in distributed storage system 120, or in another computing system and / or a storage system that is different from, but coupled to or accessible by, distributed storage system 120. Metadata store 130 and / or object store 140 comprise one or more non-transitory computer-readable media (e.g., such as those discussed with reference to memory 316 in Figure 3 In some implementations, metadata store 130 and / or object store 140 can be combined with or can be different from memory 316. In some implementations, metadata store 130 and / or object store 140 can store data associated with a database management system (DBMS), such as a DBMS included and / or controlled by storage application 124 and / or other components of system 100. In some cases, a DBMS can store data in multidimensional tables composed of rows and columns, and use programmed operations to manipulate (e.g., insert, query, update, and / or delete) data rows, although other suitable DBMS configurations are applicable.

[0055] In some embodiments, the server system 150.n can host a memory health monitor 154, such as a monitoring application used by a memory manufacturer to collect memory health data from remote storage devices (e.g., the client system 102 and the storage device 106). For example, the memory health monitor 154 can use permission-based access to receive periodic updates of memory health data from each monitored storage device. The memory health monitor 154 can be configured to process the received memory health data and provide alerts to the user 104 when a change in memory health status is detected. In some embodiments, the memory health monitor 154 can also be configured to provide memory health data and / or alerts regarding changes in memory health status to other systems and applications, such as the cloud application 152 on the server system 150.1 and / or the storage application 124 on the distributed storage system 120. In some embodiments, other systems or applications, such as the cloud application 152 on the server system 150.1 and / or the storage application 124 on the distributed storage system 120, can receive memory health data and / or alerts directly from the client system 102 and / or the storage device 106.

[0056] It should be appreciated that Figure 1 The illustrated system 100 is representative of an exemplary system, and encompasses a variety of different system environments and configurations and is within the scope of the present disclosure. For example, in some additional embodiments, various functions can move between the server system 150 and the distributed storage system 120, from the server system 150 and the distributed storage system 120 to the client, or vice versa, modules can be combined and / or split into additional components, data can be consolidated into a single data store or further split into additional data stores, and some implementations can include additional or fewer computing devices, services, and / or networks, and can implement various functions on the client or server side. Moreover, various entities of the system 100 can be integrated into a single computing device or system or additional computing devices or systems, etc.

[0057] Figure 2 Selected components of a system 200 for providing differentiated data recovery configurations based on memory health data are schematically illustrated. In some embodiments, the system 200 can be implemented in an architecture similar to the system 100 in Figure 1 The distributed storage system 220 can be configured similarly to the distributed storage system 120. The storage devices 240 can be configured similarly to the storage devices 106 and / or the internal storage devices of the client system 102. The backup application 210 and / or components thereof can be hosted by the client system 102, the distributed storage system 220, the server system 150, and / or various combinations thereof.

[0058] The backup application 210 can include one or more software and / or hardware components for providing redundant backups for one or more storage devices 240. Data 244 from the storage devices 240 can be backed up to the distributed storage system 220. In the illustrated configuration, the backup application 210 can be supported by a fast storage system embodied in the distributed storage system 220, such as described above with respect to the second protection level configured to address gradual degradation of storage reliability. More specifically, the backup application 210 can be configured to support a differentiated data recovery configuration in the distributed storage system 220 based on changes in the memory health data 246. In some embodiments, different data recovery configurations can be selected and mapped to backup data at the storage devices and / or physical storage units using the data mapping 212. For example, the backup application 210 can implement a hierarchical physical model for the storage units 242 that enables different parity levels to be set at the erase block, page, die, and / or memory device level within the storage devices 240.

[0059] The distributed storage system 220 can store a redundant set of data sets 244 from the storage devices 240. For example, the backup application 210 can store backup copies of data stored in the storage devices 240. In some embodiments, the data 244 in the distributed storage system 220 can be stored according to the data recovery configuration 214 to provide additional redundancy and error correction for recovering lost or corrupted data. The backup application 210 can determine the data recovery configuration 214 and assign it to each data unit of the data 244 as it is stored, and / or reassign data units based on changes in the memory health state of the physical memory devices associated with corresponding data 244 in the storage devices 240. In some embodiments, the distributed storage system 220 can include slow storage options and fast storage options, and the data 244 can be stored in a combination of full backups of slow storage and incremental backups of fast storage.

[0060] Storage devices 240 can include any number of edge storage devices, such as SSDs, HDDs, hybrid drives, secure digital (SD) cards, universal serial bus (USB) sticks, or flash drives in other form factors that are standalone devices or integrated into a computer system, such as a personal computer (PC), laptop, smart phone, tablet, mobile phone, wearable electronic device, smart appliance, Internet of Things (IoT) device, embedded system, server, server farm, or any other electronic device with non-volatile memory, computer processing, and network communication capabilities. In some embodiments, storage devices 240 can include a plurality of physical storage units 242 embodying non-volatile memory for storing data 244. For example, storage units 242 can include memory devices, dies, media disks, and / or their physical subunits.

[0061] Storage devices 240 can each include one or more hardware and / or software modules for collecting memory health data 264 related to the non-transitory storage media contained by the storage device. For example, storage devices 240 can aggregate various measurements, counts, parameters, and other values and store them in memory locations for use by various storage device management functions within the storage device. The memory health data 246 stored in storage devices 240 can be accessed through a secure remote access or messaging protocol, such as the Internet Protocol, Remote Memory Access (RMA), or the like. In some embodiments, storage devices 240 can host an application or service for supporting backup application 210 through an application programming interface (API). For example, storage devices 240 can send periodic backup data 250 from data 244 to backup application 210 and can include health data 252 from memory health data 246 as part of those messages or data exchanges.

[0062] As shown, the backup application 210 can include a data mapping 212 configured to map data units (files, objects, blocks, etc.) from storage devices 240 to corresponding data locations in the allocated storage system 220, a data recovery configuration 214 configured to include a plurality of differentiated data recovery configurations corresponding to different error rates processed by each configuration, a memory health monitor 216 configured to collect memory health data 246 from the storage units 240, and a reliability manager 218 configured to evaluate the collected memory health data 246 and determine which data recovery configurations 214 to apply to backup data received from particular storage devices and / or storage units. For example, the backup application 210 can periodically receive backup data 250 and memory health data 252 from the storage devices 240 in messages, sessions, or responses and evaluate the received memory health data 252 against reduced reliability thresholds to determine which data recovery configurations 214 to apply to the received backup data. Further details are described below with respect to Figure 3 Various components that can support the operation of the backup application 210 in some embodiments are further described.

[0063] Figure 3 Selected modules of a server system hosting a backup application, a distributed storage system, and / or combinations thereof as described above are schematically illustrated. The system 300 can include a bus 310 that interconnects at least one communication unit 312, at least one processor 314, and at least one memory 316. The bus 310 can include one or more conductors that permit communication among the components of the system 300. The communication unit 312 can include any transceiver-like mechanism that enables the system 300 to communicate with other devices and / or systems. For example, the communication unit 312 can include a wired or wireless mechanism for communicating with file system clients, other access systems, and / or one or more object storage systems or components such as storage nodes or controller nodes. The processor 314 can include any type of processor or microprocessor that interprets and executes instructions. The memory 316 can include a random access memory (RAM) or another type of dynamic storage device that stores information and instructions for execution by the processor 314, and / or a read only memory (ROM) or another type of static storage device that stores static information and instructions for use by the processor 314, and / or any suitable type of storage element such as a hard disk or solid state storage element.

[0064] System 300 can include or have access to one or more databases and / or specialized data repositories, such as metadata repository 380 and data repository 390. Databases can include one or more data structures for storing, retrieving, indexing, searching, filtering, and / or the like of structured and / or unstructured data elements. In some embodiments, metadata repository 380 can be structured as reference data entries and / or data fields indexed by metadata key-value entries related to data objects stored in data repository 390. Data repository 390 can include data objects composed of object data, such as host data, an amount of metadata stored as metadata tags, and a globally unique identifier (GUID). Metadata repository 380, data repository 390, and / or other databases or data structures can be maintained and managed in separate computing systems, such as server nodes, storage nodes, controller nodes, or access nodes, with separate communication, processor, memory, and other computing resources, and accessed by system 300 through data access protocols. Metadata repository 380 and data repository 390 can be shared across multiple storage systems.

[0065] Storage system 300 can include a plurality of modules or subsystems that are stored and / or instantiated in memory 316 for execution by processor 314. For example, memory 316 can include a backup interface 320 configured to receive, process, and manage backup data and related requests from one or more remote storage devices. Memory 316 can include an encoding / decoding engine 330 configured for encoding and decoding symbols corresponding to data units (files, objects, messages, and / or the like) stored in data repository 390. Memory 316 can include a memory health monitor 340 configured for receiving and managing memory health data received from remote storage devices. Memory 316 can include a reliability manager 350 configured for managing different data recovery configurations for storing backup data based on one or more memory health states of storage devices. In some embodiments, backup interface 320, encoding / decoding engine 330, memory health monitor 340, and / or reliability manager 350 can be integrated into a backup application (such as backup application 210 in Figure 2 ), a storage application (such as storage application 124 in Figure 1 ), and / or managed as separate libraries or background processes (e.g., daemons) through APIs or other interfaces.

[0066] Backup interface 320 can include a set of interface protocols or functions and parameters for storing, restoring, and otherwise managing data backup requests from remote storage devices to data store 390. For example, backup interface 320 can include functions for reading, writing, modifying, or otherwise manipulating backup data objects and / or files, and their respective client or host data and metadata, according to the protocols of an object or file storage system. In some embodiments, backup interface 320 can implement a multi-tiered backup configuration including periodic full backups for slow storage and more frequent incremental backups for fast storage. Backup interface 320 can be further configured to use a differentiated data restore configuration for storing backup data based on memory health data. For example, incremental backups for fast storage can be configured to determine a memory health status and base the selection of a data restore configuration on the memory health status and / or changes thereto.

[0067] In some embodiments, backup interface 320 can include a plurality of hardware and / or software modules configured to use processor 314 and memory 316 to process or manage defined operations of backup interface 320. For example, backup interface 320 can include a backup scheduler 322, a backup data channel 324, a storage manager 326, and a backup user interface 328. For any given storage device, backup interface 320 can use backup scheduler 322 and backup data channel 324 to receive or initiate backup requests. These backup operations can include backup operations processed by storage manager 326, including encoding operations and decoding operations. In some embodiments, backup user interface 328 can be configured to enable user-initiated backups and / or set backup configuration parameters used by backup scheduler 322.

[0068] Backup interface 320 can receive messages including backup data and parse them according to appropriate communication and storage protocols. In some embodiments, backup interface 320 can identify a transaction identifier, a client identifier, an object identifier (object name or GUID), a data operation, and additional parameters for a backup data operation, if any, from one or more received messages that make up a backup data request.

[0069] The backup scheduler 322 can include interfaces, functions, or logic to receive backup data requests from remote storage devices and / or initiate backup data requests to remote storage devices on a defined schedule, as well as related data structures. For example, according to a time- or event-based schedule, a remote storage device can send a backup data request over a network connection and the backup data request is addressed to the system 300 or a port or component thereof, such as an API for the backup data channel 324. The backup scheduler 322 can include scheduling definitions and / or scheduling logic configured to determine when data backup communications should be initiated by a remote storage device. In some embodiments, the backup scheduler 322 can be configured as an initiator of periodic backups by a remote storage device and / or as a recipient of periodic backups by a remote storage device. In some embodiments, the backup scheduler 322 can include a backup configuration 322.1 and an incremental backup mode 322.2.

[0070] For example, the backup configuration 322.1 can include a plurality of backup configuration parameters that can be modified to change the operation of the backup interface 320. In some embodiments, the backup configuration 322.1 can be stored in a configuration file or similar data structure. For example, a default backup configuration and / or one or more user-defined backup configurations can be generated and / or received through the backup user interface 328. In some embodiments, the backup configuration 322.1 can include an allocation of backup levels, such as fast storage and / or slow storage, and / or a definition of a user backup service level, such as a backup tier regarding backup frequency, a recovery term (acceptable recovery time), a security level, and / or a resource allocation such as storage space allocation and / or other cloud resource allocation (e.g., processors, dedicated hardware / software services, network resources, dedicated hardware, etc.). In some embodiments, the backup configuration 322.1 can include a data structure or function defining a periodic backup schedule that initiates and / or receives backup data requests at a predetermined interval or based on other recurring trigger conditions.

[0071] In some embodiments, the incremental backup mode 322.2 can be a default backup mode and / or can be configured for storage in the backup configuration 322.1 through the backup user interface 328. The incremental backup mode 322.2 can include logic to identify periodic full backups of a remote storage device or a portion thereof and initiate incremental or differential backups configured to store differences between the most recent full backup and the current state of the storage device. In some embodiments, a series of incremental backups can be initiated between each full backup, and the incremental backups themselves can include all changes since the previous completed backup or changes since the immediately preceding incremental backup, if any. The incremental backups can be initiated by the backup scheduler 322 and / or the remote storage device. In some embodiments, the full backups can be scheduled and initiated by the backup scheduler 322, and the incremental backups can be initiated by the remote storage device in response to changes in the data stored therein.

[0072] The backup data channel 324 can include a fixed or configurable network data path for receiving backup data from a remote storage device. The backup data channel 324 can include a combination of physical and / or logical networking configurations compatible with the communication unit 312 and compatible interfaces, protocols, and / or parameters for receiving backup data from a remote storage device. In some embodiments, the backup data channel 324 can support streaming, session-based protocols, and / or remote memory access to support efficient transfer of large amounts of backup data. In some embodiments, the backup data channel 324 can operate independently of messaging channels used to initiate and manage the backup data channel 324, such as internet communications using standard HTTPS. In some embodiments implementing incremental backups, full backups can be initiated through the high-throughput backup data channel 324, and incremental backups can be initiated through a general communication channel. In some embodiments, the backup data channel 324 can receive both backup data and memory health data from a remote storage device, store the backup data to a backup data object in the data store 390, and store the memory health data to the metadata store 380.

[0073] The storage manager 326 can include interfaces, functions, and / or parameters for reading, writing, and deleting data elements in the data store 390. For example, an object PUT command can be configured to write an object identifier, object data, and / or object tags to the object store. An object GET command can be configured to read data from the object store. An object DELETE command can be configured to delete data from the object store, or at least mark a data object for deletion until a future garbage collection or similar operation actually deletes the data, or re-allocates the physical storage location to another purpose.

[0074] In some embodiments, the storage manager 326 can oversee the writing and reading of erasure coded data elements on storage media on which the data repository 390 is stored. When a message or data unit, such as a file or data object, is received for storage, the storage manager 326 can pass the file or data object through an erasure coding engine, such as the encoding / decoding engine 330. The data unit can be divided into data symbols, and the symbols can be encoded into erasure coded data symbols 392 for storage in the data repository 390. In some embodiments, the symbols can be distributed across multiple storage nodes to aid in fault tolerance, efficiency, recovery, and other considerations.

[0075] When a data unit is to be accessed or read, the storage manager 326 can identify the storage location of each symbol, such as using the data unit / symbol mapping 382 stored in the metadata repository 580. The erasure coded data symbols 392 can be passed through an erasure decoding engine, such as the encoding / decoding engine 330, to return the original symbols that make up the data unit to the storage manager 326. The data unit can then be reassembled and used by the backup interface 320 to complete a backup data recovery operation. The storage manager 326 can work with the encoding / decoding engine 330 for storing and retrieving the erasure coded data symbols 392 in the data repository 390.

[0076] In some embodiments, the storage manager 326 can include or interface with a metadata manager for creating, modifying, deleting, accessing, and / or otherwise managing object or file metadata, such as the metadata stored in the metadata repository 380. For example, when a new object is written to the data repository 390, at least one new metadata entry can be created in the metadata repository 380 to represent parameters describing the newly created object or related thereto. The metadata manager can generate and maintain metadata that enables the storage manager 326 to locate object or file metadata within the metadata repository 380. For example, the metadata repository 380 can be organized as a key-value repository, and the object metadata can include key values for data objects and / or operations related to those objects indexed with keys including object identifiers or GUIDs for each object.

[0077] In some embodiments, the storage manager 326 can use the metadata manager to store physical unit / data unit mappings 384 and / or a memory health data repository 386 in the metadata repository 380. For example, the physical unit / data unit mappings 384 can include entries of reference values, such as logical block addresses and storage unit identifiers, that map physical storage units in a storage device to backup data units stored in those remote storage units, which can be updated and used by the memory health monitor 340 and / or the reliability manager 350. The memory health data repository 386 can include recent and / or historical memory health data related to similar populations and types of storage devices that are remote storage devices that are being backed up and / or for use by the memory health monitor 340 and / or the reliability manager 350.

[0078] The backup user interface 328 can include APIs, functions, and / or parameters for user-configurable management of the backup interface 320. For example, the backup user interface 328 can provide a graphical user interface and underlying logic and data structures for enabling a user to manage the backup scheduler 322 and / or configure the backup data channel 324. In some embodiments, the backup user interface 328 can be set up as a graphical user interface that is Internet protocol-enabled and accessible via the communication unit 312 from one or more remote computer systems. For example, a user of a remote storage device can be able to securely log into the backup user interface 328 from another computer system, such as a computer system that includes or is connected to the remote storage device or another computer system that is accessible to the same user.

[0079] The encoding / decoding engine 330 can include a collection of functions and parameters for storing, reading, and otherwise managing encoded data, such as erasure encoded data symbols 392, in the data repository 390. For example, the encoding / decoding engine 330 can include functions for encoding user data symbols into erasure encoded data symbols and decoding erasure encoded data symbols back into original user data symbols. In some embodiments, the encoding / decoding engine 330 can be included in a write path and / or a read path of the data repository 390 managed by the storage manager 326. In some embodiments, the encoding and decoding functions can be placed in separate encoding and decoding engines with redundant and / or shared functions, where similar functions are used by both encoding operations and decoding operations.

[0080] In some embodiments, the encoding / decoding engine 330 can include a plurality of hardware modules and / or software modules configured to use the processor 314 and the memory 316 to process or manage defined operations of the encoding / decoding engine 330. For example, the encoding / decoding engine 330 can include an erasure coding configuration 332, a symbol divider 334, and an encoder / decoder 336.

[0081] Erasure encoding configuration 332 can include functions, parameters, and / or logic for determining operations for dividing data units into symbols, encoding and decoding those symbols. For example, there are various erasure encoding algorithms for providing forward error correction based on converting a message of a particular number of symbols into a longer message of more symbols such that the original message can be recovered from a subset of the encoded symbols. In some embodiments, a message can be divided into a fixed number of symbols, and those symbols used as input to erasure encoding. A system erasure encoding algorithm can produce the original symbols and a fixed number of additional parity symbols. The sum of these symbols can then be stored to one or more storage locations.

[0082] In some embodiments, erasure encoding configuration 332 can enable encoding / decoding engine 330 to be configured according to available encoding algorithms 332.1 and encoding block sizes 332.2 supported by data store 390. For example, encoding algorithms 332.1 can enable selection of an algorithm type (such as parity-based, low-density parity-check code, Reed-Solomon code, etc.) and one or more algorithm parameters (such as a number of original symbols, a number of encoded symbols, an encoding rate, a reception efficiency, etc.). Encoding block sizes 332.2 can enable selection of a block size of encoded symbols. For example, an encoding block size can be selected to align with storage media considerations, such as an erase block size for a sold-state drive (SSD) and / or a symbol size that aligns with data unit and / or subunit of data operation parameters. Erasure encoding configuration 332 can also include a parity level 332.3. Parity level 332.3 can be a configurable parameter that determines a ratio between user data (such as host or backup data) and parity data. Parity level 332.3 can determine how many symbols can be corrupted or lost and still recover the entire message. In some embodiments, parity level 332.3 can be related to a maximum error rate from which backup data can be recovered, with a higher parity level corresponding to a higher allowable error rate and also a higher storage space usage for additional parity symbols. In some embodiments, erasure encoding configuration 332 can include an interface or API (such as a configuration service) to enable reliability manager 350 to select one or more parameters of encoding algorithm 332.1, encoding block size 332.2, and / or parity level 332.3 to change data recovery configurations for backup data from different remote storage devices and / or storage units therein.

[0083] Symbolizer 334 can include functions, parameters, and / or logic to receive messages to encode and partition the messages into a series of raw symbols based on data in the messages. For example, a default symbolization operation can receive a message and use a symbol size defined by erasure coding configuration 332 to partition the message into a fixed number of symbols for encoding.

[0084] Encoder / decoder 336 can include hardware and / or software encoders and decoders to implement encoding algorithm 332.1. For example, encoder / decoder 336 can include a number of register-based encoders and decoders to compute parity for symbols and return erasure coded data symbols 392. In some embodiments, encoder / decoder 336 can be integrated into the write path and the read path, respectively, such that data to be written to the storage media and data read from the storage media pass through encoder / decoder 336 for encoding and decoding according to encoding algorithm 530.1.

[0085] Memory health monitor 340 can include a collection of functions and parameters to collect memory health data from remote storage devices for use by reliability manager 350 and other system functions. For example, memory health data can be collected from remote storage devices through backup interface 320 or another data channel and stored in memory health data repository 386 in metadata store 380. Memory health monitor 340 can receive memory health data including one or more parameters related to the functioning of storage media in remote storage devices and related to memory health status. In some embodiments, memory health monitor 340 is integrated with backup interface 320 and / or reliability manager 350. In some embodiments, memory health monitor 340 can be a separate function or service operated to monitor remote devices and report memory health data parameters and / or other operational parameters for storage device management, quality, security, and other services.

[0086] In some embodiments, memory health monitor 340 can include a number of hardware modules and / or software modules configured to use processor 314 and memory 316 to process or manage defined operations of memory health monitor 340. For example, memory health monitor 340 can include memory data authority 342, memory data identifier 344, memory data handler 346, and / or memory data manager 348.

[0087] Memory data permissions 342 can include functions, parameters, and / or data structures for configuring and storing user or storage device owner permissions to access remote storage devices. For example, memory health monitor 340 can be configured to access only remote storage devices for which its explicit access permissions have been granted by a user or owner. In some embodiments, a remote storage device can include a configuration tool that is launched from the storage device or as part of a storage device registration process, and enables a user to authorize access to memory health data collected by the storage device. For example, when a user selects or configures a backup service through backup user interface 328, the user can also be given the option to enable remote memory health monitoring. Memory data access permissions from storage device owners can be collected and stored in memory data permissions 342 and / or can be selectively accessed by memory data permissions 342 (such as through a user profile or similar data structure available to system 300). In some embodiments, memory data permissions 342 can enable selection of specific data health parameters, collection frequency, allowed uses, and other configurable parameters to be used by memory health monitor 340.

[0088] Memory data identifiers 344 can include functions, parameters, and / or logic for identifying one or more types of memory health data that can be monitored by memory health monitor 340. Any storage device type can support various operational parameters that can correspond to or be used as memory health data indicative of a memory health status. Storage devices can also vary with respect to the memory health data parameters that are collected and whether they are configured for remote access. Reliability manager 350 can also be configured to support specific memory health data types. In some embodiments, memory health data types can include BER values, W / E cycle values, PLC values, ELC values, leak detection measurement values, EPD error values, bad block values, voltage margin values, and other operational values that are directly or indirectly related to a memory health status. In some storage devices, memory health data can include aggregate or derived values based on memory health data types, and these aggregate or derived values can be managed as different memory health data types. In some embodiments, memory health data identifiers 344 can include parameters for matching memory health data collected from a particular storage device or storage device type to memory health data types used by reliability manager 350 to determine a reliability condition. Memory health data identifiers 344 can also include logic for tagging or locating desired memory health data types received by memory data handler 346, and stored in memory health data repository 386.

[0089] Memory data handler 346 can include functions, parameters, and / or logic to receive or identify memory health data received by system 300. For example, memory data handler 346 can be configured to receive memory health data received from remote storage devices over backup data channel 324. In some embodiments, backup interface 320 can receive memory health data from remote storage devices as part of a message or other data transfer, and store and / or send the memory health data directly to memory data handler 346. In some embodiments, memory data handler 346 can be configured to operate independently of backup interface 320 for receiving memory health data from remote storage devices. For example, memory data handler 346 can include remote access functionality to read memory health data from remote storage devices over a network, such as the Internet. Memory data handler 346 can include an API to receive memory health data from remote storage devices over a direct addressing protocol, or after memory health data is collected and formatted by an intermediary system, such as a server system for collecting memory health data for a particular memory manufacturer.

[0090] Memory data manager 348 can include functions, parameters, and / or logic to store and retrieve memory health data in metadata store 380, such as in memory health data store 386. For example, memory data manager 348 can use memory health data types identified by memory data identifier 344 to select and format memory health data received by memory data handler 346. In some embodiments, memory data manager 348 can store memory health data in memory health data store 386 using unique storage device identifiers in device index 348.1. For example, device index can include a storage device identifier for each remote storage device monitored by memory health monitor 340 and stored in device data 386.1. Each storage device entry in device data 386.1 can include one or more memory health data types and corresponding memory health data values. Device data 386.1 can be organized as a key-value store, with multiple entries for each storage device representing different memory health data types and / or values from different collection time points. In some embodiments, memory data manager 348 can be configured to selectively return memory health data, memory health data types, and / or time points or time periods for a requested storage device, such as an array of memory health data values corresponding to desired parameters.

[0091] In some embodiments, the memory health data can be further organized by the memory health monitor 340 and the memory health data repository 386 according to physical memory sub-units of the storage device. For example, a flash storage device can include multiple flash memory dies as storage media within the storage device, and an identifier for each die can be included in the memory health data to associate the memory health data values with a particular die. In some embodiments, the memory data manager 348 can include a cell index 348.2 for organizing and indexing storage cell identifiers for storage cells within each storage device. For example, the cell index 348.2 can be an extension of the device index 348.1 that supports one or more layers of hierarchical identifiers to enable the storage device to be divided into physical storage cells for monitoring memory health data. In some embodiments, any number of hierarchical levels based on the physical hierarchy of the storage media sub-units can be assigned an identifier value, and used to manage monitoring, storage, and retrieval of memory health data based on the storage cells of the physical storage media. In some embodiments, a data structure for mapping data cells stored in the data repository 390 to the physical storage cells of the remote storage cells they came from can be included in the physical cell / data cell mapping 384 in the metadata repository 380.

[0092] The reliability manager 350 can include a collection of functions and parameters for managing different data recovery configurations for storing backup data dynamically based on memory health data of remote storage devices containing the original data being backed up. For example, as the backup interface 320 receives backup data from a remote storage device, the reliability manager 350 can evaluate the memory health data collected by the memory health monitor 340 to determine which erasure coding configuration 332 the encoding / decoding engine 330 should use when writing the corresponding backup data objects to the data repository 390. The reliability manager 350 can be configured to optimize the use of storage space and / or other system resources based on the most recent memory health data available for the actual storage device and / or its sub-units being backed up. The reliability manager 350 can change the data recovery configuration of the stored backup data from the remote storage device during the lifetime of the storage device based on changes in the memory health status and the corresponding reliability condition of the remote storage device.

[0093] In some embodiments, the reliability manager 350 can include a plurality of hardware and / or software modules configured to use the processor 314 and the memory 316 to process or manage defined operations of the reliability manager 350. For example, the reliability manager 350 can include a data recovery configuration, a service level checker 354, a reliability condition logic 356, a reliability modeler 358, and a change engine 360. In some embodiments, the reliability manager 350 can include or have access to the memory data manager 348 and / or the erasure coding configuration 332.

[0094] The data recovery configuration 352 can include functions, parameters, and / or data structures to define a plurality of different data recovery configurations that can be used to store backup data in the data store 390 with different levels of error correction capability and / or acceptable error rates. For example, each recovery configuration 352.1-352.n can correspond to a different erasure coding configuration 332 that results in a different configuration of erasure coded data symbols 392 in the data store 390. In some embodiments, the data recovery configuration 352 can include a range of configurations based on specific erasure coding parameters such as a range of parity levels 332.3 and incremental values. In some embodiments, each of the recovery configurations 352.1-352.n can correspond to a set of erasure coding configuration parameters such as different encoding algorithms 332.1, different block sizes 332.2, and / or different parity levels 332.3. In some embodiments, the recovery configurations 352.1-352.n can be organized in a table, configuration file, or similar data structure. In some embodiments, the recovery configurations 352.1-352.n can be functionally determined based on one or more functions that return desired erasure coding configuration parameters and corresponding parameters.

[0095] The service level checker 354 can include functions, parameters, and / or logic to determine a service level associated with a remote storage device and / or an associated user / owner profile or service contract. For example, a remote storage device can be configured according to a specific offering of cloud-based backup storage, and multiple service levels can be provided for the backup storage such as frequency, storage space, recovery options, etc. These service levels can correspond to pricing tiers under a service level agreement or similar terms and conditions. In some embodiments, the service level checker 354 can identify or determine one or more service level parameters that can influence the recovery configuration available to and / or selected by the reliability manager 350 for backup data related to a specific storage device and corresponding service level agreement. For example, the change engine 360 can use one or more service level parameters to determine a recovery configuration to use based on a reliability condition.

[0096] Reliability condition logic 356 can include functions, parameters, and / or logic to determine or identify a reliability condition or a change in reliability condition based on memory health data. For example, reliability condition logic 356 can use one or more memory health data values associated with a remote storage device or its storage units as an index or input value for determining whether a reliability condition has decreased and a different data recovery configuration should be implemented for the positive storage device or storage unit. In some embodiments, reliability condition logic 356 can receive one or more memory health data values as input and provide a reliability condition value or a decreased reliability condition state as output. In some embodiments, reliability condition logic 356 can determine a change in a reliability condition of a storage device and identify the change to change engine 360 to determine a response of reliability manager 350 and system 300 to the change.

[0097] In some embodiments, reliability condition logic 356 can be based on one or more reliability thresholds 356.1 corresponding to one or more memory health data values or aggregated or derived values thereof. For example, memory health data for a particular storage device can include BER values for data read from the storage device over time. Reliability condition logic 356 can implement a plurality of reliability thresholds based on the BER values. For example, below a first BER reliability threshold, a storage device can be considered to be in its prime operating condition and a high reliability condition. If the BER meets or exceeds the first BER reliability threshold, but is below a second BER reliability threshold, the storage device can be considered to be in a fair operating condition with a normal reliability condition, i.e., a decreased reliability condition relative to the high reliability condition. If the BER meets or exceeds the second BER reliability threshold, but is below a third BER reliability threshold, the storage device can be considered to be in a degraded operating condition with a risky reliability condition, i.e., a decreased reliability condition relative to the high and normal reliability conditions. If the BER meets or exceeds the third BER reliability threshold, the storage device can be considered to be in a damaged operating condition with a critical risky reliability condition, i.e., a decreased reliability condition relative to the high, normal, and risky reliability conditions. These example thresholds and tiers are provided by way of example only, and a system can implement any number of thresholds and decreased reliability condition tiers.

[0098] In some embodiments, the reliability condition logic 356 can include a reliability condition table 356.2. For example, the reliability condition table 356.2 can map a plurality of memory health data types and values to reliability thresholds, such as the reliability thresholds 356.1, for determining a reliability condition. In some embodiments, the reliability condition table 356.2 can define a multivariate process for evaluating a plurality of memory health data types against a plurality of reliability thresholds to generate a single reliability condition value or state. For example, the reliability condition table 356.2 can embody different ranges and logical combinations (and, or, or not, etc.) for evaluating different memory health data types, such as requiring both a BER and PLC or ELC value to exceed a particular threshold or fall within a particular range to correlate to a change in reliability condition, such as a reduced reliability condition.

[0099] In some embodiments, the reliability condition logic 356 can include a reliability rules engine 356.3. For example, the reliability rules engine 356.3 can be embodied in a logical rule evaluator and a plurality of logical rules defining a set of reliability rules. In some embodiments, the reliability rules engine 356.3 can evaluate the set of reliability rules for each backup data transaction and / or on a periodic basis to determine a reliability condition of a storage device or unit at a recent memory health data. In some embodiments, at least a portion of the set of reliability rules can include reliability threshold comparisons, as described above with respect to the reliability thresholds 356.1 and / or the condition table 356.2.

[0100] The reliability modeler 358 can include functions, parameters, and / or logic for modeling a plurality of memory health data parameters corresponding to changes in reliability condition. For example, the reliability modeler 358 can use historical reliability data from a population of storage devices and corresponding memory health data to generate parameters and / or rules for the reliability condition logic 356. In some embodiments, the reliability modeler 358 can use population data 386.2 for a large number of storage devices of the same storage device type to derive a reliability model 386.3. For example, the population data 386.2 can include a statistically significant number of storage devices of a particular type, manufacturer, form factor, model, etc. that can be used to determine a reliability model 386.3 corresponding to those device type parameters. In some embodiments, the reliability model 386.3 can be used to determine the reliability thresholds 356.1, parameters for the condition table 356.2, and / or reliability rules for the reliability rules engine 356.3. In some embodiments, the reliability modeler 358 can be based on a statistical analysis of a population of storage devices, failure rates, and memory health parameters.

[0101] Change engine 360 can include functions, parameters, and / or logic to determine that a reliability condition has changed for a given storage device or its physical storage units and generate a corresponding change to the data recovery configuration. For example, reliability condition logic 356 can determine that a reliability threshold or other indicator has been satisfied for a change in reliability condition and pass the changed reliability condition value to change engine 360. Change engine 360 can include an index or transfer function to determine a different data recovery configuration for storing backup data from the storage device or storage unit with the changed reliability condition than data recovery configuration 352. Change engine 360 can pass the new data recovery configuration to backup interface 320 and / or encoding / decoding engine 330 for storing backup data from the storage device in the forward direction.

[0102] In some embodiments, change engine 360 can also initiate reconfiguration of previously stored backup data from the storage device with the changed reliability condition. For example, change engine 360 can use physical unit / data unit mapping 384 to identify all data units that have been backed up from the implemented storage device or storage unit and generate one or more commands to read and store the data units previously stored with an earlier data recovery configuration (with a lower error rate tolerance) in the new data recovery configuration (with a higher error rate tolerance). In some embodiments, reconfiguration of existing backup data can be performed as a background operation or lower priority task compared to new backup requests. In some embodiments, particularly those that implement incremental backups, the incremental backup can be immediately converted to the new data recovery configuration and the new data recovery configuration can be used for the next full backup, but the previous full backup can not be reconfigured for the new data recovery configuration.

[0103] Memory 316 can include additional logic and other resources (not shown) for processing data requests, such as modules for generating, queuing, and otherwise managing object or file data requests. Processing of backup data requests by backup interface 320 can include any number of intermediate steps to produce at least one backup data request to data store 390.

[0104] As Figure 4 shown, system 300 can operate according to an example method of determining different data recovery configurations according to changes in memory health data (i.e., according to method 400 shown in blocks 402-412 of Figure 4 .

[0105] At block 402, a first data recovery configuration can be determined for data units to be stored from the remote storage device. For example, responsive to receiving a backup data request, the storage manager can determine, with assistance of the reliability manager, a first data recovery configuration, such as a standard or initial data recovery configuration used by the storage system receiving the data units.

[0106] At block 404, data corresponding to data received from the remote storage device can be stored in the storage system using the first data recovery configuration. For example, with assistance of the encoding / decoding engine, the storage manager can store data units corresponding to at least a portion of the backup data request using the first data recovery configuration.

[0107] At block 406, memory health data can be received from the remote storage device. For example, the memory health monitor can receive memory health data from the remote storage device.

[0108] At block 408, the memory health data can be evaluated for a change in reliability condition. For example, the reliability manager can compare the memory health data received at block 406 to one or more reliability thresholds to determine whether the memory health data has a significant change indicative of a substantial change in reliability of the remote storage device. If not, there is no significant change, and the method 400 can return to block 404 to store the next data unit using the first data recovery configuration. If so, there is a significant change in reliability, and the method 400 can proceed to block 410.

[0109] At block 410, a second data recovery configuration can be determined for data units to be stored from the remote storage device. For example, responsive to receiving a next backup data request, the storage manager can use a second data recovery configuration determined by the reliability manager responsive to the change determined at block 408, such as an erasure coding configuration capable of handling a higher error rate.

[0110] At block 412, data corresponding to data received from the remote storage device can be stored in the storage system using the second data recovery configuration. For example, with assistance of the encoding / decoding engine, the storage manager can store data units corresponding to at least a portion of the new backup data request using the second data recovery configuration.

[0111] As Figure 5 shown, the system 300 can operate according to an example method of managing backup configurations based on memory health data (i.e., according to the method 500 shown in FIGS. 502-522). Figure 5

[0112] ​At block 502, a periodic backup configuration can be determined for the remote storage device. For example, the backup interface can include a backup scheduler configured to determine a time- or event-based schedule for periodic backups of the remote storage device.

[0113] At block 504, an initial memory health value can be determined for the remote storage device. For example, the initial memory health value can be received by the memory health monitor when the backup configuration is set for the remote storage device.

[0114] At block 506, a service level can be determined for the remote storage device. For example, the service level can be associated with the remote storage device and / or a user / owner of the storage device as part of a backup service through the backup interface.

[0115] At block 508, a resource allocation for backups of the remote storage device can be determined. For example, the backup interface can use the service level determined at block 506 to determine storage, processing, network, and / or other resources available for backup operations of the remote storage device.

[0116] At block 510, a first data recovery configuration can be determined for the storage device. For example, based on the initial memory health value and the resource allocation, the reliability manager can determine a first data recovery configuration having a first error rate tolerance for storing backup data.

[0117] At block 512, backup data from the remote storage device can be received at the backup storage system. For example, the backup interface can receive the backup data from the remote storage device according to the periodic backup configuration determined at block 502.

[0118] At block 514, the data from the remote storage device can be encoded at a first parity level and stored in the storage system. For example, the encoding / decoding engine can encode data units corresponding to the backup data according to the parity level defined by the first data recovery configuration.

[0119] At block 516, memory health data can be received from the remote storage device. For example, the memory health monitor can receive the memory health data from the remote storage device via the same data channel and / or messages as the backup data received at block 512.

[0120] At block 518, a changed memory health value can be determined for the remote storage device. For example, the reliability manager can evaluate the memory health data received at block 516 to determine whether the memory health value has changed such that the reliability condition of the storage device has decreased.

[0121] At block 520, a second data recovery configuration can be determined for the storage device based on the changed memory health value determined at block 518. For example, the reliability manager can use the changed memory health value to identify a reduced reliability condition that corresponds to a data recovery configuration with a higher error rate allowance.

[0122] At block 522, data from the remote storage device can be encoded at the second parity level and stored in the storage system. For example, the encoding / decoding engine can encode data units corresponding to the backup data according to the parity level defined by the second data recovery configuration.

[0123] As shown, the system 300 can operate according to an example method of differentiating data recovery configurations within storage units of a storage device based on memory health data (i.e., according to the method 600 shown in blocks 602-622 of Figure 6 . Figure 6

[0124] At block 602, a plurality of physical storage units within a storage device can be determined. For example, a remote storage device can be composed of storage media that includes different dies for storing data corresponding to a particular set of physical addresses, and when those units are received for backup, the backup storage system can receive identifiers of the dies storing the original data units.

[0125] At block 604, reference values for mapping user data to physical storage units can be stored. For example, the backup interface can store the storage unit / data unit mappings in metadata based on the storage unit identifiers or reference values received from the remote storage device.

[0126] At block 606, memory health data can be received from the remote storage device for each physical storage unit. For example, the memory health monitor can receive memory health data values for each die or similar physical storage unit from which backup data has been received.

[0127] At block 608, a physical storage unit can be selected from the plurality of physical storage units. For example, the reliability engine can select a first physical storage unit to evaluate for reduced reliability and then systematically evaluate each physical storage unit in order.

[0128] At block 610, a memory health value for the selected physical storage unit can be determined. For example, the memory health monitor can store the received memory health value in a data structure that enables the physical storage unit to retrieve the value for the reliability engine.

[0129] ​At block 612, the selected physical storage unit can be evaluated for a reduced reliability condition. For example, the reliability engine can evaluate the memory health value of the selected physical storage unit for a reliability threshold. If no, the selected storage unit does not have a reduced reliability condition, and the method 600 can proceed to block 614. If yes, the selected storage unit does have a reduced reliability condition, and the method 600 can proceed to block 616.

[0130] At block 614, data units from the selected physical storage unit can continue to be stored to the backup storage system using the first data recovery configuration. For example, pending and future backup requests having a reference value or storage unit identifier that matches the selected physical storage unit can be stored by the storage manager using the first data recovery configuration.

[0131] At block 616, data units from the selected physical storage unit can be stored to the backup storage system using a second data recovery configuration. For example, pending and future backup requests having a reference value or storage unit identifier that matches the selected physical storage unit can be stored by the storage manager using the second data recovery configuration having a higher error rate tolerance.

[0132] At block 618, a determination can be made as to whether there are additional physical storage units to evaluate. For example, the reliability manager can determine whether the selected physical storage unit is not the last physical storage unit in the storage device. If yes, there are more units to evaluate, and the method 600 can return to block 108 to select the next physical storage unit. If no, there are no more units to evaluate, and the method 600 can proceed to block 620.

[0133] At block 620, previously stored data from any physical storage unit now having a reduced reliability condition can be read from the backup storage system using the first data recovery configuration. For example, the reliability manager can initiate a reconfiguration of previously stored backup data having a storage unit identifier or other reference value associated with a physical storage unit having a reduced reliability condition. In response, the storage manager can read the previously stored data according to the data recovery configuration by which the data was stored, such as the first data recovery configuration.

[0134] At block 622, the previously stored data can be reallocated in the backup storage system using a second data recovery configuration. For example, once the storage manager reads the previously stored data according to the data recovery configuration by which the data was stored, the backup data units having a reference value or storage unit identifier that matches a physical storage unit having a reduced reliability condition can be stored using the second data recovery configuration having a higher error rate tolerance.

[0135] As Figure 7As shown, system 300 can operate according to another example method of changing data recovery configuration according to reduced reliability conditions (i.e., according to the method 700 shown in blocks 702-712). Figure 7

[0136] At block 702, historical memory health data can be collected for a population of storage devices. For example, memory health data can be aggregated from a memory health data repository from quality test data, warranty and return data, memory health monitoring, and other data sources that enable memory health data to be correlated with reliability outcomes for the population of storage devices.

[0137] At block 704, the historical memory health data can be accessed for analysis of reliability patterns and memory health data indicators. For example, a reliability modeler can be configured to identify historical data sources to be used to generate data reliability models for one or more storage device types.

[0138] At block 706, at least one data reliability model can be determined. For example, a reliability modeler can use statistical modeling to analyze historical data to determine memory health data types and thresholds related to one or more layers of reduced reliability.

[0139] At block 708, at least one reduced reliability condition can be determined that can be used to evaluate individual storage devices. For example, based on the data reliability models, a reliability manager can be configured with selected memory health data types and corresponding thresholds to identify reduced reliability conditions.

[0140] At block 710, memory health data for remote storage devices can be evaluated for reduced reliability conditions. For example, a reliability manager can compare memory health data values for selected memory health data types to corresponding reliability thresholds from the data reliability models to identify storage devices that have reached a reduced reliability condition.

[0141] At block 712, a change in data recovery configuration can be initiated. For example, a reliability manager can select a different data recovery configuration with a higher error rate tolerance for incoming backup data requests from a storage device.

[0142] ​While at least one example embodiment has been presented in the foregoing detailed description of the technology, it should be appreciated that a vast number of modifications can be made. It should also be appreciated that the exemplary embodiment or embodiments are examples, and are not intended to limit the scope, applicability or configuration of the technology in any way. Rather, the foregoing detailed description will provide those skilled in the art with a convenient road map for implementing an exemplary embodiment of the technology, it being understood that various modifications can be made in the function and / or arrangement of elements described in an exemplary embodiment without departing from the scope of the technology as set forth in the appended claims and the legal equivalents thereof.

[0143] As those skilled in the art will appreciate, the various aspects of the technology can be embodied as a system, method, or computer program product. Accordingly, some aspects of the technology can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or a combination of hardware and software aspects that can all generally be referred to herein as a circuit, module, system and / or network. Furthermore, various aspects of the technology can take the form of a computer program product embodied in one or more computer readable medium(s) including computer readable program code embodied thereon.

[0144] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a physical computer readable storage medium. For example, the physical computer readable storage medium can be, but is not limited to, an electronic, magnetic, optical, crystal, polymer, electromagnetic, infrared, or semiconductor system, apparatus, or device, etc., or any suitable combination of the foregoing. Non-limiting examples of physical computer readable storage media can include, but are not limited to, an electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a compact disk read-only memory (CD-ROM), an optical processor, a magnetic processor, etc., or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain or store a program or data for use by or in connection with an instruction execution system, apparatus, or device.

[0145] Computer code embodied on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, radio frequency (RF), and the like, or any suitable combination of the foregoing. Computer code for carrying out operations for aspects of the technology can be written in any static language, such as the C programming language or other similar programming language. The computer code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network or communication system, including but not limited to a local area network (LAN) or a wide area network (WAN), an aggregated network, or a connection that can be established to an external computer (e.g., through the Internet using an Internet service provider).

[0146] Aspects of the technology can be described above with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, systems, and computer program products. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing device (processor) of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing device or other programmable data processing apparatus, create means for implementing the operations / acts specified in the flowchart and / or block diagram block or blocks.

[0147] Some computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or one or more other devices to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions which implement the operations / acts specified in the flowchart and / or block diagram block or blocks. Some computer program instructions can also be loaded onto a computing device, other programmable data processing apparatus, or one or more other devices to cause a series of operational steps to be performed on the computing device, other programmable apparatus, or one or more other devices to produce a computer-implemented process such that the instructions executed by the computer or other programmable apparatus provide one or more processes for implementing the operations / acts specified in the flowchart and / or block diagram block or blocks.

[0148] The flow diagrams and / or block diagrams in the above drawings illustrate architecture, functionality, and operation of possible implementations of apparatuses, systems, methods and / or computer program products according to aspects of the present technology. In this regard, each block in the flow diagrams and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative aspects, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending on the functionality involved. It will also be noted that one or more blocks of the block diagrams and / or flow diagrams and combination thereof can be implemented by special purpose hardware-based systems, which can be operational with one or more computer programs to perform the operations described in the blocks.

[0149] While one or more aspects of the present technology have been illustrated and described in detail, one of ordinary skill in the art will understand that modifications and / or adjustments can be made to the various aspects without departing from the scope of the present technology as set forth in the following claims.

Claims

1. A computer-implemented method comprising: storing a redundant data set from a remote storage device in a distributed storage system using a first data recovery configuration; receiving memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device; determining a change in the memory health state of the non-transitory media of the remote storage device based on the memory health data; re-allocating the redundant data set in the distributed storage system using a second data recovery configuration in response to the change in the memory health state; determining a periodic backup configuration of the remote storage device; determining at least one initial memory health value of the remote storage device; and determining the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration.

2. The computer-implemented method of claim 1, wherein: the remote storage device is a non-volatile memory device located at a site remote from the distributed storage system; and re-allocating the redundant data set in the distributed storage system comprises periodically backing up differences between a current data set stored on the remote storage device and a full copy of data stored on the remote storage device at an earlier time.

3. The computer-implemented method of claim 1, further comprising: determining a service level of at least one system resource of the distributed storage system; determining an allocation of the at least one system resource for storing the redundant data set in the distributed storage system based on the service level, wherein determining the first data recovery configuration is further based on the allocation of the at least one system resource; and determining the second data recovery configuration based on the allocation of the at least one system resource and the change in memory health.

4. The computer-implemented method of claim 1, wherein: storing the redundant data set in the distributed storage system using the first data recovery configuration comprises encoding the redundant data set in a first plurality of encoded data symbols according to a first parity level; re-allocating the redundant data set in the distributed storage system using the second data recovery configuration comprises encoding at least a portion of the redundant data set in a second plurality of encoded data symbols according to a second parity level; and the second parity level is adapted for a different error rate for recovering the portion of the redundant data set than the first parity level.

5. The computer-implemented method of claim 1, wherein the memory health data comprises at least one memory health value selected from: a bit error rate value; a write / erase cycle value; a program cycle counter value; an erase cycle counter value; a leak detection measurement value; an unstable program disturb value; a bad block value; or a voltage margin value.

6. The computer-implemented method of claim 1, further comprising: ​ receiving the redundant data set from the remote storage device according to a periodic backup schedule, wherein receiving memory health data from the remote storage device is performed in conjunction with receiving the redundant data set according to the periodic backup schedule.

7. The computer-implemented method of claim 1: further comprising: determining a plurality of physical storage units in the remote storage device; and storing reference values that associate the redundant data set stored in the distributed storage system with the plurality of physical storage units in the remote storage device that store corresponding data; wherein: receiving memory health data from the remote storage device includes receiving at least one memory health value for each physical storage unit of the plurality of physical storage units; determining the change in the memory health state includes: determining that at least one memory health value for a first physical storage unit of the plurality of physical storage units satisfies a reduced reliability condition; and determining that at least one memory health value for a second physical storage unit of the plurality of physical storage units does not satisfy the reduced reliability condition; reassigning the redundant data set in the distributed storage system using the second data recovery configuration includes storing data associated with the first physical storage unit using the second data recovery configuration in response to determining the reduced reliability condition; and data associated with the second physical storage unit remains stored using the first data recovery configuration.

8. The computer-implemented method of claim 1, wherein determining the change in the memory health state includes: determining at least one reduced reliability threshold; and evaluating the memory health data against the at least one reduced reliability threshold.

9. The computer-implemented method of claim 1, further comprising: collecting historical memory health data for a population of remote storage devices of a remote storage device type associated with the remote storage device; determining a data reliability model for the remote storage device type based on the collected historical memory health data; and determining at least one reduced reliability threshold based on the data reliability model, wherein determining the change in memory health includes evaluating the memory health data against the at least one reduced reliability threshold.

10. A system for storage, comprising: a storage system configured to store a redundant data set from a remote storage device using a first data recovery configuration; a memory health monitor configured to receive memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device; a backup interface configured to determine a periodic backup configuration for the remote storage device; a reliability manager configured to: determine a change in a memory health state of the remote storage device based on the memory health data; initiate a second data recovery configuration in response to the change in the memory health state, wherein the storage system is further configured to store redundant data from the remote storage device using the second data recovery configuration; determine at least one initial memory health value for the remote storage device; and determine the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration.

11. The system of claim 10, wherein: the remote storage device is a non-volatile storage device located at a site remote from the storage system; and the storage system is further configured to periodically store differences between a current set of data stored on the remote storage device and a full copy of data stored on the remote storage device at an earlier time.

12. The system of claim 10, wherein: the reliability manager is further configured to: determine a service level for at least one system resource of the storage system; determine an allocation of the at least one system resource for storing redundant data in the storage system based on the service level, wherein the first data recovery configuration is further based on the allocation of the at least one system resource; and determine the second data recovery configuration based on the allocation of the at least one system resource and a change in the memory health state.

13. The system of claim 10, wherein: the storage system is further configured to: encode the set of redundant data in a first plurality of encoded data symbols according to a first parity level in response to the first data recovery configuration; and encode redundant data in a second plurality of encoded data symbols according to a second parity level in response to the second data recovery configuration; and the second parity level is adapted for a different error rate for recovering data compared to the first parity level.

14. The system of claim 10: further comprising: a backup interface configured to receive backup data from the remote storage device according to a periodic backup schedule; wherein: the memory health monitor is further configured to receive memory health data from the remote storage device in conjunction with the backup interface receiving backup data according to the periodic backup schedule.

15. The system of claim 10, wherein: the memory health monitor is further configured to: determine a plurality of physical storage units in the remote storage device; store reference values associating data stored in the storage system with the plurality of physical storage units storing corresponding user data in the remote storage device; and receive at least one memory health value for each physical storage unit of the plurality of physical storage units; the reliability manager is further configured to: determine that at least one memory health value for a first physical storage unit of the plurality of physical storage units satisfies a reduced reliability condition; and determining that at least one memory health value of a second physical storage unit of the plurality of physical storage units does not satisfy the reduced reliability condition; the storage system is further configured to: store redundant data associated with the first physical storage unit using the second data recovery configuration in response to determining the reduced reliability condition; and redundant data associated with the second physical storage unit remains stored using the first data recovery configuration.

16. The system of claim 10, wherein: the reliability manager is further configured to: determine at least one reduced reliability threshold; and evaluate the memory health data against the at least one reduced reliability threshold.

17. The system of claim 10, wherein: the reliability manager is further configured to: access historical memory health data of a remote storage device population of a remote storage device type associated with the remote storage device; determine a data reliability model of the remote storage device type based on the historical memory health data; determine at least one reduced reliability threshold based on the data reliability model; and evaluate the memory health data against the at least one reduced reliability threshold.

18. A system for storage, comprising: a storage system configured to store a redundant data set from a remote storage device using a first data recovery configuration; means for receiving memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device; means for determining a periodic backup configuration of the remote storage device; means for determining a change in a memory health state of the remote storage device based on the memory health data; means for initiating a second data recovery configuration in response to the change in the memory health state, wherein the storage system is further configured to store redundant data from the remote storage device using the second data recovery configuration; means for determining at least one initial memory health value of the remote storage device; and means for determining the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration.

19. A computer-implemented method, comprising: storing a redundant data set from a remote storage device to a distributed storage system using a first data recovery configuration, wherein the distributed storage system comprises an array of non-volatile storage devices configured to store the redundant data set from the remote storage device; receiving memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device; determining a change in the memory health state of the non-transitory media of the remote storage device based on the memory health data; and reallocate the redundant data set in the distributed storage system using a second data recovery configuration in response to the change in the memory health state.

20. The computer-implemented method of claim 19, wherein: the remote storage device is a non-volatile memory device located at a site remote from the distributed storage system; and reallocate the redundant data set in the distributed storage system includes periodically backing up differences between a current data set stored on the remote storage device and a full copy of data stored on the remote storage device at an earlier time.

21. The computer-implemented method of claim 19, further comprising: determining a periodic backup configuration for the remote storage device; determining at least one initial memory health value for the remote storage device; and determining the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration.

22. The computer-implemented method of claim 21, further comprising: determining a service level for at least one system resource of the distributed storage system; determining an allocation of the at least one system resource for storing the redundant data set in the distributed storage system based on the service level, wherein determining the first data recovery configuration is further based on the allocation of the at least one system resource; and determining the second data recovery configuration based on the allocation of the at least one system resource and the change in the memory health state.

23. The computer-implemented method of claim 19, wherein: storing the redundant data set in the distributed storage system using the first data recovery configuration includes encoding the redundant data set in a first plurality of encoded data symbols according to a first parity level; reallocate the redundant data set in the distributed storage system using the second data recovery configuration includes encoding at least a portion of the redundant data set in a second plurality of encoded data symbols according to a second parity level; and the second parity level is adapted for a different error rate for recovering the portion of the redundant data set than the first parity level.

24. The computer-implemented method of claim 19, wherein the memory health data includes at least one memory health value selected from: a bit error rate value; a write / erase cycle value; a program cycle counter value; an erase cycle counter value; a leak detection measurement value; an unstable program disturb value; a bad block value; or a voltage margin value.

25. The computer-implemented method of claim 19, further comprising: receiving the redundant data set from the remote storage device according to a periodic backup schedule, wherein receiving memory health data from the remote storage device is performed in conjunction with receiving the redundant data set according to the periodic backup schedule.

26. The computer-implemented method of claim 19: further comprising: determining a plurality of physical storage units in the remote storage device; and ​ storing a reference value that associates the set of redundant data stored in the distributed storage system with the plurality of physical storage units in the remote storage device that store corresponding data; wherein: receiving memory health data from the remote storage device includes receiving at least one memory health value for each physical storage unit of the plurality of physical storage units; determining the change in the memory health state includes: determining that at least one memory health value for a first physical storage unit of the plurality of physical storage units satisfies a reduced reliability condition; and determining that at least one memory health value for a second physical storage unit of the plurality of physical storage units does not satisfy the reduced reliability condition; re-allocating the set of redundant data in the distributed storage system using the second data recovery configuration includes storing data associated with the first physical storage unit using the second data recovery configuration in response to determining the reduced reliability condition; and data associated with the second physical storage unit remains stored using the first data recovery configuration.

27. The computer-implemented method of claim 19, wherein determining the change in the memory health state includes: determining at least one reduced reliability threshold; and evaluating the memory health data against the at least one reduced reliability threshold.

28. The computer-implemented method of claim 19, further comprising: collecting historical memory health data for a population of remote storage devices of a remote storage device type associated with the remote storage device; determining a data reliability model for the remote storage device type based on the collected historical memory health data; and determining at least one reduced reliability threshold based on the data reliability model, wherein determining the change in the memory health state includes evaluating the memory health data against the at least one reduced reliability threshold.

29. A system for storage, comprising: a storage system configured to store a set of redundant data from a remote storage device using a first data recovery configuration, wherein the storage system includes an array of non-volatile storage devices configured to store the set of redundant data from the remote storage device; a memory health monitor configured to receive memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device; and a reliability manager configured to: determine a change in a memory health state of the remote storage device based on the memory health data; and initiate a second data recovery configuration in response to the change in the memory health state, wherein the storage system is further configured to store a set of redundant data from the remote storage device using the second data recovery configuration.

30. The system of claim 29, wherein: ​ The remote storage device is a non-volatile memory device located at a site remote from the storage system; and The storage system is further configured to periodically store differences between a current data set stored on the remote storage device and a full copy of data stored on the remote storage device at an earlier time.

31. The system of claim 29: further comprising: a backup interface configured to determine a periodic backup configuration of the remote storage device; wherein: the reliability manager is further configured to: determine at least one initial memory health value of the remote storage device; and determine the first data recovery configuration based on the at least one initial memory health value and the periodic backup configuration.

32. The system of claim 31, wherein: the reliability manager is further configured to: determine a service level of at least one system resource of the storage system; determine an allocation of the at least one system resource for storing redundant data in the storage system based on the service level, wherein the first data recovery configuration is further based on the allocation of the at least one system resource; and determine the second data recovery configuration based on the allocation of the at least one system resource and the change in the memory health state.

33. The system of claim 29, wherein: the storage system is further configured to: encode the set of redundant data in a first plurality of encoded data symbols according to a first parity level in response to the first data recovery configuration; and encode redundant data in a second plurality of encoded data symbols according to a second parity level in response to the second data recovery configuration; and the second parity level is configured to accommodate a different error rate for recovering data as compared to the first parity level.

34. The system of claim 29: further comprising: a backup interface configured to receive backup data from the remote storage device according to a periodic backup schedule; wherein: the memory health monitor is further configured to receive memory health data from the remote storage device in conjunction with the backup interface receiving backup data according to the periodic backup schedule.

35. The system of claim 29, wherein: the memory health monitor is further configured to: determine a plurality of physical storage units in the remote storage device; store reference values associating data stored in the storage system with the plurality of physical storage units storing corresponding user data in the remote storage device; and receive at least one memory health value for each physical storage unit of the plurality of physical storage units; the reliability manager is further configured to: determine that at least one memory health value of a first physical storage unit of the plurality of physical storage units satisfies a reduced reliability condition; and determine that at least one memory health value of a second physical storage unit of the plurality of physical storage units does not satisfy the reduced reliability condition; the storage system is further configured to: ​ in response to determining the reduced reliability condition, storing redundant data associated with the first physical storage unit using the second data recovery configuration; and redundant data associated with the second physical storage unit remains stored using the first data recovery configuration.

36. The system of claim 29, wherein: the reliability manager is further configured to: determine at least one reduced reliability threshold; and evaluate the memory health data against the at least one reduced reliability threshold.

37. The system of claim 29, wherein: the reliability manager is further configured to: access historical memory health data of a population of remote storage devices of a remote storage device type associated with the remote storage device; determine a data reliability model of the remote storage device type based on the historical memory health data; determine at least one reduced reliability threshold based on the data reliability model; and evaluate the memory health data against the at least one reduced reliability threshold.

38. A system for storage, comprising: a storage system configured to store a redundant data set from a remote storage device using a first data recovery configuration, wherein the storage system comprises an array of non-volatile storage devices configured to store the redundant data set from the remote storage device; means for receiving memory health data associated with the remote storage device, wherein the memory health data corresponds to a memory health state of a non-transitory media of the remote storage device; means for determining a change in the memory health state of the remote storage device based on the memory health data; and means for initiating a second data recovery configuration in response to the change in the memory health state, wherein the storage system is further configured to store redundant data from the remote storage device using the second data recovery configuration. ​

Citation Information

Patent Citations

  • Recovery from errors in a redundant array of disk drives

    US5278838A

  • Space-optimized backup repository grooming

    US7415585B1