Storage system and method for migrating data in a storage system
By defining a migration strategy within the storage system, the problem of data migration in read-only mode is solved by rapidly or gradually migrating data, ensuring normal system operation and improving reliability and throughput.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2026-03-27
AI Technical Summary
In multi-disk storage systems, when storage media capable of only limited erase cycles enter read-only mode, existing technologies struggle to effectively migrate data to writable drives, causing the system to malfunction.
By defining a strategy for the storage system, data can be migrated quickly or gradually from read-only extents to alternative logical or physical drives, including data copying and migration when read or write requests are received.
It enables data migration in read-only mode of the storage system, ensuring normal system operation and reducing the risk of data loss, thereby improving system reliability and throughput.
Smart Images

Figure CN113805794B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Aspects of the present disclosure relate to methods for operating migrating data in a storage system comprising drives capable of entering a read-only mode and non-transitory computer media containing instructions for operating migrating data in a storage system comprising drives capable of entering a read-only mode. BACKGROUND
[0002] The following background is provided only for technical information and is not an admission that the subject matter disclosed before the present disclosure is prior art.
[0003] Some storage devices can contain storage media (e.g., flash memory) that is only capable of a limited number of erase cycles before failing. When such storage media approaches the end of its life, the storage device can be switched to a read-only mode. This prevents further damage to the underlying storage media and enables the data held thereon to still be safely accessed.
[0004] However, this presents a problem in storage systems where the (now) read-only drive is part of a multi-disk storage system that expects or is typically operated with component drives that are capable of accepting writes, such as the read-only drive. Thus, there is a need to provide for a controlled migration of data from the read-only drive to a writeable drive and eventually replace the read-only drive with a writeable drive in the storage system. SUMMARY
[0005] Aspects of embodiments of the inventive concept relate to methods for migrating data in a storage system comprising a plurality of active extents and one or more replacement logical or physical drives. The method can comprise determining that an active extent of the storage system has entered a read-only mode; determining whether the storage system should fast migrate data or should gradually migrate data based on a storage policy; then based on determining that the storage system should fast migrate data, copying data from the read-only extent to the replacement logical or physical drive as fast as possible; or alternatively based on determining that the storage system should gradually migrate data, copying data units from the read-only drive to the replacement logical or physical drive upon receiving a read request for the data units in the extent and writing the data units to the replacement logical or physical drive upon receiving a write request for the data units in the read-only extent.
[0006] Further aspects of embodiments of the present disclosure relate to a non-transitory computer readable medium having instructions thereon to cause a processor in a computer storage system comprising a plurality of active extents and one or more replacement logical or physical drives to perform a process. The process can be as follows: determining that an active extent of the storage system has entered a read-only mode; determining, based on a storage policy, whether the storage system should fast migrate data or should gradually migrate data; then based on determining that the storage system should fast migrate data, copying data from the read-only extent to the replacement logical or physical drive as quickly as possible; or alternatively based on determining that the storage system should gradually migrate data, copying a data unit from the read-only drive to the replacement logical or physical drive upon receiving a read request for the data unit in the extent, and writing the data unit to the replacement logical or physical drive upon receiving a write request for the data unit in the read-only extent. BRIEF DESCRIPTION OF DRAWINGS
[0007] These and other features and aspects of the present disclosure will be appreciated and understood of those skilled in the art in view of the following detailed description and drawings.
[0008] Figure 1 is a schematic block diagram of an information processing system that can include an apparatus formed in accordance with some example embodiments of the present disclosure.
[0009] Figure 2A and Figure 2B shows a schematic diagram of a storage system in accordance with some example embodiments of the present disclosure.
[0010] Figure 3A and Figure 3B shows a schematic diagram of a storage system undergoing migration in accordance with some example embodiments of the present disclosure.
[0011] Figure 4 is a flowchart of a method for determining a migration technique in accordance with some example embodiments of the present disclosure.
[0012] Figure 5 is a flowchart of a method for processing a migration of a read request in accordance with some example embodiments of the present disclosure.
[0013] Figure 6 is a flowchart of a method for processing a migration of a write request in accordance with some example embodiments of the present disclosure.
[0014] Figure 7 is a flowchart of a method for fast migration in accordance with some example embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] Various example embodiments will be described hereinafter more fully with reference to the accompanying drawings. The disclosure may, however, be embodied in many different forms and should not be construed as limited to the example embodiments set forth herein. Rather, these example embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. In the drawings, the sizes and relative sizes of layers and regions can be exaggerated for clarity.
[0016] It will be understood that when an element or layer is referred to as being "on" another element or layer, "connected to" or "coupled to" another element or layer, it can be directly on, connected or coupled to the other element or layer, or intervening elements or layers can be present. When an element or layer is referred to as being "directly on," "directly connected to," or "directly coupled to" another element or layer, there are no intervening elements or layers present. Like reference numerals refer to like elements throughout. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0017] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of the present disclosure.
[0018] Spatially relative terms (such as "beneath", "below", "lower", "above", "upper", and the like) can be used herein for ease of description to describe one element's or feature's relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device is turned over, elements described as "below" or "beneath" other elements or features would then be oriented "above" the other elements or features. Thus, the term "below" can encompass both an orientation of above and below. The device can be otherwise oriented (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly.
[0019] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting of the present subject matter. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0020] Example embodiments are described herein with reference to cross-sectional illustrations that are schematic illustrations of idealized embodiments (and intermediate structures) of the present subject matter. As such, variations from the shapes of the illustrations as a result, for example, of manufacturing techniques and / or tolerances, are to be expected. Thus, examples embodiments should not be construed as limited to the particular shapes of regions illustrated herein but are to include deviations in shapes that result, for example, from manufacturing. For example, an implanted region illustrated as a rectangle will, typically, have rounded or curved features and / or a gradient of implant concentration at its edges rather than a binary change between implanted and non-implanted regions. Similarly, a buried region formed by implantation can result in some implantation in regions between the buried region and a surface through which implantation occurs. Thus, the regions illustrated in the figures are schematic and their shapes are not intended to illustrate the precise shape of a region of a device and are not intended to limit the scope of the present subject matter.
[0021] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this present subject matter belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0022] Example embodiments are explained in detail below with reference to the accompanying drawings.
[0023] Figure 1 is a schematic block diagram of an information processing system 100 that can include a semiconductor device formed in accordance with the principles of the disclosed subject matter.
[0024] Referring to Figure 1 The information processing system (or, system) 100 can include one or more devices constructed in accordance with the principles of the disclosed subject matter. In one or more other embodiments, the information processing system 100 can employ or perform one or more techniques in accordance with the principles of the disclosed subject matter.
[0025] In various embodiments, information handling system 100 can include a computing device such as a laptop computer, a desktop computer, a workstation, a server, a blade server, a personal digital assistant, a smart phone, a tablet computer, and other appropriate computer or virtual machine or virtual computing device thereof. In various embodiments, information handling system 100 can be used by a user.
[0026] Information handling system 100 according to the disclosed subject matter can also include a central processing unit (CPU), logic, and / or processor 110. In some embodiments, processor 110 can include one or more functional unit blocks (FUBs) or combinational logic blocks (CLBs) 115. In such embodiments, combinational logic blocks can include various Boolean logic operation (e.g., NAND, NOR, NOT, XOR) devices, stabilization logic devices (e.g., flip-flops, latches), other logic devices, or combinations thereof. These combinational logic operation devices can be configured in simple or complex ways to process input signals to achieve a desired result. It should be understood that while some illustrative examples of synchronous combinational logic operation are described herein, the disclosed subject matter is not so limited and can include asynchronous operation, or a mix of synchronous and asynchronous operation. In one embodiment, combinational logic blocks can include a plurality of complementary metal-oxide-semiconductor (CMOS) transistors. In various embodiments, while these CMOS transistors can be arranged to perform gates of logic operations, it should be understood that other technologies can be used and are within the scope of the disclosed subject matter.
[0027] Information handling system 100 according to the disclosed subject matter can also include volatile memory 120 (e.g., random access memory (RAM)). Information handling system 100 according to the disclosed subject matter can also include non-volatile memory 130 (e.g., a hard disk drive, an optical memory, NAND or flash memory, and / or other solid state memory). In some embodiments, volatile memory 120, non-volatile memory 130, or combinations or portions thereof can be referred to as a "storage medium." In various embodiments, volatile memory 120 and / or non-volatile memory 130 can be configured to store data in semi-permanent or substantially permanent form.
[0028] In various embodiments, the information handling system 100 can include one or more network interfaces 140 configured to allow the information handling system 100 to become part of and communicate via a communication network via wired and / or wireless and / or cellular protocols. Examples of wireless protocols can include, but are not limited to, Institute of Electrical and Electronics Engineers (IEEE) 802.11 g, IEEE 802.11η. Examples of cellular protocols can include, but are not limited to, IEEE 802.16m (also known as, WiMax (Wireless Interoperability for Microwave Access), Long Term Evolution (LTE), Enhanced Data rates for GSM (Global System for Mobile Communications) Evolution (EDGE), Evolved High Speed Packet Access (HSPA+). Examples of wired protocols can include, but are not limited to, IEEE 802.3 (also known as, Ethernet), Fibre Channel, Powerline Communication (e.g., HomePlug, IEEE 1901). It should be appreciated that the above are merely illustrative examples to which the disclosed subject matter is not limited. By connecting to a network via the network interface 140, the information handling system 100 can access other resources whether as standalone network resources or as components of externally attached systems (e.g., external volatile memory, non-volatile memory, processors / logic, and software).
[0029] The information handling system 100 according to the disclosed subject matter can also include a user interface unit 150 (e.g., a display adapter, a tactile interface, and / or a human- machine interface device). In various embodiments, the user interface unit 150 can be configured to receive input from a user and / or to provide output to the user. Other kinds of devices can be used to provide interaction with a user as well, such as, for example, external devices 160, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, and / or tactile feedback), and input from the user can be received in any form, including acoustic, speech, and / or tactile input.
[0030] In various embodiments, the information handling system 100 can include one or more other devices or hardware components (e.g., other hardware devices) 160 (e.g., a display or monitor, a keyboard, a mouse, a camera, a fingerprint reader, and / or a video processor). It should be appreciated that the above are merely illustrative examples to which the disclosed subject matter is not limited.
[0031] The information handling system 100 according to the disclosed subject matter can also include one or more system buses 105. In such embodiments, the system bus 105 can be configured to communicatively connect the processor 110, the volatile memory 120, the non-volatile memory 130, the network interface 140, the user interface unit 150, and the one or more hardware components 160. Data processed by the processor 110 or externally input from the non-volatile memory 130 can be stored in the non-volatile memory 130 or the volatile memory 120.
[0032] In various embodiments, information handling system 100 can include or execute one or more software components (e.g., software) 170. In some embodiments, software components 170 can include an operating system (OS) and / or an application. In some embodiments, the OS can be configured to provide one or more services to applications, and to manage or act as an intermediary between various hardware components of information handling system 100 (e.g., processor 110 and / or network interface 140) and applications. In such embodiments, information handling system 100 can include one or more native applications, which can be installed locally (e.g., within non-volatile memory 130) and configured to be executed directly by processor 110 and to interact directly with the OS. In such embodiments, native applications can include pre-compiled machine executable code. In some embodiments, native applications can include a script interpreter (e.g., C shell (csh), AppleScript, AutoHotkey) and / or a virtual execution machine (VM) (e.g., Java Virtual Machine, Microsoft Common Language Runtime) configured to translate source code or object code into executable code that is then executed by processor 110.
[0033] As discussed herein, an apparatus can include logic configured to perform various tasks. The logic can be implemented as hardware, software, or a combination thereof. When the logic comprises hardware, the hardware can be in the form of an application-specific circuit arrangement (e.g., ASIC), an array of programmable gate logic and memory (e.g., FPGA), or specially programmed general-purpose logic (e.g., CPU and GPU). When the logic comprises software, the software can be configured to operate specialized circuitry, or to program an array of circuitry, memory, or to operate a general-purpose processor. Logic implemented in software can be stored on any available storage media, such as DRAM, flash memory, EEPROM, resistive memory, etc.
[0034] Turning to Figure 2A and Figure 2B , an example storage system 200 is shown, in accordance with some embodiments. The storage system can be used as non-volatile memory 130 of Figure 1 , and / or it can be accessible to system 100 through network interface 140. The storage system can present a virtual disk 205 as a logical construct to a host system (not shown), such as system 100 of Figure 1 . The virtual disk can store data units. Such data units can be, for example, data blocks, shown as data blocks 210a-n in the illustrated embodiment. In some other embodiments, the data units can represent other basic storage structures, such as Key-Value values, object store objects, or blobs, etc.
[0035] In the illustrated embodiment, the logical construct of virtual disk 205 is provided from a set of two or more physical drives, each physical drive comprising an extent (such as 215a-n) of the storage system. In some instances, a single physical drive can hold one extent or many extents, however, in the following examples, drives and extents are used in a 1-to-1 mapping. Thus, as used herein, "extent" and "physical drive" are considered analogous terms, although in some alternative embodiments, a drive can comprise multiple extents. However, one skilled in the art will readily appreciate how the concepts discussed herein can be applied to the case where a single physical drive holds multiple extents. The number of storage drives / extents 215a-n is a matter of design choice in any given embodiment, according to the principles discussed below.
[0036] Data presented by virtual disk 205 is stored on and retrieved from one or more of extents 215a-n according to a mapping function 220. The nature of the mapping function, as well as the number and nature of extents 215a-n, depends on the nature of storage system 200. For example, in an embodiment where storage system 200 is configured as a RAID 0 system, all data units are divided into stripes 225a-n by the mapping function. Each stripe 225a-n has a length, which defines the number of extents 215a-n across which data is held. The portion of a stripe that lies on an extent can be referred to as a "strip." Each stripe (and strip) also has a depth, which determines the number of data units (or size of data) stored in each extent 215a-n in a predetermined (or pre-computable) order. Each extent 215a-n can hold a value of the depth's worth of data units in a particular stripe 225a-n. Thus, assuming data units are fixed-size blocks with known block numbers / logical addresses, and both the stripe length and depth are known, the location of a data unit in a set of extents 215a-n can be easily computed mathematically and / or stored in a mapping table. For example, a strip can be identified by an ordered pair of (stripe#, extent#). A block can be identified by an ordered triple of (stripe#, extent#, depth-in-stripe).
[0037] In another example embodiment, the system can be configured as a RAID 1 system. The mapping function 220 is similar in many respects to the mapping function of the above RAID 0 system, except that half of the extents 215a-n are used as direct read / write storage locations, while the other half is reserved for mirroring the contents of the first half in order to provide data redundancy.
[0038] In another example embodiment, the system can be configured as a parity-based RAID system (such as RAID 4, RAID 5, or RAID 6, etc.). Such a system can utilize one or more dedicated parity zones 230 (e.g., RAID 4), or the parity information can be algorithmically distributed among the zones 215a-n (e.g., RAID 5, 6). In a parity-based system, M > 2 zones 215a-n are used to store data, and N > 1 zones 215a-n are used to store parity computation information. If any number of zones 215a-n up to N are lost, the data can be reconstructed from the combination of the remaining data / parity data on the surviving zones 215a-n.
[0039] In embodiments using a dedicated parity system, the mapping function 220 will work very similarly to in the RAID 0 or RAID 1 above, except that one or more dedicated parity zones 230 (which can be on dedicated parity disks) will not receive direct reads or writes - instead, the parity computation 235 will be received from the data across the other zones 215a-n, which can be computed at the stripe 225a-n level or at the block level (e.g., the blocks of the first depth level in each non-parity zone in a given stripe are used to compute the parity block in the blocks of the first depth level in a given stripe in the parity disk, similarly for the second depth level, etc.).
[0040] In embodiments utilizing a distributed parity system (not shown), the parity can still be computed at the block level or at the stripe 225a-n level. In a stripe 225a-n level parity system, for each stripe 225a-n recorded, the mapping function selects N zones 215a-n to use as parity zones and M zones to use as data zones, and the data and parity are written accordingly. The selection of zones 215a-n to use as parity zones changes algorithmically (such as via a round-robin approach). In a block level parity system, each depth level of a stripe 215a-n has N zones selected as parity zones and M zones selected as data zones. The data is written one depth level at a time, and the selection of parity zones and data zones rotates programmatically one depth level at a time. When the selection of parity locations is made algorithmically in both stripe 225a-n level parity and block level parity, the mapping function 220 can mathematically determine the location of each data block.
[0041] In other example embodiments, the storage system 200 need not utilize a traditional RAID configuration. For example, the extents 215a-n can simply be copies of each other, or the data can be stored as key-values with parity.
[0042] In some embodiments, the drives providing the extents 215a-n need not be physically co-located, they can be spread across servers, chassis, and data center locations.
[0043] Turning to Figure 3A and 3B , an embodiment of an example storage system 200' is shown. The storage system 200' is largely similar to the system 200 of Figure 2A and 2B , so like components are labeled with like numbers; however, the system 200' has detected that the physical drive housing extent 2 (in this example, drive / extent 215c, but it could be any drive / extent 215 that has finally entered read-only mode) has entered read-only mode. In order for the storage system 200' to continue to operate, the data stored on extent 2 (drive 215c) must be transferred to a replacement physical drive 305 via a process 310, which will be described further below. The replacement physical drive 305 can be part of the system 200', but it is not considered by the mapping function 220 to be a drive for storing or retrieving information until drive 215c enters read-only mode.
[0044] The replacement physical drive 305 can be located anywhere. More specifically, in some examples, the replacement physical drive can be housed in a chassis with the physical drives of the extents 215a-n, in a server rack with other physical drives of the extents 215a-n, and / or attached to a network in communication with other physical drives of the extents 215a-n. In some embodiments, the physical drives of the extents 215a-n are physically dispersed (e.g., in substantially different geographic locations). In such embodiments, the replacement physical drive 305 can be selected from a location as close as possible to the read-only extent 215c of the extents 215a-n in order to gain a bandwidth / latency advantage during data transfer.
[0045] The decision to convert a physical drive housing an extent to read-only mode can be made by the particular individual physical drive, or by the storage system 200' as a whole. The decision can be made based on various factors that can indicate that a particular physical drive is approaching the end of its life. Such factors can include: reaching a threshold of raw bit error rate on its storage media, reaching a threshold number of erase cycles on its storage media, or the physical drive itself reaching a threshold age, etc.
[0046] Turning toFigure 4 Method 400 for determining whether and how data should be migrated from a read-only zone 215c to an alternate physical drive 305 that houses an alternate zone is shown. At process 405, the system detects that a drive (e.g., 215c) that houses a zone has entered read-only mode. At process 410, the policy is consulted to determine whether the migration should be "fast" (process 420) (i.e., move all data as quickly as possible) or "gradual" (process 430) (as will be further explained below). The policy can be a set of rules and guidelines, storable on a tangible computer-readable medium, that provide instructions or rules for using which form of migration in various situations, however, each situation need not provide a binary distinction, i.e., these factors can be weighed against each other.
[0047] In some embodiments, such situations and rules include the following non-limiting examples:
[0048] 1) Comparison of the number of zones in read-only mode to the number of zone failures that can be allowed before data is lost (e.g., due to parity, mirroring, RAID configuration, etc.):
[0049] a. The number of zones in read-only mode meets or exceeds the number of parity that can be allowed will weigh towards fast migration.
[0050] b. The number of writable zones in operation meets or exceeds a mirroring threshold will weigh towards gradual migration.
[0051] c. The greater the number of drives in read-only mode, the more the policy weighs towards fast migration.
[0052] 2) The level of reliability quality desired by the user of system 200':
[0053] a. Stronger data reliability guarantees will weigh more towards fast migration.
[0054] 3) The level of throughput or bandwidth desired by the user of system 200':
[0055] a. If the end user desires high throughput, then gradual migration will be used more, as immediate transfer will require a large amount of internal bandwidth usage within system 200' due to the need to move a large amount of data internally.
[0056] 4) Known or predicted workload / behavior patterns compared to current activity:
[0057] a. For example, consider the discussion above for 3), if the system 200' processes heavy workloads periodically according to a known schedule, the system can bias toward fast data transfer when light workloads are expected, and bias toward stepwise migration when heavy workloads are expected. As the expected workload schedule progresses over time, the system can decide to switch from one strategy to another. For example, if the storage system 200' intends to process a daily backup at midnight, and the drives (and associated disk zones) enter read-only mode at 11 PM, the system 200' can decide to perform immediate migration until midnight, then switch to stepwise migration during the expected heavy traffic at midnight, etc.
[0058] When a decision is made to use fast migration according to the policy (operation 420), all data on the disk zone 215c can be transferred to the replacement physical disk 305 as quickly as possible (or, immediately). With respect to Figure 7 Some example embodiments of this migration are described.
[0059] When a decision is made to use stepwise migration (process 430) according to the policy, the process can wait for I / O (input / output) to occur. The nature of the operation that occurs depends on whether the I / O received is a read operation 435 or a write operation 440, described in Figure 5 and Figure 6 Some embodiments of read operations and write operations are described.
[0060] Turning to Figure 5 , an example method 500 of handling read operations is shown when the storage system 200' has decided that "stepwise" migration can be used. In process 505, a read operation can be received that includes an identifier (e.g., a block ID) of data being sought. In process 510, a determination can be made as to whether the read request is targeted at data on a read-only ("failing") disk zone 215c, or at data on a fully operational disk zone. This determination can be made using the ID of the data. When it is determined that the read request is targeted at data on a fully operational disk zone, in process 555, the data can be read from the fully operational disk zone and returned. The process can then return to process 430 of Figure 4 If it is determined in process 510 that the data can be located on a failing disk zone 215c, an emergency table can be consulted in process 515.
[0061] The emergency table of process 515 can be a table similar to one of the example embodiments shown below, which can track which entries of a failing disk zone 215c have been migrated to a replacement disk / disk zone 305. Note that in other embodiments, the data can be tracked by block ID / block location triplets, or the tracking can be done negatively (e.g., listing which data has not been migrated).
[0062] Example Emergency Table 1 : The table can track the stripes as they are migrated. Note that the first number records the zone, and the second number records the stripe ID. In the example below, the replacement zone is labeled as zone 5.
[0063] Stripe ID New Location (2,0) (5,0) (2,2) (5,2) ...(fill as needed)... ...(fill as needed)...
[0064] Example Emergency Table 2: The table can track the blocks as they are migrated. Note that, as described above, the location of a particular block can be expressed (at this logical level) as a triple of (zone, stripe, depth level). The table can be filled in as data is migrated.
[0065] Chunk Location New Location (2,0,1) (5,0,1) (2,2,0) (5,2,0) ...(fill as needed)... ...(fill as needed)...
[0066] Example Emergency Table 3: This table lists all of the stripes (or blocks) on the read-only zone 215c, and the bits are set based on whether the stripe (or block) has been migrated to the new zone. Note that the new address of the data is implied in the table, since the corresponding location to be migrated to the replacement zone 5 is known. A similar table for tracking blocks can be drawn up as in Table 2 above.
[0067] Stripe ID Migrate to Disk Zone 5? (2,0) 1 (2,1) 0 (2,2) 0 (2,3) 1 … …
[0068] Return Figure 5 At process 515, the emergency table can be consulted to determine whether the requested data has been migrated to the replacement zone. The ID of the data can be used to make such a determination. Based on a determination that the requested data has been migrated, at process 520, the data can be read from the location indicated on the replacement zone 305, and the data can be returned to the host. In some embodiments, the data can also be read from the read-only zone 215c if it is still in operation. The process can then return to process 430 of Figure 4 .
[0069] Based on a determination that the data has not yet been migrated, at process 525, the requested data can be read from the read-only zone 215c and returned to the host. At process 530, the data can also be copied / migrated to the corresponding location on the replacement zone in the replacement 305. At process 535, the ID of the copied data and its new location can be added to the emergency table in a manner that indicates that it has been migrated. At process 540, the ID of the data that has been migrated can be removed from the migration table, or the migration table otherwise updated to indicate that the data has been migrated. Process 540 can be optional.
[0070] In some embodiments, the data migration table can be similar to the emergency table, but works in reverse. More specifically, in some embodiments, while the emergency table tracks all data that has been migrated and to which, the migration table tracks all data that has not yet been migrated. These tables can in fact be combined into one table, example table 3, for such purposes. The migration table and the emergency table can be stored in the storage system 200'. In some embodiments, a copy of the migration table and / or the emergency table can be stored on multiple extents 215a-n.
[0071] At process 545, a determination is made as to whether there is still data that needs to be migrated. This can be done by reference to the migration table and / or the emergency table. When all data has been migrated, at process 550, the read-only extent 215c can be decommissioned, and the replacement disk and extent 305 can fully replace the read-only extent 215c in the system. This can be determined by checking the migration table or by comparing the contents of the emergency table to the list of contents of the read-only extent. If there is still data that needs to be migrated, the process can return to Figure 4 process 430.
[0072] Note that for processes 525-540, the migration system can work on different sized units of data. For example, if a block of data within stripe 1 of the read-only extent 215c is read, some embodiments of the system can track the entire stripe 1 and copy the entire stripe 1 to the replacement extent 305 (even though only the block was requested), which can use, for example, an emergency table and a migration table similar to table 1. Other embodiments can choose to only track and copy the requested block; this can use an emergency table and a migration table similar to table 2.
[0073] Turning to Figure 6 , an example embodiment of a method 600 for responding to a write request in a storage system 200' that utilizes stepwise migration is shown. At process 605, a data write request including an ID (e.g., a block ID) can be received. At process 610, a determination can be made as to whether the data is to be written to a read-only (decommissioned) extent or a fully operational extent. The ID of the data can be used to make such a determination. Based on the determination that the data is to be written to a fully operational extent, at process 615, the data can be written to the fully operational extent. The process can then return to Figure 4 process 430.
[0074] Based on determining that data is to be written to read-only zone 215c, process 620 then determines whether the data (e.g., the block and / or stripe of read-only zone 215c corresponding to the ID of the data) has already been migrated, e.g., by utilizing the emergency table and / or the migration table (e.g., determining whether the block ID and / or stripe ID of read-only zone 215c corresponding to the ID of the data is in the emergency table). The ID of the data can be used to make such a determination. Based on determining that the data has already been migrated to alternate zone 305, at process 625, the data is written directly to alternate zone 305. Process can then return to Figure 4 process 430. Based on determining that the data is not in the emergency table / migration table, at process 630, the data can be added to the emergency table (or, the block ID and / or stripe ID of read-only zone 215c corresponding to the ID of the data can be added to the emergency table). At process 635, the data can be written to alternate zone 305. At process 640, a determination is made as to whether the data is found in the migration table (or, a determination is made as to whether the block ID and / or stripe ID of read-only zone 215c corresponding to the ID of the data is found in the migration table). Based on determining that the data is in the migration table, at process 645, the data can be removed from the migration table (or otherwise marked as migrated), or the block ID and / or stripe ID of read-only zone 215c corresponding to the ID of the data can be removed from the migration table.
[0075] Based on determining that the data is not in the migration table or after the data is removed from the migration table, at process 650, a determination is made as to whether there is no data to migrate from read-only zone 215c. This can be determined by checking the migration table or by comparing the contents of the emergency table to the list of contents of the read-only zone. Based on determining that there is no data to migrate, at process 655, zone 305 can fully replace (and can decommission) read-only (failed) zone 215c within storage system 200'. Based on determining that there is data to migrate, process can return to Figure 4 process 430.
[0076] Similar to the process discussed with respect to Figure 5 , Figure 6 the process can operate using units of various data sizes. For example, if a block of data within stripe 1 of read-only zone 215c is written, some embodiments of the system can track and copy the entire stripe 1 to alternate zone 305 (even though only the block was requested), which can use, e.g., an emergency table and a migration table similar to Table 1. Other embodiments can choose to track and copy only the requested block, which can use an emergency table and a migration table similar to Table 2.
[0077] Turning to Figure 7 , a method 700 for fast migration of data from read-only zone 215c to alternate zone / zone 305 is shown. When fast migration is determined to be used at process 420, Figure 4 , the following can be used:Figure 7 The processing. To achieve rapid data migration, Figure 7 The execution of this method can be given high priority relative to other tasks in system 200', so that data is migrated reasonably and as quickly as possible while maintaining system availability. In process 705, the stripe to be migrated can be selected. Selection can be made by various means (such as numerical order). In process 710, the stripe can be copied from read-only extent 215c to alternative extent 305. In process 715, the stripe (via its ID) can be added to the emergency table. In process 720, the stripe can be removed from the migration table (or marked as migrated). In process 725, a determination is made as to whether the stripe to be migrated is the last stripe on the read-only extent to require migration. Based on the determination that there are still stripes to be migrated, the process can return to 705. Based on the determination that no stripes need to be migrated, process 730 can replace read-only extent 215c with alternative extent 305 and deactivate extent 215c. This can be determined by checking the migration table or by comparing the contents of the emergency table with the contents list of the read-only extent. Note that although this process discusses copying stripes, it is consistent with the above discussion on... Figure 5 to Figure 6 The processing is the same as described, and alternative embodiments may select to track and migrate data at the block level.
[0078] Note that in some embodiments, the extent-to-drive mapping can be many-to-one. That is, an extent can be a logical subdivision (logical drive) contained within a physical drive along with other extents. Therefore, an extent selected for migration (e.g., extent 215c) can be considered a logical drive (i.e., a logical subdivision of the physical drive). Similarly, the destination extent (new extent / drive 305) can also be a logical drive (i.e., a subdivision of the physical drive). Also note that when a physical drive enters read-only mode, this can result in multiple extents entering read-only mode. This should not change the basic operational concept described herein, which simply requires using correspondingly larger (or more) emergency and migration tables to migrate more extents.
[0079] In some embodiments, Figure 4 to Figure 7 This processing can be applied to storage systems 200' that use parity checking. For example, whenever data is written to a block, the parity check for that block is also updated with the write. Therefore, if the parity check for a specific set of blocks is on read-only extent 215c, any update to a block in the set of blocks on a fully operable extent can also cause a write in read-only extent 215c and trigger... Figure 6 The processing of parity data. As another example, if a read occurs on a fully operational extent, the system may choose to treat it as also a read of parity data in order to trigger... Figure 5 The processing (this could be because parity data is rarely read, making this mechanism "speed up" the migration of parity data through reads).
[0080] In some embodiments, Figure 4 to Figure 7 The processing of the data can be adapted for use with storage systems that do not use block-based storage. For example, each entry can include a portion of a key-value store or one or more values. The key will serve as the data ID, and the data can be located through a lookup table. Thus, for example, any read or write to a value of a key-value stored at least partially on an entry located on a read-only disk region will trigger Figure 5 to Figure 6 the operations of the data.
[0081] As will be appreciated by those skilled in the art having the benefit of the present disclosure, the methods disclosed herein present various component processes that can be removed or reordered and still remain within the scope of the invention of the present disclosure.
[0082] The method processes can be stored on a tangible computer readable medium and executed by one or more programmable processors to perform a function(s) by operating on input data and generating output. The method processes can also be executed by a special purpose logic circuitry (e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit)) and the devices can be implemented as a special purpose logic circuitry (e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit)).
[0083] In various embodiments, a computer readable medium can include instructions that, when executed, cause an apparatus to perform at least a portion of a method process. In some embodiments, a computer readable medium can be included in a magnetic medium, an optical medium, other medium, or a combination thereof (e.g., a CD-ROM, a hard disk drive, a read-only memory, a flash drive). In such embodiments, the computer readable medium can be a tangible and non-transitory article of manufacture.
[0084] While the principles of the disclosed subject matter have been described above in connection with example implementations, it is to be understood that this disclosure is not limited to those implementations. On the contrary, it is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the disclosed concepts. Therefore, it is to be understood that the foregoing embodiments are merely exemplary and are not intended to limit the scope of the disclosed concepts, which is defined solely by the appended claims and their equivalents. Accordingly, it is not intended that the disclosed concepts be limited, except as by the appended claims and their equivalents.
[0085] Method processing can be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Method processing can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0086] In various embodiments, a computer-readable medium can include instructions that, when executed, cause an apparatus to perform at least a portion of a method process. In some embodiments, a computer-readable medium can include in a magnetic medium, an optical medium, other medium, or combinations thereof (e.g., CD-ROM, hard disk drive, read-only memory, flash drive). In such embodiments, a computer-readable medium can be a tangible and non-transitory article of manufacture.
[0087] Embodiments of the inventive concept can extend to the following statements without limitation:
[0088] 1) A method for migrating data in a storage system comprising a plurality of active extents and one or more replacement logical or physical drives, the method comprising:
[0089] determining that an active extent of the storage system has entered a read-only mode;
[0090] determining, based on a storage policy, whether the storage system should fast migrate data or should gradually migrate data;
[0091] based on determining that the storage system should fast migrate data, copying data from the read-only extent to the replacement logical or physical drive as soon as possible; and
[0092] based on determining that the storage system should gradually migrate data, copying data units from the read-only drive to the replacement logical or physical drive upon receiving a read request for the data units in the extent, and writing data units to the replacement logical or physical drive upon receiving a write request for the data units in the read-only extent.
[0093] 2) The method of statement 1, wherein upon receiving a read request for a data unit, an entire stripe within the read-only extent containing the data unit is copied to the replacement logical or physical drive.
[0094] 3) The method of statement 1, wherein when a data unit is copied from the read-only extent to the replacement logical or physical drive due to a read operation, or when a data unit is written to the replacement logical or physical drive due to a write operation to the read-only extent, the data unit is recorded in a panic table when it has been migrated to the replacement logical or physical drive.
[0095] 4) The method of statement 3, additionally comprising: redirecting subsequent reads or writes to data units found on the emergency table to the alternate logical or physical drive.
[0096] 5) The method of statement 4, additionally comprising: replacing the read-only extent with the alternate logical or physical drive based on determining that all data units in the read-only extent have been migrated to the alternate logical or physical drive.
[0097] 6) The method of statement 5, wherein the step of determining that all data units in the read-only extent have been migrated comprises: comparing a list of migrated data units on the emergency table to a list of data units in the read-only extent, or consulting a migration table, or determining that the migration table indicates that all data units have been migrated.
[0098] 7) The method of statement 1, wherein the read-only extent enters read-only mode due to having reached a threshold wear level.
[0099] 8) The method of statement 7, wherein the threshold wear level is determined by one or more of: an original media error rate, a number of erase cycles for a portion of the storage media, and a product life threshold being exceeded.
[0100] 9) The method of statement 1, wherein the storage policy comprises one or more of: a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement.
[0101] 10) A non-transitory computer readable medium having instructions thereon that cause a processor in a computer storage system comprising a plurality of active extents and one or more alternate logical or physical drives to perform the following operations:
[0102] determining that an active extent of the storage system has entered read-only mode;
[0103] determining, based on a storage policy, whether the storage system should fast migrate data or should gradually migrate data;
[0104] based on determining that the storage system should fast migrate data, copying data from the read-only extent to the alternate logical or physical drive as quickly as possible; and
[0105] based on determining that the storage system should gradually migrate data, copying data units from the read-only drive to the alternate logical or physical drive upon receiving a read request for a data unit in the extent, and writing data units to the alternate logical or physical drive upon receiving a write request for a data unit in the read-only extent.
[0106] 11) The non-transitory computer readable medium of statement 10, wherein upon receiving a read request for a data unit, the entire stripe within the read-only extent containing the data unit is copied to the replacement logical or physical drive.
[0107] 12) The non-transitory computer readable medium of statement 10, wherein when a data unit is copied from the read-only extent to the replacement logical or physical drive due to a read operation, or when a data unit is written to the replacement logical or physical drive due to a write operation to the read-only extent, the data unit is recorded in the emergency table when it has been migrated to the replacement logical or physical drive.
[0108] 13) The non-transitory computer readable medium of statement 12, additionally comprising: redirecting subsequent reads or writes to data units found on the emergency table to the replacement logical or physical drive.
[0109] 14) The non-transitory computer readable medium of statement 13, additionally comprising: replacing the read-only extent with the replacement logical or physical drive as an active extent based on determining that all data units in the read-only extent have been migrated to the replacement logical or physical drive, or consulting the migration table.
[0110] 15) The non-transitory computer readable medium of statement 14, wherein the step of determining that all data units in the read-only extent have been migrated comprises: comparing a list of migrated data units on the emergency table to a list of data units in the read-only extent, or determining that the migration table indicates that all data units have been migrated.
[0111] 16) The non-transitory computer readable medium of statement 10, wherein the read-only extent enters read-only mode due to having reached a threshold wear level.
[0112] 17) The non-transitory computer readable medium of statement 16, wherein the threshold wear level is determined by one or more of: an original media bit error rate, a number of erase cycles for a portion of the storage media, and a product life threshold being exceeded.
[0113] 18) The non-transitory computer readable medium of statement 10, wherein the storage policy comprises one or more of: a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement.
[0114] 19) A storage system, comprising:
[0115] a plurality of active extents;
[0116] one or more replacement drives;
[0117] a processor; and
[0118] a memory connected to the processor, wherein the memory stores instructions that, when executed by the processor, cause the processor to:
[0119] determine that an active extent of a storage system has entered a read-only mode to become a read-only extent;
[0120] determine, based on a storage policy, whether the storage system should fast migrate data or should gradually migrate data;
[0121] based on a determination that the storage system should fast migrate data, copy data from the read-only extent to a replacement logical or physical drive as quickly as possible;
[0122] based on a determination that the storage system should gradually migrate data, copy data units from the read-only drive to a replacement logical or physical drive upon receiving a read request for the data units in the extent; and write data units to the replacement logical or physical drive upon receiving a write request for the data units in the read-only extent.
Claims
1. A method for migrating data in a storage system comprising a plurality of active extents and one or more replacement logical or physical drives, the method comprising: determining that an active extent of the storage system has entered a read-only mode to become a read-only extent; determining, based on a storage policy, whether the storage system should fast migrate data or should gradually migrate data; based on determining that the storage system should fast migrate data, immediately copying data from the read-only extent to the replacement logical or physical drive; and based on determining that the storage system should gradually migrate data, upon receiving a read request for a data unit in the read-only extent, copying the data unit from the read-only extent to the replacement logical or physical drive, and upon receiving a write request for the data unit in the read-only extent, writing the data unit to the replacement logical or physical drive.
2. The method of claim 1, wherein, Upon receiving a read request for a data unit, an entire stripe within the read-only extent containing the data unit is copied to the replacement logical or physical drive.
3. The method of claim 1, wherein, When a data unit is copied from the read-only extent to the replacement logical or physical drive due to a read operation, or when a data unit is written to the replacement logical or physical drive due to a write operation to the read-only extent, the data unit is recorded in a bailout table when it has been migrated to the replacement logical or physical drive.
4. The method of claim 3, additionally comprising: Subsequent reads or writes to data units found on the bailout table are redirected to the replacement logical or physical drive.
5. The method of claim 4, additionally comprising: Based on determining that all data units in the read-only extent have been migrated to the replacement logical or physical drive, replacing the read-only extent with the replacement logical or physical drive as an active extent.
6. The method of claim 5, wherein, The step of determining that all data units in the read-only extent have been migrated comprises comparing a list of migrated data units on the bailout table to a list of data units in the read-only extent, or determining that a migration table indicates that all data units have been migrated.
7. The method of any one of claims 1 to 6, wherein, The read-only extent enters the read-only mode due to having reached a threshold wear level.
8. The method of claim 7, wherein, The threshold wear level is determined by one or more of: an original media error rate, a number of erase cycles of a portion of the storage media, and a product life threshold being exceeded.
9. The method of any one of claims 1 to 6, wherein, The storage policy comprises one or more of: a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement.
10. A non-transitory computer readable medium having instructions thereon, the instructions causing a processor in a computer storage system comprising a plurality of active extents and one or more replacement logical or physical drives to perform the following operations: determining that an active extent of the storage system has entered a read-only mode to become a read-only extent; determining, based on a storage policy, whether the storage system should fast migrate data or should gradually migrate data; based on determining that the storage system should fast migrate data, immediately copying data from the read-only extent to the replacement logical or physical drive; and based on determining that the storage system should gradually migrate data, upon receiving a read request for a data unit in the read-only extent, copying the data unit from the read-only extent to the replacement logical or physical drive, and upon receiving a write request for the data unit in the read-only extent, writing the data unit to the replacement logical or physical drive.
11. The non-transitory computer-readable medium of claim 10, wherein, Upon receiving a read request for a data unit, the entire stripe within the read-only extent containing the data unit is copied to the replacement logical or physical drive.
12. The non-transitory computer-readable medium of claim 10, wherein, When a data unit is copied from the read-only extent to the replacement logical or physical drive due to a read operation, or when a data unit is written to the replacement logical or physical drive due to a write operation to the read-only extent, the data unit is recorded in the emergency table as having been migrated to the replacement logical or physical drive.
13. The non-transitory computer-readable medium of claim 12, wherein, The operations additionally include redirecting subsequent reads or writes to the data unit found on the emergency table to the replacement logical or physical drive.
14. The non-transitory computer-readable medium of claim 13, wherein, The operations additionally include replacing the read-only extent with the replacement logical or physical drive as an active extent based on determining that all data units in the read-only extent have been migrated to the replacement logical or physical drive.
15. The non-transitory computer-readable medium of claim 14, wherein, The step of determining that all data units in the read-only extent have been migrated includes comparing a list of migrated data units on the emergency table to a list of data units in the read-only extent, or determining that the migration table indicates that all data units have been migrated.
16. The non-transitory computer-readable medium of any one of claims 10 to 15, wherein, The read-only extent enters the read-only mode due to having reached a threshold wear level.
17. The non-transitory computer-readable medium of claim 16, wherein, The threshold wear level is determined by one or more of: an original media error rate, a number of erase cycles for a portion of the storage media, and a product life threshold being exceeded.
18. The non-transitory computer-readable medium of any one of claims 10 to 15, wherein, The storage policy includes one or more of: a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement.
19. A storage system comprising: a plurality of active extents; one or more replacement drives; a processor; and a memory connected to the processor, wherein the memory stores instructions that, when executed by the processor, cause the processor to: determine that an active extent of the storage system has entered a read-only mode to become a read-only extent; determine, based on a storage policy, whether the storage system should fast migrate data or should gradual migrate data; based on determining that the storage system should fast migrate data, immediately copy data from the read-only extent to a replacement logical or physical drive; based on determining that the storage system should gradual migrate data, copy a data unit from the read-only extent to a replacement logical or physical drive upon receiving a read request for the data unit in the read-only extent, and write the data unit to the replacement logical or physical drive upon receiving a write request for the data unit in the read-only extent.