A fast and gradual migration approach
A method for data migration in storage systems addresses the challenge of transitioning from read-only to writable drives by implementing a policy-based approach that allows for quick or gradual data transfer, ensuring seamless system operation.
Patent Information
- Application Number
- JP2021092857
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-29
- Filing Date
- 2021-06-02
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-06-02
AI Technical Summary
Storage systems with read-only drives face challenges in efficiently migrating data to writable drives, particularly in multi-disk configurations where read-only drives expect and operate normally with writable components, necessitating a controlled data migration approach.
A method for migrating data from read-only extents to replacement drives based on a storage policy, allowing for either quick or gradual migration depending on system conditions, involving data copying upon read or write requests.
Enables efficient and controlled data migration from read-only drives to writable drives, ensuring data accessibility and system continuity during the transition process.
Smart Images

Figure 0007718622000004 
Figure 0007718622000005 
Figure 0007718622000006
Abstract
Description
[Technical Field]
[0001] Embodiments of the present disclosure relate to a non-transitory computer medium containing methods and instructions for effecting data migration in a storage system that includes a drive that can enter a read-only mode. [Background technology]
[0002] The following Background Art is intended merely to provide information necessary to understand the context of the inventive ideas and concepts disclosed herein. Therefore, the Background Art section contains patentable material and should not be considered a disclosure of prior art.
[0003] Some storage devices may include storage media (e.g., flash memory) that can undergo a limited number of erase cycles before failing. When such storage media nears the end of its life, the storage device may be converted to read-only mode. This prevents further damage to the underlying storage media and ensures that the data stored thereon remains safely accessible.
[0004] However, this presents a problem in storage systems that are part of a multi-disk storage system, particularly where the read-only drive currently expects and operates normally with a component drive (such as a read-only drive) that can accept writes. Therefore, there is a need to provide for controlled migration of data from the read-only drive to a writable drive, and ultimately to replace the read-only drive with a writable drive in the storage system. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] U.S. Patent No. 09710317 [Patent Document 2] U.S. Patent No. 10,089,027 [Patent Document 3] U.S. Patent No. 10,459,652 [Patent Document 4] US Patent Application Publication No. 20200020398 [Non-patent literature]
[0006] [Non-Patent Document 1] SNIA, Advancing storage & information technology, "Common RAID Disk Data Format Specification", Common RAID Disk Data Format (DDF), SNIA Technical Position, Version 2.0 , Rev. 19, pgs. 1-126, March 27, 2009, https: / / www.snia.org / sites / default / files / SNIA_DDF_Technical_Position_v2.0.pdf. Summary of the Invention [Problem to be solved by the invention]
[0007] The present disclosure has been made in consideration of the above-mentioned prior art, and an object of the present disclosure is to provide a non-transitory computer medium including a method and instructions for operating data migration in a storage system including a drive that can enter a read-only mode. [Means for solving the problem]
[0008] An embodiment of the disclosed concepts relates to a method for migrating data in a storage system including multiple active extents and one or more replacement logical or physical drives, the method including: determining that an active extent of the storage system has entered a read-only mode; determining whether the storage system should migrate data quickly or incrementally based on a storage policy; copying data from the read-only extent to a replacement logical or physical drive as quickly as possible based on the determination that the data should be migrated quickly; and copying the data from the read-only drive to the replacement logical or physical drive upon receiving a read request for a data unit of the extent and writing the data unit to the replacement logical or physical drive upon receiving a write request for the data unit of the read-only extent based on the determination that the data should be migrated incrementally.
[0009] Another embodiment of the disclosed concepts relates to a non-transitory computer-readable medium having instructions for causing a processor of a computer storage system including a plurality of active extents and one or more replacement logical or physical drives to perform a process, the process being carried out as follows: the method includes determining that an active extent of the storage system has entered a read-only mode; determining, based on a storage policy, whether the storage system should migrate data quickly or gradually; copying data from the read-only extent to a replacement logical or physical drive as quickly as possible based on the determination that the data should be migrated quickly; and copying the data from the read-only drive to the replacement logical or physical drive when a read request for a data unit of the extent is received, and writing the data unit to the replacement logical or physical drive when a write request for the data unit of the read-only extent is received, based on the determination that the data should be migrated gradually. [Effects of the Invention]
[0010] According to the present invention, when data is requested from a read-only disk in a storage system during gradual migration, the data can be quickly or gradually migrated to a replacement disk to which the data is migrated. [Brief explanation of the drawings]
[0011] These and other features and aspects of the present invention will be understood with reference to the specification, claims, and accompanying drawings.
[0012] [Figure 1]FIG. 1 is a schematic block diagram of an information processing system that may include an apparatus formed in accordance with some example embodiments of the present disclosure. [Figure 2A] FIG. 1 is a schematic diagram of a storage system according to some example embodiments of the present disclosure. [Figure 2B] FIG. 1 is a schematic diagram of a storage system according to some example embodiments of the present disclosure. [Figure 3A] FIG. 1 is a schematic diagram of a storage system undergoing migration in accordance with some example embodiments of the present disclosure. [Figure 3B] FIG. 1 is a schematic diagram of a storage system undergoing migration in accordance with some example embodiments of the present disclosure. [Figure 4] 1 is a flowchart for a method for determining a migration technique according to some example embodiments of the present disclosure. [Figure 5] 1 is a flowchart for a method for handling migration of a read request according to some example embodiments of the present disclosure. [Figure 6] 1 is a flowchart for a method for handling migration of write requests according to some example embodiments of the present disclosure. [Figure 7] 1 is a flowchart for a method for fast migration according to some example embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Various exemplary embodiments are described in detail below with reference to the accompanying drawings, in which some exemplary embodiments are shown. However, the disclosed subject matter may be embodied in many different forms and should not be construed as being limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosed subject matter to those skilled in the art. In the drawings, the sizes and relative sizes of layers and regions may be exaggerated for clarity.
[0014] When an element or layer is described as being "on," "connected," or "coupled" to another element or layer, this means that it is directly on, connected to, or coupled to the other element or layer, or that there may be other elements or layers in between. When an element is described as being "directly on," "directly connected," or "directly coupled" to another element or layer, there are no other elements or layers in between. Like reference numerals refer to like elements throughout this specification. As used herein, the term "and / or" includes any and all combinations of one or more of the associated and listed items.
[0015] Although terms such as first, second, and third may be used herein to describe various elements, components, regions, layers, and / or sections, it should be understood that these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are used to distinguish only one element, component, region, layer, or section from another region, layer, or section. Thus, a first element, component, region, layer, or section described below could be referred to as a second component, region, layer, or section without departing from the teachings of the present disclosure.
[0016] Spatially relative terms such as "below," "below," "lower," "above," and "upper" are used herein for ease of description to describe the relationship of one element or feature to other elements or features as shown in the figures. The spatially relative terms are intended to include different orientations of the device when used or operated in addition to the orientation shown in the figures. For example, if the device in the figures were turned over, an element described as "below" or "below" the other element or feature would be "above" the other element or feature. Thus, the term "below" can include both an above or below orientation. The device's orientation can be otherwise oriented (rotated 90 degrees or in other directions), and the spatially relative and descriptive terms used herein should be interpreted accordingly.
[0017] The terminology used herein is merely for the purpose of describing particular example embodiments and is not intended to limit the presently disclosed subject matter. As used herein, the singular forms "a," "an," and "the" also include the plural unless the context clearly dictates otherwise. As used herein, the terms "comprise" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components and do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0018] The exemplary embodiments herein are described with reference to cross-sectional diagrams that are schematic illustrations of idealized exemplary embodiments (and intermediate structures). As such, variations from the illustrated configurations are expected as a result, for example, of manufacturing techniques and / or tolerances. Accordingly, the exemplary embodiments should not be construed as limited to the particular shapes of regions illustrated herein and include, for example, deviations in shape that are a result of manufacturing. For example, an implanted region illustrated as a rectangle would typically have circular or curved features and / or a gradient in implant concentration at its edges, rather than a binary change from implanted to non-implanted. Similarly, a buried region formed by implantation can result in any implantation in the region between the implanted region and the surface where the implantation occurs. Therefore, regions illustrated in schematic diagrams in nature and shape are not intended to represent the actual shape of a region of a device, nor are they intended to limit the scope of the presently disclosed subject matter.
[0019] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains. It should also be understood that terms, such as those defined in commonly used dictionaries, should be interpreted to have a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless expressly defined as such herein.
[0020] In the following, exemplary embodiments will be described in detail with reference to the accompanying drawings.
[0021] FIG. 1 is a schematic block diagram of an information handling system 100 that may include semiconductor devices formed in accordance with the principles of the present disclosure.
[0022] 1, information processing system 100 may include one or more devices configured according to the principles of the present disclosure. In one or more embodiments, information processing system 100 may employ or implement one or more technologies according to the principles of the present disclosure.
[0023] In various embodiments, information handling system 100 may include a computing device such as, for example, a laptop, a desktop, a workstation, a server, a blade server, a personal digital assistant (PDA), a smartphone, a tablet, and other suitable computers, virtual machines thereof, or virtual computing devices. In various embodiments, information handling system 100 may be used by a user.
[0024] An information processing device 100 consistent with the present disclosure may further include a central processing unit (CPU), logic, or processor 110. In some embodiments, the processor 110 may include one or more functional unit blocks (FUBs) or combinatorial logic blocks (CLBs) 115. In such embodiments, the combinatorial logic blocks may include various Boolean logic operations (e.g., NAND, NOR, NOT, XOR), stabilized logic devices (flip-flops, latches), and other logic devices, or combinations thereof. These combinatorial logic operations may be configured in simple or complex ways to process input signals to achieve desired results. While the illustrated examples of several synchronized combinatorial logic operations are described, it should be understood that the present disclosure includes, but is not limited to, asynchronous operations, or a combination of synchronous and asynchronous operations. In one embodiment, the combinatorial logic operations may include multiple complementary metal oxide semiconductor (CMOS) transistors. In various embodiments, while remaining within the scope of the present disclosure, other technologies may be used, and these CMOS transistors may be arranged within gates that perform the logic operations.
[0025] Information processing system 100 consistent with the present disclosure may further include volatile memory 120 (e.g., random access memory (RAM)). Information processing system 100 consistent with the present disclosure may further include non-volatile memory 130 (e.g., a hard drive, optical memory, NAND, flash memory, and / or other solid-state memory). In some embodiments, volatile memory 120, non-volatile memory 130, or a combination or portion thereof, is referred to as a “storage medium.” In various embodiments, volatile memory 120 and / or non-volatile memory 130 are configured to store data in a semi-persistent or substantially persistent form.
[0026] In various embodiments, information handling system 100 may include one or more network interfaces 140 that enable information handling system 100 to be a part of or be configured to communicate through a communication network via wired, wireless, and / or cellular protocols. Exemplary wireless protocols may include, but are not limited to, IEEE (Institute of Electrical and Electronics Engineers) 802.11g and IEEE 802.11n. Exemplary cellular protocols may include, but are not limited to, IEEE 802.16m (aka Wireless-ME (Metropolitan Area Network) Advanced), Long Term Evolution (LTE) Advanced, Enhanced Data rates for GSM (Global System for Mobile Communications) Evolution (EDGE), and Evolved High-Speed Packet Access (HSPA+). Examples of wired protocols may include, but are not limited to, IEEE 802.3 (aka Ethernet), Fibre Channel, and Power Line communication (e.g., HomePlug, IEEE 1901). It should be understood that the foregoing are merely some illustrative examples to which the present disclosure is not limited. As a result of being coupled to a network via network interface 140, information processing system 100 may access other resources, such as external volatile memory, non-volatile memory, processors / logic, and software, whether standalone network resources or components of additional external systems.
[0027] Information processing system 100 consistent with the teachings of this disclosure may further include a user interface unit 150 (e.g., a display adapter, a haptic interface, and / or a human interface device). In various embodiments, this user interface unit 150 may be configured to receive input from a user and / or provide output to a user. Other types of devices may also be used to provide interaction with a user. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, and / or tactile feedback. And, input from the user may be received in any form, including acoustic, speech, or tactile input.
[0028] In various embodiments, information processing system 100 may include one or more devices or hardware components 160 (e.g., a display or monitor, a keyboard, a mouse, a camera, a fingerprint reader, and / or a video processor). It should be understood that the foregoing are merely some illustrative examples to which the present disclosure is not limited.
[0029] Information processing system 100 consistent with the teachings of the present disclosure may further include one or more system buses 105. In such an embodiment, system bus 105 is configured to communicatively couple processor 110, volatile memory 120, non-volatile memory 130, network interface 140, user interface unit 150, and one or more hardware components 160. Data processed by processor 110 or data input from outside non-volatile memory 130 may be stored in non-volatile memory 130 or volatile memory 120.
[0030] In various embodiments, information handling system 100 may include or execute one or more software components 170. In some embodiments, software component 170 may include an operating system (OS) and / or applications. In some embodiments, the OS is configured to provide one or more services to the applications and manage or act as an intermediary between the applications and various hardware components of information handling system 100 (e.g., processor 110 and / or network interface 140). In such embodiments, information handling system 100 includes one or more native applications located locally (e.g., in non-volatile memory 130), which may be executed directly by processor 110 and which are configured to interact directly with the OS. In such embodiments, the native applications may include pre-compiled machine-executable code. In some embodiments, a native application may include a script interpreter (e.g., C shell (csh), AppleScript, AutoHotkey) and / or a virtual execution machine (VM) (e.g., the Java Virtual Machine, the Microsoft Common Language Runtime) configured to convert source or object code into executable code executed by the processor 110.
[0031] As described herein, an apparatus may include logic configured to perform various tasks. This logic may be embodied as hardware, software, or a combination thereof. When logic includes hardware, the hardware may be in the form of special-purpose circuit arrays (e.g., ASICS), programmable arrays of gate logic and memory (e.g., FPGAs), or programmable arrays of specially programmed general-purpose logic (e.g., CPUs and GPUs). When logic includes software, the software is configured to operate special-purpose circuitry, program arrays of circuitry and memory, or operate general-purpose processors. Logic embodied in software may be stored on any available storage medium, such as DRAM, flash, EEPROM, resistive memory, and / or the like.
[0032] 2A and 2B, an exemplary storage system 200 according to some embodiments is shown. The storage system may function as non-volatile memory 130 of FIG. 1 and / or may be accessible to system 100 via network interface 140. The storage system may present a virtual disk 205 as a logical construct to a host system (not shown), such as system 100 of FIG. 1. The virtual disk may store data units. Such data units may be, for example, data blocks, illustrated as data blocks 210a-n in the illustrated embodiment. In some other embodiments, the data units may represent other underlying storage constructs, such as key-value values, objects in an object store, or blobs (binary large objects), among others.
[0033] In the illustrated embodiment, the logical organization of virtual disk 205 is provided from a set of multiple physical drives, each containing a storage system extent such as 215a-n. While in some cases a single physical drive may have a single extent or many extents, in the following examples, drives and extents are used as having a one-to-one mapping. Thus, as used herein, "extent" and "physical drive" are treated as similar terms, although in some other embodiments, a drive may contain multiple extents. However, those skilled in the art will readily understand how the concepts described herein can be applied when a single physical drive has multiple extents. In any given embodiment, the number of storage drives / extents 215a-n is a matter of design choice based on principles discussed below.
[0034] Data provided by virtual disk 205 is stored in and retrieved from one or more of extents 215a-n based on a mapping function 220. The nature of the mapping function and the number and nature of extents 215a-n depend on the nature of storage system 200. For example, in an embodiment in which storage system 200 is configured as a RAID 0 system, all data units are divided into stripes 225a-n by the mapping function. Each stripe 225a-n has a length that defines the number of extents 215a-n in which the data is stored. The portion of a stripe located in an extent is referred to as a "strip." Each stripe (and strip) also has a depth that determines the number of data units (or size of data) stored in each extent 215 in a predetermined (or pre-calculated) order. Each extent 215 can receive a depth value of data units in a particular stripe 225a-n. Thus, assuming that a data unit is a fixed-size block with a known number of blocks / logical addresses and that both the length and depth of a stripe are known, the location of a set of extents 215a-n can be easily calculated mathematically and / or stored in a mapping table. A strip may be identified by an ordered pair, for example, (stripe#, extent 3). A block may be identified by an ordered triplet, (stripe#, extent#, depth-in-stripe).
[0035] In another exemplary embodiment, the system is configured as a RAID 1 system. The mapping function 220 would be similar in many respects to that of a RAID 0 system, except that half of the extents 215a-n are used as direct read / write storage locations and the other half exist to mirror the contents of the first half to provide data redundancy.
[0036] In another exemplary embodiment, the system is configured as a RAID system utilizing parity, such as RAID 4, 5, or 6 (among others). Such a system may utilize one or more dedicated parity extents 230 (e.g., RAID 4), or parity information may be algorithmically distributed among the extents 215 a-n (e.g., RAID 5, 6). In a parity-based system, M≧2 extents 215 a-n are used to store data, and N≧1 extents 215 a-n are used to store parity calculation information. If any number of extents 215 a-n less than or equal to N is lost, data is reconstructed on the remaining extents 215 a-n from the remaining data / parity data combination.
[0037] In embodiments utilizing a dedicated parity system, mapping function 220 would operate much the same as for RAID 0 or 1, except that there are one or more dedicated parity extents 230 (which may be on a dedicated parity disk) that do not receive direct reads or writes, but rather receive parity calculations 235 from data across other extents 215a-n. Parity can be calculated at the stripe 225a-n level or at the block level (e.g., in a given stripe, blocks in the first depth level of each non-parity extent are used to calculate parity blocks for blocks in the first depth level of a given strip of the parity disk, as well as blocks in the second depth level of each non-parity extent, etc.).
[0038] In embodiments utilizing a distributed parity system (not shown), parity is still calculated at the block or stripe 225a-n level. In a stripe 225a-n level parity system, for each recorded stripe 225a-n, a mapping function selects N extents 215a-n to serve as parity extents and M extents to serve as data extents, and data and parity are written accordingly. The selection of extents 215a-n to serve as parity extents is algorithmically varied, for example, via a round-robin approach. In a block-level parity system, each depth level of a stripe 215a-n has N extents selected as parity extents and M extents selected as data extents. Data is written one depth level at a time, and the selection of parity and data extents is programmatically rotated one depth level at a time. The mapping function 220 mathematically determines the location of each data block, as the selection of the parity location is algorithmically created for both stripes 225a-n and block-level parity.
[0039] In other example embodiments, storage system 200 need not utilize a traditional RAID configuration. For example, extents 215a-n may simply be replicas of each other, or data may be stored as key-value pairs with parity.
[0040] In some embodiments, extents 215a-n need not be physically co-located, but are spread across servers, chassis, and data center locations.
[0041] Referring to FIGS. 3A and 3B, an embodiment of an example storage system 200′ is shown. Storage system 200′ is generally similar to system 200 of FIGS. 2A and 2B, and therefore similar parts are labeled with similar numbers. However, system 200′ has detected that physical drive housing extent 2 has entered read-only mode (in this example, drive / extent 215c could be any of the drives / extents 215 that ultimately enter read-only mode). In order for storage system 200′ to continue operation, the data stored on extent 2 (on drive 215c) must be transferred to replacement physical drive 305 via process 310, the operation of which is described further below. While replacement physical drive 305 may be part of system 200′, it is not considered by mapping function 220 as a drive for storing or reading information until drive 215c enters read-only mode.
[0042] The replacement physical drive 305 can be located anywhere. More specifically, in some examples, the replacement physical drive can reside in a server rack with the other physical drives of extents 215a-n, in a chassis with the physical drives of extents 215a-n, and / or be attached to a network that communicates with the other physical drives of extents 215a-n. In some embodiments, the physical drives of extents 215a-n can be physically dispersed (e.g., in substantially different geographic locations). In such embodiments, the replacement physical drive 305 is selected from a location maximally close to the read-only extent 215c of extents 215a-n to obtain bandwidth / latency benefits during data transfer.
[0043] The decision to convert a physical drive containing an extent to read-only mode is made by the particular individual physical drive or by the storage system 200' as a whole. This decision is made based on a variety of factors that can signal that a particular physical drive is nearing the end of its life. Such factors may include, among others, reaching a critical raw bit error rate on the storage medium, reaching a critical number of erase cycles on the storage medium, or the physical drive itself reaching a critical age.
[0044] Referring to FIG. 4, a method 400 is shown for determining whether and how data should be migrated from a read-only extent 215c to a replacement physical disk 305 containing a replacement extent. In process 405, the system senses whether a drive containing an extent has entered read-only mode (e.g., 215c). In process 410, a policy must be consulted to determine whether the migration should be "fast 420" (i.e., move all data as quickly as possible) or "gradual 430" (as described further below). This policy can be a set of rules and guidelines stored on a computer-readable medium of some type that provides instructions or rules to be used in various cases of migration, although each case need not provide a binary characteristic. That is, factors are weighted relative to one another. Such cases and rules, in some embodiments, include the following non-limiting examples: 1) The number of extents in read-only mode before data loss (e.g., due to parity, mirroring, RAID configurations, etc.) versus the extent errors that can be tolerated, a. Having the number of extents in read-only mode meet or exceed the allowable number of parities will have a positive impact on fast migration. b. Having the number of writable extents when operating to meet or exceed the mirroring threshold will have a positive impact on gradual migration. c. The greater the number of drives in read-only mode, the more positively the policy will affect fast migration. 2) The level of reliability desired by the user of the system 200' is: a. Stronger guarantees of data reliability will have a more positive impact on rapid migration. 3) The level of throughput or bandwidth desired by users of system 200' a. Because large amounts of data need to be moved internally, immediate transfer can require substantial internal bandwidth usage within the system 200', and therefore, if high throughput is desired by the end user, more gradual migration will be used. 4) known or predicted workload / behavior patterns to which current activity is compared; a. For example, considering the discussion regarding 3) above, if system 200' handles heavy workloads periodically on a known schedule, the system may favor rapid data transfer when the workload is expected to be lighter, favoring gradual migration when the workload is expected to be lighter. As time passes through the expected workload schedule, the system may decide to switch from one policy to another. For example, if storage system 200' is intended to process daily backups at midnight and the drives (and associated extents) enter read-only mode at 11:00 PM, system 200' may decide to perform immediate migration up until midnight, and then switch to gradual migration when heavy traffic is expected at midnight.
[0045] If, based on policy, a decision is made to use fast migration (operation 420), all data on extent 215c will be transferred as quickly as possible to replacement physical disk 305. Some example embodiments of this migration are described with respect to FIG.
[0046] If, based on policy, a decision is made to use gradual migration 430, the process may wait for I / O to occur. The nature of the operation that occurs depends on whether the incoming I / O is a read operation 435 or a write operation 440, some embodiments of which are described in Figures 5 and 6, respectively.
[0047] Referring to FIG. 5, an example method 500 for processing a read operation when storage system 200′ determines that “gradual” migration can be used is shown. In process 505, a read operation is received that includes an identifier (block ID) for the data being sought. In process 510, it is determined whether the read request targets data on a read-only (“failed”) extent 215c or on a fully operational extent. Such a determination is made using the ID of the data. If it is determined that the read request targets data on a fully operational extent, in process 555, the data can be read from the fully operational extent and returned. The process can return to process 430 of FIG. 4 to await additional I / O. If in process 510 it is determined that the data can be located on the failed extent 215c, an emergency table is consulted in process 515.
[0048] The emergency table of process 515 can be a table similar to one of the example embodiments illustrated below, which can track whether any strips of failed extent 215c have been migrated to replacement disks / extents 305. It should be noted that in other embodiments, data can be tracked by block ID / block location triplet, or this tracking can be done negatively (e.g., listing what data has not been migrated).
[0049] <Example emergency table 1> The table can track strips as they are migrated. Note that the first number records the extent and the second number records the strip ID. In the example below, the replacement extent is labeled as extent 5. [Table 1]
[0050] <Example emergency table 2> The table can track blocks as they are migrated. Note that, as mentioned above, the location of a particular block can be expressed (at this logical level) as a (Extent, Strip, Depth-level) triplet. This table can be filled in as data is migrated. [Table 2]
[0051] <Example emergency table 3> This table lists all strips (or blocks) on the read-only extent 215c, and a bit is set based on whether the strip (or block) has been migrated to a new extent. Note that the new address of the data is implicitly included in the table, since it is known to have been migrated to the corresponding location on the replacement extent 5. A similar table may be created to track blocks, such as Table 2 above. [Table 3]
[0052] Referring again to FIG. 5, in process 515, the emergency table may be consulted to determine whether the requested data has been migrated to a replacement extent. Such determination is made using the ID of the data. Based on determining that the requested data has been migrated, in process 520, the data is read from the location indicated on the replacement extent 305, and the data may be returned to the host. In some embodiments, if still active, the data may still be read from the read-only extent 215c. The process may return to process 430 of FIG. 4 to await additional I / O.
[0053] Based on determining that the data has not yet been migrated, in process 525, the requested data may be read from read-only extent 215c and returned to the host. In process 530, the data may then be copied or migrated to a corresponding location on a replacement extent on replacement disk 305. In process 535, the ID of the copied data and its new location are added to the emergency table in a manner indicating that it has been migrated. In process 540, the ID of the migrated data may be deleted from the migration table, or the migration table is otherwise updated to indicate that the data has been migrated. Process 540 is optional.
[0054] In some embodiments, a data migration table may be similar to an emergency table, but work in reverse. More specifically, in some embodiments, the emergency table tracks all data that has been migrated, while the migration table tracks all data that has yet to be migrated, where it has been migrated. The tables can effectively be merged into one table; that is, example table 3 serves this purpose. The migration table and emergency table are stored in storage system 200'. In some embodiments, copies of the migration table and / or emergency table are stored on multiple extents 215a-n.
[0055] In process 545, it is determined whether there is any other data requiring migration. This is done by consulting the migration and / or emergency tables. If all data has been migrated, in process 550, the read-only extent 215c is retired and a replacement disk and extent 305 can replace it entirely in the system. This is determined by checking the migration table or by comparing the contents of the emergency table with the contents of the read-only extent. If more data still requires migration, the process returns to process 430 of FIG. 4.
[0056] For purposes of processes 525-540, it should be noted that the migration system can operate on data units of various sizes. For example, when a data block in strip 1 of read-only extent 215c is read, some embodiments of the system can track and copy the entirety of strip 1 (but only the requested block) to replacement extent 305. This can utilize, for example, an emergency and migration table similar to Table 1. Other embodiments can choose to only track and copy the requested block. This can utilize an emergency and migration table similar to Table 2.
[0057] Referring to FIG. 6, an example embodiment of a method 600 for responding to a write request in a storage system 200′ using incremental migration is shown. In process 605, a data write request including an ID (e.g., a block ID) can be received. In process 610, a determination is made as to whether the data should be written to a read-only (failed) extent or to a fully operational extent. Such determination is made using the ID of the data. Based on determining that the data should be written to a fully operational extent, in process 615, the data is written to the fully operational extent. The processor can then return to process 430 of FIG. 4.
[0058] Based on determining that the data is written to the read-only extent 215c, process 620 determines that the data has already been migrated, for example, using an emergency table and / or a migration table. Such a determination is made using the ID of the data. Based on determining that the data has already been migrated to the replacement extent 305, in process 625, the data is written directly to the replacement extent 305. The processor may return to process 430 of FIG. 4. Based on determining that the data does not exist in the emergency table / migration table, in process 630, the data may be added to the emergency table. In process 635, the data may be written to the replacement extent 305. In process 640, a determination is made as to whether the data is found in the migration table. Based on determining that the data exists in the migration table, in process 645, the data may be deleted from the migration table (or may otherwise be marked as migrated).
[0059] Based on determining that the data is not present in the migration table, or after deleting the data from the migration table, a determination is made in process 650 as to whether any more data may be migrated from read-only extent 215c. This may be determined by checking the migration table or by comparing the contents of the emergency table with the contents of the read-only extent. Based on determining that there is no more data to migrate, in process 655, read-only (failed) extent 215c may be completely replaced (and retired) in storage system 200' by extent 305. Based on determining that data remains to be migrated, the process may return to process 430 of FIG. 4.
[0060] 5, the process of FIG. 6 can operate using multiple data size units. For example, when a data block in strip 1 of read-only extent 215c is written, some embodiments of the system can track and copy the entirety of strip 1 (but only requested blocks) to replacement extent 305. This can utilize, for example, an emergency and migration table similar to Table 1. Other embodiments can choose to only track and copy requested blocks. This can utilize an emergency and migration table similar to Table 2.
[0061] Referring to FIG. 7, a method 700 for rapidly migrating data from a read-only extent 215c to a replacement extent / disk 305 is shown. The process of FIG. 7 is used if a decision is made in process 420 of FIG. 4 to utilize rapid migration. To achieve rapid data migration, performance of the method of FIG. 7 is given a relatively high priority over other tasks in system 200′ so that data is migrated as quickly as reasonably possible while maintaining system usability. In process 705, a strip may be selected for migration. The selection process may be performed via various means, such as numerical ordering. In process 710, the strip may be copied from the read-only extent 215c to the replacement extent 305. In process 715, the strip is added to an emergency table (via its ID). In process 720, the strip may be removed from the migration table (or marked as migrated). In process 725, a determination is made as to whether the migrated strip is the last strip on the read-only extent to request migration. Based on determining that more strips request migration, the process can return to 705. Based on determining that more strips do not request migration, process 730 can switch read-only extent 215c to a replacement extent 305 and retire extent 215c. This is determined by checking the migration table or comparing the contents of the emergency table with the contents of the read-only extent. It should be noted that while this process discusses copying strips similar to the process described above with respect to Figures 5 and 6, other embodiments may choose to track and migrate data at the block level.
[0062] In some embodiments, the mapping of extents to drives is a many-to-one relationship. That is, it should be noted that an extent, along with other extents, may be a logical subdivision (logical drive) contained within a physical drive. Thus, the extent selected for migration (e.g., extent 215c) is considered a logical drive (i.e., a logical subdivision of a physical drive). Similarly, the migration destination extent (new extent / drive 305) may also be a logical drive (i.e., a subdivision of a physical drive). It should be noted that if a physical drive enters read-only mode, this will therefore cause a large number of extents to enter read-only mode. This should not change the underlying operational concepts expressed herein. This would simply require migrating more extents, utilizing correspondingly more (or numerically larger) emergency and migration tables.
[0063] In some embodiments, the processes of FIGS. 4-7 are adapted for storage system 200′ that utilizes parity. For example, whenever data is written to a block, the parity appropriate for that block is also updated by the write. Thus, if parity for a particular set of blocks resides on read-only extent 215c, any update to one of the set of blocks on a fully operational extent also results in a write to read-only extent 215c, triggering the process of FIG. 6. In another example, if a read occurs for data on a fully operational extent, the system can choose to treat it as a read for parity data as well, triggering the process of FIG. 5 (this is possible because parity data is rarely read; such a mechanism would speed up the migration of parity data via reads).
[0064] In some embodiments, the processes of Figures 4-7 are adapted to use storage systems that do not use block-based storage. For example, each strip may contain one or more values or portions of a key-value store. The key serves as an ID for the data, and the data is located through a lookup table. Thus, for example, any read or write on a key-value value stored on a strip that is at least partially located on a read-only extent will trigger the operations of Figures 5 and 6.
[0065] The methods disclosed herein propose that various component processes may be omitted and rearranged and still be within the inventive scope of the present disclosure, as will be understood by one of ordinary skill in the art upon reading this disclosure.
[0066] The processes of the method may be performed by one or more programmable processors executing a computer program stored on a tangible computer-readable medium to operate on input data to generate output and perform functions. The processes of the method may be performed by, and an apparatus may be embodied as, special purpose logic circuitry such as, for example, a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0067] In various embodiments, the computer-readable medium may include instructions that, when executed, cause a device to perform at least a portion of the processes of a method. In some embodiments, the computer-readable medium is included in a magnetic medium, an optical medium, other medium, or a combination thereof (e.g., a CD-ROM, a hard drive, a read-only memory (ROM), a flash memory). In such embodiments, the computer-readable medium is a tangibly and permanently embodied article of manufacture.
[0068] While the principles of the present disclosure have been described with reference to exemplary embodiments, it will be apparent to those skilled in the art that numerous changes and modifications can be made without departing from the spirit and scope of the disclosed concepts. Accordingly, the above-described embodiments should be understood to be merely illustrative, not limiting. The scope of the disclosed concepts should therefore be determined by the broadest permissible interpretation of the following claims and their equivalents, and should not be limited or restricted by the above description. It is therefore intended that the appended claims encompass all such modifications and variations within the scope of the present embodiments.
[0069] The processes of the method are performed by one or more programmable processors executing a computer program to operate on input data to generate output and perform functions. The processes of the method may also be performed by, and an apparatus may be embodied as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0070] In various embodiments, the computer-readable medium may include instructions that, when executed, cause a device to perform at least a portion of the processes of a method. In some embodiments, the computer-readable medium is included in a magnetic medium, an optical medium, other medium, or a combination thereof (e.g., a CD-ROM, a hard drive, a read-only memory (ROM), a flash memory). In such embodiments, the computer-readable medium is a tangibly and permanently embodied article of manufacture.
[0071] Embodiments of the present invention extend, without limitation, to the following statements.
[0072] Statement 1) A method of migrating data in a storage system including a plurality of active extents and one or more replacement logical or physical drives, the method comprising: determining that an active extent of the storage system has entered a read-only mode to become a read-only extent; determining whether the storage system should migrate data quickly or gradually based on a storage policy; copying data from the read-only extent to a replacement logical or physical drive as quickly as possible based on the storage system determining that the data must be migrated quickly; and copying the data unit from the read-only drive to the replacement logical or physical drive when the storage system receives a read request for the data unit of the read-only extent based on the determination that data must be incrementally migrated, and writing the data unit to the replacement logical or physical drive when the storage system receives a write request for the data unit of the read-only extent.
[0073] Statement 2) The method of statement 1, wherein upon receiving a read request for a data unit, all strips within the read-only extent that contain the data unit are copied to the replacement logical or physical drive.
[0074] Statement 3) In the method of statement 1, when a data unit is copied from the read-only extent to the replacement logical or physical drive as a result of a read operation, or when a data unit is written to the replacement logical or physical drive as a result of a write operation to the read-only extent, the data unit is recorded in an emergency table as having been migrated to the replacement logical or physical drive.
[0075] Statement 4) The method of statement 3, further comprising the step of redirecting subsequent reads or writes to data units found in the emergency table to the replacement logical or physical drive.
[0076] Statement 5) The method of statement 4, further comprising: substituting the read-only extent for the replacement logical or physical drive as an active extent based on determining that all data units in the read-only extent have been migrated to the replacement logical or physical drive.
[0077] Statement 6) In the method of statement 5, the step of determining that all data units of the read-only extent have been migrated includes a step of comparing a list of migrated data units on the emergency table with a list of data units of the read-only extent, a step of referring to a migration table, or a step of the migration table making a determination that all data units have been migrated.
[0078] Statement 7) The method of statement 1, wherein the read-only extent enters read-only mode because it has reached a critical level of wear.
[0079] Statement 8) The method of statement 7, wherein the criticality level of the data is determined by one or more of a raw-media error rate, a number of erase cycles for a portion of the storage media, and a critical value for the product life to be exceeded.
[0080] Statement 9) The method of statement 1, wherein the storage policy includes one or more of a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement.
[0081] Statement 10) A processor of a computer storage system including a non-transitory computer-readable medium, the computer storage system including a plurality of active extents and one or more replacement logical or physical drives, determining that an active extent of the storage system has entered read-only mode; Based on the storage policy, the storage system determines whether data should be migrated quickly or gradually; the storage system, based on determining that the data must be migrated quickly, copies the data from the read-only extent to a replacement logical or physical drive as quickly as possible; a non-transitory computer-readable medium having instructions that cause the storage system to perform the steps of: copying the data unit from the read-only drive to the replacement logical or physical drive when a read request is received for the data unit of the read-only extent based on a determination that data should be gradually migrated; and writing the data unit to the replacement logical or physical drive when a write request is received for the data unit of the read-only extent.
[0082] Statement 11) The non-transitory computer-readable medium of statement 10, wherein upon receiving a read request for a data unit, all strips within the read-only extent that contain the data unit are copied to the replacement logical or physical drive.
[0083] Statement 12) The non-transitory computer-readable medium of statement 10, wherein when a data unit is copied from the read-only extent to the replacement logical or physical drive as a result of a read operation, or when a data unit is written to the replacement logical or physical drive as a result of a write operation to the read-only extent, the data unit is recorded in an emergency table as having been migrated to the replacement logical or physical drive.
[0084] Statement 13) The non-transitory computer-readable medium of statement 12, further comprising: redirecting subsequent reads or writes to data units found on the emergency table to the replacement logical or physical drive.
[0085] Statement 14) The non-transitory computer-readable medium of Statement 13, further comprising substituting the read-only extent onto the replacement logical or physical drive as an active extent based on determining that all data units within the read-only extent have been migrated to the replacement logical or physical drive.
[0086] Statement 15) The non-transitory computer-readable medium of statement 14, wherein the step of determining that all data units of the read-only extent have been migrated includes a step of comparing a list of migrated data units on the emergency table with a list of data units of the read-only extent, or a step of making a determination that a migration table indicates that all data units have been migrated.
[0087] Statement 16) The non-transitory computer-readable medium of statement 10, wherein the read-only extent enters a read-only mode because a critical level of wear has been reached.
[0088] Statement 17) The non-transitory computer-readable medium of Statement 16, wherein the criticality level of the storage media is determined by one or more of a raw-media error rate, a number of erase cycles for a portion of the storage media, and a product life criticality that is exceeded.
[0089] Statement 18) The non-transitory computer-readable medium of statement 10, wherein the storage policy includes one or more of a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement.
[0090] Statement 19) A storage system, comprising: Multiple active extents; one or more replacement drives; a processor; a memory coupled to the processor; The memory, when executed by the processor, causes the processor to: determining that an active extent of the storage system has entered read-only mode to become a read-only extent; Based on the storage policy, the storage system determines whether data should be migrated quickly or gradually; the storage system, based on determining that the data must be migrated quickly, copies the data from the read-only extent to a replacement logical or physical drive as quickly as possible; The storage system stores instructions that, based on determining that data must be migrated incrementally, copy the data unit from the read-only drive to the replacement logical or physical drive when a read request for the data unit of the extent is received, and write the data unit to the replacement logical or physical drive when a write request for the data unit of the read-only extent is received. [Explanation of symbols]
[0091] 110: Processor and / or logic 120: Volatile memory 130: Non-volatile memory 140: Network interface 150: User Interface Unit 160: Other hardware devices 170: Software
Claims
1. 1. A method of migrating data in a storage system including a plurality of active extents and one or more replacement logical or physical drives, comprising: determining that an active extent of the storage system has entered a read-only mode to become a read-only extent; determining whether the storage system should migrate data quickly or gradually based on a storage policy; the storage system, based on determining that the data must be migrated quickly, copying the data from the read-only extent to a replacement logical or physical drive as quickly as possible; and copying the data unit from a read-only drive to the replacement logical or physical drive when the storage system receives a read request for the data unit of the read-only extent based on a determination that data must be incrementally migrated, and writing the data unit to the replacement logical or physical drive when the storage system receives a write request for the data unit of the read-only extent.
2. Upon receiving a read request for a data unit, all strips within the read-only extent that contain the data unit are copied to the replacement logical or physical drive. The method of claim 1.
3. When a data unit is copied from the read-only extent to the replacement logical or physical drive as a result of a read operation, or when a data unit is written to the replacement logical or physical drive as a result of a write operation to the read-only extent, the data unit is recorded in an emergency table as being migrated to the replacement logical or physical drive. The method of claim 1.
4. and redirecting subsequent reads or writes to data units found in the emergency table to the replacement logical or physical drive. The method of claim 3.
5. and switching the read-only extent on the replacement logical or physical drive as an active extent based on determining that all data units in the read-only extent have been migrated to the replacement logical or physical drive. The method of claim 4.
6. Determining that all data units of the read-only extent have been migrated includes comparing a list of migrated data units on the emergency table with a list of data units of the read-only extent, or a migration table makes a determination indicating that all data units have been migrated. The method of claim 5.
7. The read-only extent enters read-only mode because it has reached a critical level of wear. The method of claim 1.
8. The criticality level of the wear is determined by one or more of the raw-media error rate, the number of erase cycles for a portion of the storage media, and a product life criticality that is exceeded. The method of claim 7.
9. The storage policy includes one or more of a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement. The method of claim 1.
10. 1. A non-transitory computer-readable medium, comprising: a processor in a computer storage system including a plurality of active extents and one or more replacement logical or physical drives; determining that an active extent of the storage system has entered a read-only mode to become a read-only extent; determining whether the storage system should migrate data quickly or gradually based on a storage policy; the storage system, based on determining that the data must be migrated quickly, copying the data from the read-only extent to a replacement logical or physical drive as quickly as possible; and upon receiving a read request for the data unit of the read-only extent based on the determination that the data should be migrated incrementally, copying the data unit from a read-only drive to the replacement logical or physical drive, and upon receiving a write request for the data unit of the read-only extent, writing the data unit to the replacement logical or physical drive. Non-transitory computer-readable medium.
11. Upon receiving a read request for a data unit, all strips within the read-only extent that contain the data unit are copied to the replacement logical or physical drive. The non-transitory computer-readable medium of claim 10.
12. When a data unit is copied from the read-only extent to the replacement logical or physical drive as a result of a read operation, or when a data unit is written to the replacement logical or physical drive as a result of a write operation to the read-only extent, the data unit is recorded in an emergency table as being migrated to the replacement logical or physical drive. The non-transitory computer-readable medium of claim 10.
13. and redirecting any subsequent reads or writes to the data units found in the emergency table to the replacement logical or physical drive. The non-transitory computer-readable medium of claim 12.
14. and substituting the read-only extent to the replacement logical or physical drive as an active extent based on determining that all data units in the read-only extent have been migrated to the replacement logical or physical drive.
14. The non-transitory computer-readable medium of claim 13.
15. Determining that all data units of the read-only extent have been migrated includes comparing a list of migrated data units on the emergency table with a list of data units of the read-only extent, or a migration table makes a determination indicating that all data units have been migrated.
15. The non-transitory computer-readable medium of claim 14.
16. The read-only extent enters read-only mode because it has reached a critical level of wear. The non-transitory computer-readable medium of claim 10.
17. The criticality level of the wear is determined by one or more of a raw-media error rate, a number of erase cycles for a portion of the storage media, and a product life criticality that is exceeded.
17. The non-transitory computer-readable medium of claim 16.
18. The storage policy includes one or more of a minimum throughput requirement, a minimum bandwidth requirement, a durability requirement, a RAID level requirement, and a parity level requirement. The non-transitory computer-readable medium of claim 10.
19. 1. A storage system, comprising: Multiple active extents; one or more replacement drives; a processor; a memory coupled to the processor; The memory, when executed by the processor, causes the processor to: determining that an active extent of the storage system has entered read-only mode to become a read-only extent; Based on the storage policy, the storage system determines whether to migrate data quickly or gradually; the storage system, based on determining that the data must be migrated quickly, copies the data from the read-only extent to a replacement logical or physical drive as quickly as possible; copying the data unit from the read-only drive to the replacement logical or physical drive upon receiving a read request for the data unit of the read-only extent based on the storage system determining that data should be incrementally migrated; and stores a command to write the data unit to the replacement logical or physical drive when a write request for the data unit of the read-only extent is received. Storage system.
Citation Information
Patent Citations
Computer system
JP2000194508A
Data storage device, data migration method and program
JP2012133436A
Data migration method, script generation method to be used therein, and information processor
JP2014215686A
Server device, data management system, data management method, and program
JP2015138334A
Information processing system
US10089027B2