Storage systems, computing systems and methods

By classifying and processing user data, parity data, and log data in a RAID storage system and storing them in a multi-stream mode, the write hole problem is solved, and the durability and performance of the storage system are improved.

CN109582219BActive Publication Date: 2025-10-31INTEL CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201810985334.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-28
Filing Date
2018-08-28
Publication Date
2025-10-31
Estimated Expiration
2038-08-28

AI Technical Summary

Technical Problem

Existing RAID storage systems are prone to parity inconsistencies during write operations, leading to write gaps that affect data recovery and storage efficiency.

Method used

By classifying and processing user data, parity data, and log data, and distributing them across multiple storage devices using a multi-stream model, frequent updates of parity data and log data are ensured, reducing the occurrence of write gaps.

Benefits of technology

It improves the durability and performance of the storage system, reduces the write amplification factor, and enhances the reliability and efficiency of data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109582219B_ABST
    Figure CN109582219B_ABST
Patent Text Reader

Abstract

According to various aspects, a storage system is provided, comprising a plurality of storage devices and one or more processors configured to: store user data on the plurality of storage devices, the stored user data being distributed among the plurality of storage devices along with redundant data and log data; generate classifications associated with the redundant data and log data to provide classified redundant data and classified log data; and write the classified redundant data and classified log data to the respective storage devices among the plurality of storage devices according to the classifications associated with the redundant data and log data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The various aspects generally involve storage systems and the methods used to operate storage systems. Background Technology

[0002] Efficient data processing (e.g., including storing, updating, and / or retrieving data) is becoming increasingly important due to the increasing volume and flow of data with modern technologies. In one or more applications, data is stored using RAID (Redundant Array of Independent Disks or Redundant Array of Independent Drives) technology. RAID storage technology can be implemented in hardware (also known as hardware RAID), in software (also known as software RAID), or in both hardware and software (also known as hybrid RAID or host RAID). RAID storage technology can be offered in various types or modifications. These types or modifications can differ from one another regarding the number of storage devices used, the type of partitioning, and / or the addressing of the respective storage devices and / or the embedded functionality to prevent data loss in the event of one or more failures of the storage devices used. Different types of RAID storage technology can be referred to as RAID levels. Currently, several standard RAID levels and non-standard RAID levels are available, such as RAID-0, RAID-1, RAID-5, and RAID-6. However, various combinations or modifications of standard and non-standard RAID levels are possible, resulting in a large number of possible RAID levels, such as RAID-01, RAID-05, RAID-10, RAID-1.5, RAID-15, RAID-1E, RAID-1E0, RAID-30, RAID-45, RAID-50, RAID-51, RAID-53, RAID-55, RAID-5E, RAID-5EE, RAID-5DE, RAID-60, Matrix-RAID, RAID-S, RAID-TP, RAID-Z, etc. Attached Figure Description

[0003] Throughout the accompanying drawings, it should be noted that the same reference numerals are used to depict the same or similar elements, features, and structures. The drawings are not necessarily drawn to scale, but generally focus on illustrating aspects of this disclosure. In the following description, some aspects of this disclosure are described with reference to the following drawings, in which:

[0004] Figure 1 A schematic diagram illustrates a storage system based on various aspects;

[0005] Figure 2 The diagram illustrates multiple storage devices in a storage system based on various aspects;

[0006] Figure 3A schematic diagram illustrates one or more processors in a storage system based on various aspects;

[0007] Figures 4A to 4C The various writing strategies for unclassified and classified data are illustrated based on various aspects.

[0008] Figure 5 A schematic flowchart illustrating methods for operating a storage system from various perspectives is shown;

[0009] Figure 6 A schematic flowchart illustrating methods for operating a storage system from various perspectives is shown;

[0010] Figures 7A to 7C The writing strategies for categorized data are illustrated based on various aspects.

[0011] Figure 8A The diagram illustrates various aspects of storage systems and how to access them.

[0012] Figure 8B A schematic diagram illustrates a storage system based on various aspects;

[0013] Figure 8C A schematic flowchart illustrating methods for operating a storage system according to various aspects is shown; and

[0014] Figure 9A and Figure 9B Write amplification measurements for storage systems are shown from various perspectives. Detailed Implementation

[0015] The following detailed description refers to the accompanying drawings, which illustrate, by way of illustration, specific details and aspects in which this disclosure may be practiced. These aspects are described in sufficient detail to enable those skilled in the art to practice this disclosure. Other aspects may be utilized, and structural, logical, and electrical changes may be made without departing from the scope of this disclosure. The various aspects are not necessarily mutually exclusive, as some aspects may be combined with one or more other aspects to form new aspects. The various aspects are described in conjunction with methods, and the various aspects are described in conjunction with devices. However, it is understood that aspects described in conjunction with methods can be similarly applied to devices, and vice versa.

[0016] The terms "at least one" and "one or more" can be understood to include any integer greater than or equal to one, i.e., one, two, three, four [...], etc. The term "multiple" can be understood to include any integer greater than or equal to two, i.e., two, three, four, five [...], etc.

[0017] The phrase “at least one” regarding a group of elements may be used herein to mean at least one element from a group of elements. For example, the phrase “at least one” regarding a group of elements may be used herein to mean a choice of: one of the listed elements, one of the plural listed elements, a plural of individual listed elements, or a plural of multiple listed elements.

[0018] The terms “plural” and “multiple” in the specification and claims explicitly refer to a quantity greater than one. Therefore, any phrase explicitly referencing the above words (e.g., “plural [objects]”, “multiple [objects]”) referring to a certain number of objects explicitly refers to more than one of the objects. The terms “group of,” “set of,” “collection of,” “series of,” “sequence of,” “group of,” etc. (if any) in the specification and claims refer to a quantity equal to or greater than one, i.e., one or more.

[0019] Although the terms first, second, third, etc., may be used herein to describe various elements, components, regions, layers, and / or parts, these elements, components, regions, layers, and / or parts should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, and / or part from another. Unless the context clearly indicates otherwise, terms such as “first,” “second,” and other numerical terms, when used herein, do not imply sequence or order. Therefore, without departing from the teachings of the examples, the first element, component, region, layer, or part discussed below may be referred to as the second element, component, region, layer, or part.

[0020] As used herein, the term "data" can be understood to include information in any suitable analog or digital form, such as as a file, part of a file, a set of files, a signal or stream, a set of signals or streams, etc. Furthermore, the term "data" can also be used to indicate a reference to information, for example, in the form of a pointer.

[0021] For example, the terms "processor" or "controller" as used herein can be understood as any kind of entity that allows the processing of data. Data can be processed according to one or more specific functions performed by the processor or controller. Furthermore, as used herein, "processor" or "controller" can be understood as any kind of circuit, such as any kind of analog or digital circuit. For example, the terms "disposal" or "processing" (referring to data processing, document processing, or request processing) as used herein can be understood as any kind of operation (e.g., I / O operation) or any kind of logical operation. I / O (input / output) operations can be, for example, storage (also known as writing) and reading.

[0022] Therefore, a processor or controller can be or includes analog circuits, digital circuits, mixed-signal circuits, logic circuits, processors, microprocessors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), integrated circuits, application-specific integrated circuits (ASICs), and any combination thereof. Any other kind of implementation of the corresponding functions, which will be described in more detail below, can also be understood as a processor, controller, or logic circuit. It should be understood that any two (or more) of the processors, controllers, or logic circuits detailed herein can be implemented as a single entity with equivalent functionality, and conversely, any single processor, controller, or logic circuit detailed herein can be implemented as two (or more) separate entities with equivalent functionality.

[0023] In current technology, the distinction between software and hardware implementations of data processing can be blurred. Therefore, it must be understood that the processors, controllers, or circuits detailed herein can be implemented in software, hardware, or as a hybrid implementation that includes both software and hardware.

[0024] The term “system” (e.g., storage system, RAID system, computing system, etc.) as detailed in this article can be understood as a set of interacting elements; by way of example rather than limitation, these elements can be one or more mechanical components, one or more electronic components, one or more instructions (e.g., encoded in storage media), one or more processors, etc.

[0025] The term "storage" (e.g., storage device, storage system, etc.) as used in detail herein can be understood as any suitable type of memory or storage device, such as hard disk drives (HDDs), solid-state drives (SSDs), and any other suitable storage device. The term "storage" may also be used herein to refer to, for example, permanent data storage based on non-volatile memory.

[0026] As used herein, the terms “memory,” “memory device,” etc., are understood to refer to a non-transitory computer-readable medium in which data or information can be stored for retrieval. Therefore, the reference to “memory” as used herein is understood to refer to volatile or non-volatile memory, including random access memory (RAM), read-only memory (ROM), flash memory, solid-state storage devices, magnetic tape, hard disk drives, optical drives, and 3D XPoint. TMTechnologies, etc., or any combination thereof. Furthermore, it should be recognized that registers, shift registers, processor registers, data buffers, etc., are also included herein by the term memory. It should be understood that a single component referred to as "memory" or "a memory" can consist of more than one different type of memory, and therefore can refer to a collective component including one or more types of memory. It is readily understood that any single memory component can be divided into multiple collectively equivalent memory components, and vice versa. Furthermore, while memory may be depicted as separate from one or more other components (e.g., in the figures), it should be understood that memory can be integrated within another component, for example, on a common integrated chip.

[0027] Volatile memory can be a storage medium that uses electricity to maintain the state of data stored by the medium. Non-limiting examples of volatile memory can include various types of RAM, such as Dynamic Random Access Memory (DRAM) or Static Random Access Memory (SRAM). One particular type of DRAM that can be used in memory modules is Synchronous Dynamic Random Access Memory (SDRAM). In some aspects, the DRAM of a memory component can conform to standards issued by the Joint Electronic Equipment Committee (JEDEC), such as JESD79F for Double Data Rate (DDR) SDRAM, JESD79-2F for DDR2 SDRAM, JESD79-3F for DDR3 SDRAM, JESD79-4A for DDR4 SDRAM, JESD209 for Low Power DDR (LPDDR), JESD209-2 for LPDDR2, JESD209-3 for LPDDR3, and JESD209-4 for LPDDR4 (these standards are available at www.jedec.org). These standards (and similar standards) can be referred to as DDR-based standards, and the communication interfaces of storage devices that implement such standards can be referred to as DDR-based interfaces.

[0028] Various aspects can be applied to any memory device, including non-volatile memory. In one aspect, the memory device is a block-addressable memory device, such as those based on NAND or NOR logic technologies. Memory can also include next-generation non-volatile devices, such as 3D XPoint. TM Technical memory devices, or other byte-addressable in-situ write-able non-volatile memory devices. 3D XPoint TM Technical memories may include transistorless stackable cross-point architectures, where memory cells are located at the intersection of word lines and bit lines and are individually addressable, and where bit storage is based on variations in bulk resistance.

[0029] In some aspects, a memory device can be or may include memory devices using chalcogenide glass, multi-threshold NAND flash memory, NOR flash memory, single-level or multi-level phase-change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), antiferroelectric memory, magnetoresistive random access memory (MRAM) incorporating memristor technology, including metal oxide-based, oxygen vacancy-based resistive memory and bridged random access memory (CB-RAM), or spin-transfer torque (STT)-MRAM, spintronic junction-based devices, magnetic tunnel junction (MTJ-based devices), DW (domain wall) and SOT (spin-orbit transfer)-based devices, thyristor-based memory devices, or combinations of any of the above devices or other memories. The term memory or memory device can refer to the die itself and / or the packaged memory product.

[0030] Depending on various aspects, computing systems and / or storage systems may be provided, including RAID based on multiple member storage devices and one or more strategies for efficiently processing data written to the respective member storage devices. Illustratively, a RAID system may be provided, or in other words, a computing system and / or storage system implementing one or more RAID functions may be provided. Furthermore, corresponding methods for operating the RAID system may be provided.

[0031] Depending on various aspects, the RAID system described herein can be based on striping utilizing distributed redundancy information (e.g., striping at the bit, byte, or block level). Depending on various aspects, the RAID system described herein can be RAID-5, RAID-6, or a RAID level based thereon. Depending on various aspects, the RAID system described herein can be based on block-level striping utilizing distributed parity blocks. Depending on various aspects, parity information can be used as redundancy information for error correction, for example, to prevent information loss in the event of a failure of one member storage device in the RAID system. Various aspects involve RAID (e.g., 5 / 6 / EC) with parity for physical data placement on storage devices (e.g., for NAND physical data placement) and early write log classification.

[0032] As an example, RAID-5 can utilize distributed parity to achieve block-level striping. Parity data can be distributed across the storage devices used in the RAID system (also known as member storage devices or member drives). During the operation of the RAID system, member storage devices may fail or other errors may occur. In this case, data on one member storage device may be lost. This lost data can be recalculated based on the distributed parity data and data from the remaining storage devices. Depending on various aspects, a RAID system based on striping using distributed parity data can include at least three storage devices. In this case, at least three individual physical drives can be used, such as at least three hard disk drives or at least three solid-state drives, or some combination thereof. Alternatively, for example, in virtualization, at least three logical storage devices can be used, independent of the number of underlying physical drives.

[0033] Parity data (or other suitable redundancy information) can be used for fault tolerance in RAID systems, depending on various factors. Parity data can be provided through calculations based on data from two or more member storage devices (referred to herein as user data). The parity data (as a result of the calculation) can be stored on another member storage device. This calculation can be based on a logical XOR (Exclusive Orb) operation, which descriptively provides parity information. However, redundancy data can be provided in a similar manner using other logical operations besides XOR.

[0034] Typically, parity inconsistency can occur in the event of a RAID system crash. As an example, a system crash or another interruption of I / O operations may end up in a corrupted state where parity information may be inconsistent with the corresponding user data. From this corrupted state, the parity information may be insufficient for data recovery. A corrupted state can occur if one of the user data or parity data is written without a corresponding write operation to the other. In RAID systems, parity inconsistency problems are also known as write holes or RAID write holes. Therefore, a RAID write hole can be a data corruption problem caused, for example, by an interruption in write operations to the used data and its corresponding parity data.

[0035] Figure 1 A storage system 100 is illustrated schematically according to various aspects. As mentioned above, the storage system 100 may be a RAID storage system. The storage system 100 may include multiple storage devices 101, also referred to as member storage devices. As an example, Figure 1The diagram shows three storage devices 101a, 101b, and 101c. However, more than three storage devices can be used in a similar manner; for example, four storage devices 101 can be used in a RAID-5 configuration, five storage devices can be used in a RAID-6 configuration, and so on.

[0036] According to various embodiments, storage system 100 may include one or more processors 102. One or more processors 102 may be configured to receive user data 111 to be stored 104. One or more processors 102 may be configured to distribute the user data 111 to be stored across multiple storage devices 101. The user data 111 may be distributed across the multiple storage devices 101 and stored therein along with corresponding redundant data 113 for data recovery. Through this interaction, user data can be recovered, for example, in the event of a failure of one of the multiple storage devices. Furthermore, log data 115 (e.g., corresponding to redundant data 113 and / or user data 111) may be distributed among the multiple storage devices 101 for write hole protection. Data (e.g., user data 111 and / or redundant data 113) may be distributed across the multiple storage devices 101 via striping (e.g., at the bit, byte, or block level).

[0037] Depending on various aspects, one or more processors 102 can be configured to generate classifications 124 associated with redundant data 113 and / or log data 115. Illustratively, classified redundant data 113 and classified log data 115 may be provided. Classification 124 may be implemented based on classification labels associated with the corresponding data or data type, based on lookup tables associated with the corresponding data or data type, etc.

[0038] Furthermore, depending on various aspects, one or more processors 102 may be configured to write 114 redundant data 113 and log data 115 to corresponding storage devices in a plurality of storage devices according to a classification 124 associated with redundant data 113 and log data 115.

[0039] In a similar manner, user data 111 can be categorized. In this case, one or more processors 102 can be configured to write user data 111 to corresponding storage devices in a plurality of storage devices according to the categorization 124 associated with user data 111.

[0040] Depending on various aspects, redundant data 113 can be parity data. Parity data can be calculated, for example, using an XOR operation. Various examples of parity data 113 are provided below with reference to it; however, any other redundant data 113 suitable for data recovery can be used in a similar manner (e.g., in the event of a failure of one of multiple storage devices). Depending on various aspects, log data 115 used for write hole protection can also be referred to as pre-written log data.

[0041] Depending on various aspects, in addition to the corresponding distribution of data 111, 113, 115 among the multiple storage devices 101, classification 124 can provide a basis for placing (e.g., physically arranging) data 111, 113, 115 on each of the multiple storage devices 101 according to classification 124.

[0042] Figure 2 Multiple storage devices 101 are shown according to various aspects. The multiple storage devices 101 may be part of a storage system 100 as described herein.

[0043] According to various embodiments, multiple storage devices 101 can be configured in a striped configuration (also known as RAID striping, disk striping, drive striping, etc.). In this case, user data 111 to be stored can be distributed across multiple physical drives. However, in some aspects, the data 111 to be stored can be distributed across multiple logical drives. Each of the multiple storage devices 101 can be divided into stripes 203. Across the multiple storage devices 101, stripes 203 can form multiple stripes 201. In other words, each stripe among the multiple stripes 201 can include multiple stripes 203. In other words, storage devices 101a, 101b, 101c among the multiple storage devices 101 can be divided into stripes 203. One stripe 203 from each of the multiple storage devices 101 can provide stripes 201 across the multiple storage devices 101.

[0044] According to various embodiments, data (e.g., user data 111 and / or redundant data 113) can be written into stripes 203 along corresponding stripes 201. Each stripe 203 can be associated with a storage block (e.g., in block-level striping). Each storage block in the storage block can have a predefined block size, e.g., 128 kiB. However, other block sizes can be used in a similar manner.

[0045] Depending on various aspects, one or more processors 102 as described herein can be configured to distribute parity data 113 across multiple storage devices 101, for example, Figure 1 and Figure 2As shown in the diagram. As an example, parity data 113 of the first stripe 201a can be written into the corresponding stripe 203 of the third storage device 101c among the multiple storage devices 101. The calculation of these parity data 113 can be based on user data 111 written into the corresponding stripe 203 of the first storage device 101a and the second storage device 101b among the multiple storage devices 101. Furthermore, parity data 113 of the second stripe 201b can be written into the corresponding stripe 203 of the first storage device 101a among the multiple storage devices 101. The calculation of these parity data 113 can be based on user data 111 written into the corresponding stripe 203 of the second storage device 101b and the third storage device 101c among the multiple storage devices 101. Furthermore, the parity data 113 of the third stripe 201c can be written to the corresponding stripe 203 of the second storage device 101b among the multiple storage devices 101. The calculation of these parity data 113 can be based on the user data 111 written to the corresponding stripe 203 of the first storage device 101a and the third storage device 101c among the multiple storage devices 101. And so on. In addition, the log data 115 for each write operation in the write operation can be written to the corresponding storage device accordingly.

[0046] For each instance of writing user data 111 to the corresponding stripe 203, the parity data 113 of stripe 201 can be updated accordingly to maintain parity consistency. Therefore, since two or more user data stripes 203 are included in each stripe of stripe 201 corresponding to a parity data stripe, the parity data 113 can be written more frequently than the user data 111.

[0047] Furthermore, log data 115 for each write operation can be written more frequently than user data 111. For example, for each parity data 113 write operation, a corresponding log data 115 write operation can be performed, see example. Figure 8A .

[0048] Depending on various aspects, providing a classification 124 for user data 111, parity data 113, and / or log data 115 can allow these data to be stored efficiently on the corresponding storage devices 101a, 101b, and 101c. For example, the corresponding storage devices 101a, 101b, and 101c may include erase block sizes for deleting or updating data. The erase block size can be larger than the write block size used for writing data, as is the case with SSDs (e.g., NAND SSDs). In this case, garbage collection can be used to separate valid data that should be retained from invalid data within a single erase block, which should be deleted to provide storage space for new write operations.

[0049] Based on the classification 124 provided for user data 111, parity data 113, and / or log data 115, this data can be stored such that only data with similar statistical lifetimes is stored in a single erase block. This allows for the avoidance or reduction of garbage collection efforts, since all data in a single erase block will statistically become invalid substantially simultaneously (i.e., after its statistical lifetime or after a predefined lifetime). This allows data in the erase block to be erased without prior separation of valid data, see, for example... Figure 4A and Figure 4B .

[0050] Depending on various aspects, when multiple storage devices have a first granularity for writing data and a second granularity larger than the first granularity for updating or deleting data, utilizing the classification 124 of data can be an efficient way to reduce the write amplification (WA) factor (WAF). This is, for example, if NAND (“NAND”) memory (e.g., NAND SSD or NAND flash memory device) can be used as storage device 101.

[0051] The term write amplification (WA) or write amplification factor (WAF) refers to the effect that occurs when the amount of physical data actually written is greater than the amount of logical data written by the host computer (see Figure 8). Figure 9A and Figure 9B As an example, NAND flash memory may include storage elements that must be erased before they can be rewritten. Furthermore, NAND flash memory may allow writing a single page at a time (e.g., with page sizes of, for example, 4kiB to 64kiB or larger), and, however, erasing only one block at a time (also called a NAND block or erase block). An erase block may include multiple pages, such as hundreds of pages. Therefore, internal movement of user data 111 can be performed, for example, in a background thread (also known as garbage collection), removing user data 111 from the block to be erased (or, in other words, the data that should be stored). Therefore, the total number of write operations performed on this type of storage device can typically be greater than the number of write operations intended to be written. The write amplification factor (WAF) is the mathematical representation of this phenomenon and represents the ratio of physical write operations to logical write operations. Small block random writes can result in a higher WAF and more drive wear compared to large block sequential writes, for example... Figure 9A and Figure 9B As shown in the diagram. Similarly, a "full" drive typically has a higher WAF compared to an "empty" drive.

[0052] Figure 3The configuration of one or more processors 102 is illustrated schematically according to various aspects. One or more processors 102 may be part of the storage system 100 described herein. According to various aspects, one or more processors 102 may also be configured to write user data 111 and parity data 113 to each of the plurality of storage devices 101 (exemplarily shown for the first storage device 101a) via a first data stream 313a directed to a first storage region 301a of storage device 101a and a second data stream 313b directed to a second storage region 301b of the respective storage device 101a, which is different from the first storage region 301a.

[0053] Depending on various aspects, based on classification 124 as described above, one or more processors 102 can be configured such that the first data stream 313a may include only parity data, and the second data stream may include both user data and parity data. Alternatively, based on classification 124 as described above, one or more processors 102 can be configured such that the first data stream 313a may include only parity data 113, and the second data stream may include only user data 111. Furthermore, based on classification 124 as described above, one or more processors 102 can be configured such that the first data stream 313a may include only parity data 113 and log data 115 (e.g., only data with statistically shorter lifetimes), and the second data stream may (e.g., only) include user data 111 (e.g., only data with statistically longer lifetimes).

[0054] Illustratively, each storage device in storage device 101 can be written in a multi-stream mode. Depending on various aspects, at least two data streams 313a, 313b can be used to write data to different blocks of the respective storage device 101 according to the classification associated with the data. This can provide storage device 101 with better durability properties, improved performance of storage system 100, and / or consistent latency. It is expected that all data associated with the first stream 313a can be invalidated (e.g., updated, disposed of) substantially simultaneously.

[0055] Furthermore, as described above, striping can be performed when the storage devices 101 are arranged in a RAID system. Using an N-drive array (in other words, a storage system 100 with N storage devices 101 in a RAID configuration) as an example, the first bit, byte, or block (e.g., a write block) can be written to the first drive, the second bit, byte, or block can be written to the second drive, and so on; until the (N-1)th bit, byte, or block, and the parity bit, byte, or block are written to the Nth drive. Then, the (N+1)th bit, byte, or block is written to the first drive again, and the entire process restarts, for example, by arranging the parity bit, byte, or block on different drives.

[0056] Depending on various factors, when writing files to a RAID storage system 100 using block striping and the file size is larger than the block size, the file is split into portions with a block size, which can be, for example, 128 kB or any other suitable size. This block size can also be referred to as the stripe size. The stripe size can be predefined. Depending on various factors, the stripe size can be smaller than the erase block size.

[0057] For example, such as Figure 3 As shown, the first storage area 301a may include another erase block different from the second storage area 301b, such that data streams 313a and 313b write corresponding data to different erase blocks to avoid data with substantially different lifetimes being mixed in a single erase block.

[0058] Figure 4A The example shown for sequential write operations illustrates an exemplary distribution of data within two erase blocks 440 on a third storage device 101c of multiple storage devices 101, as described above, in the case where no data classification is provided and data is written without any specific distribution strategy. In this scenario, after the lifetime of parity data 113 and / or log data 115 expires, the still valid user data 111 must be reassigned to another block before the corresponding erase block 440 can be erased and a new write operation can be performed within that erase block 440. Depending on various aspects, an erase block 440 may include more than one write block 430. In other words, each storage device in storage device 101 may have a first granularity for writing data (e.g., writing user data 111, parity data 113, and log data 115) and a second granularity, larger than the first granularity, for updating or deleting the written data.

[0059] Depending on various aspects, one or more processors 102 of the storage system 100 can be configured to write corresponding data in write block 430. Write block 430 has a predefined write block size. A corresponding erase block 440 has an erase block size for erasing or updating data. Depending on various aspects, the erase block size can be larger than the write block size.

[0060] Figure 4B The example shown for sequential write operations illustrates the distribution of data within two erase blocks 440 on a third storage device 101c out of multiple storage devices 101, as described above, provided that data classification 124 is provided and data is written according to a specific distribution strategy. The distribution strategy can be for data distribution on a single storage device within storage device 101. In this case, the distribution strategy could include, for example, storing all parity data 113 and all log data 115 into one erase block of erase block 440 via a first data stream 313a; and, for example, storing all user data 111 into another erase block of erase block 440 via a second data stream 313b. In this scenario, after the lifespan of parity data 113 and log data 115 expires (which can be expected before the expiration of user data 111), the erase block containing parity data 113 and log data 115 can be erased, and new write operations can be performed within the erase block 440. Meanwhile, the erase block 440 containing user data 111 can be retained, as user data 111 may still be valid.

[0061] In this scenario, depending on various aspects, the first data stream 313a may include a low ratio of user data 111 to parity data 113 and / or log data 115, for example, a ratio less than 15% (e.g., less than 10%, less than 5%, less than 1%). Depending on various aspects, the first data stream 313a may have no user data 111. Depending on various aspects, the second data stream 313b may include a high ratio of user data 111 to parity data 113 and / or log data 115, for example, a ratio greater than 70% (e.g., greater than 80%, greater than 90%, greater than 99%). Depending on various aspects, the second data stream 313b may have no parity data 113 and / or log data 115.

[0062] As described above, one or more processors 102 of storage system 100 can be configured to write user data 111, classified parity data 113, and classified log data 115 to corresponding storage devices 101a, 101b, and 101c via a first data stream 313a and a second data stream 313b. In other words, one or more processors 102 of storage system 100 can be configured to write these data to corresponding storage devices 101a, 101b, and 101c via the first data stream 313a according to a classification 124 provided for the parity data 113 and the log data 115. The first data stream 313a may include only the classified parity data 113 and the classified log data 115, and the second data stream may (e.g., only) include user data 111.

[0063] Figure 4C The example shown for sequential write operations illustrates an exemplary distribution of data within three erase blocks 440 on a third storage device 101c out of multiple storage devices 101, as described above, provided that data classification 124 is provided and data is written according to a specific distribution strategy. The distribution strategy can be for data distribution on a single storage device within storage device 101. In this case, the distribution strategy could include, for example, storing all parity data 113 into the first erase block 440 via a first data stream 313a; storing all user data 111 into the second erase block 440 via a second data stream 313b; and storing all log data 115 into the third erase block 440 via a third data stream 413c. In this scenario, after the lifespan of parity data 113 expires (which is expected to happen before the expiration of user data 111), the erase block 440 containing parity data 113 can be erased, and new write operations can be performed within this erase block 440. Simultaneously, the erase block 440 containing user data 111 can be retained, as user data 111 may still be valid. Similarly, after the lifespan of log data 115 expires (which is expected to happen before the expiration of user data 111), the erase block 440 containing log data 115 can be erased, and new write operations can be performed within this erase block 440. Simultaneously, the erase block 440 containing user data 111 can be retained, as user data 111 may still be valid.

[0064] In this scenario, depending on various aspects, the first data stream 313a may include a low ratio of user data 111 and log data 115 to parity data 113, for example, a ratio less than 15% (e.g., less than 10%, less than 5%, less than 1%). Depending on various aspects, the first data stream 313a may have no user data 111 and log data 115. Depending on various aspects, the second data stream 313b may include a high ratio of user data 111 to parity data 113 and log data 115, for example, a ratio greater than 70% (e.g., greater than 80%, greater than 90%, greater than 99%). Depending on various aspects, the second data stream 313b may have no parity data 113 and log data 115. Depending on various aspects, the third data stream 443c may include a high ratio of log data 115 to parity data 113 and user data 111, for example, a ratio greater than 85% (e.g., greater than 90%, greater than 95%, greater than 99%). Depending on various factors, the third data stream 413c may be without parity data 113 and user data 111.

[0065] As described above, one or more processors 102 of storage system 100 can be configured to write classified parity data 113 to corresponding storage devices 101a, 101b, and 101c via a first data stream 313a, to write user data 111 to corresponding storage devices 101a, 101b, and 101c via a second data stream 313b, and to write log data 115 to corresponding storage devices 101a, 101b, and 101c via a third data stream 413c. In other words, one or more processors 102 of storage system 100 can be configured to write these data to corresponding storage devices 101a, 101b, and 101c via two different data streams 313a and 413c according to a classification 124 provided for parity data 113 and log data 115.

[0066] Depending on various aspects, the log data 115 used for writing hole protection may include information about write operations on user data 111 and corresponding write operations on parity data 113 associated with user data 111, as described above.

[0067] Depending on various factors, storage system 100 can be host-based RAID.

[0068] As described above, the storage system may include a plurality of storage devices 101 and one or more processors 102, the one or more processors 102 being configured to distribute 114 user data, along with corresponding redundant data 113 (e.g., for data recovery) and corresponding log data 115, among the plurality of storage devices 101, generate at least a classification 124 associated with the redundant data 113 and the log data 115, and write the redundant data 113 and the log data 115 into different storage areas within each of the plurality of storage devices according to the classification.

[0069] Figure 5 A schematic flowchart of a method 500 for operating a storage system is shown, according to various aspects. Method 500 can be performed in a manner similar to that described above with respect to the configuration of storage system 100, and vice versa. According to various aspects, method 500 may include: in 510, distributing user data 111 along with redundant data 113 (e.g., for data recovery) and log data 115 (e.g., for writing hole protection) across multiple storage devices 101; in 520, generating a classification 124 associated with user data 111, redundant data 113, and log data 115 (e.g., providing classified user data, redundant data, and classified log data); and in 530, writing the (classified) user data 111, (classified) redundant data 113, and (classified) log data 115 into different storage regions (e.g., different storage regions 301a, 301b) within each of the multiple storage devices 101a, 101b, 101c, according to classification 124.

[0070] In a similar manner, only one of the redundant data 113 or the log data is classified. In this case, method 500 can be performed in a manner similar to that described above, including, for example: distributing user data 111 together with redundant data 113 (e.g., for data recovery) or log data 115 (e.g., for writing hole protection) across multiple storage devices 101; generating a classification 124 associated with user data 111 and with redundant data 113 or log data 115; and writing (classified) user data 111 to and (classified) redundant data 113 or (classified) log data 115 to different storage areas (e.g., different storage areas 301a, 301b) within each of the multiple storage devices 101a, 101b, 101c according to classification 124.

[0071] Depending on various factors, different storage areas can be different erase blocks as described above, see [link to relevant documentation]. Figure 3 and Figures 4A to 4C The storage area mentioned in this article can be the physical area of ​​the corresponding storage device.

[0072] Depending on various aspects, redundant data 113 can be, for example, parity data calculated using an XOR operation. However, any other redundant data 113 suitable for data recovery can be used in a similar manner (e.g., in the event of a failure of one of multiple storage devices). Depending on various aspects, log data 115 used for write hole protection can also be referred to as pre-written log data.

[0073] Figure 6 A schematic flowchart of a method 600 for operating a storage system is shown according to various aspects. Method 600 can be performed in a manner similar to that described above with respect to the configuration of storage system 100, and vice versa. According to various aspects, method 600 may include: in 610, dividing at least three storage devices 101a, 101b, 101c into stripes 203 and providing a plurality of stripes 201, each of the plurality of stripes 201 including at least three stripes 203; in 620, receiving user data 111 and distributing the user data 111 along with parity data 113 corresponding to the user data 111 along the stripes 201, such that each stripe 203 includes at least two user data stripes (each user data stripe including user data 111) and at least one parity stripe (parity data 113) associated with the at least two user data stripes. The parity data stripe includes parity data; and in 630, user data 111 and parity data 113 are written to one or more storage devices in the storage devices by a first data stream 313a directed to a first storage area 301a of the corresponding storage devices 101a, 101b, 101c and a second data stream 313b directed to a second storage area 301b of the corresponding storage devices 101a, 101b, 101c, which is different from the first storage area 301a, such that the first data stream 313a (e.g., only) includes parity data 113 and the second data stream (e.g., only) includes user data 111.

[0074] Depending on various aspects, the methods described herein, such as 500, 600, or storage system 100, can be used in Rapid Storage Technology Enterprise (Intel RSTe).

[0075] Depending on various aspects, RAID-5, RAID-6, or RAID-EC systems may include both parity classification and write-ahead log classification, as described herein. Classification can be used to control the physical data placement on the respective storage device 101. Physical data placement may include two or more streams for I / O operations.

[0076] Depending on various aspects, based on classification, one or more stream-guided strategies (also referred to as distribution strategies in this paper) can be provided, for example, for host-based RAID systems, where the RAID implementation (e.g., in software, hardware, or both) can place data for optimal durability and performance.

[0077] As described in more detail below, storage system 100 and the methods presented herein can avoid the problems that may occur with parity RAID.

[0078] One potential problem is reduced durability. To provide data reliability in parity RAID, parity data can be used in addition to user data. Partial stripe writes can result in parity updates being written once per stripe compared to data stripes. For an N-drive RAID system, parity stripes can be written up to N-1 times more than user data stripes. From the perspective of individual member drives (also called member storage devices), mixing shorter-lived parity stripe writes with longer-lived data stripe writes similarly leads to fragmentation within the member drives, which in turn reduces the overall durability of the member drives. Depending on various aspects, RAID systems can be used that are configured to place parity writes into their own erase units (also called erase blocks), which are automatically invalidated without causing any fragmentation. Alternatively, RAID systems can be configured to prevent write holes by writing an advance write log, which can be stored at the end of each member drive. The advance write log can be stored within a predefined LBA range, which the RAID system can write to in a circular buffer. RAID systems may overwrite the LBA range corresponding to daily data more frequently than the remaining LBAs on member drives. In some cases, daily writes can account for approximately 50% of total writes. On member drives, when this shorter-lived data (which is frequently overwritten) is mixed with longer-lived data, it can lead to fragmentation within the SSD. To defragment, the SSD must reposition valid data blocks (also known as garbage collection), which results in write amplification and reduced durability. Depending on various aspects, RAID systems like those described in this article can place daily writes into their own erase units, which are automatically invalidated without causing any fragmentation.

[0079] Another issue can be performance degradation. For example, NAND SSDs must be erased before new data can be placed in the same physical location (also called a storage area). Erase units (also referred to herein as erase blocks) can have much larger granularity than program (e.g., write) units (also referred to herein as write blocks), so all valid data must be moved to a new location before an erase block can be erased. Such relocation operations use the same SSD machinery that would otherwise be available to the host. This can significantly degrade the performance observed by the host. RAID systems, which reduce garbage collection, are provided for various reasons; and therefore, performance is significantly improved.

[0080] Another issue could be reduced I / O determinism. For example, when an SSD is undergoing its background garbage collection, the host can observe higher latency during these periods. "Tail" latency can be on the order of the average or 50th percentile. RAID systems offer significant improvements in tail latency for one or more workloads, depending on various factors.

[0081] As described above, to prevent the problems mentioned, the RAID engine can use knowledge about the generated parity and the lifespan of the log. The RAID engine can provide hints about the update frequency along with the data to the member drives of the RAID system (e.g., to SSDs). This can be applied to all RAID levels, such as those with parity (e.g., RAID-5, RAID-6, etc.).

[0082] Depending on various factors, there may be two master hint generators, or in other words, two classifiers: a parity classifier and an early write log classifier.

[0083] For sequential workloads, the parity classifier can be configured to assign parity data that is assumed to be invalid to a data stream with data that is updated very frequently (e.g., the first data stream 313a as described above). Valid parity data can be classified as the same stream as user data, depending on various factors.

[0084] For random workloads, the parity classifier can be configured to write each parity data point to a separate data stream (e.g., the first data stream 313a as described above). Statistically, parity data 113 can be updated more frequently than user data 111. The data stream assigned to parity data 113 can contain other data with similar update frequencies.

[0085] The pre-write log classifier can be configured to classify log data 115 that is updated very frequently.

[0086] Due to categorization, as described in this article, data is placed more efficiently on member drives. Efficient data placement reduces background activities on the drives, such as garbage collection or wear leveling processes.

[0087] The following section illustrates RAID technology for a RAID-5 configuration with four member drives 401, as shown in the diagram. Figures 7A to 7C As shown in the diagram. However, this RAID technique can be applied to other RAID levels in the same or similar manner.

[0088] Figure 7A The parity data 113 and user data 111 illustrate a RAID-5 data placement 700 for stripes 1 to 4 and member drives 1 to 4. Each member drive can be divided into stripes 203 of equal size. Figure 7A Each cell in the table shown represents one of stripes 203. Each row in the table represents stripe 201. In this case, each stripe in stripe 201 includes a parity stripe and multiple data stripes.

[0089] For sequential workloads, such as when a user writes 128kiB of data to LBA0 of a RAID-5 volume with a stripe size equal to 128kiB, the following process flow can be executed:

[0090] 1. Map LBA0 to drive 1

[0091] 2. Read data from drivers 2 and 3 to calculate parity (assuming, for example, using a different algorithm).

[0092] 3. Calculate parity (e.g., perform XOR on data from drive 2, drive 3, and drive 1).

[0093] 4. Write the data to drive 1 and write the parity check to drive 4.

[0094] The following section illustrates a write sequence for a user writing data sequentially to a RAID-5 volume. In this case, the RAID system will generate the following write requests to the member drives:

[0095] 1. Write data to drive 1, and write parity data to drive 4.

[0096] 2. Write data to drive 2, and write parity data to drive 4.

[0097] 3. Write the data to drive 3, and write the parity check to drive 4.

[0098] etc.

[0099] Parity will be written once and rewritten twice on drive 4. From the perspective of one drive (e.g., from the perspective of drive 4), data will be placed on NAND blocks for this sequential workload, as described in more detail below.

[0100] Figure 7B The diagram illustrates data placement 703 for drive 4. P1 represents parity from stripe 1, D2 represents data from stripe 2, and so on. Parity 713i can be rewritten very quickly and is no longer valid. However, when the firmware decides that a NAND block should be erased, some data must be moved to another NAND block before erasure. As mentioned above, this consumes time and increases write amplification. Ideally, the entire block should contain only invalid data, as data movement would not be necessary in this case.

[0101] Depending on various aspects, by allocating traffic to parity 713i units identified as potentially becoming invalid soon, NAND-based drivers can be able to place the parity 713i unit into a separate NAND block. Depending on various aspects, parity classifiers for sequential workloads, such as... Figure 7C The following diagram illustrates and describes the following:

[0102] 1. Assume that the parity checks are invalid (in the example above, this means that two first parity checks are written to P1, P1, P5, P5, P9, P9, etc.) and assign them to a different stream 705 than the user data (e.g., assign them to stream number 1, where all user data D2, D3, D4, D9, D7, D8, etc. are assigned to stream number 0).

[0103] 2. Assign the last written stripe (e.g., valid) parity block to the same stream as the user data (e.g., assign it to stream number 0).

[0104] Based on Stream 705, NAND-based drivers can be configured to place invalid parity blocks into one NAND block, but can place valid parity blocks and user data into another NAND block.

[0105] This RAID technology has the added benefit of transforming data writes and parity writes to member drives into sequential (at the physical layer) rather than random operations. This further improves WAF and performance.

[0106] For random workloads, data with two different lifetimes is generated: longer-lived user data 111 and shorter-lived parity data 113, as described herein. This is, for example, due to the fact that a parity block is "allocated" to multiple user data blocks. For example, for stripe 1, when data is written to drive 1, the parity on drive 4 must be updated. When data is written to drive 2 or drive 3 in stripe 1, the parity on drive 4 must also be updated. For RAID-5 systems or similar RAID systems, the parity will be updated up to N-1 times more frequently than the user data, where N represents the number of member drives.

[0107] Depending on various aspects, the parity classifier for random workloads can be configured to assign parity data 113 to a separate stream (e.g., different from the stream with user data 111), as described herein. Depending on various aspects, the stream assigned to parity data 113 can be used solely for random parity checking.

[0108] Depending on various factors, write-ahead classification can be provided. Parity-based RAID levels may suffer from, for example, a silent data corruption condition known as RAID write holes (RWH). A necessary condition for an RWH to occur is a drive failure followed by an unintended shutdown, and vice versa.

[0109] Figure 8A A computing system 800a is schematically illustrated according to various aspects. The computing system 800a may include a host system 802 and a storage system 100. The storage system 100 may be configured as described above. The host system 802 may be communicatively coupled to the storage system 100. The host system 802 may include one or more host processors configured to send user data to be stored to the storage system 100. The host system 802 may be configured to send I / O requests (see reference numerals 1 and 5) to a RAID controller 804 (e.g., to one or more processors of a RAID system, as described herein). As described above, the RAID controller 804 (also referred to as a RAID engine or RAID system) may control four member drives D1, D2, D3, and D4. However, the RAID controller may be configured to control any other number of member drives, e.g., three or more, as described herein. Depending on various aspects, the host system 802 may be part of the storage system 100. Alternatively, host system 802 and storage system 100 can be configured remotely, wherein host system 802 communicates with remote storage system 100 via any suitable communication network.

[0110] also, Figure 8AThe diagram shows reads 2a, 2b and writes 4a, 4b to member drives D1 and D4, including parity calculation 3 (e.g., controlled by RAID controller 804).

[0111] In the event of an unexpected shutdown or power failure of the host system 802 during the write step, data corruption may occur, causing the first write operation 4a (e.g., writing user data to drive 1, D1) to complete but the corresponding second write operation 4b (e.g., writing parity data to drive 4, D4) to fail. This will result in inconsistent stripes, meaning that an XOR operation on any three stripes will not produce a fourth stripe, and parity will be ensured in other ways. Upon system restart, if, for example, bystander drive 2 (D2) fails, a stripe reconstructed from D2 cannot be provided because the XOR calculations on D1 and D4 (new and old) will result in garbage data due to data corruption.

[0112] To close RAID write holes, write-ahead logging (WAL) can be used on each RAID-5 write based on the log data, as described in this article. The write-ahead log data can be placed in the metadata section at the end of the member drive. The log can be stored on the member drive storing parity. Because the write-ahead log area at the end of each member drive can be a very small LBA range, this small LBA range experiences frequent writes as the RAID controller 804 (e.g., RSTe) writes the write-ahead log in a circular buffer before performing write operations 4a, 4b. From the member drive's perspective, the write-ahead log's lifetime is much shorter than user data and / or parity. Furthermore, in some applications, the write-ahead log can account for more than 50% of the total GB of writes.

[0113] Figure 8BA computing system 800b is illustrated schematically according to various aspects. The computing system 800b may include a RAID controller 804 and three or more member drives 101. The RAID controller 804 may be configured as described above. The computing system 800b may include one or more processors 802p configured to run an operating system (OS) and one or more applications communicating with the RAID controller 804. Alternatively, the one or more processors 802p of the computing system 800b may be configured (e.g., remotely) to communicate with an external system running an operating system (OS) and one or more applications communicating with the RAID controller 804. The computing system 800b may be, for example, a server (e.g., as part of a cloud system), a remote storage server, a desktop computer, a laptop computer, a tablet computer, etc. The one or more processors 802p of the computing system 800b may be configured to perform one or more functions of the RAID controller 804. The RAID controller 804 may be implemented as software executed by the RAID controller 804. Alternatively, the RAID controller 804 may be implemented as software executed by the RAID controller 804 and by one or more additional processors (e.g., one or more processors 102 as described above). The computing system 800b may include one or more interfaces 802i for sending and receiving data requests associated with data stored or to be stored on member drives 101 via the RAID controller 804.

[0114] Figure 8C A schematic flowchart of a method 800c for operating a storage system is shown, according to various aspects. Method 800c can be performed in a manner similar to that described above with respect to the configuration of storage system 100, and vice versa. According to various aspects, method 800c may include: in 810, operating multiple storage devices as an independent drive redundant array; in 820, receiving user data and storing the received user data on multiple storage devices in a striped configuration; in 830, calculating parity data associated with the received user data (e.g., for data recovery); and in 840, generating at least a parity classification associated with the parity data, which specifies the physical data placement of the parity data and user data on each of the multiple storage devices according to the parity classification.

[0115] Depending on various factors, separating sequential or random I / O at different speeds leads to a reduction in WAF. An advance write log classifier can be configured to assign write log data 115 to separate streams (see example...). Figure 4C This generates WAF benefits.

[0116] Figure 9A and Figure 9B WAF measurements for various classification types are shown. All measurements were performed using a RAID-5 system, which was created using three NVMe drives (Fast NVM (NVMe), also known as the Non-Volatile Memory Host Controller Interface Specification (NVMHCI), which describes a logical device interface specification for accessing non-volatile storage media attached via a Fast PCI (PCIe) bus).

[0117] One of the factors measured is the write amplification factor (WAF), calculated as follows:

[0118]

[0119] A lower WAF value indicates less writing to the memory device (e.g., flash memory). This is associated with higher drive durability and better performance.

[0120] Figure 9A The WAF 900y relative time 900x for a sequential write 900s (e.g., 128kiB) via a RAID-5 system with RAID write hole protection is shown. With both the parity classifier and the early write log classifier enabled (e.g., as depicted by curve 902c), the WAF is approximately twice as low as without classification (e.g., as depicted by curve 902). With only the early write log classifier enabled (e.g., as depicted by curve 902w), the WAF is also low as without classification (e.g., as depicted by curve 902).

[0121] Figure 9B The diagram illustrates the WAF 900y relative time 900x of a random write 900r (e.g., 4kiB) performed via a RAID-5 system with RAID write hole protection, for example using a random parity classifier as described above. With both an early write log classifier and a parity classifier used, the WAF benefit can be approximately 12%, for example, as depicted by curve 904c. With only one of the early write log classifier (e.g., as depicted by curve 904w) or parity classifier (e.g., as depicted by curve 904p) used, the WAF can still be lower than with no classifier (e.g., as depicted by curve 904).

[0122] Another factor being measured is latency (not shown in the figure), such as average latency and latency for quality of service (QoS).

[0123] Based on various factors, the average latency can be improved compared to unclassified latency as follows:

[0124] 128kiB Sequential Write 4kiB random write Parity check classification 104.0% 89.5% Pre-writing log categories 100.7% 94.9% Parity checking and early write log classification 45.5% 83.4%

[0125] As shown above, for sequential workloads and two active classifiers, the average latency is more than twice as low as without classification. For random workloads, it is approximately 17% lower than without classification. Therefore, as discussed in this article, users can access data in RAID volumes much faster with classification.

[0126] Based on various factors, the average latency service quality (99.99%) can be improved compared to no classification as follows:

[0127] 128kiB Sequential Write 4kiB random write Parity checking and early write log classification 40.4% 97.8%

[0128] As shown above, for sequential writes, the maximum RAID volume response time will be more than twice as low as without a classifier (within 99.99% of the total usage time).

[0129] Another factor measured is throughput in the classified case compared to no classification, as described in this paper.

[0130] Based on various factors, the throughput can be improved compared to unclassified throughput as follows:

[0131] 128kiB Sequential Write 4kiB random write Parity checking and early write log classification 221.2% 115.2%

[0132] As shown above, the parity classifier and the write-ahead log classifier allow users to access data more than twice as fast as without classification. For random workloads, it's about 15% faster.

[0133] Depending on various aspects, NVMe / SATA protocol analyzers can be used to analyze write commands to member drives. Parity writes can be identified by analyzing LBA or update frequency. Stream identifiers can be read from write command frames.

[0134] Depending on various aspects, storage devices (e.g., SSDs) may have one or more implementations of stream booting and may be paired with the RAID system described herein.

[0135] According to various aspects, the storage system 100 described herein (or in other words, the RAID system, RAID engine, RAID controller, etc.) can provide improved durability, improved performance, improved QoS, and potential power savings due to minimized garbage collection.

[0136] Depending on various aspects, log data 115 as described herein may be part of a write-ahead logging (WAL) implemented in one or more processors 102 of storage system 100. Write-ahead logging may include writing all modifications to user data 111 and / or parity data 113 to the log before applying the modifications. Depending on various aspects, log data 115 may include redo and undo information. Using WAL allows modifications to be applied in-situ, which can reduce the extent to which indexes and / or block lists are modified.

[0137] Various examples are provided below with reference to the above aspects.

[0138] Example 1 is a storage system. This storage system may include multiple storage devices and one or more processors configured to distribute user data, along with redundant data and log data, across the multiple storage devices to generate classifications associated with the user data, redundant data, and log data, and to write the user data, redundant data, and log data to different storage areas within each of the multiple storage devices according to their respective classifications.

[0139] In Example 2, the subject of Example 1 may optionally include: redundant data including parity data.

[0140] In Example 3, the subject of either Example 1 or 2 may optionally include: log data including write hole protection data.

[0141] In Example 4, the subject of any of Examples 1 to 3 may optionally include: each of the plurality of storage devices includes non-volatile memory, and the non-volatile memory is configured to store user data, redundant data, and log data.

[0142] In Example 5, the subject of any one of Examples 1 to 4 may optionally include: each of a plurality of storage devices includes a first granularity associated with written data, and a second granularity greater than the first granularity associated with updated or deleted data.

[0143] In Example 6, the subject of any of Examples 1 through 5 may optionally include: each of the multiple storage devices includes a solid-state drive.

[0144] In Example 7, the subject of any of Examples 1 through 6 may optionally include: one or more processors are configured to distribute redundant data across multiple storage devices.

[0145] In Example 8, the subject of any of Examples 1 through 7 may optionally include: one or more processors are configured to distribute log data across multiple storage devices.

[0146] In Example 9, the subject of any one of Examples 1 to 8 may optionally include: one or more processors being configured to write at least one of redundant data or log data to a specific storage device among a plurality of storage devices via a first data stream, and to write user data to the specific storage device via a second data stream.

[0147] In Example 10, the subject of any one of Examples 1 to 8 may optionally include: one or more processors being configured to write redundant data to a specific storage device among a plurality of storage devices via a first data stream, write user data to a specific storage device via a second data stream, and write log data to a specific storage device via a first data stream or a third data stream.

[0148] In Example 11, the subject of Example 10 may optionally include: the third data stream does not actually contain user data.

[0149] In Example 12, the subject of either Example 10 or 11 may optionally include: the third data stream has no redundant data.

[0150] In Example 13, the subject of any of Examples 9 through 12 may optionally include: the first data stream does not actually contain user data.

[0151] In Example 14, the subject of any of Examples 9 through 13 may optionally include: the first data stream does not actually contain log data.

[0152] In Example 15, the subject of any of Examples 9 through 14 may optionally include: the second data stream has no redundant data.

[0153] In Example 16, the subject of any of Examples 9 through 15 may optionally include: the second data stream does not actually contain log data.

[0154] In Example 17, the subject matter of any one of Examples 9 to 16 may optionally include: a first data stream being directed to a first storage region of a respective storage device among a plurality of storage devices, and a second data stream being directed to a second storage region of a respective storage device among a plurality of storage devices, the second storage region being different from the first storage region.

[0155] In Example 18, the subject matter of any one of Examples 10 to 12 may optionally include: a first data stream being directed to a first storage area of ​​a respective storage device among a plurality of storage devices, a second data stream being directed to a second storage area of ​​a respective storage device among a plurality of storage devices, the second storage area being different from the first storage area, and a third data stream being directed to an additional storage area of ​​a respective storage device among a plurality of storage devices, the additional storage area being different from the first and second storage areas.

[0156] In Example 19, the subject of any of Examples 1 through 18 may optionally include: one or more processors are configured to write redundant data, log data, and user data as blocks of write block size.

[0157] In Example 20, the subject of Example 19 may optionally include: one or more of a plurality of storage devices include an erase block size associated with erasing data, and the erase block size is larger than the write block size.

[0158] In Example 21, the subject of any of Examples 1 through 20 may optionally include: log data including information about write operations to user data and write operations to redundant data associated with user data.

[0159] In Example 22, the subject of Example 21 may optionally include: one or more processors are configured to write user data and redundant data associated with the log data after the log data has been written for a write operation.

[0160] In Example 23, the subject of any of Examples 1 through 22 may optionally include: one or more processors are configured to write log data or redundant data into a circular buffer.

[0161] Example 24 is a computing system. This computing system may include a storage system of any of Examples 1 through 23. The computing system may also include a host system communicatively coupled to the storage system, the host system including one or more host processors configured to send user data to be stored to the storage system.

[0162] Example 25 is a method for operating a storage system. The method may include distributing user data, along with redundant data and log data, across multiple storage devices; generating categories associated with the user data, redundant data, and log data; and writing the user data, redundant data, and log data to different storage areas within each of the multiple storage devices according to the categories.

[0163] In Example 26, the subject of Example 25 may optionally include: redundant data including parity data.

[0164] In Example 27, the subject of either Example 25 or 26 may optionally include: log data includes write hole protection data.

[0165] In Example 28, the subject of any of Examples 25 through 27 may optionally include: writing redundant data across multiple storage devices.

[0166] In Example 29, the subject of any of Examples 25 through 28 may optionally include: writing log data across multiple storage devices.

[0167] In Example 30, the subject matter of any one of Examples 25 to 29 may optionally include: writing at least one of redundant data or log data to a specific storage device among a plurality of storage devices via a first data stream, and writing user data to the specific storage device via a second data stream.

[0168] In Example 31, the subject matter of any one of Examples 25 to 29 may optionally include: writing redundant data to a specific storage device among a plurality of storage devices via a first data stream, writing user data to a specific storage device via a second data stream, and writing log data to a specific storage device via a first data stream or a third data stream.

[0169] In Example 32, the subject of Example 31 may optionally include: the third data stream does not actually contain user data.

[0170] In Example 33, the subject of either Example 31 or 32 may optionally include: the third data stream has no redundant data.

[0171] In Example 34, the subject of any of Examples 30 to 33 may optionally include: the first data stream does not actually contain user data.

[0172] In Example 35, the subject of any of Examples 30 to 34 may optionally include: the first data stream does not actually contain log data.

[0173] In Example 36, the subject of any of Examples 30 to 35 may optionally include: the second data stream has no redundant data.

[0174] In Example 37, the subject of any of Examples 30 through 36 may optionally include: the second data stream does not actually contain log data.

[0175] In Example 38, the subject matter of any one of Examples 30 to 37 may optionally include: the method further includes directing a first data stream to a first storage area of ​​the corresponding storage device and directing a second data stream to a second storage area of ​​the corresponding storage device that is different from the first storage area.

[0176] In Example 39, the subject matter of any one of Examples 31 to 33 may optionally include: the method further includes directing a first data stream to a first storage area of ​​the corresponding storage device, directing a second data stream to a second storage area of ​​the corresponding storage device that is different from the first storage area, and directing a third data stream to an additional storage area of ​​the corresponding storage device that is different from the first and second storage areas.

[0177] In Example 40, the subject of any of Examples 25 through 39 may optionally include: writing redundant data and log data at a write block size smaller than the minimum erase block size.

[0178] In Example 41, the subject of any of Examples 25 to 40 may optionally include: log data including information about write operations to user data and write operations to redundant data associated with user data.

[0179] In Example 42, the subject of Example 41 may optionally include: the method further includes writing log data associated with the write operation before the user data and the redundant data associated therewith.

[0180] Example 43 is a storage system. The storage system may include one or more processors configured to: divide at least three storage devices into stripes, providing a plurality of stripes, each of the plurality of stripes including at least three stripes; receive user data, such that the user data, together with parity data associated with the user data, is distributed along the plurality of stripes, such that each of the plurality of stripes includes at least two data stripes and at least one parity stripe associated with the at least two data stripes; and write the user data and parity data to each of the at least three storage devices based on a first data stream directed to a first storage region of a corresponding storage device in the at least three storage devices and a second data stream directed to a second storage region of a corresponding storage device in the at least three storage devices. The second storage region is different from the first storage region. The first data stream includes parity data, and the second data stream includes user data.

[0181] In Example 44, the subject of Example 43 may optionally include: at least two data stripes comprising user data, and the at least one parity strip comprising parity data.

[0182] In Example 45, the subject of Example 44 may optionally include: one or more processors are also configured to write log data associated with the writing of user data and parity data to at least three storage devices.

[0183] In Example 46, the subject of Example 45 may optionally include: log data including write hole protection data.

[0184] In Example 47, the subject of any of Examples 45 or 46 may optionally include: one or more processors are configured to distribute log data across at least three storage devices.

[0185] In Example 48, the subject matter of any one of Examples 45 through 47 may optionally include: one or more processors are further configured to write log data to each of the at least three storage devices based on a first data stream directed to a first storage region or a third data stream directed to a third storage region corresponding to one of the at least three storage devices. The third storage region is different from the first and second storage regions.

[0186] In Example 49, the subject of Example 48 may optionally include a third data stream that does not actually contain user data.

[0187] In Example 50, the subject of any of Examples 48 or 49 may optionally include: the third data stream does not actually contain parity data.

[0188] In Example 51, the subject of any of Examples 43 to 50 may optionally include: the first data stream does not actually contain user data.

[0189] In Example 52, the subject of any of Examples 43 through 51 may optionally include: the first data stream does not actually contain log data.

[0190] In Example 53, the subject of any of Examples 43 through 52 may optionally include: the second data stream does not actually contain parity data.

[0191] In Example 54, the subject of any of Examples 43 to 53 may optionally include: the second data stream does not actually contain log data.

[0192] In Example 55, the subject matter of any of Examples 43 to 54 may optionally include: each of the at least three storage devices includes a non-volatile memory configured to store user data and parity data.

[0193] In Example 56, the subject matter of any one of Examples 43 to 55 may optionally include: each of at least three storage devices having a first granularity associated with written data and a second granularity, larger than the first granularity, associated with updated or deleted data.

[0194] In Example 57, the subject of any of Examples 43 through 56 may optionally include: each of at least three storage devices includes a solid-state drive.

[0195] In Example 58, the subject of any of Examples 43 through 57 may optionally include: one or more processors are configured to distribute parity data across at least three storage devices.

[0196] In Example 59, the subject matter of any one of Examples 43 to 58 may optionally include: one or more processors are configured to write user data and parity data respectively with corresponding write block sizes, and erase user data and parity data respectively with corresponding erase block sizes, wherein the erase block size is larger than the write block size.

[0197] In Example 60, the subject of any of Examples 43 to 59 may optionally include: the log data includes information about write operations on user data and write operations on parity data associated with the user data.

[0198] In Example 61, the subject of Example 60 may optionally include: one or more processors are configured to write user data and redundant data associated with the log data after writing log data for a write operation.

[0199] In Example 62, the subject of any of Examples 43 through 61 may optionally include: one or more processors are configured to write log data or redundant data into a circular buffer.

[0200] Example 63 is a method for operating a storage system. The method may include dividing at least three storage devices into stripes and providing a plurality of stripes, each of the plurality of stripes including at least three stripes; receiving user data and distributing the user data along the stripes together with parity data associated with the user data, such that each stripe includes at least two data stripes and at least one parity stripe associated with the at least two data stripes; and writing the user data and parity data into one or more storage devices via a first data stream directed to a first storage region of the respective storage device and a second data stream directed to a second storage region of the respective storage device that is different from the first storage region, such that the first data stream includes parity data and the second data stream includes user data.

[0201] In Example 64, the subject of Example 63 may optionally include: the method further includes writing log data to each of at least three storage devices via a first data stream directed to a first storage area or via a third data stream directed to a third storage area different from the first and second storage areas.

[0202] In Example 65, the subject of Example 64 may optionally include: the third data stream consists only of log data.

[0203] In Example 66, the subject of any of Examples 63 through 65 may optionally include: the second data stream includes only user data.

[0204] Example 67 is a non-transitory computer-readable medium that stores instructions which, when executed by a processor, cause the processor to perform a method according to any one of Examples 25 to 42 and any one of Examples 63 to 66.

[0205] Example 68 is a storage system. The storage system may include multiple storage devices and one or more processors configured to operate the multiple storage devices as a redundant array of independent drives, receive user data and store the received user data on the multiple storage devices in a striped configuration, calculate parity data corresponding to the received user data, and generate at least a parity classification associated with the parity data, the parity classification corresponding to the physical data placement of the parity data and user data on each of the multiple storage devices according to the parity classification.

[0206] In Example 69, the subject of Example 68 may optionally include: one or more processors configured to provide a first stream and a second stream to write parity data and user data to each of a plurality of storage devices according to parity classification. The first stream includes parity data, and the second stream includes user data.

[0207] In Example 70, the subject of any of Examples 68 or 69 may optionally include: one or more processors are further configured to provide log data associated with received user data and computed parity data, and to generate at least a log classification associated with the log data, which corresponds to the physical data placement of the log data on each of a plurality of storage devices according to the log data classification.

[0208] In Example 71, the subject of Example 70 may optionally include: one or more processors configured to provide a first stream, a second stream, and a third stream to write parity data, user data, and log data to each of a plurality of storage devices according to parity classification and log data classification. The first stream includes parity data, the second stream includes user data, and the third stream includes log data.

[0209] In Example 72, the subject matter of any one of Examples 68 to 71 may optionally include: one or more processors are configured to provide a first stream and a second stream to write parity data, user data, and log data to each of a plurality of storage devices according to parity classification and log data classification. The first stream includes parity data and log data, and the second stream includes user data.

[0210] In Example 73, the subject of any of Examples 68 through 72 may optionally include: each of a plurality of storage devices includes a solid-state drive.

[0211] In Example 74, the subject of any of Examples 68 through 73 may optionally include: each erase block includes multiple write blocks associated with writing data to the erase block.

[0212] In Example 75, the subject of any of Examples 68 to 74 may optionally include: each of the plurality of storage devices includes an erase block size and a write block size, wherein the erase block size is greater than the write block size.

[0213] In Example 76, the subject of any one of Examples 68 to 75 may optionally include: each of a plurality of storage devices includes a first granularity associated with written data and a second granularity, which is larger than the first granularity, associated with updated or deleted data.

[0214] In Example 77, the subject of any of Examples 68 to 76 may optionally include: one or more processors are configured to distribute parity data across multiple storage devices.

[0215] Example 78 is a method for operating a storage system. The method may include operating multiple storage devices as a redundant array of independent drives, receiving user data and storing the received user data on the multiple storage devices in a striped configuration, calculating parity data corresponding to the received user data, and generating a parity classification at least associated with the parity data, the parity classification corresponding to the physical data placement of the parity data and user data on each of the multiple storage devices according to the parity classification.

[0216] In Example 79, the subject of Example 78 may optionally include: the method further includes writing parity data to each of the plurality of storage devices via a first stream, and writing user data to each of the plurality of storage devices via a second stream according to parity classification.

[0217] In Example 80, the subject matter of any of Examples 78 or 79 may optionally include: the method further includes providing log data associated with the received user data and the generated parity data, and generating at least a log classification associated with the log data, the log classification corresponding to the physical data placement of the log data on each of a plurality of storage devices according to the log data classification.

[0218] In Example 81, the subject of Example 80 may optionally include: the method further includes writing parity data to each of the plurality of storage devices via a first stream, writing user data to each of the plurality of storage devices via a second stream according to parity classification, and writing log data to each of the plurality of storage devices via a third stream.

[0219] Example 82 is a storage system. The storage system may include multiple storage devices and one or more processors configured to receive user data, generate auxiliary data associated with the user data, generate classifications associated with the user data and the auxiliary data, and distribute the user data and auxiliary data across the multiple storage devices such that the user data and auxiliary data are stored in different storage areas within each of the multiple storage devices according to their respective classifications.

[0220] In Example 83, the subject of Example 82 may optionally include: auxiliary data including redundant data or log data.

[0221] In Example 84, the subject of Example 82 may optionally include: auxiliary data including redundant data and log data.

[0222] In Example 85, the subject of any of Examples 83 or 84 may optionally include: redundant data includes parity data.

[0223] In Example 86, the subject of any of Examples 83 through 85 may optionally include: log data including write hole protection data.

[0224] While this disclosure has been specifically shown and described with reference to particular aspects, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of this disclosure as defined by the appended claims. Therefore, the scope of this disclosure is indicated by the appended claims and is thus intended to cover all changes falling within the meaning and scope of equivalents of the claims.

Claims

1. A storage system, comprising: One or more processors, which are configured to Stripe at least three storage devices. Multiple stripes are provided across the at least three storage devices, each of the multiple stripes comprising at least three bands. Receive user data, The user data, along with parity data associated with the user data, is distributed along the plurality of stripes, such that each of the plurality of stripes includes at least two data stripes and at least one parity stripe associated with the at least two data stripes. Based on a first data stream directed to a first storage region of a corresponding storage device among the at least three storage devices and a second data stream directed to a second storage region of a corresponding storage device among the at least three storage devices, the user data and the parity data are written to each of the at least three storage devices. The second storage area is different from the first storage area. The first data stream includes the parity data, and The second data stream includes the user data.

2. The storage system according to claim 1, The at least two data stripes include user data, and the at least one parity stripe includes parity data.

3. The storage system according to claim 2, in, The one or more processors are also configured to write log data associated with writes to the user data and the parity data to the at least three storage devices.

4. The storage system according to claim 3, The log data includes data written to the hole protection database.

5. The storage system according to claim 3, in, The one or more processors are configured to distribute the log data across the at least three storage devices.

6. The storage system according to claim 3, in, The one or more processors are also configured to Based on a third data stream directed to a third storage region of one of the at least three storage devices, the log data is written to each of the at least three storage devices. The third storage area is different from the first storage area and the second storage area.

7. The storage system according to claim 1, The first data stream does not actually contain the user data.

8. The storage system according to claim 1, The second data stream does not actually contain the parity check data.

9. The storage system according to claim 1, in, Each of the at least three storage devices has a first granularity associated with written data and a second granularity, which is larger than the first granularity, associated with updated or deleted data.

10. The storage system according to claim 1, in, Each of the at least three storage devices includes a solid-state drive.

11. The storage system according to claim 1, in, The one or more processors are configured to write the user data and the parity data respectively, with corresponding write block sizes. The user data and the parity data are erased respectively according to the corresponding erase block size, wherein the erase block size is larger than the write block size.

12. A storage system, comprising: Multiple storage devices; as well as One or more processors, which are configured to Receive user data, Generate auxiliary data associated with the user data. Generate classifications associated with the user data and the auxiliary data, and The user data and the auxiliary data are distributed across the multiple storage devices, such that the user data and the auxiliary data are stored in different storage areas within each of the multiple storage devices according to their respective classifications.

13. The storage system according to claim 12, The auxiliary data includes redundant data or log data.

14. The storage system according to claim 12, The auxiliary data includes redundant data and log data.

15. The storage system according to claim 13, The redundant data includes parity check data.

16. The storage system according to claim 13 or 14, The log data includes data written to the hole protection database.

17. A method for operating a storage system, the method comprising: Multiple storage devices are operated as a redundant array of independent drives. Receive user data and store the received user data on the plurality of storage devices configured in a striped manner. Calculate the parity check data corresponding to the received user data. Generate at least one parity classification associated with the parity data, the parity classification corresponding to the physical data placement of the parity data and the user data on each of the plurality of storage devices according to the parity classification. The parity data is written to each of the plurality of storage devices via a first stream, and According to the parity classification, the user data is written to each of the plurality of storage devices via a second stream.

18. A method for operating a storage system, the method comprising: Multiple storage devices are operated as a redundant array of independent drives. Receive user data and store the received user data on the plurality of storage devices configured in a striped manner. Calculate the parity check data corresponding to the received user data. Generate at least one parity classification associated with the parity data, the parity classification corresponding to the physical data placement of the parity data and the user data on each of the plurality of storage devices according to the parity classification. Provide log data associated with the received user data and the generated parity data, and Generate at least one log category associated with the log data, the log category corresponding to the physical data placement of the log data on each of the plurality of storage devices according to the log data category.

19. The method of claim 18, further comprising: The parity data is written to each of the plurality of storage devices via a first stream. Based on the parity check classification, the user data is written to each of the plurality of storage devices via a second stream. The log data is written to each of the plurality of storage devices via the first stream or the third stream.

Citation Information

Patent Citations

  • System and Method of Write Hole Protection for a Multiple-Node Storage Cluster

    US20150135006A1