On-ssd erasure coding with unidirectional commands
By offloading EC computation to SSD member drives and utilizing one-way commands and a specific NVMe command set, the high data transfer and computation overhead in existing storage systems is addressed, thereby improving the scalability and performance of the storage system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTEL CORP
- Filing Date
- 2020-02-20
- Publication Date
- 2026-04-28
AI Technical Summary
Existing EC-based storage systems require a large amount of data transfer and computation during random write operations, resulting in high interface bandwidth consumption, which limits storage performance. Furthermore, bidirectional commands are difficult to implement in scalable storage systems.
By employing one-way commands and EC offload technology, complex mathematical operations are offloaded to SSD member drives, reducing network traffic and computing overhead on the centralized storage controller. Distributed processing of EC computation is achieved through WriteAndXor, SaveEC, and ECwrite commands.
It reduces the amount of data transmission and computational complexity, improves the scalability and performance of the storage system, and reduces the burden on the centralized storage controller.
Smart Images

Figure CN115428074B_ABST
Abstract
Description
Technical Field
[0001] The embodiments generally relate to electronic storage systems. More specifically, the embodiments relate to solid-state drives (SSDs) erasure coding utilizing one-way commands. Background Technology
[0002] Some storage systems use erasure coding (EC) technology and corresponding data layouts for volume member drives to improve data reliability and durability. Examples of EC systems include redundant arrays of independent disks (RAID) systems, such as RAID5 and RAID6, which have one or two additional drives in the volume, and M+P (e.g., 6+3) EC utilized in some cloud storage systems for the purpose of paving redundant data. More generally, in an M+P EC configuration, M+P drives are used to encode data that would otherwise be intended to be stored on M drives by using P additional drives. The data stored across drives includes parity data for the P drives, which is used to recover the system from the failure of up to P drives. Attached Figure Description
[0003] The materials described herein are illustrated in the accompanying drawings by way of example rather than limitation. For the sake of simplicity and clarity, the elements shown in the drawings are not necessarily drawn to scale. For example, the scale of some elements may be exaggerated relative to others for clarity. Furthermore, reference numerals are repeated in the drawings where deemed appropriate to indicate corresponding or similar elements. In the drawings:
[0004] Figure 1 This is a block diagram of an example of an electronic storage device according to an embodiment;
[0005] Figure 2 This is a block diagram of an example electronic device according to an embodiment;
[0006] Figures 3A to 3E This is a flowchart illustrating an example of a method for controlling a storage device according to an embodiment;
[0007] Figure 4A This is a block diagram of an example electronic storage system according to an embodiment;
[0008] Figure 4B This is a table of examples of EC controller operation and corresponding drive operation according to the embodiments;
[0009] Figure 5 This is a block diagram of another example of an electronic storage system according to an embodiment;
[0010] Figure 6 This is a block diagram of another example of an electronic storage system according to an embodiment;
[0011] Figure 7 This is an illustrative diagram illustrating an example of a command format according to an embodiment;
[0012] Figure 8 This is a block diagram of an example computing system according to an embodiment; and
[0013] Figure 9 This is a block diagram of an example of a solid-state drive (SSD) according to an embodiment. Detailed Implementation
[0014] One or more embodiments or implementations are now described with reference to the accompanying drawings. While specific configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of this specification. It will be apparent to those skilled in the art that the techniques and / or arrangements described herein can also be used in various other systems and applications besides those described herein.
[0015] While the following description illustrates various implementations that can be embodied in architectures such as System-on-a-Chip (SoC) architectures, the implementations of the technologies and / or arrangements described herein are not limited to any particular architecture and / or computing system and can be implemented by any architecture and / or computing system for similar purposes. For example, various architectures employing, for instance, multiple integrated circuit (IC) chips and / or packages and / or various computing devices and / or consumer electronics (CE) devices (such as set-top boxes, smartphones) can implement the technologies and / or arrangements described herein. Furthermore, while the following description may elaborate on many specific details, such as the logical implementation, type, and interrelationships of system components, logical partitioning / integration choices, etc., the claimed subject matter can be practiced without such specific details. In other cases, some materials (such as, for example, control structures and complete sequences of software instructions) may not be shown in detail to avoid obscuring the materials disclosed herein.
[0016] The materials disclosed herein can be implemented in hardware, firmware, software, or any combination thereof. The materials disclosed herein can also be implemented as instructions stored on a machine-readable medium, which can be read and executed by one or more processors. A machine-readable medium can include any medium and / or mechanism for storing or transmitting information in a machine-readable (e.g., computing device) form. For example, a machine-readable medium can include read-only memory (ROM); random access memory (RAM); disk storage media; optical storage media; flash memory devices; electrical, optical, acoustic, or other forms of propagation signals (e.g., carrier waves, infrared signals, digital signals, etc.).
[0017] References to phrases such as "an implementation," "an implementation method," and "an exemplary implementation method" in the specification indicate that the described implementation may include specific features, structures, or characteristics, but each embodiment may not necessarily include all such features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same implementation method. Additionally, when a specific feature, structure, or characteristic is described in connection with an embodiment, it is believed that implementing such a feature, structure, or characteristic in conjunction with other implementation methods, whether or not explicitly described herein, is within the knowledge of those skilled in the art.
[0018] The various embodiments described herein may include memory components and / or interfaces to memory components. Such memory components may include volatile and / or non-volatile (NV) memory. Volatile memory can be a storage medium that requires power to maintain the state of data stored by the medium. Non-limiting examples of volatile memory may include various types of random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). One particular type of DRAM that can be used in a memory module is synchronous dynamic RAM (SDRAM). In certain embodiments, the DRAM of the memory component may conform to standards issued by the Joint Electronic Equipment Committee (JEDEC), such as JESD79F for Double Data Rate (DDR) SDRAM, JESD79-2F for DDR2 SDRAM, JESD79-3F for DDR3 SDRAM, JESD79-4A for DDR4 SDRAM, JESD209 for Low Power DDR (LPDDR), JESD209-2 for LPDDR2, JESD209-3 for LPDDR3, and JESD209-4 for LPDDR4 (these standards are available at jedec.org). Such standards (and similar standards) may be referred to as DDR-based standards, and the communication interface of a memory device implementing such a standard may be referred to as a DDR-based interface.
[0019] NV memory (NVM) can be a storage medium that does not require power to maintain the state of data stored by the medium. In one embodiment, the memory device may include block-addressable memory devices, such as those based on NAND or NOR technologies. The memory device may also include next-generation non-volatile devices, such as three-dimensional (3D) cross-point memory devices, or other byte-addressable, write-in-place non-volatile memory devices. In one embodiment, a memory device may be or may include memory devices using the following: chalcogenide glass, multi-threshold level NAND flash memory, NOR flash memory, single-level or multi-level phase-change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor RAM (FeTRAM), antiferroelectric memory, magnetoresistive RAM (MRAM) memory employing memristor technology, resistive memory including metal oxide-based, oxygen vacancy-based, and conductive bridge RAM (CB-RAM), or spin-transfer torque (STT)-MRAM, devices based on spintronic magnetic junction memory, devices based on magnetic tunnel junction (MTJ), devices based on DW (domain wall) and SOT (spin-orbit transfer), thyristor-based memory devices, or any combination thereof described above, or other memories. A memory device may refer to a memory product that is the die itself and / or packaged. In certain embodiments, memory components having non-volatile memory may conform to one or more standards issued by JEDEC, such as JESD218, JESD219, JESD220-1, JESD223B, JESD223-1, or other suitable standards (the JEDEC standards cited herein are available at jedec.org).
[0020] refer to Figure 1Embodiments of the electronic storage device 10 may include a persistent storage medium 12 and a device controller 11 communicatively coupled to the persistent storage medium 12. The device controller 11 may include a logic unit 13 (e.g., a command logic unit, control logic unit, device logic unit, etc. for the device 10 itself) for controlling local access to the persistent storage medium 12. For example, in response to one or more commands, the logic unit 13 may be configured to determine an intermediate parity value based on a first local parity calculation and to locally store the intermediate parity value. For example, the logic unit 13 may also be configured to determine a final parity value based on the intermediate parity value and a second local parity calculation in response to one or more commands. In some embodiments, in response to a first one-way command, the logic unit 13 may be configured to read an old data value from a first address indicated in the first one-way command, perform an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value, and locally store the intermediate parity value at a first location associated with a first index indicated in the first one-way command. For example, in response to the first one-way command, logic unit 13 may also be configured to write the new data at the first address. In some embodiments, in response to a second one-way command, logic unit 13 may be configured to read an intermediate parity value from a second location associated with a second index indicated in the second one-way command, and store the intermediate parity value at the second address indicated in the second one-way command.
[0021] In some embodiments, in response to a third one-way command, logic unit 13 may also be configured to read an old parity data value from a third address indicated in the third one-way command and locally store the old parity data value at a third location associated with a third index indicated in the third one-way command. For example, in response to the third one-way command, logic unit 13 may also be configured to perform the second parity calculation based on the old parity data value, an intermediate parity value indicated in the third one-way command, and a coefficient value indicated in the third one-way command to determine a final parity value, and write the final parity value to the third address.
[0022] In some embodiments, in response to a fourth one-way command, logic unit 13 may also be configured to: read an old data value from a fourth address indicated in the fourth one-way command; perform an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine an intermediate parity value; perform a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command to determine a final parity value; and write the final parity value to a fifth address indicated in the fourth one-way command. In any embodiment of the embodiments herein, persistent storage medium 12 may include a solid-state drive (SSD), a hard disk drive (HDD), or the like.
[0023] Embodiments of each of the device controller 11, persistent storage medium 12, logic unit 13, and other device components described above can be implemented in hardware, software, or any suitable combination thereof. For example, hardware implementations may include configurable logic units, such as programmable logic arrays (PLAs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), or fixed-function logic hardware using circuitry techniques, such as application-specific integrated circuits (ASICs), complementary metal-oxide-semiconductor (CMOS) or transistor-transistor (TTL) technology, or any combination thereof. Embodiments of the device controller 11 may include general-purpose controllers, dedicated controllers, memory controllers, storage controllers, microcontrollers, general-purpose processors, dedicated processors, central processing units (CPUs), execution units, etc. In some embodiments, persistent storage medium 12 and / or logic unit 13 may be located within or co-located with various components including the device controller 11 (e.g., on the same die, in the same package, in the same housing, etc.).
[0024] Alternatively or additionally, all or part of these components may be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc.) for execution by a processor or computing device. For example, the computer program code for performing the operations of the components may be written in any combination of programming languages applicable to or suitable for one or more operating systems (OS), including object-oriented programming languages (such as Python, Perl, Java, Smalltalk, C++, C#, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). For example, persistent storage medium 12, other persistent storage media, or other device memory may store a set of instructions that, when executed by device controller 11, cause device 10 to implement one or more components, features, or aspects of device 10 (e.g., logic unit 13, determining an intermediate parity value based on a first local parity calculation, locally storing the intermediate parity value, determining a final parity value based on the intermediate parity value and a second local parity calculation, etc.).
[0025] Now go to Figure 2 Embodiments of the electronic device 15 may include one or more substrates 16 and a logic unit 17 coupled to one or more substrates 16. The logic unit 17 may be configured to: control local access to persistent storage media, and in response to one or more commands, determine an intermediate parity value based on a first local parity calculation and locally store the intermediate parity value. For example, the logic unit 17 may also be configured to, in response to the one or more commands, determine a final parity value based on the intermediate parity value and a second local parity calculation. In some embodiments, in response to a first one-way command, the logic unit 17 may be configured to read an old data value from a first address indicated in the first one-way command, perform an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value, and locally store the intermediate parity value at a first location associated with a first index indicated in the first one-way command. For example, in response to the first one-way command, the logic unit 17 may also be configured to write the new data to the first address. In some embodiments, in response to a second one-way command, logic unit 17 may be configured to read an intermediate parity value from a second location associated with a second index indicated in the second one-way command, and store the intermediate parity value at a second address indicated in the second one-way command.
[0026] In some embodiments, in response to a third one-way command, logic unit 17 may also be configured to read an old parity data value from a third address indicated in the third one-way command and locally store the old parity data value at a third location associated with a third index indicated in the third one-way command. For example, in response to the third one-way command, logic unit 17 may also be configured to perform the second parity calculation based on the old parity data value, an intermediate parity value indicated in the third one-way command, and a coefficient value indicated in the third one-way command to determine a final parity value, and write the final parity value to the third address.
[0027] In some embodiments, in response to a fourth one-way command, logic unit 17 may also be configured to read an old data value from a fourth address indicated in the fourth one-way command, perform an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine an intermediate parity value, perform a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command to determine a final parity value, and write the final parity value to a fifth address indicated in the fourth one-way command. In any embodiment of the embodiments herein, the persistent storage medium may include an SSD.
[0028] Embodiments of logic unit 17 may be implemented in systems, devices, computers, apparatuses, etc., such as those described herein. More specifically, hardware implementations of logic unit 17 may include configurable logic units, such as PLA, FPGA, CPLD, or fixed-function logic hardware using circuit technologies (such as, for example, ASIC, CMOS, or TTL technologies, or any combination thereof). Alternatively or additionally, logic unit 17 may be implemented in one or more modules as a set of logical instructions stored in a machine- or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.) for execution by a processor or computing device. For example, computer program code for performing the operations of the component may be written in any combination of one or more OS-suitable / appropriate programming languages, including object-oriented programming languages (such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages).
[0029] For example, logic cell 17 may be implemented on a semiconductor device that may include one or more substrates 16, wherein logic cell 17 is coupled to one or more substrates 16. In some embodiments, logic cell 17 may be at least partially implemented in one or more configurable logic cells and fixed-function hardware logic on one or more semiconductor substrates (e.g., silicon, sapphire, gallium arsenide, etc.). For example, logic cell 17 may include a transistor array and / or other integrated circuit components coupled to one or more substrates 16, wherein transistor channel regions are located within one or more substrates 16. The interface between logic cell 17 and one or more substrates 16 may not be abrupt. Logic cell 17 may also be considered as an epitaxial layer grown on an initial wafer of one or more substrates 16.
[0030] Turn now Figures 3A to 3E An embodiment of method 20 for controlling a storage device may include controlling local access to a persistent storage medium at block 21, and in response to one or more commands: determining an intermediate parity value based on a first local parity calculation at block 22, and locally storing the intermediate parity value at block 23. Method 20 may also include determining a final parity value based on the intermediate parity value and a second local parity calculation at block 24 in response to the one or more commands. Some embodiments of method 20 may include: reading an old data value from a first address indicated in the first one-way command at block 26 in response to a first one-way command at block 25, performing an XOR operation on the old data value and a new data value indicated in the first one-way command at block 27 to determine an intermediate parity value, and locally storing the intermediate parity value at a first location associated with a first index indicated in the first one-way command at block 28. For example, in response to the first one-way command, method 20 may also include writing the new data at the first address at block 29. Some embodiments of method 20 may further include: in response to a second one-way command at block 30, reading an intermediate parity value from a second location associated with a second index indicated in the second one-way command at block 31, and storing the intermediate parity value at a second address indicated in the second one-way command at block 32.
[0031] Some embodiments of method 20 may further include: in response to a third one-way command at block 33, reading an old parity data value from a third address indicated in the third one-way command at block 34, and storing the old parity data value at a third location associated with a third index indicated in the third one-way command at block 35. For example, in response to the third one-way command, method 20 may further include: performing a second parity calculation at block 36 based on the old parity data value, an intermediate parity value indicated in the third one-way command, and a coefficient value indicated in the third one-way command to determine a final parity value, and writing the final parity value at the third address at block 37.
[0032] Some embodiments of method 20 may further include: in response to a fourth one-way command at block 38, reading an old data value from a fourth address indicated in the fourth one-way command at block 39, performing an XOR operation on the old data value and a new data value indicated in the fourth one-way command at block 40 to determine an intermediate parity value, performing a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command at block 41 to determine a final parity value, and writing the final parity value to a fifth address indicated in the fourth one-way command at block 42. In any embodiment of the embodiments herein, at block 43, the persistent storage medium may include an SSD.
[0033] Embodiments of method 20 may be implemented in systems, apparatuses, computers, devices, etc., such as those described herein. More specifically, hardware implementations of method 20 may include configurable logic units, such as, for example, PLA, FPGA, CPLD, or fixed-function logic hardware using circuit technologies (such as, for example, ASIC, CMOS, or TTL technologies), or any combination thereof. Alternatively or additionally, method 20 may be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.) for execution by a processor or computing device. For example, computer program code performing the operations of components may be written in any combination of one or more OS-appropriate / suitable programming languages, including object-oriented programming languages (such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages).
[0034] For example, method 20 may be implemented on a computer-readable medium, as described below in conjunction with Examples 25 through 32. Embodiments or portions of method 20 may be implemented in firmware, an application (e.g., via an application programming interface (API)), or driver software running on an operating system (OS). Additionally, the logic instructions may include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, state setting data, configuration data of integrated circuits, state information of individual electronic circuits, and / or other structural components derived from hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.).
[0035] Some implementations can advantageously provide techniques for on-SSD erasure coding (EC) without using bidirectional commands. EC-based storage technologies require additional parity calculations. This is accomplished by reading old data and old parity, recalculating parity given new data, and storing new data and new parity. The central component of a typical EC-based storage system may include a storage controller entity (e.g., a RAID host bus adapter (HBA) or similar software), which is a centralized entity that exposes one or more storage volumes to other entities, coordinates all EC flows, and performs the necessary calculations. For the purpose of random write operations (which become partial striped writes), the entity generates a large number of read and write operations relative to member drives and performs the necessary data calculations using the central storage controller's computation engine.
[0036] The centralized entity must first obtain all the data required to perform the parity calculation, typically including both data from the requesting side and data from the member drives. After performing the calculation, the centralized entity needs to write all the data and the calculated parity data to the member drives. For example, in a typical EC-based storage system, this requires P+1 data transfers to read the old parity and data, and P+1 data transfers to save the update, where the parity calculation is performed at the centralized storage controller. For example, in an exemplary 8+3 EC system, each random write becomes eight (8) data transfers (e.g., 3+1+3+1, where P=3), which corresponds to 800% data transfer overhead.
[0037] The high data transfer overhead to and from the centralized storage controller consumes local (e.g., PCIe) and / or remote (e.g., Ethernet) interface bandwidth and can ultimately limit the maximum storage performance of a typical EC-based storage system. All computations are also performed by the centralized storage controller entity, which can also be a bottleneck for large systems. Some other EC-based systems can utilize bidirectional commands to reduce overhead (e.g., a single command to move data from host memory to a member drive and vice versa). However, implementing such bidirectional commands in scalable EC-based storage systems is difficult or impractical. Some other EC-based systems may include techniques for offloading XOR operations, but the scope of such operations is limited in supporting more complex EC-based storage systems.
[0038] Advantageously, some embodiments can provide advanced EC offloading techniques to overcome one or more of the aforementioned problems of conventional EC-based storage systems. Some embodiments can advantageously support generalized EC levels. Some embodiments can also advantageously utilize one-way commands (e.g., a single command involving only a single data transfer between host memory and member drives) to simplify the implementation of scalable EC-based storage systems. Some embodiments can offload complex mathematical operations (e.g., multiplication) from the EC controller to member drives to advantageously reduce network traffic and computational overhead from the centralized storage controller.
[0039] Examples of WriteAndXor, SaveEC, and ECwrite
[0040] Not limited to a specific implementation, the embodiment may introduce the following three (3) SSD commands to move complex mathematical operations to the SSD member drive: 1) WriteAndXor(Logical Block Address (LBA) L, Data D, Descriptor CD) idx 1) This command XORs the old and new data and places them at a temporary location indexed by the specified descriptor; 2) SaveEC(descriptor CD) idx1) `Memory Address M`: This command reads the temporary location specified by the descriptor and places the result in memory address M. The space at memory address M is allocated by the central memory controller and can be allocated in host memory, persistent memory region (PMR), internal memory cache (IMB), control memory cache (CMB), etc.; and 2) `ECwrite(LBAL, data X, data g)`: This command reads the old (parity) data at LBA L and stores it at the temporary location specified by the `CDidx` descriptor. This command also updates the data at L using, for example, the parity calculation result calculated using the parameter g and the intermediate parity value X (e.g., from memory address M) using Galois field (GF) mathematics.
[0041] Embodiments of the central storage controller may include techniques for detecting the availability of these commands on the constituent SSDs (e.g., queries for device capabilities or compatibility) and coordinating operations. The central storage controller may also include techniques for maintaining a list of currently used indexes and selecting currently unused indexes when initiating new command flows for random writes and / or partial stripe writes. Advantageously, some embodiments reduce the number of I / O operations against member drives, offloading the CPU / hardware utilization of the central storage controller and improving the scalability of the overall storage solution. For the exemplary 8+3EC, some embodiments eliminate XOR operations and GF multiplications on the central storage controller and reduce the number of data transfers on the bus / network from 8 to 5 (e.g., a reduction of approximately 40%; see Table 1 below).
[0042] Table 1 illustrates examples comparing XOR operations, GF multiplications, and data transfers for a baseline conventional storage system versus EC-offloaded storage, based on embodiments for a general M+P configuration and an exemplary 8+3 (M=8, P=3) configuration. The rightmost two columns summarize improvements provided according to some embodiments (e.g., reductions in data transfers and operations at the central storage controller).
[0043]
[0044] Table 1
[0045] Not limited to a specific implementation, the following commands can be implemented on an SSD that provides logical units to implement commands as indicated by the following pseudocode (and as described in further detail herein):
[0046]
[0047]
[0048] refer to Figures 4A to 4BEmbodiments of the electronic storage system 45 may include an EC controller 46 (e.g., a RAID controller, such as a RAID HBA, Intel Virtual RAID (VROC) on a CPU, etc.) communicatively coupled to multiple member drives 47 (e.g., a 2+2 EC configuration including SSD1 to SSD4, or a 4-drive RAID6 configuration) via a bus and / or network interface 48. According to some embodiments, the EC controller 46 includes a logical unit for coordinating commands, such as in… Figure 4A The flowcharts numbered 1 to 7 are shown, which further describe the operation of the EC controller 46 and member driver 47. Figure 4B As shown in the image.
[0049] The commands WriteAndXor and SaveEC use command descriptor indexes (CDs). idx They are interconnected, among which CD idx Generated by the host / server node for a given RAID operation, and the host guarantees that no two incomplete RAID operations on the same drive have the same CD. idx The intermediate result (Y) of the WriteAndXor command idx ) is temporarily stored on the SSD drive (e.g., as in Figure 4A The cache of the SSD1 shown in the diagram. The ECwrite parameter g can be a coefficient used for GF mathematical calculations. The value of g can be set by the RAID logical unit and can depend on the RAID level (e.g., the number of parity drives in the stripe) and the parity drive index in a given RAID stripe.
[0050] At arrow 1, EC controller 46 receives the Write(L',D1) command. At point 2, EC controller 46 allocates IDX1 as the command descriptor index for SSD1 and allocates space at memory address M1 in memory to EC controller 46 to store the intermediate parity value X1. For example, EC controller 46 can be a software entity running on a host computer, and EC controller 46 can allocate a portion of host memory specifically for EC controller to store the intermediate parity value X1. EC controller 46 can also map LBA L' to LBA L1 of SSD1.
[0051] At arrow 3, the EC controller issues the command WriteAndXor(L1,D1,IDX1) to SSD1. On member drive 47, the WriteAndXor command is responsible for writing the updated data to the target drive, while placing the intermediate parity calculation in a temporary cache, the location of which is indicated by a specified descriptor. For example, member drives 47 may each include an EC offload logic unit to implement the command. Therefore, on SSD1, the WriteAndXor(L1,D1,IDX1) command causes the EC offload logic unit to read the old data OLD from LBA L1. D1 Perform an XOR operation on the old data and the new data D1 (Y IDX1 =D1 XOR OLD D1 ), and the intermediate parity check value Y IDX1 The data is stored in a buffer at the location indicated by IDX1. Then, the WriteAndXor(L1,D1,IDX1) command causes the EC unload logic unit to write the new data D1 to LBAL1.
[0052] At arrow 4, EC controller 46 sends the command SaveEC(IDX1, M1) to SSD1. The SaveEC command causes the EC unloading logic unit to obtain the calculated intermediate parity value Y. IDX1 This data is placed in memory address M1, which can be in host memory (e.g., a portion of host memory allocated to EC controller 46, or in PRM, CMB, IMB, etc.). At arrows 5 and 6, the EC controller performs an intermediate parity calculation X1 (e.g., read from memory address M1) and issues an ECwrite command to one or more parity drives, having LBA L1, data value X1, and corresponding appropriate GF coefficients g0 and g1. Internally, member drives SSD3 and SSD4 perform the parity calculation and store the result at the designated LBA. At arrow 7, the application is notified of completion. Advantageously, in this embodiment, all EC solution calculations for partial stripe write purposes are offloaded to member drives. Compared to conventional RSTe technology, some embodiments can advantageously demonstrate substantial improvements in input / output per second (IOP), CPU utilization per I / O (relative percentage), average latency, and 99% QoS, as well as various configuration aspects (e.g., RAID5, RAID6, etc.).
[0053] refer to Figure 5Embodiments of the electronic storage system 50 may include a first storage node 51 communicatively coupled to a second storage node 52 in a distributed RAID configuration. Solid arrows may indicate local data transfers, while dashed arrows indicate network data transfers. The first storage node 51 includes M data SSDs, while the second storage mode 52 includes P parity SSDs. The first storage node 51 includes a RAID logic unit 53 as described herein to offload EC calculations, and each of the member SSDs includes an offload logic unit to perform EC calculations. Advantageously, some embodiments reduce network bandwidth in the case of a distributed EC system. For example, the first storage node 51 does not need to read data from the parity SSDs on the second storage node 52, thereby eliminating P data transfers (e.g., saving 50% of network data transfers in a system with two parity drives). The illustrated embodiment also eliminates parity calculations on the first storage node 51, advantageously reducing the computational complexity and overhead for the RAID logic unit 53.
[0054] ECrmw Example
[0055] In some EC systems, partial update penalties may exist. Partial updates occur when a stored file or block is partially modified. This results in updating not only a corresponding source vector but also several corresponding parity blocks. For example, some conventional EC systems might read all unmodified source vectors from different storage nodes in the storage system to a single storage node, recalculate all parity vectors, and then replace the stale parity vectors with the new parity vectors from the different storage nodes. Data updates involve all relevant storage nodes. The more storage nodes involved, the more complex the interaction, the higher the latency, the greater the data overhead over the network, and the greater the computational workload at each storage node.
[0056] Some embodiments can overcome one or more of the aforementioned problems by utilizing vendor-specific NVMe commands to perform EC operations on the storage device. Instead of calculating parity information on a central storage server, some embodiments include techniques for performing EC operations on the storage device, including reads, GF calculations, and write-backs. In some embodiments, the storage device may be an NVMe drive or an NVMe over Fabric (NVMf) target. In some embodiments, commands for the storage device may indicate the LBA for stale parity data and a new expected LBA for new parity data, as well as multiplication factors (e.g., for GF calculations). Embodiments of appropriately configured NVMe storage devices can decode commands to perform partial EC update calculations using a common XOR engine on the device.
[0057] Some embodiments may integrate vendor-specific NVMe commands into the firmware of the NVMe flash drive. Some embodiments may be configured to extend NVMe target-related software. Some embodiments may be integrated into a Smart Network Interface (NIC) card. Storage system embodiments may be configured to determine whether EC offload code is deployed in the storage system device, and if so, to send only one command to complete the EC partial update, rather than through regular interoperation between the host server and the storage device. Advantageously, some embodiments may improve throughput and latency, and may simplify the associated software logic units.
[0058] refer to Figure 6 An embodiment of electronic system 60 may include a requester 61, which is communicatively coupled to NVMe subsystem 62 via storage server node 63. For example, storage server node 63 may be communicatively coupled to NVMe subsystem 62 via a bus or network interface (such as an NVMe interface, NVMf interface, etc.). At arrow 1, storage server node 63 receives incremental data from requester 61 (e.g., another storage node, application, software agent, etc.). At arrow 2, RAID logic unit 65 on storage server node 63 issues an ECrmw command filled with appropriate data to NVMe subsystem 62. Offload logic unit 64 (e.g., firmware) on NVMe subsystem 62 directly transfers the incremental data to storage device(s). Internally, the offload logic unit processes the data as follows: reads out stale data (arrow A), calculates a new parity with the input incremental data (point B), and writes back the new data (arrow C).
[0059] Advantageously, the operation initiated by the ECrmw command occurs within the NVMe subsystem 62, thus requiring no host server involvement from the storage server node 63. Therefore, some embodiments alleviate the burden on the host in multiple data manipulations and computations, while simultaneously reducing operational latency. After the NVMe subsystem 62 receives both the ECrmw command and the incremental data, the storage server node 63 can proceed to indicate a successful response to the requester 61 at arrow 3. A typical read-modify-write EC operation involves several steps between the server node and the storage node, so it is not an atomic operation. Advantageously, embodiments of the ECrmw command can perform read-modify-write as a single atomic operation on the drive (e.g., a one-way command). Therefore, for upper-layer applications, it is possible to remove or significantly simplify complex software logic units used to prevent data loss or inconsistency.
[0060] Not limited to a specific implementation, some embodiments of EC storage systems can perform EC-related operations based on the following equation:
[0061] Pn'=αn (Dm'XOR Dm)XOR Pn [Equation 1]
[0062] refer to Figure 7 An embodiment of command format 70 for the exemplary ECrmw command may include a plurality of 32-bit data words (Dwords) arranged as illustrated. The incremental data is the data to be transferred to the storage device. The multiplication factor should also be implemented by the device. Similarly, unlike a typical read / write command that specifies only one LBA, in the ECrmw command, both the previous parity data LBA and the next new parity data LBA are specified to the storage device. The data length may be specified by the number of blocks, and there may be specific metadata for the new parity. Generally, the format of the ECrmw command may be similar to that of a typical NVMe read / write command, with the following differences: reserved segments in Dword2 and 3 are used to contain the LBA for new parity data; a segment in Dword10 and 11 is used to contain the LBA for old parity data; coefficients for GF multiplication are contained in Byte2 and Dword12; the meta pointer segments in Dword4 and 5 point to the metadata cache for new parity; and incremental data pointers are referenced by physical page regions (PRPs) PRP1 and PRP2 (e.g., following NVMe Spec 1.3). Advantageously, the ECrmw command provides information from the host shared to the storage device to perform read-modify-write operations against the EC as atomic operations.
[0063] Some implementations involve collaboration between the host and the storage device. For example, NAND flash memory provides a read / write / erase interface. Within a NAND package, the storage medium is organized into a hierarchy of dies, planes, blocks, and pages. Within each plane, NAND is organized in the form of blocks and pages. Each plane contains the same number of blocks, and each block contains the same number of pages. A page is the smallest unit for reading and writing, while the unit for erasing is a block. There are fewer restrictions on read operations; each read operation can be considered random. However, for write operations, also known as procedural operations, there is a major restriction: within each block, procedural operations must be appended sequentially within each block. In other words, blocks only allow purely sequential write access. To accommodate the characteristics of NAND, a file translation layer (FTL) within the storage device is responsible for exposing the regular block I / O interface.
[0064] When performing data overwrite, a common action of FTL is to find a new physical address for the new data, mark the original data's physical address as obsolete, and remap the LBA using the new physical address. Furthermore, considering data reliability in distributed systems with synchronization issues, advanced storage systems aim to retain stale parity data to prevent data rollback due to failures on other partner nodes. This is especially true for EC M+P (where M is the amount of source data and P is the amount of parity data, providing redundancy for P node failures while ensuring data reliability). To guarantee redundancy for any P node failures during read-modify-write operations, if each parity node directly overwrites the stale parity with the new parity, any unsuccessful node write will leave those nodes in an inconsistent state, causing M+PEC to lose its redundancy guarantee. In other words, retaining stale parity data along with the new parity data is a critical requirement for EC data updates.
[0065] Therefore, the data retention behavior of some embodiments satisfies both the functional requirements of the host software and the inherent characteristics of NAND flash memory as the primary medium for NVMe / NVMf.
[0066] The technologies discussed in this article can be provided in a variety of computing systems (e.g., non-mobile computing devices such as desktops, workstations, servers, rack systems, etc.; mobile computing devices such as smartphones, tablets, ultra-mobile personal computers (UMPCs), laptops, ultrabook computing devices, smartwatches, smart glasses, smart bracelets, etc.; and / or client / edge devices such as Internet of Things (IoT) devices (e.g., sensors, cameras, etc.).
[0067] Now go to Figure 8 Embodiments of computing device 100 may include one or more processors 102-1 to 102-N (generally referred to herein as "multiple processors 102" or "processor 102"). Processors 102 may communicate via interconnect or bus 104. Each processor 102 may include a variety of components, and for clarity, only some of these components are discussed with reference to processor 102-1. Therefore, each of the remaining processors 102-2 to 102-N may include the same or similar components discussed with reference to processor 102-1.
[0068] In some embodiments, processor 102-1 may include one or more processor cores 106-1 to 106-M (referred to herein as "multiple cores 106" or more generally as "core 106"), cache 108 (which may be a shared cache or a private cache in various embodiments), and / or router 110. Processor core 106 may be implemented on a single integrated circuit (IC) chip. Furthermore, the chip may include one or more shared and / or private caches (such as cache 108), buses or interconnects (such as bus or interconnect 112), logic units 170, memory controllers, or other components.
[0069] In some embodiments, router 110 can be used to communicate between various components of processor 102-1 and / or device 100. Furthermore, processor 102-1 may include more than one router 110. Additionally, multiple routers 110 may be in communication to enable data routing between various components inside or outside processor 102-1.
[0070] Cache 108 can store data (e.g., instructions) used by one or more components of processor 102-1 (such as core 106). For example, cache 108 can locally cache data stored in memory 114 for faster access by components of processor 102. Figure 8 As shown, memory 114 can communicate with processor 102 via interconnect 104. In some embodiments, cache 108 (which may be shared) can have various levels; for example, cache 108 may be a mid-level cache and / or a last-level cache (LLC). Similarly, each core in core 106 may include a level 1 (L1) cache (116-1) (generally referred to herein as "L1 cache 116"). Various components of processor 102-1 can communicate directly with cache 108 via a bus (e.g., bus 112) and / or a memory controller or hub.
[0071] As in Figure 8 As shown, memory 114 can be coupled to other components of device 100 via memory controller 120. Memory 114 may include volatile memory and may be interchangeably referred to as main memory. Even though memory controller 120 is shown coupled between interconnect 104 and memory 114, memory controller 120 may be located elsewhere in device 100. For example, in some embodiments, memory controller 120 or a portion thereof may be located within one of the processors 102.
[0072] Device 100 can communicate with other devices / systems / networks via network interface 128 (e.g., it communicates with a computer network and / or cloud 129 via a wired or wireless interface). For example, network interface 128 may include an antenna (not shown) to communicate wirelessly with network / cloud 129 (e.g., via an IEEE 802.11 interface (including IEEE 802.11a / b / g / n / ac, etc.), cellular interface, 3G, 4G, LTE, Bluetooth, etc.).
[0073] Device 100 may also include a storage device, such as an SSD device 130 coupled to interconnect 104 via SSD controller logic unit 125. Therefore, logic unit 125 can control access to the SSD device 130 by various components of device 100. Furthermore, even if logic unit 125 is shown as being directly coupled to... Figure 8 Interconnect 104 and logic unit 125 can also alternatively communicate with one or more other components of device 100 (e.g., where the storage bus is coupled to interconnect 104 via other logic units such as bus bridges, chipsets, etc.) via storage bus / interconnect (such as SATA (Serial Advanced Technology Attachment) bus, Peripheral Component Interconnect (PCI) (or PCI Express (PCIe) interface), NVM Express (NVMe), etc.). Additionally, logic unit 125 can be incorporated into memory controller logic units (such as reference...) Figure 9 Those discussed) or are disposed on the same integrated circuit (IC) device in various embodiments (e.g., on the same circuit board device as SSD device 130 or in the same housing as SSD device 130).
[0074] Furthermore, logic unit 125 and / or SSD device 130 may be coupled to one or more sensors (not shown) to receive information (e.g., in the form of one or more bits or signals) to indicate a state or value detected by one or more sensors. These sensors may be provided close to components of device 100 (or other computing systems discussed herein), including core 106, interconnect 104 or 112, components external to processor 102, SSD device 130, SSD bus, SATA bus, logic unit 125, logic unit 160, logic unit 170, etc., to sense changes in various factors affecting the power / thermal behavior of the system / platform, such as temperature, operating frequency, operating voltage, power consumption, and / or inter-core communication activity.
[0075] Figure 9 A block diagram illustrating various components of an SSD device 130 according to an embodiment is shown. (As in...) Figure 9As illustrated, the logic unit 160 can be located in various locations, such as inside the SSD device 130 or the controller 382, and can include components that are combined with... Figure 8 Similar technologies discussed. SSD device 130 includes a controller 382 (which in turn includes one or more processor cores or processor 384 and a memory controller logic unit 386), a cache 138, RAM 388, firmware storage device 390, and one or more memory devices 392-1 to 392-N (collectively referred to as memory 392, which may include NAND flash, NOR flash, or other types of non-volatile memory). Memory 392 is coupled to memory controller logic unit 386 via one or more memory channels or buses. Similarly, SSD device 130 communicates with logic unit 125 via an interface (such as SATA, SAS, PCIe, NVMe, etc.). Reference Figure 1-7 One or more features / aspects / operations discussed can be derived from... Figure 9 One or more of the components are used to execute the operation. The processor 384 and / or controller 382 can compress / decompress (or otherwise cause compression / decompression) data written to or read from memory devices 392-1 to 392-N. Similarly, Figure 1-7 One or more of the features / aspects / operations can be programmed into firmware 390. Additionally, SSD controller logic unit 125 may also include logic unit 160.
[0076] As in Figure 8 and Figure 9 As illustrated, SSD device 130 may include logic unit 160, which may be housed in the same enclosure as SSD device 130 and / or fully integrated on the printed circuit board (PCB) of SSD device 130. Device 100 may also include logic unit 170 external to SSD device 130. Advantageously, logic unit 160 and / or logic unit 170 may include components for implementing device 10, means 15, method 20. Figures 3A to 3E This includes technologies for systems 45, 50, 60, command format 70, and / or any of the features discussed herein. For example, logic unit 170 may include technologies for implementing host server / device / agent / RAID logic / EC controller aspects of the various embodiments described herein (e.g., sending commands / parameters to SSD device 130, including WriteAndXor, SaveEC, ECwrite, and ECrmw commands and associated data as described herein).
[0077] For example, logic unit 160 may include technologies for controlling local access to SSD device 130 (e.g., command logic unit, control logic unit, device logic unit, etc., for SSD device 130 itself). In response to one or more commands (e.g., commands such as WriteAndXor, SaveEC, ECwrite, and ECrmw as described herein), logic unit 160 may be configured to determine an intermediate parity value based on a first local parity calculation and to locally store the intermediate parity value (e.g., in RAM 388). For example, logic unit 160 may also be configured to determine a final parity value based on the intermediate parity value and a second local parity calculation in response to one or more commands.
[0078] In some embodiments, in response to a WriteAndXor command, logic unit 160 may be configured to read an old data value from the address indicated in the WriteAndXor command, perform an XOR operation on the old data value and the new data value indicated in the WriteAndXor command to determine an intermediate parity value, and locally store the intermediate parity value at a location associated with the index indicated in the WriteAndXor command. For example, in response to a WriteAndXor command, logic unit 160 may also be configured to write new data to the address indicated in the WriteAndXor command.
[0079] In some embodiments, in response to a SaveEC command, logic unit 160 may be configured to read an intermediate parity value from a location associated with an index indicated in the SaveEC command and store the intermediate parity value at the address indicated in the SaveEC command.
[0080] In some embodiments, in response to an ECwrite command, logic unit 160 may be configured to read an old parity data value from the address indicated in the ECwrite command and locally store the old parity data value at a location associated with the index indicated in the ECwrite command. For example, in response to an ECwrite command, logic unit 160 may also be configured to perform a second parity calculation based on the old parity data value, an intermediate parity value indicated in the ECwrite command, and a coefficient value indicated in the ECwrite command to determine a final parity value, and write the final parity value to the address indicated in the ECwrite command.
[0081] In some embodiments, in response to an ECrmw command, logic unit 160 may be configured to read an old data value from a first LBA indicated in the ECrmw command, perform an XOR operation on the old data value and a new data value indicated in the ECrmw command to determine an intermediate parity value, perform a multiplication operation (e.g., XOR with the old parity data value according to Equation 1) based on the intermediate parity value and a coefficient value indicated in the ECrmw command to determine a final parity value, and write the final parity value to a second LBA indicated in the ECrmw command.
[0082] In other embodiments, SSD device 130 may utilize any suitable storage / memory technology / medium instead. In some embodiments, logic cells 160 / 170 may be coupled to one or more substrates (e.g., silicon, sapphire, gallium arsenide, printed circuit board (PCB), etc.) and may include transistor channel regions located within one or more substrates. In other embodiments, SSD device 130 may include two or more types of storage media. For example, the body of the storage device may be NAND and may also include some faster, finer-grained accessible (e.g., byte-addressable) NVM, such as Intel 3DXP media. SSD device 130 may alternatively or additionally include persistent volatile memory (e.g., battery or capacitor-backed DRAM or SRAM). For example, SSD device 130 may include impending power-off (PLI) technology with energy storage capacitors. The energy storage capacitors can provide sufficient energy (power) to complete any commands being executed and ensure that any data in the DRAM / SRAM is committed to the non-volatile NAND media. The capacitors can serve as backup batteries for persistent volatile memory. Figure 8 As shown, the features or aspects of logic unit 160 and / or logic unit 170 may be distributed throughout the device 100 and / or co-located / integrated with various components of the device 100.
[0083] Additional notes and examples
[0084] Example 1 includes an electronic device comprising: one or more substrates; and a logic unit coupled to the one or more substrates, the logic unit controlling local access to a persistent storage medium and responding to one or more commands for: determining an intermediate parity value based on a first local parity calculation, locally storing the intermediate parity value, and determining a final parity value based on the intermediate parity value and a second local parity calculation.
[0085] Example 2 includes the apparatus of Example 1, wherein, in response to a first one-way command, the logic unit is further configured to: read an old data value from a first address indicated in the first one-way command; perform an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value; and locally store the intermediate parity value at a first location associated with a first index indicated in the first one-way command.
[0086] Example 3 includes the apparatus of Example 2, wherein, in response to the first one-way command, the logic unit is further configured to: write the new data at the first address.
[0087] Example 4 includes an apparatus of any one of Examples 1 to 3, wherein, in response to a second one-way command, the logic unit is further configured to: read an intermediate parity value from a second location associated with a second index indicated in the second one-way command; and store the intermediate parity value at a second address indicated in the second one-way command.
[0088] Example 5 includes an apparatus of any one of Examples 1 to 4, wherein, in response to a third one-way command, the logic unit is further configured to: read an old parity data value from a third address indicated in the third one-way command, and locally store the old parity data value at a third location associated with a third index indicated in the third one-way command.
[0089] Example 6 includes the apparatus of Example 5, wherein, in response to the third one-way command, the logic unit is further configured to perform the second parity calculation based on the old parity data value, the intermediate parity value indicated in the third one-way command, and the coefficient value indicated in the third one-way command to determine the final parity value, and write the final parity value at the third address.
[0090] Example 7 includes an apparatus of any one of Examples 1 to 6, wherein, in response to a fourth one-way command, the logic unit is further configured to read an old data value from a fourth address indicated in the fourth one-way command, perform an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine the intermediate parity value, perform a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command to determine the final parity value, and write the final parity value to a fifth address indicated in the fourth one-way command.
[0091] Example 8 includes an apparatus of any one of Examples 1 to 7, wherein the persistent storage medium includes a solid-state drive.
[0092] Example 9 includes an electronic storage system comprising: a persistent storage medium; and a controller communicatively coupled to the persistent storage medium, the controller including logic units for: controlling local access to the persistent storage medium, and in response to one or more commands, determining an intermediate parity value based on a first local parity calculation, locally storing the intermediate parity value, and determining a final parity value based on the intermediate parity value and a second local parity calculation.
[0093] Example 10 includes the system of Example 9, wherein, in response to a one-way command, the logic unit is further configured to: read an old data value from a first address indicated in the first one-way command, perform an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value, and locally store the intermediate parity value at a first location associated with a first index indicated in the first one-way command.
[0094] Example 11 includes the system of Example 10, wherein, in response to the first one-way command, the logic unit is further configured to write the new data at the first address.
[0095] Example 12 includes a system of any one of Examples 9 to 11, wherein, in response to a second one-way command, the logic unit is further configured to read an intermediate parity value from a second location associated with a second index indicated in the second one-way command, and to store the intermediate parity value at a second address indicated in the second one-way command.
[0096] Example 13 includes a system of any one of Examples 9 to 12, wherein, in response to a third one-way command, the logic unit is further configured to read an old parity data value from a third address indicated in the third one-way command and locally store the old parity data value at a third location associated with a third index indicated in the third one-way command.
[0097] Example 14 includes the system of Example 13, wherein, in response to the third one-way command, the logic unit is further configured to perform the second parity calculation based on the old parity data value, the intermediate parity value indicated in the third one-way command, and the coefficient value indicated in the third one-way command to determine the final parity value, and write the final parity value at the third address.
[0098] Example 15 includes a system of any one of Examples 9 to 14, wherein, in response to a fourth one-way command, the logic unit is further configured to read an old data value from a fourth address indicated in the fourth one-way command, perform an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine the intermediate parity value, perform a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command to determine the final parity value, and write the final parity value to a fifth address indicated in the fourth one-way command.
[0099] Example 16 includes a system of any one of Examples 9 to 15, wherein the persistent storage medium includes a solid-state drive.
[0100] Example 17 includes a method of controlling a storage device, comprising: controlling local access to a persistent storage medium, and in response to one or more commands: determining an intermediate parity value based on a first local parity calculation, locally storing the intermediate parity value, and determining a final parity value based on the intermediate parity value and a second local parity calculation.
[0101] Example 18 includes the method of Example 17, further comprising, in response to a first one-way command: reading an old data value from a first address indicated in the first one-way command; performing an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value; and locally storing the intermediate parity value at a first location associated with a first index indicated in the first one-way command.
[0102] Example 19 includes the method of Example 18, and further includes, in response to the first one-way command: writing the new data at the first address.
[0103] Example 20 includes a method of any of Examples 17 to 19, further comprising, in response to a second one-way command: reading an intermediate parity value from a second location associated with a second index indicated in the second one-way command; and storing the intermediate parity value at a second address indicated in the second one-way command.
[0104] Example 21 includes the method of Example 20, and further includes, in response to a third one-way command: reading an old parity data value from a third address indicated in the third one-way command, and locally storing the old parity data value at a third location associated with a third index indicated in the third one-way command.
[0105] Example 22 includes the method of any one of Examples 17 to 21, further comprising, in response to the third one-way command: performing the second parity calculation to determine the final parity value based on the old parity data value, the intermediate parity value indicated in the third one-way command, and the coefficient value indicated in the third one-way command, and writing the final parity value at the third address.
[0106] Example 23 includes the method of any one of Examples 17 to 22, further comprising, in response to a fourth one-way command: reading an old data value from a fourth address indicated in the fourth one-way command, performing an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine the intermediate parity value, performing a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command to determine the final parity value, and writing the final parity value to a fifth address indicated in the fourth one-way command.
[0107] Example 24 includes the method of any one of Examples 17 to 23, wherein the persistent storage medium includes a solid-state drive.
[0108] Example 25 is at least one non-transitory machine-readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to control local access to a persistent storage medium and in response to one or more commands for: determining an intermediate parity value based on a first local parity calculation; locally storing the intermediate parity value; and determining a final parity value based on the intermediate parity value and a second local parity calculation.
[0109] Example 26 includes at least one non-transitory machine-readable medium of Example 25, including a plurality of additional instructions that are executed on the computing device and in response to a first one-way command, such that the computing device is configured to: read an old data value from a first address indicated in the first one-way command, perform an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value, and locally store the intermediate parity value at a first location associated with a first index indicated in the first one-way command.
[0110] Example 27 includes at least one non-transitory machine-readable medium of Example 26, including a plurality of additional instructions that are executed on the computing device and in response to the first one-way command, causing the computing device to: write the new data at the first address.
[0111] Example 28 includes at least one non-transitory machine-readable medium of any of Examples 25 to 27, including a plurality of additional instructions that are executed on the computing device and in response to a second one-way command, causing the computing device to read an intermediate parity value from a second location associated with a second index indicated in the second one-way command, and to store the intermediate parity value at a second address indicated in the second one-way command.
[0112] Example 29 includes at least one non-transitory machine-readable medium of any of Examples 25 to 28, including a plurality of additional instructions that are executed on the computing device and in response to a third one-way command, causing the computing device to read an old parity data value from a third address indicated in the third one-way command and to locally store the old parity data value at a third location associated with a third index indicated in the third one-way command.
[0113] Example 30 includes at least one non-transitory machine-readable medium of Example 29, including a plurality of additional instructions that are responsive to being executed on a de facto computing device and in response to a third one-way command, causing the computing device to perform a second parity calculation based on the old parity data value, an intermediate parity value indicated in the third one-way command, and a coefficient value indicated in the third one-way command to determine the final parity value, and to write the final parity value at the third address.
[0114] Example 31 includes at least one non-transitory machine-readable medium of any of Examples 25 to 30, comprising a plurality of additional instructions that are executed on the computing device and in response to a fourth one-way command, causing the computing device to read an old data value from a fourth address indicated in the fourth one-way command, perform an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine the intermediate check value, perform a multiplication operation based on the intermediate check value and a coefficient value indicated in the fourth one-way command to determine the final check value, and write the final check value to a fifth address indicated in the fourth one-way command.
[0115] Example 32 includes at least one non-transitory machine-readable medium of any of Examples 25 to 31, wherein the persistent storage medium includes a solid-state drive.
[0116] Example 33 includes a storage device controller apparatus including a unit for controlling local access to a persistent storage medium, and, in response to one or more commands, a unit for determining an intermediate parity value based on a first local parity calculation, a unit for locally storing the intermediate parity value, and a unit for determining a final parity value based on the intermediate parity value and a second local parity calculation.
[0117] Example 34 includes the apparatus of Example 33, further comprising: a unit for reading an old data value from a first address indicated in the first one-way command in response to a first one-way command; a unit for performing an XOR operation on the old data value and a new data value indicated in the first one-way command to determine the intermediate parity value; and a unit for locally storing the intermediate parity value at a first location associated with a first index indicated in the first one-way command.
[0118] Example 35 includes the apparatus of Example 34, and further includes a unit for writing the new data at the first address in response to the first one-way command.
[0119] Example 36 includes an apparatus comprising any one of Examples 33 to 35, further comprising, in response to a second one-way command: a unit for reading an intermediate parity value from a second location associated with a second index indicated in the second one-way command, and a unit for storing the intermediate parity value at a second address indicated in the second one-way command.
[0120] Example 37 includes an apparatus of any of Examples 33 to 36, further comprising, in response to a third one-way command: a unit for reading an old parity data value from a third address indicated in the third one-way command, and a unit for locally storing the old parity data value at a third location associated with a third index indicated in the third one-way command.
[0121] Example 38 includes the apparatus of Example 37, and further includes, in response to the third one-way command: a unit for performing the second parity calculation to determine the final parity value based on the old parity data value, the intermediate parity value indicated in the third one-way command, and the coefficient value indicated in the third one-way command; and a unit for writing the final parity value at the third address.
[0122] Example 39 includes an apparatus comprising any one of Examples 33 to 38, further comprising, in response to a fourth one-way command: a unit for reading an old data value from a fourth address indicated in the fourth one-way command; a unit for performing an XOR operation on the old data value and a new data value indicated in the fourth one-way command to determine the intermediate parity value; a unit for performing a multiplication operation based on the intermediate parity value and a coefficient value indicated in the fourth one-way command to determine the final parity value; and a unit for writing the final parity value at a fifth address indicated in the fourth one-way command.
[0123] Example 40 includes an apparatus of any of Examples 33 to 39, wherein the persistent storage medium includes a solid-state drive.
[0124] The term "coupling" may be used herein to refer to any type of direct or indirect relationship between the components under discussion, and may be applied to electrical, mechanical, fluid, optical, electromagnetic, electromechanical, or other connections. Furthermore, unless otherwise stated, the terms "first," "second," etc., may be used herein for ease of discussion only and do not have a specific temporal or chronological meaning.
[0125] As used in this application and claims, a list of items connected by the term "one or more" can represent any combination of the listed terms. For example, the phrase "one or more of A, B, and C" and the phrase "one or more of A, B, or C" can both represent A; B; C; A and B; A and C; B and C; or A, B, and C. Various components of the systems described herein can be implemented in software, firmware, and / or hardware and / or any combination thereof. For example, various components of the systems or devices discussed herein can be provided at least in part by the hardware of a computing SoC, such as that found in a computing system like a smartphone. Those skilled in the art will recognize that the systems described herein can include additional components not depicted in the corresponding figures. For example, the systems discussed herein can include additional components such as bitstream multiplexers or demultiplexers, etc., which are not depicted for clarity.
[0126] Although implementations of the exemplary processes discussed herein may include performing all the operations shown in the order shown, this disclosure is not limited in this respect, and in various examples, implementations of the exemplary processes herein may include only a subset of the operations shown, operations performed in a different order than shown, or additional operations.
[0127] Furthermore, any one or more of the operations discussed herein can be performed in response to instructions provided by one or more computer program products. Such program products may include signal-bearing media that provide instructions, which, when executed by, for example, a processor, can provide the functionality described herein. Computer program products may be provided in any form of one or more machine-readable media. Thus, for example, a processor including one or more graphics processing units or processor cores may perform one or more blocks of exemplary processes in response to program code and / or instructions or instruction sets transmitted to the processor via one or more machine-readable media. Generally, machine-readable media may deliver software in the form of program code and / or instructions or instruction sets that can enable any device and / or system described herein to implement at least a portion of the operations discussed herein and / or any part of a device, system, or any module or component as discussed herein.
[0128] As used in any implementation described herein, the term "module" refers to software logic, firmware logic, hardware logic, and / or any combination of circuitry configured to provide the functionality described herein. Software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware" as used in any implementation described herein may, for example, individually or in any combination, include hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, and / or firmware storing instructions executed by the programmable circuitry. Modules may be embodied collectively or individually as circuitry forming part of a larger system, such as integrated circuits (ICs), system-on-a-chip (SoCs), etc.
[0129] Various embodiments can be implemented using hardware components, software components, or a combination of both. Examples of hardware components may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application programming interfaces (APIs), instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. The determination of whether to use hardware components and / or software components to implement an embodiment can vary based on any number of factors, such as desired computational speed, power level, thermal tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.
[0130] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium representing various logic units within a processor, which, when read by a machine, cause the machine to manufacture logic to perform the techniques described herein. This representation, referred to as an IP core, can be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities for loading into manufacturing machines that actually manufacture the logic or processor.
[0131] Although the specific features set forth herein have been described with reference to various implementations, this description is not intended to be limiting. Therefore, various modifications to the implementations described herein, as well as other embodiments that will be obvious to those skilled in the art to which this disclosure pertains, are considered to fall within the spirit and scope of this disclosure.
[0132] It should be understood that the embodiments are not limited to those described herein, but can be practiced with modifications and variations without departing from the scope of the appended claims. For example, the embodiments described above may include specific combinations of features. However, the embodiments described above are not limited in this respect, and in various implementations, the embodiments described above may include only a subset of these features, different orders of such features, different combinations of such features, and / or more features than those expressly listed. Therefore, the scope of the embodiments should be determined by reference to the appended claims and the full scope of their equivalents.
Claims
1. An erasure coding controller, comprising: One or more substrates; EC controller memory; as well as A logical unit, coupled to the one or more substrates, is configured to control access to persistent storage media and, in response to one or more commands, is further configured to: An intermediate parity value is determined based on a first parity check calculation by sending a one-way command to the persistent storage medium. The intermediate parity value is stored in the EC controller memory, and The final parity value is determined based on the intermediate parity value and the second parity value calculation.
2. The EC controller according to claim 1, wherein, In order to determine the intermediate parity value, the logic unit is configured to determine the intermediate parity value based on the first parity calculation of the erasure coding offloading logic unit of the persistent storage medium.
3. The EC controller according to claim 2, wherein, To determine the intermediate parity value, the logic unit is configured to send a one-way instruction to the EC offload logic unit, causing the EC offload logic unit to: Read the old data value from the first address indicated in the one-way command; Perform an XOR operation on the old data value and the new data value indicated in the one-way command to determine the intermediate parity value; and In the persistent storage medium, the intermediate parity value is stored at a first location associated with the first index indicated in the one-way command.
4. The EC controller according to claim 3, wherein, The one-way command causes the EC unloading logic unit to: Write the new data value at the first address.
5. The EC controller according to claim 2, wherein, The intermediate parity value is a first intermediate parity value, and the logic unit is further configured to send a second one-way command to cause the EC to unload the logic unit for: Read the second intermediate parity value from the second position associated with the second index indicated in the second one-way command; and The second intermediate parity value is stored at the second address indicated in the second one-way command.
6. The EC controller according to claim 5, wherein, The second intermediate parity check value is the first intermediate parity check value.
7. The EC controller according to claim 1, wherein, The persistent storage medium includes a solid-state drive.
8. An electronic storage system, comprising: Persistent storage media; as well as An erasure coding controller, communicatively coupled to the persistent storage medium, includes an EC controller memory and a logic unit configured to control access to the persistent storage medium and in response to one or more commands. The logic unit is further configured to: An intermediate parity value is determined based on a first parity check calculation by sending a one-way command to the persistent storage medium. The intermediate parity value is stored in the EC controller memory, and The final parity value is determined based on the intermediate parity value and the second parity value calculation.
9. The system according to claim 8, wherein, The persistent storage medium includes an erasure coding offload logic unit, which is configured as follows: Perform the first parity check calculation; and Receive the one-way command.
10. The system according to claim 9, wherein, In response to the one-way command, the EC unloading logic unit is further configured to: Read the old parity data value from the address indicated in the one-way command; and The old parity data value is stored at the location associated with the index indicated in the one-way command.
11. The system according to claim 10, wherein, The intermediate parity value is a first intermediate parity value, and in response to the one-way command, the EC unloading logic unit is further configured to: The second parity calculation is performed based on the old parity data value, the second intermediate parity value indicated in the one-way command, and the coefficient value indicated in the one-way command to determine the final parity value; as well as Write the final parity value at the address.
12. The system according to claim 11, wherein, The second intermediate parity check value is the first intermediate parity check value.
13. The system according to claim 9, wherein, In response to the second one-way command, the EC unloading logic unit is further configured to: Read the old data value from the second address indicated in the second one-way command; Perform an XOR operation on the old data value and the new data value indicated in the second one-way command to determine the intermediate parity value; as well as In the persistent storage medium, the intermediate parity value is stored at a second location associated with the second index indicated in the second one-way command.
14. The system according to claim 9, wherein, The persistent storage medium includes a solid-state drive.
15. A method for controlling a storage device, comprising: Access to persistent storage is controlled by an erasure coding controller, and responds to one or more commands: An intermediate parity value is determined based on a first parity check calculation by sending a one-way command to the persistent storage medium. The intermediate parity value is stored in the memory of the EC controller, and The final parity value is determined based on the intermediate parity value and the second parity value calculation.
16. The method according to claim 15, wherein, Determining the intermediate parity value includes having the erasure coding offload logic unit of the persistent storage medium calculate the intermediate parity value based on the first parity check.
17. The method of claim 16, further comprising, in response to sending the one-way command to the persistent storage medium: Read the old data value from the first address indicated in the one-way command; Perform an XOR operation on the old data value and the new data value indicated in the one-way command to determine the intermediate parity value; and In the persistent storage medium, the intermediate parity value is stored at a first location associated with the first index indicated in the one-way command.
18. The method of claim 17, further comprising, in response to sending the one-way command to the persistent storage medium: Write the new data value at the first address.
19. The method of claim 15, further comprising sending a second one-way command to the persistent storage medium, and in response to sending the second one-way command to the persistent storage medium: Read the old parity data value from the second address indicated in the second one-way command; and The old parity data value is stored at a second location associated with the second index indicated in the second one-way command.
20. The method according to claim 19, wherein, The intermediate parity value is a first intermediate parity value, and the method further includes, in response to sending the second one-way command to the persistent storage medium: The second parity calculation is performed based on the old parity data value, the second intermediate parity value indicated in the second one-way command, and the coefficient value indicated in the second one-way command to determine the final parity value. as well as Write the final parity value at the second address.
21. The method according to claim 20, wherein, The second intermediate parity check value is the first intermediate parity check value.
22. At least one non-transitory machine-readable medium, comprising a plurality of instructions responsive to being executed on an EC controller including an erasure coding controller memory, such that the EC controller controls access to a persistent storage medium, and responsive to one or more commands for: An intermediate parity value is determined based on a first parity check calculation by sending a one-way command to the persistent storage medium. The intermediate parity value is stored in the EC controller memory; and The final parity value is determined based on the intermediate parity value and the second parity value calculation.
23. The at least one non-transitory machine-readable medium according to claim 22, wherein, In order to determine the intermediate parity value, the EC controller is configured to determine the intermediate parity value based on the first parity calculation of the erasure coding offload logic unit of the persistent storage medium.
24. The at least one non-transitory machine-readable medium of claim 23, further comprising a plurality of additional instructions responsive to being executed on the EC controller and responsive to sending the one-way command to the persistent storage medium, such that the EC offload logic unit is configured to: Read the old data value from the first address indicated in the one-way command; Perform an XOR operation on the old data value and the new data value indicated in the one-way command to determine the intermediate parity value; A multiplication operation is performed based on the intermediate parity value and the coefficient value indicated in the one-way command to determine the final parity value; and Write the final parity value to the second address indicated in the one-way command.
25. The at least one non-transitory machine-readable medium according to claim 23, wherein, The intermediate parity value is a first intermediate parity value, and the at least one non-transitory machine-readable medium includes a plurality of additional instructions that are executed on the EC controller and, in response to sending a second one-way command to the persistent storage medium, cause the EC offload logic unit to: Read the second intermediate parity value from the second position associated with the second index indicated in the second one-way command; and The second intermediate parity value is stored at the third address indicated in the second one-way command.
26. The at least one non-transitory machine-readable medium according to claim 25, wherein, The second intermediate parity check value is the first intermediate parity check value.
27. The at least one non-transitory machine-readable medium according to claim 22, wherein, The persistent storage medium includes a solid-state drive.
Citation Information
Patent Citations
Information processing system, storage apparatus and storage device
US20170322845A1