System and method for parity-based fault protection of a storage device

By performing XOR operations in the SSD's storage device controller, the problem of inefficient parity protection in SSD is solved, efficient data protection and recovery is achieved, system cost and power consumption are reduced, and system performance is improved.

CN113971104BActive Publication Date: 2025-07-18KIOXIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110838581.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-14
Filing Date
2021-07-23
Publication Date
2025-07-18
Estimated Expiration
2041-07-23

AI Technical Summary

Technical Problem

The existing parity-based protection schemes are difficult to adapt to the evolution of system architecture in solid-state drives (SSDs), resulting in inefficient performance, and traditional dedicated disk array controllers (DACs) cannot effectively utilize the high-performance features of SSDs, increasing storage costs and computing overhead.

Method used

In the storage system, by performing XOR operations in the controller of the storage device, the parity data is updated and restored directly on the SSD, reducing the calculation and data movement at the host level, and optimizing interface connections with the NVMe interface to achieve data protection and recovery.

Benefits of technology

It improves input/output efficiency, host CPU efficiency, memory resource efficiency and data transmission efficiency, reduces overall system cost and power consumption, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971104B_ABST
    Figure CN113971104B_ABST
Patent Text Reader

Abstract

The present invention generally relates to systems and methods for parity-based fault protection for storage devices. The various embodiments described herein relate to systems and methods for providing data protection and recovery against drive failures, including: receiving a write request by a storage device from a host operatively coupled to the storage device; and determining an XOR result by the storage device rather than the host by performing an XOR operation on new data and existing data. The new data is received from the host. The existing data is stored in a non-volatile storage device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 056,411, filed Jul. 24, 2020, the entire contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present invention generally relates to systems, methods, and non-transitory processor-readable media for data protection and recovery from drive failures in data storage devices. Background Art

[0004] Redundant Array of Inexpensive Disks (RAID) can be implemented on non-volatile memory device-based drives to achieve protection against drive failures. Various forms of RAID can be broadly classified based on whether data is replicated or parity is protected. In terms of storage cost, replication is more expensive because replication doubles the number of required devices.

[0005] On the other hand, parity protection generally requires a lower storage cost than replication. In the case of RAID 5, one additional device is needed to provide protection against a single device failure at a given time by maintaining parity data for at least two data devices. When RAID 5 parity protection is employed, the additional storage cost, which is a certain percentage of the total cost, generally decreases as the number of devices protected in the RAID group increases.

[0006] In the case of RAID 6, which provides protection against up to two device failures simultaneously, two additional devices are needed to maintain parity data for at least two data devices. Similarly, when RAID 6 parity protection is employed, the additional storage cost, which is a certain percentage of the total cost, decreases as the number of devices protected in the RAID group increases. To mitigate the risk of failure of the drive storing the parity data, the drive storing the parity data is rotated.

[0007] Other variations of parity protection include combining replication and parity protection (e.g., in RAID 51 and RAID 61), varying the stripe size used between devices to match a given application, etc. Summary of the Invention

[0008] In some arrangements, a storage device includes a non-volatile storage device and a controller configured to receive a write request from a host operatively coupled to the storage device and determine an XOR result by performing an XOR operation on new data and existing data. The new data is received from the host. The existing data is stored in the non-volatile storage device.

[0009] In some arrangements, a method includes: receiving, by a controller of a storage device operatively coupled to a host via an interface, a write request from the host; and determining, by the controller, an XOR result by performing an XOR operation on new data and existing data. The new data is received from the host. The existing data is stored in a non-volatile storage device.

[0010] In some arrangements, a non-transitory computer-readable medium includes computer-readable instructions that, when executed, cause a processor to perform operations including: receiving a write request from a host operatively coupled to the storage device and determining an XOR result by performing an XOR operation on new data and existing data, where the new data is received from the host and the existing data is stored in a non-volatile storage device. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A block diagram showing an example of a system including a storage device and a host according to some embodiments.

[0012] Figure 2A is a block diagram illustrating an example method for performing parity update according to some embodiments.

[0013] Figure 2B is a flowchart illustrating an example method for performing parity update according to some embodiments.

[0014] Figure 3A is a block diagram illustrating an example method for performing data update according to some embodiments.

[0015] Figure 3B is a flowchart illustrating an example method for performing data update according to some embodiments.

[0016] Figure 4A is a block diagram illustrating an example method for performing data recovery according to some embodiments.

[0017] Figure 4B is a flowchart illustrating an example method for performing data recovery according to some embodiments.

[0018] Figure 5Ais a block diagram illustrating an example method for bringing a spare storage device into use according to some embodiments.

[0019] Figure 5B is a flowchart illustrating an example method for bringing a spare storage device into use according to some embodiments.

[0020] Figure 6 is a process flowchart illustrating an example method for providing data protection and recovery in the event of a drive failure according to some embodiments.

[0021] Figure 7 is a schematic diagram illustrating a host-side view for updating data according to some embodiments.

[0022] Figure 8 is a schematic diagram illustrating the placement of parity data according to some embodiments. DETAILED DESCRIPTION

[0023] Parity-based protection faces various challenges. Currently, the vast majority of embodiments are achieved through a dedicated disk array controller (DAC). The DAC calculates parity data by performing an exclusive OR (XOR) operation on each data stripe of each data disk in a given RAID group and stores the resulting parity data on one or more parity disks. The DAC is typically attached to the main central processing unit (CPU) via a high-speed peripheral component interconnect (PCIe) bus or network, while the DAC uses storage-specific interconnects and protocols (interfaces) (such as, but not limited to, AT attachment (ATA), Small Computer System Interface (SCSI), Fibre Channel, and Serial Attached SCSI (SAS)) to connect to and communicate with the disks. With storage-specific interconnects, a dedicated hardware controller is required to translate between the PCIe bus and the storage interface (such as SCSI or Fibre Channel).

[0024] The applicant has recognized that the evolution of non-volatile memory-based storage devices (such as solid state drives (SSDs)) has fundamentally changed the system architecture in that the storage device is directly attached to the PCIe bus via a high-speed non-volatile memory (NVMe) interface, thus eliminating inefficiencies in the path and optimizing cost, power, and performance. Although the functionality of the DAC is still required for SSD fault protection, the DAC functionality has migrated from a dedicated hardware controller to software running on a general-purpose CPU.

[0025] When the access time in a hard disk drive (HDD) is in milliseconds, the inefficiency of the DAC is not exposed. The emergence of SSDs has reduced data access times, thus placing stricter requirements on the DAC to derive and aggregate the performance of a large number of SSDs. The applicant has recognized that as the access time to SSDs is reduced by several orders of magnitude to reach tens of microseconds, conventional implementations of the DAC become inefficient as the performance of the DAC translates to the performance of the SSDs.

[0026] Due to the rise of NVMe, the interface for HDDs needs to be upgraded to be compatible with SSDs. Such an interface defines the way to deliver commands, return status, and exchange data between the host and the storage device. The interface can optimize and simplify the connection directly to the CPU without being troubled by an intermediate interface converter.

[0027] In addition, as the adoption of SSDs has started to increase, the cost (per GB) of HDDs has decreased significantly, partly due to the increased capacity provided in each disk of the HDD to differentiate from SSDs in the market. However, the increased capacity comes at the cost of performance. Therefore, DAC vendors have shifted from using parity-based protection for HDDs to replication-based protection. The data storage and access times of HDDs have been slower than those of SSDs, so packing more capacity makes the average performance of HDDs worse. In this regard, DAC vendors do not want to slow down HDDs further by using parity-based protection. Therefore, replication-based protection has in fact been almost universally used in standard DACs for HDDs. When SSDs come with requirements for several orders of magnitude greater improvement that the DAC has to address, it is also opportunistic and timely for the DAC to simply reuse replication-based protection for SSDs.

[0028] Therefore, parity-based protection for SSDs has not kept up with the architectural changes occurring at the system level. Additionally, cost barriers and access barriers have been established in the form of special stock keeping units (SKUs) on the main CPU, with limited availability for selecting customers. DAC vendors have lost the freedom to provide parity-based protection for SSDs, the same freedom that DAC vendors have for replication-based SSDs.

[0029] Therefore, it has become even more difficult to implement parity-based protection such as RAID 5 and RAID 6 for SSDs.

[0030] Conventionally, for storage devices (e.g., SSDs), RAID (e.g., RAID 5 and RAID 6) redundancy is created by relying on the host to perform XOR calculations and update parity data on the SSD. The SSD performs its normal functions of reading data from the storage medium (e.g., memory array) or writing data to the storage medium (e.g., memory array) without knowing whether the data is parity data. Thus, in RAID 5 and RAID 6, the computational overhead and extra data generation and movement typically become performance bottlenecks on the storage medium.

[0031] The arrangements disclosed herein relate to parity-based protection schemes, which are cost-effective solutions for SSD fault protection without compromising the need to deliver business requirements more quickly. The present invention addresses the challenges with respect to parity-based protection while creating a solution suitable for current system architectures and evolutionary changes. In some arrangements, the present invention relates to collaboratively performing data protection and recovery operations between two or more elements of a storage system. While non-volatile memory devices are presented herein as examples, the disclosed schemes can be implemented on any storage system or device that is connected to a host via an interface and stores data temporarily or permanently for later retrieval by the host.

[0032] To help illustrate the current embodiments, Figure 1 FIG. 1 shows a block diagram of a system including a storage device 100 coupled to a host 101 according to some embodiments. In some instances, the host 101 may be a user device operated by a user. The host 101 may include an operating system (OS) configured to provide a file system and applications that use the file system. The file system communicates via a suitable wired or wireless communication link or network with the storage device 100 (e.g., the controller 110 of the storage device 100) to manage data storage in the storage device 100.

[0033] In this regard, the file system of host 101 uses a suitable interface 140 to a communication link or network to send data to and receive data from storage device 100. Interface 140 allows the software of host 101 (e.g., the file system) to communicate with storage device 100 (e.g., controller 110). Storage device 100 (e.g., controller 110) is directly operatively coupled to the PCIe bus of host 101 via interface 140. Although interface 140 is conceptually shown as a dashed line between host 101 and storage device 100, the interface may include one or more controllers, one or more namespaces, ports, transport mechanisms, and their connectivity. To send and receive data, the software or file system of host 101 uses a storage data transfer protocol running on the interface to communicate with the storage device. Examples of the protocol include but are not limited to SAS, SATA, and NVMe protocols. Interface 140 includes hardware (e.g., a controller) implemented on host 101, storage device 100 (e.g., controller 110), or another device operatively coupled to host 101 and / or storage device 100 via one or more suitable networks. Interface 140 and the storage protocol running thereon also include software and / or firmware executed on the hardware.

[0034] In some instances, storage device 100 is located in a data center (not shown for simplicity). The data center may include one or more platforms, each of which supports one or more storage devices (e.g., but not limited to, storage device 100). In some embodiments, the storage devices within a platform are connected to a top-of-rack (TOR) switch and can communicate with each other via the TOR switch or another suitable in-platform communication mechanism. In some embodiments, at least one router may facilitate communication among storage devices in different platforms, racks, or cabinets via suitable networking fiber. Examples of storage device 100 include non-volatile devices such as, but not limited to, solid state drives (SSD), non-volatile dual in-line memory modules (NVDIMM), universal flash storage devices (UFS), secure digital (SD) devices, etc.

[0035] The storage device 100 includes at least a controller 110 and a memory array 120. Other components of the storage device 100 are not shown for simplicity. The memory array 120 includes NAND flash memory devices 130a to 130n. Each of the NAND flash memory devices 130a to 130n includes one or more individual NAND flash dies, which are non-volatile memories (NVMs) capable of retaining data without power. Thus, the NAND flash memory devices 130a to 130n refer to multiple NAND flash memory devices or dies within the storage device 100. Each of the NAND flash memory devices 130a to 130n includes one or more dies, and each of the one or more dies has one or more planes. Each plane has multiple blocks, and each block has multiple pages.

[0036] Although the NAND flash memory devices 130a to 130n are shown as examples of the memory array 120, other examples of non-volatile memory technologies for implementing the memory array 120 include, but are not limited to, dynamic random access memory (DRAM), magnetic random access memory (MRAM), phase change memory (PCM), ferroelectric RAM (FeRAM), etc. The arrangements described herein can be similarly implemented on memory systems using such memory technologies and other suitable memory technologies.

[0037] Examples of the controller 110 include, but are not limited to, an SSD controller (e.g., a client SSD controller, a data center SSD controller, an enterprise SSD controller, etc.), a UFS controller, or an SD controller, etc.

[0038] The controller 110 can combine raw data storage devices in the multiple NAND flash memory devices 130a to 130n such that those NAND flash memory devices 130a to 130n are used as a single storage device. The controller 110 can include a processor, a microcontroller, buffers (e.g., buffer 112), an error correction system, a data encryption system, a flash translation layer (FTL), and a flash interface module. These functions can be implemented in hardware, software, and firmware or any combination thereof. In some arrangements, the software / firmware of the controller 110 can be stored in the memory array 120 or any other suitable computer-readable storage medium.

[0039] The controller 110 includes suitable processing and memory capabilities for performing the functions described herein and other functions. As described, the controller 110 manages various features of the NAND flash memory devices 130a to 130n, including but not limited to, I / O handling, reading, writing / programming, erasing, monitoring, logging, error handling, garbage collection, wear leveling, logical to physical address mapping, data protection (encryption / decryption), etc. Thus, the controller 110 provides visibility to the NAND flash memory devices 130a to 130n.

[0040] As shown, the controller 110 includes a buffer 112 (driver buffer). The buffer 112 is local memory of the controller 110. In some instances, the buffer 112 is a volatile storage device. In some instances, the buffer 112 is a non-volatile persistent storage device. Examples of the buffer 112 include but are not limited to, RAM, DRAM, static RAM (SRAM), magnetic RAM (MRAM), phase change memory (PCM), etc. The buffer 112 may refer to multiple buffers each configured to store different types of data. Different types of data include new data received from the host 101 to be written to the memory array 120, transient XOR data received from the host 101, old data read from the memory array 120, saved data read from the memory array 120, XOR results, etc.

[0041] In one instance, in response to receiving data (e.g., new data, transient XOR data, etc.) from the host 101 (via the host interface 140), the controller 110 confirms the write command to the host 101 after writing the data to the buffer 112. The controller 110 may write the data stored in the buffer 112 to the memory array 120 (e.g., the NAND flash memory devices 130a to 130n). Once the write to the physical address in the memory array 120 is complete, the FTL updates the mapping between the logical address used by the host to associate with the data and the physical address used by the controller to identify the physical location of the data. In another instance, the buffer 112 may also store data read from the memory array 120.

[0042] The host 101 includes a buffer 102 (e.g., host buffer). The buffer 102 is local memory of the host 101. In some instances, the buffer 102 is a volatile storage device. Examples of the buffer 102 include but are not limited to RAM, DRAM, SRAM, etc. The buffer 102 may refer to multiple buffers each configured to store different types of data. Different types of data include new data to be sent to the storage device 100, transient XOR data to be sent to the storage device 100, previous data to be sent to the storage device 100, etc.

[0043] Although non-volatile memory devices (e.g., NAND flash memory devices 130a to 130n) are presented herein as examples, the disclosed solutions may be implemented on any storage system or device connected to a host 101 via an interface, where this system stores data for the host 101 either temporarily or permanently for later retrieval.

[0044] In some arrangements, the host 101 and the storage device 100 cooperate to perform XOR calculations not only for parity data reads and writes but also for non-parity data reads and writes. Specifically, instead of receiving the XOR result from the host 101, the storage device 100 (e.g., the controller 110) may be configured to perform the XOR calculation. Thus, the host 101 does not need to consume additional memory resources for such operations, does not need to consume CPU cycles to send additional commands for performing the XOR calculation, does not need to allocate hardware resources for the associated direct memory access (DMA) transfers, does not need to consume submission and completion queues for additional commands, and does not need to consume additional bus / network bandwidth.

[0045] Given the improvements obtained by having the storage device 100 perform the XOR calculation internally within the storage device 100, not only is the overall system cost (for a system comprising the host 101 and the storage device 100) reduced, but also the system performance is improved. Thus, compared to conventional RAID redundancy mechanisms, the present invention involves allowing the storage device 100 to offload functionality and repartition functionality, thereby achieving fewer operations and less data movement.

[0046] In some arrangements, the XOR calculation may be performed within the controller 110 and via the interface 140 (e.g., an NVMe interface) to update parity data stored in the memory array 120. The storage device 100 is a parity drive.

[0047] Traditionally, to update parity data (or parity) on a parity drive in a RAID 5 group, the following steps are performed. Host 101 submits an NVMe read request to controller 110 via interface 140. In response, controller 110 performs a NAND read into the drive buffer. In other words, controller 110 reads the data (old existing parity data) requested in the read request from memory array 120 (one or more of NAND flash memory devices 130a to 130n) and stores the data in buffer 112. Controller 110 transfers the data from buffer 112 across interface 140 into buffer 102 (e.g., old data buffer). The old data buffer of host 101 thus stores the old data read from memory array 120. Host 101 then performs an XOR operation between: (i) the data (referred to as transient XOR data) that host 101 had previously calculated and that resides in buffer 102 (e.g., transient XOR buffer) and (ii) the old data read from memory array 120 and stored in the old data buffer of host 101. The result (referred to as new data) is then stored in buffer 102 (e.g., new data buffer). In some cases, the new data buffer may potentially be the same as the old data buffer or the transient XOR buffer, as the new data may replace the existing content in those buffers to conserve memory resources. Host 101 then submits an NVMe write request to controller 110 and presents the new data from the new data buffer to controller 110. In response, controller 110 then performs a data transfer to obtain the new data from the new data buffer of host 101 across interface 140 and stores the new data in drive buffer 112. Controller 110 then updates the old existing data by writing the new data into memory array 120 (e.g., one or more of NAND flash memory devices 130a to 130n). The new data shares the same logical address (e.g., the same logical block address (LBA)) as the old data and has a different physical address (e.g., stored in different NAND pages of NAND flash memory devices 130a to 130n), due to the nature of the operation of NAND flash memory. Controller 110 also updates the logical-to-physical mapping table to record the new physical address.

[0048] On the other hand, Figure 2A is a block diagram illustrating an example method 200a for performing parity update according to some embodiments. Referring to Figures 1 to 2A, compared with the conventional parity check update method described above, method 200a provides improved input / output (I / O) efficiency, host CPU efficiency, memory resource efficiency, and data transfer efficiency. Method 200a can be executed by host 101 and storage device 100. Communication between host 101 and storage device 100 (e.g., data and command communication) can be executed via interface 140.

[0049] In Figure 2A , the boxes and processes shown above the dashed line representing interface 140 relate to host 101, while the boxes and processes shown below the dashed line relate to storage device 100. Host buffer (new data) 201 is a specific implementation of buffer 102. NAND page (old data) 203 and NAND page (XOR result) 206 are different pages in NAND flash memory devices 130a to 130n.

[0050] In method 200a, at 211, host 101 submits an NVMe write request to controller 110 via interface 140. Host 101 presents host buffer (new data) 201 to be written to controller 110. In response, controller 110 performs a data transfer to obtain new data (parity data) from host buffer (new data) 201 across interface 140 and stores the new data in drive buffer (new data) 202. The write request includes the logical address (e.g., LBA) of the new data.

[0051] At 212, controller 110 performs a NAND read into drive buffer (old data) 204. In other words, controller 110 reads old and existing data corresponding to the logical address in the read request from memory array 120 (e.g., one or more NAND pages (old data) 203) and stores the old data in drive buffer (old data) 204. One or more NAND pages (old data) 203 are pages in one or more of NAND flash memory devices 130a to 130n. The new data and the old data are parity data (e.g., one or more parity bits). In other words, the old data (old parity data) is updated to the new data (new parity data).

[0052] At 213, the controller 110 performs an XOR operation between the new data stored in the drive buffer (new data) 202 and the old data stored in the drive buffer (old data) 204 to determine the XOR result and stores the XOR result in the drive buffer (XOR result) 205. In some arrangements, the drive buffer (new data) 202, the drive buffer (old data) 204, and the drive buffer (XOR result) 205 are separate buffers of the buffer 112 and specific embodiments. In other arrangements, to conserve memory resources, the drive buffer (XOR result) 205 can be the same as the drive buffer (old data) 204 or the drive buffer (new data) 202, such that the XOR result can overwrite the content of the drive buffer (old data) 204 or the drive buffer (new data) 202.

[0053] At 214, the controller 110 then updates the old data with the new data by writing the XOR result into the NAND page (XOR result) 206. The controller 110 (e.g., the FTL) updates the logical-to-physical addressing mapping table to make the physical address of the NAND page (XOR result) 206 correspond to the logical address. The controller 110 marks the physical address of the NAND page (old data) 203 for garbage collection.

[0054] Figure 2B is a flowchart illustrating an example method 200b for performing parity update according to some embodiments. Refer to Figures 1 to 2B , method 200b corresponds to method 200a. Method 200b can be executed by the controller 110 of the storage device 100.

[0055] At 221, the controller 110 receives a write request from a host 101 operatively coupled to the storage device 100. At 222, in response to receiving the write request, the controller 110 transfers the new data (new parity data) from the host 101 (e.g., from the host buffer (new data) 201) to the new data drive buffer (e.g., the drive buffer (new data) 202) across the interface 140. Thus, the controller 110 receives the new data corresponding to the logical address identified in the write request from the host 101. At 223, the controller 110 performs a read operation to read the existing (old) data (existing old parity data) from the non-volatile storage device (e.g., from the NAND page (old data) 203) into the existing data drive buffer (e.g., the drive buffer (old data) 204).

[0056] At 224, the controller 110 determines an XOR result by performing an XOR operation on new data and existing data. At 225, after determining the XOR result, the controller 110 temporarily stores the XOR result in an XOR result driver buffer (e.g., driver buffer (XOR result) 205). At 226, the controller 110 writes the XOR result stored in the XOR result driver buffer to a non-volatile storage device (e.g., NAND page (XOR result) 206). As described, the new data and the existing data correspond to the same logical address. The existing data is located at a first physical address of the non-volatile storage device (e.g., located at NAND page (old data) 204). Writing the XOR result to the non-volatile storage device includes: writing the XOR result to a second physical address of the non-volatile storage device (e.g., located at NAND page (XOR result) 206) and updating the logical-to-physical (L2P) mapping so that the logical address corresponds to the second physical address.

[0057] Methods 200a and 200b improve conventional parity data update methods by performing XOR operations in drive hardware (e.g., as described in the hardware of storage device 100) to completely avoid any need for XOR operations at the host level. Compared with conventional parity update methods, methods 200a and 200b improve I / O efficiency, host CPU efficiency, memory resource efficiency, and data transfer efficiency.

[0058] Regarding I / O performance efficiency, the host 101 only needs to submit one request (write request, at 211 and 221) to update parity data instead of two requests. In some instances, the work involved in each request includes: 1) the host 101 writes a command to the submission queue; 2) the host 101 writes an updated submission queue tail pointer to a doorbell register; 3) the storage device 100 (e.g., controller 110) extracts the command from the submission queue; 4) the storage device 100 (e.g., controller 110) processes the command; 5) the storage device 100 (e.g., controller 110) writes details about the completion status to the completion queue; 6) the storage device 100 (e.g., controller 110) notifies the host 101 that the command has been completed; 7) the host 101 processes the completion; and 8) the host 101 writes an updated completion queue head pointer to the doorbell register.

[0059] Therefore, the host 101 does not need to read the existing parity data and perform the XOR operation of the transient XOR data and the existing parity data. Excluding all the elapsed time for extracting commands, processing commands, extracting data from the storage medium (e.g., the memory array 120), and completing the XOR operation within the storage device 100, the mechanism disclosed herein consumes nearly 10% of the total elapsed time to extract 4KB of data from the storage device 100. Thus, the current arrangement can reduce the number of host requests by half (from two to one), which represents a significant efficiency improvement.

[0060] Regarding host CPU efficiency, host computations are more expensive than those of the storage device 100 because the cost of the host CPU is significantly higher than that of the CPU of the storage device 100. Therefore, saving computational cycles for the host CPU results in higher efficiency. Research has estimated that the number of CPU clocks required per NVMe request is approximately 34,000. Thus, whenever parity updates need to be performed, a CPU savings of 34,000 clocks occurs. For comparison purposes, a 12Gb SAS interface request for an SSD consumes approximately 79,000 clocks / request. With NVMe interface technology, this requirement can be reduced to approximately 34,000, saving approximately 45,000 clock cycles. Considering the elimination of XOR computations and the reduction in requests at the host level, the efficiency improvement can be comparable to that provided by the NVMe interface compared to the SAS interface.

[0061] Regarding memory resource efficiency, in addition to the savings at the CPU level for the host 101, there are also savings to be had in terms of memory consumption. DRAM memory remains a precious resource for the host 101, not only because only a limited amount can be added to the host 101 due to the limited dual in-line memory module (DIMM) slots, but also due to the capacity scaling limitations of the DRAM technology itself. Additionally, modern applications such as machine learning, in-memory databases, big data analytics, etc. increase the need for additional memory at the host 101. Thus, given that DRAM cannot meet this increased memory need, a new class of devices called storage class memory (SCM) has emerged to bridge the gap. Although this technology is still in its infancy, the vast majority of existing systems still seek solutions that can help reduce the consumption of memory resources without compromising other attributes such as cost or performance. The current arrangement reduces memory consumption by eliminating the need to potentially allocate up to two buffers in the host 101 (a savings of up to 200% / request), thus reducing costs.

[0062] Regarding data transfer efficiency, the number of data transfers across the NVMe interface for copying data from a drive buffer (e.g., buffer 112) to a host buffer (e.g., buffer 102) can be reduced by half, which reduces the consumption of the desired hardware resources for DMA transfers and the utilization of the PCIe bus / network. This resource reduction not only reduces power consumption but also improves performance.

[0063] In some arrangements, XOR calculations can be performed within controller 110 and via interface 140 (e.g., the NVMe interface) to update data stored in memory array 120. Storage device 100 is a data drive.

[0064] Traditionally, to update data (regular, non-parity data) on a data drive in a RAID 5 group, the following steps are performed. Host 101 submits an NVMe read request via interface 140 to controller 110. In response, controller 110 performs a NAND read into the drive buffer. In other words, controller 110 reads the data requested in the read request from memory array 120 (one or more of NAND flash memory devices 130a to 130n) and stores the data in buffer 112. Controller 110 transfers the data from buffer 112 across interface 140 into buffer 102 (e.g., the old data buffer). The old data buffer of host 101 thus stores the old data read from memory array 120. Host 101 then submits an NVMe write request to controller 110 and presents the new data to be written by controller 110 in the new data buffer of host 101. In response, controller 110 performs a data transfer to transfer the new data from the new data buffer of host 101 across the NVMe interface into the drive buffer (e.g., buffer 112). Controller 110 then updates the old existing data by writing the new data into memory array 120 (e.g., one or more of NAND flash memory devices 130a to 130n). The new data shares the same logical address (e.g., LBA) as the old data and has a different physical address (e.g., stored in different NAND pages of NAND flash memory devices 130a to 130n). Host 101 then performs an XOR operation between: (i) the new data that has been resident in the new data buffer of host 101 and (ii) the existing data read from storage device 100 and resident in the old data buffer of host 101. Host 101 stores the result of the XOR operation (referred to as the transient XOR data) in the transient XOR host buffer of host 101. In some cases, the transient XOR buffer can potentially be the same as the old data buffer or the new data buffer, as the transient XOR data can replace the existing content in those buffers to conserve memory resources.

[0065] On the other hand, Figure 3A is a block diagram illustrating an example method 300a for performing data updates according to some embodiments. Referring to Figures 1 to 3A , similar to methods 200a and 200b, method 300a provides improved I / O efficiency, host CPU efficiency, and memory resource efficiency compared to the conventional data update methods described above. Method 300a can be executed by host 101 and storage device 100. Communication between host 101 and storage device 100 (e.g., data and command communication) can be performed via interface 140.

[0066] At Figure 3A , the boxes and processes shown above the dashed line representing interface 140 pertain to host 101, while the boxes and processes shown below the dashed line pertain to storage device 100. Buffer 102 includes host buffer (new data) 301 and host buffer (transient XOR) 307. NAND page (old data) 303 and NAND page (new data) 305 are different pages in NAND flash memory devices 130a to 130n.

[0067] In method 300a, at 311, host 101 submits an NVMe write request to controller 110 via interface 140. Host 101 presents host buffer (new data) 301 to be written to controller 110. In response, controller 110 performs a data transfer to obtain new data (conventional, non-parity data) from host buffer (new data) 301 across interface 140 and stores the new data in drive buffer (new data) 302. The write request includes the logical address of the new data (e.g., LBA).

[0068] At 312, controller 110 performs a NAND read into drive buffer (old data) 304. In other words, controller 110 reads old and existing data corresponding to the logical address in the read request from memory array 120 (e.g., one or more NAND pages (old data) 303) and stores the old data in drive buffer (old data) 304. One or more NAND pages (old data) 303 are pages in one or more of NAND flash memory devices 130a to 130n. The new data and the old data are data (e.g., conventional, non-parity data). In other words, the old data is updated to the new data.

[0069] At 313, the controller 110 then updates the old data with the new data by writing the new data from the drive buffer (new data) 302 into the NAND page (new data) 305. The controller 110 (e.g., FTL) updates the addressing mapping table so that the physical address of the NAND page (new data) 305 corresponds to the logical address. The controller 110 marks the physical address of the NAND page (old data) 303 for garbage collection.

[0070] At 314, the controller 110 performs an XOR operation between the new data stored in the drive buffer (new data) 302 and the old data stored in the drive buffer (old data) 304 to determine an instantaneous XOR result and stores the instantaneous XOR result in the drive buffer (instantaneous XOR) 306. In some arrangements, the drive buffer (new data) 302, the drive buffer (old data) 304, and the drive buffer (instantaneous XOR) 306 are separate buffers of the buffer 112 and specific embodiments. In other arrangements, to conserve memory resources, the drive buffer (instantaneous XOR) 306 may be the same as the drive buffer (old data) 304 or the drive buffer (new data) 302, such that the instantaneous XOR result may overwrite the content of the drive buffer (old data) 304 or the drive buffer (new data) 302.

[0071] At 315, the controller 110 transfers the instantaneous XOR result from the drive buffer (instantaneous XOR) 306 to the host buffer (instantaneous XOR) 307 across the interface 140. For example, the host 101 may submit an NVMe read request to the controller 110 for the instantaneous XOR result, and in response, the controller 110 transfers the instantaneous XOR result from the drive buffer (instantaneous XOR) 306 to the host buffer (instantaneous XOR) 307 across the interface 140.

[0072] Figure 3B is a flowchart illustrating an example method 300b for performing data updates according to some embodiments. Refer to Figure 1 、 2A and 3B, the method 300b corresponds to the method 300a. The method 300b may be executed by the controller 110 of the storage device 100.

[0073] At 321, the controller 110 receives a write request from the host 101 operatively coupled to the storage device 100. At 322, in response to receiving the write request, the controller 110 transfers new data (new normal, non-parity data) from the host 101 (e.g., from the host buffer (new data) 301) to the new data driver buffer of the storage device 100 (e.g., the driver buffer (new data) 302) across the interface 140. Thus, the controller 110 receives new data corresponding to the logical address identified in the write request from the host 101. At 323, the controller 110 performs a read operation to read existing (old) data from the non-volatile storage device (e.g., from the NAND page (old data) 303) into the existing data driver buffer (e.g., the driver buffer (old data) 304). The existing data has the same logical address as the new data identified in the write request.

[0074] At 324, the controller 110 writes the new data stored in the new data driver buffer of the storage device 100 to the non-volatile storage device (e.g., the NAND page (new data) 305). As described, the new data and the existing data correspond to the same logical address. The existing data is located at a first physical address of the non-volatile storage device (e.g., at the NAND page (old data) 303). Writing the new data to the non-volatile storage device includes: writing the new data to a second physical address of the non-volatile storage device (e.g., at the NAND page (new data) 305) and updating the L2P mapping so that the logical address corresponds to the second physical address. Blocks 323 and 324 may be performed in any suitable order or simultaneously.

[0075] At 325, the controller 110 determines an XOR result by performing an XOR operation on the new data and the existing data. The XOR result is referred to as the instantaneous XOR result. At 326, after determining the instantaneous XOR result, the controller 110 temporarily stores the instantaneous XOR result in the instantaneous XOR result driver buffer (e.g., the driver buffer (instantaneous XOR) 306). At 327, the controller 110 transfers the instantaneous XOR result from the instantaneous XOR result driver buffer to the host 101 (e.g., to the host buffer (instantaneous XOR) 307) across the interface 140.

[0076] Since updating data on the data drive (e.g., methods 300a and 300b) is subsequently followed by corresponding updates of the parity on the parity drive (e.g., methods 200a and 200b) to maintain the integrity of RAID 5 group protection, the efficiency is actually the sum of the two processes. In other words, each write in the conventional mechanism results in four I / O operations (reading old data, reading old parity, writing new data, and writing new parity). In the mechanism described herein, the number of I / O operations is reduced by half to two (writing new data, writing new parity).

[0077] In some arrangements, the XOR calculation can be performed within the controller 110 and via the interface 140 (e.g., NVMe interface) to recover data of a failed device in the RAID 5 group.

[0078] Traditionally, to recover data from a failed device in a RAID 5 group, the following steps are performed. The host 101 submits an NVMe read request via the interface 140 to the controller 110 of the first storage device in a series of storage devices of the RAID 5 group. In the RAID 5 group, the nth storage device is the failed device, and the first storage device to the (n - 1)th storage device are functional devices. Each storage device in the RAID 5 group can be, for example but not limited to, the storage device 100. In response, the controller 110 of the first storage device performs a NAND read into the drive buffer of the first storage device. In other words, the controller 110 reads the data requested in the read request from the memory array 120 (one or more of the NAND flash memory devices 130a to 130n) of the first storage device and stores the data in the buffer 112 of the first storage device. The controller 110 of the first storage device transfers the data from the buffer 112 of the first storage device across the interface 140 into the buffer 102 (e.g., previous data buffer) of the host 101.

[0079] Next, host 101 submits an NVMe read request via interface 140 to the controller 110 of the second storage device of the RAID 5 group. In response, the controller 110 of the second storage device performs a NAND read into the drive buffer of the second storage device. In other words, the controller 110 of the second storage device reads the data requested in the read request from the memory array 120 (one or more of the NAND flash memory devices 130a to 130n) of the second storage device and stores the data in the buffer 112 of the second storage device. The controller 110 of the second storage device transfers the data from the buffer 112 of the second storage device across interface 140 into the buffer 102 (e.g., the current data buffer) of host 101. Host 101 then performs an XOR operation between: (i) the data in the previous data buffer of host 101 and (ii) the data in the current data buffer of host 101. The result (the instantaneous XOR result) is then stored in the instantaneous XOR buffer of host 101. In some cases, the instantaneous XOR buffer of host 101 may potentially be the same as the previous data buffer or the current data buffer, as the instantaneous XOR data may replace the existing content in those buffers to conserve memory resources.

[0080] Next, host 101 submits an NVMe read request via interface 140 to the controller 110 of the next storage device of the RAID 5 group to read the current data of the next storage device in the manner described with respect to the second storage device. Host 101 then performs an XOR operation between: (i) the data in the previous data buffer, which is the instantaneous XOR result determined in the previous iteration involving the previous storage device; and (ii) the current data of the next storage device. Such a process is repeated until host 101 determines the recovered data by performing an XOR operation between the current data of the (n - 1)th storage device and the instantaneous XOR result determined in the previous iteration involving the (n - 2)th storage device.

[0081] On the other hand, Figure 4A is a block diagram illustrating an example method 400a for performing data recovery according to some embodiments. Referring to Figure 1 and 4A , similar to methods 200a and 200b, method 400a provides improved host CPU efficiency and memory resource efficiency compared to the conventional data update methods described above. Method 400a can be executed by host 101 and storage device 100. Communication between host 101 and storage device 100 (e.g., data and command communication) can be performed via interface 140.

[0082] In Figure 4AIn it, the boxes shown above the dashed line representing interface 140 and the processes involve host 101, while the boxes and processes shown below the dashed line involve storage device 100. Buffer 102 includes host buffer (previous data) 401 and host buffer (instantaneous XOR) 406. NAND page (stored data) 403 refers to one or more pages in NAND flash memory devices 130a to 130n.

[0083] Figure 4A Show an iteration of the recovery data of the failed nth device in a RAID 5 group. Each storage device in the RAID 5 group can be a storage device such as, but not limited to, storage device 100. Regarding the first storage device in the RAID 5 group, host 101 submits an NVMe read request for a logical address to the controller 110 of the first storage device of the RAID 5 group via interface 140. In response, the controller 110 of the first storage device performs a NAND read into the drive buffer of the first storage device. In other words, the controller 110 reads the start data corresponding to the logical address requested in the read request from the memory array 120 (one or more of NAND flash memory devices 130a to 130n) of the first storage device and stores the start data in buffer 112 of the first storage device. The controller 110 of the first storage device transfers the start data from buffer 112 of the first storage device to buffer 102 (e.g., host buffer (previous data) 401) of host 101 across interface 140.

[0084] Method 400a can be executed by one of the second to (n - 1) storage devices in the RAID 5 group. The storage device that executes method 400a is simply referred to as the current storage device 100. At 411, host 101 submits an NVMe read request to the controller 110 of the current storage device 100 via interface 140. Host 101 presents host buffer (previous data) 401 to be written to the controller 110. In response, the controller 110 performs a data transfer to obtain the previous data from host buffer (previous data) 401 across interface 140 and stores the previous data in drive buffer (new data) 402. The write request includes the logical address (e.g., LBA) of the previous data.

[0085] In an instance where the current storage device 100 is the second storage device of a RAID 5 group, the previous data is the starting data stored in the host buffer (previous data) 401 obtained from the first storage device. In an instance where the current storage device 100 is the third to the (n - 1)th storage device, the previous data refers to the transient XOR data stored in the host buffer (transient XOR) 406 of the previous storage device 100. The previous storage device immediately precedes the current storage device in the RAID 5 group.

[0086] At 412, the controller 110 performs a NAND read into the drive buffer (stored data) 404. In other words, the controller 110 reads the stored data corresponding to the logical address in the read request from the memory array 120 (e.g., one or more NAND pages (stored data) 403) and stores the stored data in the drive buffer (stored data) 404. The one or more NAND pages (stored data) 403 are pages in one or more of the NAND flash memory devices 130a to 130n.

[0087] At 413, the controller 110 performs an XOR operation between the previous data stored in the drive buffer (new data) 402 and the stored data stored in the drive buffer (stored data) 404 to determine a transient XOR result and stores the transient XOR result in the drive buffer (transient XOR) 405. In some arrangements, the drive buffer (new data) 402, the drive buffer (stored data) 404, and the drive buffer (transient XOR) 405 are separate buffers of the buffer 112 and specific embodiments. In other arrangements, to conserve memory resources, the drive buffer (transient XOR) 405 may be the same as the drive buffer (stored data) 404 or the drive buffer (new data) 402, such that the transient XOR result may overwrite the content of the drive buffer (stored data) 404 or the drive buffer (new data) 402.

[0088] At 414, the controller 110 transfers the instantaneous XOR result from the drive buffer (instantaneous XOR) 405 to the host buffer (instantaneous XOR) 406 across the interface 140. For example, the host 101 may submit an NVMe read request for the instantaneous XOR result to the controller 110, and in response, the controller 110 transfers the instantaneous XOR result from the drive buffer (instantaneous XOR) 405 to the host buffer (instantaneous XOR) 406 across the interface 140. The iteration for the current storage device 100 is completed at this time, and the instantaneous XOR result corresponds to the previous data of the next storage device after the current storage device in the next iteration. In the case where the current storage device is the (n - 1)th storage device, the instantaneous XOR result is in fact the recovered data of the failed nth storage device.

[0089] Figure 4B is a flowchart illustrating an example method 400b for performing data recovery according to some embodiments. Refer to Figure 1 、 4A 、and 4B, method 400b corresponds to method 400a. Method 400b may be executed by the controller 110 of the storage device 100.

[0090] At 421, the controller 110 receives a write request from the host 101 operatively coupled to the storage device 100. At 422, in response to receiving the write request, the controller 110 transfers the previous data from the host 101 (e.g., from the host buffer (new data) 401) to the new data drive buffer of the storage device 100 (e.g., the drive buffer (new data) 402) across the interface 140. Thus, the controller 110 receives the previous data corresponding to the logical address identified in the write request from the host 101. At 423, the controller 110 performs a read operation to read the existing (saved) data from the non-volatile storage device (e.g., from the NAND page (saved data) 403) into the existing data drive buffer (e.g., the drive buffer (saved data) 404). The saved data has the same logical address as the previous data.

[0091] At 424, the controller 110 determines the XOR result by performing an XOR operation on the previous data and the saved data. The XOR result is referred to as the instantaneous XOR result. At 425, after determining the instantaneous XOR result, the controller 110 temporarily stores the instantaneous XOR result in the instantaneous XOR result drive buffer (e.g., the drive buffer (instantaneous XOR) 405). At 426, the controller 110 transfers the instantaneous XOR result from the instantaneous XOR result drive buffer to the host 101 (e.g., to the host buffer (instantaneous XOR) 406) across the interface 140.

[0092] In some arrangements, the XOR calculation can be performed within the controller 110 and via the interface 140 (e.g., the NVMe interface) to put the spare storage device in the RAID 5 group into use.

[0093] Traditionally, to put the spare storage device in the RAID 5 group into use, the following operations are performed. The host 101 submits an NVMe read request via the interface 140 to the controller 110 of the first storage device in a series of storage devices in the RAID 5 group. In the RAID 5 group, the nth storage device is the spare device, and the first storage device to the (n - 1)th storage device are the currently functional devices. Each storage device in the RAID 5 group can be, for example but not limited to, the storage device 100. In response, the controller 110 of the first storage device performs a NAND read into the drive buffer of the first storage device. In other words, the controller 110 reads the data requested in the read request from the memory array 120 (one or more of the NAND flash memory devices 130a to 130n) of the first storage device and stores the data in the buffer 112 of the first storage device. The controller 110 of the first storage device transfers the data from the buffer 112 of the first storage device across the interface 140 into the buffer 102 (e.g., the previous data buffer) of the host 101.

[0094] Next, the host 101 submits an NVMe read request via the interface 140 to the controller 110 of the second storage device in the RAID 5 group. In response, the controller 110 of the second storage device performs a NAND read into the drive buffer of the second storage device. In other words, the controller 110 of the second storage device reads the data requested in the read request from the memory array 120 (one or more of the NAND flash memory devices 130a to 130n) of the second storage device and stores the data in the buffer 112 of the second storage device. The controller 110 of the second storage device transfers the data from the buffer 112 of the second storage device across the interface 140 into the buffer 102 (e.g., the current data buffer) of the host 101. The host 101 then performs an XOR operation between: (i) the data in the previous data buffer of the host 101 and (ii) the data in the current data buffer of the host 101. The result (the instantaneous XOR result) is then stored in the instantaneous XOR buffer of the host 101. In some cases, the instantaneous XOR buffer of the host 101 can potentially be the same as the previous data buffer or the current data buffer, since the instantaneous XOR data can replace the existing content in those buffers to conserve memory resources.

[0095] Next, the host 101 submits an NVMe read request via the interface 140 to the controller 110 of the next storage device in the RAID 5 group to read the current data of the next storage device in the manner described with respect to the second storage device. The host 101 then performs an XOR operation between: (i) the data in the previous data buffer, which is the instantaneous XOR result determined in the previous iteration involving the previous storage device; and (ii) the current data of the next storage device. Such a process is repeated until the host 101 determines the recovered data by performing an XOR operation between the current data of the (n - 1)th storage device and the instantaneous XOR result determined in the previous iteration involving the (n - 2)th storage device. The host 101 stores the recovered data in the recovered data buffer of the host 101.

[0096] The recovered data will be written to the spare nth storage device at the logical address. For example, the host 101 submits an NVMe write request to the nth device and presents the recovered data buffer of the host 101 to be written. In response, the nth storage device performs data transfer to obtain the recovered data from the host 101 by transferring the recovered data from the recovered data buffer of the host 101 to the drive buffer of the nth storage device across the NVMe interface. The controller 110 of the nth storage device then updates the old data stored in the NAND pages of the nth storage device with the recovered data by writing the recovered data from the drive buffer of the nth storage device into one or more new NAND pages. The controller 110 (e.g., FTL) updates the addressing mapping table to make the physical address of the new NAND page correspond to the logical address. The controller 110 marks the physical address of the NAND page storing the old data for garbage collection.

[0097] On the other hand, Figure 5A is a block diagram illustrating an example method 500a for bringing a spare storage device into use according to some embodiments. Refer to Figure 1 and 5A Similar to methods 200a and 200b, compared with the conventional data update method described above, method 500a provides improved host CPU efficiency and memory resource efficiency. Method 500a can be executed by the host 101 and the storage device 100. The communication (e.g., data and command communication) between the host 101 and the storage device 100 can be performed via the interface 140.

[0098] In Figure 5AIn the figure, the boxes and processes shown above the dashed line representing interface 140 involve host 101, while the boxes and processes shown below the dashed line involve storage device 100. Buffer 102 includes host buffer (previous data) 501 and host buffer (transient XOR) 506. NAND page (stored data) 503 refers to one or more pages in NAND flash memory devices 130a to 130n.

[0099] Figure 5A Shows an iteration that brings a spare nth device in a RAID 5 group into use. Each storage device in the RAID 5 group can be a storage device such as, but not limited to, storage device 100. Regarding the first storage device in the RAID 5 group, host 101 submits an NVMe read request for a logical address to the controller 110 of the first storage device via interface 140. In response, the controller 110 of the first storage device performs a NAND read into the drive buffer of the first storage device. In other words, the controller 110 reads start data corresponding to the logical address requested in the read request from the memory array 120 of the first storage device (one or more of NAND flash memory devices 130a to 130n) and stores the start data in buffer 112 of the first storage device. The controller 110 of the first storage device transfers the start data from buffer 112 of the first storage device to buffer 102 of host 101 (e.g., host buffer (previous data) 504) across interface 140.

[0100] Method 500a can be executed by one of the second to (n - 1) storage devices in the RAID 5 group. The storage device that executes method 500a is simply referred to as the current storage device 100. At 511, host 101 submits an NVMe read request to the controller 110 of the current storage device 100 via interface 140. Host 101 presents host buffer (previous data) 501 to be written to the controller 110. In response, the controller 110 performs a data transfer to obtain previous data from host buffer (previous data) 501 across interface 140 and stores the previous data in drive buffer (new data) 502. The write request includes the logical address (e.g., LBA) of the previous data.

[0101] In an example where the current storage device 100 is the second storage device of the RAID 5 group, the previous data is the start data stored in host buffer (previous data) 501 obtained from the first storage device. In an example where the current storage device 100 is the third to (n - 1) storage devices, the previous data refers to the transient XOR data stored in host buffer (transient XOR) 506 of the previous storage device 100. The previous storage device is immediately before the current storage device in the RAID 5 group.

[0102] At 512, the controller 110 performs a NAND read into the drive buffer (saved data) 504. In other words, the controller 110 reads the saved data corresponding to the logical address in the read request from the memory array 120 (e.g., one or more NAND pages (saved data) 503) and stores the saved data in the drive buffer (saved data) 504. The one or more NAND pages (saved data) 503 are pages in one or more of the NAND flash memory devices 130a to 130n.

[0103] At 513, the controller 110 performs an XOR operation between the previous data stored in the drive buffer (new data) 502 and the saved data stored in the drive buffer (saved data) 504 to determine an instantaneous XOR result and stores the instantaneous XOR result in the drive buffer (instantaneous XOR) 505. In some arrangements, the drive buffer (new data) 502, the drive buffer (saved data) 504, and the drive buffer (instantaneous XOR) 505 are separate buffers of the buffer 112 and specific implementations. In other arrangements, to conserve memory resources, the drive buffer (instantaneous XOR) 505 can be the same as the drive buffer (saved data) 504 or the drive buffer (new data) 502, such that the instantaneous XOR result can overwrite the content of the drive buffer (saved data) 504 or the drive buffer (new data) 502.

[0104] At 514, the controller 110 transfers the instantaneous XOR result from the drive buffer (instantaneous XOR) 505 to the host buffer (instantaneous XOR) 506 across the interface 140. For example, the host 101 can submit an NVMe read request for the instantaneous XOR result to the controller 110, and in response, the controller 110 transfers the instantaneous XOR result from the drive buffer (instantaneous XOR) 505 to the host buffer (instantaneous XOR) 506 across the interface 140. The iteration for the current storage device 100 is completed at this time, and the instantaneous XOR result corresponds to the previous data of the next storage device after the current storage device in the next iteration. In the case where the current storage device is the (n - 1)th storage device, the instantaneous XOR result is in fact the recovered data of the failed nth storage device.

[0105] Figure 5B is a flowchart illustrating an example method 500b for bringing a spare storage device into use according to some embodiments. Refer to Figure 1 、 5A and 5B, the method 500b corresponds to the method 500a. The method 500b can be executed by the controller 110 of the storage device 100.

[0106] At 521, the controller 110 receives a write request from the host 101 operatively coupled to the storage device 100. At 522, in response to receiving the write request, the controller 110 transfers the previous data from the host 101 (e.g., from the host buffer (new data) 501) to the new data driver buffer of the storage device 100 (e.g., the driver buffer (new data) 502) across the interface 140. Thus, the controller 110 receives the previous data corresponding to the logical address identified in the write request from the host 101. At 523, the controller 110 performs a read operation to read the existing (stored) data from the non-volatile storage device (e.g., from the NAND page (stored data) 503) into the existing data driver buffer (e.g., the driver buffer (stored data) 504). The stored data has the same logical address as the previous data.

[0107] At 524, the controller 110 determines the XOR result by performing an XOR operation on the previous data and the stored data. The XOR result is referred to as the instantaneous XOR result. At 525, after determining the instantaneous XOR result, the controller 110 temporarily stores the instantaneous XOR result in the instantaneous XOR result driver buffer (e.g., the driver buffer (instantaneous XOR) 505). At 526, the controller 110 transfers the instantaneous XOR result from the instantaneous XOR result driver buffer to the host 101 (e.g., to the host buffer (instantaneous XOR) 506) across the interface 140.

[0108] Figure 6 is a process flow diagram illustrating an example method 600 for providing data protection and recovery against drive failures according to some embodiments. Refer to Figures 1 to 6 , the method 600 is executed by the controller 110 of the storage device 100.

[0109] At 610, the controller 110 receives a write request from the host 101 operatively coupled to the storage device 100. The host 101 is operatively coupled to the storage device 100 through the interface 140. At 620, the controller 110 determines the XOR result by performing an XOR operation on the new data and the existing data. The new data is received from the host 101. The existing data is stored in the non-volatile storage device (e.g., stored in the memory array 120).

[0110] In some arrangements, in response to receiving a write request, the controller 110 transfers new data from the host 101 (e.g., from the buffer 102) across the interface 140 to the new data drive buffer of the storage device 100. The controller 110 performs a read operation to read existing data from the non-volatile storage device (e.g., in the memory array 120) into the existing data drive buffer.

[0111] As described with reference to updating parity check data (e.g., methods 200a and 200b), the controller 110 is further configured to store the XOR result in the XOR result drive buffer (e.g., drive buffer (XOR result) 205) and write the XOR result to the non-volatile storage device (e.g., write to NAND page (XOR result) 206) after determining the XOR result. The new data and the existing (old) data correspond to the same logical address (the same LBA). The existing data is located at a first physical address of the non-volatile storage device (e.g., NAND page (old data) 203). The controller 110 writing the XOR result to the non-volatile storage device includes: writing the XOR result to a second physical address of the non-volatile storage device (e.g., NAND page (XOR result) 206) and updating the L2P mapping to correspond the logical address to the second physical address. The existing data and the new data are parity bits.

[0112] As described with reference to updating conventional non-parity data (e.g., methods 300a and 300b), the XOR result corresponds to an instantaneous XOR result, and the controller 110 is further configured to transfer the instantaneous XOR result from the instantaneous XOR result drive buffer across the interface 104 to the host 101. The existing data and the new data are data bits. The new data and the existing (old) data correspond to the same logical address (the same LBA). The existing data is located at a first physical address of the non-volatile storage device (e.g., NAND page (old data) 303). The controller 110 writes the new data from the new data drive buffer to the non-volatile storage device by: writing the new data to a second physical address of the non-volatile storage device (e.g., NAND page (new data) 305) and updating the L2P mapping to correspond the logical address to the second physical address.

[0113] As described with reference to performing data recovery (e.g., methods 400a and 400b), the XOR result corresponds to an instantaneous XOR result, and the controller 110 is further configured to transfer the XOR result from the instantaneous XOR result drive buffer across the interface 140 to the host 101 to be transferred by the host 101 to another storage device as previous data. The another storage device is the next storage device after the storage device in a series of storage devices of a RAID group.

[0114] As described in connection with bringing a spare storage device into use (e.g., methods 500a and 500b), the XOR result corresponds to an instantaneous XOR result, and the controller 110 is further configured to transfer the XOR result from the instantaneous XOR result driver buffer across the interface 140 to the host 101 to be transferred by the host 101 to another storage device as recovered data. The another storage device is the spare storage device being brought into use. The spare storage device is the next storage device in a series of storage devices of a RAID group after the storage device.

[0115] Figure 7 is a schematic diagram illustrating a host side view 700 for updating data according to some embodiments. Refer to Figures 1 to 7 , the file system of the host 101 includes logical blocks 701, 702, 703, 704, and 705. The logical blocks 701 to 704 contain regular non-parity data. The logical block 705 contains parity data for the data in the logical blocks 701 to 704. As shown, in response to determining that new data 711 is to be written to the logical block 702 (originally containing old data) in the storage device 100, instead of performing two XOR operations (as conventionally done), the host 101 only needs to update the logical block 705 to the new data 711. As described, the controller 110 performs the XOR operation. Both the new data 711 and the old data 712 are appended to the parity data in the logical block 705. In the next garbage collection (GC) cycle, the new and old parity data is compressed.

[0116] Figure 8 is a schematic diagram illustrating the placement of parity data according to some embodiments. Refer to Figure 1 and 8 , the RAID group 800 (e.g., a RAID 5 group) includes four drives - drive 1, drive 2, drive 3, drive 4. An instance of each drive is the storage device 100. Each of drives 1 to 4 stores data and parity data in its respective memory array 120. Parity A is generated by performing an XOR operation on data A1, A2, and A3 and stored on drive 4. Parity B is generated by performing an XOR operation on data B1, B2, and B3 and stored on drive 3. Parity C is generated by performing an XOR operation on data C1, C2, and C3 and stored on drive 2. Parity D is generated by performing an XOR operation on data D1, D2, and D3 and stored on drive 1.

[0117] Conventionally, if A3 is to be modified (updated) to A3', host 101 may read A1 from drive 1 and A2 from drive 2 and perform an XOR operation on A1, A2, and A3' to generate parity A', and write A3' to drive 3 and parity A' to drive 4. Alternatively and conventionally, to avoid having to reread all the other drives (especially when there are more than four drives in a RAID group), host 101 may also generate parity A' by reading A3 from drive 3, reading parity A from drive 4, and performing an XOR operation on A3, A3', and parity A, and then write A3' to drive 3 and parity A' to drive 4. In both conventional scenarios, modifying A3 would require host 101 to perform at least two reads from drives and two writes to drives.

[0118] The arrangement disclosed herein can eliminate the need for host 101 to read parity A and generate parity A' itself by enabling storage device 100 (e.g., its controller 110) to perform the XOR operation internally. In some instances, storage device 100 may support a new vendor-unique command (VUC) to calculate and store the result of the XOR operation.

[0119] Host 101 may send a VUC containing the LBA of: (1) data or parity data stored on storage device 100; and (2) data to be XORed with the data or parity data corresponding to the LBA to storage device 100. Storage device 100 may read the data or parity data corresponding to the LBA sent by host 101. The read is an internal read and the read data is not sent back to host 101. Storage device 100 performs an XOR operation on the data sent by host 101 and the data or parity data read internally, stores the result of the XOR operation in the LBA corresponding to the data or parity data that has been read, and acknowledges successful command completion to host 101. This command can be executed no more times than would be required for a read and write command of comparable size.

[0120] Host 101 may send a command containing the LBA of parity A and the XOR result of A3 and A3' to drive 4, and in response, drive 4 calculates parity A' and stores parity A' in the same LBA that previously contained parity A.

[0121] In some instances, the XOR operation may be performed within the controller 110 via a command. The command may be implemented via any interface used to communicate with the storage device. In some instances, NVMe interface commands may be used. For example, the host 101 may send a XFER command (with an indicated LBA) to the controller 110 to cause the controller 110 to perform a compute function (CF) on data corresponding to the indicated LBA from the memory array 120. The type of CF to be performed, when the CF will be performed, and on what data the CF will be performed are CF specific. For example, the CF operation may call read data from the memory array 120 to perform an XOR operation with write data transferred from the host 101 (before the data is written to the memory array 120).

[0122] In some instances, the CF operation is not performed on the metadata. The host 101 may specify protection information to be included as part of the CF operation. In other instances, the XFER command causes the controller 110 to operate on the data and metadata according to the CF specified for the logical block indicated in the command. The host 101 may similarly specify protection information to be included as part of the CF operation.

[0123] In some arrangements, the host 101 may call the CF on data sent to the storage device 100 (e.g., for a write operation) or requested from the storage device 100 (e.g., for a read operation). In some instances, the CF is applied before the data is saved to the memory array 120. In some instances, the CF is applied after the data is saved to the memory array 120. In some instances, the storage device 100 sends the data to the host 101 after performing the CF. Examples of CFs include XOR operations as described herein, for example, for RAID5.

[0124] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Those skilled in the art will readily appreciate various modifications to these aspects, and the generic principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, where the singular forms of elements are not intended to mean "one and only one" (unless specifically stated otherwise), but rather "one or more". The term "some", unless specifically stated otherwise, means one or more. All structural and functional equivalents of the elements of the various aspects described throughout this specification of the invention that are known or later come to be known to those skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is expressly recited in the claims. An element of a claim should not be considered a means-plus-function element unless the element is expressly recited using the phrase "means for".

[0125] It should be understood that the specific order or hierarchy of steps in the disclosed processes is an example of an illustrative approach. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged while remaining within the scope of the foregoing description. The appended method claims present elements of the various steps in exemplary order and are not intended to be limited to the specific order or hierarchy presented.

[0126] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the disclosed subject matter. Those skilled in the art will readily appreciate various modifications to these embodiments, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the foregoing description. Thus, the foregoing description is not intended to be limited to the embodiments shown herein, but is to be accorded the broadest scope consistent with the principles and novel features disclosed herein.

[0127] The various examples illustrated and described are provided only as examples to illustrate the various features of the claims. However, the features shown and described with respect to any given example need not be limited to the associated example and can be used or combined with the other examples shown and described. In addition, the claims are not intended to be limited by any example.

[0128] The foregoing method descriptions and process flow diagrams are provided only as illustrative examples and are not intended to require or imply that the steps of the various examples must be performed in the order presented. As will be appreciated by those skilled in the art, the steps in the foregoing examples may be performed in any order. For example, words such as "thereafter," "then," "next," etc. are not intended to limit the order of the steps; these words are merely used to guide the reader through the description of the method. Additionally, any reference to a technical solution element in the singular using, for example, the articles "a," "an," or "the" should not be construed as limiting the element to the singular.

[0129] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, the functionality of the various illustrative components, blocks, modules, circuits, and steps has been generally described above in terms of their functionality. Whether the functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0130] Hardware for implementing the various illustrative logics, logical blocks, modules, and circuits described in connection with the examples disclosed herein may be implemented using a general purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some steps or methods may be performed by circuitry specific to a given function.

[0131] In some illustrative examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium that can be accessed by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, a drive and an optical disk include a compact disk (CD), a laser disk, an optical disk, a digital versatile disk (DVD), a floppy disk drive, and a Blu-ray disk, where the drive typically reproduces data magnetically, and the optical disk reproduces data optically by laser. Combinations of the above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of code and / or instructions in a non-transitory processor-readable storage medium and / or a computer-readable storage medium that may be incorporated into a computer program product.

[0132] The foregoing description of the disclosed examples is provided to enable a person skilled in the art to make or use the present invention. Those skilled in the art will readily appreciate various modifications to these examples, and the generic principles defined herein may be applied to some examples without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the examples shown herein but is to be accorded the widest scope consistent with the appended claims and the principles and novel features disclosed herein.

Claims

1. A storage device, comprising: A non-volatile storage device; And A controller configured to: Receive a write request from a host operatively coupled to the storage device; Determine an XOR result by performing an XOR operation on new data and existing data, the new data being received from the host and the existing data being stored in the non-volatile storage device; And Buffer the new data, the existing data, and the XOR result in respective buffers of the storage device.

2. The apparatus according to claim 1, wherein In response to receiving the write request, the controller transfers the new data from the host to a new data driver buffer of the storage device across an interface; and The controller performs a read operation to read the existing data from the non-volatile storage device into an existing data driver buffer.

3. The apparatus according to claim 2, wherein the controller is further configured to: Store the XOR result in an XOR result driver buffer after determining the XOR result; and Write the XOR result to the non-volatile storage device.

4. The apparatus according to claim 3, wherein The new data and the existing data correspond to the same logical address; The existing data is located at a first physical address of the non-volatile storage device; And Writing the XOR result to the non-volatile storage device includes: Writing the XOR result to a second physical address of the non-volatile storage device; And Updating the logical-to-physical mapping to correspond to the logical address to the second physical address.

5. The apparatus according to claim 3, wherein the existing data and the new data are parity bits.

6. The apparatus according to claim 2, wherein The XOR result corresponds to an instantaneous XOR result; and The controller is further configured to transfer the instantaneous XOR result from an instantaneous XOR result driver buffer to the host across the interface and transfer the new data driver buffer to the non-volatile storage device.

7. The apparatus according to claim 6, wherein the existing data and the new data are data bits.

8. The apparatus according to claim 6, wherein The new data and the existing data correspond to the same logical address; The existing data is located at a first physical address of the non-volatile storage device; and The controller is further configured to write the new data from the new data driver buffer to the non-volatile storage device by: Writing the new data to a second physical address of the non-volatile storage device; and Updating the logical-to-physical mapping to correspond to the logical address to the second physical address.

9. The apparatus according to claim 2, wherein The XOR result corresponds to an instantaneous XOR result; and The controller is further configured to transfer the XOR result from the transient XOR result driver buffer across the interface to the host for being transmitted by the host as previous data to another storage device, where the another storage device is the next storage device in a series of storage devices after the storage device.

10. The apparatus according to claim 2, wherein the XOR result corresponds to a transient XOR result; and the controller is further configured to transfer the XOR result from the transient XOR result driver buffer across the interface to the host for being transmitted by the host as recovered data to another storage device, where the another storage device is a standby storage device put into use and the recovered data is stored by a controller of the another storage device in a non-volatile memory of the another storage device.

11. A method for data processing, comprising: receiving, by a controller of a storage device operatively coupled to a host through an interface, a write request from the host, where the storage device includes a non-volatile storage device; determining, by the controller, an XOR result by performing an XOR operation on new data and existing data, where the new data is received from the host and the existing data is stored in the non-volatile storage device; and buffering the new data, the existing data, and the XOR result in respective buffers of the storage device.

12. The method according to claim 11, further comprising: transferring, in response to receiving the write request, the new data from the host across the interface to a new data driver buffer of the storage device; and performing a read operation to read the existing data from the non-volatile storage device into an existing data driver buffer.

13. A non-transitory computer-readable medium, comprising computer-readable instructions that, when executed, cause a processor to perform the following operations: receiving a write request from a host operatively coupled to a storage device, where the storage device includes a non-volatile storage device; determining an XOR result by performing an XOR operation on new data and existing data, where the new data is received from the host and the existing data is stored in the non-volatile storage device; and buffering the new data, the existing data, and the XOR result in respective buffers of the storage device.

14. The non-transitory computer-readable medium according to claim 13, wherein the processor is further caused to perform the following operations: transferring, in response to receiving the write request, the new data from the host across the interface to a new data driver buffer of the storage device; and performing a read operation to read the existing data from the non-volatile storage device into an existing data driver buffer.

15. The non-transitory computer-readable medium according to claim 14, wherein the processor is further caused to perform the following operations: storing the XOR result in an XOR result driver buffer after determining the XOR result; and Write the XOR result to the non-volatile storage device.

16. The non-transitory computer-readable medium of claim 15, wherein the new data and the existing data correspond to the same logical address; the existing data is located at a first physical address of the non-volatile storage device; and writing the XOR result to the non-volatile storage device includes: writing the XOR result to a second physical address of the non-volatile storage device; and updating the logical-to-physical mapping to correspond to the logical address to the second physical address.

17. The non-transitory computer-readable medium of claim 14, wherein the XOR result corresponds to an instantaneous XOR result; and the processor is further caused to transfer the instantaneous XOR result from the instantaneous XOR result driver buffer to the host across the interface and transfer the new data driver buffer to the non-volatile storage device.

18. The non-transitory computer-readable medium of claim 17, wherein the new data and the existing data correspond to the same logical address; the existing data is located at a first physical address of the non-volatile storage device; and writing the new data from the new data driver buffer to the non-volatile storage device includes: writing the new data to a second physical address of the non-volatile storage device; and updating the logical-to-physical mapping to correspond to the logical address to the second physical address.

19. The non-transitory computer-readable medium of claim 14, wherein the XOR result corresponds to an instantaneous XOR result; and the processor is further caused to transfer the XOR result from the instantaneous XOR result driver buffer to the host across the interface for use as previous data by the host to transfer to another storage device, the another storage device being the next storage device in a series of storage devices after the storage device.

20. The non-transitory computer-readable medium of claim 14, wherein the processor is caused to execute a compute function CF in response to receiving a request from the host, and operations associated with the CF include the XOR operation.

Citation Information

Patent Citations

  • Independent disk redundant array parity calculation in operation

    CN109426583A