Reconfigurable storage device centered on latency and throughput and its operation method

The storage device with a reconfigurable integrated circuit offloads data-intensive operations to reduce host device load, minimizing latency and expanding capacity while maintaining performance.

JP7698425B2Active Publication Date: 2025-06-25SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021010312
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2021-01-26
Publication Date
2025-06-25
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

Existing storage systems experience performance degradation due to over-utilization of host device resources for data-intensive operations, leading to increased latency and reduced throughput.

Method used

A storage device equipped with a reconfigurable integrated circuit (RIC) that includes static and dynamic logic blocks, allowing for dynamic reconfiguration of operations to offload data-intensive tasks closer to the storage destination, reducing the need for host device processing.

Benefits of technology

This approach minimizes latency and expands data capacity without compromising performance by offloading resource-intensive operations to the storage device, thereby optimizing CPU usage and bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698425000001
    Figure 0007698425000001
  • Figure 0007698425000002
    Figure 0007698425000002
  • Figure 0007698425000003
    Figure 0007698425000003
Patent Text Reader

Abstract

To provide a storage device for accelerating a data intensive operation of a host device and an operation method thereof.SOLUTION: A storage device comprises: a storage controller for receiving data from a host device and storing the data in a storage memory; and a reconfigurable integrated circuit which is connected to the storage controller so as to be able to perform communicate, and accelerates a logic operation executed on the data stored in the storage memory. The reconfigurable integrated circuit includes: a first logic block for executing a static logic operation out of logic operations; a second logic block for executing one or more dynamic logic operations out of logic operations; and a plurality of memory buffers for storing input and output of the first and second logic blocks.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a storage device, and more particularly, to a storage device and an operation method thereof for accelerating near-storage of latency-critical and throughput-oriented data-intensive operations.

Background Art

[0002] A storage system generally includes a host device and one or more storage devices. Such storage devices include, for example, magnetic storage devices (e.g., hard disk drives (HDDs), etc.), optical storage devices (e.g., Blu-ray Disc (registered trademark) drives, compact disc (CD) drives, digital versatile disc (DVD) drives, etc.), flash memory devices (e.g., USB flash drives, solid state drives (SSDs), etc.). Generally, in order to process data stored in a storage device, a host device first reads the data from the storage device and transfers the data from the storage device to the main memory of the host device. The host device (e.g., a host device including a host processor such as a central processing unit (CPU)) processes the data transferred from the storage device to the main memory of the host device.

[0003] For example, with respect to a database management system, a host device outputs a response to an input database query by performing various data aggregation-type operations on data stored in a storage device. As an example, the host device first reads data elements from the storage device, and then processes the data elements to identify and output a subset of the data elements from a table corresponding to the input database query, thereby performing various operations (e.g., filtering, sorting, grouping, aggregating, etc.) on the table of data elements stored in the storage device. Such operations are data-intensive because when processed by the host device, they require a large amount of data (e.g., a table of data elements) to be transferred from the storage device to the host device. When the host device processes data aggregation-type operations so as to transfer a large amount of data between the storage device and the host device for processing, it over-utilizes the resources of the host device (e.g., CPU usage, bandwidth, etc.), resulting in latency and degrading the performance of the storage system.

[0004] Therefore, a storage device is needed to accelerate data aggregation-type operations close to the storage destination.

[0005] The above information disclosed in the background art is only for enhancing the understanding of the background of the present invention and includes information that does not constitute the prior art.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] The present invention has been made in view of the above prior art, and an object of the present invention is to provide a storage device for accelerating data-intensive operations of a host device and an operation method thereof.

Means for Solving the Problems

[0008] The present invention made to achieve the above object relates to a storage device for accelerating near storage of latency-critical and throughput-oriented data-intensive operations and a method including the same.

[0009] A storage device according to an aspect of the present invention made to achieve the above object includes a storage controller configured to receive data from a host device and store the data in a storage memory, and a reconfigurable integrated circuit communicably connected to the storage controller and configured to accelerate logical operations executed on the data stored in the storage memory. The reconfigurable integrated circuit includes a first logic block configured to execute static logical operations among the logical operations, a second logic block configured to execute one or more dynamic logical operations among the logical operations, and a plurality of memory buffers configured to store inputs and outputs of the first and second logic blocks.

[0010] The logical operations correspond to a pipeline workflow, the first logic block is statically configured with the static logical operations for the pipeline workflow, and the second logic block can be dynamically reconfigured with the one or more dynamic logical operations for at least one stage of the pipeline workflow. The one or more dynamic logical operations include a first dynamic logical operation and a second dynamic logical operation, the second logic block is configured with the first dynamic logical operation during a first stage of the pipeline workflow, and can be dynamically reconfigured with the second dynamic logical operation during a second stage of the pipeline workflow. The plurality of memory buffers may include an input / output (I / O) buffer configured to store the inputs and outputs of the first and second logic blocks, an intermediate I / O buffer configured to store the intermediate outputs of the second logic block while the second logic block is being reconfigured, and a configuration buffer configured to store a configuration file for reconfiguring the second logic block. The second logic block may be dynamically reconfigured by loading one configuration file from among the configuration files stored in the configuration buffer into the second logic block. The output of the second logic block is stored in the intermediate I / O buffer during a first stage, the second logic block is reconfigured with different dynamic logic instructions for a second stage, and the intermediate I / O buffer may be designated as the input buffer of the second logic block during the second stage. The static logic operation may correspond to a latency-critical operation, and the one or more dynamic logic operations may correspond to a throughput-oriented operation. The latency-critical operation may be an operation having a completion time shorter than the reconfiguration time of the second logic block. The storage device may be a solid state drive. The reconfigurable integrated circuit may be a Field Programmable Gate Array (FPGA).

[0011] A method for accelerating operations in a storage device according to an aspect of the present invention made to achieve the above object is a method for accelerating operations in a storage device including a storage controller, a storage memory, and a reconfigurable integrated circuit including a first logic block, a second logic block, and a buffer, the method comprising: executing, by the first logic block, a first logical operation on input data stored in the storage memory; storing, by the first logic block, an output of the first logical operation in an intermediate output buffer of the buffer; configuring, by the reconfigurable integrated circuit, a second logical operation in the second logic block; designating, by the reconfigurable integrated circuit, the intermediate output buffer as an input buffer for the second logical operation; and executing, by the second logic block, the second logical operation on the output of the first logical operation stored in the intermediate output buffer.

[0012] While the first logical operation is being executed in the first logic block, the second logical operation can be configured in the second logic block. The step of configuring the second logical operation in the second logic block may include monitoring a value of the intermediate output buffer, determining whether the value of the intermediate output buffer exceeds a threshold value, and configuring the second logical operation in the second logic block when it is determined that the value of the intermediate output buffer exceeds the threshold value. The threshold value may be a high water mark of the intermediate output buffer. The buffer may include a configuration buffer configured to store a configuration file for configuring the second logic block. The step of configuring the second logical operation in the second logic block may include loading, from among configuration files stored in the configuration buffer, a bit file corresponding to the second logical operation into the second logic block. The step of designating the intermediate output buffer as the input buffer for the second logical operation may include a step of determining whether the first logical operation has been interrupted, a step of designating the intermediate output buffer as the input buffer for the second logical operation when it is determined that the first logical operation has been interrupted, and a step of designating the input buffer of the first logical operation as the output buffer for the second logical operation. The step of determining whether the first logical operation has been interrupted may include a step of determining whether the end of the intermediate output buffer has been reached. The method may further include a step of determining whether the second logical block has processed all the outputs of the first logical operation stored in the intermediate output buffer, and a step of designating the output buffer of the second logical operation as the final output buffer. The storage device may be a solid state drive, and the reconfigurable integrated circuit may be an FPGA.

Advantages of the Invention

[0013] According to the storage system of the present invention, it is possible to reduce or minimize the latency that occurs when transferring data stored in the storage memory via a long distance and / or an external interface, and it is possible to expand to hold a large amount of data without losing the performance per storage memory capacity unit.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Embodiments for Carrying Out the Invention

[0015] Hereinafter, specific examples of embodiments for carrying out the present invention will be described in detail with reference to the drawings. Here, the same reference numerals refer to the same components. However, the present invention should not be construed as being limited to the embodiments described in this specification and can be implemented in various different forms. Rather, these embodiments are provided as examples so that the present invention becomes thorough and complete, and aspects and features of the present invention are sufficiently conveyed to those skilled in the art. Therefore, unnecessary processes, elements, and technologies for those skilled in the art are not described for a complete understanding of the aspects and features of the present invention. Unless otherwise specified, the same reference numerals refer to the same elements throughout the description given in the drawings and the specification, and thus the description will not be repeated.

[0016] One embodiment of the present invention relates to a storage device for accelerating data-intensive operations of a host device close to the storage destination (e.g., close to or within the storage destination). For example, the host device off-loads data-intensive operations to the storage device so that the storage device processes the data stored in response to the data-intensive operations. In this case, in one embodiment, the storage device processes the stored raw data and outputs a reduced amount of data, and instead of transferring the raw data (e.g., the entire raw data) to be processed by the host device, transfers the reduced amount of data to the host device. Therefore, instead of the host device reading data from the storage device and processing the newly fetched data, most of the operations to be performed on the newly fetched data are off-loaded to the storage device by the host device, and the resources of the host device (e.g., CPU usage, bandwidth, etc.) are used for, for example, cross-device operations (e.g., combining information from tables stored in multiple storage devices). Therefore, the performance of the storage system is improved, for example, by reducing the amount of traffic between the host device and the storage device.

[0017] In one embodiment, as a data-intensive operation, when offloading to a storage device, the scalability of the storage device is improved, for example, by reducing the resources of the host device. Otherwise, the host device is used to process the fetched data stored in the storage device. For example, when the host device processes a data-intensive operation, the host device may become an efficient scalability bottleneck. As an example, a scale-out cluster used in a modern data processing system generally uses a server including one or two low to moderate core CPUs that process data from 4 to 8 SSDs before reaching the upper limit of the interface. In this case, to expand the storage of such a data processing system, instead of expanding the number of SSDs that the core CPU of the existing server in the cluster processes, additional servers are generally added to the scale-out cluster to process additional data processing from additional SSDs. On the other hand, the storage device according to one embodiment accelerates the data-intensive operation of the host device so that each server processes more data. For example, when data is first filtered by the storage device and the filtered minimum table is sent to the host device and combined with information from other tables (stored, for example, in the same or other storage devices within the server), the overall performance of a given decision support benchmark is improved without adding servers to the cluster.

[0018] In one embodiment, the storage device is at least partially and dynamically (e.g., in real time or near real time) reconfigurable (e.g., reprogrammable) to process data stored therein. For example, in one embodiment, the storage device includes a plurality of logic blocks configured to perform data-intensive operations offloaded to the storage device. In one embodiment, the logic blocks include static logic blocks and dynamic logic blocks. The static logic blocks correspond to logic blocks that are statically configured in the storage device at least for the entire pipeline workflow. The dynamic logic blocks correspond to logic blocks that are dynamically reconfigured as needed or desired for one or more stages of the pipeline workflow. As described above, the pipeline workflow refers to a series of operations (e.g., processes) performed on the data of the stage (e.g., simultaneously and / or sequentially) such that the data read from the storage device becomes an input to the first operation of the first stage of the pipeline workflow. The output of the first operation of the first stage becomes an input to the second operation of the second stage of the pipeline workflow, and the output of the final operation of the final stage of the pipeline workflow continues until it becomes the final result of the series of operations.

[0019] In one embodiment, the operations corresponding to a given pipeline workflow include one or more operations that are latency-critical operations and / or one or more operations that are throughput-oriented operations. As described above, a latency-critical operation means an operation that attempts to optimize or shorten the time taken from the start of a data read operation to the end of the operation performed on the read data, whereas a throughput-oriented operation means an operation that optimizes or increases a speed parameter, such as, for example, the number of operations performed per unit time or the amount of data processed per unit time, but not necessarily the latency of any operation. In this case, the latency-critical operations (which do not tolerate the time taken to reconfigure the storage device) correspond to the static logic blocks, and the throughput-oriented operations (which tolerate the time taken to reconfigure the storage device) correspond to the dynamic logic blocks.

[0020] For example, in one embodiment, reconfiguring a dynamic logic block requires a reconfiguration time (e.g., about 1 millisecond (ms)), but user requirements (e.g., a service level agreement (SLA)) require certain latency-critical operations to be performed in a time shorter than the reconfiguration time (e.g., about 25 microseconds (μs)). In this case, the latency-critical operations do not tolerate the time required to reconfigure the dynamic logic block (e.g., if the operation has a completion time shorter than the reconfiguration time), and thus, the latency-critical operations are composed of static logic blocks. On the other hand, when the storage device includes only static logic blocks, the operations offloaded to the storage device are limited according to the fixed resources of the storage device. For example, in this case, data-intensive operations are configured simultaneously (e.g., all at once or concurrently) in the storage device, and thus, the amount of data to be processed and / or the type of operations configured simultaneously in the storage device are limited by the fixed resources of the storage device.

[0021] In one embodiment, the storage device is configured (e.g., reconfigured or reprogrammed) at start time and / or runtime (e.g., real-time or near real-time) as needed or desired according to various user requirements (e.g., SLAs, etc.), available resources of the storage device (e.g., available memory, number of available look-up tables (LUTs), etc.), pipeline workflows, acceleration performance, data size, selectivity of data reduction operations, etc. For example, in one embodiment, latency-critical and / or throughput-centric operations are composed of static and dynamic logic blocks as needed or desired considering SLAs (e.g., operations considered latency-critical), reconfiguration time of logical blocks, available resources of the storage device (e.g., reconfigurable integrated circuits), pipeline workflows, etc. In other examples, the storage device operates in various modes according to acceleration performance and / or selectivity of data reduction operations. For example, in one embodiment, if the size of the data returned to the host device for data-intensive operations offloaded to the storage device is not reduced, the data-intensive operations offloaded to the storage device are instead performed by the host device, and thus the storage device is dynamically reconfigured to operate in a normal mode (e.g., a mode in which data is read and processed by the host device instead of being offloaded to the storage device).

[0022] These and other aspects and features of the present invention will be described in more detail below with reference to the drawings.

[0023] FIG. 1 is a system diagram of a storage system according to an embodiment of the present invention.

[0024] A storage system 100 according to an embodiment of the present invention includes a host device 102 (e.g., a host computer) and a storage device 104. The host device 102 offloads various data-intensive operations to the storage device 104 so that the storage device 104 can accelerate the data-intensive operations of the host device 102. For example, the host device 102 is communicably connected to the storage device 104 and transfers data to the storage device 104 for storing the data in the storage device 104. The host device 102 transmits various commands to the storage device 104, and the storage device 104 processes the data stored therein according to the commands instead of transmitting all the data to be processed by the host processor 106 (e.g., a CPU) to the host memory 108 (e.g., a main memory). For example, instead of transmitting most of the large amount of raw data stored in the storage device 104 to the host memory 108 so that most of the raw data is filtered by the host processor 106, the storage device 104 processes the raw data stored therein in response to a command and outputs a reduced amount of processed data (e.g., a subset of the raw data) to the host device 102. Therefore, most of the operations performed on the newly fetched data are offloaded to the storage device 104, and the resources of the host device 102 (e.g., CPU usage, I / O bus bandwidth, CPU cache capacity, cache for memory bandwidth, memory capacity, etc.) that are performed close to the storage destination (e.g., close to or within the storage destination) are used for other operations such as operations in memory and cross-device operations (e.g., combining data stored in multiple storage devices).

[0025] More specifically, as shown in FIG. 1, the host device 102 includes a host processor 106 and a host memory 108. The host processor 106 is a general-purpose processor such as, for example, a CPU core of the host device 102. The host memory 108 is regarded as the high-performance main memory (e.g., primary memory) of the host device 102. For example, in one embodiment, the host memory 108 includes (or is) a volatile memory such as Dynamic Random Access Memory (DRAM). However, the present invention is not limited thereto, and the host memory 108 includes (or is instead of) any suitable high-performance main memory (e.g., primary memory) for the host device 102, as is known to those skilled in the art. For example, in other embodiments, the host memory 108 is a relatively high-performance non-volatile memory such as NAND flash memory, PCM (Phase Change Memory), resistive RAM, STTRAM (Spin-Transfer Torque RAM), any suitable memory based on PCM technology, memristor technology, and / or Resistive Random Access Memory (ReRAM), and includes, for example, chalcogenides and the like.

[0026] Storage device 104 is regarded as a secondary memory that persistently stores data accessible by host device 102. In this regard, storage device 104 includes (or is) a relatively slow memory compared to the high-performance memory of host memory 108. For example, in one embodiment, storage device 104 is a secondary memory of host device 102 such as an SSD. However, the present invention is not limited thereto, and in other embodiments, storage device 104 includes (or is) any suitable storage device such as an HDD, a USB flash drive, a Blu-ray Disc (registered trademark) drive, etc. In one embodiment, storage device 104 conforms to a large form factor standard (e.g., 3.5-inch hard drive form factor), a small form factor standard (e.g., 2.5-inch hard drive form factor), an M.2 form factor, etc. In other embodiments, storage device 104 conforms to any suitable or desired derivative of these form factors.

[0027] In one embodiment, the storage device 104 includes a storage interface 110, a storage controller 112, a storage memory 114, a reprogrammable integrated circuit (RIC) device 116, a direct (or individual) wiring 118 between the storage controller 112 and the RIC device 116, and a RIC extended memory 120. The storage interface 110 facilitates communication between the host device 102 and the storage device 104 (e.g., using a connector and a protocol). For example, in one embodiment, the storage interface 110 is used for data communication with the host device 102, the storage controller 112, and / or the RIC device 116. In one embodiment, the storage interface 110 facilitates the exchange of storage requests and responses between the host device 102 and the storage device 104. In one embodiment, the storage interface 110 facilitates data transfer from the host memory 108 of the host device 102 to the storage device 104 and data transfer to the host memory 108. For example, in one embodiment, the storage interface 110 (e.g., a connector and its protocol) includes (or is analogous to) Peripheral Component Interconnect Express (PCIe), Remote Direct Memory Access (RDMA) via Ethernet (registered trademark), Serial Advanced Technology Attachment (SATA), Fibre Channel, Serial Attached SCSI (SAS), Non-Volatile Memory Express (NVMe), etc. In other embodiments, the storage interface 110 (e.g., a connector and its protocol) includes (or is analogous to) various general-purpose interfaces such as Ethernet (registered trademark), Universal Serial Bus (USB), etc.In another embodiment, the storage interface 110 supports additional acceleration or coherence protocols such as CCIX, CAPI, OpenCAPI, nvLink, or CXL over the connector itself and related protocols (such as PCIe or Ethernet).

[0028] Storage controller 112 is connected to storage interface 110 and responds to input / output (I / O) requests received from host device 102 via storage interface 110. Storage controller 112 controls storage memory 114 and provides an interface for access from and to storage memory 114. For example, storage controller 112 includes at least one processing circuit built therein to interface with host device 102 and storage memory 114. The processing circuit includes a digital circuit (e.g., microcontroller, microprocessor, digital signal processor, FPGA, application-specific integrated circuit (ASIC), etc.) that executes data access instructions to provide access from and to data stored in storage memory 114, for example, according to data access instructions. For example, data access instructions include any suitable data storage and retrieval algorithm (e.g., read / write) instructions, encryption / decryption algorithm instructions, compression algorithm instructions, etc. Storage memory 114 persistently stores data received from host device 102. For example, with respect to a database management system, storage memory 114 stores data in any suitable self-describing columnar format such as AVRO, ORC, PARQUET, etc. However, the present invention is not limited thereto, and storage memory 114 stores data in any suitable format according to the application of storage system 100. For example, with respect to a media system, storage memory 114 stores data in any suitable media format such as H.264, H.265, MPEG, AVI, etc. In one embodiment, storage memory 114 stores data received from host device 102 in an encrypted and / or compressed format. Storage memory 114 includes a non-volatile memory such as, for example, NAND flash memory.However, the present invention is not limited thereto, and the storage memory 114 includes any suitable memory according to the type of the storage device 104 such as a phase change memory, a magnetic memory, a ferroelectric memory, etc.

[0029] The RIC device 116 processes the data stored in the storage memory 114 by commands from the host device 102. For example, in one embodiment, the RIC device 116 is communicatively connected to the storage controller 112 (e.g., via the direct wiring 118) to access (e.g., read) the data stored in the storage memory 114, and processes the read data to transmit a reduced amount of the processed data (e.g., a subset of the search data stored in the storage memory 114) to the host device 102 (e.g., perform reduction, filtering, sorting, grouping, aggregation, duplicate removal, etc.). In this case, the RIC device 116 has a plurality of logic blocks having various suitable configurations for processing the data stored in the storage memory 114 by commands from the host device 102. As described herein, the logic blocks include the logical components of the RIC device 116 and have connections configured (e.g., as defined in a LUT in some FPGAs) to perform various logical operations (e.g., filtering, sorting, aggregation, duplicate removal, etc.). Since the RIC device 116 includes logic blocks for performing various operations on the newly fetched data instead of the host device 102, the resource utilization of the host device 102 (e.g., CPU usage, PCI bandwidth, etc.) is reduced.

[0030] Accordingly, the RIC device 116 is regarded as a different processor separate from the processor of the host device 102 (e.g., from the host processor 106). For example, in one embodiment, the RIC device 116 is implemented as an integrated circuit (IC). In one embodiment, the RIC device 116 is implemented on the storage device 104 (e.g., built into the same substrate or the same circuit board as the storage device 104). For example, the RIC device 116 is implemented (e.g., attached or mounted) on the storage device 104 as a System On Chip (SOC). In this case, since the RIC device 116 is implemented on the storage device 104, the data stored in the storage device 104 is processed closer to the storage memory 114. Accordingly, the latency that occurs when transferring the data stored in the storage memory 114 via a long distance and / or an external interface can be reduced or minimized. The storage system 100 further benefits from the additional internal data transfer bandwidth between the RIC device 116 introduced by each additional storage device 104 and the storage controller 112. Accordingly, the pure data transfer throughput of the storage system 100 is no longer limited by the host with respect to the storage interface. However, the present invention is not limited thereto, and in other embodiments, the RIC device 116 is implemented on a substrate different from the storage device 104 (e.g., a different circuit board) and communicably connected to the storage device 104. In one embodiment, the RIC device 116 includes (or is) an FPGA configured to support dynamic partial reconfiguration, at least a part of which is dynamically reconfigured as needed or desired, but the present invention is not limited thereto. For example, in other embodiments, the RIC device 116 includes (or is) an ASIC, a Graphical Processing Unit (GPU), a Complex Programmable Logic Device (CPLD), a Coarse-Grained Reconfigurable Array (CGRA), etc.

[0031] In one embodiment, the RIC device 116 is regarded as an auxiliary processor of a different storage device 104 separate from the storage controller 112. For example, in one embodiment, unlike the storage controller 112 which is not easily reprogrammable, the RIC device 116 supports a DPR (e.g., dynamically reprogrammable) that can be at least partially and dynamically reconfigured (e.g., dynamically reprogrammed) as needed or desired by commands from the host device 102. However, the present invention is not limited thereto, and in other embodiments, the RIC device 116 is implemented as part of the storage controller 112 when, for example, all or part of the storage controller 112 is reprogrammable (e.g., configured to support DPR). As will be described in more detail with reference to FIG. 2, in one embodiment, the RIC device 116 includes static logic blocks and dynamically reconfigurable logic blocks (e.g., dynamic logic blocks) to perform various operations on the data stored in the storage memory 114 by commands from the host device 102.

[0032] Referring to FIG. 1, in one embodiment, the RIC device 116 is communicatively connected to the storage controller 112 via a direct (or individual) wiring 118. For example, in one embodiment, the RIC device 116 reads data stored in the storage memory 114 by directly communicating with the storage controller 112 via the direct wiring 118 using peer-to-peer (P2P) communication without having the host device 102. For example, instead of first loading data from the storage memory 114 to the host memory 108 and then transferring the data to the RIC device 116 for additional processing, the RIC device 116 directly communicates with the storage controller 112 to access or receive data from the storage memory 114 without having the host device 102. The P2P communication via the direct wiring 118 between the RIC device 116 and the storage controller 112 can further reduce or eliminate the read / write overhead from the host memory 108 and reduce the operation latency that occurs when communicating data via the host device 102. The data transfer bandwidth of each direct wiring 118 is added to the total data transfer throughput of the storage system 100 in proportion to the amount of data in the storage memory 114, even when an additional storage device 104 is arranged in the storage system 100. This scalability advantage of the present invention is that the storage system 100 can be expanded to hold much more data without losing performance per storage memory capacity unit compared to a comparative system. After the data is processed by the RIC device 116, the processed data is provided to the host device 102. The processed data becomes smaller through the filtering operation performed by the RIC device 116, or the host processing becomes easier through the reformatting operation performed by the RIC device 116, or is more suitable for viewing by the clients of the storage system 100 through the transcoding operation performed by the RIC device 116. For example, by integrating the processing capabilities of the RIC device 116 into the storage device 104, additional performance and utility advantages are generated for the consumers of the functions realized by the storage system 100.However, the present invention is not limited to this.

[0033] The RIC extended memory 120 is communicably connected to the RIC device 116 and is implemented as a memory chip (e.g., a DRAM chip) connected to a channel (e.g., a DDR (Double Data Rate) memory interface) of the RIC device 116 on the storage device 104. For example, in one embodiment, the RIC extended memory 120 is incorporated into the storage device 104 as a plurality of memory devices (e.g., a plurality of DRAM memory chips) connected to the DDR port of the RIC device 116. As used herein, a "memory device" means the smallest functionally replaceable unit of memory that stores data. For example, a DRAM memory device has 36 billion bits of data, and each bit is implemented by a capacitor for storing charge and a transistor for selectively filling the capacitor with 1 bit of data. However, the present invention is not limited to this, and the RIC extended memory 120 can have any suitable type of memory for expanding the main memory (e.g., internal memory) of the RIC device 116. For example, in other embodiments, the RIC extended memory 120 includes (or is) any suitable volatile memory or non-volatile memory known to those skilled in the art, such as SRAM, MRAM, NAND, tightly-coupled memory (TCM), PCM, resistive RAM, STTRAM, any suitable memory based on PCM technology, memristor technology, and / or ReRAM, and includes, for example, chalcogenides.

[0034] In one embodiment, the RIC extended memory 120 is a relatively slow memory compared to the main memory (e.g., internal memory) of the RIC device 116 (e.g., see FIG. 2), but in an embodiment where the RIC device 116 is a Xilinx UltraScale+ FPGA, it has a larger capacity (e.g., larger storage space) than the main on-chip memory of the RIC device 116 such as any block RAM or integrated RAM. In this case, as will be described in more detail with reference to FIG. 3, the off-chip RIC extended memory 120 is used not only as a staging memory into which the RIC extended memory 120 is divided to store intermediate input / outputs, but also to store configuration files for dynamically reconfiguring dynamic logic blocks as needed or desired. For example, in one embodiment, the RIC device 116 includes gates and flip-flops (e.g., logic blocks), and the functions and / or connections between the gates and / or flip-flops (e.g., LUTs in the case of an FPGA) are configured by loading configuration data (e.g., object file or bit file in the case of an FPGA) into the RIC device 116, which is a configuration file stored in the RIC extended memory 120 and quickly retrieved as needed or desired. However, the present invention is not limited thereto, and in other embodiments, the RIC extended memory 120 can be omitted, for example, when the main memory of the RIC device 116 has sufficient capacity to perform the functions of the RIC extended memory 120 described herein (e.g., has sufficient capacity to be divided for an intermediate staging memory).

[0035] FIG. 2 is a block diagram showing in more detail the RIC device of the storage device of FIG. 1 according to an embodiment of the present invention. FIG. 3 is a block diagram showing in more detail the extended memory 120 of the RIC device of the storage device of FIG. 1 according to an embodiment of the present invention. Hereinafter, for convenience, the RIC device 116 will be described more specifically with respect to an FPGA, but the present invention is not limited thereto.

[0036] Referring to FIGS. 1-3, the RIC device 116 processes data stored in the storage memory 114 by commands of the host device 102. For example, in one embodiment, the RIC device 116 includes a RIC accelerator 202 and a RIC memory 204 (e.g., main memory or internal memory). The RIC device 116 receives data from the storage memory 114 via the direct wiring 118 and processes the read data according to the configuration of the RIC accelerator 202. The input / output data processed by the RIC accelerator 202 is stored in the RIC memory 204 (and / or the RIC extended memory 120). When the data is completely processed (e.g., by the RIC accelerator 202), the processed data is transferred to the host device 102. In one embodiment, the RIC accelerator 202 is at least partially and dynamically reconfigured (e.g., in real time or substantially in real time) according to the available resources of the RIC device 116, user requirements (e.g., SLA), pipeline workflow, the size of the data transferred between stages, acceleration performance, selectivity of data reduction operations, etc.

[0037] For example, the RIC accelerator 202 includes a static logic block 206 and a dynamic logic block 208. The static logic block 206 corresponds to the logic blocks configured in the RIC accelerator 202 for at least the entire pipeline workflow, and the dynamic logic block 208 corresponds to the logic blocks that are dynamically reconfigured as needed or desired for one or more stages corresponding to the pipeline workflow. For example, the pipeline workflow is divided into a plurality of stages, and each stage performs (e.g., simultaneously for maximum throughput or minimum latency or sequentially for maximum throughput per RIC accelerator resource) one or more operations on data (e.g., data read from the storage memory 114 or data output from the previous stage). For each stage of the pipeline workflow, the RIC accelerator 202 maintains the internally configured static logic block 206, but for any particular one or more stages, the RIC accelerator 202 dynamically reconfigures the dynamic logic block 208 as needed or desired. For example, as will be described in more detail with reference to FIGS. 4-5B, the static logic block 206 and the dynamic logic block 208 are configured in the RIC accelerator 202 depending on (e.g., in response to) the critical latency requirements of the operations and / or the amount of available resources on the RIC accelerator 202 that are configured simultaneously (e.g., all at once or concurrently) to process the operations.

[0038] The RIC memory 204 is regarded as the main memory (e.g., internal memory) of the RIC device 116. The RIC memory 204 has an I / O buffer 210, a first memory 212, and a second memory 214. The I / O buffer 210 is divided between the first and second memories (212, 214) and serves as a buffer for the inputs and outputs of logical blocks (e.g., static logical blocks and / or dynamic logical blocks) executed by the RIC accelerator 202. The RIC extended memory 120 is the extended memory (e.g., external memory or secondary memory) of the RIC device 116. In one embodiment, the RIC extended memory 120 serves as the staging memory of the RIC device 116. As shown in FIG. 3, in one embodiment, the RIC extended memory 120 has a read / write buffer 302, an intermediate I / O buffer 304, a config buffer 306, and a third memory 308. The read / write buffer 302, the intermediate I / O buffer 304, and the config buffer 306 are divided over the third memory 308.

[0039] The read / write buffer 302 stores data read from and written to the storage memory 114. The intermediate I / O buffer 304 serves as an intermediate buffer for the inputs and outputs of logical blocks between the stages of the pipeline workflow. For example, when a dynamic logical block is reconfigured between the stages of the pipeline workflow, the output of the dynamic logical block of the previous stage is stored in the intermediate I / O buffer 304, the dynamic logical block is reconfigured for the current stage, and then the intermediate I / O buffer 304 is designated as the input buffer for the dynamically reconfigured logical block for the current stage.

[0040] The configuration buffer 306 stores configuration files (e.g., object files or bit files in the case of FPGAs) with various and different configurations for the dynamic logic blocks. In this case, the dynamic logic blocks are reconfigured by loading different configuration files from the configuration buffer 306 to the RIC accelerator 202 (e.g., corresponding to the desired operation) as needed or desired. The reconfiguration time of the dynamic logic blocks when storing the configuration files in the configuration buffer 306 of the RIC extended memory 120 shortens the reconfiguration time of the dynamic logic blocks (e.g., about 1 ms) compared to other cases where the configuration files are stored externally and / or provided from other devices (e.g., host devices). However, the present invention is not limited thereto, and in other embodiments, the configuration buffer 306 is omitted. In this case, the configuration files are stored, for example, in the storage memory 114 or the RIC memory 204 or provided from an external device (e.g., a host device, etc.).

[0041] In one embodiment, the first memory 212 is the earliest available memory of the RIC device 116 but has a low capacity (e.g., low storage space). The second memory 214 has a higher capacity than the first memory 212 but may be slower than the first memory 212. The third memory 308 has the largest capacity (e.g., the largest storage space) but is the slowest available memory of the RIC device 116. For example, with respect to an FPGA, the first memory 212 has block random access memory (BRAM), the second memory 214 has unified random access memory (URAM), and the third memory 308 has DRAM. However, the present invention is not limited thereto, and in other embodiments, one of the first and second memories (212, 214) is omitted, or the first and second memories (212, 214) include any suitable type of memory according to the type of the RIC device 116. For example, in other embodiments, with respect to an FPGA, the second memory 214 (e.g., URAM) is omitted. In one embodiment, the third memory 308 has a 4GB DRAM chip or an 8GB DRAM chip (e.g., is a 4GB DRAM chip or an 8GB DRAM chip), but the present invention is not limited to this.

[0042] In one embodiment, the RIC accelerator 202 stores the inputs / outputs of the static logic block 206 and the dynamic logic block 208 in the first, second, and / or third memories (212, 214, 308) according to the size of the data transferred between stages and / or the desired data rate. For example, if the amount of data transferred between stages is relatively small and / or the operations performed by the logic blocks of the RIC accelerator 202 (e.g., the static logic block 206) are latency critical, the inputs / outputs of such logic blocks are stored in the first memory 212 or the second memory 214. On the other hand, if the data transferred between stages is relatively large and / or the operations performed by the logic blocks (e.g., the dynamic logic block 208) are throughput oriented, the inputs / outputs of such logic blocks are stored in the third memory 308.

[0043] In one embodiment, when the data transferred between stages is relatively large, the output of the operations performed by a logic block (e.g., dynamic logic block 208) is initially stored in the first or second memory (212, 214), and when reconfiguring the dynamic logic block 208 (e.g., between stages), the output is transferred to a third memory 308 (e.g., intermediate I / O buffer 304) to reconfigure the dynamic logic block 208. The output stored in the third memory 308 is designated as an input buffer for the reconfigured dynamic logic block 208. In this case, the output of the reconfigured dynamic logic block 208 is stored in any one of the appropriate ones among the first, second, and third memories (212, 214, 308) (e.g., according to speed, data size, etc.). In other embodiments, when the data transferred between stages is relatively large, the output of the operations performed by a logic block (e.g., dynamic logic block 208) is initially stored in the third memory 308 (e.g., intermediate I / O buffer 304), and when reconfiguring the dynamic logic block 208, the output stored in the third memory 308 from the previous stage is designated as the input to the reconfigured dynamic logic block 208 of the current stage. However, the present invention is not limited to these examples, and any appropriate combination of the static logic block 206 and the dynamic logic block 208 consumes any appropriate ones of the resources of the first, second, and third memories (212, 214, 308) as required or desired according to the amount of data transferred between stages, the desired data speed, etc.

[0044] FIG. 4 is a diagram showing an example of a pipeline workflow according to an embodiment of the present invention. FIG. 5(a) is a comparative example of statically configuring a storage device using operations related to the pipeline workflow of FIG. 4, and FIG. 5(b) is an example of configuring a storage device using operations related to the pipeline workflow of FIG. 4. For convenience, the pipeline workflow will be described with respect to an exemplary database query in a database application, but the present invention is not limited thereto.

[0045] As shown in FIGS. 1 to 5, a typical pipeline workflow 400 that responds to a database query has a plurality of stages (402 to 412). For example, the stages include a first stage 402, a second stage 404, a third stage 406, a fourth stage 408, a fifth stage 410, and a sixth stage 412. Each stage (402 to 412) is stored in the storage memory 114 and starts with the operation of the first stage 402 that processes the data received from the storage controller 112 as the first step of processing the database query. Then, as the second step of processing the database query, one or more operations (e.g., simultaneously or sequentially) are performed, such as the operator of the second stage 404 processing the output of the first stage 402. According to one embodiment of the present invention, the operations related to any combination of the stages (402 to 412) are offloaded to the storage device 104 instead of being performed by the host device 102. Thus, in this case, as indicated by arrows having apertures of different widths from each other between each stage (402 to 412), data of different sizes is transferred between different components of the storage device 104 in order to perform the operations related to the stages. For example, the operations related to the first stage 402 are performed by the storage controller 112, and the operations related to the second to sixth stages (404 to 412) are performed by the RIC device 116, but the present invention is not limited thereto. For example, in other embodiments, all the operations related to the first stage to the sixth stage (402 to 412) are performed by the RIC device 116, or a part of the operations related to the second to sixth stages (404 to 412) (e.g., a part of the latency-critical operations) are performed by the storage controller 112.

[0046] As shown in FIG. 4, for a database application, data tables are generally stored in a storage device 104 (e.g., storage memory 114) in a compressed and encrypted format. Thus, one or more operations related to the first stage 402 include an operation of decrypting data. According to an embodiment of the present invention, one or more operations related to the first stage 402 are performed, for example, by a storage controller 112. As a result, the compressed data is decrypted by the storage controller 112, and the decrypted compressed data is transmitted to the RIC device 116 (e.g., via a direct wiring 118) for additional processing. For example, as shown in FIG. 4, the decrypted compressed data is transmitted from the storage controller 112 to the RIC device 116 at a speed of about 3.2 GB / s to about 6.4 GB / s, but the present invention is not limited thereto.

[0047] The decrypted data is parsed, for example, to identify a desired compressed column within a data table, and the desired compressed column of the data table is decompressed during the second stage 404. As an example, a database query corresponds to an operation of "identifying all male smokers residing in zip code 95134 sorted by age group" in one or more data tables stored in the storage device 104, where a column of the data table corresponds to a zip code, one column corresponds to gender, one column corresponds to age, and one column corresponds to smoker / non-smoker, etc. Thus, one or more operations related to the second stage 404 include an operation of analyzing the stored data format (e.g., a self-describing column format with respect to a database application), and an operation of decompressing the data analyzed using the reverse of the algorithm used to compress the originally stored data (e.g., decompressing the compressed columns corresponding to zip code, gender, age, smoker / non-smoker, etc.). Thus, for example, if the stored data is compressed using the gzip algorithm, when reading the compressed data, the second stage 404 includes an operation of decompressing the data analyzed using the corresponding gunzip decompression algorithm. As a result, as indicated by an increase in the width of the arrow between the second stage 404 and the third stage 406 in FIG. 4, since the compression ratio generally has a factor of 2 to 2.5, the size of the data increases from about 3.2 to 6.4 GB / s of the compressed data to about 6.4 to 16 GB / s of the uncompressed data.

[0048] The uncompressed data is filtered according to one or more conditions defined in the database query. Therefore, one or more operations related to the third stage 406 include operations of filtering the decompressed data according to the conditions defined in the database query. For example, as conditions corresponding to an exemplary database query, the postal code is 95134, the gender is male, and smokers rather than non-smokers are mentioned. In this case, for example, the RIC device 116 selects all rows corresponding to 95134 from the column corresponding to the postal code, and fetches the remaining data of other columns (for example, gender, smoker, age, etc.) of these matching rows. Next, after the RIC device 116 selects all rows corresponding to male from the column corresponding to the gender of the matching rows, it fetches the remaining data of other columns (for example, smoker / non-smoker, age, etc.) of the rows that match the postal code 95134 and male. Similarly, the RIC device 116 selects all rows corresponding to smokers rather than non-smokers from the column corresponding to smoker / non-smoker of the matching rows until all filtering conditions are applied.

[0049] As a result, as shown by the decrease in the width of the arrow between the third stage 406 and the fourth stage 408 in FIG. 4, the size of the data decreases from the size of the uncompressed data to the size of the filtered data. The size of the filtered data varies according to the selectivity of the conditions used to filter the data. For example, if the condition only matches some items of the uncompressed data out of millions of items of the uncompressed data, the size of the filtered resulting data is much smaller than the size of the uncompressed data. As a result, the traffic to the host device 102 is substantially reduced. On the other hand, if the condition is not very selective and most of the uncompressed data remains (for example, not filtered), the size of the filtered data is substantially the same as the size of the uncompressed data. In this case, since the amount of the filtered data is the same as or substantially the same as the uncompressed data, the acceleration by the storage device 104 is not very beneficial.

[0050] Thus, in one embodiment, depending on the selectivity of the data reduction operation (e.g., the filtering operation in the example of FIG. 4), if the conditions used to reduce the data (e.g., filtering of the data) are not very selective, control returns to the host device to perform the remaining operations. For example, in one embodiment, the selectivity of the data reduction operation (e.g., filtering operation) for a given pipeline workflow is not known in advance (e.g., during the planning stage). In this case, during execution time (e.g., during runtime), the selectivity of the data reduction operation being performed by one or more logic blocks is monitored (e.g., by the host device 102 or another device or system communicatively coupled to the host device 102, such as a runtime service). If the reduction in data size is less than the threshold reduction, control returns to the host device 102 (e.g., along with the reduced data) to control the host device 102 to perform the remaining operations on the data.

[0051] After the data is filtered, the filtered data for an exemplary database query is sorted at a fourth stage 408, grouped at a fifth stage 410, and aggregated at a sixth stage 412. For example, the filtered data is sorted by age at the fourth stage 408, grouped into different age groups at the fifth stage 410, and the grouped data is aggregated at the sixth stage 412. Thus, one or more operations for the fourth stage 408 include an operation of sorting the filtered data, one or more operations for the fifth stage 410 include an operation of grouping the sorted data, and one or more operations for the sixth stage 412 aggregate the grouped data. As indicated by the constant width of the arrows between the fourth through sixth stages (408 - 412), the operations of sorting and grouping do not affect the data size, but aggregation sporadically reduces the size. Next, the aggregated data is transmitted to the host device 102 as indicated by the last arrow. Thus, the reduction in the size of the data processed by the RIC device 116 varies according to the selectivity of the filter condition during the filtering stage 406, the number of separate groups formed during the grouping stage 410, and / or the degree to which the operations of the aggregation stage 412 summarize the data for each group.

[0052] As shown in FIG. 5(a), the resources required to statically configure the RIC accelerator 202 for all operations related to the pipeline workflow 400 of FIG. 4 exceed the amount of resources available in the RIC device 116. For example, with respect to an FPGA, the connection of gates and flip - flops of logic blocks (e.g., static logic block 206 and dynamic logic block 208) that configure the operation of the logic blocks is defined by a LUT, such as a truth table. However, the number of LUTs configured in the FPGA at a given time is limited to the maximum total number of LUTs of the FPGA. For example, a small FPGA is limited to 300K, which is the maximum total number of LUTs. In this case, the total number of LUTs used for the operation implementation of each stage (404 - 412) exceeds the maximum total number of LUTs of the FPGA.

[0053] For example, as shown in FIG. 5(a), the number of LUTs used to analyze the stored data format (e.g., in the second stage 404) is about 25K, and the number of LUTs used to decompress the analyzed data (in the second stage 404) is about 12K. Assuming a compression factor of 2, the number of LUTs used to filter the decompressed data flowing at twice the speed of the stored data (e.g., in the third stage 406) is much larger (e.g., 90K). (Assuming, for example, that 90% of the data is filtered for the purpose of explanation, in the fourth stage 408), the number of LUTs used to sort the reduced data is about 100K. (For example, in the fifth stage 410 and the sixth stage 412), the number of LUTs used to group and aggregate the data is about 100K (assuming, for example, for the purpose of explanation, that the data is composed of 10 groups). In this comparative example, the total number of LUTs used to process the data by the pipeline workflow 400 of FIG. 4 is 327K, which exceeds the maximum total number of LUTs of the FPGA (e.g., 300K in this example). Therefore, all operations related to the pipeline workflow 400 may not simultaneously (e.g., all at once or simultaneously) fit into the resources of the FPGA, and thus may not be statically configured in the FPGA at one time. In this case, the number of operations related to the pipeline workflow 400 offloaded to the FPGA decreases or is limited according to the available resources of the FPGA.

[0054] On the one hand, as shown in FIG. 5(b), when at least a part of the operations of the pipeline workflow 400 are dynamically configured according to necessity or desire, the operations related to the pipeline workflow 400 are offloaded to the FPGA. For example, the operations 502 related to the analysis, decompression, and filtering stages (e.g., the second stage 404 and the third stage 406) are statically configured, and when the other remaining operations (504, 506) related to the sorting, grouping, and aggregation stages (e.g., the fourth stage 408, the fifth stage 410, and the sixth stage 412) are dynamically configured according to necessity or desire, the maximum number of LUTs used at any time is 227K (e.g., 127K for the statically configured logic blocks and 100K for the dynamically configured logic blocks). Therefore, when at least a part of the operations are dynamically configured according to necessity or desire, the number of operations offloaded to the FPGA increases.

[0055] In one embodiment, since the configuration file is stored in the configuration buffer 306 (see, e.g., FIG. 3) for quick search according to necessity or desire, the reconfiguration time of the dynamic logic block is reduced (e.g., to about 1 ms). In this case, when another operation is performed by one of the dynamic logic blocks 208, the corresponding configuration file is loaded from the configuration buffer 306 to reconfigure the dynamic logic block within 1 ms. However, in this case as well, there are latency-critical operations that do not allow the time taken to reconfigure the dynamic logic block. Therefore, in one embodiment, the operations of the pipeline workflow corresponding to the latency-critical operations are configured in the static logic blocks so as not to add the reconfiguration time to such operations. The dynamic logic blocks are configured with other operations of the pipeline workflow (e.g., throughput-oriented operations) that allow the time taken to reconfigure the dynamic logic block, improving the resource utilization of the RIC device 116. However, the present invention is not limited thereto. For example, referring to FIGS. 6A and 6B, in one embodiment, the dynamic logic block is reconfigured while other logical operations are being executed so that the reconfiguration time of the dynamic logic block is hidden.

[0056] FIG. 6A and FIG. 6B are flowcharts showing a method 600 for accelerating data-intensive operations by a storage device according to an embodiment of the present invention. However, the present invention is not limited to the order or number of operations of method 600 shown in FIGS. 6A and 6B, and can be changed to any desired order or number of operations as recognized by those skilled in the art. For example, in one embodiment, the order can be changed and the method can include fewer or more operations.

[0057] Referring to FIGS. 6A and 6B, the method begins when the storage device 104 receives one or more commands from the host device 102 for processing data stored in the storage memory 114 (see, e.g., FIG. 1). Since the commands are related to a specific pipeline workflow, the pipeline workflow is divided into multiple stages, and each stage corresponds to one or more data-intensive operations related to the commands. For each stage, one or more logical operations are dynamically configured in the dynamic logic block 208 of the RIC device 116 (see, e.g., FIG. 2) to execute the operations. For example, the first logical operation is composed of a logical block (e.g., a dynamic logical block), and the input data (e.g., stored in the storage memory 114) is sent to the logical block that actively executes the first logical operation at stage 605. The logical block is configured to store its output in an intermediate output buffer (e.g., the intermediate I / O buffer 304 of FIG. 3) at stage 610.

[0058] When the output of the logic block fills the intermediate output buffer, it is monitored at step 615 to determine whether the intermediate output buffer has reached a threshold value (e.g., a high water mark (HWM)). If it is determined at step 615 that the HWM has not been reached (NO), it is determined at step 620 whether the first logical operation has been completed. If it is determined at step 620 that the first logical operation has not been completed (NO), the first logical operation is continued until it reaches the HWM at step 615 or until the first logical operation is completed at step 620. On the other hand, if it is determined at step 620 that the first logical operation has been completed (YES), the process continues at step 625 (A), which will be described with reference to FIG. 6B below.

[0059] If it is determined at step 615 that the HWM has been reached (YES), while the first logical operation is continued at step 630, a second logical operation is configured (e.g., in a second dynamic logic block). In this case, for example, while the second logical operation is being configured, the first logical operation continues to be executed, so the reconfiguration time of the second logical operation in the second logic block is hidden (e.g., not significant). In this case, in one embodiment, the second logical operation is an extension of the first operation. Method 600 works well when the first logical operation and the second logical operation form a throughput-oriented pipeline (e.g., a minimum version of pipeline workflow 400), but the present invention is not limited thereto.

[0060] If the first logical operation is interrupted, it continues to be executed until the end of the intermediate output (e.g., until the intermediate output buffer is full), or if the first logical operation is considered to be completed, the input data ends. Therefore, it is determined at step 635 whether the first logical operation has been interrupted. If the first logical operation has not been interrupted at step 635 (NO), it is executed until the first logical operation is completed, and in this case, the intermediate output buffer of the first logical operation is designated as the final output buffer at step 640. The data stored in the final output buffer is transmitted to the host device 102.

[0061] On the one hand, when it is determined in step 635 that the first logical operation has been interrupted (YES), in step 645, the intermediate output buffer of the first logical operation is designated as the input buffer for the second logical operation, and in step 650, the input buffer of the first logical operation is designated as the output buffer for the second logical operation. In step 655, the data in the intermediate output buffer (currently designated as the input buffer for the second logical operation) is processed by the second logical operation. Method 600 is repeated until there is no input, and the entire pipeline workflow of the operation is performed for all inputs.

[0062] Referring to FIG. 6B, when it is determined in step 620 that the first logical operation has been completed (YES), in step 625, the process continues (A), where it is determined whether there is an additional logical operation in the configured pipeline workflow. In step 625, if it is determined that there is no additional logical operation configured for the pipeline workflow (NO), in step 660, the intermediate output buffer of the first logical operation is designated as the final output buffer. The data stored in the final output buffer is transmitted to the host device 102.

[0063] On the other hand, in step 625, if it is determined that there is an additional logical operation configured for the pipeline workflow (YES), in step 665, the next logical operation is configured (e.g., in the second logic block). In step 670, the intermediate output buffer of the first logical operation is designated as the input buffer for the next logical operation, and in step 675, the input buffer of the first logical operation is designated as the output buffer for the next logical operation. In step 680, the data in the intermediate output buffer (currently designated as the input buffer for the next logical operation) is processed by the next logical operation, and method 600 is repeated until there is no input, and the entire pipeline workflow of the operation is performed for all inputs.

[0064] While one embodiment has been described with reference to the drawings, the present invention may be implemented in various different forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided as examples to thoroughly and completely disclose the present invention and to fully convey the aspects and features of the present invention to those skilled in the art. Therefore, descriptions of aspects and features of the present invention should generally be considered applicable to other similar aspects and features of other exemplary embodiments unless otherwise specified.

[0065] In this specification, terms such as "first," "second," "third," etc. are used to describe various elements, components, regions, layers, and / or sections, but these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are used to distinguish one element, component, region, layer, or section from another element, component, region, layer, or section. Therefore, a first element, component, region, layer, or section described hereinafter may refer to a second element, component, region, layer, or section without departing from the spirit and scope of the present invention.

[0066] The terms used in this specification are for the purpose of describing particular embodiments and are not intended to limit the present invention. The singular form "a" includes the plural representation unless the context clearly dictates otherwise. In this specification, terms such as "including," "comprising," "having," and "possessing" specify the presence of the disclosed features, numbers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, elements, components, and / or combinations thereof. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items, and expressions such as "at least one," when preceding a list of elements, modify the entire list of elements and not the individual elements of the list.

[0067] As used herein, the terms "substantially", "about", and similar terms are used as terms of approximation rather than degree, and are for explaining the inherent variations in measured or calculated values that would be recognized by those skilled in the art. Further, the term "can" used when describing embodiments of the present invention means "one or more embodiments of the present invention". The terms "use" and "used" as used herein are considered synonyms of the terms "utilize" and "utilized", respectively.

[0068] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Also, terms defined in commonly used dictionaries are to be interpreted as having a meaning consistent with the context of the relevant art and / or this specification, and are not to be interpreted in an idealized or overly formal sense unless clearly defined herein.

[0069] Figures 1-3 show exemplary packaging, but various functions and components can be arranged in other suitable ways by those skilled in the art by applying them to large-scale data processing or by applying semiconductor packaging, printed circuit board design, integrated circuit design, system design, and rack or system cluster design according to the number of required components.

[0070] Furthermore, any of the wirings shown in Figures 1-3 can be replaced by any suitable wired or wireless connection, from something as simple as conductive or optical connections within an integrated circuit, to silicon through vias between dies, packages, or chiplets, or other non-silicon optical, inductive, conductive, or capacitive connections, and printed circuit board traces, wire bonds, switches, or direct cable or wire connections between chips, packages, and / or systems, or as complex as those on the scale of an entire data center or rack-scale fabric.

[0071] The concept of the present invention in FIGS. 1 to 3 can be applied to systems of any suitable scale, from a single-core host processor to a multi-core host processor, from a single-channel host memory to a multi-channel including multiple devices such as DIMMs, from a single host to hundreds of thousands or more hosts, from one storage device to a fabric connected to multiple devices of each host or multiple hosts, from a device with one storage controller to a device with multiple controllers, from a device including one RIC device to a device including many different types of multiple devices.

[0072] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the technical idea of the present invention.

Explanation of Reference Numerals

[0073] 100 Storage system 102 Host device 104 Storage device 106 Host processor 108 Host memory 110 Storage interface 112 Storage controller 114 Storage memory 116 Reprogrammable integrated circuit (RIC) device 118 Wiring 120 RIC extended memory 202 RIC accelerator 204 RIC memory 206 Static logic block 208 Dynamic logic block 210 Input / output (I / O) buffer 212 First memory 214 Second memory 302 Input / output (read / write) buffer 304 Intermediate I / O buffer 306 Configuration buffer 308 Third Memory 400 Pipeline Workflow 402 First Stage 404 Second Stage 406 Third Stage 408 Fourth Stage 410 Fifth Stage 412 Sixth Stage 600 Acceleration Method

Claims

1. A storage device comprising: a storage controller configured to receive data from a host device and store the data in a storage memory; and a reconfigurable integrated circuit communicably connected to the storage controller and configured to accelerate logical operations performed on the data stored in the storage memory. The reconfigurable integrated circuit includes: a first logic block configured to execute static logical operations among the logical operations; a second logic block configured to execute one or more dynamic logical operations among the logical operations; and a plurality of memory buffers configured to store inputs and outputs of the first and second logic blocks, wherein the plurality of memory buffers include: an input / output (I / O) buffer configured to store inputs and outputs of the first and second logic blocks; an intermediate I / O buffer configured to store intermediate outputs of the second logic block while the second logic block is being reconfigured; and a configuration buffer configured to store a configuration file for reconfiguring the second logic block. [[ / ID=10]]

2. The logical operations correspond to a pipeline workflow, the first logic block is statically configured with the static logical operations for the pipeline workflow, and the second logic block is dynamically reconfigured with the one or more dynamic logical operations for at least one stage of the pipeline workflow. [[ / ID=14]]

3. The one or more dynamic logical operations include a first dynamic logical operation and a second dynamic logical operation, and the second logic block is configured with the first dynamic logical operation during a first stage of the pipeline workflow and is dynamically reconfigured with the second dynamic logical operation during a second stage of the pipeline workflow. [[ / ID=17]]

4. The second logic block is dynamically reconfigured by loading one configuration file from the configuration files stored in the configuration buffer into the second logic block. [[ / ID=19]]

5. Output of the second logic block is stored in the intermediate I / O buffer during a first stage. [[ / ID=21]] The second logic block is reconfigured with different dynamic logic instructions for the second stage, The intermediate I / O buffer is designated as an input buffer of the second logic block during the second stage, and the storage device according to claim 1 is characterized in that.

6. The static logic operation corresponds to a latency-critical operation, The one or more dynamic logic operations correspond to throughput-oriented operations, and the storage device according to claim 1 is characterized in that.

7. The latency-critical operation is an operation whose completion time is shorter than the reconfiguration time of the second logic block, and the storage device according to claim 6 is characterized in that.

8. The storage device is a solid state drive, and the storage device according to claim 1 is characterized in that.

9. The reconfigurable integrated circuit is a field programmable gate array (FPGA), and the storage device according to claim 8 is characterized in that.

10. A method for accelerating operations in a storage device including a storage controller, a storage memory, and a reconfigurable integrated circuit including a first logic block, a second logic block, and a buffer, Executing, by the first logic block, a first logic operation on input data stored in the storage memory; Storing, by the first logic block, the output of the first logic operation in an intermediate output buffer of the buffer; Configuring, by the reconfigurable integrated circuit, a second logic operation in the second logic block; Designating, by the reconfigurable integrated circuit, the intermediate output buffer as an input buffer for the second logic operation; Executing, by the second logic block, the second logic operation on the output of the first logic operation stored in the intermediate output buffer, and having, While the first logic operation is being executed in the first logic block, the second logic operation is being configured in the second logic block, The buffer includes a configuration buffer configured to store a configuration file for configuring the second logic block, and a method for accelerating operations in a storage device is characterized in that.

11. The step of configuring the second logic operation in the second logic block includes Monitoring the value of the intermediate output buffer; Determining whether the value of the intermediate output buffer exceeds a threshold value; When it is determined that the value of the intermediate output buffer exceeds a threshold value, the method for accelerating the operation of the storage device according to claim 10, comprising: the step of configuring the second logical operation in the second logical block.

12. The method for accelerating the operation of the storage device according to claim 11, wherein the threshold value is the high water mark of the intermediate output buffer.

13. The step of configuring the second logical operation in the second logical block includes the step of loading a bit file corresponding to the second logical operation from the configuration file stored in the configuration buffer into the second logical block, the method for accelerating the operation of the storage device according to claim 10.

14. The step of designating the intermediate output buffer as an input buffer for the second logical operation includes: the step of determining whether the first logical operation has been interrupted; when it is determined that the first logical operation has been interrupted, the step of designating the intermediate output buffer as an input buffer for the second logical operation; the step of designating the input buffer of the first logical operation as an output buffer for the second logical operation, the method for accelerating the operation of the storage device according to claim 10.

15. The step of determining whether the first logical operation has been interrupted includes: the step of determining whether the end of the intermediate output buffer has been reached, the method for accelerating the operation of the storage device according to claim 14.

16. the step of determining whether the second logical block has processed all the outputs of the first logical operation stored in the intermediate output buffer; the method for accelerating the operation of the storage device according to claim 10, further comprising: the step of designating the output buffer of the second logical operation as the final output buffer.

17. The storage device is a solid state drive, The method for accelerating the operation of the storage device according to claim 10, wherein the reconfigurable integrated circuit is an FPGA.

Citation Information

Patent Citations

  • Intelligent data storage and processing using FPGA device

    JP2012014705A

  • Solid state drive

    JP2018005686A

  • Data-centric computing architecture based on storage server in NDP server data center

    JP2019185764A

  • Configurable Logical Platform

    JP2019530099A