Semiconductor device, computing system and data computing method
By adopting semiconductor devices and computing systems in heterogeneous computing systems and using PRP or SGL to indicate memory addresses, the data bottleneck problem is solved, the bus utilization rate is improved and the latency is reduced, and efficient data transmission and storage is achieved.
Patent Information
- Application Number
- CN202410951043.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2024-07-16
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, heterogeneous computing systems have data bottlenecks in data transmission and storage, resulting in problems of low bus utilization and high latency.
Using semiconductor devices and computing systems, data computing and memory management can be achieved by using chip-to-chip interconnect protocols, combined with physical area pages (PRP) or distributed aggregation lists (SGL) to indicate memory addresses, enabling efficient management of data computing and memory, improving bus utilization and reducing latency.
Improves the bus utilization of the computing system, reduces power consumption, and optimizes the efficiency of data transmission and storage.
Smart Images

Figure CN120371197A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority and benefit of Korean Patent Application No. 10 - 2024 - 0011774, filed with the Korean Intellectual Property Office on January 25, 2024, the entire content of which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to a semiconductor device, a computing system, and a data computing method. Background Art
[0004] The rapid growth of data and the use of specialized workloads such as compression, encryption, and artificial intelligence are driving the demand for heterogeneous computing, in which specialized accelerators work in cooperation with general - purpose processors.
[0005] These accelerators need to have high - performance connections with processors, and sharing storage space is an ideal choice for reducing overhead and latency. Accordingly, there is research on chip - to - chip interconnection protocols for connecting processors to various accelerators to maintain storage and cache coherence. Summary of the Invention
[0006] The present disclosure is intended to provide a semiconductor device, a computing system, and a data computing method that can improve bus utilization by alleviating data bottlenecks in a scenario using a chip - to - chip interconnection protocol.
[0007] A semiconductor device may include: a memory configured to store data; a controller configured to monitor input / output data of the memory, determine a calculation to be performed on first data among the input / output data in response to receiving a calculation command and address information from a host, the address information including an instruction for an address of the memory using at least one of a physical region page (PRP) or a scatter - gather list (SGL), and the calculation being determined based on the calculation command and the address information; and a calculation engine configured to perform the determined calculation on the first data.
[0008] A computing system may include: a host configured to generate and output a data command, a calculation command, and address information; a plurality of Compute Express Link (CXL) storage devices configured to perform data operations based on the data command and the address information; and a semiconductor device configured to perform data calculations based on the calculation command and the address information; wherein the computing system may be configured to use at least one of a physical region page (PRP) or a scatter - gather list (SGL) to generate address information indicating a memory address of the semiconductor device.
[0009] A data calculation method may include: receiving a command and address information of at least one of a physical region page (PRP) or a scatter-gather list (SGL) from a host; monitoring input / output data; determining whether a device physical address of the address information and first data among the input / output data match each other; and when the device physical address and the address information match each other, performing a calculation on the first data according to the command. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a block diagram showing a computing system according to at least one embodiment.
[0011] Figure 2 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0012] Figure 3 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0013] Figure 4 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0014] Figure 5 is a block diagram showing a CXL storage device according to at least one embodiment.
[0015] Figure 6 is a diagram for explaining a physical region page according to at least one embodiment.
[0016] Figure 7 is a diagram for explaining a physical region page according to at least one embodiment.
[0017] Figure 8 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0018] Figure 9 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0019] Figure 10 is a block diagram of a semiconductor device according to at least one embodiment.
[0020] Figure 11 is a block diagram of a semiconductor device according to at least one embodiment.
[0021] Figure 12 is a diagram for explaining monitoring performed by a controller based on address information according to at least one embodiment.
[0022] Figure 13 and Figure 14 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0023] Figure 15 is a flowchart of a data calculation method according to at least one embodiment.
[0024] Figure 16 is a block diagram of a computer system according to at least one embodiment.
[0025] Figure 17 is a block diagram of a computer system according to at least one embodiment.
[0026] Figure 18 is a block diagram of a data center of an application computer system according to at least one embodiment. Detailed implementation manners
[0027] In the following detailed description, only some embodiments of the present disclosure are shown and described by way of illustration. As those skilled in the art will recognize, the described embodiments can be modified in various different ways, all of which do not depart from the spirit or scope of the present invention.
[0028] Therefore, the drawings and the description should be considered illustrative in nature and not restrictive. Throughout the specification, like reference numerals represent like elements. In the flowcharts described with reference to the drawings, the order of operations can be changed, multiple operations can be combined, some operations can be divided, and specific operations can be not executed. Additionally, in this specification, functional elements and / or devices including units having and / or configured to have at least one function or operation (e.g., “controller” and / or “… engine”) can be implemented by a processing circuit including hardware, software, or a combination of hardware and software. For example, more specifically, the processing circuit can include, but is not limited to, a central processing unit (CPU), a neural processing unit (NPU), a graphics processing unit (GPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application specific integrated circuit (ASIC), etc.
[0029] Any or all of the elements described with reference to the drawings can communicate with any or all of the other elements described with reference to the corresponding drawings. For example, any element in the drawings (e.g., Figure 10 ) can communicate unidirectionally and / or bidirectionally and / or broadcast communicate with any or all of the other elements in the same drawing (e.g., Figure 10 ) to transmit, and / or exchange, and / or receive information (e.g., but not limited to, data and / or commands) in a manner such as serial and / or parallel via a bus such as a wireless and / or wired bus (not shown). The information can be encoded in various formats (e.g., analog format and / or digital format).
[0030] In addition, unless an explicit expression such as "a" or "a single" is used, an expression written in the singular may be interpreted as singular or plural. Terms including ordinal numbers such as first, second, etc. will only be used to describe various components and should not be construed as limiting these components. These terms may be used for the purpose of distinguishing one component from other components.
[0031] Figure 1 is a block diagram showing a computing system according to at least one embodiment.
[0032] Referring to Figure 1 , the computing system 10 according to at least one embodiment may be included in, for example, a user device (e.g., a personal computer, a laptop computer, a server, a data center, a media player, a digital camera, etc.) and / or an automotive device (e.g., a navigation device, a black box, automotive electronics, etc.) and / or a mobile system (e.g., a portable communication terminal (e.g., a mobile phone), a smart phone, a tablet personal computer (tablet), a wearable device, a healthcare device, an Internet of Things (IoT) device, etc.).
[0033] The computing system 10 includes a plurality of hosts 11 to 14 and a plurality of Compute Express Link (CXL) devices 31 to 33. The plurality of hosts 11 to 14 and the plurality of CXL devices 31 to 33 may be connected to the cache coherence interface 20 through different physical ports. That is, since the plurality of CXL devices 31 to 33 are connected to the cache coherence interface 20, the storage area managed by the plurality of hosts 11 to 14 can increase in capacity, and the plurality of CXL devices 31 to 33 can also exchange data. The plurality of hosts 11 to 14 may include a first host to an Mth host (M is an integer greater than 1). The plurality of CXL devices 31 to 33 may include a first CXL device to an Nth CXL device (N is an integer greater than 1).
[0034] The cache coherence interface 20 may be configured to support a dynamic protocol mix of coherence, memory access, and input / output protocol (IO protocol), and indicate low-latency and high-bandwidth links, thereby enabling various connections between accelerators, memory devices, and / or various electronic devices.
[0035] In at least one embodiment, the cache coherence interface 20 may be implemented as a CXL interface. For example, the CXL interface may include sub-protocols (e.g., CXL.io protocol, CXL.mem protocol, CXL.cache protocol, etc.). The CXL.io protocol may include I / O semantics that may be similar to the Peripheral Component Interconnect Express (PCIe) standard. The CXL.cache protocol may include cache semantics, the CXL.mem protocol may include memory semantics, and both the cache semantics and the memory semantics may be optional. For example, in at least one embodiment, multiple hosts 11 to 14 may be configured to send command signals to multiple CXL devices 31 to 33 via the CXL.io protocol and may be configured to receive data corresponding to the command signals via the CXL.io protocol; and the multiple CXL devices 31 to 33 may be configured to exchange data with each other via the CXL.mem protocol.
[0036] The cache coherence interface 20 is not necessarily limited to the CXL interface, and the multiple hosts 11 to 14 and the multiple CXL devices 31 to 33 may communicate with each other based on various computing interfaces (e.g., GEN-Z protocol, NVLink protocol, Accelerator Cache Coherence Interconnect (CCIX) protocol, Open Coherent Accelerator Processor Interface (CAPI) protocol, etc.).
[0037] In at least one embodiment, the multiple hosts 11 to 14 may be configured to control the overall operation of the computing system 10. In at least one embodiment, the multiple hosts 11 to 14 may be (and / or include) processing circuitry. In some example embodiments, the multiple hosts 11 to 14 may be one of various processors such as a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), and / or a data processing unit (DPU). In at least one embodiment, the multiple hosts 11 to 14 may include a single-core processor or a multi-core processor. Each of the hosts 11 to 14 may be connected to at least one memory device via a double data rate (DDR) interface and exchange data. The multiple hosts 11 to 14 may include memory controllers configured to control at least one memory device. For example, in at least some embodiments, at least one of the multiple hosts 11 to 14 may include a memory controller, and / or each of the multiple hosts 11 to 14 may include a memory controller. However, the scope of the present disclosure is not limited thereto, and the multiple hosts 11 to 14 may communicate with at least one memory device via various interfaces.
[0038] In at least one embodiment, multiple hosts 11 to 14 may be configured to send read requests to multiple CXL devices 31 to 33, and the multiple CXL devices 31 to 33 may be configured to send data to the multiple hosts 11 to 14 based on the read requests. In at least one embodiment, the multiple hosts 11 to 14 may divide tasks, distribute the tasks to the multiple CXL devices 31 to 33, and collect the results. The multiple hosts 11 to 14 may distribute tasks including data and programs (e.g., workloads) for processing the data to the multiple CXL devices 31 to 33.
[0039] Each of the multiple CXL devices 31 to 33 may be a memory device (or memory module) or a storage device (or storage module). The memory device may be a dynamic random access memory (DRAM) device and may have various form factors (e.g., dual in-line memory module (DIMM), high bandwidth memory (HBM), etc.). However, the scope of the present disclosure is not limited thereto, and the memory device may include non-volatile memory (e.g., flash memory, phase change random access memory (PRAM), resistive random access memory (RRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), etc.). Additionally, the storage device may be implemented as various types of storage devices (e.g., solid state drive (SSD), embedded multimedia card (eMMC), universal flash storage (UFS), compact flash (CF), secure digital (SD), micro secure digital (micro-SD), mini secure digital (Mini-SD), extreme digital (xD), and / or memory stick).
[0040] In at least one example of the computing system 10, at least a first CXL device 31 among the multiple CXL devices 31 to 33 may be configured for data calculation. For example, the first CXL device 31 may be configured to perform calculations on input data or output data. For example, the first CXL device 31 may receive data from the multiple hosts 11 to 14, or the second CXL device 32 to the Nth CXL device 33; may output the data to the multiple hosts 11 to 14, or the second CXL device 32 to the Nth CXL device 33; and may perform calculations on this data. Examples of data calculation will be further explained in detail below.
[0041] According to an embodiment, the first CXL device 31 may perform data calculations (e.g., parity generation, data reconstruction, encryption, decryption, compression, decompression, deduplication, pattern search (and / or replacement), data integrity check, structured query language (SQL), statistics, etc.).
[0042] In at least one embodiment, data can be recorded in a compressed manner so that at least some of the multiple hosts 11 to 14 can more widely utilize the storage space. The first CXL device 31 can perform decompression on the compressed data.
[0043] In at least one embodiment, as a data integrity check, the first CXL device 31 can perform end-to-end (E2E) verification on the data generated by the second CXL device 32 to the Nth CXL device 33. For example, the second CXL device 32 to the Nth CXL device 33 can store data blocks in the data format of data integrity field (DIF) / data integrity extension (DIX). The DIF / DIX data format can include a data block field, a reference tag field, an application tag field, and a protection field. The first CXL device 31 can verify the DIF / DIX data format.
[0044] When the first CXL device 31 performs data calculation, it is not necessary to send the data for calculation to the multiple hosts 11 to 14. Accordingly, the computing system 10 can improve the bus utilization of the cache coherence interface 20 and reduce power consumption.
[0045] Figure 2 is a block diagram for explaining the operation of a computing system according to at least one embodiment.
[0046] Reference Figure 2 , a computing system 100 according to at least one embodiment can include a host 110, a memory 111, a CXL switch 120, multiple CXL storage devices 130, and a semiconductor device 150. Although for better understanding and ease of description, Figure 2 shows that the computing system 100 includes one host 110, but the embodiment is not limited thereto, and the computing system 100 can include one or more hosts. For example, the host 110 can be one of the multiple hosts 11 to 14 described in reference Figure 1 . Additionally, the semiconductor device 150 and the multiple CXL storage devices 130 can be the multiple CXL devices 31 to 33 described in reference Figure 1 (and / or included in the multiple CXL devices 31 to 33 described in reference Figure 1 ). For example, the semiconductor device 150 can represent the first CXL device 31 configured for data calculation.
[0047] The host 110 can send commands to at least one of the multiple CXL storage devices 130 and the semiconductor device 150. For example, the host 110 can include a CXL controller, and the CXL controller can communicate with the multiple CXL storage devices 130 and the semiconductor device 150 through the CXL switch 120.
[0048] The host 110 may be connected to a memory 111 that stores data. In at least one embodiment, the memory 111 may be configured as the main memory and / or system memory of the computing system 100. In at least one embodiment, the memory 111 may operate as a buffer. In at least one embodiment, the memory 111 may be a dynamic random access memory (DRAM) device and may have various form factors (e.g., dual in-line memory module (DIMM), high bandwidth memory (HBM), etc.). However, the scope of the present disclosure is not limited thereto, and the memory 111 may include non-volatile memory (e.g., flash memory, phase change RAM (PRAM), resistive RAM (RRAM), magnetoresistive RAM (MRAM), ferroelectric RAM (FRAM), etc.).
[0049] In at least one embodiment, the memory 111 may be configured to communicate with the host 110 through an interface. For example, the memory 111 may be configured to directly communicate with the host 110 through a DDR interface. In at least one embodiment, the host 110 may include a memory controller configured to control the memory 111. However, the scope of the present disclosure is not limited thereto, and the memory 111 may communicate with the host 110 through various interfaces. The host 110 may be connected to one or more memories 111 and may communicate with one or more memories 111.
[0050] The memory 111 may be configured to store data received from at least one of a plurality of CXL storage devices 130 and semiconductor devices 150. According to an embodiment, the memory 111 may be implemented as a memory located inside the host 110.
[0051] At least one of the plurality of CXL storage devices 130 and semiconductor devices 150 may receive commands from the host 110. The plurality of CXL storage devices 130 may be implemented as various types of storage devices (e.g., solid state drive (SSD), embedded multimedia card (eMMC), universal flash storage (UFS), compact flash (CF), secure digital (SD), micro-SD, mini-SD, extreme digital (xD), memory stick, etc.). When the plurality of CXL storage devices 130 are SSDs, the SSDs may be devices that comply with the non-volatile memory express (NVMe) standard. In at least one embodiment, the SSD may be a DRAM-less SSD. Additionally, when the plurality of CXL storage devices 130 are embedded memories or external memories, the embedded memories or external memories may be devices that comply with the UFS standard or the eMMC standard. The host 110 and the plurality of CXL storage devices 130 may each generate and send data packets according to the standard protocols employed.
[0052] A plurality of CXL storage devices 130 may include a first CXL storage device 130_1 to a Gth CXL storage device 130_G (G is an integer greater than 1). For example, each of the plurality of CXL storage devices 130 may include at least one of non-volatile memories (NVMs) 131_1 to 131_G and at least one of controllers (CTRLs) 132_1 to 132_G. The non-volatile memories 131_1 to 131_G may store data. The non-volatile memories 131_1 to 131_G may record, output, and / or erase data based on the control of the controllers 132_1 to 132_G. The controllers 132_1 to 132_G may control the operation of the plurality of CXL storage devices 130 based on commands from the host 110. For example, the controllers 132_1 to 132_G may record data in the non-volatile memories 131_1 to 131_G based on commands from the host 110, and / or may output or erase the data in the non-volatile memories 131_1 to 131_G.
[0053] In at least one embodiment, the controllers 132_1 to 132_G may be configured to record, read, and / or erase data in other devices. For example, the controller 132_1 of the first CXL storage device 130_1 may record data in the non-volatile memory 131_G of the Gth CXL storage device 130_G, and / or may output or erase the data in the non-volatile memory 131_G. Additionally, the controller 132_1 may record data in the semiconductor device 150, and / or may output or erase the data recorded in the semiconductor device 150.
[0054] Controllers 132_1 to 132_G can be connected to the CXL switch 120. Controllers 132_1 to 132_G can communicate with the host 110 and / or other CXL devices via the CXL switch 120. Controllers 132_1 to 132_G can include a PCIe 5.0 (or other version) architecture for the CXL.io path and can add CXL-specific CXL.cache and CXL.mem paths. In another embodiment, controllers 132_1 to 132_G can be configured to be backward compatible with previous cache coherence protocols (e.g., CXL 1.1, CXL 2.0, etc.). Controllers 132_1 to 132_G can be configured to implement CXL.io, CXL.mem, and CXL.cache protocols, or other suitable cache coherence protocols. Controllers 132_1 to 132_G can be configured to support different CXL device types (e.g., type 1, type 2, and / or type 3 CXL devices). Controllers 132_1 to 132_G can be configured to support PCIe protocols (e.g., PCIe 5.0 protocol, PCIe 6.0 protocol, etc.). Controllers 132_1 to 132_G can be configured to support the PIPE 5.x protocol by using any suitable PHY interface for the PCI Express (PIPE) interface width (capable of configuring PIPE interface widths such as 8 bits, 16 bits, 32 bits, 64 bits, and 128 bits).
[0055] In some embodiments, at least one of controllers 132_1 to 132_G can include intellectual property (IP) circuits designed to implement application-specific integrated circuits (ASICs) and / or field-programmable gate arrays (FPGAs). In various embodiments, controllers 132_1 to 132_G can be implemented to support the CXL interface (e.g., the CXL 3.0 specification or any other version).
[0056] In some embodiments, the multiple CXL storage devices 130 can also include components for storage operations (e.g., direct memory access (DMA) engines, address engines, etc.). The address engine can read recorded data based on address information defined in the NVMe standard (e.g., physical region pages (PRPs), scatter-gather lists (SGLs), etc.).
[0057] The semiconductor device 150 can be implemented as a CXL memory device including a computing engine, persistent memory (PMEM), a CXL storage device, etc. The semiconductor device 150 can perform computations on input data or output data. The semiconductor device 150 can monitor data and, when specific conditions are met, can perform computations on the data.
[0058] For example, the semiconductor device 150 may receive address information from the host 110. The host 110 may send the address information together with a write command to a plurality of CXL storage devices 130. The plurality of CXL storage devices 130 may read data from the semiconductor device 150 in response to the write command. When the address for reading the data matches the address information of the host 110, the semiconductor device 150 may perform calculations on the data.
[0059] The host 110, the plurality of CXL storage devices 130, and the semiconductor device 150 may communicate with each other via a CXL switch 120. The CXL switch 120 may be a component included in the CXL interface. The CXL switch 120 may be used to implement a memory cluster through one-to-many and many-to-one switching between the semiconductor device 150 and the connected plurality of CXL storage devices 130. For example, the CXL switch 120 may: (i) connect a plurality of routing ports to one endpoint, (ii) connect one routing port to a plurality of endpoints, and / or (iii) connect a plurality of routing ports to a plurality of endpoints.
[0060] The CXL switch 120 may provide a packet switching function for CXL packets. The CXL switch 120 may be used to connect a plurality of CXL storage devices 130 to one or more hosts 110. The CXL switch 120 may enable the plurality of CXL storage devices 130 to: (i) include various types of memories with different characteristics, (ii) virtualize the memories of the plurality of CXL storage devices 130, and store data with different attributes (e.g., access frequency) in an appropriate type of memory, and / or (iii) support Remote Direct Memory Access (RDMA). Here, "virtualizing" the memory means performing memory address translation between the processing circuit and the memory.
[0061] The CXL switch 120 may be configured to mediate communication between the host 110, the plurality of CXL storage devices 130, and the semiconductor device 150. For example, when the host 110 communicates with the plurality of CXL storage devices 130 and the semiconductor device 150, the CXL switch 120 may be configured to transmit information (e.g., requests, data, responses, and / or signals, etc.) transmitted from the host 110, the plurality of CXL storage devices 130, and / or the semiconductor device 150 via the CXL.io protocol.
[0062] When the plurality of CXL storage devices 130 and the semiconductor device 150 communicate with each other, the CXL switch 120 may be configured to transmit information (e.g., requests, data, responses, signals, etc.) between the plurality of CXL storage devices 130 and the semiconductor device 150 via the CXL.mem protocol.
[0063] The plurality of CXL memory devices 130 and semiconductor devices 150 may communicate with each other through the CXL.mem protocol and operate independently of the CXL.io protocol of the host 110. Therefore, the communication between the plurality of CXL memory devices 130 and semiconductor devices 150 may have limited impact on the input / output delay of the host 110.
[0064] Figure 3 , Figure 4 , Figure 8 and Figure 9 is a block diagram for explaining the operation of a computing system according to at least one embodiment. Figure 5 is a block diagram illustrating a CXL storage device according to at least one embodiment. Figure 6 and Figure 7 is a diagram for explaining a physical area page according to at least one embodiment.
[0065] refer to Figure 3 ,refer to Figure 2 The description of can be equally applied to the computing system 100 according to at least one embodiment. Therefore, redundant description will be omitted.
[0066] The host 110 may generate address information INF_ADR. The address information INF_ADR may include an address of data stored in the semiconductor device 150. The address information INF_ADR may be transmitted to the plurality of CXL storage devices 130 together with a command (eg, write). In an embodiment, the address information INF_ADR may include PRP or SGL of the NVMe standard.
[0067] The host 110 may generate address information INF_ADR for each of the plurality of CXL storage devices 130. For example, the host 110 may generate first to Gth address information corresponding to the first to Gth CXL storage devices 130_1 to 130_G, respectively. At this time, the address information INF_ADR may include the first to Gth address information.
[0068] In at least one embodiment, the plurality of CXL storage devices 130 may form a redundant array of inexpensive disks (RAID) group to implement the RAID technology. According to an embodiment, some of the plurality of CXL storage devices 130 may form a RAID group. That is, each storage device forming the RAID group among the plurality of CXL storage devices 130 may store divided data blocks. At this time, the logical block address (LBA) of each storage device storing the data block may be the same. The host 110 may generate address information INF_ADR for the storage devices forming the RAID group to store the data blocks.
[0069] The host 110 may transmit a command CMD1 to the semiconductor device 150. The host 110 may transmit the command CMD1 through the CXL switch 120, and the command CMD1 may indicate data computation (e.g., parity generation, data reconstruction, encryption, decryption, compression, decompression, pattern search, data integrity check, Structured Query Language (SQL), statistics, etc.). In at least one embodiment, the command CMD1 may indicate parity generation. Additionally, the command CMD1 may include address information INF_ADR. In at least one embodiment, the command CMD1 may be a mailbox command of the CXL standard. The semiconductor device 150 may store the address information INF_ADR and may perform computations based on the command CMD1.
[0070] In at least one embodiment, the host 110 may generate address information of the memory 111, and the address information may be transmitted to multiple CXL storage devices 130 together with commands such as read, write, etc.
[0071] Reference Figure 4 Referring to, the host 110 may transmit commands CMD2 and CMD3 to multiple CXL storage devices 130. The controllers 132_1 to 132_G may be configured to control the operations of the multiple CXL storage devices 130 based on the commands CMD2 and CMD3. In at least one embodiment, the commands CMD2 and CMD3 may be write commands, and the controllers 132_1 to 132_G may record data in the non-volatile memories 131_1 to 131_G.
[0072] In at least one embodiment, the commands CMD2 and CMD3 may follow the command format of the NVMe standard. For example, the command format may include "Command Identifier (CID)", "Namespace Identifier (NSID)", "Reserved Field", "Metadata", "PRP1", "PRP2", "Initial LBA (SLBA)", "Length", etc.
[0073] "PRP1" and "PRP2" may indicate the locations storing the data to be accessed. "PRP1" may indicate the first physical region page. "PRP2" may indicate the second physical region page. The first physical region page and the second physical region page may indicate the addresses of the memory (e.g., memory 111 or semiconductor device 150) for DMA communication.
[0074] In at least one embodiment, when two or more physical region pages are required to indicate data, the PRP list may include pointers indicating the physical region pages. For example, "PRP1" may indicate the first PRP list, "PRP2" may indicate the second PRP list, and the last one of the second PRP list may indicate the third PRP list.
[0075] In at least one embodiment, when implementing two or more scattered gather list (SGL) segments to indicate data, the SGL segments may include pointers to subsequent SGL segments.
[0076] The host 110 may transmit the address information INF_ADR together with the commands CMD2 and CMD3. The host 110 may transmit the commands CMD2 and CMD3 and the address information INF_ADR through the CXL switch 120. For example, the host 110 may transmit the command CMD2 and the first address information to the first CXL storage device 130_1, and may transmit the command CMD3 and the Gth address information to the Gth CXL storage device 130_G. The PRP (or SGL) of the command CMD2 may indicate the first address information, and the PRP (or SGL) of the command CMD3 may indicate the Gth address information. That is, the multiple storage devices 130 may access the semiconductor device 150 based on the commands CMD2 and CMD3 and the address information INF_ADR. When the commands CMD2 and CMD3 are write commands, the multiple storage devices 130 may record data in the specified LBAs of the non-volatile memories 131_1 to 131_G.
[0077] Reference Figure 5 , the CXL storage device 200 according to at least one embodiment may include a memory 210, a direct memory access engine (DMA) 220, an address engine (ADDR) 230, and a non-volatile memory (NVM) 240. Figure 4 At least one of the multiple CXL storage devices 130 of Figure 5 may be implemented as
[0078] The memory 210 may be a device memory managed by the host 110 (e.g., host-managed device memory (HDM)) (see Figure 4 ). For example, the memory 210 may be a volatile memory, but is not limited thereto, and may be implemented as a non-volatile memory (e.g., single-level cell (SLC) type flash memory). The memory 210 may receive and store the calculation results from the semiconductor device 150 (see, for example, Figure 4 ).
[0079] The memory 210 may be configured to operate as a buffer. For example, the memory 210 may temporarily store the write data received from the outside (e.g., the data obtained from the semiconductor device 150), and the write data may be moved to the non-volatile memory 240 and recorded in the non-volatile memory 240.
[0080] The DMA engine 220 can be configured to fetch or transfer data based on the source address and destination address set by the address engine 230. For example, the DMA engine 220 can record the fetched data in the non-volatile memory 240, and / or transfer the data in the non-volatile memory 240 to the semiconductor device 150.
[0081] The address engine 230 can be configured to determine the source address and destination address based on the commands and address information of the host 110. The address engine 230 can parse the entries of the address information based on the commands of the host 110, and can detect the locations referred to by each entry within the semiconductor device 150.
[0082] In at least some embodiments, the address engine 230 can be implemented as a PRP engine or an SGL engine. For example, when the address information is in the PRP form, the address engine 230 can be implemented as a PRP engine, and when the address information is in the SGL form, the address engine 230 can be implemented as an SGL engine. The DMA engine 220 and the address engine 230 can be included in the controllers 132_1 to 132_G (see Figure 4 ).
[0083] The non-volatile memory 240 can store data. For example, the non-volatile memory 240 can store data blocks, calculation results (such as parity, etc.), etc. The non-volatile memory 240 can store the data of the memory 210.
[0084] According to an embodiment, the CXL storage device 200 may further include components required for storage operations.
[0085] Figure 6 and Figure 7 Addresses indicated by PRP1 and PRP2 included in the command of the host 110 according to at least one embodiment are explained.
[0086] Reference Figure 6 , according to at least one embodiment, the command of the host 110 may include "PRP1", and "PRP1" may indicate the first PRP list PRP LIST 1. The first PRP list may include PRP entries PRP Entry#0 to PRP Entry#2. The PRP entries PRP Entry#0 to PRP Entry#2 may include a page base address and an offset. In at least one embodiment, the offset may be zero.
[0087] The PRP entries PRP Entry#0 to PRP Entry#2 may indicate memory pages (the first memory page to the third memory page). The memory pages may represent Figure 4memory pages of the semiconductor device 150, but the embodiments are not limited thereto. For example, the PRP entry PRP Entry#0 may indicate the first memory page, the PRP entry PRP Entry#1 may indicate the second memory page, and the PRP entry PRP Entry#2 may indicate the third memory page.
[0088] The first to third memory pages may be discontinuous with each other. In at least some embodiments, each of the first to third memory pages may be 4 KiB (64 × 64 bytes). That is, 4 KiB of data of the first memory page can be obtained based on the PRP entry PRP Entry#0, 4 KiB of data of the second memory page can be obtained based on the PRP entry PRP Entry#1, and 4 KiB of data of the third memory page can be obtained based on the PRP entry PRP Entry#2. In this way, 12 KiB of data that is discontinuously present in the semiconductor device 150 can be sequentially obtained based on "PRP1".
[0089] In at least one embodiment, "PRP1" may indicate a memory page. That is, the CXL storage device can detect the position of the memory page indicated by "PRP1" without using the PRP list.
[0090] Reference Figure 7 , according to at least one embodiment, the command of the host 110 may include "PRP2", and "PRP2" may indicate the second PRP list PRP LIST 2. The second PRP list may include PRP entries PRP Entry#0 to PRP Entry#62 and a PRP list pointer. The PRP entries PRP Entry#0 to PRP Entry#62 may include a page base address and an offset. In at least one embodiment, the offset may be zero. In at least one embodiment, when more addresses than the number of addresses indicated by using the second PRP list are to be indicated, the host 110 may include the PRP list pointer in the second PRP list. That is, when the number is less than or equal to the number of addresses indicated by using the second PRP list, the host 110 may indicate the addresses by using the second PRP list.
[0091] The PRP list pointer can be located at the tail of the second PRP list and can indicate the third PRP list, PRP LIST3. The third PRP list can include PRP entries from PRP Entry#63 to PRP Entry#65. The PRP entry PRP Entry#0 of the second PRP list can indicate the Jth memory page, and the PRP entry PRP Entry#62 can indicate the Kth memory page. Additionally, the PRP entry PRP Entry#63 of the third PRP list can indicate the Pth memory page, the PRP entry PRP Entry#64 can indicate the Qth memory page, and the PRP entry PRP Entry#65 can indicate the Rth memory page. J, K, P, Q, and R can represent integers greater than 1. Figure 7 The memory pages can be discontinuous. In this way, 264 KiB of data that exists discontinuously in the semiconductor device 150 can be sequentially obtained based on "PRP2".
[0092] Reference Figure 8 , multiple CXL storage devices 130 can obtain the data DATA1 and DATA2 of the semiconductor device 150 based on the commands CMD2 and CMD3 and the address information INF_ADR. The multiple CXL storage devices 130 can obtain the data DATA1 and DATA2 in DMA format. For example, the commands CMD2 and CMD3 can be write commands, and the first CXL storage device 130_1 can obtain the data DATA1 by sending a read command and the address corresponding to the first address information to the semiconductor device 150 based on the commands CMD2 and the first address information. In the same way, the Gth storage device 130_G can obtain the data DATA2 by sending a read command and the address corresponding to the Gth address information to the semiconductor device 150 based on the commands CMD3 and the Gth address information. The semiconductor device 150 can transmit the data DATA1 and DATA2 based on the read command and the address.
[0093] In at least one embodiment, the multiple CXL storage devices 130 can include an address engine and a DMA engine. For example, the controllers 132_1 to 132_G can include an address engine and a DMA engine. The address engine can be implemented as a PRP engine and / or an SGL engine. The multiple CXL storage devices 130 can obtain data by using the address engine and the DMA engine. Each address engine can set the source address and destination address of the DMA engine based on the address information. The address engine can parse the entries of the address information based on the commands of the host 110 and can detect the locations referred to by each entry within the semiconductor device 150.
[0094] The address engine can access the semiconductor device 150 based on the address information and set the source address and the destination address. In at least one embodiment, when the command of the host 110 indicates a write, the source address can be the address of the semiconductor device 150, and the destination address can be the address of the CXL storage device 130. In at least one embodiment, when the command of the host 110 indicates a read, the source address can be the address of the CXL storage device 130, and the destination address can be the address of the semiconductor device 150. In at least one embodiment, the address of the semiconductor device 150 can be set based on PRP or SGL, and the address of the CXL storage device 130 can be set based on the logical address. That is to say, the address of the CXL storage device 130 can be the logical block address.
[0095] The DMA engine can retrieve data from the semiconductor device 150 based on the source address and the destination address, and / or transfer the data to the semiconductor device 150. In at least one embodiment, when the command of the host 110 indicates a write, the DMA engine can read data from the semiconductor device 150 and can write the data into a buffer in the memory (e.g., the CXL storage device 130) or into the non-volatile memories 131_1 to 131_G. In at least one embodiment, when the command of the host 110 indicates a read, the DMA engine can read data from the memory and can write the data into the semiconductor device 150.
[0096] The semiconductor device 150 can monitor the acquired data DATA1 and DATA2. For example, the semiconductor device 150 can monitor the data DATA1 and DATA2 based on the address information INF_ADR received from the host 110 (see Figure 3 ). The semiconductor device 150 can determine whether to output the data DATA1 and DATA2 according to the address information INF_ADR. That is to say, the semiconductor device 150 can determine whether the data DATA1 and DATA2 are output in the order indicated by the address information INF_ADR.
[0097] The semiconductor device 150 can monitor the data input / output based on the device physical address obtained by using the mapping table. The semiconductor device 150 can determine whether the device physical address is sequentially activated in the address information INF_ADR.
[0098] When outputting the data DATA1 and DATA2 according to the address information INF_ADR, the semiconductor device 150 can perform calculations on the data DATA1 and DATA2. The semiconductor device 150 can perform according to the command CMD1 (see Figure 3)Perform calculations on DATA1 and DATA2. For example, the semiconductor device 150 may perform calculations such as generating a parity check from DATA1 and DATA2. The semiconductor device 150 may store the calculation result in an internal memory. According to an embodiment, the internal memory may be, for example, DRAM and / or static RAM (SRAM). The semiconductor device 150 may report the calculation result to the host 110.
[0099] Reference Figure 9 , the semiconductor device 150 may report DATA3 as the calculation result to the host 110. In at least one embodiment, the semiconductor device 150 may report DATA3 as a response to the command CMD1. The DATA3 may be transmitted to the host 110 through the CXL switch 120.
[0100] After receiving DATA3, the host 110 may generate address information. The host 110 may send the address information to the CXL storage device to store the calculation result for executing the write command. The host 110 may send a command to the corresponding CXL storage devices 130_1 to 130_G through the CXL switch 120. The CXL storage device may obtain and store DATA3 from the semiconductor device 150 based on the write command and address information of the host 110. For example, when the CXL storage devices 130 form a RAID group and the second CXL storage device among the CXL storage devices 130 is responsible for parity check, the second CXL storage device may obtain and store the parity check generated from the semiconductor device 150 as the calculation result.
[0101] Figure 10 is a block diagram of a semiconductor device according to at least one embodiment.
[0102] Reference Figure 10 , according to at least one embodiment, the semiconductor device 300 may be a CXL memory device and may represent, for example, Figures 2 to 4 and / or Figures 8 to 9 the semiconductor device 150. According to at least one embodiment, the semiconductor device 300 may be a globally structured attached memory (GFAM) device defined in the CXL specification. The semiconductor device 300 may include an interface circuit (I / F) 310, a controller 320, a calculation engine 330, a memory controller (MC) 340, and a memory 350.
[0103] The interface circuit 310 can be configured to implement communication between the semiconductor device 300 and other components. The interface circuit 310 can be responsible for the PCIe physical layer (PHY) and the CXL protocol, and can implement communication with respect to the CXL switch. For example, the interface circuit 310 can receive data from the CXL switch and / or transmit data to the CXL switch.
[0104] The controller 320 can be configured to control the overall operation of the semiconductor device 300. The controller 320 can manage the mapping table. The mapping table can represent the mapping relationship between the host physical address and the device physical address. The controller 320 can convert the host physical address and the device physical address into each other by using the mapping table. When the semiconductor device 300 is connected to multiple hosts, the controller 320 can manage the mapping table for each host.
[0105] When receiving a data request from a CXL device (e.g., Figure 2 the CXL storage device 130) connected to the CXL switch, the controller 320 can obtain the device physical address from the mapping table. The data request can include the address of the memory 350 and commands such as write, read, etc.
[0106] The controller 320 can manage the address information INF_ADR of the host (see Figure 3 ). The controller 320 can compare the device physical address with the address information. The controller 320 can perform the comparison by using an index (or identifier). The index can be used to identify the address within the address information. For example, the controller 320 can assign "index 0" to the initial address of the address information, and can sequentially assign subsequent addresses by incrementing the index value by 1. That is, the controller 320 can assign "index 1" to the address after the initial address. In this way, the index being assigned can be referred to as the "current index", and the subsequent index (or indices) can be referred to as the "subsequent index" (or "multiple subsequent indices").
[0107] When the device physical address of the first data for input / output matches the initial address, the controller 320 can increment the index and perform the comparison. The controller 320 can check whether the address of "index 1" is activated. In this way, when inputting / outputting data according to the address information, the controller 320 can instruct the computing engine 330 to perform a calculation. The controller 320 can manage the index separately for each individual address information of the multiple CXL storage devices 130.
[0108] The computing engine 330 can perform calculations on the data output from the memory 350. The computing engine 330 can be based on the command of the host (e.g., Figure 3Execute calculations by means of CMD1). For example, the host can instruct the semiconductor device 300 to perform parity generation. Therefore, the computing engine 330 can generate a parity by performing an exclusive OR (XOR) calculation on the data output from the memory 350. For example, the computing engine 330 can include an XOR engine. The computing engine 330 can snoop on the data output from the memory 350 and can generate a parity by performing an XOR calculation on the data using the XOR engine.
[0109] The memory controller 340 can control the operation of the memory 350. For example, the memory controller 340 can control the input / output operation, refresh operation, etc. of the memory 350. The memory controller 340 can output the data of the memory 350 or record data in the memory 350 based on the instructions of the controller 320. The controller 320 can indicate to the memory controller 340 the address to be targeted for data operations based on the mapping table.
[0110] The memory controller 340 can be implemented in multiple numbers and control each memory 350. In at least one embodiment, the number of memory controllers 340 can be multiple and the same as the number of memories included in the memory 350, and there can be a one-to-one relationship between the memory controller 340 and the memory 350.
[0111] The memory 350 can be a volatile memory. The memory 350 can include multiple DRAMs and store data. However, the embodiment is not limited thereto.
[0112] In at least one embodiment, the semiconductor device 300 can further include a power supply, SRAM, etc. In at least one embodiment, the power supply can continuously supply power to the semiconductor device 300 so that the memory 350 can store data non-volatily. The power supply can be implemented as a capacitor module, etc. In at least one embodiment, the SRAM can store the calculation results of the computing engine 330. In at least one embodiment, the SRAM can store the address information of the host. In at least one embodiment, the SRAM can store the mapping table.
[0113] Figure 11 is a block diagram of a semiconductor device according to at least one embodiment.
[0114] Refer to Figure 11 According to at least one embodiment, the semiconductor device 400 can be a storage device and can represent, for example Figures 2 to 4 and / or Figures 8 to 9The semiconductor device 150. The semiconductor device 400 may include an interface circuit (I / F) 410, a controller 420, a computing engine 430, a DMA engine 440, a first memory controller (MC1) 450, a volatile memory (VM) 460, a second memory controller (MC2) 470, and a non-volatile memory (NVM) 480. The configuration of the semiconductor device 400 may be different from Figure 2 the configuration of the CXL storage device 130.
[0115] The interface circuit 410 may be configured to implement communication between the semiconductor device 400 and other components. The interface circuit 410 may be responsible for the PCIe physical layer (PHY) and the CXL protocol, and may implement communication with respect to the CXL switch. For example, the interface circuit 410 may receive data from the CXL switch and / or transmit data to the CXL switch.
[0116] The controller 420 may be configured to control the overall operation of the semiconductor device 400. The controller 420 may manage a mapping table. The mapping table may represent the mapping relationship between the host physical address and the device physical address. The controller 420 may convert the host physical address and the device physical address into each other by using the mapping table. When the semiconductor device 400 is connected to multiple hosts, the controller 420 may manage the mapping table for each host.
[0117] When receiving a data request from a CXL device (e.g., Figure 2 the CXL storage device 130) connected to the CXL switch, the controller 420 may obtain the device physical address from the mapping table. The data request may include the address of the volatile memory 460 and commands such as read, write, etc.
[0118] The controller 420 may manage the address information INF_ADR of the host (see Figure 3 ). The controller 420 may compare the device physical address with the address information. The controller 420 may perform the comparison by using an index (or identifier). The index may be used to identify the address within the address information. For example, the controller 420 may assign "index 0" to the initial address of the address information, and may sequentially assign subsequent addresses by incrementing the index value by 1. That is, the controller 420 may assign "index 1" to the address after the initial address.
[0119] When the device physical address of the first data for input / output matches the initial address, the controller 420 may increment the index and perform a comparison. The controller 420 may check whether the address of "Index 1" is activated. In this way, when inputting / outputting data according to the address information, the controller 420 may instruct the computing engine 430 to perform a calculation. The controller 420 may manage the index separately for each individual address information of the plurality of CXL storage devices 130.
[0120] The computing engine 430 may perform a calculation on the data output from the volatile memory 460. The computing engine 430 may perform a calculation based on a command from the host (e.g., Figure 3 CMD1). For example, the host may instruct the semiconductor device 400 to perform parity generation. Thus, the computing engine 430 may generate a parity by performing an exclusive OR (XOR) calculation on the data output from the volatile memory 460. For example, the computing engine 430 may include an XOR engine. The computing engine 430 may monitor the data output from the volatile memory 460 and may generate a parity by performing an XOR calculation on the data using the XOR engine.
[0121] In at least one embodiment, the semiconductor device 400 according to at least one embodiment may be connected to a DRAM-less storage device via a CXL switch. The controller 420 may monitor the data input / output between the DRAM-less storage device and the semiconductor device 400. The controller 420 may force the computing engine 430 to analyze the input / output data in order to calculate the lifespan of the memory (e.g., NAND flash) of the DRAM-less storage device.
[0122] In at least one embodiment, the computing engine 430 may perform data processing (e.g., encryption, decryption, compression, decompression, deduplication, etc.). In at least one embodiment, the computing engine 430 may perform data calculations (e.g., pattern search, data integrity check, filtering, SQL, statistics, etc.).
[0123] The direct memory access (DMA) engine 440 may be configured to acquire data to transfer it to the non-volatile memory 480, and / or may transfer the data of the non-volatile memory 480 to the outside. For example, the DMA engine 440 may acquire data from the host or a CXL device via a CXL switch, and / or may transfer the data to the host or a CXL device.
[0124] The first memory controller 450 may be configured to control the operation of the volatile memory 460. For example, the first memory controller 450 may control input / output operations, refresh operations, etc. of the volatile memory 460. The first memory controller 450 may output data of the volatile memory 460 or record data in the volatile memory 460 based on instructions from the controller 420. The controller 420 may indicate to the first memory controller 450, based on a mapping table, the address to be targeted for data operations.
[0125] Multiple first memory controllers 450 may be implemented and each volatile memory 460 may be controlled. In at least one embodiment, the number of first memory controllers 450 may be multiple and the same as the number of memories included in the volatile memory 460, and there may be a one-to-one relationship between the first memory controller 450 and the volatile memory 460.
[0126] The volatile memory 460 may be a device memory managed by a host (e.g., high density (HDM) memory). The volatile memory 460 may include multiple DRAMs and store data. The volatile memory 460 may operate as a buffer memory for the non-volatile memory 480. However, this embodiment is not limited thereto.
[0127] The second memory controller 470 may control the operation of the non-volatile memory 480. For example, the second memory controller 470 may control input / output operations, wear leveling operations, garbage collection operations, flash translation layer (FTL), etc. of the non-volatile memory 480.
[0128] The non-volatile memory 480 may store data non-temporarily and may be implemented as flash memory, PRAM, RRAM, MRAM, FRAM, etc.
[0129] In at least one embodiment, the semiconductor device 400 may further include a power supply, SRAM, etc. In at least one embodiment, the power supply may continuously supply power to the semiconductor device 400 such that the volatile memory 460 can store data non-volatily. The power supply may be implemented as a capacitor module, etc. In at least one embodiment, the SRAM may store the calculation results of the computing engine 430. In at least one embodiment, the SRAM may store the address information of the host. In at least one embodiment, the SRAM may store the mapping table.
[0130] Figure 12 It is a diagram for explaining monitoring performed by a controller based on address information according to at least one embodiment.
[0131] Reference Figure 12, according to at least one embodiment, a host can generate address information INF_ADR for each CXL storage device included in a RAID set. The address information INF_ADR can be represented as a PRP or an SGL. For example, the first to the Gth CXL storage devices can be included in the RAID set, and the host can assign RAID IDs from 1 to G to the first to the Gth CXL storage devices. The host can generate first address information INF_ADR1 and send it to the first CXL storage device. In the same way, the host can generate second address information INF_ADR2 to Gth address information INF_ADRG and can send the corresponding information to the second CXL storage device to the Gth CXL storage device. The first address information INF_ADR1 can indicate addresses ADR1_1 to ADR1_A (where A is an integer greater than 3), the second address information INF_ADR2 can indicate addresses ADR2_1 to ADR2_B (where B is an integer greater than 3), and the Gth address information INF_ADRG can indicate addresses ADRG_1 to ADRG_C (where C is an integer greater than 3). The first CXL storage device to the Gth CXL storage device can perform data communication with a semiconductor device based on a command from the host and the first address information INF_ADR1 to the Gth address information INF_ADRG.
[0132] In addition, the host can send the address information INF_ADR including the first address information INF_ADR1 to the Gth address information INF_ADRG to the semiconductor device. The semiconductor device can receive the address information INF_ADR from the host. The semiconductor device can monitor whether data is input / output based on the address information INF_ADR. When data is input / output based on the address information INF_ADR, the semiconductor device can perform data calculation.
[0133] The semiconductor device can monitor the input / output data by using indexes INDEX_1 to INDEX_G. The semiconductor device can manage the indexes INDEX_1 to INDEX_G for each storage device. The indexes INDEX_1 to INDEX_G can represent the addresses of the data to be input / output through the storage devices among the first address information INF_ADR1 to the Gth address information INF_ADRG, and when data is input or output at the corresponding addresses, the semiconductor device can update the values of the indexes INDEX_1 to INDEX_G.
[0134] For example, when acquiring data of address ADR1_2 (whose index INDEX_1 is 1) in the first address information INF_ADR1, the semiconductor device may update the index INDEX_1 to 2. Thereafter, the semiconductor device may check whether to acquire data of address ADR1_3 with index INDEX_1 being 2. In the same manner, when acquiring data of address ADR2_1 (whose index INDEX_2 is 0) in the second address information INF_ADR2, the semiconductor device may update the index INDEX_2 to 1. The same description may equally apply to the Gth address information INF_ADRG. Accordingly, the semiconductor device may manage which data ranges have been acquired and which data ranges have been calculated by using the address information INF_ADR1 to INF_ADRG and the indexes INDEX_1 to INDEX_G.
[0135] When inputting / outputting the address information INF_ADR sequentially, the semiconductor device may perform data calculation. For example, the semiconductor device may receive a command indicating parity generation from the host. The semiconductor device may receive first data of address ADR1_1 and receive second data of address ADR2_1. The semiconductor device may generate first result data by performing an XOR calculation on the first data and the second data. Thereafter, the semiconductor device may receive third data of address ADR1_2. The semiconductor device may generate second result data by performing an XOR calculation on the third data and the first result data. In this way, when acquiring data of a predetermined size, the semiconductor device may store the final result data generated by the XOR calculation as parity. The host may preset the size of the data for which parity is to be generated.
[0136] By allowing the semiconductor device (e.g., semiconductor devices 150, 300, and / or 400) to perform operations on behalf of the host or the accelerator, data bottlenecks can be alleviated, interface bus utilization can be improved, and power consumption can be reduced.
[0137] Figure 13 and Figure 14 are block diagrams for explaining the operations of a computing system according to at least one embodiment.
[0138] Refer to Figure 13 and, refer to Figure 2 The description may equally apply to the computing system 100 according to at least one embodiment. Therefore, redundant descriptions will be omitted.
[0139] The host 110 may generate address information INF_ADR. The address information INF_ADR may include the address where the semiconductor device 150 will store data. The address information INF_ADR may be transmitted to the plurality of CXL storage devices 130 together with a command (e.g., read). In some embodiments, the address information INF_ADR may include PRP and / or SGL of the NVMe standard.
[0140] The host 110 may generate address information INF_ADR for each of the plurality of CXL storage devices 130. For example, the host 110 may generate first address information to G-th address information corresponding to the first CXL storage device 130_1 to the G-th CXL storage device 130_G, respectively. At this time, the address information INF_ADR may include the first address information to the G-th address information.
[0141] In at least one embodiment, the plurality of CXL storage devices 130 may form a RAID set. According to an embodiment, some of the plurality of CXL storage devices 130 may form a RAID group. That is, each storage device forming a RAID group among the plurality of CXL storage devices 130 may store a divided data block. At this time, the logical block address (LBA) where each storage device stores the data block may be the same. The host 110 may generate address information INF_ADR for data reconstruction of the storage devices forming the RAID set. The plurality of CXL storage devices 130 may transmit the data block to the semiconductor device 150 based on the address information INF_ADR, and the semiconductor device 150 may recover the erroneous data block.
[0142] The host 110 may transmit a command CMD1 to the semiconductor device 150. The host 110 may transmit the command CMD1 through the CXL switch 120, and the command CMD1 may indicate data calculation (e.g., parity generation, data reconstruction, encryption, decryption, compression, decompression, pattern search, data integrity check, SQL, statistics, etc.). In at least one embodiment, the command CMD1 may indicate data reconstruction. Additionally, the command CMD1 may include the address information INF_ADR. In at least one embodiment, the command CMD1 may be a mailbox command of the CXL standard. The semiconductor device 150 may store the address information INF_ADR and may perform calculations based on the command CMD1.
[0143] In at least one embodiment, the host 110 may generate address information of the memory 111, and the address information may be transmitted to the plurality of CXL storage devices 130 together with commands such as read, write, etc.
[0144] The host 110 may transmit commands CMD2 and CMD3 to a plurality of CXL storage devices 130. The controllers 132_1 to 132_G may control the operations of the plurality of CXL storage devices 130 based on commands CMD2 and CMD3. In at least one embodiment, commands CMD2 and CMD3 may be read commands, and the controllers 132_1 to 132_G may read data from the non-volatile memories 131_1 to 131_G.
[0145] In at least one embodiment, commands CMD2 and CMD3 may follow the command format of the NVMe standard. For example, the command format may include "Command Identifier (CID)", "Namespace Identifier (NSID)", "Reserved Field", "Metadata", "PRP1", "PRP2", "Starting LBA (SLBA)", "Length", etc.
[0146] "PRP1" and "PRP2" may indicate the location to which the data is to be transferred. "PRP1" may indicate the first physical region page. "PRP2" may indicate the second physical region page. The first physical region page and the second physical region page may indicate the address of the memory (e.g., memory 111 or semiconductor device 150) for DMA communication.
[0147] In at least one embodiment, when two or more physical region pages are applied to indicate data, the PRP list may include pointers indicating the physical region pages. For example, "PRP1" may indicate the first PRP list, "PRP2" may indicate the second PRP list, and the last one of the second PRP list may indicate the third PRP list.
[0148] In at least one embodiment, when two or more Scattered Gather List (SGL) segments are applied to indicate data, the SGL segments may include pointers to subsequent SGL segments.
[0149] The host 110 may transmit the address information INF_ADR together with the commands CMD2 and CMD3. The host 110 may transmit the commands CMD2 and CMD3 and the address information INF_ADR through the CXL switch 120. For example, the host 110 may transmit the command CMD2 and the first address information to the first CXL storage device 130_1, and may transmit the command CMD3 and the Gth address information to the Gth CXL storage device 130_G. The PRP (or SGL) of the command CMD2 may indicate the first address information, and the PRP (or SGL) of the command CMD3 may indicate the Gth address information. That is, the multiple storage devices 130 may access the semiconductor device 150 based on the commands CMD2 and CMD3 and the address information INF_ADR. When the commands CMD2 and CMD3 are read commands, the multiple storage devices 130 may transmit the data of the specified LBA of the non-volatile memories 131_1 to 131_G to the semiconductor device 150.
[0150] Reference Figure 14 , the multiple CXL storage devices 130 may transmit the data DATA1 and DATA2 to the semiconductor device 150 based on the commands CMD2 and CMD3 and the address information INF_ADR. The multiple CXL storage devices 130 may transmit the data DATA1 and DATA2 in the DMA format. For example, the commands CMD2 and CMD3 may be read commands, and the first CXL storage device 130_1 may transmit the data DATA1 by sending a read command and an address corresponding to the first address information to the semiconductor device 150 based on the command CMD2 and the first address information. In the same way, the Gth storage device 130_G may transmit the data DATA2 by sending a write command and an address corresponding to the Gth address information to the semiconductor device 150 based on the command CMD3 and the Gth address information. The semiconductor device 150 may record the data DATA1 and DATA2 based on the write command and the address.
[0151] In at least one embodiment, the multiple CXL storage devices 130 may include an address engine and a DMA engine, and may transmit data by using the address engine and the DMA engine. The address engine may set the source address and the destination address of the DMA engine based on the address information. The address engine may parse the entries of the address information based on the command of the host 110, and may detect the locations referred to by each entry in the semiconductor device 150. The DMA engine may retrieve data from the semiconductor device 150 based on the source address and the destination address, and / or transmit data to the semiconductor device 150. Reference Figure 8 's description may equally apply to the address engine and the DMA engine. Therefore, redundant descriptions will be omitted.
[0152] The semiconductor device 150 can monitor the transmitted data DATA1 and DATA2. For example, the semiconductor device 150 can monitor the data DATA1 and DATA2 based on the address information INF_ADR received from the host 110 (see Figure 13 ). The semiconductor device 150 can determine whether the data DATA1 and DATA2 are received according to the address information INF_ADR. That is, the semiconductor device 150 can determine whether the data DATA1 and DATA2 are received in the order indicated by the address information INF_ADR.
[0153] The semiconductor device 150 can monitor data input / output based on the device physical address obtained by using a mapping table. The semiconductor device 150 can determine whether the device physical address is sequentially activated in the address information INF_ADR.
[0154] When the data DATA1 and DATA2 are received according to the address information INF_ADR, the semiconductor device 150 can perform calculations on the data DATA1 and DATA2. The semiconductor device 150 can perform calculations on the data DATA1 and DATA2 according to the command CMD1 (see Figure 13 ). For example, the semiconductor device 150 can perform calculations such as generating recovery data from the data DATA1 and DATA2. The semiconductor device 150 can store the calculation result in an internal memory. According to an embodiment, the internal memory can be a DRAM or an SRAM. The semiconductor device 150 can report the calculation result to the host 110.
[0155] The host 110 can generate new address information based on the calculation result. The host 110 can transmit a write command and the new address information to the CXL storage device. The CXL storage device can retrieve and store the recovery data of the semiconductor device 150 based on the write command and the new address information.
[0156] Figure 15 is a flowchart of a data calculation method according to at least one embodiment.
[0157] Refer to Figure 15 , the data calculation method according to at least one embodiment can be performed by a semiconductor device. The semiconductor device can be connected to at least one host and at least one CXL device through a CXL switch. In some embodiments, the semiconductor device can be implemented as a memory device, a storage device, etc.
[0158] A semiconductor device can receive command and address information from a host (S1010). The command can indicate data calculations (e.g., parity generation, data reconstruction, encryption, decryption, compression, decompression, pattern search, data integrity check, SQL, statistics, etc.). The address information can correspond to the command. For example, the address information can indicate the address of the data that is the target of the command. In at least one embodiment, after receiving the command and address information, the semiconductor device can perform data calculations based on the command and address information.
[0159] The host can generate address information for each CXL device. For example, the host can generate multiple address information corresponding to multiple CXL devices and can send the multiple address information to the semiconductor device.
[0160] The semiconductor device can monitor input / output data (S1020). For example, at least one CXL device can obtain data from the semiconductor device based on a command from the host and / or record data in the semiconductor device. The semiconductor device can manage the input / output of data by using a mapping table. The mapping table can represent the mapping relationship between the host physical address and the device physical address. The semiconductor device can convert the host physical address to the device physical address by using the mapping table and can record or output data by using the device physical address.
[0161] The semiconductor device can determine whether the device physical addresses of the address information and the data match each other (S1030). In at least one embodiment, when the host indicates parity generation, the semiconductor device can determine whether to output the data corresponding to the address information. In at least one embodiment, when the host indicates data reconstruction, the semiconductor device can determine whether to input (or record) data at the position of the address information.
[0162] The semiconductor device can manage the address information by using an index. For example, the semiconductor device can increment the index of the address that matches the device physical address of the currently input / output data. Accordingly, the semiconductor device can know which data range has been input / output (or calculated) in which address information.
[0163] When the device physical addresses of the address information and the data match, the semiconductor device can perform data calculations according to the command (S1040). In at least one embodiment, when the command is parity generation, the semiconductor device can perform an XOR calculation when outputting the data corresponding to the address information. In at least one embodiment, when the command is data reconstruction, the semiconductor device can perform an XOR calculation when outputting the data corresponding to the address information.
[0164] When the address information does not match the address of the data, the semiconductor device can monitor the input / output data and return to operation S1020.
[0165] Figure 16 is a block diagram of a computer system according to at least one embodiment. Hereinafter, for ease of description, redundant descriptions of redundant components are not included here.
[0166] Reference Figure 16 , the computing system 500 may include a host 510, a memory 511, a CXL switch 520, a plurality of CXL storage devices 530_1 to 530_P (where P is an integer greater than 1), and a plurality of CXL memory devices 550_1 to 550_Q (where Q is an integer greater than 1). Although for better understanding and ease of description, Figure 16 the computing system 500 is shown to include one host 510, the embodiments are not limited thereto, and the computing system 500 may include one or more hosts.
[0167] The host 510 may be directly connected to the memory 511. The host 510, the plurality of CXL storage devices 530_1 to 530_P, and the plurality of CXL memory devices 550_1 to 550_Q may be connected to the CXL switch 520 and may communicate with each other through the CXL switch 520, respectively.
[0168] In at least one embodiment, the host 510 may manage the plurality of CXL storage devices 530_1 to 530_P as one storage cluster and may manage the plurality of CXL memory devices 550_1 to 550_Q as one memory cluster. For one storage cluster, the host 510 may allocate a partial area of the memory cluster as a dedicated area (for example, an area for storing mapping data of the storage cluster). Alternatively, for the plurality of CXL storage devices 530_1 to 530_P, the host 510 may allocate each area of the plurality of CXL memory devices 550_1 to 550_Q as a dedicated area.
[0169] The host 510 may manage the memory cluster by using interleaving techniques. The plurality of CXL memory devices 550_1 to 550_Q may have separate interleaving granularities. The plurality of CXL memory devices 550_1 to 550_Q may perform data calculations based on the sizes of the separate interleaving granularities. That is, the CXL memory devices 550_1 to 550_Q may perform data calculations separately. The host 510 may receive the calculation results of the plurality of CXL memory devices 550_1 to 550_Q and determine the final calculation result.
[0170] In at least one embodiment, the plurality of CXL storage devices 530_1 to 530_P may have the same as the referenceFigures 1 to 15 A structure similar to the described CXL device (or CXL storage device). That is, multiple CXL storage devices 530_1 to 530_P can form a RAID set. Multiple CXL storage devices 530_1 to 530_P can receive command and address information from the host 510. The host 510 can generate and transmit separate address information for each of the multiple CXL storage devices 530_1 to 530_P. Multiple CXL storage devices 530_1 to 530_P can access the memory cluster based on the command and address information and can perform data operations.
[0171] Each of the CXL memory devices 550_1 to 550_Q included in the memory cluster can have a structure similar to the semiconductor device described in the reference Figures 1 to 15 That is, multiple CXL memory devices 550_1 to 550_Q can receive commands and the address information of multiple CXL storage devices 530_1 to 530_P from the host 510. Multiple CXL memory devices 550_1 to 550_Q can determine whether the device physical address of the input / output data matches the address information. At this time, multiple CXL memory devices 550_1 to 550_Q can manage which data range has been input / output (or calculated) by using the index in the address information. When the device physical address of the input / output data matches the address information, multiple CXL memory devices 550_1 to 550_Q can perform data calculations based on the commands of the host 510.
[0172] Figure 17 is a block diagram of a computer system according to at least one embodiment. For ease of description, detailed descriptions of previously described components are not included here.
[0173] Reference Figure 17 , a computer system 1000 according to at least one embodiment may include a first CPU 1310a, a second CPU 1310b, a GPU 1330, an NPU 1340, a CXL switch SW_CXL, a CXL memory 1350, multiple CXL storage devices 1352_1 to 1352_z (where z is an integer greater than 1), a PCIe device 1354, and an accelerator (CXL device) 1356. Each of the first CPU 1310a, the second CPU 1310b, the GPU 1330, and the NPU 1340 can be directly connected to separate memories 1320a, 1320b, 1320c, 1320d, and 1320e.
[0174] The first CPU 1310a, the second CPU 1310b, the GPU 1330, the NPU 1340, the CXL memory 1350, the multiple CXL storage devices 1352_1 to 1352_z, the PCIe device 1354, and the accelerator (CXL device) 1356 can be commonly connected to the CXL switch SW_CXL and can communicate with each other through the CXL switch SW_CXL respectively.
[0175] In at least one embodiment, the multiple CXL storage devices 1352_1 to 1352_z can have a structure similar to the CXL device (or CXL storage device) described with reference to Figures 1 to 16 That is, the multiple CXL storage devices 1352_1 to 1352_z can form a RAID set. The multiple CXL storage devices 1352_1 to 1352_z can receive command and address information from a host (e.g., the first CPU 1310a). The host can generate and transmit separate address information for each of the multiple CXL storage devices 1352_1 to 1352_z. The multiple CXL storage devices 1352_1 to 1352_z can access the CXL memory 1350 based on the command and address information and can perform data operations.
[0176] The CXL memory 1350 can have a structure similar to the semiconductor device described with reference to Figures 1 to 16 That is, the CXL memory 1350 can receive a command and the address information of the multiple CXL storage devices 1352_1 to 1352_z from the host. The CXL memory 1350 can determine whether the device physical address of the input / output data matches the address information. At this time, the CXL memory 1350 can manage which data range has been input / output (or calculated) by using the index in the address information. When the device physical address of the input / output data matches the address information, the CXL memory 1350 can perform data calculation based on the host's command.
[0177] Although Figure 17 it shows the first CPU 1310a communicating with the multiple CXL storage devices 1352_1 to 1352_z, the embodiment is not limited thereto, and at least one of the second CPU 1310b, the GPU 1330, and the NPU 1340 can be implemented as a host to communicate with the multiple CXL storage devices 1352_1 to 1352_z.
[0178] At least a portion of the memory (1362_1 - 1362_z) of the multiple CXL storage devices 1352_1 to 1352_z can be allocated to at least one cache buffer of the first CPU 1310a, the second CPU 1310b, the GPU 1330, the NPU 1340, the CXL memory 1350, the multiple CXL storage devices 1352_1 to 1352_z, the PCIe device 1354, and the accelerator 1356, through one or more of the first CPU 1310a, the second CPU 1310b, the GPU 1330, and the NPU 1340.
[0179] In at least one embodiment, at least a portion of the memory 1360 of the CXL memory 1350 can be allocated to at least one cache buffer of the first CPU 1310a, the second CPU 1310b, the GPU 1330, the NPU 1340, the CXL memory 1350, the multiple CXL storage devices 1352_1 to 1352_z, the PCIe device 1354, and the accelerator 1356, through one or more of the first CPU 1310a, the second CPU 1310b, the GPU 1330, and the NPU 1340. That is, the CXL memory 1350 and the multiple CXL storage devices 1352_1 to 1352_z can serve as the storage space STR of the computer system 1000.
[0180] In at least one embodiment, the CXL switch SW_CXL can be connected to the PCIe device 1354 or the accelerator 1356 configured to support various functions, and the PCIe device 1354 or the accelerator 1356 can communicate with each of the first CPU 1310a, the second CPU 1310b, the GPU 1330, and the NPU 1340 through the CXL switch SW_CXL, and / or alternatively, can access the storage space STR including the multiple CXL storage devices 1352_1 to 1352_z and the CXL memory 1350.
[0181] In at least one embodiment, the CXL switch SW_CXL can be connected to an external network or fabric, and can be configured to communicate with an external server through the external network or fabric.
[0182] Figure 18 It is a block diagram of a data center of a computer system according to at least one embodiment. For ease of description, detailed descriptions of previously described components are not included here.
[0183] Reference Figure 18, the data center 1400, as a facility for collecting various data and providing services, can also be referred to as a data storage center. The data center 1400 can be a system for operating a search engine and a database, and can be a computer system used in an enterprise or a government agency such as a bank. The data center 1400 can include application servers 1410a to 1410h and storage servers 1420a to 1420h. The number of application servers and the number of storage servers can be differently selected according to at least one embodiment and can be different from each other.
[0184] In the following, the configuration of the first storage server 1420a will be mainly described. Each of the application servers 1410a to 1410h and the storage servers 1420a to 1420h can have a similar structure, and the application servers 1410a to 1410h and the storage servers 1420a to 1420h can communicate with each other through the network NT.
[0185] The first storage server 1420a can include a processor 1421, a memory 1422, a switch 1423, a CXL memory 1424, a storage device 1425, and a network interface card (NIC) 1426. The processor 1421 can control the overall operation of the first storage server 1420a, can access the memory 1422, can execute commands loaded into the memory 1422, and / or can process data. The memory 1422 can be a double data rate synchronous DRAM (DDR SDRAM), a hybrid memory cube (HMC), a DIMM, an HBM, an Optane DIMM, and / or a non-volatile DIMM (NVDIMM). The processor 1421 and the memory 1422 can be directly connected, and various selections can be made for the number of processors 1421 and the number of memories 1422 included in one storage server 1420a.
[0186] In at least one embodiment, the processor 1421 and the memory 1422 can provide a processor-memory pair. In at least one embodiment, the number of processors 1421 and the number of memories 1422 can be different. The processor 1421 can include a single-core processor or a multi-core processor. The above description of the storage server 1420a can be similarly applied to each of the application servers 1410a to 1410h.
[0187] The switch 1423 can be configured to mediate or route communications between the components included in the first storage server 1420a. In at least one embodiment, the switch 1423 can be the interface or CXL switch described with reference to Figures 1 to 16 The switch 1423 can be a switch implemented based on the CXL protocol.
[0188] The CXL memory 1424 can be connected to the switch 1423. The CXL memory 1424 can have a structure similar to the semiconductor device described with reference to Figures 1 to 16 That is, the CXL memory 1424 can receive a command and address information of the storage device 1425 from the host. The CXL memory 1424 can determine whether the device physical address of the input / output data matches the address information. At this time, the CXL memory 1424 can manage which data range has been input / output (or calculated) by using the index in the address information. When the device physical address of the input / output data matches the address information, the CXL memory 1424 can perform data calculation based on the command of the host.
[0189] In at least one embodiment, the CXL memory 1424 can be used as a memory expander for the processor 1421. In at least one embodiment, the CXL memory 1424 can be allocated as a dedicated memory or buffer memory for the storage device 1425.
[0190] The storage device 1425 can include a CXL interface circuit CXL_IF, a controller CTRL, and a NAND flash (NAND). The storage device 1425 can store data or output stored data according to the request of the processor 1421.
[0191] In at least one embodiment, the storage device 1425 can be the storage device described with reference to Figures 1 to 16 described. In at least one embodiment, similar to that described with reference to Figures 1 to 16 described, the storage device 1425 in the storage server 1420a can be implemented as multiple (e.g., as one of multiple), and form a RAID set. The number of storage devices 1425 included in the data center 1400 can be selected in various ways according to the embodiment.
[0192] The storage device 1425 can receive a command and address information from the processor 1421. The processor 1421 can generate and transmit separate address information for each storage device 1425. The storage device 1425 can access the CXL memory 1424 based on the command and address information, and can perform data operations.
[0193] The network interface card (NIC) 1426 can be connected to the switch 1423. The NIC 1426 can communicate with other storage servers 1420b to 1420h or other application servers 1410a to 1410h through the network NT.
[0194] In at least one embodiment, NIC 1426 may include a network interface card, a network adapter, etc. NIC 1426 may be connected to network NT through a wired interface, a wireless interface, a Bluetooth interface, an optical fiber interface, etc. NIC 1426 may include an internal memory, a digital signal processor (DSP), a host bus interface, etc., and may be connected to processor 1421 and / or switch 1423 through the host bus interface. In at least one embodiment, NIC 1426 may be integrated with at least one of processor 1421, switch 1423, and / or storage device 1425.
[0195] In at least one embodiment, network NT may be implemented using Fibre Channel (FC), Ethernet, etc. In this case, FC, as a medium for relatively high-rate data transmission, may use an optical switch that provides high performance and high availability. A storage server may be provided according to the access method of network NT, such as file storage, block storage, and / or object storage.
[0196] In at least one embodiment, network NT may be a storage-only network (e.g., a storage area network (SAN)). For example, a SAN may be an FC-SAN implemented using an FC network and according to the FC protocol (FCP). As another example, a SAN may be an IP-SAN implemented using a TCP / IP network and according to the iSCSI (SCSI over TCP / IP or Internet SCSI) protocol. In at least one embodiment, network NT may be a general network such as a TCP / IP network. For example, network NT may be implemented according to protocols such as Ethernet Fibre Channel (FCoE), network-attached storage (NAS), and NVMe fabric (NVMe-oF).
[0197] In at least one embodiment, at least one of application servers 1410a to 1410h may store, through network NT, data requested by a user or a client in one of storage servers 1420a to 1420h. At least one of application servers 1410a to 1410h may obtain, through network NT, data requested by a user or a client to be read from one of storage servers 1420a to 1420h. For example, at least one of application servers 1410a to 1410h may be implemented as a web server or a database management system (DBMS).
[0198] In at least one embodiment, at least one of application servers 1410a to 1410h may access a memory, a CXL memory, and / or a storage device included in another application server via network NT, and / or may access a memory, a CXL memory, and / or a storage device included in storage servers 1420a to 1420h via network NT. Accordingly, at least one of application servers 1410a to 1410h may perform various operations on data stored in other application servers and / or storage servers. For example, at least one of application servers 1410a to 1410h may execute a command to move or copy data between other application servers and / or storage servers. In this case, the data may be directly moved to the memory or CXL memory of the application server, or may be moved from the storage device of the storage server to the memory or CXL memory of the application server through the memory or CXL memory of the storage server. Data moved through the network may be encrypted to ensure security or privacy.
[0199] In at least one embodiment, a storage device included in at least one of application servers 1410a to 1410h and storage servers 1420a to 1420h may be allocated a CXL memory included in at least one of application servers 1410a to 1410h and storage servers 1420a to 1420h as a dedicated area, and the storage device may use the allocated dedicated area as a buffer memory (e.g., may store mapping data). For example, a storage device 1425 included in storage server 1420a may be allocated a CXL memory included in another storage server (e.g., 1420h), and may access the CXL memory included in another storage server (e.g., 1420h) through switch 1423 and NIC 1426. In this case, mapping data for the storage device 1425 of the first storage server 1420a may be stored in the CXL memory of another storage server 1420h. That is, the storage devices and CXL memories of data center 1400 according to the present disclosure may be connected and implemented in various ways.
[0200] In some embodiments, each component or a combination of two or more components referred to Figures 1 to 18 to may be implemented as a digital circuit, a programmable or non-programmable logic device or array, an application specific integrated circuit (ASIC), etc.
[0201] Although the present disclosure has been described in connection with what are presently considered to be practical embodiments, it is to be understood that the invention is not limited to the disclosed embodiments, but on the contrary, is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A semiconductor device, comprising: A memory configured to store data; A controller configured to: Monitor input / output data of the memory, Determine a calculation to be performed on first data among the input / output data in response to receiving a calculation command and address information from a host, wherein the address information includes an instruction for an address of the memory using at least one of a physical region page PRP or a scatter-gather list SGL, and the calculation is determined based on the calculation command and the address information; And A calculation engine configured to perform the determined calculation on the first data.
2. The semiconductor device according to claim 1, wherein, The controller is further configured to determine whether a device physical address of the input / output data matches the address information, and when a first device physical address of the first data matches the address information, transmit a calculation instruction corresponding to the calculation command to the calculation engine.
3. The semiconductor device according to claim 2, wherein, The controller is configured to: Manage the address information using an index, and Determine whether an address of the current index in the device physical address and the address information matches each other.
4. The semiconductor device according to claim 3, wherein, The controller is configured to increase a value of the current index when the address of the current index is determined to match the device physical address.
5. The semiconductor device according to claim 2, wherein, The controller is configured to listen to the input / output data to perform a calculation based on the calculation command.
6. The semiconductor device according to claim 1, wherein, The controller is configured to, in response to receiving a host physical address from the host, convert the host physical address to a device physical address using a mapping table and monitor the input / output data based on the device physical address.
7. The semiconductor device according to claim 6, further comprising: A static random access memory SRAM storing the mapping table and the address information.
8. The semiconductor device according to claim 1, wherein, The controller is configured to receive multiple address information from the host and manage the multiple address information using separate indexes.
9. The semiconductor device according to claim 1, wherein, The calculation engine is further configured to: when the calculation command is one of parity generation or data reconstruction, perform an exclusive OR XOR calculation using the first data.
10. The semiconductor device according to claim 1, wherein, The calculation engine is further configured to transmit a calculation result of the calculation to the host.
11. A computing system, comprising: A host configured to generate and output a data command, a calculation command, and address information; Multiple Compute Express Link CXL storage devices configured to perform data operations based on the data command and the address information; And A semiconductor device configured to perform data calculations based on the calculation command and the address information; Wherein, the host is configured to use at least one of a physical region page PRP or a scatter-gather list SGL to generate the address information indicating a memory address of the semiconductor device.
12. The computing system according to claim 11, wherein, The host is configured to: Transmit first address information to a first CXL storage device among the multiple CXL storage devices, and transmit second address information to a second CXL storage device; And Transmit the first address information and the second address information to the semiconductor device.
13. The computing system according to claim 12, wherein: The first CXL storage device is configured to access the semiconductor device based on a first data command and the first address information; The second CXL storage device is configured to access the semiconductor device based on a second data command and the second address information; and The semiconductor device is configured to monitor input / output data based on the first address information and the second address information.
14. The computing system according to claim 13, wherein, The semiconductor device is configured to: manage the first address information using a first index, manage the second address information using a second index, and listen to the input / output data to perform the data calculation when the address of the input / output data is determined to match the address of the first index or the second index.
15. The computing system according to claim 14, wherein, The semiconductor device is configured to: when the address of the input / output data is determined to match the address of the first index or the second index, increase the corresponding value of the first index or the second index.
16. The computing system according to claim 11, wherein: The semiconductor device includes at least one of: a memory device including a volatile memory, or a storage device including a non-volatile memory; and The host, the plurality of CXL storage devices, and the semiconductor device are connected through a CXL switch.
17. The computing system according to claim 11, wherein: The plurality of CXL storage devices form a redundant array of inexpensive disks (RAID) group; and The semiconductor device is configured to perform at least one of parity generation or data reconstruction based on the computing command.
18. The computing system according to claim 11, wherein, The semiconductor device includes a plurality of CXL memory devices configured to perform the data calculation based on separate interleaving granularities.
19. The computing system according to claim 11, wherein: The semiconductor device is configured to transfer the result of the data calculation to the host; and The host is configured to generate new address information based on the result of the data calculation and transfer the new address information and a write command to one of the plurality of CXL storage devices.
20. A data calculation method, comprising: receiving, from a host, a command and address information of at least one of a physical region page (PRP) or a scatter-gather list (SGL); monitoring input / output data; determining whether the address information matches the device physical address of the first data among the input / output data; and when the device physical address matches the address information, performing a calculation on the first data according to the command.
Citation Information
Patent Citations
Process for producing high-grade tricaprylin
KR1020240011774A