Memory expander, heterogeneous computing device, and method of operating a heterogeneous computing device
Patent Information
- Application Number
- CN202111239813.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-29
- Filing Date
- 2021-10-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-10-25
Smart Images

Figure CN114428587B_ABST
Abstract
Description
[0001] This application claims priority to Korean Patent Application No. 10-2020-0141710, filed on October 29, 2020, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference. Technical Field
[0002] The embodiments of this disclosure described herein relate to computing systems, and more specifically, to memory expanders, heterogeneous computing devices using the memory expanders, and methods of operating the heterogeneous computing devices. Background Technology
[0003] Computing systems provide various information technology (IT) services to users. As these IT services are offered, the amount of data processed by computing systems increases. For this reason, there is a need to improve the speed of data processing. Computing systems are evolving into heterogeneous computing environments that provide various IT services. Currently, various technologies are being developed for processing data at high speeds within heterogeneous computing environments. Summary of the Invention
[0004] Embodiments of this disclosure provide a memory expander with improved performance, a heterogeneous computing device using the memory expander, and a method of operating the heterogeneous computing device.
[0005] According to one embodiment, the memory expander includes: a memory device for storing a plurality of task data; and a controller for controlling the memory device. The controller receives metadata and management requests from an external central processing unit (CPU) via a compute fast link (CXL) interface, and operates in management mode in response to the management requests. In management mode, the controller receives a read request and a first address from an accelerator via the CXL interface, and in response to the read request, sends one of the plurality of task data to the accelerator based on the metadata.
[0006] According to one embodiment, a heterogeneous computing device includes: a central processing unit (CPU); a memory for storing data under the control of the CPU; an accelerator for repeatedly performing computations on multiple task data and generating multiple result data; and a memory expander that operates in a management mode in response to management requests from the CPU, and manages the multiple task data to be provided to the accelerator and the multiple result data provided from the accelerator in the management mode. The CPU, accelerator, and memory expander communicate with each other through a heterogeneous computing interface.
[0007] According to one embodiment, a method of operating a heterogeneous computing device including a central processing unit (CPU), an accelerator, and a memory expander connected via a compute fast link (CXL) interface includes: the CPU sending metadata to the memory expander; the CPU sending a management request to the memory expander; the CPU sending a task request to the accelerator; the accelerator sending a read request to the memory expander in response to the task request; the memory expander sending first task data from a plurality of task data to the accelerator based on metadata in response to the read request; the accelerator performing a first computation on the first task data to generate first result data; the accelerator sending a write request and the first result data to the memory expander; and the memory expander storing the first result data in response to the write request. Attached Figure Description
[0008] The above and other objects and features of this disclosure will become clear from the detailed description of embodiments thereof with reference to the accompanying drawings.
[0009] Figure 1 This is a diagram illustrating a computing device according to an embodiment of the present disclosure.
[0010] Figure 2 It is shown Figure 1 A flowchart of the operation of the computing device.
[0011] Figure 3 It is shown Figure 1 A block diagram of the configuration of the memory expander.
[0012] Figure 4 It is shown by Figure 3 A diagram illustrating the metadata managed by the controller's metadata manager.
[0013] Figure 5A It is shown Figure 1 A flowchart of the operation of a computing device.
[0014] Figure 5B It is shown Figure 5A A diagram illustrating the header information of the management request in operation S103.
[0015] Figure 6 It is shown Figure 1 The flowchart shows the operation of the memory expander.
[0016] Figure 7 It is shown by Figure 1 A diagram illustrating a memory expander based on a dataset managed by cell size.
[0017] Figure 8 It is shown by Figure 1 A diagram illustrating a delimiter-managed dataset for a memory expander.
[0018] Figure 9 It is shown by Figure 1 A diagram illustrating the dataset managed by the memory expander.
[0019] Figure 10 It is used to describe according to Figure 9 A flowchart of the operation of an embodiment.
[0020] Figure 11A It is shown Figure 1 A diagram illustrating the operation of a computing device.
[0021] Figure 11B It is shown Figure 11A A diagram showing the completed header information.
[0022] Figure 12 It is shown Figure 1 A flowchart of the operation of a computing device.
[0023] Figure 13 It is shown Figure 1 A flowchart of the operation of a computing device.
[0024] Figure 14 It is shown Figure 1 The flowchart shows the operation of the memory expander.
[0025] Figure 15 This is a block diagram illustrating a computing device according to an embodiment of the present disclosure.
[0026] Figure 16 It is shown Figure 15 A flowchart of the operation of a computing device.
[0027] Figure 17 It is shown Figure 15 The flowchart shows the operation of the memory expander.
[0028] Figure 18A This is a flowchart illustrating the operation of a memory expander according to an embodiment of the present disclosure.
[0029] Figure 18B It is shown Figure 18A A diagram illustrating the header information included in the status request in operation S1010.
[0030] Figure 19 This is a block diagram illustrating a solid-state drive (SSD) system suitable for a memory expander according to this disclosure.
[0031] Figure 20 This is a circuit diagram illustrating a three-dimensional structure of a memory device included in a memory expander according to an embodiment of the present disclosure.
[0032] Figure 21 This is a block diagram illustrating a data center with a server system according to an embodiment of the present disclosure.
[0033] Figure 22 This is a diagram illustrating an example (e.g., a CXL interface) of a heterogeneous computing interface applied to embodiments of this disclosure. Detailed Implementation
[0034] Hereinafter, embodiments of the present disclosure will be described in detail and clearly to the extent that those skilled in the art can readily implement the present disclosure.
[0035] Figure 1 This is a diagram illustrating a computing device according to an embodiment of the present disclosure. (Refer to...) Figure 1 The computing device 100 may include a central processing unit (CPU) 101, memory 102, an accelerator 103, and a memory expander 110. In one embodiment, the computing device 100 may include a heterogeneous computing device or a heterogeneous computing system. A heterogeneous computing system can be a system that includes different types of computing devices organically connected to each other and configured to perform various functions. For example, such as Figure 1 As shown, computing device 100 may include CPU 101 and accelerator 103. CPU 101 and accelerator 103 may be different types of computing devices. Hereinafter, for ease of description, the simple term "computing device" is used, but this disclosure is not limited thereto. For example, computing device may refer to heterogeneous computing device.
[0036] CPU 101 may be a processor core configured to control the overall operation of computing device 100. For example, CPU 101 may be configured to decode instructions from the operating system or various programs driven on computing device 100 and process data based on the decoding results. CPU 101 may communicate with memory 102. Data processed by CPU 101 or data required during the operation of CPU 101 may be stored in memory 102. In one embodiment, memory 102 may be a dual in-line memory module (DIMM) based memory and may communicate directly with CPU 101. Memory 102 may be used as a buffer memory, cache memory, or system memory of CPU 101.
[0037] Accelerator 103 may include a processor core or computer (or calculator) configured to perform specific computations. For example, accelerator 103 may be a computer or processor configured to perform artificial intelligence (AI) operations, such as a graphics processing unit (GPU) or a neural processing unit (NPU). In one embodiment, accelerator 103 may perform computational operations under the control of CPU 101.
[0038] The memory expander 110 can operate under the control of the CPU 101. For example, the memory expander 110 can communicate with the CPU 101 via a heterogeneous computing interface. In one embodiment, the heterogeneous computing interface may include an interface based on the compute express link (CXL) protocol. Hereinafter, for ease of description, it is assumed that the heterogeneous computing interface is an interface based on the CXL protocol (i.e., a CXL interface), but this disclosure is not limited thereto. For example, the heterogeneous computing interface can be implemented using at least one of various computing interfaces based on protocols such as Gen-Z, NVLink, CCIX (Cache Coherent Interconnect for Accelerators), and the open CAPI (Coherent Accelerator Processor Interface) protocol.
[0039] The memory expander 110 can be controlled by the CPU 101 via the CXL interface to store or output stored data. That is, the CPU 101 can use the memory expander 110 as a memory region with functions similar to those of the memory 102. In one embodiment, the memory expander 110 may correspond to a Type 3 memory device as defined by the CXL standard.
[0040] The memory expander 110 may include a controller 111 and a memory device 112. The controller 111 may store data in the memory device 112 or may read data stored in the memory device 112.
[0041] In one embodiment, accelerator 103 can be connected to a CXL interface. Accelerator 103 can receive task commands from CPU 101 via the CXL interface, and can receive data from memory expander 110 via the CXL interface in response to the received task commands. Accelerator 103 can perform computational operations on the received data, and can store the results of the computational operations in memory expander 110 via the CXL interface.
[0042] In one embodiment, the data to be processed by accelerator 103 may be managed by CPU 101. In this case, whenever computation is processed by accelerator 103, data allocation may need to be handled by CPU 101, resulting in performance degradation. Memory expander 110 according to embodiments of the present disclosure may be configured to manage data allocated or to be allocated to accelerator 103 based on metadata from CPU 101. The operation of memory expander 110 according to embodiments of the present disclosure will now be described more fully with reference to the accompanying drawings.
[0043] Figure 2 It is shown Figure 1 A flowchart of the operation of a computing device. (Refer to...) Figure 2 This description refers to the operation of CPU 101 directly managing data to be processed by accelerator 103. In the following, unless otherwise defined, it is assumed that communication between components is performed on a packet-based communication basis using the CXL protocol. That is, CPU 101, accelerator 103, and memory expander 110 can communicate with each other via a CXL interface, and requests or data exchanged between CPU 101, accelerator 103, and memory expander 110 can have a packet-based communication structure using the CXL protocol. However, this disclosure is not limited thereto. For example, CPU 101, accelerator 103, and memory expander 110 can communicate with each other via at least one of various computing interfaces based on protocols such as Gen-Z, NVLink, CCIX, and the Open CAPI protocol.
[0044] For ease of description, it is assumed that the data to be processed by accelerator 103 is stored in memory expander 110. That is, CPU 101 can store the data to be processed by accelerator 103 in memory expander 110. However, this disclosure is not limited thereto.
[0045] Reference Figure 1 and Figure 2 In operation S1, CPU 101 can send a task request RQ_task and a first address AD1 to accelerator 103. The task request RQ_task can be a task start command for the data corresponding to the first address AD1. The first address AD1 can be a memory address managed by CPU 101, and the data corresponding to the first address AD1 can be stored in the memory device 112 of the memory expander 110.
[0046] In operation S2, accelerator 103 may send a read request RQ_rd and a first address AD1 to memory expander 110 in response to task request RQ_task. In operation S3, memory expander 110 may send first data DT1 corresponding to the first address AD1 to accelerator 103. In operation S4, accelerator 103 may perform a calculation on the first data DT1. In operation S5, accelerator 103 may send the calculation result of the first data DT1 (i.e., first result data RST1) and a write request RQ_wr to memory expander 110. In operation S6, memory expander 110 may store the first result data RST1 in memory device 112. In operation S7, accelerator 103 may generate an interrupt to CPU 101. In one embodiment, the interrupt may be information providing the following notification: the first data DT1 corresponding to the first address AD1 has been fully calculated, and the first result data RST1 has been stored in memory expander 110.
[0047] In operation S8, CPU 101 may send a task request RQ_task and a second address AD2 to accelerator 103 in response to an interrupt. In operation S9, accelerator 103 may send a read request RQ_rd and a second address AD2 to memory expander 110 in response to task request RQ_task. In operation S10, memory expander 110 may send second data DT2 corresponding to the second address AD2 to accelerator 103. In operation S11, accelerator 103 may perform a calculation on the second data DT2. In operation S12, accelerator 103 may send the calculation result of the second data DT2 (i.e., second result data RST2) and a write request RQ_wr to memory expander 110. In operation S13, memory expander 110 may store the second result data RST2 in memory device 112. In operation S14, accelerator 103 may generate an interrupt for CPU 101. In operation S15, CPU 101 may send a task request RQ_task and a third address AD3 to accelerator 103 in response to an interrupt. CPU 101, accelerator 103, and memory expander 110 may repeatedly perform the above process until all data included in the task has been fully computed.
[0048] As described above, when the data to be processed by accelerator 103 is managed by CPU 101, an interrupt can be generated for CPU 101 each time a computation is completed by accelerator 103, and CPU 101 repeatedly (e.g., again) sends an address to accelerator 103. This repetitive operation and interruption can result in low utilization of CPU 101.
[0049] Figure 3 It is shown Figure 1 A block diagram of the configuration of the memory expander. Figure 4 It is shown by Figure 3 A diagram illustrating the metadata managed by the controller's metadata manager. (See also...) Figure 1 , Figure 3 and Figure 4 The controller 111 may include a processor 111a, an SRAM 111b, a metadata manager 111c, a data manager 111d, a progress manager 111e, a host interface circuit 111f, and a memory interface circuit 111g.
[0050] Processor 111a can be configured to control the overall operation of controller 111 or memory expander 110. SRAM (Static Random Access Memory) 111b can operate as a buffer memory or system memory of controller 111. In one embodiment, components described below (such as metadata manager 111c, data manager 111d, and progress manager 111e) can be implemented in software, hardware, or a combination thereof. Components implemented in software can be stored in SRAM 111b and can be driven by processor 111a.
[0051] Metadata manager 111c can be configured to manage metadata MDT stored in memory expander 110. For example, metadata MDT may be information provided from CPU 101. Metadata MDT may include information needed by memory expander 110 to allocate data to accelerator 103, store result data, or manage reference data.
[0052] like Figure 4 As shown, the metadata MDT can include task metadata MDT_t and reference metadata MDT_r. Task metadata MDT_t can include at least one of the following: task number TN, data management mode dataMNGMode, data start address dAD_s, data end address dAD_e, task data cell size tUS, task data delimiter tDL, accelerator identifier ACID, result data start address rAD_s, and result data end address rAD_e.
[0053] The task number TN can be a number or identifier used to distinguish tasks assigned by CPU 101. The data management mode dataMNGMode can indicate the method of managing data or result data associated with the relevant task. In one embodiment, the data management method can include various management methods (such as heap, queue, and key-value).
[0054] The data start address dAD_s can indicate the start address of the memory region storing the dataset to be processed in the relevant task, and the data end address dAD_e can indicate the end address of the memory region storing the dataset to be processed in the relevant task.
[0055] The unit size tUS of the task data can indicate the unit used to divide the dataset between the data start address dAD_s and the data end address dAD_e. For example, if the dataset between the data start address dAD_s and the data end address dAD_e is 16 kilobytes and the unit size tUS of the task data is set to 2 kilobytes, the dataset can be divided into 8 (=16 / 2) units, and the accelerator 103 can receive one unit and perform computation on the received unit.
[0056] The delimiter tDL for task data can indicate the character or marker used to divide the dataset between the start address dAD_s and the end address dAD_e. For example, when the delimiter tDL for task data is set to ",", the dataset between the start address dAD_s and the end address dAD_e can be divided into cell data based on the delimiter ",". In one embodiment, the size of the cell data divided by the task data delimiter tDL can be different.
[0057] An accelerator identifier (ACID) can be information used to specify an accelerator that will process a related task. In one embodiment, as described with reference to the accompanying drawings, the computing device may include multiple accelerators. In this case, multiple individual tasks may be assigned to multiple accelerators respectively, and the CPU may set the accelerator identifier (ACID) of metadata based on the tasks assigned to each accelerator.
[0058] The starting address rAD_s of the result data can indicate the starting address of the memory region storing the result data calculated by accelerator 103. The ending address rAD_e of the result data can indicate the ending address of the memory region storing the result data calculated by accelerator 103.
[0059] In one embodiment, a particular computation performed by accelerator 103 may require reference to reference data or different data included in memory expander 110, as well as task data. In this case, the reference data may be selected based on reference metadata MDT_r.
[0060] For example, the reference metadata MDT_r may include at least one of the following: reference number RN, opcode, reference start address rAD_s, reference end address rAD_e, reference data cell size rUS, reference data separator rDL, tag, lifetime, and reservation information.
[0061] Reference number RN can be information used to identify data that will be referenced.
[0062] Opcodes can indicate information that defines computational operations or internal calculations that utilize reference data.
[0063] The reference start address rAD_s can indicate the start address of the memory region storing reference data, and the reference end address rAD_e can indicate the end address of the memory region storing reference data.
[0064] The reference data cell size rUS indicates the size of the data used to partition the reference data, and the reference data delimiter rDL can be a character or marker used to partition the reference data. Except for the target data, the reference data cell size rUS and reference data delimiter rDL are functionally and operationally similar to the task data cell size tUS and task data delimiter tDL; therefore, additional descriptions will be omitted to avoid redundancy.
[0065] Tags may include tag information associated with computational operations or internal computations that utilize reference data. Lifetime may include information about the time during which a computational operation is performed (e.g., the end time of the computational operation). Retention information may include any other information associated with the metadata MDT.
[0066] Reference data can be selected based on the aforementioned reference metadata MDT_r, and various computational operations (such as arithmetic operations, merge operations, and vector operations) can be performed on the task data and the selected reference data. In one embodiment, the information in the reference metadata MDT_r may correspond to a portion of the information in the task metadata MDT_t.
[0067] For better understanding, it is assumed hereafter that the metadata MDT is the task metadata MDT_t. That is, in the embodiments described below, it is assumed that accelerator 103 performs computations on task data. However, this disclosure is not limited thereto. For example, reference data can be selected using reference metadata MDT_r, and various computations can be performed on the task data and the selected reference data.
[0068] Metadata Manager 111c can manage the aforementioned metadata MDT.
[0069] Data manager 111d can be configured to manage data stored in memory device 112. For example, data manager 111d can identify task data stored in memory device 112 by using cell size or delimiter, and can sequentially output the identified cell data. Optionally, data manager 111d can be configured to sequentially store result data from accelerator 103 in memory device 112. In one embodiment, data manager 111d may include a read counter and a write counter, the read counter being configured to manage the output of task data and the write counter being configured to manage the input of result data.
[0070] The progress manager 111e can provide information about the task progress to the CPU 101 in response to a progress check request from the CPU 101. For example, the progress manager 111e can check the status of the task data being calculated and the task data that has been calculated based on the values of the read counter and write counter managed by the data manager 111d, and can provide the checked information to the CPU 101.
[0071] The controller 111 can communicate with the CPU 101 and the accelerator 103 via the host interface circuit 111f. The host interface circuit 111f can be an interface circuit based on the CXL protocol. The host interface circuit 111f can be configured to support at least one of various heterogeneous computing interfaces such as the Gen-Z protocol, NVLink protocol, CCIX protocol, and open CAPI protocol, as well as heterogeneous computing interfaces (such as the CXL protocol).
[0072] The controller 111 can be configured to control the memory device 112 via the memory interface circuitry 111g. The memory interface circuitry 111g can be configured to support various interfaces depending on the type of memory device 112. In one embodiment, the memory interface circuitry 111g can be configured to support memory interfaces such as a toggle interface or a double data rate (DDR) interface.
[0073] Figure 5A It is shown Figure 1 A flowchart of the operation of a computing device. Figure 5B It is shown Figure 5A A diagram illustrating the header information of the management request in operation S103. The embodiments described below are provided to readily illustrate the technical features of this disclosure, but this disclosure is not limited thereto. For example, in the flowcharts described below, specific operations are shown as independent operations, but this disclosure is not limited thereto. For example, several operations may be integrated by a single request.
[0074] For ease of description, it is assumed that the task data for the computation of accelerator 103 is pre-stored in memory expander 110 by CPU 101.
[0075] Unless otherwise defined, it is assumed that communication between components is performed based on the CXL protocol. That is, communication between components can be performed on the basis of communication packets based on the CXL protocol. However, this disclosure is not limited thereto. For example, components can communicate with each other based on one of the interfaces described above.
[0076] Reference Figure 1 , Figure 3 , Figure 5A and Figure 5B In operation S101, CPU 101 can send write request RQ_wr and metadata MDT to memory expander 110.
[0077] In operation S102, memory expander 110 can store metadata MDT. For example, metadata manager 111c of controller 111 can store and manage metadata MDT in memory device 112 or SRAM 111b.
[0078] In one embodiment, operations S101 and S102 can be executed during the initialization operation of the computing device 100. That is, the metadata MDT required by the memory expander 110 to perform management operations on task data can be loaded during the initialization operation. Alternatively, the memory expander 110 can store the metadata MDT as separate firmware. In this case, instead of operations S101 and S102, the CPU 101 can provide a request for loading firmware to the memory expander 110, and the memory expander 110 can store the metadata MDT by loading the firmware.
[0079] In operation S103, CPU 101 may send a management request RQ_mg to memory expander 110. The management request RQ_mg may be a message or communication packet used to request task data management operations associated with computation of accelerator 103. For example, the management request RQ_mg may correspond to an M2S RwD (Master to Subordinate Request with Data) message or communication packet of the CXL protocol. In this case, the management request RQ_mg may include... Figure 5B The first CXL header, CXL_header1, is shown in the image.
[0080] The first CXL header, CXL_header1, may include the following fields: Valid, MEMopcode, MetaField, MetaValue, SNP Type, Address, Tag, TC (traffic class), Poison, RSVD (reserved), and Position.
[0081] The Valid field can include information about whether the relevant request is valid.
[0082] The memory opcode field MEM opcode may include information about memory operations. In one embodiment, the memory opcode field MEM opcode of the management request RQ_mg according to an embodiment of this disclosure may include information about the initiation of data management operations of the memory expander 110 (e.g., "1101", "1110", or "1111").
[0083] The MetaField may include information indicating whether an update to the metadata is needed. The MetaValue field may include the value of the metadata. In one embodiment, the metadata MDT according to embodiments of this disclosure and the metadata described above may be different.
[0084] The SNP Type field can include information about the listener type.
[0085] The Address field may include information about the physical address of the host, which is associated with the memory opcode field. In one embodiment, the Address field of the management request RQ_mg according to an embodiment of this disclosure may include a first address information “Address[1]” about the starting memory address of the data to be managed and a second address information “Address[2]” about the ending memory address of the data to be managed.
[0086] The tag field may include tag information for identifying a pre-allocated memory region. In one embodiment, the tag field of the management request RQ_mg according to an embodiment of this disclosure may include an associated task number (refer to task metadata MDT_t).
[0087] The Service Category field (TC) can include information that defines the Quality of Service (QoS) associated with the request.
[0088] The Poison field can include information indicating whether errors exist in the data associated with the request.
[0089] The Position field may include the address of the metadata MDT. The Reserved RSVD field may include any other information associated with the request. In one embodiment, the Position field may be a newly added field compared to the M2S RwD field defined by the CXL standard (version 1.1). In one embodiment, the Position field may be omitted, and the information associated with the Position field may be included in the Reserved RSVD field or the Business Category field TC.
[0090] As described above, the management request RQ_mg according to embodiments of this disclosure can be generated by modifying some fields of the M2S RwD message defined by the CXL protocol.
[0091] In one embodiment, the write request RQ_wr in operation S101 can be replaced by the management request RQ_mg in operation S103. In this case, CPU 101 can send the management request RQ_mg and the metadata MDT to memory expander 110 simultaneously or through one operation. In this case, in response to the management request RQ_mg, memory expander 110 can store the metadata MDT and can begin management operations for task data.
[0092] In one embodiment, in response to a management request RQ_mg, the memory expander 110 may perform management operations on task data as follows.
[0093] In operation S104, CPU 101 can send the first task request RQ_task1 and the first task address dAD_s1 to accelerator 103. The first task address dAD_s1 can be the starting address or initial address of the dataset corresponding to the task assigned to accelerator 103.
[0094] In operation S105, accelerator 103 can send the first task address dAD_s1 and read request RQ_rd to memory expander 110.
[0095] In operation S106, the memory expander 110 can send the first data DT1 corresponding to the first address AD1 to the accelerator 103. For example, in response to the management request RQ_mg in operation S103, the memory expander 110 can be in a mode that performs management operations for task data. In this case, in response to the read request RQ_rd from the accelerator 103 and the first task address dAD_s1, the memory expander 110 can look up the task corresponding to the first task address dAD_s1 based on the metadata MDT, and can check the read count corresponding to the found task. The read count can indicate the number of cell data sent. The memory expander 110 can provide the data at the address corresponding to the sum of the first task address dAD_s1 and the read count to the accelerator 103. When the read count is "0" (i.e., in the case of the first task data output), in operation S106, the first data DT1 corresponding to the first address AD1 can be sent to the accelerator 103.
[0096] In operation S107, accelerator 103 can perform calculations on the first data DT1. In operation S108, accelerator 103 can send the calculation result of the first data DT1 (i.e., the first result data RST1), the first result address rAD_s1, and the write request RQ_wr to memory expander 110.
[0097] In operation S109, memory expander 110 can store first result data RST1. For example, in response to management request RQ_mg in operation S103, memory expander 110 can be in a mode that performs management operations on task data. In this case, in response to write request RQ_wr from accelerator 103 and first result address rAD_s1, memory expander 110 can look up the task corresponding to the first result address rAD_s1 based on metadata MDT, and can check the write count corresponding to the found task. The write count can indicate the amount of result data received from accelerator 103. Memory expander 110 can store the received task data at an address corresponding to the sum of the first result address rAD_s1 and the write count. When the write count is "0" (i.e., in the case of first result data input), in operation S109, first result data RST1 can be stored in the memory area corresponding to the first result address rAD_s1.
[0098] In operation S110, accelerator 103 may send a read request RQ_rd and a first task address dAD_s1 to memory expander 110. In one embodiment, accelerator 103 may execute operation S110 in response to a response to the request in operation S108 being received from memory expander 110. Operation S110 may be the same as operation S105.
[0099] In operation S111, the memory expander 110 can send the second data DT2 corresponding to the second address AD2 to the accelerator 103 in response to the read request RQ_rd and the first task address dAD_s1. For example, as described above, operation S110 can be the same as operation S105. However, because the memory expander 110 is in a mode performing management operations for task data, the memory expander 110 can output different data based on the read count for read requests for the same address. That is, in operation S111, the read count can be "1" (for example, because the first data DT1 was output in operation S106). In this case, the memory expander 110 can provide the second data DT2 of the second address AD2 corresponding to the sum of the first task address dAD_s1 and the read count "1" to the accelerator 103. In other words, when the memory expander 110 operates in management mode, even if read requests are received from the accelerator 103 along with the same address, different cell data can be output based on the read count.
[0100] In operation S112, accelerator 103 can perform calculations on the second data DT2. In operation S113, accelerator 103 can send the calculation result of the second data DT2 (i.e., the second result data RST2), the first result address rAD_s1, and the write request RQ_wr to memory expander 110.
[0101] In operation S114, memory expander 110 can store the second result data RST2. In one embodiment, memory expander 110 can store the second result data RST2 in a manner similar to that described in operation S109. For example, memory expander 110 can look up the task corresponding to the first result address rAD_s1 based on metadata MDT, and can check the write count corresponding to the found task. Memory expander 110 can store the received task data at an address corresponding to the sum of the first result address rAD_s1 and the write count. When the write count is "1" (i.e., in the case of second result data input), in operation S114, the second result data RST2 can be stored in the memory region corresponding to the sum of the first result address rAD_s1 and the write count "1". In other words, when memory expander 110 operates in management mode, even if a write request is received from accelerator 103 along with the same address, the result data can be stored in different memory regions according to the write count.
[0102] Then, accelerator 103 and memory expander 110 can perform operations S115 to S117. Operations S115 to S117 are similar to the operations described above, therefore, additional descriptions will be omitted to avoid redundancy.
[0103] As described above, according to one embodiment of this disclosure, the memory expander 110 can perform management operations on task data in response to a management request RQ_mg from the CPU 101. In this case, the accelerator 103 can receive multiple task data from the memory expander 110, perform computations on the received task data, and store multiple result data in the memory expander 110 without additional intervention from the CPU 101. That is, as referred to Figure 2 In contrast to the description, when accelerator 103 repeatedly performs multiple calculations, because interrupts are not generated from accelerator 103, CPU 101 can perform any other operations. Therefore, the utilization of CPU 101 can be improved.
[0104] Figure 6 It is shown Figure 1 A flowchart illustrating the operation of the memory expander. (Refer to...) Figure 1 , Figure 5A and Figure 6 In operation S201, the memory expander 110 can receive metadata MDT from the CPU 101. The memory expander 110 can store and manage the received metadata MDT.
[0105] In operation S202, the memory expander 110 can receive a management request RQ_mg from the CPU 101. The management request RQ_mg may include references Figure 6 The first CXL header is described. Memory expander 110 can perform management operations on task data in response to management request RQ_mg.
[0106] In operation S203, the memory expander 110 can receive a read request RQ_rd and a first task address dAD_s1 from the accelerator 103.
[0107] In operation S204, the memory expander 110 can check the read count RC. For example, the memory expander 110 can look up the task corresponding to the first task address dAD_s1 based on the metadata MDT, and can check the read count RC corresponding to the found task.
[0108] In operation S205, memory expander 110 can send data DT stored in the memory region corresponding to the sum of read count RC and first task address dAD_s1 (i.e., dAD_s1+RC) to accelerator 103. In one embodiment, after operation S205, memory expander 110 can increase the value of read count RC corresponding to the found task by "1".
[0109] In operation S206, memory expander 110 can receive result data RST, first result address rAD_s1 and write request RQ_wr from accelerator 103.
[0110] In operation S207, the memory expander 110 can check the write count WC. For example, the memory expander 110 can look up the task corresponding to the first result address rAD_s1 based on the metadata MDT, and can check the write count WC corresponding to the found task.
[0111] In operation S208, memory expander 110 can store the result data RST in the memory region corresponding to the sum of the write count WC and the first result address rAD_s1 (i.e., rAD_s1+WC). In one embodiment, after operation S208, memory expander 110 can increase the value of the write count WC corresponding to the found task by "1".
[0112] As described above, when the memory expander 110 performs a management operation, the memory expander 110 may output task data based on the metadata MDT in response to a read request RQ_rd received from the accelerator 103. Optionally, when the memory expander 110 performs a management operation, the memory expander 110 may store result data based on the metadata MDT in response to a write request RQ_wr received from the accelerator 103.
[0113] Figure 7 It is shown by Figure 1 A diagram illustrating a memory expander based on a cell size-managed dataset. (See reference...) Figure 1 and Figure 7 The memory device 112 can store the dataset for the first task, Task_1. The dataset for Task_1 can be stored in a memory area between the start address dAD_s1 and the end address dAD_e1 of the first task. If the dataset for Task_1 is divided based on a first unit size US_1, the dataset for Task_1 can be divided into first data DT1 to fourth data DT4, and each of the first data DT1 to fourth data DT4 can have a first unit size US_1.
[0114] In one embodiment, the result dataset for the first task Task_1 can be stored in a memory region between the first result start address rAD_s1 and the first result end address rAD_e1. The memory region between the first task start address rAD_s1 and the first task end address rAD_e1 and the memory region between the first result start address rAD_s1 and the first result end address rAD_e1 can be different memory regions, or they can at least partially overlap.
[0115] If the dataset used for Task 1 is partitioned based on the first cell size US_1, the result dataset used for Task 1 can also be partitioned based on the first cell size US_1. In this case, the result dataset used for Task 1 can be divided into first result data RST1 to fourth result data RST4.
[0116] In one embodiment, first result data RST1 can indicate the calculation result of first data DT1, second result data RST2 can indicate the calculation result of second data DT2, third result data RST3 can indicate the calculation result of third data DT3, and fourth result data RST4 can indicate the calculation result of fourth data DT4. That is, the order in which the result data is stored sequentially can be the same as the order in which the unit data is output sequentially. However, this disclosure is not limited to this. For example, the result data can be stored non-sequentially. For example, first result data RST1 can indicate the calculation result of third data DT3, second result data RST2 can indicate the calculation result of first data DT1, third result data RST3 can indicate the calculation result of fourth data DT4, and fourth result data RST4 can indicate the calculation result of second data DT2. The order in which the result data is stored can be changed or modified depending on the characteristics of the calculation operation.
[0117] Figure 8 It is shown by Figure 1 A diagram illustrating the delimiter-managed dataset of the memory expander. (See also...) Figure 1 and Figure 8 The memory device 112 can store the dataset for the second task, Task_2. The dataset for Task_2 can be stored in the memory area between the start address dAD_s2 and the end address dAD_e2 of the second task. If the dataset for Task_2 is divided based on the second delimiter DL_2, the dataset for Task_2 can be divided into fifth data DT5 to eighth data DT8, and the fifth data DT5 to the eighth data DT8 can have different sizes.
[0118] The result dataset for the second task, Task_2, can be stored in the memory region between the starting address rAD_s2 and the ending address rAD_e2 of the second result. The memory region between the starting address rAD_s2 and the ending address rAD_e2 of the second task, and the memory region between the starting address rAD_s2 and the ending address rAD_e2 of the second result, can be different memory regions, or they can at least partially overlap.
[0119] The result dataset for Task_2 can be divided into fifth result data RST5 to eighth result data RST8, and the sizes of fifth result data RST5 to eighth result data RST8 can correspond to the sizes of fifth data DT5 to eighth data DT8, respectively. In one embodiment, the order in which the result data is stored can be as follows: Figure 7 The descriptions have been changed or modified, therefore, additional descriptions will be omitted to avoid redundancy.
[0120] Figure 9 It is shown by Figure 1 A diagram illustrating the dataset managed by the memory expander. Figure 10 It is used to describe according to Figure 9 A flowchart illustrating the operation of an embodiment. See also... Figure 1 , Figure 9 and Figure 10 The memory device 112 can store the dataset for the third task, Task_3. The dataset for Task_3 can be stored in the memory area between the start address dAD_s3 and the end address dAD_e3 of the third task. If the dataset for Task_3 is partitioned based on the third delimiter DL_3, the dataset for Task_3 can be divided into ninth data DT9 to twelfth data DT12.
[0121] In one embodiment, a portion of the ninth data DT9 through twelfth data DT12 (e.g., the tenth data DT10) may not be stored in memory device 112. In this case, the tenth data DT10 may be stored in memory 102 directly connected to CPU 101, and memory expander 110 may include the address point ADP of the tenth data DT10 instead of the tenth data DT10 itself. The address point ADP may indicate information about the location where the tenth data DT10 is actually stored (i.e., information about the address of memory 102).
[0122] When the tenth data DT10 is output, the memory expander 110 can provide the address point ADP corresponding to the tenth data DT10 to the accelerator 103. The accelerator 103 can receive the tenth data DT10 from the CPU 101 based on the address point ADP.
[0123] For example, such as Figure 10 As shown, CPU 101, accelerator 103, and memory expander 110 can execute operations S301 to S305. Except that the task request is the third task request RQ_task3 and the initial address of the task data is the third task start address dAD_s3, operations S301 to S305 are similar to... Figure 5A Operations S101 to S105 are similar, therefore, additional descriptions will be omitted to avoid redundancy.
[0124] In operation S306, the memory expander 110 can determine whether address point ADP is stored in the memory region corresponding to the task data to be sent to accelerator 103. When it is determined that address point ADP is not stored in the memory region corresponding to the task data to be sent to accelerator 103, in operation S307, the memory expander 110 and accelerator 103 can perform operations associated with sending and calculating the task data. Operation S307 is related to reference... Figure 5A The operations described (i.e., the operation of exchanging task data between memory expander 110 and accelerator 103 and the operation of performing calculations on the data) are similar, so additional descriptions will be omitted to avoid redundancy.
[0125] When the address point ADP is determined to be stored in the memory area corresponding to the task data to be sent to the accelerator 103, in operation S308, the memory expander 110 can send information about the address point ADP to the accelerator 103.
[0126] In operation S309, accelerator 103 may send address point ADP and read request RQ_rd to CPU 101 in response to address point ADP received from memory expander 110.
[0127] In operation S310, CPU 101 may send read command RD and address point ADP to memory 102 in response to read request RQ_rd; in operation S311, memory 102 may send tenth data DT10 corresponding to address point ADP to CPU 101. In one embodiment, operations S310 and S311 may be executed based on the communication interface (e.g., DDR interface) between CPU 101 and memory 102.
[0128] In operation S312, CPU 101 can send the tenth data DT10 to accelerator 103. Accelerator 103 and memory expander 110 can execute operations S313 and S314. Operations S313 and S314 are related to reference. Figure 5A The data computation operations and result data storage operations described are similar; therefore, additional descriptions will be omitted to avoid redundancy.
[0129] As described above, the memory expander 110 can store address point information of the actual data, rather than storing the actual data corresponding to a portion of the task dataset. In this case, even if the entire task dataset is not stored in the memory expander 110, the processing of receiving the task dataset from the CPU 101 can be omitted before the management operation is executed, thus improving the overall performance of the computing device 100.
[0130] Figure 11A It is shown Figure 1 A diagram illustrating the operation of a computing device. Figure 11B It is shown Figure 11A A diagram illustrating the completed header information. (Refer to...) Figure 1 , Figure 11A and Figure 11B The CPU 101, accelerator 103, and memory expander 110 can execute operations S401 to S414. Operations S401 to S414 are related to... Figure 5A The metadata storage operations, management operation initiation requests, task data reading operations, task data calculation operations, and result data storage operations described are similar; therefore, additional descriptions will be omitted to avoid redundancy.
[0131] In one embodiment, the nth data DTn sent from memory expander 110 to accelerator 103 by operations S410 to S411 may be the last task data or the last cell data associated with the assigned task.
[0132] In this scenario, after the result data RSTn, which is the calculation result of the nth data DTn, is stored in the memory expander 110 (i.e., after operation S414), in operation S415, the memory expander 110 can send the completion associated with the assigned task to the CPU 101. In response to the received completion, the CPU 101 can recognize that the assigned task has been completed and can assign the next task.
[0133] In one embodiment, after all assigned tasks have been completed, in operation S416, memory expander 110 may send the completion notification to accelerator 103. Upon receiving the completion notification, accelerator 103 can recognize that the assigned tasks have been completed.
[0134] In one embodiment, after sending completion to CPU 101, memory expander 110 may cease management operations. In this case, when a new management request RQ_mg is sent from CPU 101 to memory expander 110, memory expander 110 may resume management operations. Alternatively, after sending completion to CPU 101, memory expander 110 may continue performing management operations until an explicit request to cease management operations is received from CPU 101.
[0135] In one embodiment, the completion associated with the assigned task may have the structure of an S2M DRS (Subordinate to Master Data Response) message or communication packet based on the CXL protocol. For example, the completion may include... Figure 11B The second CXL header, CXL_header2, is shown in the image.
[0136] The second CXL header, CXL_header2, may include the Valid field, the MEMopcode field, the MetaField field, the MetaValue field, the Tag field, the TC field, the Poison field, the RSVD field, and the Position field. Each field of the second CXL header, CXL_header2, is referenced... Figure 5B Since it is already described, additional descriptions will be omitted to avoid redundancy.
[0137] In one embodiment, the Memory Opcode field (MEM opcode) included in the completion according to this disclosure may include information about the processing result of the assigned task (e.g., information indicating "normal", "abnormal", "error", or "interruption"). The Tag field (Tag) included in the completion according to this disclosure may include information about the task number of the completed task. The Position field (Position) included in the completion according to this disclosure may include information about the address of the result data. The Reserved field (RSVD) included in the completion according to this disclosure may include various statistics about the assigned task (e.g., throughput, processing time, and processing error count).
[0138] In one embodiment, as referenced above Figure 5B In the given description, Figure 11B The Position field can be newly added to the S2M DRS fields defined by the CXL standard (version 1.1). Optionally, Figure 11BThe Position field can be omitted, and the information associated with the Position field can be included in the Reserved field RSVD or the Business Category field TC.
[0139] Figure 12 It is shown Figure 1 A flowchart of the operation of the computing device. (Refer to...) Figure 1 and Figure 12 The CPU 101, accelerator 103, and memory expander 110 can execute operations S501 to S514. Operations S501 to S514 are related to... Figure 5A The metadata storage operations, management operation initiation requests, task data reading operations, task data calculation operations, and result data storage operations described are similar; therefore, additional descriptions will be omitted to avoid redundancy.
[0140] In one embodiment, the nth data DTn sent from memory expander 110 to accelerator 103 by operations S510 to S511 may be the last task data or the last cell data associated with the assigned task.
[0141] In this case, after the result data RSTn, which is the calculation result of the nth data DTn, is stored in the memory expander 110 (i.e., after operation S514), in operation S515, the memory expander 110 can send the completion associated with the assigned task to the accelerator 103. In operation S516, the accelerator 103 can send the completion to the CPU 101 in response to the completion from the memory expander 110. Besides the completion being provided from the memory expander 110 to the CPU 101 via the accelerator 103, the grouping structure of the completion is consistent with the reference... Figure 11A and Figure 11B The grouping structures described are similar, therefore, additional descriptions will be omitted to avoid redundancy.
[0142] Figure 13 It is shown Figure 1 A flowchart of the operation of the computing device. (Refer to...) Figure 1 and Figure 13 The CPU 101, accelerator 103, and memory expander 110 can execute operations S601 to S614. Operations S601 to S614 are related to... Figure 5A The metadata storage operations, management operation initiation requests, task data reading operations, task data calculation operations, and result data storage operations described are similar; therefore, additional descriptions will be omitted to avoid redundancy.
[0143] In one embodiment, the nth data DTn sent from memory expander 110 to accelerator 103 by operations S610 to S611 may be the last task data or the last cell data associated with the assigned task.
[0144] In this case, the result data RSTn, which is the calculation result of the nth data DTn, can be stored in the memory expander 110 (operation S614). Then, in operation S615, the read request RQ_rd and the first task address dAD_s1 can be received from the accelerator 103. The memory expander 110 can check that the last cell data was sent to the accelerator 103 based on the metadata MDT and the read count (or the value of the read counter). In this case, in operation S616, the memory expander 110 can send the end data EoD to the accelerator 103. In response to the end data EoD, the accelerator 103 can recognize that the assigned task has been completed. In operation S617, the accelerator 103 can send the completion to the CPU 101.
[0145] As described above, a task can include computational operations targeting multiple units of data. Upon completion of a task, the completion can be sent from the accelerator 103 or the memory expander 110 to the CPU 101. That is, in accordance with reference... Figure 2 Compared to the described method, this can prevent frequent interruptions, thus improving the overall performance of the computing device 100.
[0146] Figure 14 It is shown Figure 1 A flowchart illustrating the operation of the memory expander. (Refer to...) Figure 1 and Figure 14 In operation S701, the memory expander 110 may receive a request RQ. In one embodiment, the request RQ may be a request received from the CPU 101, accelerator 103, or any other component via the CXL interface.
[0147] In operation S702, the memory expander 110 can determine whether the current operating mode is a management mode (i.e., a mode for performing management operations). For example, as described above, the memory expander 110 can perform management operations on task data in response to a management request RQ_mg from the CPU 101.
[0148] When it is determined that the memory expander 110 is in management mode, in operation S703, the memory expander 110 can be based on a reference Figures 3 to 13 The described method (i.e., management mode) handles request RQ. When it is determined that the memory expander 110 is not in management mode, in operation S704, the memory expander 110 can handle request RQ based on normal mode.
[0149] For example, suppose request RQ is a read request, and a first address is received along with request RQ. In this case, when memory expander 110 is in management mode, memory expander 110 can look up the task corresponding to the first address based on metadata MDT, and can output data of the memory region corresponding to the first address and the sum of the read counts corresponding to the found task. Conversely, when memory expander 110 is not in management mode (i.e., in normal mode), memory expander 110 can output data of the memory region corresponding to the first address.
[0150] In the above embodiments, a configuration was described in which CPU 101 provides a task request RQ_task and a task start address dAD_s to accelerator 103, and accelerator 103 sends a read request RQ_rd and a task start address dAD_s to memory expander 110, but this disclosure is not limited thereto. For example, the task start address dAD_s can be replaced with a task number TN, and memory expander 110 can receive the task number from accelerator 103 and can manage the task data and result data corresponding to the received task number based on metadata MDT.
[0151] Figure 15 This is a block diagram illustrating a computing device according to an embodiment of the present disclosure. (Refer to...) Figure 15 The computing device 1000 may include a CPU 1010, a memory 1011, multiple accelerators 1210 to 1260, and a memory expander 1100. The CPU 1010, memory 1011, and memory expander 1100 of the computing device 1000 are referenced to... Figure 1 Therefore, additional descriptions will be omitted to avoid redundancy. Each of the multiple accelerators 1210 to 1260 can perform operations with reference to... Figure 1 The operation of the accelerator 103 described is similar to that of the accelerator 103, therefore, additional descriptions will be omitted to avoid redundancy.
[0152] The CPU 1010, multiple accelerators 1210 to 1260, and memory expander 1100 can communicate with each other via a CXL interface. Each of the multiple accelerators 1210 to 1260 can be configured as shown in reference. Figures 1 to 14 The described execution allocates computation from CPU 1010. That is, multiple accelerators 1210 to 1260 can be configured to perform parallel computation. Memory expander 1100 can be configured as described in reference... Figures 1 to 14The described management involves task data or result data that will be computed at each of the multiple accelerators 1210 to 1260. The memory expander 1100 may include a controller 1110 and a memory device 1120. In addition to the computing device 1000 including multiple accelerators, communication and reference between the CPU 1010, the multiple accelerators 1210 to 1260, and the memory expander 1100 are also described. Figures 1 to 14 The described communications are similar; therefore, additional descriptions will be omitted to avoid redundancy. (Refer to...) Figure 16 An embodiment of parallel computing utilizing multiple accelerators is described more fully.
[0153] In one embodiment, the number of accelerators 1210 to 1260 included in the computing device 1000 may be changed or modified differently.
[0154] Figure 16 It is shown Figure 15 A flowchart of the operation of the computing device is provided. For ease of description, parallel computing utilizing the first accelerator 1210 and the second accelerator 1220 will be described. However, this disclosure is not limited thereto. See also... Figure 15 and Figure 16 CPU 1010 and memory expander 1100 can execute operations S801 to S803. Operations S801 to S803 are related to reference. Figure 5A Operations S101 to S103 are described similarly, therefore, additional descriptions will be omitted to avoid redundancy.
[0155] In operation S804, CPU 1010 may send task request RQ_task to first accelerator 1210. In operation S805, CPU 1010 may send task request RQ_task to second accelerator 1220. In one embodiment, the task request RQ_task provided to first accelerator 1210 and the task request RQ_task provided to second accelerator 1220 may be associated with the same task. Optionally, the task request RQ_task provided to first accelerator 1210 and the task request RQ_task provided to second accelerator 1220 may be associated with different tasks. Optionally, the task request RQ_task provided to first accelerator 1210 and the task request RQ_task provided to second accelerator 1220 may not include information about the task (e.g., task number or task starting address).
[0156] Hereinafter, it is assumed that the task request RQ_task provided to the first accelerator 1210 and the task request RQ_task provided to the second accelerator 1220 do not include information about the task (e.g., task number or task start address). However, this disclosure is not limited thereto. When the task request RQ_task provided to the first accelerator 1210 and the task request RQ_task provided to the second accelerator 1220 include information about the task (e.g., task number or task start address), the task number or task start address can be provided to the memory expander 1100 in a read request for task data or a write request for result data. In this case, the memory expander 1100 can be as described with reference to Figures 1 to 14 The described operation is to be performed.
[0157] In operation S806, the first accelerator 1210 may send a read request RQ_rd to the memory expander 1100. In operation S807, the memory expander 1100 may send first data DT1 for the first task to the first accelerator 1210. For example, in response to the read request RQ_rd, the memory expander 1100 may search in the metadata MDT for a first task corresponding to the accelerator identifier of the first accelerator 1210 that sent the read request RQ_rd. The memory expander 1100 may then send the task data (i.e., the first data DT1) corresponding to the first task found therein to the first accelerator 1210. In one embodiment, as referenced... Figures 1 to 14 The first data DT1 may be data output from a memory region corresponding to the sum of the task start address and the read counts corresponding to the first task found therefrom.
[0158] In operation S808, the first accelerator 1210 can perform calculation operations on the first data DT1.
[0159] In operation S809, the second accelerator 1220 may send a read request RQ_rd to the memory expander 1100. In operation S810, in response to the read request RQ_rd from the second accelerator 1220, the memory expander 1100 may send second data DT2 for the second task to the second accelerator 1220. In one embodiment, operation S810 may be similar to operation S807, except that the accelerator and the data sent are different; therefore, additional descriptions will be omitted to avoid redundancy.
[0160] In operation S811, the second accelerator 1220 can perform computational operations on the second data DT2.
[0161] In operation S812, the first accelerator 1210 may send first result data RST1, which is the result of a computation operation on first data DT1, and a write request RQ_wr to the memory expander 1100. In operation S813, the memory expander 1100 may store the first result data RST1 in response to the write request RQ_wr from the first accelerator 1210. In one embodiment, the memory region for storing the first result data RST1 may be determined based on the accelerator identifier of the first accelerator 1210 and the metadata MDT. For example, the memory expander 1100 may search in the metadata MDT for a task corresponding to the accelerator identifier of the first accelerator 1210, and may determine the memory region for storing the first result data RST1 based on the starting address of the result data corresponding to the found task and the write count (or the value of the write counter). Except for the operation of searching for the task number corresponding to the accelerator identifier, the remaining components are similar to the components described above (e.g., similar in operation), therefore, additional descriptions will be omitted to avoid redundancy.
[0162] In operation S814, the second accelerator 1220 can send the write request RQ_wr and the second result data RST2 to the memory expander 1100. In operation S815, the memory expander 1100 can store the second result data RST2. Operation S815 is similar to operation S814, therefore, additional descriptions will be omitted to avoid redundancy.
[0163] In one embodiment, communication between the first accelerator 1210 and the second accelerator 1220 and the memory expander 1100 can be performed in parallel. For example, when the first accelerator 1210 performs a computation operation (i.e., operation S808), the second accelerator 1220 and the memory expander 1100 can perform operations such as sending a read request RQ_rd and sending task data. Optionally, when the second accelerator 1220 performs a computation operation (i.e., operation S811), the first accelerator 1210 and the memory expander 1100 can perform operations such as sending a write request RQ_wr and sending result data.
[0164] As described above, the memory expander 1100 according to embodiments of the present disclosure can be configured to manage task data to be processed by multiple accelerators and result data processed by multiple accelerators.
[0165] In one embodiment, depending on the task allocation method of CPU 1010, multiple accelerators 1210 to 1260 can be configured to process the same task in parallel, or to process different tasks.
[0166] Figure 17 It is shown Figure 15A flowchart illustrating the operation of the memory expander. (Refer to...) Figure 15 and Figure 17 In operation S911, the memory expander 1100 may receive a read request RQ_rd from a first accelerator 1210 among a plurality of accelerators 1210 to 1260. In one embodiment, the read request RQ_rd may include a reference Figure 5A The description includes the starting address of the task. Optionally, the read request RQ_rd may include information about the task to be processed by the first accelerator 1210 (e.g., task number). Optionally, the read request RQ_rd may include information about the first accelerator 1210 (e.g., accelerator identifier).
[0167] In operation S912, the memory expander 1100 can look up the task corresponding to the first accelerator 1210 based on the metadata MDT. For example, the memory expander 1100 can search for the relevant task based on at least one of the information included in the read request RQ_rd (e.g., task start address, task number, and / or accelerator identifier).
[0168] In operation S913, the memory expander 1100 can check the read count of the found task. In operation S914, the memory expander 1100 can send the data corresponding to the read count to the first accelerator 1210. Operations S913 and S914 (i.e., the read count check operation and the data transmission operation) are consistent with reference to... Figure 6 The read count check operation and data sending operation are similar, so additional descriptions will be omitted to avoid redundancy.
[0169] In one embodiment, after data is sent to the first accelerator 1210, the memory expander 1100 can increase the read count of the found task by "1".
[0170] For ease of description, the configuration in which task data is sent to the first accelerator 1210 is described, but this disclosure is not limited thereto. For example, the task data sending operation associated with each of the remaining accelerators can be performed in a manner similar to that described with reference to operations S911 to S914.
[0171] In operation S921, memory expander 1100 may receive write request RQ_wr and second result data RST2 from second accelerator 1220. In one embodiment, as described above with reference to operation S911, write request RQ_wr may include the starting address of the result, information about the task being processed (e.g., task number), or information such as an accelerator identifier.
[0172] In operation S922, the memory expander 1100 can search for the task corresponding to the second accelerator 1220 based on the metadata MDT. The search operation (i.e., operation S922) is similar to operation S912, therefore, additional descriptions will be omitted to avoid redundancy.
[0173] In operation S923, the memory expander 1100 can check the write count of the found task. In operation S924, the memory expander 1100 can store the result data in the area corresponding to the write count. Operations S923 and S924 are consistent with reference to... Figure 5A The write count check operation and the result data storage operation are similar, so additional descriptions will be omitted to avoid redundancy.
[0174] For ease of description, the configuration in which the result data from the second accelerator 1220 is stored is described, but this disclosure is not limited thereto. For example, result data received from each of the remaining accelerators may be stored in a manner similar to that described with reference to operations S921 to S924.
[0175] Figure 18A This is a flowchart illustrating the operation of a memory expander according to an embodiment of the present disclosure. Figure 18B It is shown Figure 18A A diagram illustrating the header information included in the status request during operation S1010. For ease of description, it will be based on... Figure 18A The operation of the flowchart was Figure 15 The memory expander 1100 is used to provide a description. However, this disclosure is not limited thereto.
[0176] Reference Figure 15 , Figure 18A and Figure 18B In operation S1010, the memory expander 110 can receive status requests from the CPU 1010. For example, the CPU 1010 can assign various tasks to multiple accelerators 1210 to 1260 and can request the memory expander 1100 to perform management operations on task data to be processed by the multiple accelerators 1210 to 1260. The memory expander 1100 can manage task data and result data based on metadata MDT without intervention from the CPU 1010. When various tasks are executed, the CPU 1010 can check the progress status of the executing task. In this case, the CPU 1010 can send a status request to the memory expander 1100.
[0177] In one embodiment, a status request can be an M2S Req (Master to Subordinate Request) message or communication packet from the CXL protocol. For example, a status request may include... Figure 18B The third CXL header, CXL_header3, is shown below. The third CXL header, CXL_header3, may include the Valid field, the MEMopcode field, the MetaField field, the MetaValue field, the SNP Type field, the Address field, the Tag field, the TC field, and the Reserved field RSVD. Each field of the third CXL header, CXL_header3, has been described above; therefore, additional descriptions will be omitted to avoid redundancy.
[0178] In one embodiment, the memory opcode field MEMopcode for a status request according to embodiments of the present disclosure can be set to various values (e.g., "1101", "1110", or "1111") depending on the type of command used to query the processing status of data included in the memory expander 1100. In one embodiment, the type of command used to query the processing status can include the following types: simple query, query interrupt, and query wait.
[0179] In one embodiment, the Tag field of the status request according to an embodiment of this disclosure may include information about the task number.
[0180] In one embodiment, the Address field of the status request according to embodiments of this disclosure may indicate the range of query request data. In one embodiment, the query request data may indicate task data for which processing status will be checked.
[0181] In one embodiment, the reserved field RSVD of the status request according to embodiments of this disclosure may include information such as query request unit (or query unit) or query request time (or query time).
[0182] The header information of the above query request is provided as an example, and this disclosure is not limited thereto.
[0183] In operation S1020, the memory expander 1100 can check the read count and write count. For example, based on the read count or write count, the memory expander 1100 can determine whether the data corresponding to the address field included in the status request has been processed. Specifically, the read count "10" associated with the first task can indicate that the first to tenth task data among the multiple task data corresponding to the first task has been sent to the accelerator. Optionally, the write count "10" associated with the first task can indicate that the first to tenth result data among the multiple result data corresponding to the first task has been stored in the memory expander 1100. That is, the status of the currently processed task can be checked based on the read count (or the value of the read counter) and the write count (or the value of the write counter).
[0184] In operation S1030, the memory expander 1100 can provide information about read counts and write counts to the CPU 1010. The CPU 1010 can check the progress status of the current task based on the information received from the memory expander 1100.
[0185] As described above, according to this disclosure, the memory expander can perform data management operations for each of the multiple accelerators. In this case, the CPU does not need to control the multiple accelerators individually, and each of the multiple accelerators does not need to generate a separate interrupt for the CPU until the task of the specific unit is completed. Therefore, the performance of the computing device can be improved, and the utilization of the CPU can be increased.
[0186] Figure 19 This is a block diagram illustrating a solid-state drive (SSD) system suitable for a memory expander according to this disclosure. (Refer to...) Figure 19 The SSD system 2000 may include a host 2100 and a storage device 2200. The storage device 2200 can exchange signals SIG with the host 2100 through a signal connector 2201 and can be powered by a power connector 2202. The storage device 2200 includes an SSD controller 2210, multiple non-volatile memories (NVMs) 2221 to 222n, an auxiliary power supply 2230, and a cache memory 2240.
[0187] SSD controller 2210 can control multiple non-volatile memories 2221 to 222n in response to a SIG signal received from host 2100. The multiple non-volatile memories 2221 to 222n can operate under the control of SSD controller 2210. Auxiliary power supply 2230 is connected to host 2100 via power connector 2202. Auxiliary power supply 2230 can be charged by power PWR supplied from host 2100. When power PWR is not stably supplied from host 2100, auxiliary power supply 2230 can supply power to storage device 2200. Buffer memory 2240 can be used as buffer memory for storage device 2200.
[0188] In one embodiment, host 2100 may include, as referenced Figures 1 to 18B The CPU and multiple accelerators described. In one embodiment, storage device 2200 may be a reference... Figures 1 to 18B The described memory expander. The host 2100 and storage device 2200 can communicate with each other via a CXL interface, and can be referenced... Figures 1 to 18B The described embodiments are operated.
[0189] Figure 20 This is a circuit diagram illustrating a three-dimensional structure of a memory device included in a memory expander according to an embodiment of the present disclosure. In one embodiment, the memory device may be implemented based on various types of memory. (Refer to...) Figure 20 This disclosure describes the configuration of a memory device based on a specific memory architecture, but is not limited thereto. For example, the memory device may be implemented based on at least one of a variety of memories.
[0190] Reference Figure 20 The memory device can be implemented in a three-dimensional stacked structure. For example, the memory device includes a first memory cell array layer MCA1 to a fourth memory cell array layer MCA4. The first memory cell array layer MCA1 to the fourth memory cell array layer MCA4 may include a plurality of memory cells MC1, MC2, MC3 and MC4.
[0191] First memory cell array layers MCA1 to fourth memory cell array layers MCA4 can be stacked along a third direction D3. Conductive lines CL1 and CL2 extending along a first direction D1 and a second direction D2 can be alternately formed between the first memory cell array layers MCA1 and fourth memory cell array layers MCA4. For example, the first conductive line CL1 can extend along the first direction D1, and the second conductive line CL2 can extend along the second direction D2. The first memory cell array layer MCA1 can be formed above the first conductive line CL1, and the second conductive line CL2 can be formed between the first memory cell array layer MCA1 and the second memory cell array layer MCA2. The first conductive line CL1 can be formed between the second memory cell array layer MCA2 and the third memory cell array layer MCA3, and the second conductive line CL2 can be formed between the third memory cell array layer MCA3 and the fourth memory cell array layer MCA4. The first conductive line CL1 can be formed above the fourth memory cell array layer MCA4. The first conductive line CL1 and the second conductive line CL2 can be electrically connected to adjacent memory cells on the third direction D3.
[0192] In one embodiment, target bit lines and target word lines can be determined based on the location of the target memory cell MC. For example, if the first memory cell MC1 of the first memory cell array layer MCA1 is the target memory cell MC, conductive lines CL1a and CL2a can be selected as target lines. If the second memory cell MC2 of the second memory cell array layer MCA2 is the target memory cell MC, conductive lines CL2a and CL1b can be selected as target lines. If the third memory cell MC3 of the third memory cell array layer MCA3 is the target memory cell MC, conductive lines CL1b and CL2b can be selected as target lines. That is, target lines can be selected based on the location of the target memory cell MC, and the selected target lines can be used as bit lines and word lines or as word lines and bit lines, depending on the location of the target memory cell. However, this disclosure is not limited thereto.
[0193] Figure 21 This is a block diagram illustrating a data center employing a server system according to an embodiment of the present disclosure. (Refer to...) Figure 21A data center 3000, serving as a facility for maintaining diverse data and providing various data-related services, can be referred to as a "data storage center." The data center 3000 can be a system used for search engines or database management, or it can be a computing system used by various organizations. The data center 3000 may include multiple application servers 3100_1 to 3100_n and multiple storage servers 3200_1 to 3200_m. The number of application servers 3100_1 to 3100_n and the number of storage servers 3200_1 to 3200_m can be varied or modified.
[0194] For ease of description, an example of the first storage server 3200_1 will be described below. Each of the remaining storage servers 3200_2 to 3200_m and the plurality of application servers 3100_1 to 3100_n may have a structure similar to that of the first storage server 3200_1.
[0195] The first storage server 3200_1 may include a processor 3210_1, a memory 3220_1, a switch 3230_1, a network interface connector (NIC) 3240_1, a storage device 3250_1, and a compute fast link (CXL) interface controller 3260_1. The processor 3210_1 may perform (e.g., control) the overall operation of the first storage server 3200_1. The memory 3220_1 may store various instructions or data under the control of the processor 3210_1. The processor 3210_1 may be configured to access the memory 3220_1 to execute various instructions or process data. In one embodiment, the memory 3220_1 may include at least one of various memory devices such as DDR SDRAM (Double Data Rate Synchronous DRAM), HBM (High Bandwidth Memory), HMC (Hybrid Memory Cube), DIMM (Dual In-line Memory Module), Optane DIMM, and NVDIMM (Non-Volatile DIMM).
[0196] In one embodiment, the number of processors 3210_1 and the number of memories 3220_1 included in the first storage server 3200_1 can be changed or modified differently. In one embodiment, the processors 3210_1 and memories 3220_1 included in the first storage server 3200_1 can form processor-memory pairs, and the number of processor-memory pairs included in the first storage server 3200_1 can be changed or modified differently. In one embodiment, the number of processors 3210_1 and the number of memories 3220_1 included in the first storage server 3200_1 can be different. The processor 3210_1 can include a single-core processor or a multi-core processor.
[0197] Under the control of processor 3210_1, switch 3230_1 can selectively connect processor 3210_1 and storage device 3250_1, or selectively connect NIC 3240_1, storage device 3250_1 and CXL interface controller 3260_1.
[0198] NIC 3240_1 connects the first storage server 3200_1 to the network NT. NIC 3240_1 may include a network interface card, network adapter, etc. NIC 3240_1 can connect to the network NT via a wired interface, wireless interface, Bluetooth interface, or optical interface. NIC 3240_1 may include internal memory, a DSP (Digital Signal Processor), a host bus interface, etc., and can connect to processor 3210_1 or switch 3230_1 via the host bus interface. The host bus interface may include at least one of the following interfaces: ATA (Advanced Technology Attachment) interface, SATA (Serial ATA) interface, e-SATA (External SATA) interface, SCSI (Small Computer System Interface) interface, SAS (Serial Attached SCSI) interface, PCI (Peripheral Component Interconnect) interface, PCIe (PCI Fast) interface, NVMe (NVM Fast) interface, IEEE 1394 interface, USB (Universal Serial Bus) interface, SD (Secure Digital) card interface, MMC (Multimedia Card) interface, eMMC (Embedded Multimedia Card) interface, UFS (Universal Flash Memory) interface, eUFS (Embedded Universal Flash Memory) interface, and CF (Compact Flash Memory) card interface. In one embodiment, NIC3240_1 may be integrated with at least one of processor 3210_1, switch 3230_1, and storage device 3250_1.
[0199] Under the control of processor 3210_1, storage device 3250_1 can store data or output stored data. Storage device 3250_1 may include controller (CTRL) 3251_1, non-volatile memory 3252_1 (e.g., NAND flash memory), DRAM 3253_1, and interface (I / F) 3254_1. In one embodiment, storage device 3250_1 may also include a security element (SE) for security or privacy.
[0200] Controller 3251_1 can control the overall operation of storage device 3250_1. In one embodiment, controller 3251_1 may include SRAM. In response to signals received through interface 3254_1, controller 3251_1 may store data in non-volatile memory 3252_1, or may output data stored in non-volatile memory 3252_1. In one embodiment, controller 3251_1 may be configured to control non-volatile memory 3252_1 based on a switching interface or ONFI (Open NAND Flash Interface).
[0201] DRAM 3253_1 can be configured to temporarily store data to be stored in or read from non-volatile memory 3252_1. DRAM 3253_1 can be configured to store various data (e.g., metadata and mapping data) required for the operation of storage controller 3251_1. Interface 3254_1 can provide physical connectivity between controller 3251_1 and processor 3210_1, switch 3230_1, or NIC 3240_1. In one embodiment, the interface can be implemented to support DAS (Direct-Attached Storage) mode, allowing storage device 3250_1 to be directly connected via a dedicated cable. In one embodiment, interface 3254_1 can be implemented via a host interface bus based on at least one of the various interfaces described above.
[0202] The above-described components of the first storage server 3200_1 are provided as examples, and this disclosure is not limited thereto. The above-described components of the first storage server 3200_1 can be applied to each of the remaining storage servers 3200_2 to 3200_m or each of the application servers 3100_1 to 3100_n. In one embodiment, each of the storage devices 3150_1 to 3150_n of the application servers 3100_1 to 3100_n can be selectively omitted.
[0203] Multiple application servers 3100_1 to 3100_n and multiple storage servers 3200_1 to 3200_m can communicate with each other via a network NT. The network NT can be implemented using Fibre Channel (FC), Ethernet, etc. In this case, FC can be the medium for high-speed data transmission, and a high-performance / high-availability optical switch can be used. Depending on the access method of the network NT, the storage servers 3200_1 to 3200_m can be configured as file storage devices, block storage devices, or object storage devices.
[0204] In one embodiment, the network NT can be a storage-specific network (such as a storage area network (SAN)). For example, the SAN can be an FC-SAN implemented using an FC network and according to the FC protocol (FCP). Alternatively, the SAN can be an IP-SAN implemented using a TCP / IP network and according to the iSCSI (or SCSI over TCP / IP, or Internet SCSI) protocol. In one embodiment, the network NT can be a general network (such as a TCP / IP network). For example, the network NT can be implemented according to protocols such as FCoE (FC over Ethernet), NAS (Network Attached Storage), or NVMe-oF (NVMe over Fibre).
[0205] In one embodiment, at least one of the plurality of application servers 3100_1 to 3100_n can be configured to access at least one of the remaining application servers 3200_1 to 3200_m via network NT.
[0206] For example, the first application server 3100_1 can store data requested by a user or client in at least one of multiple storage servers 3200_1 to 3200_m via the network NT. Optionally, the first application server 3100_1 can obtain the data requested by a user or client from at least one of the multiple storage servers 3200_1 to 3200_m via the network NT. In this case, the first application server 3100_1 can be implemented using a web server, a database management system (DBMS), etc.
[0207] In other words, the processor 3110_1 of the first application server 3100_1 can access the memory (e.g., 3120_n) or storage device (e.g., 3150_n) of another application server (e.g., 3100_n) via the network NT. Optionally, the processor 3110_1 of the first application server 3100_1 can access the memory 3220_1 or storage device 3250_1 of the first storage server 3200_1 via the network NT. Thus, the first application server 3100_1 can perform various operations on data stored in the remaining application servers 3100_2 to 3100_n or the multiple storage servers 3200_1 to 3200_m. For example, the first application server 3100_1 can execute or issue instructions for moving or copying data between the remaining application servers 3100_2 to 3100_n or between the multiple storage servers 3200_1 to 3200_m. In this scenario, data destined for movement or copying can be moved from storage devices 3250_1 to 3250_m of storage servers 3200_1 to 3200_m to storage devices 3220_1 to 3220_m of application servers 3100_1 to 3100_n via storage devices 3220_1 to 3220_m, or directly moved from storage devices 3250_1 to 3250_m of storage servers 3200_1 to 3200_m to storage devices 3120_1 to 3120_n of application servers 3100_1 to 3100_n. Data transmitted over the network NT can be encrypted for security or privacy purposes.
[0208] In one embodiment, multiple storage servers 3200_1 to 3200_m and multiple application servers 3100_1 to 3100_n can be connected to a memory expander 3300 via a CXL interface. The memory expander 3300 can serve as extended storage for each of the multiple storage servers 3200_1 to 3200_m and multiple application servers 3100_1 to 3100_n. The multiple storage servers 3200_1 to 3200_m and multiple application servers 3100_1 to 3100_n can be based on a reference... Figures 1 to 18B The described method allows communication between the CXL interface and the memory expander 3300.
[0209] Figure 22 This is a diagram illustrating an example (e.g., a CXL interface) of a heterogeneous computing interface applied to embodiments of this disclosure. Figure 22The heterogeneous computing interface connected to a memory expander according to embodiments of the present disclosure will be described with reference to the CXL interface, but the present disclosure is not limited thereto. For example, the heterogeneous computing interface can be implemented by at least one of a variety of computing interfaces based on protocols such as Gen-Z protocol, NVLink protocol, CCIX protocol and open CAPI protocol.
[0210] Reference Figure 22 The heterogeneous computing system 4000 may include multiple CPUs 4100 and 4200, multiple memories 4110 and 4210, accelerators 4120 and 4220, optional memory 4130 and 4230, and a memory expander 4300. Each of the multiple CPUs 4100 and 4200 may be a processor configured to perform various operations / operations / computations. The multiple CPUs 4100 and 4200 may communicate with each other via separate links. In one embodiment, the separate links may include coherence links between CPUs.
[0211] Multiple CPUs 4100 and 4200 can communicate with multiple memories 4110 and 4210, respectively. For example, a first CPU 4100 can communicate directly with a first memory 4110, and a second CPU 4200 can communicate directly with a second memory 4210. Each of the first memory 4110 and the second memory 4210 may include DDR memory. In one embodiment, virtual memory allocated to different virtual machines according to embodiments of this disclosure may be memory allocated from DDR memories 4110 and 4210.
[0212] Multiple CPUs 4100 and 4200 can communicate with accelerators 4120 and 4220 via a flex bus. Accelerators 4120 and 4220 can be calculators or processors that perform operations independently of the multiple CPUs 4100 and 4200. Accelerator 4120 can operate under the control of its corresponding CPU 4100, and accelerator 4220 can operate under the control of its corresponding CPU 4200. Accelerators 4120 and 4220 can be connected to optional memories 4130 and 4230, respectively. In one embodiment, the multiple CPUs 4100 and 4200 can be configured to access optional memories 4130 and 4230 via the flex bus and accelerators 4120 and 4220.
[0213] Multiple CPUs 4100 and 4200 can communicate with the memory expander 4300 via the flex bus. Multiple CPUs 4100 and 4200 can use the memory space of the memory expander 4300.
[0214] In one embodiment, the flex bus can be a bus or port configured to select either the PCIe or CXL protocol. That is, the flex bus can be configured to select either the PCIe or CXL protocol based on the characteristics or communication type of the device connected to it. In one embodiment, the memory expander 4300 can be like... Figures 1 to 18B It operates in the same way as the described memory expander and can communicate with multiple CPUs 4100 and 4200 based on the CXL protocol.
[0215] In one embodiment, the communication structure based on the flex bus is... Figure 22 The components are shown to be independent of each other, but this disclosure is not limited thereto. For example, Figure 22 CXL communication between the various components shown can be performed via the same bus or the same link.
[0216] According to this disclosure, a memory expander connected to a CPU and multiple accelerators via a heterogeneous computing interface can be configured to manage data to be provided to or received from multiple accelerators. This reduces the CPU's burden associated with data management. Therefore, a memory expander with improved performance, a heterogeneous computing device using the memory expander, and a method of operating the heterogeneous computing device are provided.
[0217] As is conventional in the art, embodiments can be described and illustrated in the form of blocks that perform one or more described functions. These blocks (which may be referred to herein as cells or modules, etc.) are physically implemented using analog and / or digital circuitry (such as logic gates), integrated circuits, microprocessors, microcontrollers, memory circuitry, passive electronic components, active electronic components, optical components, hardwired circuitry, etc., and may optionally be driven by firmware and / or software. For example, the circuitry may be implemented in one or more semiconductor chips, or on a substrate support (such as a printed circuit board, etc.). The circuitry constituting a block may be implemented using dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware for performing some functions of the block and a processor for performing other functions of the block. Without departing from the scope of disclosure, each block of an embodiment may be physically separated into two or more interactive and independent blocks. Similarly, without departing from the scope of disclosure, the blocks of an embodiment may be physically combined into more complex blocks. One aspect of an embodiment may be implemented by instructions stored in a non-transitory memory medium and executed by a processor.
[0218] Although this disclosure has been described with reference to embodiments thereof, it will be apparent to those skilled in the art that various changes and modifications may be made thereto without departing from the spirit and scope of this disclosure as set forth in the appended claims.
Claims
1. A memory expander, comprising: A memory device configured to store multiple task data; as well as A controller is configured to control a memory device, wherein the controller is configured to: receive metadata and management requests from an external central processing unit via a compute fast link interface; operate in management mode in response to a management request; and in management mode, receive a read request and a first address from an accelerator via a compute fast link CXL interface, and in response to the read request, send one of the plurality of task data to the accelerator based on the metadata. The controller is also configured, in management mode, to send completion data to the central processing unit via the CXL interface when multiple result data associated with the multiple task data are stored in the memory device. The controller is also configured to, in management mode: Receive status requests from the central processing unit via the CXL interface; and In response to a status request, information about the read counters associated with the plurality of task data and information about the write counters associated with the plurality of result data are sent to the central processing unit.
2. The memory expander according to claim 1, wherein, The management request is a master-to-slave request M2S RwD message with data based on the CXL protocol.
3. The memory expander according to claim 2, wherein: The management request includes a first CXL header, and The first CXL header includes: The memory opcode field indicates the activation of management mode; The first address field indicates the starting memory address of the plurality of task data; The second address field indicates the end memory address of the multiple task data; The label field indicates the task number associated with the plurality of task data; and The location field indicates the storage address of the metadata.
4. The memory expander according to claim 1, wherein, The controller is also configured to, in management mode: Receive write requests and first result data from the accelerator; and In response to a write request, the first result data is stored in the memory device based on metadata.
5. The memory expander according to claim 1, wherein, Metadata includes at least one of the following: The task number is associated with the task corresponding to the plurality of task data; Data management mode, indicating the management method associated with the data of the multiple tasks; The starting memory address of the multiple task data; The end memory address of the multiple task data; The unit size is associated with the multiple task data; A separator is used to distinguish the multiple task data. Accelerator identifier, corresponding to the mission number; The starting memory address of multiple result data, which are associated with the multiple task data; as well as The end memory address of the multiple result data.
6. The memory expander according to claim 1, wherein, The completion is based on the CXL protocol's slave-to-master data response S2M DRS message.
7. The memory expander according to claim 6, wherein: The completion includes a second CXL header, and The second CXL header includes: The memory opcode field indicates the type of calculation result for the multiple task data; The label field indicates the task number associated with the multiple task data; The location field indicates the memory address of the plurality of result data; and Reserved fields include statistical information about the data from the multiple tasks.
8. The memory expander according to claim 1, wherein, Status requests are master-to-slave request M2S Req messages based on the CXL protocol.
9. The memory expander according to claim 8, wherein: The status request includes a third CXL header, and The third CXL header includes: The memory opcode field indicates how to query the processing status associated with the multiple task data; The label field indicates the task number associated with the multiple task data; The address field indicates the range of task data to be queried from the plurality of task data; and Reserved fields indicate the query unit or query time of the processing status.
10. A heterogeneous computing device, comprising: CPU; The memory is configured to store data under the control of the central processing unit; The accelerator is configured to repeatedly perform computations on multiple task data and generate multiple result data. as well as The memory expander is configured to operate in management mode in response to management requests from the central processing unit, and in management mode, manage the plurality of task data to be provided to the accelerator and the plurality of result data provided from the accelerator, wherein... The central processing unit, accelerator, and memory expander communicate with each other through a heterogeneous computing interface. The memory expander is configured to, upon receiving all the multiple result data from the accelerator, transmit the completed data to the central processing unit via a heterogeneous computing interface. The memory expander is also configured to, in management mode: Receive status requests from the central processing unit via the CXL interface; and In response to a status request, information about the read counters associated with the plurality of task data and information about the write counters associated with the plurality of result data are sent to the central processing unit.
11. The heterogeneous computing device according to claim 10, wherein, The heterogeneous computing interface is an interface based on the Compute Fast Link (CXL) protocol, the Gen-Z protocol, the NVlink protocol, the Cache Coherent Interconnect (CCIX) protocol for accelerators, and the Open Coherent Accelerator Processor Interface (CAPI) protocol.
12. The heterogeneous computing device according to claim 10, wherein, The memory expander is also configured to: Receive metadata from the central processing unit; and The multiple task data and multiple result data are managed based on metadata.
13. The heterogeneous computing device according to any one of claims 10 to 12, wherein, When the accelerator performs the calculations on the multiple task data respectively, the accelerator does not generate an interrupt to the central processing unit.
14. A method of operating a heterogeneous computing device, the heterogeneous computing device comprising a central processing unit, an accelerator, and a memory expander connected via a compute fast link (CXL) interface, the method comprising: The central processing unit sends metadata to the memory expander; The central processing unit sends management requests to the memory expander; The central processing unit sends the task request to the accelerator; The accelerator sends a read request to the memory expander in response to a task request; The memory expander, in response to a read request, sends the first task data from a set of multiple task data to the accelerator based on metadata. The accelerator performs a first computation on the first task data to generate the first result data; The accelerator sends the write request and the first result data to the memory expander. The memory expander stores the first result data in response to a write request; as well as When the memory expander receives all the result data from the accelerator, it sends the completed data to the central processing unit via the CXL interface. The method further includes: The memory expander receives status requests from the central processing unit via the CXL interface; and In response to a status request, the memory expander sends information about read counters associated with the plurality of task data and write counters associated with the plurality of result data to the central processing unit.
15. The method according to claim 14, wherein, The accelerator does not generate an interrupt to the central processing unit after performing the first calculation.
16. The method of claim 14, further comprising, after the first result data is stored in the memory expander: The accelerator sends read requests to the memory expander; and The memory expander sends the second task data from multiple task data sets to the accelerator based on metadata in response to a read request.
17. The method of claim 16, further comprising: The accelerator performs a second computation on the second task data to generate second result data; The accelerator sends the write request and the second result data to the memory expander. as well as The memory expander stores the second result data based on metadata in response to a write request, wherein, The first result data and the second result data are stored in different memory areas.
Citation Information
Patent Citations
Core drill device with changing operated mode
KR1020200141710A
Storage device and method thereof and controller
CN111198839A
Flexible on-die fabric interface
US20200327084A1