Computing system for reducing latency between serially connected electronic devices
By introducing a doorbell parser and intermediate buffer into the computing system and optimizing the switch design, the communication delay problem between serially connected electronic devices is solved, and the communication speed and data transmission efficiency are improved.
Patent Information
- Application Number
- CN202010363031.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-14
- Filing Date
- 2020-04-30
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-04-30
AI Technical Summary
In computing systems, communication delays between serially connected electronic devices can significantly reduce the communication speed between high-speed electronic devices and a host computer.
By introducing a doorbell parser and an intermediate buffer into the computing system, the design of the switch is optimized, the switch delay is reduced, the parallel processing of doorbells and commands is achieved, the memory buffer is directly accessed, and the communication delay is reduced.
The communication delay between serially connected electronic devices is effectively reduced, and the communication speed is improved, especially the data transmission efficiency between high-speed electronic devices and the host.
Smart Images

Figure CN112395235B_ABST
Abstract
Description
[0001] This application claims priority from Korean Patent Application No. 10-2019-0099851 filed on August 14, 2019, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] Embodiments of the inventive concepts described herein relate to a computing system, and more particularly, a computing system for reducing delay between serially connected electronic devices. Background Art
[0003] In a computing system, multiple electronic devices are connected to each other for communication. Multiple electronic devices can be connected serially or sequentially. Serially connected electronic devices can communicate with the host computer of the computing system.
[0004] An electronic device corresponding to an endpoint device or terminal device among multiple electronic devices can communicate with a host through at least one or more serially connected electronic devices. The communication speed between the endpoint device and the host can be slowed by the delay of at least one or more intermediate electronic devices inserted between the endpoint device and the host. In particular, in configurations where the endpoint device is a high-speed electronic device, the communication speed between the high-speed electronic device and the host via the intermediate electronic device can be significantly reduced. Summary of the Invention
[0005] Embodiments of the inventive concept provide a computing system for reducing delay between serially connected electronic devices.
[0006] According to an embodiment, a computing system is provided, comprising: a host; a first electronic device coupled to the host; and a second electronic device coupled to the first electronic device, the second electronic device being configured to communicate with the host through the first electronic device. The first electronic device is configured to request the host to transmit a write command written to a submission queue of the host based on a doorbell signal received from the host, store the write command received from the host, request the host to transmit write data of the write command stored in a data buffer of the host, and store the write data received from the host.
[0007] According to an embodiment, a computing system is provided, comprising: a host; a first electronic device coupled to the host; and a second electronic device coupled to the first electronic device, the second electronic device being configured to communicate with the host through the first electronic device. The first electronic device is configured to receive a write command for the second electronic device from the host and transmit the write command to the second electronic device, receive a doorbell associated with a submission queue in which the write command is written from the host, transmit the doorbell to the second electronic device, request write data of the write command stored in a data buffer of the host, and store the write data received from the host.
[0008] According to an embodiment, a computing system is provided, comprising: a host; a first electronic device coupled to the host, the first electronic device including a submission queue controller memory buffer (CMB) and a write CMB; and a second electronic device coupled to the first electronic device, the second electronic device configured to communicate with the host via the first electronic device. The first electronic device is configured to receive a write command from the host written to a submission queue of the host and store the write command in the submission queue CMB, receive write data of the write command stored in a data buffer of the host from the host, store the write data in the write CMB, receive a doorbell regarding the submission queue sent from the host, and send the doorbell to the second electronic device. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other objects and features of the inventive concept will become apparent by describing in detail embodiments of the inventive concept with reference to the accompanying drawings.
[0010] Figure 1 shows a block diagram of a computing system according to an embodiment;
[0011] Figure 2 Shown Figure 1 Detailed block diagram of the computing system;
[0012] Figure 3 Shown Figure 1 A block diagram of a computing system;
[0013] Figure 4 Shown Figure 3 Timing diagram of the computing system;
[0014] Figure 5 Shown Figure 1 A block diagram of a computing system;
[0015] Figure 6 Shown Figure 5 Timing diagram of the computing system;
[0016] Figure 7 Shows the operation Figure 5Methods of computing systems;
[0017] Figure 8 Shown Figure 1 A block diagram of a computing system;
[0018] Figure 9 Shown Figure 8 Timing diagram of the computing system;
[0019] Figure 10 Shows the operation Figure 8 Methods of computing systems;
[0020] Figure 11 Shown Figure 1 A block diagram of a computing system;
[0021] Figure 12 Shown Figure 11 Timing diagram of the computing system;
[0022] Figure 13 Shows the operation Figure 11 Methods of computing systems;
[0023] Figure 14 Shown Figure 1 A block diagram of a computing system;
[0024] Figure 15 Shown Figure 14 Timing diagram of the computing system;
[0025] Figure 16 Shown Figure 1 A block diagram of a computing system;
[0026] Figure 17 Shown Figure 16 Timing diagram of the computing system;
[0027] Figure 18 Shows the operation Figure 16 Methods of computing systems;
[0028] Figure 19 A block diagram showing a storage device according to an embodiment; and
[0029] Figure 20 A block diagram of a computing device according to an embodiment is shown. DETAILED DESCRIPTION
[0030] Figure 1 A block diagram of a computing system according to an embodiment is shown.
[0031] like Figure 1As shown in FIG, the computing system 10 may include a host 11, a first electronic device 12, and a second electronic device 13. The host 11, the first electronic device 12, and the second electronic device 13 of the computing system 10 may be any of various electronic devices (such as desktop computers, laptop computers, tablet computers, video game consoles, workstations, servers, computing devices, and electric vehicles), or may be any of various electronic devices located on a motherboard of an electronic device.
[0032] The host 11 may be implemented as a system on chip (SoC), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). The control circuit of the host 11 may also include a general-purpose processor, a central processing unit, a dedicated processor, or an application processor. The host 11 may be implemented as a processor itself, or the host 11 may be an electronic device or system including a processor. The host 11 may control all communications of the computing system 10, including communications between the host 11, the first electronic device 12, and the second electronic device 13.
[0033] The first electronic device 12 may be directly or serially (or sequentially) connected to the host 11. The second electronic device 13 may be directly or serially connected to the first electronic device 12. Therefore, the host 11, the first electronic device 12, and the second electronic device 13 may be connected in a sequential manner. In other words, the second electronic device 13 may be connected to the host 11 through the first electronic device 12. For example, the host 11 may communicate directly with the first electronic device 12, and may communicate with the second electronic device 13 through the first electronic device 12. Therefore, the second electronic device 13 may correspond to an endpoint device, and the first electronic device 12 may be an intermediate device that connects the endpoint electronic device (i.e., the second electronic device) 13 to the host 11.
[0034] and Figure 1 Unlike the example shown in FIG, an additional electronic device may be further connected in the computing system 10 anywhere between the host 11 and the first electronic device 12 and between the first electronic device 12 and the second electronic device 13. Of course, the additional electronic device may be connected only to the second electronic device 13, so that the additional electronic device may be an endpoint device of the computing system 10.
[0035] The first electronic device 12 and the second electronic device 13 may be implemented as the same electronic device or different electronic devices. For example, the first electronic device 12 may correspond to a switch or a switch device that connects the second electronic device 13 to the host 11, and the second electronic device 13 may correspond to an endpoint device different from the switch device.
[0036] Figure 2 Shown Figure 1 Detailed block diagram of the computing system. Figure 1The computing system 100 corresponding to the computing system 10 may include a processor 110, a root complex 120, a memory 130, and electronic devices 141, 142, 151 to 154, and 161 to 163. Here, the electronic devices may also be referred to as "input / output (I / O) devices", and the electronic devices 141, 142, 151 to 154, and 161 to 163 may be connected to the processor 110, the root complex 120, the memory 130, and the electronic devices 141, 142, 151 to 154, and 161 to 163. Figure 1 The first electronic device 12 and the second electronic device 13 shown in FIG. 1 correspond to each other. The processor 110, the root complex 120 and the memory 130 may be Figure 1 Components of the host 11 are shown in FIG.
[0037] The processor 110 may perform various types of arithmetic or logical operations. For example, the processor 110 may include an internal cache memory and at least one or more cores (e.g., homogeneous multi-core or heterogeneous multi-core) that control operations. The processor 110 may execute program code, software, applications, etc. loaded from the memory 130.
[0038] Root complex 120 can mediate communications between processor 110, memory 130, and electronic devices 141, 142, 151 to 154, and 161 to 163. For example, root complex 120 can manage the timing, sequence, and environment of communications between processor 110, memory 130, and electronic devices 141, 142, 151 to 154, and 161 to 163. Root complex 120 can be implemented in hardware, software, or a combination of hardware and software, and can be implemented on the motherboard of computing system 100. Root complex 120 can be the root of an I / O hierarchy that communicatively connects processor 110 and memory 130 to electronic devices 141, 142, 151 to 154, and 161 to 163. Root complex 120 can include one or more downstream ports (DPs). Electronic devices 141 and 142 can be connected to the downstream ports (DPs) of root complex 120. The electronic devices 151 to 154 may be connected to the downstream port DP of the root complex 120. In addition, the electronic devices 161 to 163 may be connected to the downstream port DP of the root complex 120. The number of the downstream ports DP is not limited to Figure 2 The number of electronic devices connected to each downstream port DP may be one or more.
[0039] In one embodiment, communication between the root complex 120 and the electronic devices 141, 142, 151 to 154 and 161 to 163 can be performed in accordance with various communication interface protocols (such as the Peripheral Component Interconnect Express (PCIe) protocol, the Mobile PCIe (M-PCIe) protocol, the Non-Volatile Memory Express (NVMe) protocol, the Universal Serial Bus (USB) protocol, the Small Computer System Interface (SCSI) protocol, the Advanced Technology Attachment (ATA) protocol, the Parallel ATA (PATA), the Serial ATA (SATA), the Serial Attached SCSI (SAS) protocol, the Integrated Drive Electronics (IDE) protocol, the Universal Flash Storage (UFS) protocol, and the FireWire protocol).
[0040] The memory 130 may store data used for the operation of the computing system 100. The memory 130 may store data processed by the processor 110 or to be processed by the processor 110. For example, the memory 130 may include a volatile memory (such as a static random access memory (SRAM) or a dynamic RAM (DRAM)) or a non-volatile memory (NVM). Applications, file systems, or device drivers that can be executed by the processor 110 may be loaded onto the memory 130. The programs and software layers loaded onto the memory 130 may be executed under the control of the processor 110, and the information loaded into the memory 130 is not limited to Figure 2 The example shown in FIG. Memory 130 may include a host memory buffer (HMB). A portion of the entire storage area of memory 130 may be allocated to the host memory buffer HMB.
[0041] The processor 110 can communicate with Figure 1 The processor 110 and the root complex 120 can be connected to Figure 1 Furthermore, the processor 110, the root complex 120 and the memory 130 may all be connected to Figure 1 The processor 110, the root complex 120, and the memory 130 may be implemented as a system on a chip (SoC) to constitute the host 11. Alternatively, the processor 110, the root complex 120, and the memory 130 may be implemented as one or more separate components to constitute the host 11.
[0042] Each of the electronic devices 142, 152, 154, and 163 may be configured as an endpoint device. Each of the electronic devices 142, 152, 154, and 163 may include an endpoint port EP. Each of the remaining electronic devices 141, 151, 153, 161, and 162 may correspond to an intermediate device. An intermediate device may be connected to an endpoint device, another intermediate device, or the root complex 120. Each of the electronic devices 141, 151, 153, 161, and 162 may include an upstream port UP and a downstream port DP. For example, the upstream port UP of the electronic devices 141, 151, 153, 161, and 162 may be provided on the upstream side of the electronic devices 141, 151, 153, 161, and 162 toward the root complex 120. The downstream ports DP of the electronic devices 141, 151, 153, 161, and 162 may be provided on the downstream side of the electronic devices 141, 151, 153, 161, and 162 toward the endpoint. The endpoint port EP of the endpoint device may be connected to the downstream port DP of the intermediate device or the root complex 120. The endpoint port EP may also be referred to as an "upstream port UP." Figure 2 In the configuration of Figure 1 Each of the electronic devices 142, 152, 154, and 163 may correspond to the first electronic device 12. Figure 1 corresponding to the second electronic device 13.
[0043] In one embodiment, Figure 1 The electronic devices 141, 151, 153, 161 and 162 corresponding to the first electronic device 12 may be a PCIe switch, a PCIe device, an NVMe device, a storage device or a solid-state drive (SSD). Figure 1 The electronic devices 142, 152, 154, and 163 corresponding to the second electronic device 13 can be PCIe switches, PCIe devices, NVMe switches, NVMe devices, storage devices, or SSDs. As described above, any other endpoint devices connected to the electronic devices 142, 152, 154, and 163 can also be included in the computing system 100.
[0044] Figure 3 Shown Figure 1 Block diagram of a computing system. Figure 4 Shown Figure 3 Timing diagram of the computing system.
[0045] like Figure 3 As shown in FIG, the computing system 200 may include a host 210, a switch 220, and a storage device 230. For example, the computing system 200, the host 210, the switch 220, and the storage device 230 may be respectively connected to Figure 1The computing system 10, the host 11, the first electronic device 12 and the second electronic device 13 correspond to each other.
[0046] The host 210 may include a submission queue (SQ) 211, a completion queue (CQ) 212, and a data buffer 213. The submission queue 211, the completion queue 212, and the data buffer 213 may be implemented in Figure 2 The host 210 may perform input / output (I / O) operations on the storage device 230 through the switch 220 based on the submission queue 211 and the completion queue 212 .
[0047] The switch 220 may be provided between the host 210 and the storage device 230 and may transmit electrical signals from the host 210 (or the storage device 230) to the storage device 230 (or the host 210). Thus, the switch 220 acts as an intermediary device between the host 210 and the storage device 230. The storage device 230 may receive commands from the host 210, process the received commands, and transmit the processing results to the host 210 via the switch 220. The storage device 230 may include a non-volatile memory 239 and a controller 231 that controls the non-volatile memory 239. The non-volatile memory 239 may include NAND flash memory, NOR flash memory, phase change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), ferroelectric RAM (FeRAM), and the like.
[0048] The host 210 may input or write a command for the storage device 230 in an entry (or slot) of the submission queue 211 and may update a tail pointer of the submission queue 211 (i.e., a doorbell of the submission queue 211). For example, the doorbell may include an address indicating the submission queue 211. A pair of the submission queue 211 and the completion queue 212 may be provided to each core of the processor 110, and the number of the pair may be one or more. Figure 3 and Figure 4, the host 210 may submit or send the doorbell of the submission queue 211 to the storage device 230 through the switch 220 (①). When the switch delay of the switch 220 has passed after the doorbell is sent from the host 210 (①), the storage device 230 may receive the doorbell from the switch 220 (①). The delay may be referred to as "time". For example, the entire (total) delay of the switch 220 may be divided into a host-side switch delay and a storage device-side switch delay. In detail, the switch 220 may include one or more transmission circuits for sending the doorbell and a transmission path including a physical path in the switch 220. For example, the time it takes for the doorbell to pass through a portion of the transmission path of the switch 220 may correspond to the host-side switch delay, and the time it takes for the doorbell to pass through the remaining portion of the transmission path of the switch 220 may correspond to the storage device-side switch delay.
[0049] The host 210 can update the doorbell register (SQTDBL) 232 of the controller 231 of the storage device 230 by sending a doorbell (①). The storage device 230 can read the doorbell stored in the doorbell register 232 and can confirm (verify) that the command is input or written into the submission queue 211 by the host 210. The storage device 230 can extract or read the command of the submission queue 211 of the host 210 based on the doorbell (②, ③). For example, the storage device 230 can request a command of the submission queue 211 of the host 210 and can send the request to the host 211 through the switch 220 (②). Then, the storage device 230 can read or receive the command of the submission queue 211 from the host 210 through the switch 220 (③). When the switch delay has passed after the request for the command is sent from the storage device 230 to the switch 220 (②), the request can reach the host 210 (②). Furthermore, after the command is sent from host 210 to switch 220 (③), the switch delay elapses, and the command in submission queue 211 can reach storage device 230 (③). As in the case of doorbell transmission, switch delays occur in each of sending a request for a command (②) and sending a command (③).
[0050] The controller 231 of the storage device 230 may include a direct memory access (DMA) engine 233 configured to send requests or data directly to the switch 220 or the host 210. The DMA engine 233 may receive commands from the switch 220 and may store the received commands in a submission queue buffer 234 of the controller 231.
[0051] The storage device 230 can process the command stored in the submission queue buffer 234. For example, the command can be a write command. The storage device 230 can decode the command and can read the write data stored in the data buffer 213 of the host 210 based on the command (e.g., the decoded command) (④, ⑤). For example, the storage device 230 can request write data from the data buffer 213 of the host 210 and can send the request to the host 210 through the switch 220 (④). Then, the storage device 230 can receive the write data from the data buffer 213 from the host 210 through the switch 220 (⑤). When the switch delay has elapsed after the request for write data is sent from the storage device 230 to the switch 220 (④), the request can arrive at the host 210 (④). In addition, when the switch delay has elapsed after the write data is sent from the data buffer 213 to the switch 220 (⑤), the write data can arrive at the storage device 230 (⑤). As in the transmission of the doorbell, switch delays occur in each of the cases of sending the request for write data (④) and sending the write data (⑤).
[0052] The DMA engine 233 may request write data, may receive the write data, and may store the received write data in the write buffer 235 of the controller 231. Although the above operation has been described in the context of requesting and sending write data under the assumption that the command is a write command, when the command is not a write command (for example, a read command), sending of the write data may not be performed.
[0053] The storage device 230 can process the command of the host 210 and can send completion information about the command to the host 210 through the switch 220 (⑥). For example, the completion information may include whether the command was successfully completed or unsuccessfully completed, the result of processing the command, etc. When the switch delay has elapsed after the completion information is sent from the storage device 230 to the switch 220 (⑥), the completion information can arrive at the host 210 (⑥). The completion information can be stored or written in the completion queue 212 of the host 210. When the switch delay has elapsed after the completion information is sent from the storage device 230 to the switch 220 (⑥), the completion information can arrive at the host 210 (⑥). As in the transmission of the doorbell, switch delay may occur even when the completion information (⑥) is sent.
[0054] Taking into account switch delays caused by the switch 220 between the host 210 and the storage device 230, the embodiments described herein may provide a plurality of computing systems 300 to 700 and a computing device 1000, which are used to reduce the time taken to process commands generated (or issued) from the host 210, or the time taken for commands generated from the host 210 or write data to reach an endpoint device (i.e., the storage device 230).
[0055] Figure 5 Shown Figure 1 Block diagram of a computing system. Figure 6 Shown Figure 5 Timing diagram of the computing system. Figure 7 Shows the operation Figure 5 Methods of computing systems.
[0056] like Figure 5 As shown in FIG, computing system 300 may include host 310, switch 320, and storage device 330. Host 310 may include submission queue 311, completion queue 312, and data buffer 313. Components 311 to 313 of host 310 may be similar to components 211 to 213 of host 210, respectively. Figure 3 Compared with the switch 220, Figure 5 The switch 320 may include a doorbell parser 321, a command parser 322, an intermediate submission queue buffer 323, and an intermediate write buffer 324. The storage device 330 may include a controller 331 and a non-volatile memory 339. The controller 331 may include a doorbell register 332, a DMA engine 333, a submission queue buffer 334, and a write buffer 335. Figure 5 The components 331 to 335 and 339 of the storage device 330 may be respectively Figure 3 The components 231 to 235 and 239 of the storage device 230 are similar. Redundant descriptions of the above components will be omitted, and the following description will focus on Figure 5 The computing system 300 and Figure 3 The differences between the computing systems 200 are discussed.
[0057] Combined reference Figure 5 、 Figure 6 and Figure 7In operation S303, the host 310 may write or store a command for the storage device 330 in the submission queue 311. In operation S306, the host 310 may send the doorbell of the submission queue 311 to the switch 320 (①), and the switch 320 may receive the doorbell (①). In operation S309, the switch 320 may send (forward, relay, resend, etc.) the doorbell to the storage device 330 (①). In operation S313, the switch 320 may read or parse the doorbell. The doorbell parser 321 of the switch 320 may read or parse the doorbell. In operation S316, based on the doorbell, the doorbell parser 321 may request a command for the submission queue 311 of the host 310 and may send the request to the host 310 (①). For example, the doorbell parser 321 may access the submission queue 311 indicated by the address of the doorbell (corresponding to the address of the doorbell) among one or more submission queues of the host 310. In operation S319 , the controller 331 of the storage device 330 may read or parse the doorbell of the doorbell register 332 .
[0058] Because the switch 320 can parse the doorbell in operation S313 after sending the doorbell to the storage device 330 in operation S309, the switch 320 and the storage device 330 can parse the doorbell simultaneously. Alternatively, upon receiving the doorbell from the host 310 in operation S306, the switch 320 can immediately begin parsing the doorbell in operation S313 before sending the doorbell to the storage device 330 in operation S309. In this regard, the switch 320 can generate a copy of the doorbell to parse the doorbell and send the doorbell to the storage device 330, or perform any other operation to complete parsing the doorbell before the storage device 330 completes parsing the doorbell. As a result, the communication delay between the host 310 and the storage device 330 caused by the presence of the switch 320 can be reduced.
[0059] In operation S323, the storage device 330 may request a command from the submission queue 311 of the host 310 and may send the request to the switch 320 (②). For example, operations S313 and S316 are similar to operations S319 and S323, respectively. In operation S326, the switch 320 may receive the command from the submission queue 311 of the host 310 (③). The switch 320's receipt of the command from the submission queue 311 of the host 310 may be in response to the switch 320's command request in operation S316, thereby preemptively avoiding the delay of the storage device 330's command request in operation S323. The intermediate submission queue buffer 323 of the switch 320 may store the received command. In operation S329, the switch 320 may send the command from the intermediate submission queue buffer 323 to the storage device 330 in response to the request in operation S323 (③). In a case where a command of the intermediate submission queue buffer 323 is received through the switch 320 in operation S326 before a command request is received from the memory device 330 in operation S323, the switch 320 may then buffer the command of the intermediate submission queue buffer 323 until a command request is received from the memory device 330. The DMA engine 333 of the controller 331 may receive the command and may store the received command in the submission queue buffer 334.
[0060] During at least a portion of a period in which the received doorbell is parsed and sent to the storage device 330, the doorbell parser 321 of the switch 320 may write a command in the submission queue 311 of the host 310 based on the doorbell request. During at least a portion of a period in which the doorbell is sent to the storage device 330 and a request for a command is received from the storage device 330, the switch 320 may receive a command written in the submission queue 311 of the host 310 from the host 310.
[0061] The order of operations S309 to S326 is not limited to Figure 7 For example, operation S309 and operation S313 may be performed simultaneously. For example, operation S326 may be performed prior to operation S323. In any case, when sending the received doorbell to the storage device 330, the switch 320 may, for example, simultaneously write a command in the submission queue 311 of the host 310 based on the doorbell request. Figure 6 As shown by the shading in FIG, at least a portion of the delay required for the request of operation S316 sent by the doorbell parser 321 to reach the host 310 may overlap with the delay required for the doorbell of operation S309 sent by the switch 320 to reach the storage device 330 or the delay required for the request of operation S323 sent by the storage device 330 to reach the switch 320. Figure 6As shown by the shading in FIG, at least a portion of the delay required for the command of operation S326 sent from the host 310 to reach the switch 320 may overlap with the delay required for the doorbell of operation S309 sent by the switch 320 to reach the storage device 330, or the delay required for the request of operation S323 sent by the storage device 330 to reach the switch 320. Compared to the switch 220, the switch 320 can mix the delays required for the request and the command corresponding to the doorbell by using the doorbell parser 321.
[0062] In operation S333, the switch 320 may determine whether the command received in operation S326 is a write command. When the command received in operation S326 is a write command (Y), the switch 320 may parse the command in operation S336 to obtain a physical region page (PRP) or a scatter gather table (SGL). Here, the PRP or SGL may be an address included in the command received in operation S326, which may indicate a specific data storage area (location) of the data buffer 313 or a specific data storage area of the storage device 330. In operation S339, the switch 320 may request write data of the data buffer 313 of the host 310 and may send the request to the host 310 (③). For example, the command parser 322 of the switch 320 may perform operations S333 to S339. In operation S343, the storage device 330 may determine whether the command received in operation S329 is a write command. When the command received in operation S329 is a write command (Y), the storage device 330 may parse the command in operation S346 to obtain a PRP or SGL. In operation S349, the storage device 330 may request write data from the data buffer 313 of the host 310 and may send the request to the switch 320 (④). The DMA engine 333 may request write data. In operation S353, the switch 320 may receive write data from the data buffer 313 of the host 310 (⑤). The intermediate write buffer 324 of the switch 320 may store the received write data. In operation S356, the switch 320 may send the write data stored in the intermediate write buffer 324 to the storage device 330 in response to the request received in operation S349 (⑤). The DMA engine 333 may receive the write data and may store the received write data in the write buffer 335. For example, operation S333, operation S336, operation S339, and operation S353 may be similar to operation S343, operation S346, operation S349, and operation S356, respectively.
[0063] As similarly described above with respect to the doorbell, before or during at least a portion of the period between parsing a received command and sending the received command to the storage device 330, the command parser 322 of the switch 320 can request write data stored in the data buffer 313 of the host 310 based on the command or the address contained in the command. Before or during at least a portion of the period between sending the command to the storage device 330 and receiving the request for write data from the storage device 330, the switch 320 can receive the write data stored in the data buffer 313 of the host 310 from the host 310. As a result of preemptively processing the command, a delay in communication between the host 310 and the storage device 330 caused by the presence of the switch 320 can be reduced.
[0064] The order of operations S329 to S356 is not limited to Figure 7 For example, operation S329 and operation S333 may be performed simultaneously. For example, operation S353 may be performed prior to operation S349. In any case, when sending the received write command to the storage device 330, the switch 320 may, for example, simultaneously request the write data stored in the data buffer 313 of the host 310 based on the write command. Figure 6 As shown by the shading in FIG, at least a portion of the delay required for the request of operation S339 sent by the command parser 322 to reach the host 310 may overlap with the delay required for the write command of operation S329 sent by the switch 320 to reach the storage device 330 or the delay required for the request of operation S349 sent by the storage device 330 to reach the switch 320. Figure 6 As shown by the shading in FIG, at least a portion of the delay required for the write data of operation S353 sent from the host 310 to reach the switch 320 may overlap with the delay required for the write command of operation S329 sent by the switch 320 to reach the storage device 330, or the delay required for the request of operation S349 sent by the storage device 330 to reach the switch 320. Compared to the switch 220, the switch 320 can mix the delay required for the write data corresponding to the request and the write command by using the command parser 322.
[0065] In operation S343, when the command received in operation S329 is not a write command (N) or after operation S356, the storage device 330 may process the command in operation S359. For example, when the command received in operation S329 is a write command, the controller 331 may store the write data in the non-volatile memory 339. Operations S343, S346, S349, and S356 may be included in operation S359. For example, when the command received in operation S329 is a read command, the controller 331 may transmit the read data stored in the non-volatile memory 339 to the switch 320. The switch 320 may receive the read data and may transmit the received read data to the host 310. For example, the read data may be stored in the data buffer 313 of the host 310.
[0066] In operation S363, the storage device 330 may transmit completion information regarding the command to the switch 320 (⑥), and the switch 320 may receive the completion information (⑥). If the command received in operation S326 is not a write command (N) or after operation S356, the switch 320 may transmit the completion information to the host 310 in operation S366 (⑥). The completion information may be stored or written in the completion queue 312 of the host 310.
[0067] In one embodiment, components 321 to 324 of switch 320 may be implemented in hardware, software, or a combination of hardware and software. In the case of hardware, components 321 to 324 may be implemented differently by using registers, latches, flip-flops, logic circuits, logic gates, etc. Intermediate submission queue buffer 323 and intermediate write buffer 324 may correspond to areas allocated on an on-chip memory included in switch 320. In one embodiment, components 332 to 335 of controller 331 may be implemented in hardware, software, or a combination of hardware and software.
[0068] In one embodiment, the host 310 may not directly access the intermediate submission queue buffer 323 and the intermediate write buffer 324 of the switch 320 and the submission queue buffer 334 and the write buffer 335 of the controller 331 of the storage device 330. The host 310 can directly access the doorbell register 332 of the controller 331 of the storage device 330 through the switch 320 without the request of the switch 320 or the storage device 330. When the host 310 updates the doorbell of the doorbell register 332, by executing Figure 7In operations S309 to S356 , the switch 320 and the storage device 330 may obtain commands written in the submission queue 311 of the host 310 or write data stored in the data buffer 313 , and may store the commands or write data in the above components 323 , 324 , 334 , and 335 .
[0069] Figure 8 Shown Figure 1 Block diagram of a computing system. Figure 9 Shown Figure 8 Timing diagram of the computing system. Figure 10 Shows the operation Figure 8 The redundant description of the above components will be omitted, and the following description will focus on Figure 8 The computing system 400 and Figure 5 The differences between the computing systems 300 are discussed.
[0070] like Figure 8 As shown in , computing system 400 may include a host 410, a switch 420, and a storage device 430. Host 410 may include a submission queue 411, a completion queue 412, and a data buffer 413. Components 411 to 413 of host 410 may be similar to components 311 to 313 of host 310, respectively. Switch 420 may include a command parser 422 and an intermediate write buffer 424. Components 422 and 424 of switch 420 may be similar to components 322 and 324 of switch 320, respectively.
[0071] The storage device 430 may include a controller 431 and a non-volatile memory 439. The controller 431 may include a doorbell register 432, a DMA engine 433, a submission queue controller memory buffer (SQ CMB) 434, and a write buffer 435. The components 431, 432, 433, 435, and 439 of the storage device 430 may be similar to the components 331, 332, 333, 335, and 339 of the storage device 330, respectively. However, the submission queue controller memory buffer 434 may be different from the submission queue buffer 334. The host 410 may directly write or store a command in the submission queue controller memory buffer 434. In detail, the host 410 may write a command to the submission queue 411, and may directly write the command written in the submission queue 411 to the submission queue controller memory buffer 434 without a request from the switch 420 or the storage device 430. For example, Figure 5Submission queue 311 can be placed on submission queue controller memory buffer 434 as a submission queue and can also be placed on host memory buffer HMB as submission queue 411. The same command can be stored in submission queue 411 and all submission queues on submission queue controller memory buffer 434. The size of submission queue 411 can be smaller than the size of submission queue 311, and the size of the submission queue on submission queue controller memory buffer 434 can be the same as the size of submission queue 311. Conversely, host 310 may not directly write or store commands of submission queue 311 in submission queue buffer 334. Instead, the commands may be stored in submission queue buffer 334 only after controller 331 performs operations S309, S319, S323, and S329. Because controller 431 includes or supports submission queue controller memory buffer 434 that can be directly accessed by host 410, switch 420 may not include doorbell parser 321 and intermediate submission queue buffer 323.
[0072] Combined reference Figure 8 、 Figure 9 and Figure 10 , the host 410 may store a command for the storage device 430 in the submission queue 411. In operation S403, the host 410 may send the command written in the submission queue 411 to the switch 420 (①), and the switch 420 may receive the command (①). In operation S406, the switch 420 may send the command to the storage device 430 (①). In operation S409, which may be located after operation S403, the host 410 may send a doorbell of the submission queue placed on the submission queue controller memory buffer 434 to the switch 420 (②), and the switch 420 may receive the doorbell (②). In operation S413, the switch 420 may send the doorbell to the storage device 430 (②). The delay required for the storage device 430 to receive the command and the doorbell may be shorter than the delay required for the storage device 230 / 330 to receive both the command and the doorbell.
[0073] Figure 10 Operations S416 to S449 in Figure 7The operations S333 to S366 in are similar. The switch 420 may perform operations S416 to S423, request write data of the data buffer 413, and send the request to the host 410 (②). The storage device 430 may perform operations S426 to S433, request write data of the data buffer 413, and send the request to the switch 420 (③). The switch 420 may perform operation S436, and may receive write data of the data buffer 413 from the host 410 (④). The switch 420 may perform operations S436 and S439, store the write data in the intermediate write buffer 424, and send the write data stored in the intermediate write buffer 424 to the storage device 430 (④). The storage device 430 may perform operations S439 to S446, and may send completion information about the command to the switch 420 (⑤). In operation S449, the switch 420 may send the completion information to the host 410 (⑤).
[0074] Figure 11 Shown Figure 1 Block diagram of a computing system. Figure 12 Shown Figure 11 Timing diagram of the computing system. Figure 13 Shows the operation Figure 11 The redundant description of the above components will be omitted, and the following description will focus on Figure 11 The computing system 500 and Figure 5 The differences between the computing systems 300 are discussed.
[0075] like Figure 11 As shown in FIG, computing system 500 may include a host 510, a switch 520, and a storage device 530. Host 510 may include a submission queue 511, a completion queue 512, and a data buffer 513. Figure 11 The components 511 to 513 of the host 510 can be respectively Figure 5 The components 311 to 313 of the host 310 are similar. The switch 520 may include a doorbell resolver 521 and an intermediate submission queue buffer 523. Figure 11 The components 521 and 523 of the switch 520 can be respectively Figure 5 Components 321 and 323 of switch 320 are similar.
[0076] The storage device 530 may include a controller 531 and a non-volatile memory 539 . The controller 531 may include a doorbell register 532 , a DMA engine 533 , a submission queue buffer 534 , and a write controller memory buffer (write CMB) 535 . Figure 11The components 531, 532, 533, 534 and 539 of the storage device 530 can be respectively Figure 5 Components 331, 332, 333, 334, and 339 of the storage device 330 are similar. However, the write controller memory buffer 535 may be different from the write buffer 335. The host 510 may store the write data directly in the write controller memory buffer 535. In detail, the host 510 may store the write data in the data buffer 513, and may directly store the write data stored in the data buffer 513 in the write controller memory buffer 535 through the switch 520 without a request from the switch 520 or the storage device 530. For example, Figure 5 The data buffer 313 can be set on the write controller memory buffer 535 as a data buffer, and can also be set on the host memory buffer HMB as the data buffer 513. The same data can be stored in all data buffers on the data buffer 513 and the write controller memory buffer 535. The size of the data buffer 513 can be smaller than the size of the data buffer 313, and the size of the data buffer on the write controller memory buffer 535 can be the same as the size of the data buffer 313. In contrast, the host 310 may not directly store the write data of the data buffer 313 in the write buffer 335. Instead, the data may be stored in the write buffer 335 only after the controller 331 performs operations S343, S346, S349, and S356. Since the controller 531 includes the write controller memory buffer 535 that can be directly accessed by the host 510, the switch 520 may not include the command parser 322 and the intermediate write buffer 324.
[0077] Combined reference Figure 11 、 Figure 12 and Figure 13 , operations S503, S506, and S509 may be similar to operations S303, S353, and S356, respectively. In operation S503, the host 510 may store a command for the storage device 530 in the submission queue 511. When the command received in operation S503 is a write command, in operation S506, the host 510 may send the write data to the switch 520 (①). In operation S509, the switch 520 may send the write data to the storage device 530 (①). The delay required for the storage device 530 to receive the write data may be shorter than the delay required for the storage device 230 / 330 / 430 to receive the write data. When the command is not a write command, operations S506 and S509 may be omitted.
[0078] Operations S513 to S536 are similar to operations S306 to S329. The host 510 may perform operation S513 and may send the doorbell of the submission queue 511 to the switch 520 (②). The switch 520 may perform operations S513 to S523, may send the doorbell to the storage device 530 (②), may request a command of the submission queue 511, and may send the request to the host 510 (②). The storage device 530 may perform operations S516, S526, and S529, may request a command of the submission queue 511, and may send the request to the switch 520 (③). The switch 520 may perform operations S533 and S536, may receive the command of the submission queue 511 from the host 510 (④), and may send the command stored in the intermediate submission queue buffer 523 to the storage device 530 (④). The storage device 530 may perform operation S536 and may receive the command (④). The storage device 530 may perform operations S539 and S543 and may transmit completion information on the command to the switch 520 (⑤). In operation S546, the switch 520 may transmit the completion information to the host 510 (⑤).
[0079] Figure 14 Shown Figure 1 Block diagram of a computing system. Figure 15 Shown Figure 14 The redundant description of the above components will be omitted, and the following description will focus on Figure 14 The computing system 600 and Figure 5 computing system 300, Figure 8 The computing system 400 and Figure 11 The differences between the computing systems 500 are discussed.
[0080] The computing system 600 may include a host 610, a switch 620, and a storage device 630. The host 610 may include a submission queue 611, a completion queue 612, and a data buffer 613. Figure 14 The components 611, 612 and 613 of the host 610 can be respectively Figure 8 Components 411 and 412 of the host 410 and Figure 11 The components 513 of the host 510 are similar. Figure 14 The switch 620 can be connected with Figure 3 Similar to switch 220.
[0081] The storage device 630 may include a controller 631 and a non-volatile memory 639. The controller 631 may include a doorbell register 632, a submission queue controller memory buffer 634, and a write controller memory buffer 635. Figure 14The components 631, 632, 634, 635 and 639 of the storage device 630 can be respectively Figure 5 Components 331 and 332 of the storage device 330, Figure 8 Component 434 of the storage device 430, Figure 11 The components 535 of the storage device 530 and Figure 5 339 of the storage device 330. Although not shown in the drawings, the controller 631 may also include a DMA engine. Since the controller 631 includes the submission queue controller memory buffer 634 and the write controller memory buffer 635 that can be directly accessed by the host 610, the switch 620 may not include the components 321 to 324 of the switch 320.
[0082] Combined reference Figure 14 and Figure 15 , the host 610 may store a command for the storage device 630 in the submission queue 611. The host 610 may send the command of the submission queue 611 to the switch 620 (①). The switch 620 may send the command to the storage device 630 (①). When the command is a write command, the host 610 may send write data to the switch 620 (②). The switch 620 may send write data to the storage device 630 (②). When the command is not a write command, the sending of the write data may be omitted. Figure 14 and 15 Unlike the example shown in , write data may be sent prior to the command, or write data and the command may be sent simultaneously. After sending the command and write data to the switch 620, the host 610 may send a doorbell of the submission queue 611 to the switch 620 (③). The switch 620 may send the doorbell to the storage device 630 (③). The delay required for the storage device 630 to receive the command, write data, and doorbell may be shorter than the delay required for the storage device 230 / 330 / 530 to receive all of the command, write data, and doorbell. The storage device 630 may process the command. The storage device 630 may send completion information about the command to the switch 620 (④). The switch 620 may send the completion information to the host 610 (④).
[0083] Figure 16 Shown Figure 1 Block diagram of a computing system. Figure 17 Shown Figure 16 Timing diagram of the computing system. Figure 18 Shows the operation Figure 16 The redundant description of the above components will be omitted, and the following description will focus on Figure 16 The computing system 700 and Figure 5 The computing system 300 and Figure 14 The differences between the computing systems 600 in FIG.
[0084] The computing system 700 may include a host 710, a switch 720, and a storage device 730. The host 710 may include a submission queue 711, a completion queue 712, and a data buffer 713. Figure 16 The components 711 to 713 of the host 710 can be respectively Figure 14 The components 611 to 613 of the host 610 are similar. The storage device 730 may include a controller 731 and a non-volatile memory 739. The controller 731 may include a doorbell register 732, a DMA engine 733, a submission queue buffer 734, and a write buffer 735. Figure 16 The components 731, 732, 733, 734, 735 and 739 of the storage device 730 can be respectively Figure 5 Components 331, 332, 333, 334, 335 and 339 of the storage device 330 are similar.
[0085] The switch 720 may include a submission queue controller memory buffer 723 and a write controller memory buffer 724. The operation of the submission queue controller memory buffer 723 may be similar to the operation of the submission queue controller memory buffer 434. The host 710 may directly write or store commands in the submission queue controller memory buffer 723 of the switch 720. In contrast, the host 710 may not directly write or store commands in the submission queue buffer 734 of the controller 731 of the storage device 730. The operation of the write controller memory buffer 724 may be similar to the operation of the write controller memory buffer 535. The host 710 may directly store write data in the write controller memory buffer 724 of the switch 720. In contrast, the host 710 may not directly store write data in the write buffer 735 of the controller 731 of the storage device 730.
[0086] Combined reference Figure 16 、 Figure 17 and Figure 18 In operation S703, the host 710 may send the command of the submission queue 711 to the switch 720 (①). The submission queue controller memory buffer 723 may store the command received in operation S703. When the command is a write command, in operation S706, the host 710 may send the write data of the data buffer 713 to the switch 720 (②). The write controller memory buffer 724 may store the write data received in operation S706. Figures 16 to 18Unlike the example shown in FIG, write data may be sent prior to commands, or both write data and commands may be sent simultaneously. After operations S703 and S706, in operation S709, the host 710 may send a doorbell of the submission queue 711 to the switch 720 (③). Operation S713 may be similar to operation S309. The switch 720 may perform operation S713 and send the doorbell to the storage device 730 (③).
[0087] Figure 18 Operation S716, operation S719, operation S723, operation S726, operation S729, operation S733, operation S736, operation S739, operation S743 and operation S746 in Figure 7 Operations S319, S323, S329, S343, S346, S349, S356, S359, S363, and S366 are similar to those in the preceding text. The storage device 730 may perform operations S716 and S719, request a command to submit the queue 711, and send the request to the switch 720 (④). The switch 720 may perform operation S723, and send a command to submit the queue controller memory buffer 723 to the storage device 730 (⑤). The storage device 730 may perform operations S726, S729, and S733, request write data to the data buffer 713, and send the request to the switch 720 (⑥). The switch 720 may perform operation S736, and send the write data written to the controller memory buffer 724 to the storage device 730 (⑦). The storage device 730 may perform operations S739 and S743 and may transmit completion information on the command to the switch 720 (⑧). In operation S746, the switch 720 may transmit the completion information to the host 710 (⑧).
[0088] Figure 19 A block diagram of a storage device according to an embodiment is shown.
[0089] Reference Figures 3 to 18 describe Figure 19 The storage device 830 may be one of the storage devices 230 to 730. The storage device 830 may include a controller 831, a memory buffer 838, and a non-volatile memory 839.
[0090] The controller 831 may be implemented using an SoC, an ASIC, or an FPGA. The controller 831 may include a processor 831-1, an on-chip memory 831-2, a non-volatile memory interface circuit 831-3, an external interface circuit 831-4, a DMA engine 831-5, and a buffer interface circuit 831-6. The processor 831-1 may control the components 831-2 to 831-6 of the controller 831. The processor 831-1 may include an internal cache memory and at least one or more cores (e.g., homogeneous multi-core or heterogeneous multi-core). The processor 831-1 may execute program code, software, applications, etc. loaded onto the on-chip memory 831-2 or the memory buffer 838.
[0091] The on-chip memory 831-2 may include latches, registers, SRAM, DRAM, thyristor random access memory (TRAM), tightly coupled memory (TCM), etc. A flash translation layer (FTL) may be loaded onto the on-chip memory 831-2. The FTL may manage the mapping between externally provided logical addresses and physical addresses of the non-volatile memory 839. The FTL may also perform garbage collection operations, wear leveling operations, and address mapping operations.
[0092] The nonvolatile memory interface circuit 831-3 can communicate with the nonvolatile memory 839 according to an interface protocol such as toggle double data rate (DDR). The nonvolatile memory interface circuit 831-3 can communicate with one or more nonvolatile memories 839 through a channel CH1, and can communicate with one or more nonvolatile memories 839 through a channel CHn (n is an integer of 2 or greater). The number of channels CH1 to CHn between the controller 831 and the nonvolatile memory 839 can be one or more, the number of nonvolatile memories assigned to one channel can be one or more, and each nonvolatile memory 839 can be a reference. Figures 3 to 18 Under the control of the processor 831-1, the nonvolatile memory interface circuit 831-3 can transmit data from the outside (for example, the host 11, the root complex 120, the electronic devices 141, 151, 153, 161 and 162, and the switches 220 to 720: see Figures 2 to 18 ) is sent to the nonvolatile memory 839, and the write data may be stored in the nonvolatile memory 839. The nonvolatile memory interface circuit 831-3 may receive read data sent from the nonvolatile memory 839 under the control of the processor 831-1.
[0093] The external interface circuit 831-4 can communicate with the outside (eg, the host 11, the root complex 120, the electronic devices 141, 151, 153, 161, and 162, and the switches 220 to 720: see Figures 2 to 18 For example, the interface protocol may be USB, SCSI, PCIe, M-PCIe, NVMe, ATA, PATA, SATA, SAS, IDE, UFS, FireWire, etc.
[0094] Under the control of the processor 831-1, the DMA engine 831-5 can directly access devices (eg, the host 11, the root complex 120, the electronic devices 141, 151, 153, 161, and 162, and the switches 220 to 720: see Figures 2 to 18 ) various memory buffers (e.g., HMB, 211, 213, 323, 324, 424, 523, 723, and 724 of the storage device 830). For example, the DMA engine 831-5 can access one of the above memory buffers, can receive commands, can receive write data, and can send read data of the storage device 830. The DMA engine 831-5 can access the various memory buffers 234 to 734 and 235 to 735 of the storage devices 230 to 830, the on-chip memory 831-2, and the memory buffer 838, and can exchange data with them.
[0095] The buffer interface circuit 831-6 can communicate with the memory buffer 838 according to an interface protocol (such as the DDR standard). The buffer interface circuit 831-6 can exchange data with the memory buffer 838 under the control of the processor 831-1. The memory buffer 838 may include latches, registers, SRAM, DRAM, TRAM, TCM, etc. For example, the memory buffer 838 may be provided outside the controller 831 and may be placed inside the memory device 830. For another example, the memory buffer 838 may not be included in the memory device 830. In a configuration in which the memory buffer 838 is included in the memory device 830, the processor 831-1 can use the memory buffer 838 and the on-chip memory 831-2 as a cache memory.
[0096] In one embodiment, the controller 831 may execute the Figures 5 to 18The following operations are associated with the storage devices 330 to 730 described in the following examples: S309, S319, S323, S329, S343, S346, S349, S356, S359, S363, S406, S413, S426, S429, S433, S439, S443, S446, S509, S516, S526, S529, S536, S539, S543, S713, S716, S719, S723, S726, S729, S733, S736, S739, and S743. In another embodiment, the operation of the storage device 830, which is not an endpoint device but an intermediate device, can be the same as that of the reference device 830. Figures 5 to 18 The operations of the switches 320 to 720 described above are similar. In this case, the controller 831 may also include the components 321 to 324, 422, 424, 521, 523, 723 and 724 of the switches 320 to 720. The controller 831 may perform the same operations as described above. Figures 5 to 18 The following operations are associated with the described switches 320 to 720: S306, S309, S313, S316, S323, S326, S329, S333, S336, S339, S349, S353, S356, S363, S366, S403, S406, S409, S413, S416, S419, S423 , S433, S436, S439, S446, S449, S506, S509, S513, S516, S519, S523, S529, S533, S536, S543, S546, S703, S706, S709, S713, S719, S723, S733, S736, S743 and S746.
[0097] Figure 20 A block diagram of a computing device according to an embodiment is shown. Figures 1 to 19 In the described computing systems 10 and 100 to 700, various embodiments of the inventive concept may be applied to a computing device 1000. The computing device 1000 may include a main processor 1100, a memory 1200, a user interface 1300, a storage device 1400, a communication block 1500, and a graphics processor 1600. For example, the computing device 1000 may be referred to as a "mobile device."
[0098] The main processor 1100 can control all operations of the computing device 1000. The main processor 1100 can be configured to process various arithmetic operations or logical operations. The main processor 1100 can be implemented using a dedicated logic circuit including one or more processor cores, an FPGA, an ASIC, an SoC, etc. The main processor 1100 can be implemented using a central processing unit, a microprocessor, a general-purpose processor, a dedicated processor, or an application processor. For example, each of the hosts 11 and 210 to 710 and the processor 110 can correspond to the main processor 1100.
[0099] The memory 1200 may temporarily store data used for the operation of the computing device 1000. The memory 1200 may store data processed by the main processor 1100 or to be processed by the main processor 1100. For example, the memory 130 may correspond to the memory 1200.
[0100] The user interface 1300 may perform communication mediation between the user and the computing device 1000 under the control of the main processor 1100. For example, the user interface 1300 may process input from a keyboard, a mouse, a keypad, buttons, a touch panel, a touch screen, a touch pad, a touch ball, a camera, a gyro sensor, a vibration sensor, etc. In addition, the user interface 1300 may process output to be provided to a display device, a speaker, a motor, etc.
[0101] The storage device 1400 may include a storage medium capable of storing data regardless of whether power is supplied. For example, the storage device 1400 may be a reference Figures 1 to 19 The computing device 1000 may include one of the electronic devices 12, 13, 141, 142, 151 to 154, and 161 to 163, the switches 220 to 720, and the storage devices 230 to 730. The storage device 1400 may be an intermediate device, and another intermediate device and another endpoint device connected to the storage device 1400 may also be included in the computing device 1000.
[0102] The communication block 1500 may communicate with external devices / systems of the computing device 1000 under the control of the main processor 1100. For example, the communication block 1500 may communicate with external devices / systems of the computing device 1000 based on at least one of various wired communication protocols (such as Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Universal Serial Bus (USB), and FireWire) and / or at least one of various wireless communication protocols (such as Long Term Evolution (LTE), Worldwide Interoperability for Microwave Access (WiMAX), Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Bluetooth, Near Field Communication (NFC), Wireless Fidelity (Wi-Fi), and Radio Frequency Identification (RFID)).
[0103] The graphics processor 1600 may be a graphics processing unit (GPU) and may include multiple processor cores (e.g., multiple graphics processing units). The processor cores included in the graphics processor 1600 may quickly process graphics data in parallel. For example, the graphics processor 1600 may process various graphics operations such as pixel shaders, supersampling, and color space conversion by using the processor cores.
[0104] Each of the main processor 1100, the memory 1200, the user interface 1300, the storage device 1400, the communication block 1500, and the graphics processor 1600 may be implemented using a circuit-level, chip-level, and / or package-level device to be installed in the computing device 1000. Alternatively, each of the main processor 1100, the memory 1200, the user interface 1300, the storage device 1400, the communication block 1500, and the graphics processor 1600 may be implemented using an independent electronic device to be assembled in the computing device 1000. The installed or assembled components may be connected to each other via a bus 1700.
[0105] The bus 1700 may provide a communication path of wires between components of the computing device 1000. The components of the computing device 1000 may exchange data with each other based on the bus format of the bus 1700. For example, the bus format may include one or more of various protocols such as PCIe, NVMe, SCSI, ATA, PATA, SATA, SAS, and UFS.
[0106] According to an embodiment of the inventive concept, a communication speed between an endpoint device and a host may be improved by using a buffer of an electronic device placed between the endpoint device and the host.
[0107] While the inventive concept has been described with reference to the embodiments of the inventive concept, it will be apparent to those skilled in the art that various changes and modifications can be made thereto without departing from the spirit and scope of the inventive concept as set forth in the appended claims.
Claims
1. A computing system comprising: Host; a first electronic device coupled to the host; as well as a second electronic device coupled to the first electronic device, the second electronic device being configured to communicate with the host through the first electronic device, The first electronic device is configured as follows: Based on the doorbell received from the host, request the host to send a write command that is written into the host's submission queue; Stores write commands received from the host; requesting the host to send write data corresponding to the write command stored in the host's data buffer; and Stores write data received from the host, The first electronic device includes a command parser configured to request write data stored in a data buffer of the host during at least a portion of a period of parsing a write command and sending the write command to the second electronic device.
2. The computing system of claim 1, wherein: The first electronic device is a PCI Express switch, and The second electronic device is a non-volatile memory flash device.
3. The computing system of claim 1 , wherein: The first electronic device further includes: an intermediate submission queue buffer configured to store write commands; and An intermediate write buffer is configured to store write data.
4. The computing system of claim 3, wherein: The host does not directly access the intermediate submission queue buffer and the intermediate write buffer of the first electronic device.
5. The computing system of claim 3, wherein: The first electronic device is further configured to: In response to a request from the second electronic device receiving the doorbell, sending a write command stored in the intermediate submission queue buffer to the second electronic device; as well as In response to a request from the second electronic device that receives the write command, the write data stored in the intermediate write buffer is transmitted to the second electronic device. The computing system according to claim 5 , wherein: The first electronic device is further configured to: The completion information of the write command is received from the second electronic device that received the write command, and the completion information is transmitted to the host.
7. The computing system of claim 3, wherein: The first electronic device further includes: The doorbell parser is configured to request a write command in a submission queue of the host during at least a portion of a period of parsing the doorbell and sending the doorbell to the second electronic device.
8. The computing system of claim 7, wherein: The doorbell resolver is also configured to access the host's submission queue indicated by the doorbell's address.
9. The computing system of claim 1, wherein: The first electronic device is further configured to: Request the host to send a read command that is written to the host's submission queue; Stores read commands received from the host; receiving read data corresponding to the read command sent from the second electronic device that has received the read command; as well as Sends read data to the host.
10. A computing system comprising: Host; a first electronic device coupled to the host; as well as a second electronic device coupled to the first electronic device, the second electronic device being configured to communicate with the host through the first electronic device, The first electronic device is configured as follows: receiving a write command for a second electronic device from the host and sending the write command to the second electronic device; receiving, from the host, a doorbell associated with a submission queue in which a write command is written; Sending the doorbell to a second electronic device; requesting write data corresponding to the write command stored in a data buffer of the host; and Stores write data received from the host, The first electronic device includes a command parser configured to request write data stored in a data buffer of the host during at least a portion of a period of parsing a write command and sending the write command to the second electronic device.
11. The computing system of claim 10, wherein: The second electronic device includes: a submission queue controller memory buffer configured to store write commands sent from a host via the first electronic device; and The doorbell register is configured to store the doorbell signal sent from the host through the first electronic device.
12. The computing system of claim 11, wherein: The host is configured to directly access the doorbell register and the submission queue controller memory buffer of the second electronic device without a request from the second electronic device.
13. The computing system of claim 10, wherein: The first electronic device further includes: An intermediate write buffer is configured to store write data.
14. The computing system of claim 13, wherein: The first electronic device is further configured to transmit the write data stored in the intermediate write buffer to the second electronic device in response to a request of the second electronic device that receives the write command.
15. The computing system of claim 14, wherein: The first electronic device is further configured to: The completion information of the write command is received from the second electronic device that received the write command, and the completion information is transmitted to the host.
16. The computing system of claim 10, wherein: The host is configured to send a doorbell to the first electronic device after sending a write command to the first electronic device.
17. A computing system comprising: Host; a first electronic device coupled to the host, the first electronic device including a submission queue controller memory buffer and a write controller memory buffer; as well as a second electronic device coupled to the first electronic device, the second electronic device being configured to communicate with the host through the first electronic device, The first electronic device is configured as follows: receiving a write command from a host that is written to a submission queue of the host and storing the write command in a submission queue controller memory buffer; receiving, from the host, write data corresponding to the write command stored in a data buffer of the host; storing the write data in a write controller memory buffer; Receive doorbells from the host regarding submission queues; and Send the doorbell to the second electronic device, The first electronic device further includes: a command parser configured to request write data stored in a data buffer of the host during at least a portion of a period of parsing the write command and sending the write command to the second electronic device.
18. The computing system of claim 17, wherein: The first electronic device is further configured to: sending a write command stored in a submission queue controller memory buffer to the second electronic device in response to receiving a request from the second electronic device of the doorbell; as well as In response to a request from the second electronic device that receives the write command, the write data stored in the write controller memory buffer is transmitted to the second electronic device.
19. The computing system of claim 18, wherein: The first electronic device is further configured to: receiving completion information of the write command from the second electronic device that received the write command; and Sends completion information to the host.
Citation Information
Patent Citations
Air-wiper for cleanning a observation window
KR1020190099851A
System and method for processing and arbitrating submission and completion queues
CN110088723A