High speed peripheral component interconnect device and method of operation thereof
By introducing throughput calculation and latency management components into PCIe devices, the problem of unstable performance of PCIe devices is solved, and more efficient resource utilization and accurate command acquisition are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SK HYNIX INC
- Filing Date
- 2021-10-18
- Publication Date
- 2026-04-28
AI Technical Summary
In the existing technology, PCIe devices lack effective control over throughput and latency management, resulting in unstable performance and low resource utilization efficiency.
By employing components such as a throughput calculator, a throughput analysis information generator, a latency information generator, a command lookup table storage device, and a command acquirer, the system achieves precise performance control of PCIe devices by calculating and managing the throughput and latency of each function.
It enables precise management of PCIe device throughput and latency, improves resource utilization efficiency and performance stability, and ensures the accuracy and efficiency of command acquisition.
Smart Images

Figure CN115114013B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to Korean Patent Application No. 10-2021-0035522, filed with the Korean Intellectual Property Office on March 18, 2021, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] The various embodiments of this disclosure generally relate to an electronic device, and more particularly to a high-speed peripheral component interconnect (PCIe) device and a method of operating the PCIe device. Background Technology
[0004] Peripheral Component Interconnect (PCI) defines a bus protocol for connecting input / output devices to host devices. High-speed PCI (PCIe) incorporates the programming concept defined in the PCI standard and defines the physical communication layer as a high-speed serial interface.
[0005] A storage device is a means of storing data under the control of a host device such as a computer or smartphone. A storage device may include a memory device for storing data and a memory controller for controlling the memory device. Memory devices can be classified as volatile memory devices and non-volatile memory devices.
[0006] A volatile memory device is a memory device that stores data only when powered on and loses the stored data when power is interrupted. Examples of volatile memory devices can include static random access memory (SRAM) and dynamic random access memory (DRAM).
[0007] Non-volatile memory devices are memory devices that retain stored data even when power is interrupted. Examples of non-volatile memory devices include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), and flash memory. Summary of the Invention
[0008] Various embodiments of this disclosure relate to a PCIe device capable of limiting the performance of each function and a method of operating the PCIe device.
[0009] Embodiments of this disclosure may provide a high-speed peripheral component interconnect (PCIe) device. The PCIe device may include: a throughput calculator configured to calculate the throughput of each of a plurality of functions; a throughput analysis information generator configured to generate, for each of the plurality of functions, throughput analysis information indicating the result of a comparison between a throughput limit and the calculated throughput, the throughput limit being set for each of the plurality of functions; a latency information generator configured to generate, based on the throughput analysis information, a latency time for delaying a command fetch operation for each of the plurality of functions; a command lookup table storage device configured to store command-related information and the latency time of the function corresponding to the target command, the command-related information including information related to the target command to be fetched from the host; and a command fetcher configured to fetch the target command from the host based on the command-related information and the latency time of the corresponding function.
[0010] Embodiments of this disclosure may provide a method for operating a high-speed peripheral component interconnect (PCIe) device. The method may include: calculating the throughput of each of a plurality of functions; generating throughput analysis information for each of the plurality of functions, the throughput analysis information indicating the result of a comparison between a throughput limit set for each of the plurality of functions and the calculated throughput; generating a latency time for delaying a command retrieval operation for each of the plurality of functions based on the throughput analysis information; retrieving command-related information, the command-related information including information related to a target command to be retrieved from a host; and retrieving the target command from the host based on the command-related information and the latency time of the function corresponding to the target command among the plurality of functions. Attached Figure Description
[0011] Figure 1 A computing system according to an embodiment of the present disclosure is shown.
[0012] Figure 2 It shows Figure 1 The host.
[0013] Figure 3 It shows Figure 1 PCIe devices.
[0014] Figure 4 It shows Figure 3 The structure of layers included in a PCIe interface device.
[0015] Figure 5 A PCIe device according to an embodiment of the present disclosure is shown.
[0016] Figure 6 This is a graph illustrating the operation of generating delay time information according to embodiments of the present disclosure.
[0017] Figure 7 The command acquisition operation according to an embodiment of the present disclosure is illustrated.
[0018] Figure 8 This is a flowchart illustrating a method of operating a PCIe device according to an embodiment of the present disclosure.
[0019] Figure 9 This is a flowchart illustrating a method for obtaining a target command according to an embodiment of the present disclosure. Detailed Implementation
[0020] The specific structural or functional descriptions of embodiments of this disclosure as presented in this specification or application are used as examples to describe embodiments based on the concept of this disclosure. Embodiments based on the concept of this disclosure may be practiced in various forms and should not be construed as limited to the embodiments described in the specification or application.
[0021] Figure 1 A computing system 100 according to an embodiment of the present disclosure is shown.
[0022] Reference Figure 1 The computing system 100 may include a host 1000 and a high-speed peripheral component interconnect (PCIe) device 2000. For example, the computing system 100 may be a mobile phone, smartphone, MP3 player, laptop computer, desktop computer, game console, television, tablet computer, in-vehicle infotainment system, etc.
[0023] The host 1000 can control the data processing and operation of the computing system 100. The host 1000 can store the data, commands and / or program code required for the operation of the computing system 100.
[0024] The host 1000 may include an input / output control module that connects input / output devices to each other. For example, the input / output control module may include one or more of the following: a Universal Serial Bus (USB) adapter, a Peripheral Component Interconnect (PCI) or High-Speed PCI (PCIe) adapter, a Small Computer System Interface (SCSI) adapter, a Serial Atomic (SATA) adapter, a High-Speed Non-Volatile Memory (NVMe) adapter, etc. The host 1000 can communicate information with devices connected to the computing system 100 through the input / output control module.
[0025] PCI defines a bus protocol for connecting input / output devices to each other. PCIe has the concept of programming defined in the PCI standard and defines the physical communication layer as a high-speed serial interface.
[0026] The PCIe device 2000 can communicate with the host 1000 using PCIe. For example, the PCIe device 2000 can be implemented as various I / O device types such as network and storage devices.
[0027] In an embodiment, the PCIe device 2000 may be defined as an endpoint or a device including an endpoint.
[0028] An endpoint represents the type of functionality that can be a requester or completer of a PCIe transaction. Endpoints can be classified as traditional endpoints, PCIe endpoints, or Root Complex Integrated Endpoints (RCiEPs).
[0029] Traditional endpoints can be functions with a configuration space header of type 00h. Traditional endpoints can act as completers to support configuration requests. Traditional endpoints can act as completers to support input / output (I / O) requests. Traditional endpoints can accept I / O requests for either or both of the 80h and 84h positions, regardless of the corresponding endpoint's I / O decoding configuration. Traditional endpoints can generate I / O requests. Traditional endpoints should not issue lock requests. Traditional endpoints can implement extended configuration space capabilities.
[0030] Traditional endpoints, not required for requesting memory transactions, do not generate addresses larger than or equal to 4GB. When requesting interruptible resources, traditional endpoints are needed to support either or both Message Signaled Interrupts (MSI) and MSI-X. When implementing MSI, traditional endpoints can support 32-bit or 64-bit message address versions of the MSI functional architecture. Traditional endpoints can support 32-bit address allocation for the base address register of the requested memory resource. Traditional endpoints can appear in any hierarchical domain originating from the root union.
[0031] A PCIe endpoint can be a function with a configuration space header of type 00h. A PCIe endpoint can support configuration requests as a completer. A PCIe endpoint should not depend on the operating system (OS) allocation of I / O resources requested via the base address register (BAR). A PCIe endpoint cannot generate I / O requests. A PCIe endpoint can neither support lock requests as a completer nor generate lock requests as a requester. PCIe-compliant software drivers and applications can be created to avoid using locking semantics when accessing PCIe endpoints.
[0032] PCIe endpoints used as requesters of memory transactions can generate addresses larger than 4GB. When requesting interruptible resources, PCIe endpoints may be required to support either or both of MSI and MSI-X. When implementing MSI, PCIe endpoints can support the 64-bit message address version of the MSI functional architecture. The minimum memory address range requested by the base address register can be 128 bytes. PCIe endpoints can appear in any hierarchical domain originating from the root union.
[0033] A Root Union Integrated Endpoint (RCiEP) can be implemented with the internal logic of a root union that includes the root port. An RCiEP can be a function with a configuration space header of type 00h. An RCiEP can support configuration requests as a completer. An RCiEP may not require I / O resources requested via the base address register. An RCiEP may not generate I / O requests. An RCiEP may neither support lock requests as a completer nor generate lock requests as a requester. PCIe-compliant software drivers and applications can be created to avoid using locking semantics when accessing an RCiEP. An RCiEP acting as a requester of memory transactions can generate addresses with a capacity equal to or greater than the capacity of addresses that can be processed by host 1000 as a completer.
[0034] When requesting interruption resources, RCIEP is required to support either or both of MSI and MSI-X. When implementing MSI, RCIEP is allowed to support 32-bit or 64-bit message address versions of the MSI functional architecture. RCiEP can support 32-bit address allocation for the base address register of the requested memory resource. RCiEP cannot implement the link capacity, link state, link control, link capacity 2, link state 2, and link control 2 registers in the High-Speed PCI extension capabilities. RCiEP may not implement active state power management. RCiEP may not be hot-pluggable completely and independently of the root federation. RCiEP may not appear in the hierarchical domain exposed by the root federation. RCiEP may not appear in switches.
[0035] In this embodiment, the PCIe device 2000 can generate one or more virtual devices. For example, the PCIe device 2000 can store program code for generating one or more virtual devices.
[0036] In this embodiment, the PCIe device 2000 can generate a physical function (PF) device or a virtual function (VF) device based on a virtualization request received from the host 1000. For example, a physical function device can be configured as a virtual device that can be accessed by a virtualization intermediary of the host 1000. A virtual function device can be configured as a virtual device assigned to a virtual machine of the host 1000.
[0037] Figure 2 It shows Figure 1 The host.
[0038] In an embodiment, Figure 2 The PCIe available in host 1000 is shown.
[0039] Reference Figure 2 The host 1000 may include multiple system images 1010-1 to 1010-n, a virtualization intermediary 1020, a processor 1030, a memory 1040, a root union 1050, and a switch 1060, where n is a positive integer.
[0040] In this embodiment, each of the plurality of PCIe devices 2000-1 to 2000-3 may correspond to Figure 1 PCIe devices 2000.
[0041] System images 1010-1 to 1010-n can be software components that can run on a virtual system with PCIe functionality. In embodiments, system images 1010-1 to 1010-n can be referred to as virtual machines. System images 1010-1 to 1010-n can be software, such as an operating system for running applications or trusted services. For example, system images 1010-1 to 1010-n may include a guest operating system, shared or non-shared I / O device drivers, etc. To improve the efficiency of hardware resource utilization without modifying the hardware, multiple system images 1010-1 to 1010-n can run on computing system 100.
[0042] In an embodiment, a PCIe function may be a separate operating unit that provides physical resources included in PCIe devices 2000-1 to 2000-3. In this specification, the terms "PCIe function" and "function" may have the same meaning.
[0043] Virtualization intermediary 1020 may be a software component supporting multiple system images 1010-1 to 1010-n. In embodiments, virtualization intermediary 1020 may be referred to as a hypervisor or virtual machine monitor (VMM). Virtualization intermediary 1020 may be situated between hardware such as processor 1030 and memory 1040 and system images 1010-1 to 1010-n. Input / output (I / O) operations (inbound or outbound I / O operations) in computing system 100 may be intercepted and processed by virtualization intermediary 1020. Virtualization intermediary 1020 may present individual system images 1010-1 to 1010-n with their own virtual systems by abstracting hardware resources. The actual hardware resources available in the respective system images 1010-1 to 1010-n may vary depending on workload or client-specific policies.
[0044] The processor 1030 may include circuitry, interfaces, or program code that performs data processing and control operations on the components of the computing system 100. For example, the processor 1030 may include a central processing unit (CPU), an advanced RISC machine (ARM), an application-specific integrated circuit (ASIC), etc.
[0045] The memory 1040 may include volatile memory, such as SRAM, DRAM, etc., that stores data, commands, and / or program code required for the operation of the computing system 100. Furthermore, the memory 1040 may include non-volatile memory. In embodiments, the memory 1040 may also store program code operable to run one or more operating systems (OS) and virtual machines (VMs), as well as program code that runs a virtualization intermediary (VI) 1020 to manage virtual machines.
[0046] The processor 1030 can run one or more operating systems and virtual machines by executing program code stored in memory 1040. Furthermore, the processor 1030 can run a virtualization intermediary 1020 for managing virtual machines. In this way, the processor 1030 can control the operation of components of the computing system 100.
[0047] The root union 1050 indicates the root of the I / O hierarchy that connects the processor 1030 / memory 1040 to the I / O ports.
[0048] The computing system 100 may include one or more root unions. Further, each root union 1050 may include one or more root ports, such as 1051 and 1052. Root ports 1051 and 1052 represent separate hierarchies. Root union 1050 can communicate with switch 1060 or PCIe devices 2000-1 to 2000-3 via root ports 1051 and 1052.
[0049] The ability to route peer transactions between hierarchical domains via the root union 1050 is optional. Each hierarchical domain can be implemented as a sub-hierarchy comprising a single endpoint or one or more switches and endpoints.
[0050] When routing peer transactions between hierarchical domains, the root union 1050 may split packets into smaller packets. For example, the root union 1050 may split a single packet with a 256-byte payload into two packets, each with a 128-byte payload. An exception is that the root union 1050, which supports peer routing for vendor-defined messages (VDMs), is not allowed to split each vendor-defined message packet into smaller packets, except at 128-byte boundaries (i.e., the payload size of all resulting packets except the last one should be an integer multiple of 128 bytes).
[0051] Root union 1050, as the requester, should support generating configuration requests. Root union 1050, as the requester, can support generating I / O requests.
[0052] Root union 1050, as a completer, should not support locking semantics. Root union 1050, as a requester, can support generating lock requests.
[0053] The switch 1060 can be defined as a logical component of various virtual PCI-PCI bridge devices. The switch 1060 can communicate with the PCIe devices 2000-2 and 2000-3 connected to it.
[0054] The switch 1060 is indicated by two or more logical PCI-PCI bridges in the configuration software.
[0055] The 1060 switch can use the PCI bridge mechanism to transmit transactions. The 1060 switch can transmit all types of Transaction Layer Packets (TLPs) between all port sets. The 1060 switch can support locking requests.
[0056] The 1060 switch cannot break packets into smaller packets.
[0057] When contention occurs within the same virtual channel, arbitration between the ingress ports of the switch 1060 can be implemented in a round-robin or weighted round-robin manner.
[0058] Endpoints should not be represented as peers indicating downstream ports of the switch in the configuration software of the internal bus of the 1060 switch.
[0059] Figure 3 It shows Figure 1 PCIe devices.
[0060] Reference Figure 3 The PCIe device 2000 may include a PCIe interface device 2100 and multiple direct memory access (DMA) devices 2200-1 to 2200-n.
[0061] PCIe interface device 2100 can receive transaction layer data packets from multiple functions operating in multiple DMA devices 2200-1 to 2200-n. PCIe interface device 2100 can transmit the transaction layer data packets received from each function to... Figure 1 The host is 1000.
[0062] The types of DMA devices 2200-1 to 2200-n can include high-speed non-volatile memory (NVMe) devices, solid-state drive (SSD) devices, artificial intelligence central processing units (AI CPUs), artificial intelligence system-on-a-chip (AI SoCs), Ethernet devices, sound cards, graphics cards, etc. The types of DMA devices 2200-1 to 2200-n are not limited to these and can include other types of electronic devices using a PCIe interface. Functions can run on DMA devices 2200-1 to 2200-n and can be software or firmware that processes transaction layer data packets.
[0063] Functions can run on each of the DMA devices 2200-1 to 2200-n. For example, each of the DMA devices 2200-1 to 2200-n may include one or more functions running thereon. Here, the number of functions running on each of the DMA devices 2200-1 to 2200-n may vary depending on the embodiment. The PCIe device 2000 may generate physical or virtual functions in response to a virtualization request received from the host 1000. The PCIe device 2000 may assign functions to the respective DMA devices 2200-1 to 2200-n. The number of functions assigned to and running on each of the DMA devices 2200-1 to 2200-n can be set individually. Thus, one or more functions may be assigned to a DMA device (e.g., one of 2200-1 to 2200-n), and each function may operate as an independent unit of operation.
[0064] Figure 4 The structure of the layers included in a PCIe interface device is shown.
[0065] Reference Figure 4The diagram illustrates a first PCIe interface device 2100a and a second PCIe interface device 2100b. Each of the first PCIe interface device 2100a and the second PCIe interface device 2100b can correspond to... Figure 3 The PCIe interface device 2100.
[0066] The PCIe layer included in each of the first PCIe interface devices 2100a and 2100b may include three separate logical layers. For example, the PCIe layer may include a transaction layer, a data link layer, and a physical layer. Each layer may include two parts. One of these parts can handle outbound information (or information to be sent), and the other can handle inbound information (or received information). Furthermore, the first PCIe interface device 2100a and 2100b may use transaction layer data packets for information communication.
[0067] In each of the first PCIe interface device 2100a and the second PCIe interface device 2100b, the transaction layer can assemble and decompose transaction layer packets. Furthermore, the transaction layer can implement split transactions, which are transactions used to transmit other traffic to the link while the target system is collecting data required for a response. For example, the transaction layer can implement time-separated requests and responses. In embodiments, the four transaction address spaces may include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions may include one or more of the following: read requests and write requests for sending / receiving data to / from a memory-mapped location. In embodiments, memory space transactions may use two different address formats, such as a short address format like a 32-bit address and a long address format like a 64-bit address. Configuration space transactions can be used to access the configuration space of the PCIe system. Transactions targeting the configuration space may include read requests and write requests. Message space transactions (or messages) can be defined to support in-band communication between PCIe systems.
[0068] The transaction layer can store link configuration information, etc. In addition, the transaction layer can generate transaction layer packets (TLPs), or convert TLPs received from external devices into payloads or status information.
[0069] The data link layer performs link management functions and data integrity functions, including error detection and correction. Specifically, the sending side of the data link layer can accept a TLP assembled by the transaction layer, assign a data protection code to the TLP, and calculate the TLP sequence number. Furthermore, the sending side of the data link layer can send the data protection code and the TLP sequence number to the physical layer to transmit the corresponding information through the link. The receiving side of the data link layer can check the data integrity of the TLP received from the physical layer and send the TLP to the transaction layer for further processing.
[0070] The physical layer may include all circuitry used to perform interface operations. Here, all circuitry may include drivers, input buffers, serial-to-parallel conversion circuits, parallel-to-serial conversion circuits, phase-locked loops (PLLs), and impedance matching circuitry.
[0071] Furthermore, the physical layer may include logical sub-blocks and electronic blocks for physically transmitting data packets to an external PCIe system. Here, the logical sub-blocks may play a role necessary for the "digital" functions of the physical layer. For this purpose, the logical sub-blocks may include a transmitting section for preparing outgoing information to be transmitted by the electronic block, and a receiving section for identifying and preparing received information before passing it to the data link layer.
[0072] The physical layer may include transmitters and receivers. A transmitter can receive symbols from logical sub-blocks, serialize the symbols, and send the serialized symbols to an external device (e.g., an external PCIe system). Furthermore, a receiver can receive serialized symbols from an external device and convert the received symbols into a bit stream. The bit stream can be deserialized and supplied to logical sub-blocks. In other words, the physical layer can convert TLPs received from the data link layer into a serialized format and can convert data packets received from external devices into a deserialized format. Additionally, the physical layer may include logical functions related to interface initialization and maintenance.
[0073] although Figure 4 The structures of the first PCIe interface device 2100a and the second PCIe interface device 2100b are shown, but the first PCIe interface device 2100a and the second PCIe interface device 2100b may include any form, such as a fast path interconnect structure, a next-generation high-performance computing interconnect structure or any other hierarchical structure.
[0074] Figure 5 A PCIe device 500 according to an embodiment of the present disclosure is shown.
[0075] PCIe device 500 can correspond to Figures 1 to 3Any one of the PCIe devices 2000, 2000-1, 2000-2, and 2000-3 shown.
[0076] Reference Figure 5 The PCIe device 500 may include a throughput calculator 510, a throughput analysis information generator 520, a latency information generator 530, a command lookup table storage device 540, and a command retriever 550.
[0077] Throughput Calculator 510 can calculate the throughput of each of several functions running on multiple DMA devices. Throughput can be a metric indicating the performance of each function. Throughput Calculator 510 can periodically calculate the throughput of each function.
[0078] In one embodiment, the throughput calculator 510 can calculate throughput based on the occupancy of multiple functions for a data path shared among the multiple functions. In another embodiment, the data path can be a path used to connect a PCIe interface device to multiple DMA devices.
[0079] For example, the throughput calculator 510 can calculate the utilization rate of a function based on the number of transaction layer data packets processed per unit time through the data path for each of the multiple functions. Each of the multiple functions can send transaction layer data packets containing identification information of the corresponding function through the data path. Therefore, the throughput calculator 510 can calculate the utilization rate of each of the multiple functions based on the function identification information included in the transaction layer data packets. The throughput calculator 510 can calculate the throughput of multiple functions based on the calculated utilization rate of multiple functions. The throughput calculator 510 can provide the calculated throughput to the throughput analysis information generator 520.
[0080] In an embodiment, the throughput calculator 510 can calculate the read throughput corresponding to a read operation and the write throughput corresponding to a write operation for each of the multiple functions. Here, the read throughput corresponding to a read operation can be the throughput calculated during the read operation of the corresponding function, and the write throughput corresponding to a write operation can be the throughput calculated during the write operation of the corresponding function. Therefore, the throughput of each of the multiple functions can include the read throughput corresponding to a read operation and the write throughput corresponding to a write operation.
[0081] Throughput analysis information generator 520 can generate throughput analysis information for each of the multiple functions based on the throughput limits set for each function and the throughput calculated for each function. For example, throughput analysis information generator 520 can periodically generate throughput analysis information based on the throughput provided from throughput calculator 510.
[0082] Here, the throughput limit can be a threshold set to limit the throughput of each function. For example, the throughput analysis information generator 520 can receive information from the host 1000 regarding the throughput limit of each of the multiple functions. The throughput analysis information generator 520 can set the throughput limit for each of the multiple functions based on the received information about the throughput limit.
[0083] Here, throughput analysis information can be information indicating the result of a comparison between a throughput limit and the calculated throughput. In embodiments, throughput analysis information may include at least one of the following: information on whether the calculated throughput exceeds the throughput limit, the excess ratio of the calculated throughput to the throughput limit, the residual ratio of the calculated throughput to the throughput limit; information related to whether each function is idle; and information related to whether the calculated throughput is below a minimum performance threshold set for each function. Throughput analysis information may further include any of a variety of types of information that can be obtained through comparative analysis of throughput.
[0084] In an embodiment, when the throughput calculated for a specific function exceeds the throughput limit set for that specific function, the excess ratio of the calculated throughput to the throughput limit can be calculated. For example, the excess ratio of the calculated throughput to the throughput limit can be represented by the following equation (1).
[0085] Excess ratio = (Calculated throughput - Throughput limit) / Throughput limit (1)
[0086] In an embodiment, when the throughput calculated for a specific function does not exceed the throughput limit set for that specific function, the residual ratio of the calculated throughput to the throughput limit can be calculated. For example, the residual ratio of the calculated throughput to the throughput limit can be represented by the following equation (2).
[0087] Residual ratio = (throughput limit - calculated throughput) / throughput limit (2)
[0088] In this embodiment, the throughput analysis information generator 520 can generate read throughput analysis information corresponding to read operations and write throughput analysis information corresponding to write operations. For example, the throughput analysis information generator 520 can generate read throughput analysis information corresponding to read operations based on the comparison between the throughput corresponding to a read operation and the throughput limit. Further, the throughput analysis information generator 520 can generate write throughput analysis information corresponding to write operations based on the comparison between the throughput corresponding to a write operation and the throughput limit. Therefore, the throughput analysis information can include read throughput analysis information corresponding to read operations and write throughput analysis information corresponding to write operations.
[0089] In this embodiment, the minimum performance threshold may be a threshold used to prevent latency during the operation of a specific function. The throughput analysis information generator 520 may set a minimum performance threshold for each of the multiple functions.
[0090] Throughput analysis information generator 520 can provide throughput analysis information to delay time information generator 530.
[0091] The delay time information generator 530 can generate the delay time for each of the multiple functions based on throughput analysis information. Here, the delay time can be information used to delay the command acquisition operation corresponding to each function.
[0092] In an embodiment, when the delay time information generator 530 generates a delay time for one of a plurality of functions whose calculated throughput exceeds the throughput limit, the delay time information generator 530 may increase the delay time of the function based on the excess ratio of the calculated throughput to the throughput limit. For example, the delay time information generator 530 may calculate the delay time increment by multiplying a first constant value by the excess ratio. Here, the first constant value may be set differently by the host 1000 according to the settings. The delay time information generator 530 may calculate the value of the delay time increment added from the previous delay time of a previously generated function, as the current delay time corresponding to that function.
[0093] In an embodiment, when the delay time information generator 530 generates a delay time for one of a plurality of functions whose delay time is greater than an initial value and whose calculated throughput does not exceed a throughput limit, the delay time information generator 530 can reduce the delay time of that function based on the residual ratio of the calculated throughput to the throughput limit. In an embodiment, the initial value of the delay time can be "0". For example, the delay time information generator 530 can calculate the delay time reduction value by multiplying a second constant value by the residual ratio. Here, the second constant value can be set differently by the host 1000 according to the settings. The delay time information generator 530 can calculate the value by which the delay time reduction value is reduced from the previous delay time of a previously generated function, as the current delay time of that function.
[0094] In this embodiment, the delay time information generator 530 can set the delay time of the following functions among a plurality of functions to an initial value: functions that are in an idle state and functions whose calculated throughput is below a minimum performance threshold. Therefore, the delay time of those functions can be set to "0".
[0095] In this embodiment, the latency may include a read latency corresponding to a read operation and a write latency corresponding to a write operation. For example, the latency information generator 530 may generate a read latency corresponding to a read operation based on read throughput analysis information corresponding to a read operation. Furthermore, the latency information generator 530 may generate a write latency corresponding to a write operation based on write throughput analysis information corresponding to a write operation.
[0096] The delay time information generator 530 can provide the delay time to the command lookup table storage device 540.
[0097] The command lookup table storage device 540 may include a command lookup table. Here, the command lookup table may store: command-related information, including information related to a target command to be obtained from the host 1000; and the delay time of the function corresponding to the target command among a plurality of functions. The command lookup table may store command-related information for each of the plurality of target commands. In an embodiment, the command-related information may include the address of each target command stored in the host 1000, information indicating whether the corresponding target command is a read command or a write command, identification information of the function assigned to the corresponding target command, etc.
[0098] Command lookup table storage device 540 can receive command-related information about a target command from host 1000. For example, host 1000 can update the submission queue head doorbell to request PCIe device 500 to execute a target command. Here, command lookup table storage device 540 can receive command-related information about the requested target command from host 1000.
[0099] In an embodiment, the command lookup table storage device 540 can store delay elapsed information by associating delay elapsed information with command-related information. Here, delay elapsed information can be information indicating whether the delay time corresponding to the function of the target command has elapsed since the time point when the command-related information of the target command was stored in the command lookup table. For example, when the target command is a read command, delay elapsed information can be generated based on the read delay time corresponding to the read operation. When the target command is a write command, delay elapsed information can be generated based on the write delay time corresponding to the write operation.
[0100] In this embodiment, the command lookup table storage device 540 can start counting time from the point when command-related information is stored in the command lookup table, and then check whether the delay time has elapsed. For example, when the function's delay time has elapsed, the delay time elapsed information may include information indicating that the function's delay time has expired. On the other hand, when the function's delay time has not yet elapsed, the delay time elapsed information may include information indicating that the function's delay time has not yet expired.
[0101] Command acquirer 550 can acquire target commands from host 1000 based on command-related information of the target command and the delay time of the function corresponding to the target command.
[0102] In this embodiment, the command acquirer 550 can determine whether to acquire the target command based on the elapsed delay time information. For example, when it is determined, based on the elapsed delay time information, that the corresponding function's delay time has elapsed since the time point when the command-related information was stored in the command lookup table, the command acquirer 550 can send an acquisition command to the host 1000 to acquire the target command from the host 1000. On the other hand, when it is determined, based on the elapsed delay time information, that the corresponding function's delay time has not yet elapsed since the time point when the command-related information was stored in the command lookup table, the command acquirer 550 can delay the command acquisition operation for the target command. In this case, the command acquirer 550 can skip the command acquisition operation for the target command and can perform a command acquisition operation for another target command for which the delay time has already elapsed.
[0103] According to embodiments of this disclosure, command retrieval operations for target commands can be controlled based on the delay time assigned to each function, thereby enabling the rapid and accurate execution of performance limits for each function.
[0104] According to embodiments of this disclosure, the components of the PCIe device 500 may be implemented using one or more processors and memory or registers.
[0105] Figure 6 This is a graph illustrating the operation of generating delay time information according to embodiments of the present disclosure.
[0106] Figure 6 The upper part can indicate how the delay time of function i changes over elapsed time. The delay time of function i can be determined by... Figure 5 The delay time information generator 530 generates it. Figure 6 The lower part can indicate how the throughput of function i changes over time. The throughput of function i can be determined by... Figure 5 The throughput calculator 510 was generated.
[0107] Figure 6 The function i described in the text can indicate Figure 3 One of the many functions shown. Figure 6 In this example, assume that the throughput limit of function i is set to 1Gb / s and the minimum performance threshold of function i is set to 200Mb / s. Assume that the initial value of the latency of function i is "0".
[0108] Before time T0, the throughput of function i is lower than the throughput limit, so the delay time of function i can remain at the initial value.
[0109] During the time interval from time T0 to time T1, the throughput of function i exceeds the throughput limit. Therefore, the delay time information generator 530 can calculate the delay time increment value based on the excess ratio of the calculated throughput of function i to the throughput limit. Thus, the delay time of function i can be increased by the delay time increment value.
[0110] During the time interval from time T1 to time T2, the throughput of function i did not exceed the throughput limit but was higher than the minimum performance threshold. Therefore, the delay time information generator 530 can calculate the delay time reduction value based on the residual ratio of the calculated throughput of function i to the throughput limit. Thus, the delay time of function i can be reduced by the delay time reduction value.
[0111] During the time period from time T2 to time T3, the throughput of function i exceeds the throughput limit. Therefore, the delay time information generator 530 can calculate the delay time increment value based on the excess ratio of the calculated throughput of function i to the throughput limit. Thus, the delay time of function i can be increased again by the delay time increment value.
[0112] During the time interval from time T3 to time T4, it is assumed that the delay time of function i is repeatedly increased and decreased, thereby keeping the delay time of function i at a constant value. In this way, according to embodiments of this disclosure, command fetching operations can be controlled based on the delay time of each function, thereby enabling rapid and accurate performance limiting of each function.
[0113] During the time interval from time T4 to time T5, the throughput of function i did not exceed the throughput limit and was above the minimum performance threshold. Therefore, the delay time information generator 530 can calculate the delay time reduction value based on the residual ratio of the calculated throughput of function i to the throughput limit. Thus, the delay time of function i can be further reduced by the delay time reduction value.
[0114] At time T5, the throughput of function i is below the minimum performance threshold, so the delay time information generator 530 can set the delay time of function i to an initial value. Therefore, the delay time of function i can be "0".
[0115] During the time interval from time T5 to time T6, when the throughput of function i is below the throughput limit and therefore the delay time of function i has an initial value, the delay time of function i can be maintained at the initial value. That is, when the delay time is the initial value, the delay time will not increase until the throughput of function i exceeds the throughput limit.
[0116] At time T6, the throughput of function i exceeds the throughput limit, so the delay time information generator 530 can calculate the delay time increment value based on the excess ratio of the calculated throughput of function i to the throughput limit. Therefore, the delay time of function i can be increased again by the delay time increment value.
[0117] Figure 7 The command acquisition operation according to an embodiment of the present disclosure is illustrated.
[0118] Reference Figure 7 The command lookup table can store command-related information for multiple target commands, as well as elapsed delay information associated with that command-related information. Figure 7 In this context, assume that the command lookup table stores command-related information for five target commands, CMD1 to CMD5, from CMD1 INFO to CMD5 INFO.
[0119] Figure 5 The command acquirer 550 can determine whether to perform a command acquisition operation on the target command based on a command lookup table. The command acquirer 550 can check the command-related information and delay time elapsed information stored in the command lookup table at the time of execution of the command acquisition operation. Based on the results of the checks, the command acquirer 550 can perform the command acquisition operation on the target command when the delay time corresponding to the target command has expired, and can skip the command acquisition operation on the target command if the delay time has not expired.
[0120] For example, refer to Figure 7 The elapsed delay information stored in association with command-related information CMD1 INFO, CMD4 INFO, and CMD5 INFO may include information indicating that the corresponding delay time has expired. In this case, the command acquirer 550 may send an acquire command to the host 1000 to acquire the first target command CMD1, the fourth target command CMD4, and the fifth target command CMD5 from the host 1000.
[0121] Unlike these target commands, the elapsed delay information stored in association with command-related information CMD2 INFO and CMD3 INFO may include information indicating that the corresponding delay time has not yet expired. In this case, the command acquirer 550 can skip the command acquisition operation for the second target command CMD2 and the third target command CMD3.
[0122] Figure 8 This is a flowchart illustrating a method of operating a PCIe device according to an embodiment of the present disclosure.
[0123] Figure 8 The method shown can be derived from, for example... Figure 5 The PCIe device 500 shown is used to perform this.
[0124] Reference Figure 8 At S801, PCIe device 500, such as throughput calculator 510, can calculate the throughput of multiple functions.
[0125] Here, the PCIe device 500, such as the throughput calculator 510, can calculate the occupancy of a function on a data path based on the number of transaction layer packets processed per unit time through a data path shared among the multiple functions. The PCIe device 500, such as the throughput calculator 510, can calculate throughput based on the occupancy.
[0126] At S803, PCIe device 500, such as throughput analysis information generator 520, can generate throughput analysis information for each of the multiple functions based on the throughput limits set for each function in the function and the throughput calculated for each function in the function.
[0127] At S805, PCIe device 500, such as latency information generator 530, can generate latency for each of the multiple functions based on throughput analysis information.
[0128] Here, the PCIe device 500, such as the latency information generator 530, can increase the latency of one of the multiple functions whose calculated throughput exceeds the throughput limit based on the calculated excess ratio of throughput to throughput limit.
[0129] Furthermore, the PCIe device 500, such as the latency information generator 530, can reduce the latency of functions whose latency exceeds the initial value but whose calculated throughput does not exceed the throughput limit, based on the residual ratio of the calculated throughput to the throughput limit.
[0130] In addition, the PCIe device 500, such as the latency information generator 530, can set the latency of the following functions among a plurality of functions to an initial value: functions that are in an idle state and functions whose calculated throughput is below a minimum performance threshold.
[0131] At S807, the PCIe device 500, such as the command lookup table storage device 540, can obtain command-related information including information related to the target command to be obtained from the host.
[0132] At S809, the PCIe device 500, such as the command lookup table storage device 540, can store command-related information and the delay time of the function corresponding to the target command.
[0133] At S811, PCIe device 500, such as command retriever 550, can retrieve a target command from the host based on command-related information and the delay time of the function corresponding to the target command.
[0134] Figure 9 This is a flowchart illustrating a method for obtaining a target command according to an embodiment of the present disclosure.
[0135] Figure 9 The method shown can be implemented Figure 8 It is obtained by S809 and S811 as shown.
[0136] Figure 9 The method shown can be derived from, for example... Figure 5The PCIe device 500 shown is used to perform this.
[0137] Reference Figure 9 At S901, the PCIe device 500, such as the command lookup table storage device 540, can store command-related information.
[0138] At S903, the PCIe device 500, such as the command lookup table storage device 540, can store delay time elapsed information in association with command-related information.
[0139] At S905, the PCIe device 500, such as the command lookup table storage device 540, can determine whether the delay time for the function corresponding to the target command has elapsed or expired based on the delay time elapsed information. When it is determined at S905 that the delay time has elapsed, the PCIe device 500, such as the command retriever 550, can execute S907.
[0140] At S907, a PCIe device 500, such as a command retriever 550, can retrieve a target command from the host.
[0141] Conversely, when it is determined at S905 that the delay time has not yet elapsed, the PCIe device 500, such as the command retriever 550, can execute S909.
[0142] At S909, the PCIe device 500, such as the command retriever 550, can delay the command retrieval operation for the target command.
[0143] According to this disclosure, a PCIe device capable of limiting the performance of each function and a method for operating the PCIe device are provided.
Claims
1. A high-speed peripheral component interconnect device, i.e., a PCIe device, comprising: Throughput calculator, calculates the throughput of each of multiple functions running on multiple direct memory access devices, i.e., DMA devices. A throughput analysis information generator generates throughput analysis information for each of the plurality of functions, indicating the result of a comparison between a throughput limit and a calculated throughput, wherein the throughput limit is set for each of the plurality of functions. A delay time information generator generates a delay time for each of the plurality of functions for delaying command acquisition operations, based on the throughput analysis information. A command lookup table storage device stores command-related information and the delay time of the function corresponding to the target command. The command-related information includes information related to the target command to be obtained from the host. as well as The command acquirer retrieves the target command from the host based on the command-related information and the corresponding function's delay time.
2. The PCIe device of claim 1, wherein the throughput calculator calculates the throughput of a function based on the occupancy rate of each of the plurality of functions for the data path shared among the plurality of functions.
3. The PCIe device of claim 2, wherein the throughput calculator calculates the utilization rate of a function based on the number of transaction layer packets processed per unit time through the data path for each of the plurality of functions.
4. The PCIe device of claim 1, wherein the throughput analysis information generator receives from the host information relating to the throughput limit of each of the plurality of functions, and sets the throughput limit of each of the plurality of functions based on the received information relating to the throughput limit.
5. The PCIe device of claim 1, wherein for each of the plurality of functions, the throughput analysis information includes at least one of the following: information indicating whether the calculated throughput exceeds the throughput limit, the excess ratio of the calculated throughput to the throughput limit, and the residual ratio of the calculated throughput to the throughput limit; Information relating to whether each of the plurality of functions is in an idle state; And information relating to whether the calculated throughput is below the minimum performance threshold set for each of the plurality of functions.
6. The PCIe device of claim 5, wherein when the calculated throughput exceeds the throughput limit, the latency information generator increases the latency of a given function based on the excess ratio, the given function being one of the plurality of functions.
7. The PCIe device of claim 5, wherein when the calculated throughput does not exceed the throughput limit, the latency information generator reduces the latency of a given function based on the residual ratio, the given function being one of the plurality of functions whose latency is higher than the initial value.
8. The PCIe device of claim 5, wherein the latency information generator sets the latency of the following functions among the plurality of functions to an initial value: functions in an idle state, and functions whose calculated throughput is below a minimum performance threshold.
9. The PCIe device of claim 1, wherein the command lookup table storage device stores delay time elapsed information, the delay time elapsed information indicating whether a delay time for a corresponding function has elapsed since the time point at which the command-related information was stored in the command lookup table storage device, the delay time elapsed information being stored in association with the command-related information.
10. The PCIe device of claim 9, wherein, based on the elapsed delay information, when the delay time of the corresponding function has elapsed since the time point when the command-related information was stored, the command acquirer acquires the target command from the host.
11. The PCIe device of claim 9, wherein, based on the elapsed delay information, when the delay time of the corresponding function has not elapsed since the time point at which the command-related information was stored, the command acquirer delays the command acquisition operation for the target command.
12. The PCIe device according to claim 1, wherein: The calculated throughput includes the read throughput corresponding to the read operation of each of the plurality of functions and the write throughput corresponding to the write operation of each of the plurality of functions. The throughput analysis information includes read throughput analysis information corresponding to read operations and write throughput analysis information corresponding to write operations, and The delay time includes the read delay time corresponding to the read operation and the write delay time corresponding to the write operation.
13. A method for operating a high-speed peripheral component interconnect device, i.e., a PCIe device, the method comprising: Calculate the throughput of each of the multiple functions running on multiple direct memory access devices, i.e., DMA devices. For each of the plurality of functions, throughput analysis information is generated, which indicates the result of a comparison between the throughput limit set for each of the plurality of functions and the calculated throughput. Based on the throughput analysis information, a delay time for delaying command retrieval operations is generated for each of the plurality of functions; Obtain command-related information, which includes information related to the target command to be obtained from the host; as well as The target command is obtained from the host based on the command-related information and the delay time of the function corresponding to the target command among the multiple functions.
14. The method of claim 13, wherein calculating the throughput comprises: The occupancy rate of a function on the data path is calculated based on the number of transaction layer data packets processed per unit time through the data path shared among the multiple functions for each of the multiple functions. as well as The throughput is calculated based on the occupancy rate.
15. The method of claim 13, wherein for each of the plurality of functions, the throughput analysis information includes at least one of the following: information indicating whether the calculated throughput exceeds a throughput limit, the excess ratio of the calculated throughput to the throughput limit, and the residual ratio of the calculated throughput to the throughput limit; Information relating to whether each of the plurality of functions is in an idle state; And information relating to whether the calculated throughput is below the minimum performance threshold set for each of the plurality of functions.
16. The method of claim 15, wherein generating the delay time comprises: When the calculated throughput exceeds the throughput limit, the latency of a given function, which is one of the plurality of functions, is increased based on the excess ratio.
17. The method of claim 15, wherein generating the delay time comprises: When the calculated throughput does not exceed the throughput limit, the latency of a given function is reduced based on the residual ratio, where the given function is one of the functions whose latency is higher than the initial value.
18. The method of claim 15, wherein generating the delay time comprises: The initial values are set for the delay times of the following functions among the plurality of functions: functions that are in an idle state, and functions whose calculated throughput is below the minimum performance threshold.
19. The method of claim 13, further comprising: Store the information related to the command; as well as The system stores delay time elapsed information, which indicates whether the delay time corresponding to the function of the target command has elapsed since the time point when the command-related information was stored. The delay time elapsed information is stored in association with the command-related information.
20. The method of claim 19, wherein obtaining the target command comprises: Based on the delay time information, when the delay time corresponding to the target command has elapsed since the time point when the command-related information was stored, the target command is obtained from the host. as well as Based on the delay time information, if the delay time corresponding to the function of the target command has not elapsed since the time point when the command-related information was stored, the command retrieval operation for the target command is delayed.
Citation Information
Patent Citations
Battery module, battery rack and energy storage system comprising the battery module
KR1020210035522A
System and method for processing and arbitrating submission and completion queues
CN110088725A
METHOD, SYSTEM, AND COMPUTER PROGRAM PRODUCT FOR CONTROLLING FLOW OF PCIe TRANSPORT LAYER PACKETS
US20140281099A1