Apparatus and method for managing packet transfer across a memory fabric physical layer interface

The device and method address traffic bottlenecks in memory fabrics by prioritizing atomic requests and offloading CPU optimization, resulting in reduced latency and improved energy efficiency for data center operations.

JP7691422B2Active Publication Date: 2025-06-11ADVANCED MICRO DEVICES INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022532022
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-03
Filing Date
2020-10-02
Publication Date
2025-06-11
Estimated Expiration
2040-10-02

AI Technical Summary

Technical Problem

Traffic bottlenecks occur in memory fabrics due to differences in data rate and link width between physical layer interfaces, such as PCIe and Gen-Z, leading to inefficiencies in packet traffic management.

Method used

A device and method for managing packet transfer across cross-physical layer interfaces, which prioritizes atomic requests over other packet types by using a hierarchical priority buffer structure and offloads optimization from the host CPU, thereby optimizing traffic flow in data centers.

Benefits of technology

The solution effectively reduces latency and improves energy efficiency by prioritizing atomic requests and offloading packet management from the CPU, enhancing overall performance in data center environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691422000001
    Figure 0007691422000001
  • Figure 0007691422000002
    Figure 0007691422000002
  • Figure 0007691422000003
    Figure 0007691422000003
Patent Text Reader

Abstract

An apparatus and method for managing packet transfers between a memory fabric having a physical layer interface with a higher data rate than the data rate of a physical layer interface of another device receives incoming packets from the memory fabric physical layer interface, at least some of the packets including a different instruction type. The apparatus and method determine a packet type of the incoming packets received from the memory fabric physical layer interface, and if the determined incoming packet type is a packet type that includes an atomic request, the method and apparatus prioritizes forwarding of the incoming packets with the atomic request to memory access logic that accesses local memory within the apparatus over incoming packets of other packet types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Government license right) This invention was made with government support under the PathForward Project by Lawrence Livermore National Security, LLC, under Contract No. DE-AC52-07NA27344 and subcontract No. B620717, awarded by the Department of Energy (DOE). The government has certain rights in this invention.

Background Art

[0002] A system that uses a memory-semantic fabric to extend the central processing unit (CPU) memory byte addressable load-store model to the entire system such as a data center is adopted. The memory fabric is a type of point-to-point communication switch, also called a Gen-Z fabric, that is outside other types of devices that enable interfaces between pools of external memory modules and devices through the memory fabric within systems such as processor system-on-chips (SoCs), media modules, and data centers. For example, some processor SoCs include a processor with multiple processing cores that communicate with local memory such as dynamic random access memory (DRAM) or other suitable memory through local memory access logic such as a data fabric. Processor SoCs and other devices also need to interface with the memory fabric so that they can use, for example, a fabric-attached memory (FAM) module that can be an external (e.g., non-local) memory directly attached to the data center memory fabric. In some systems, the FAM module has memory access logic to process load and store requests but has no or very little computing power. In addition, the memory fabric attaches the FAM module as an addressable part of the entire host memory. The use case of the FAM module enables a disaggregated memory pool within a cloud data center. With the FAM module, the host is not restricted by the memory capacity limitations of the local server. Instead, the host obtains access to a vast pool of memory that is not attached to any host. The host partitions the memory among them or cooperates to share the FAM module.Gen-Z fabric has emerged as a high-performance, low-latency memory semantic fabric that can be used to communicate with each device within a system.

[0003] There is a need for improved apparatus and methods for managing traffic across the physical layer interfaces of a memory fabric that employs fabric-attached memory.

[0004] Embodiments may be more readily understood by considering the following description in conjunction with the following drawings, in which like reference numerals represent like elements.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Best Mode for Carrying Out the Invention

[0006] Traffic bottlenecks can occur due to the memory fabric. The physical layer interface of the memory fabric, also referred to as the memory fabric physical layer (PHY) interface, has higher performance operations than the physical layer interfaces associated with system-on-chips (e.g., host SoCs) or other devices connected to the memory fabric PHY interface. For example, the signaling standard used by the memory fabric to enable access to and messaging across the memory fabric may be on the order of 56 GT / s as compared to 16 or 32 GT / s using a peripheral component interconnect such as a PCIe interface on SoCs. In addition, the link width for the memory fabric is also designed to be larger. Some current processor SoC devices that interface with the PCI-e bus use a first-in first-out (FIFO) buffer to queue packet traffic, but the differences in data rate and link width on cross PHY interfaces such as the PCI-e physical layer (PHY) interface to the memory fabric PHY interface are still potential bottlenecks for packet traffic.

[0007] In some embodiments, the device serves as an interface for managing traffic priorities at the junction between multiple physical layer interfaces, such as between a PCIe PHY interface and a native memory fabric PHY interface like a Gen-Z 802.3 type memory fabric interface. In some embodiments, the device provides hardware assisted automated prioritization of packet traffic on a cross-PHY interface for data center workloads. In some embodiments, the device offloads cross-PHY interface optimization from a host CPU within a data center or other system that employs a memory fabric physical layer interface.

[0008] In certain embodiments, an apparatus and method for managing packet transfer between memory fabrics having a physical layer interface with a higher data rate than the data rate of a physical layer interface of another device receives incoming packets from the memory fabric physical layer interface, and at least a portion of the packets includes different instruction types. The apparatus and method determine the packet type of the incoming packets received from the memory fabric physical layer interface, and when the determined incoming packet type is a packet type that includes an atomic request, the method and apparatus prioritize the transfer of the incoming packets having an atomic request to memory access logic that accesses local memory within the apparatus over incoming packets of other packet types.

[0009] In some examples, the method includes queuing incoming packets determined to include atomic requests in a first priority buffer and queuing other packet types in a second priority buffer. The method also includes prioritizing the output of packets from the first priority buffer over the output of packets from the second priority buffer. In a particular example, the method includes queuing incoming packets determined to include store requests in a buffer and providing incoming packets having atomic requests to memory access logic.

[0010] In some examples, the method includes accessing data such as from one or more configuration registers and defining at least a portion of the local memory of the device as a priority memory region, where each memory region has a maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval, and maintaining a count of the number of memory accesses made to the memory regions defined by the memory fabric physical layer interface over a time interval. The method also includes storing read packets in a second priority buffer when the maximum permitted number of accesses is exceeded and, in a next time interval, providing the stored packets from the second priority buffer to memory access logic.

[0011] In a particular example, the method includes allocating the second priority buffer to include a plurality of second priority buffers, where each of the plurality of second priority buffers corresponds to a different defined memory region that stores incoming packets of a type determined to include read requests in each of the plurality of second priority buffers based on an address associated with the incoming packet.

[0012] According to some embodiments, the apparatus includes one or more processors and memory access logic coupled to a local memory, where the local memory is configurable as an addressable portion of addressable memory through a memory fabric physical layer interface. In some embodiments, the physical layer interface receives incoming packets from a memory fabric physical layer interface having a data rate higher than the data rate of the physical layer interface, and at least a portion of the packets includes different instruction types. In certain embodiments, the controller determines the packet type of the incoming packet received from the memory fabric physical layer interface, and when the determined incoming packet type is a packet type that includes an atomic request, the controller prioritizes the transfer of the incoming packet having the atomic request to the memory access logic over incoming packets of other packet types.

[0013] In a particular example, the apparatus includes a first priority buffer and a second priority buffer having a lower priority than the first priority buffer, and the controller prioritizes the transfer of the incoming packet having the atomic request over incoming packets of other packet types by queuing the incoming packet determined to include the atomic request in the first priority buffer and queuing other packet types in the second priority buffer. In some examples, the controller prioritizes the output of packets from the first priority buffer over the output of packets from the second priority buffer.

[0014] In some examples, the apparatus includes a buffer, and the controller prioritizes the transfer of the incoming packet having the atomic request over incoming packets of other packet types by queuing the incoming packet determined to include a store request in the buffer and providing the incoming packet having the atomic request to the memory access logic.

[0015] In a particular example, the apparatus includes a configuration register that stores data defining at least a portion of the local memory as a priority memory region, and each memory region has a maximum unrestricted number of memory fabric physical layer interface accesses permitted per time interval. In some embodiments, the controller maintains a count of the number of memory accesses made to a memory region defined by the memory fabric physical layer interface over a time interval, and stores a read packet in a second priority buffer when the maximum permitted number of accesses is exceeded, and provides the stored packet from the second priority buffer to the memory access logic in the next time interval.

[0016] In some examples, the apparatus includes a plurality of second priority buffers, and each of the plurality of second priority buffers corresponds to a different defined memory region. In a particular example, the controller stores an incoming packet of a type determined to include a read request in each second priority buffer based on an address associated with the incoming packet. In a particular example, the apparatus includes a bridge circuit that enables communication of packets between a physical layer interface and a memory fabric physical layer interface.

[0017] According to some embodiments, the apparatus includes a local memory and memory access logic, and the local memory can be configured as an addressable portion of addressable memory through a memory fabric physical layer interface. In certain embodiments, the physical layer interface receives incoming packets from a memory fabric physical layer interface having a data rate higher than the data rate of the physical layer interface, and at least a portion of the packets includes different instruction types. In some embodiments, the apparatus includes an incoming packet buffer structure including a hierarchically ordered priority buffer structure including at least a first priority buffer and a second priority buffer having a lower priority than the first priority buffer. In some examples, the apparatus includes a controller that determines the packet type of incoming packets from the memory fabric physical layer interface, and when the determined packet type indicates that an atomic request is in the incoming packet, the controller stores the incoming packet in the first priority buffer. When the determined packet type indicates that a load instruction is in the incoming packet, the controller stores the incoming packet in the second priority buffer and provides the stored incoming packet to the memory access logic in a hierarchical order according to the priority buffer order.

[0018] In a particular example, the apparatus includes a buffer, and the controller queues incoming packets determined to include a store request in the buffer and provides incoming packets having an atomic request to the memory access logic without buffering the atomic request type packets in a high priority buffer.

[0019] In some examples, the apparatus includes a configuration register that includes data defining at least a portion of the local memory as a priority memory region, where each memory region has a maximum unrestricted number of memory fabric physical layer interface accesses permitted per time interval. In certain embodiments, the controller maintains a count of the number of memory accesses made to the memory region defined by the memory fabric physical layer interface over a time interval and stores packets in a second priority buffer when the maximum permitted number of accesses is exceeded. In some embodiments, the controller provides the stored packets from the second priority buffer to the memory access logic in a next time interval.

[0020] In a particular example, the second priority buffer includes a plurality of second priority buffers, each of the plurality of second priority buffers corresponding to a different defined memory region, and the controller stores incoming packets of a type determined to include a read request in respective second priority buffers based on an address associated with the incoming packet.

[0021] According to some embodiments, the system includes a memory fabric physical layer interface operable to interconnect a plurality of distributed non-volatile memories with a first device and a second device. The first device and the second device may include a server, SoCs, or other devices. Each of the first device and the second device has a physical layer interface for receiving memory access requests from each other via the memory fabric physical layer interface. In some embodiments, the second device includes local memory such as DRAM operable to be coupled to memory access logic. The local memory is configured as an addressable portion of the distributed non-volatile memory that can be addressed through the memory fabric physical layer interface. In a particular example, the physical layer interface receives incoming packets from a memory fabric physical layer interface having a data rate higher than the data rate of the physical layer interface, and at least a portion of the packets are packets of different instruction types. In a particular example, the controller determines the packet type of the incoming packet received from the memory fabric physical layer interface, and when the determined incoming packet type is a packet type including an atomic request, the controller gives priority to the transfer of the incoming packet having the atomic request to the memory access logic over incoming packets of other packet types.

[0022] In a particular example, the second device includes a first priority buffer and a second priority buffer having a lower priority than the first priority buffer. In some examples, the controller gives priority to the transfer of the incoming packet having the atomic request over incoming packets of other packet types by queuing the incoming packet determined to include the atomic request in the first priority buffer and queuing other packet types in the second priority buffer. In some embodiments, the controller gives priority to the output of packets from the first priority buffer to the memory access logic over the output of packets from the second priority buffer.

[0023] In some examples, the second device includes a buffer, and the controller queues incoming packets determined to include a store request in the buffer and provides incoming packets having an atomic request to the memory access logic, thereby prioritizing the transfer of incoming packets having an atomic request over incoming packets of other packet types.

[0024] In certain examples, it includes a configuration register containing data that defines at least a portion of the local memory as a priority memory region, and each memory region has a maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval. In some examples, the controller maintains a count of the number of memory accesses made to the memory region defined by the memory fabric physical layer interface over a time interval, and stores read packets in a second priority buffer when the maximum permitted number of accesses is exceeded, and in the next time interval, provides the stored packets from the buffer to the memory access logic.

[0025] In some examples, the second priority buffer includes a plurality of second priority buffers, and each of the plurality of second priority buffers corresponds to a different defined memory region. In certain examples, the controller stores incoming packets of a type determined to include a read request in respective second priority buffers based on an address associated with the incoming packet.

[0026] FIG. 1 shows an example of a system 100, such as a cloud-based computing system within a data center or other system, including a plurality of devices 102 and 104 each coupled to a memory fabric physical layer interface 106 via a memory fabric bridge 108. In one example, the memory fabric physical layer interface 106 is implemented as a memory fabric switch including a memory fabric physical layer interface, such as a Gen-Z fabric (e.g., 802.3 PHY), or any other suitable memory fabric physical layer interface. The memory fabric physical layer interface 106 provides access to a fabric-attached memory 112, such as a media module having a media controller for DRAM, storage class memory (SCM), or other suitable memory, which is part of a fabric-attached memory pool addressable by devices 102 and 104. Also, the fabric-attached memory 112 may include a graphics processing unit, a field programmable gate array having local memory, or any other suitable module providing addressable memory addressable by devices 102 and 104 as part of a system-wide memory fabric. In some embodiments, the local memory on devices 102 and 104 is also a fabric-attached memory and is addressable and accessible by other processor-based devices within the system. In some embodiments, devices 102 and 104 are physical servers within a data center system.

[0027] In this example, device 102 is shown as including a processor system-on-chip that, in combination with a memory fabric bridge 108, serves as a type of fabric-attached memory module 114. However, any suitable implementation may be employed, including but not limited to integrated circuits, multiple package devices having multiple SoCs, fabric-attached memory modules having no processor or limited computing capabilities, physical servers or servers, or any other suitable device. It will be recognized that the various blocks may be combined or distributed in any suitable manner. For example, in one example, the memory fabric bridge 108 is an integrated circuit separate from the SoC, and in other embodiments, it is integrated as part of the SoC. In this example, device 102 serves as a type of fabric-attached memory module that includes local memory 116 accessible by device 104, and thus receives incoming packets from device 104 or other devices connected to the memory fabric physical layer interface 106, and the memory fabric physical layer interface 106 uses load requests, store requests, and atomic requests in relation to local memory 116.

[0028] Memory fabric bridge 108 can be any suitable bridge circuit. In one example, it includes an interface for connecting to the memory fabric physical layer interface 106 and another interface for connecting to the physical layer interface 120. Memory fabric bridge 108 can be an individual integrated circuit, integrated as part of a system-on-chip, or in another package. Also, it should be recognized that many variations of device 102 may be employed, including devices incorporating multiple SoCs, individual integrated circuits, multiple packages, or any other suitable configuration. For example, as described above, device 102 may alternatively be configured as a media module without a processor such as processor 128 or other computing logic. However, alternatively, it may include a media controller that serves as a memory controller for local memory 116 so that external devices such as other fabric-attached memories or other processor system-on-chip devices can access local memory 116 as part of the fabric-addressable memory.

[0029] In this example, fabric-attached memory module 114 includes a physical layer interface 120 that communicates with a controller 122 that serves as a memory fabric interface priority buffer controller. In a particular example, controller 122 is implemented as an integrated circuit such as a field programmable gate array (FPGA), but any suitable structure including an application specific integrated circuit (ASIC), a programmed processor that executes executable instructions stored in memory, a state machine, or any other suitable structure may be used. Controller 122 provides packets including atomic requests, load requests, and store requests to memory access logic 126.

[0030] In one example, the memory access logic 126 includes one or more memory controllers that process memory access requests provided by the controller 122 to access the local memory 116. The controller 122 serves data from the local memory 116 in response to incoming requests from remote computing units (e.g., the memory fabric physical layer interface 106, the fabric-attached memory 112, or the remote device 104 (e.g., remote node)) associated with the fabric connected to the fabric, and also serves data from the fabric-attached memory 112 in response to requests issued to the fabric from local computing units such as the processor 128. In this example, the fabric-attached memory module 114 includes one or more processors 128 that also use the local memory 116 via the memory access logic 126 and use the fabric-attached memory 112 via the memory fabric physical layer interface 106. In certain embodiments, the local memory 116 includes non-volatile random access memory (RAM) such as, but not limited to, DDR4 type dynamic random access memory and / or non-volatile dual in-line memory module (NVDIMM). However, any suitable memory may be used. The local memory 116 can be configured as an addressable portion of a distributed non-volatile memory that can be addressed via the memory fabric physical layer interface 106. Thus, the local memory 116 serves as a type of fabric-attached memory when accessed by remote computing units on the fabric.

[0031] The memory fabric physical layer interface 106 interconnects a plurality of other distributed non-volatile memories that are fabric-attached memory modules, and the fabric-attached memory modules are, in one example, but not limited to, DDR4 type dynamic random access memory, non-volatile dual in-line memory modules (NVDIMMs), NAND flash memory, storage class memory (SCM), or media modules configured as fabric-attached memory each including non-volatile random access memory (RAM) such as any other suitable memory.

[0032] The physical layer interface 120 receives incoming packets from the memory fabric physical layer interface 10 having a data rate higher than the data rate of the physical layer interface 120. Packets received by the physical layer interface include packets of different instruction types. In this example, the packets may be of a type having atomic requests, load requests, and / or store requests. Other packet types are also contemplated. The physical layer interface 120 in this example is described as a PCI Express (PCIe) physical layer interface, but any suitable physical layer interface may be employed. In this example, the memory fabric physical layer interface 106 has a data rate higher than the data rate of the physical layer interface 120. In one example, the data rate may be higher due to the rate of data transfer and / or the number of data link widths used for the interconnection. In this example, the physical layer interface 120 receives memory access requests from the device 104. Similarly, the device 102 may also make memory access requests to the device 104.

[0033] The fabric-attached memory module 114 employs an incoming packet buffer structure 130 that includes a hierarchically ordered priority buffer structure that includes different priority buffers, such as a FIFO buffer or other suitable buffer structure, etc., and some of the buffers have a higher priority than others. The fabric-attached memory module 114 in this example includes a buffer 132, such as a local buffer, that queues incoming packets determined to include store requests. Buffer 132 is considered a lower priority buffer in that writes are not considered important in many cases for the application being executed. Thus, the controller 122 stores store requests (e.g., writes) in buffer 132. The physical layer interface 120 has a lower data rate and a smaller link width than the memory fabric bridge 108, and thus buffering is performed before the physical layer interface 120 and after the memory fabric bridge 108. Packets having atomic requests and load requests are given a higher priority. For example, data is held in buffer 132 for a longer period of time while many other instructions can complete. If an error occurs when the write is finally executed, the controller 122 causes an asynchronous retry operation. In some examples, the incoming packet buffer structure 130 and buffer 132 are included in an integrated circuit that includes the controller 122.

[0034] In some embodiments, the memory access logic 126 may be implemented as a local data fabric that includes the CPU, its cache, and various other components for interfacing with the local memory 116. For example, the memory access logic 126 connects to one or more memory controllers, such as a DRAM controller, to access the local memory 116. Although some of the components are not shown, any suitable memory access logic, such as any suitable memory controller configuration for processing memory requests such as atomic requests, load requests, and store requests, or any other suitable logic, etc., may be employed.

[0035] FIG. 2 is a flowchart showing one example of a method 200 for managing packet transfer executed by device 102. In this example, the operations are executed by controller 122. In one embodiment, controller 122 is implemented as a field programmable gate array (FPGA) operating as described herein. However, it will be recognized that any suitable logic may be employed. Controller 122 receives incoming packets from memory fabric physical layer interface 106 via physical layer interface 120. As shown in block 202, the method includes determining the packet type of the incoming packets received from memory fabric physical layer interface 106. In one embodiment, the packet includes packet identification data within the packet that identifies the packet as having an atomic request, a load request, a store request, or any suitable combination thereof. Other packet types are also contemplated. In other examples, the packet may include an index or other data representing the packet type. In other embodiments, controller 122 determines the packet type based on data determined to be within the packet itself, added to the packet, or otherwise associated with the packet. In one embodiment, controller 122 evaluates the packet for a packet type identifier, which may be one or more bits, and determines the packet type of each incoming packet from the packet identifier.

[0036] As shown in block 204, when the determined incoming packet type is a packet type that includes an atomic request, the method includes prioritizing the transfer of incoming packets having an atomic request over incoming requests of other packet types such that the atomic request is given the highest priority. For example, when an incoming packet is determined to be an atomic request, the packet is passed directly to the memory access logic 126 for processing without buffering if the memory access logic has sufficient bandwidth to handle the packet. In another example, the controller 122 queues incoming atomic request type packets in a higher priority buffer, and the higher priority buffer is read out and provided to the memory access logic 126 before other packet types are provided to the memory access logic. The packet is ultimately provided to the memory access logic 126, but priority buffering is performed before the traffic enters the physical layer interface 120 and after it exits the memory fabric bridge 108.

[0037] For example, referring also to FIG. 1, in some embodiments, prioritizing the transfer of incoming packets having atomic requests includes queuing incoming packets determined to include atomic requests in the high-priority buffer 150 and queuing other packet types in the medium-priority buffer 152, where the first priority buffer has a higher priority than the second priority buffer. The controller 122 prioritizes the output of packets from the high-priority buffer 150 over the output of packets from the medium-priority buffer 152. In this example, this is done through a multiplexing operation. In some embodiments, prioritizing the transfer of incoming packets having atomic requests over incoming packets of other packet types includes queuing incoming packets determined to include store requests in buffer 132 and providing incoming packets having atomic requests to the memory access logic 126 without the need to store the atomic requests in the high-priority buffer 150. As described above, in some embodiments, this is done when the memory access logic 126 has the bandwidth capacity to receive incoming packets without further queuing.

[0038] The figures in FIG. 3 are examples that do not unduly limit the scope of the claims. Those skilled in the art will recognize many variations, alternatives, and modifications. There may be many alternatives, modifications, and variations in which the selected group of processings described above are used. For example, a portion of a processing may be extended and / or combined. Other processings may be inserted into those described above. Depending on the embodiment, the sequence of processings may be interchanged with the replaced ones.

[0039] FIG. 3 is a block diagram illustrating another example of a device that manages packet transfer by a memory fabric physical layer interface. In this example, an incoming packet buffer structure 130 is shown as having a plurality of medium priority buffers 300 and 302, each corresponding to a different defined memory region within local memory 116. A high priority buffer 150 serves as a higher priority buffer for storing atomic requests, and a multiplexer 304 (MUX) prioritizes the transfer of packets from the high priority buffer 150 over read requests stored in the medium priority buffers 300 and 302. The controller 122 may control the multiplexing operation of the multiplexer 304 as indicated by the dashed arrow 306, or the multiplexer 304 may include logic appropriate for outputting data in a high priority buffer entry before an entry from a lower priority buffer. The multiplexer 304 selects from higher priority buffers having entries and provides packets from those entries in a hierarchical priority scheme. The multiplexer 304 is considered to be part of the controller 122 in a particular example.

[0040] In one example, the local memory 116 is partitioned into memory regions by the controller 122, processor 128 under the control of an operating system or driver executing on the processor 128, or any other suitable mechanism. In one example, the controller 122 includes a configuration register 308 that contains data defining the start addresses and sizes of various memory regions within the local memory 116. In other examples, data representing start and end addresses is stored in the configuration register 308, or any other suitable data that defines one or more memory regions of the local memory 116. The controller 122 determines the various memory regions based on the data from the configuration register and organizes medium priority buffer 152 into each of the plurality of medium priority buffers 300 and 302 corresponding to the memory regions defined by the configuration register 308.

[0041] In this example, the physical layer interface 120 includes an uplink cross physical layer interface 310 and a downlink cross physical layer interface 312. The transmit packet buffer 314 enables the memory access logic 126 to output transmit packets to the memory fabric bridge 108 and thus to the memory fabric physical layer interface (e.g., switch) 106. The received packets shown as 318 are handled by a controller as described herein. In certain embodiments, the throughput may be reversed, and thus the downlink structure has similar buffering for transmit packets as shown for arriving traffic.

[0042] FIG. 4 is a flowchart illustrating one example of a method 400 for managing packet transfer according to one embodiment. As described above, the controller 122 determines whether packet traffic entering the fabric-attached memory module 114 from the memory fabric physical layer interface 106 is permitted to enter the memory access logic 126 such as the data fabric of the SoC. In this example, a priority level is selected based on the packet type and the destination memory region entering the SoC from the memory fabric physical layer. As shown in block 402, the controller 122 receives an incoming packet from the memory fabric physical layer interface 106 via the memory fabric bridge 108. As shown in block 404, the controller 122 determines the packet type and the destination memory region of the incoming packet. In one example, this is done by evaluating the data within the packet, and the data identifies the packet type and the destination memory region within the local memory 116 by the destination address or other appropriate information. As shown in block 406, if the packet type is an atomic type instruction packet, the controller 122 stores the atomic packet in a higher priority buffer (in this example, the high priority buffer 150). The controller allocates a medium priority buffer to include a plurality of priority buffers, and each of the plurality of priority buffers corresponds to a different defined memory region. The controller stores an incoming packet of the type determined to include a read request in the corresponding region priority buffer based on the memory address associated with the incoming packet.

[0043] However, as shown in block 408, when the packet is a store request packet such as a write request, the controller stores the packet in a local buffer such as buffer 132. As shown in block 410, when the packet is a load request and is determined to mean a load instruction or a read request, the controller 122 gives the load packet a priority lower than that of an atomic request but higher than that of a store instruction, stores the packet in the medium priority buffer 152, the high priority buffer 150 is a higher priority buffer, and buffer 132 is considered a low priority buffer.

[0044] Referring to FIG. 3, the controller 122 determines which of the medium priority buffers 300 and 302 stores the incoming packet based on the destination address of the packet. In one example, if the incoming packet has a destination address within an address region identified by data representing a data region defined by the configuration register 308, the controller places the incoming packet in the appropriate medium priority buffer for the incoming load request. In one example, the high priority buffer 150 and the medium priority buffer 152 are first in first out (FIFO) buffers, but any suitable buffer mechanism may be employed. Thus, the controller 122 accesses data defining the memory region, such as data stored in the configuration register 308 or other memory location. In some embodiments, an operating system and / or driver running on one or more processors may specify the memory region via the configuration register exposed by the controller. In other embodiments, the controller includes some predetermined designated memory regions such as those provided by firmware or microcode. However, any suitable embodiment may be employed.

[0045] Referring to FIG. 4, as shown in block 412, the method includes dispersing the packets in the high-priority buffer 150 before the medium-priority buffers 300 and 302 and other packets in buffer 132. In one example, this is performed by a multiplexer 304 that outputs packets from buffer entries according to priority. As shown in block 414, if the packet is a store request, the method includes determining whether the cross physical layer interface has sufficient bandwidth for a write access to pass through. In other words, the controller 122 determines whether the physical layer interface 120 (with a lower data rate and smaller link width than the memory fabric bridge 108) can pass the write packet through it. For example, if the physical layer interface 120 is saturated by the receiver of the MUX 304, the physical layer interface 120 starts to reject the packet. When this occurs, buffering of read requests as described above is started. If there is sufficient bandwidth for a write access to pass through, as shown in block 416, the method includes passing the write access through to the local memory 116. If there is not sufficient bandwidth for a write access to pass through, as shown in block 418, the method includes waiting until there is sufficient bandwidth and then passing the write operation according to the write request packet.

[0046] Referring to the medium-priority buffers 300 and 302, as shown in block 420, if the incoming packet is a load request (e.g., a read request) and there are no other high-priority incoming packets in the high-priority buffer 150, the method includes performing a load operation from memory by outputting the packet from the appropriate medium-priority buffers 300, 302 to the memory access logic 126. The method continues for the incoming packet.

[0047] As shown by block 422, when the upper limit set quality of the service mechanism is adopted, the method includes determining whether the corresponding memory area in the local memory 116 having the address of the packet is upper limit set during a time interval when the packet is a read type packet. If there are too many accesses to that memory area in the local memory 116 occurring within the time interval, the controller 122 stores the incoming packet in the medium priority buffer 152 until the next time interval occurs.

[0048] FIG. 5 is a block diagram showing another example of a device that manages packet transfer by a memory fabric physical layer interface. In this example, the controller 122 includes a packet analyzer 500, a priority buffer packet router 502, and a per memory area counter 504. The packet analyzer 500 analyzes the incoming packet to determine the packet type, for example, by scanning the content of the packet, interpreting the content, or evaluating the packet type data in the packet to determine whether the packet is a packet for an atomic request (for example, atomic instruction execution), a load request (for example, read instruction), or a store request (for example, write instruction), and outputs the packet 506 to the priority buffer packet router 502, and the priority buffer packet router 502 places the packet in an appropriately prioritized buffer. The packet 506 may be directly routed to an appropriate buffer or include a routing identifier indicating which priority buffer to place the packet in. As described above, the configuration register 308 defines at least some memory areas of the local memory 116 of the device as priority memory areas, and each memory area has a maximum unrestricted number of memory fabric physical layer interface accesses permitted per time interval. For example, the controller 122 is permitted to record some memory areas in the local memory 116 as priority memory areas via a configuration register (for example, a model specific register (MSR)).

[0049] Between the recorded memory regions and all other memory fabric-mapped memories, the controller 122 employs a type of service ceiling setting technique. For example, each of the recorded memory regions has a maximum unlimited number of fabric accesses that the memory region can receive per time interval, and the total ceiling for the entire local memory 116 mapped to the memory fabric space does not exceed the available cross-PHY capacity (e.g., limited by the PCIe channel width and data rate). Due to the fact that the allocated ceiling does not exceed the available channel capacity (excessive packets wait until the next time interval), the level of isolation from other memory regions within the fabric is high. Larger memory regions receive larger ceilings.

[0050] Each memory area counter 504 is incremented for each memory access per area, and the controller 122 compares the current count value with a threshold corresponding to the maximum unlimited number of fabric accesses per area. The threshold may be stored as part of configuration register information or other mechanism. Thus, the controller 122 maintains a count of the number of memory accesses made to the memory areas defined by the memory fabric physical layer interface over a time interval. The controller 122 stores the read packet (load packet) in the medium priority buffer 152 when the maximum allowable number of accesses is exceeded during the defined time interval. The controller stores the packet in the area buffer each time the MUX issues a higher priority packet (such as atomic). The upper limit determines whether to issue a packet from a given area buffer to the MUX. When the upper limit is reached, issuing to the MUX is temporarily stopped until the next time interval begins (a configurable controller parameter). In other words, the recorded memory areas defined by the configuration register are memory areas within the local memory 116. Thus, any recorded memory area has a buffer. In this example, the medium priority buffers 300 and 302 are associated with the defined memory areas, each storing accesses from the fabric to that memory area in the local memory 116. The upper limit determines whether to issue an access from a given buffer to the MUX, and thus the memory access can ultimately find its path to that memory area within the local memory 116. When the upper limit is reached, issuing to the MUX is temporarily stopped until the next time interval begins.

[0051] In the next time interval, the stored packets are provided from the medium priority buffer 152 to the memory access logic 126. In this example, the controller 122 stores incoming packets of the type determined to include read requests in each of the plurality of medium priority buffers 300 and 302 based on the address associated with the incoming packet, and also tracks the number of memory accesses within a particular memory area so as to set an upper limit on the number of accesses occurring within the memory area.

[0052] Figure 6 is a flowchart showing one example of a method 600 for managing packets according to one example. As shown in block 602, the method includes defining at least a portion of the memory area of the local memory as a priority memory area, where each memory area has a maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval. As shown in block 604, the method includes maintaining a count of the number of memory accesses made through the defined memory area of the local memory by the memory fabric physical layer interface over a time interval. As shown in block 606, the method includes storing a packet in a priority buffer when the maximum allowable number of accesses is exceeded, and as shown in block 608, the method includes providing the stored packet from the buffer to the memory access logic within the next time interval.

[0053] Among other technical advantages, the adoption of a controller and buffer hierarchy reduces the latency of access to addressable local memory as part of the memory fabric and improves energy efficiency. Also, when the fabric-attached memory module includes one or more processors, the controller offloads packet management from the one or more processors, thus increasing the speed of processor operations. Other advantages will be recognized by those of ordinary skill in the art.

[0054] Features and elements have been described in specific combinations, but each feature and element may be used alone or in various combinations with or without other characteristics and elements. In some embodiments, the apparatus described herein is manufactured using a computer program, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of computer-readable storage media include magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks, and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs).

[0055] In the foregoing detailed description of various embodiments, reference has been made to the accompanying drawings which form a part hereof, and in which are shown by way of illustration specific preferred embodiments in which the invention may be practiced. These embodiments have been described in sufficient detail to enable those skilled in the art to practice the invention, and it is to be understood that other embodiments may be utilized and that logical, mechanical, and electrical changes may be made without departing from the scope of the invention. To avoid unnecessary detail that is not necessary for those skilled in the art to practice the invention, the description may omit certain information known to those skilled in the art. Further, various embodiments incorporating the teachings of the present disclosure may be readily constructed by those skilled in the art. Accordingly, the invention is not intended to be limited to the specific forms shown herein but rather is intended to cover alternatives, modifications, and equivalents as may be reasonably included within the scope of the invention. Accordingly, the foregoing detailed description is not to be construed in a limiting sense, and the scope of the invention is defined only by the appended claims. The above detailed description of the embodiments and the examples described therein are presented for purposes of illustration and explanation only and are not limiting. For example, the operations described may be performed in any suitable order or manner. Accordingly, the invention encompasses any modification, variation, or equivalent falling within the scope of the basic principles disclosed herein and claimed hereinbelow.

Claims

1. A method for managing packet transfer executed by a device coupled to a memory fabric physical layer interface having a data rate higher than the data rate of a physical layer interface of the device, the method comprising: Receiving incoming packets from the memory fabric physical layer interface, wherein at least a portion of the incoming packets includes different instruction types; Determining a packet type of the incoming packets received from the memory fabric physical layer interface; Prioritizing the transfer of the incoming packets having the atomic request to memory access logic that accesses local memory within the device over incoming packets of other packet types when the determined incoming packet type is a packet type that includes an atomic request. A method.

2. Prioritizing the transfer of the incoming packets having the atomic request over incoming packets of other packet types comprises: Queueing the incoming packets determined to include the atomic request in a first priority buffer; Queueing incoming packets of other packet types in a second priority buffer; Prioritizing the output of packets from the first priority buffer over the output of packets from the second priority buffer. The method of Claim 1.

3. Accessing data that defines at least a plurality of memory regions of the local memory of the device as priority memory regions, each memory region having a maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval; Maintaining a count of the number of memory accesses made to the memory regions defined by the memory fabric physical layer interface over the time interval; Storing read packets in the second priority buffer when the maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval is exceeded; Further comprising providing the stored packets from the second priority buffer to the memory access logic in a next time interval. The method of Claim 2.

4. Prioritizing the transfer of the incoming packets having the atomic request over incoming packets of other packet types comprises: Queueing in a buffer the incoming packets determined to include a store request and providing the incoming packets having the atomic request to the memory access logic The method of claim 1 **Claim 5** Allocating the second priority buffer to include a plurality of second priority buffers, each of the plurality of second priority buffers corresponding to a different defined memory area Based on the address associated with the incoming packet, storing the incoming packet of the type determined to include a read request in each of the second priority buffers corresponding to the different defined memory areas The method of claim 3 **Claim 6** One or more processors Memory access logic operably coupled to the one or more processors A local memory operably coupled to the memory access logic and configurable as an addressable portion of the addressable memory via a memory fabric physical layer interface A physical layer interface operably coupled to the memory access logic, operable to receive incoming packets from the memory fabric physical layer interface having a data rate higher than the data rate of the physical layer interface, at least a portion of the incoming packets including different instruction types A controller operably coupled to the physical layer interface The controller Determining the packet type of the incoming packet received from the memory fabric physical layer interface When the determined incoming packet type is a packet type including an atomic request, prioritizing the transfer of the incoming packet having the atomic request to the memory access logic over incoming packets of other packet types Configured to perform Device **Claim 7** Comprising a first priority buffer and a second priority buffer having a lower priority than the first priority buffer The controller By queuing the incoming packet determined to include the atomic request in the first priority buffer, prioritizing the transfer of the incoming packet having the atomic request over incoming packets of other packet types queuing other packet types in the second priority buffer; prioritizing the output of packets from the first priority buffer over the output of packets from the second priority buffer; configured to perform; the apparatus of claim 6. **Claim 8** comprising a buffer operably coupled to the controller, wherein the controller is configured to queue incoming packets determined to include a store request in the buffer and provide the incoming packets having the atomic request to the memory access logic, thereby prioritizing the transfer of the incoming packets having the atomic request over incoming packets of other packet types; the apparatus of claim 6. **Claim 9** comprising a configuration register configured to include data defining at least a plurality of memory regions of the local memory of the apparatus as priority memory regions, each memory region having a maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval, wherein the controller is configured to maintain a count of the number of memory accesses made to the memory regions defined by the memory fabric physical layer interface over the time interval, store read packets in the second priority buffer when the maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval is exceeded, and provide the stored packets from the second priority buffer to the memory access logic in a next time interval; configured to perform; the apparatus of claim 7. **Claim 10** wherein the second priority buffer includes a plurality of second priority buffers, each of the plurality of second priority buffers corresponding to a different defined memory region, and the controller is configured to store incoming packets of a type determined to include a read request in respective second priority buffers based on an address associated with the incoming packets; the apparatus of claim 9. **Claim 11** comprising a memory fabric bridge circuit operably coupled to the controller and the memory fabric physical layer interface and operable to communicate packets between the physical layer interface and the memory fabric physical layer interface; the apparatus of claim 6.

12. A local memory operable to a memory access logic and configurable as an addressable portion of an addressable memory via a memory fabric physical layer interface, and A physical layer interface operable to receive incoming packets from the memory fabric physical layer interface having a data rate higher than the data rate of the physical layer interface, wherein at least a part of the incoming packets includes different instruction types, and A hierarchically ordered incoming packet buffer structure including at least a first priority buffer and a second priority buffer having a lower priority than the first priority buffer, and A controller operably coupled to the physical layer interface and the incoming packet buffer structure, comprising The controller is Determining a packet type of the incoming packet from the memory fabric physical layer interface; When the determined packet type indicates that an atomic request is in the incoming packet, storing the incoming packet in the first priority buffer; When the determined packet type indicates that a load instruction is in the incoming packet, storing the incoming packet in the second priority buffer; Providing the stored incoming packet to the memory access logic in a hierarchical order according to the priority buffer order; Is configured to perform Device.

13. Comprising a buffer operably coupled to the controller, The controller is configured to prioritize the transfer of the incoming packet having the atomic request over incoming packets of other packet types by queuing the incoming packet determined to include a store request in the buffer and providing the incoming packet having the atomic request to the memory access logic, The device of claim 12.

14. Comprising a configuration register configured to include data defining at least a plurality of memory regions of the local memory of the device as priority memory regions, each memory region having a maximum unlimited number of memory fabric physical layer interface accesses permitted per time interval, The controller is Maintaining a count of the number of memory accesses made to a memory area defined by the memory fabric physical layer interface over the time interval; Storing a packet in the second priority buffer when the maximum unrestricted number of memory fabric physical layer interface accesses permitted per time interval is exceeded; Providing, in a next time interval, the stored packet in the second priority buffer from the buffer to the memory access logic; is configured to perform; The apparatus of claim 13. **Claim 15** The second priority buffer includes a plurality of second priority buffers, each of the plurality of second priority buffers corresponding to a different defined memory area, The controller is configured to store an incoming packet of a type determined to include a read request in respective second priority buffers based on an address associated with the incoming packet. The apparatus of claim 12. **Claim 16** A memory fabric operable to interconnect a plurality of distributed non-volatile memories; A first device operably coupled to the memory fabric; A second device operably coupled to the memory fabric, the first device and the second device having a physical layer interface for receiving memory access requests from each other via the memory fabric; The second device, A local memory operably coupled to the memory access logic of the second device, the local memory being configurable as an addressable portion of the distributed non-volatile memory addressable via the memory fabric; A physical layer interface operably coupled to the memory access logic, the physical layer interface being operable to receive incoming packets from the memory fabric having a data rate higher than the data rate of the physical layer interface, at least a portion of the incoming packets including different instruction types; A controller operably coupled to the physical layer interface; The controller, Determining the packet type of the incoming packet received from the memory fabric; When the determined incoming packet type is a packet type including an atomic request, transfer the incoming packet having the atomic request to the memory access logic with higher priority than incoming packets of other packet types, configured to perform, a system.

17. The second device includes a first priority buffer and a second priority buffer having a lower priority than the first priority buffer, The controller, by queuing an incoming packet determined to include an atomic request in the first priority buffer, transfer the incoming packet having the atomic request with higher priority than incoming packets of other packet types, queuing other packet types in the second priority buffer, prioritizing the output of packets from the first priority buffer over the output of packets from the second priority buffer, configured to perform, the system of claim 16.

18. The second device includes a buffer operably coupled to the controller, The controller is configured to queue an incoming packet determined to include a store request in the buffer and provide the incoming packet having the atomic request to the memory access logic, thereby prioritizing the transfer of the incoming packet having the atomic request over incoming packets of other packet types, the system of claim 16.

19. The second device, includes a configuration register configured to include data defining at least a plurality of memory regions of the local memory of the second device as priority memory regions, each memory region having a maximum unrestricted number of memory fabric accesses permitted per time interval, maintaining a count of the number of memory accesses made to the memory regions defined by the memory fabric over the time interval, when exceeding the maximum unrestricted number of memory fabric accesses permitted per time interval, storing a read packet in the second priority buffer, in a next time interval, providing the stored packet from the second priority buffer to the memory access logic, configured to perform, the system of claim 17.

20. The second priority buffer includes a plurality of second priority buffers, and each of the plurality of second priority buffers corresponds to a different defined memory area. The controller is configured to store incoming packets of a type determined to include a read request in respective second priority buffers based on an address associated with the incoming packet. The system of claim 19.

Citation Information

Patent Citations

  • Assisted coherent shared memory

    JP2015127949A

  • Virtual channel for data transfers between devices

    US20140344488A1

  • Technologies for scalable remotely accessible memory segments

    US20160314073A1

  • Simultaneous execution of compute and graphics applications

    US9665920B1