Apparatus and method for managing packet transfer across a memory fabric physical layer interface

By managing packet transmission between physical layer interfaces of the memory architecture, and employing hardware-assisted automated priority partitioning and hierarchical ordered buffer structures, the traffic bottleneck caused by the difference in data rate and link width between the PCIe interface and the Gen-Z memory architecture interface in the memory architecture is resolved, thereby improving the data transmission efficiency and atomic request response speed of the data center system.

CN115004163BActive Publication Date: 2026-04-21ADVANCED MICRO DEVICES INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ADVANCED MICRO DEVICES INC
Filing Date
2020-10-02
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing memory architecture physical layer interfaces suffer from traffic bottlenecks in data center systems, particularly packet traffic bottlenecks caused by differences in data rates and link widths between PCIe interfaces and Gen-Z memory architecture interfaces, which affect data transmission efficiency.

Method used

An apparatus and method are employed to manage packet transmission between physical layer interfaces of memory architectures, utilize hardware-assisted automated priority partitioning technology to prioritize packets containing atomic requests, and ensure high priority for atomic requests by queuing and transmitting different types of packets through a hierarchical ordered priority buffer structure.

Benefits of technology

It improves the data transmission efficiency of the memory architecture in the data center system, reduces packet transmission latency, ensures a fast response to atomic requests, and enhances the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004163B_ABST
    Figure CN115004163B_ABST
Patent Text Reader

Abstract

An apparatus and method for managing packet transfer between memory fabrics having physical layer interfaces receives incoming packets from the memory fabric physical layer interfaces, the physical layer interfaces having a higher data rate than the data rate of the physical layer interfaces of another device, wherein at least some of the packets include different instruction types. The apparatus and method determines the packet type of the incoming packets received from the memory fabric physical layer interfaces, and when the determined incoming packet type is a type containing an atomic request, the method and apparatus causes the incoming packet having the atomic request to be transferred to memory access logic, which accesses local memory within the apparatus, in preference to other packet types of incoming packets.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Government licensing rights

[0002] This invention was carried out with government support under the PathForward project (basic contract number DE-AC52-07NA27344, subcontract number B620717) awarded by the Lawrence Livermore National Security Agency to the U.S. Department of Energy (DOE). The government enjoys certain rights in this invention. Background of the Invention

[0004] Systems employing memory-semantic architectures are being adopted that extend the byte-addressable load-store model of central processing unit (CPU) memory to the entire system, such as data centers. A memory architecture is a type of point-to-point communication switch (also known as the Gen-Z architecture) located outside the processor system-on-a-chip (SoC), media modules, and other types of devices that allow devices to interface with pools of external memory modules through the memory architecture in systems such as data centers. For example, some processor SoCs include a processor containing multiple processing cores that communicate with local memory, such as dynamic random access memory (DRAM) or other suitable memory, via local memory access logic, such as a data architecture. The processor SoC and other devices also need to interface with the memory architecture to use architecture-attached memory (FAM) modules, which may be external (e.g., non-local) memory, for example, directly attached to the data center memory architecture. In some systems, the FAM module has memory access logic for handling load and store requests but has little or no computational power. Furthermore, the memory architecture attaches the FAM module as an addressable portion of the entire main memory. FAM module use cases enable decomposed memory pools in cloud data centers. With FAM modules present, hosts are not constrained by the storage capacity limitations of their local servers. Instead, hosts gain access to a large pool of memory not attached to any other host. Hosts coordinate to partition the memory among themselves or share FAM modules. The Gen-Z architecture has emerged as a high-performance, low-latency memory-semantic architecture that can be used to communicate with every device in the system.

[0005] Improved devices and methods are needed for managing traffic across the physical layer interface of memory architectures that employ architecture-attached memory. Attached Figure Description

[0006] With the accompanying drawings, the implementation will be more readily understood from the following description, wherein the same reference numerals denote the same elements, and in the drawings:

[0007] Figure 1 is a block diagram illustrating a system employing a device for managing packet transmission over a cross-physical layer interface using a memory architecture, according to an example set forth in this disclosure.

[0008] Figure 2 is a flowchart illustrating a method for managing packet transmissions performed by means of a device coupled to a physical layer interface of a memory architecture, according to an example set forth in this disclosure.

[0009] Figure 3 is a block diagram illustrating an example of a device for managing packet transmission according to the present disclosure;

[0010] Figure 4 is a flowchart illustrating a method for managing packet transmissions performed by means of a device coupled to a physical layer interface of a memory architecture, according to an example set forth in this disclosure.

[0011] Figure 5 is a block diagram illustrating an example of a device for managing packet transmission according to this disclosure; and

[0012] Figure 6 is a flowchart illustrating a method for managing packet transmissions performed by means of a device coupled to a physical layer interface of a memory architecture, according to an example set forth in this disclosure. Detailed Implementation

[0013] Memory architectures can become traffic bottlenecks. The physical layer interface of a memory architecture (also known as the memory architecture physical layer (PHY) interface) operates at higher performance than physical layer interfaces associated with on-chip systems (e.g., host SoCs) or other devices connected to the memory architecture PHY interface. For example, the signaling standards used for access to the FAM and message passing through the memory architecture can reach approximately 56 Gt / s, compared to 16 or 32 Gt / s for peripheral interconnects such as PCIe interfaces used on the SoC. Furthermore, link widths for memory architectures are designed to be larger. While some current processor SoCs that interface to the PCI-e bus use first-in-first-out (FIFO) buffers to queue packet traffic, the difference in data rates and link widths across PHY interfaces, such as the PCI-e physical layer (PHY) interface to the memory architecture PHY interface, remains a potential bottleneck for packet traffic.

[0014] In some implementations, a device serves as an interface to manage traffic prioritization at connection points between multiple physical layer interfaces, such as PCIe PHY interfaces and local memory architecture PHY interfaces (such as Gen-Z802.3 type memory architecture interfaces). In some implementations, the device provides hardware-assisted automated prioritization of packet traffic across PHY interfaces for data center workloads. In some implementations, the device offloads cross-PHY interface optimization from host CPUs in data centers or other systems employing memory architecture physical layer interfaces.

[0015] In some implementations, an apparatus and method for managing packet transfers between memory architectures having physical layer interfaces receive incoming packets from a memory architecture physical layer interface having a higher data rate than the physical layer interface of another device, wherein at least some of the packets include different instruction types. The apparatus and method determine the packet type of the incoming packets received from the memory architecture physical layer interface, and when the determined incoming packet type is one containing an atomic request, the method and apparatus cause the incoming packet with the atomic request to be transferred to memory access logic, which accesses local memory within the device, prioritizing other packet types of incoming packets.

[0016] In some instances, the method includes queuing incoming packets determined to contain atomic requests in a first priority buffer and queuing other packet types in a second priority buffer. The method also includes prioritizing the output of packets from the first priority buffer over the output of packets from the second priority buffer. In some instances, the method includes queuing incoming packets determined to contain memory requests in a buffer while simultaneously providing incoming packets with atomic requests to the memory access logic.

[0017] In some instances, the method includes: accessing data, such as from one or more configuration registers, which defines at least some memory regions of the device's local memory as priority memory regions, wherein each memory region allows an unlimited maximum number of memory architecture physical layer interface (MIBMI) accesses per time interval; and maintaining a count of the number of memory accesses made to the defined memory regions via the MIBMI within the time interval. The method further includes: storing a read packet in a second priority buffer when the maximum allowed number of accesses is exceeded; and providing the stored packet from the second priority buffer to memory access logic in the next time interval.

[0018] In some instances, the method includes allocating a second priority buffer to include a plurality of second priority buffers, each of the plurality of second priority buffers corresponding to a different defined memory region, wherein an incoming packet, determined to be of the type containing a read request based on the address associated with the incoming packet, is stored in the corresponding of the plurality of second priority buffers.

[0019] According to some implementations, a device includes one or more processors and memory access logic also coupled to local memory, wherein the local memory is configurable as an addressable portion of memory addressable via a memory architecture physical layer interface. In some implementations, the physical layer interface receives incoming packets from a memory architecture physical layer interface having a higher data rate than the physical layer interface itself, at least some of the packets including different instruction types. In some implementations, a controller determines the packet type of the incoming packets received from the memory architecture physical layer interface, and when the determined incoming packet type is one containing an atomic request, the controller causes the incoming packet with the atomic request to be transmitted to the memory access logic with priority over other packet types of incoming packets.

[0020] In some instances, the device includes a first priority buffer and a second priority buffer with a lower priority than the first priority buffer, and the controller prioritizes incoming packets with atomic requests over other packet types by queuing incoming packets determined to contain atomic requests in the first priority buffer and queuing other packet types in the second priority buffer. In some instances, the controller prioritizes the output of packets from the first priority buffer over the output of packets from the second priority buffer.

[0021] In some instances, the device includes a buffer, and the controller prioritizes incoming packets with atomic requests over other packet types by queuing incoming packets identified as containing storage requests in the buffer while providing incoming packets with atomic requests to the memory access logic.

[0022] In some instances, the device includes a configuration register storing data that defines at least some memory regions of local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture physical layer interface accesses per time interval. In some implementations, the controller maintains a count of the number of memory accesses to the defined memory regions via the memory architecture physical layer interface within the time interval, and when the maximum allowed number of accesses is exceeded, a read packet is stored in a second priority buffer and the stored packet is provided from the second priority buffer to the memory access logic in the next time interval.

[0023] In some instances, the device includes multiple second-priority buffers, each corresponding to a different defined memory region. In some instances, the controller stores incoming packets determined to contain a read request type in a corresponding second-priority buffer based on the address associated with the incoming packet. In some instances, the device includes a bridge circuit that allows packet transfer between the physical layer interface and the memory architecture physical layer interface.

[0024] According to some implementations, a device includes local memory and memory access logic, wherein the local memory is configurable as an addressable portion of the memory addressable via a memory architecture physical layer interface. In some implementations, the physical layer interface receives incoming packets from a memory architecture physical layer interface having a higher data rate than the physical layer interface itself, at least some of the packets including different instruction types. In some implementations, the device includes an incoming packet buffer structure comprising a hierarchically ordered priority buffer structure, the priority buffer structure including at least a first priority buffer and a second priority buffer having a priority lower than the first priority buffer. In some instances, the device includes a controller that determines the packet type of incoming packets from the memory architecture physical layer interface, and when the determined packet type indicates that an atomic request exists in the incoming packet, the controller stores the incoming packet in the first priority buffer. When the determined packet type indicates that a load instruction exists in the incoming packet, the controller stores the incoming packet in the second priority buffer and provides the stored incoming packets to the memory access logic in hierarchical order according to the priority buffer order.

[0025] In some instances, the device includes a buffer, and the controller will determine that incoming packets containing storage requests are queued in the buffer, while incoming packets with atomic requests are provided to the memory access logic, without buffering atomic request type packets in a high-priority buffer.

[0026] In some instances, the device includes a configuration register containing data that defines at least some memory regions of local memory as priority memory regions, where each memory region is allowed an unlimited maximum number of memory architecture physical layer interface accesses per time interval. In some implementations, the controller maintains a count of the number of memory accesses to the defined memory regions via the memory architecture physical layer interface within the time interval, and stores packets in a second priority buffer when the maximum allowed number of accesses is exceeded. In some implementations, the controller provides the packets stored in the second priority buffer from the buffer to the memory access logic in the next time interval.

[0027] In some instances, the second priority buffer includes multiple second priority buffers, each of which corresponds to a different defined memory region, and the controller stores incoming packets that are determined to be of the type containing a read request in the corresponding second priority buffer based on the address associated with the incoming packet.

[0028] According to some implementations, a system includes a memory architecture physical layer interface (PLI) that operates to interconnect a plurality of distributed non-volatile memories with a first device and a second device. The first and second devices may include servers, SoCs, or other devices. Each of the first and second devices has a physical layer interface to receive memory access requests from each other via the PLI. In some implementations, the second device includes local memory, such as DRAM operatively coupled to memory access logic. The local memory may be configured as an addressable portion of the distributed non-volatile memory addressable via the PLI. In some instances, the physical layer interface receives incoming packets from a PLI having a higher data rate than the physical layer interface, at least some of which are different instruction types. In some instances, a controller determines the packet type of the incoming packets received from the PLI, and when the determined incoming packet type is one containing an atomic request, the controller causes the incoming packet with the atomic request to be transmitted to the memory access logic prior to other packet types of incoming packets.

[0029] In some instances, the second device includes a first priority buffer and a second priority buffer having a lower priority than the first priority buffer. In some instances, the controller prioritizes incoming packets with atomic requests over other packet types by queuing incoming packets determined to contain atomic requests in the first priority buffer and queuing other packet types in the second priority buffer. In some implementations, the controller operates to prioritize the output of packets from the first priority buffer over the output of packets from the second priority buffer before proceeding to memory access logic.

[0030] In some instances, the second device includes a buffer, and the controller prioritizes incoming packets with atomic requests over other packet types by queuing incoming packets identified as containing storage requests in the buffer while providing incoming packets with atomic requests to the memory access logic.

[0031] In some instances, the second device includes a configuration register containing data that defines at least some memory regions of the local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture physical layer interface accesses per time interval. In some instances, the controller maintains a count of the number of memory accesses made to the defined memory regions via the memory architecture physical layer interface within the time interval, and when the maximum allowed number of accesses is exceeded, stores a read packet in a second priority buffer and provides the stored packet from the second priority buffer to the memory access logic in the next time interval.

[0032] In some instances, the second priority buffer comprises multiple second priority buffers, each corresponding to a different defined memory region. In some instances, the controller stores incoming packets determined to contain a read request type in the appropriate second priority buffer based on the address associated with the incoming packet.

[0033] Figure 1 illustrates an example of system 100, such as a cloud-based computing system or other system in a data center, comprising multiple devices 102 and 104, each coupled to a memory architecture physical layer interface 106 via a memory architecture bridge 108. In one example, the memory architecture physical layer interface 106 is implemented as a memory architecture switch, which includes a memory architecture physical layer interface such as a Gen-Z architecture (e.g., 802.3PHY) or any other suitable memory architecture physical layer interface. The memory architecture physical layer interface 106 provides access to architecture-attached memory 112, such as a media module having a media controller and other suitable memory such as DRAM, storage-class memory (SCM), or as part of a pool of architecture-attached memory addressable by devices 102 and 104, which is a non-volatile memory. The architecture-attached memory 112 may also include a graphics processing unit, a field-programmable gate array with local memory, or any other suitable module that provides addressable memory addressable by devices 102 and 104 as part of the overall system memory architecture. In some implementations, the local memory on devices 102 and 104 is also architecture-attached memory and can be addressed and accessed by other processor-based devices in the system. In some implementations, devices 102 and 104 are physical servers in a data center system.

[0034] In this example, device 102 is shown as a combination of a processor system-on-a-chip (SoC) serving as a type of architecture-attached memory module 114 and a memory architecture bridge 108. However, any suitable implementation may be employed, such as, but not limited to: integrated circuits, multiple packaged devices with multiple SoCs, architecture-attached memory modules without processors or with limited computing power, one or more physical servers, or any other suitable device. It will also be appreciated that various frames can be combined or allocated in any suitable manner. For example, the memory architecture bridge 108 is an integrated circuit separate from the SoC in one instance, while in other implementations it is integrated as part of the SoC. In this example, device 102 serves as a type of architecture-attached memory module including local memory 116 accessible by device 104, and therefore receives incoming packets from device 104 or other devices connected to the memory architecture physical layer interface 106 using load requests, storage requests, and atomic requests related to local memory 116.

[0035] The memory architecture bridge 108 can be any suitable bridge circuit, and in one instance includes an interface connected to the memory architecture physical layer interface 106 and another interface connected to the physical layer interface 120. The memory architecture bridge 108 can be a separate integrated circuit, or it can be integrated as part of a system-on-a-chip (SoC) or a separate package. Furthermore, it will be appreciated that many variations of the device 102 can be employed, including devices incorporating multiple SoCs, separate integrated circuits, multiple packages, or any other suitable configuration as needed. For example, as described above, the device 102 can alternatively be configured not to include a processor or other computing logic such as processor 128, but instead may include a media module, such as a media controller, which acts as a memory controller for local memory 116, allowing external devices such as other architecture-attached memory or other processor SoC devices to access local memory 116 as part of architecture-addressable memory.

[0036] In this example, the architecture-attached memory module 114 includes a physical layer interface 120 that communicates with a controller 122, which acts as a memory architecture interface priority buffer controller. In some instances, the controller 122 is implemented as an integrated circuit, such as a field-programmable gate array (FPGA), but any suitable architecture can be used, including an application-specific integrated circuit (ASIC), a programmable processor that executes executable instructions stored in memory, a state machine, or any suitable architecture. The controller 122 provides the memory access logic 126 with groups of requests, including atomic requests, load requests, and store requests.

[0037] In one example, memory access logic 126 includes one or more memory controllers that process memory access requests provided by controller 122 to access local memory 116. Controller 122 provides data from local memory 116 in response to requests from remote computing units on the architecture (e.g., associated with memory architecture physical layer interface 106 or architecture-attached memory 112) or remote devices 104 (e.g., remote nodes) connected to the architecture, and also provides data from architecture-attached memory 112 in response to requests made to the architecture by local computing units such as processor 128. In this example, architecture-attached memory module 114 also includes one or more processors 128 that also use local memory 116 via memory access logic 126 and architecture-attached memory 112 via memory architecture physical layer interface 106. In some implementations, local memory 116 includes non-volatile random access memory (RAM), such as, but not limited to, DDR4-based dynamic random access memory and / or non-volatile dual in-line memory modules (NVDIMMs). However, any suitable memory can be used. Local memory 116 can be configured as an addressable portion of distributed non-volatile memory addressable via memory architecture physical layer interface 106. Therefore, local memory 116 serves as a type of architecture-attached memory when accessed by remote computing units on the architecture.

[0038] The memory architecture physical layer interface 106 interconnects multiple other distributed non-volatile memories as architecture-attached memory modules, which in one instance are media modules configured as architecture-attached memories, each of which includes non-volatile random access memory (RAM), such as, but not limited to, DDR4-based dynamic random access memory, non-volatile dual in-line memory modules (NVDIMMs), NAND flash memory, storage-class memory (SCM), or any other suitable memory.

[0039] Physical layer interface 120 receives incoming packets from memory architecture physical layer interface 106, which has a higher data rate than physical layer interface 120. Packets received by the physical layer interface include packets of different instruction types. In this example, packets may be of types with atomic requests, load requests, and / or store requests. Other packet types are also contemplated. Physical layer interface 120 will be described as a PCI Express (PCIe) physical layer interface in this example, but any suitable physical layer interface may be used. In this example, memory architecture physical layer interface 106 has a higher data rate than physical layer interface 120. In one example, the higher data rate is due to the data transfer rate and / or the amount of data link width used for interconnection. In this example, physical layer interface 120 receives memory access requests from device 104. Similarly, device 102 may also issue memory access requests to device 104.

[0040] The architecture-attached memory module 114 employs an incoming packet buffer structure 130, which includes a hierarchically ordered priority buffer structure comprising different priority buffers, such as FIFO buffers or other suitable buffer structures, where some buffers have higher priorities than others. In this example, the architecture-attached memory module 114 also includes a buffer 132, such as a local buffer, for queuing incoming packets determined to contain storage requests. Buffer 132 is considered a lower priority buffer because writes are not considered critical for the executing application in many cases. Therefore, the controller 122 holds storage requests (e.g., writes) in buffer 132. The physical layer interface 120 has a smaller data rate and smaller link width than the memory architecture bridge 108, thus buffering occurs before the physical layer interface 120 and after the memory architecture bridge 108. Packets with atomic requests and load requests are given higher priority. While many other instructions can be completed, write data is held in buffer 132 for, for example, a longer period of time. If an error occurs during the final write operation, controller 122 generates an asynchronous retry operation. In some instances, incoming packet buffer structure 130 and buffer 132 are included in an integrated circuit that includes controller 122.

[0041] In some implementations, memory access logic 126 may be implemented as a local data architecture that includes various other components for interfacing with the CPU, its cache, and local memory 116. For example, memory access logic 126 may be connected to one or more memory controllers, such as a DRAM controller, to access local memory 116. Although some components are not shown, any suitable memory access logic may be used, such as any suitable memory controller configuration or any other suitable logic for handling memory requests such as atomic requests, load requests, and store requests.

[0042] Figure 2 is a flowchart illustrating an example of a method 200 for managing packet transmissions performed by device 102. In this example, the operation is performed by controller 122. In one implementation, controller 122 is implemented as a field-programmable gate array (FPGA) operating as described herein. However, it will be appreciated that any suitable logic may be employed. Controller 122 receives incoming packets from memory architecture physical layer interface 106 via physical layer interface 120. As shown in block 202, the method includes determining the packet type of the incoming packets received from memory architecture physical layer interface 106. In one implementation, the packet includes packet identification data within the packet that identifies the packet as having an atomic request, load request, storage request, or any suitable combination thereof. Other packet types are also contemplated. In other instances, the packet may include an index or other data indicating the packet type. In other implementations, controller 122 determines the packet type based on whether the data is within the packet itself, appended to the packet, or separately determined to be associated with the packet. In one implementation, controller 122 evaluates packets for a packet type identifier that may be one or more bits, and determines the packet type of each incoming packet based on the packet identifier.

[0043] As shown in box 204, when the determined incoming packet type is one containing an atomic request, the method includes prioritizing incoming packets with atomic requests over other packet types with incoming requests, such that atomic requests are given the highest priority. For example, if an incoming packet is determined to be an atomic request, the packet is passed directly to the memory access logic 126 for processing without buffering, provided that the memory access logic 126 has sufficient bandwidth to process the packet. In another instance, the controller 122 queues incoming atomic request type packets in a higher priority buffer, which are read and provided to the memory access logic 126 before other packet types are provided. The packet is eventually provided to the memory access logic 126, but priority buffering occurs before the traffic enters the physical layer interface 120 and after the traffic leaves the memory architecture bridge 108.

[0044] For example, and referring simultaneously to FIG1, in some implementations, prioritizing the transmission of incoming packets with atomic requests involves queuing incoming packets identified as containing atomic requests in a high-priority buffer 150 and queuing other packet types in a medium-priority buffer 152, where the first priority buffer has a higher priority than the second priority buffer. Controller 122 prioritizes the output of packets from the high-priority buffer 150 over the output of packets from the medium-priority buffer 152. In one instance, this is accomplished through multiplexing. In some implementations, prioritizing the transmission of incoming packets with atomic requests over other packet types involves queuing incoming packets identified as containing storage requests in buffer 132 while providing the incoming packets with atomic requests to memory access logic 126 without storing the atomic requests in the high-priority buffer 150. As described above, this occurs in some implementations when memory access logic 126 has the bandwidth capacity to accept incoming packets without further queuing.

[0045] Figure 3 is an example, and this example should not unduly limit the scope of the claims. Those skilled in the art will recognize many variations, alternatives, and modifications. While a selected set of processes using the described method illustrates the foregoing, many alternatives, modifications, and variations are possible. For example, some processes may be extended and / or combined. Other processes may be inserted into the above-described processes. Depending on the implementation, the order of processes may be interchanged where other processes are replaced.

[0046] Figure 3 is a block diagram illustrating another example of a device for managing packet transmission to a physical layer interface of a memory architecture, according to one example. In this example, the incoming packet buffer structure 130 is shown as having a plurality of medium-priority buffers 300 and 302, each corresponding to a different defined memory region within local memory 116. A high-priority buffer 150 serves as a higher-priority buffer for storing atomic requests, and a multiplexer 304 (MUX) prioritizes packets from the high-priority buffer 150 over read requests stored in the medium-priority buffers 300 and 302. The controller 122 may control the multiplexing operation of the multiplexer 304 as indicated by dashed arrow 306, or the multiplexer 304 may include suitable logic to output data from high-priority buffer entries before entries from lower-priority buffers. The multiplexer 304 selects from the higher-priority buffers containing entries and provides packets from those entries in a hierarchical priority manner. In some instances, the multiplexer 304 is considered as part of the controller 122.

[0047] In one instance, local memory 116 is divided into memory regions by controller 122, processor 128 under the control of an operating system or driver executing on processor 128, or any other suitable mechanism. In one instance, controller 122 includes configuration register 308 containing data that defines the start address and size of each memory region in local memory 116. In other instances, data representing start and end addresses, or any other suitable data defining one or more memory regions of local memory 116, is stored in configuration register 308. Controller 122 determines each memory region based on data from the configuration register and organizes medium-priority buffers 152 into a plurality of medium-priority buffers 300 and 302, each buffer corresponding to a memory region defined by configuration register 308.

[0048] In this example, physical layer interface 120 includes an uplink cross-physical layer interface 310 and a downlink cross-physical layer interface 312. Outbound packet buffer 314 allows memory access logic 126 to output outbound packets to memory architecture bridge 108, and thus to memory architecture physical layer interface (e.g., switch) 106. Incoming packets, shown as 318, are processed by a controller as described herein. In some implementations, throughput can be inverted, thus the downlink structure produces a buffer for outbound packets similar to that shown for inbound traffic.

[0049] Figure 4 is a flowchart illustrating an example of a method 400 for managing packet transmission according to one embodiment. As previously described, controller 122 determines when to allow packet traffic from memory architecture physical layer interface 106 into architecture attached memory module 114 to enter memory access logic 126, such as the data architecture of a SoC. In this example, a priority level is selected based on the packet type and destination memory region entering the SoC from the memory architecture physical layer interface. As shown in block 402, controller 122 receives incoming packets from memory architecture physical layer interface 106 via memory architecture bridge 108. As shown in block 404, controller 122 determines the packet type and destination memory region of the incoming packet. In one example, this is done by evaluating the data within the packet, where the data identifies the packet type, and evaluating the destination memory region within local memory 116 by the destination address or other suitable information. As shown in block 406, if the packet type is an atomic type instruction packet, controller 122 stores the atomic packet in a higher priority buffer (high priority buffer 150 in this example). The controller allocates medium-priority buffers to include multiple priority buffers, each corresponding to a different defined memory region. The controller stores incoming packets, determined to be of the type containing a read request, into the corresponding region priority buffer based on the memory address associated with the incoming packet.

[0050] However, if the packet is a storage request packet as shown in box 408, such as a write request, the controller saves the packet to a local buffer such as buffer 132. As shown in box 410, if the packet is determined to be a load request representing a load instruction or a read request, the controller 122 gives the load packet a higher priority than a storage instruction but a lower priority than an atomic request, and stores the packet in a medium-priority buffer 152, while the high-priority buffer 150 is a higher-priority buffer and buffer 132 is treated as a low-priority buffer.

[0051] Referring to Figure 3, controller 122 determines which of the medium-priority buffers 300 and 302 to store the incoming packet based on the packet's destination address. In one instance, if the destination address of the incoming packet is within an address region identified by data representing a data region defined by configuration register 308, the controller places the incoming packet in the appropriate medium-priority buffer in response to the incoming load request. In one instance, high-priority buffer 150 and medium-priority buffer 152 are first-in-first-out (FIFO) buffers, but any suitable buffer mechanism may be used. Therefore, controller 122 accesses data defining memory regions, such as data stored in configuration register 308 or other memory locations. In some implementations, an operating system and / or driver executing on one or more processors may specify memory regions via configuration registers revealed by the controller. In other implementations, the controller includes multiple predetermined specified memory regions, such as those provided by firmware or microcode. However, any suitable implementation may be used.

[0052] Referring to Figure 4, as shown in box 412, the method includes allocating packets to high-priority buffer 150, and then allocating other packets to medium-priority buffers 300 and 302, as well as buffer 132. In one instance, this is accomplished by multiplexing multiplexer 304 that outputs packets from buffer entries according to priority. As shown in box 414, in the case that the packet is a storage request, the method includes determining whether there is sufficient bandwidth across the physical layer interface to deliver the write access. In other words, controller 122 determines whether physical layer interface 120 (which has a smaller data rate and smaller link width than memory architecture bridge 108) can deliver write packets. For example, when MUX 304 saturates the receiver of physical layer interface 120, physical layer interface 120 begins to reject packets. When this occurs, buffering of read requests begins as described above. If there is sufficient bandwidth to deliver the write access, then as shown in box 416, the method includes delivering the write access to local memory 116. If there is not enough bandwidth to pass write accesses, as shown in box 418, the method includes waiting until sufficient bandwidth is available, and then passing write operations to memory in groups according to write requests.

[0053] Regarding the medium-priority buffers 300 and 302, as shown in box 420, if the incoming packet is a load request (e.g., a read request) and no other high-priority incoming packets are present in the high-priority buffer 150, the method includes performing a load operation from memory by outputting the packet from the appropriate medium-priority buffers 300, 302 to memory access logic 126. The method continues for the incoming packet.

[0054] As shown in box 422, when a capped QoS mechanism is employed, if the packet is a read-type packet, the method includes determining whether the memory region in local memory 116 corresponding to the packet's address is capped during the time interval. For example, if excessive access to that memory region in local memory 116 occurs during the time interval, controller 122 stores the incoming packet in a medium-priority buffer 152 until the next time interval occurs.

[0055] Figure 5 is a block diagram illustrating another example of an apparatus for managing packet transmission to a memory architecture physical layer interface. In this example, controller 122 includes a packet analyzer 500, a priority buffer packet router 502, and a per-memory region counter 504. The packet analyzer 500 analyzes incoming packets to determine packet type by, for example, scanning and interpreting the contents of the packet, or evaluating packet type data within the packet to determine whether the packet is about an atomic request (e.g., atomic instruction execution), a load request (e.g., a read instruction), or a storage request (e.g., a write instruction), and outputs packet 506 to the priority buffer packet router 502, which places the packet in an appropriate priority buffer. Packet 506 may be directly routed to an appropriate buffer or may include a routing identifier to indicate which priority buffer the packet should be placed in. As described, configuration register 308 defines at least some memory regions of the apparatus's local memory 116 as priority memory regions, where each memory region allows an unlimited maximum number of memory architecture physical layer interface accesses per time interval. For example, controller 122 allows some memory regions in local memory 116 to be registered as priority memory regions via configuration registers (e.g., model-specific registers (MSR)).

[0056] In the registered memory regions, and in all other memory regions mapped to the memory architecture, controller 122 employs a certain type of quality of service (QoS) capping technique. For example, each registered memory region has an unlimited maximum number of architecture accesses that the memory region can receive per time interval, and the sum of the upper limits mapped to the memory architecture space for the entire local memory 116 does not exceed the available cross-PHY capacity (e.g., limited by PCIe channel width and data rate). Because the allocated upper limits do not exceed the available channel capacity (excess packets will wait until the next time interval), the level of isolation from other memory regions in the architecture is high. Larger memory regions receive larger upper limits.

[0057] For each memory access in each region, the per-memory-region counter 504 is incremented, and the controller 122 compares the current count value with a threshold corresponding to the maximum number of architecture accesses for each region without limitation. The threshold may be stored as part of configuration register information or other mechanisms. Therefore, the controller 122 maintains a count of the number of memory accesses made to the defined memory regions via the memory architecture physical layer interface within the defined time interval. When the maximum allowed number of accesses is exceeded during the defined time interval, the controller 122 stores a read packet (load packet) in a medium-priority buffer 152. Whenever a higher-priority packet (e.g., an atom) is to be issued to the MUX, the controller stores the packet in the region buffer. An upper limit determines whether a packet is issued to the MUX from a given region buffer. If the upper limit is reached, issuance to the MUX is temporarily suspended until the start of the next time interval (which is a configurable controller parameter). In other words, the registered memory regions defined by the configuration registers are memory regions in local memory 116. Each memory region thus registered has a buffer. In this example, medium-priority buffers 300 and 302 are associated with defined memory regions, and each stores accesses from the architecture to that memory region in local memory 116. An upper limit determines whether an access is issued to the MUX from a given buffer; thus, a memory access can eventually enter its memory region in local memory 116. If the upper limit is reached, issuance to the MUX is temporarily suspended until the next time interval begins.

[0058] The stored packets are provided from the medium-priority buffer 152 to the memory access logic 126 in the next time interval. In this example, the controller 122 stores incoming packets that are determined to be of the type containing a read request in each of the multiple medium-priority buffers 300 and 302 based on the address associated with the incoming packet, and also tracks the number of memory accesses in the memory region to cap the number of accesses occurring in a particular memory region.

[0059] Figure 6 is a flowchart illustrating an example of a method 600 for managing packets according to one instance. As shown in block 602, the method includes defining at least some memory regions of local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture physical layer interface accesses per time interval. As shown in block 604, the method includes maintaining a count of the number of memory accesses made to the defined memory regions of local memory via the memory architecture physical layer interface within the time interval. As shown in block 606, the method includes storing packets in a priority buffer when the maximum allowed number of accesses is exceeded, and as shown in block 608, the method includes providing the stored packets from the buffer to memory access logic in the next time interval.

[0060] Among other technical advantages, the adoption of a controller and buffer hierarchy reduces latency and improves energy efficiency when accessing local memory, which is addressable as part of the memory architecture. Additionally, in cases where the architecture-attached memory module includes one or more processors, the controller offloads packet management from one or more processors, thereby increasing the speed of processor operation. Other advantages will be recognized by those skilled in the art.

[0061] Although features and elements have been described above in specific combinations, each feature or element may be used alone without other features and elements, or in various combinations with or without other features and elements. The devices described herein in some implementations are manufactured using computer programs, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM discs and digital versatile optical discs (DVDs)).

[0062] In the preceding detailed description of various embodiments, reference has been made to the accompanying drawings, which form a part thereof, and specific preferred embodiments in which the invention may be practiced are illustrated by way of example. These embodiments have been described in sufficient detail to enable those skilled in the art to practice the invention, and it should be understood that other embodiments may be utilized and logical, mechanical, or electrical changes may be made without departing from the scope of the invention. Certain information known to those skilled in the art may be omitted in this specification to avoid details unnecessary for enabling those skilled in the art to practice the invention. Furthermore, many other different embodiments incorporating the teachings of this disclosure can be readily conceived by those skilled in the art. Therefore, the invention is not intended to be limited to the specific forms set forth herein, but rather, it is intended to cover such alternatives, modifications, and equivalents that can be reasonably included within the scope of the invention. Therefore, the preceding detailed description should not be construed as limiting, and the scope of the invention is defined only by the appended claims. The detailed description of the embodiments and examples described above is presented for illustrative and descriptive purposes only, and not for limiting purposes. For example, the described operations are performed in any suitable order or manner. Therefore, it is contemplated that the invention covers any and all modifications, variations, or equivalents falling within the scope of the basic principles disclosed above and claimed herein.

Claims

1. A method for managing packet transmissions performed by a device, the method comprising: The device receives incoming packets for local memory within its physical layer via a memory architecture physical layer interface. The local memory is configured as architecture-attached memory and is addressable by other devices. At least some of the packets include different instruction types. The memory architecture physical layer interface has a higher data rate than the physical layer interface of the device. The controller of the device determines the packet type of the incoming packet received from the memory architecture physical layer interface via the physical layer interface; as well as When the type of an incoming packet from the physical layer interface of the memory architecture is determined to be a type containing an atomic request, the incoming packet containing the atomic request is queued in a first priority buffer, and the incoming packet containing other types of requests is queued in one or more priority buffers with lower priority than the first priority buffer. The controller then causes the incoming packet with the atomic request to be transmitted to the memory access logic, which accesses the local memory within the device, with priority over other packet types of incoming packets.

2. The method of claim 1, wherein transmitting the incoming packet having the atomic request over other packet types of incoming packets comprises: Incoming packets identified as containing the atomic request will be queued in the first priority buffer; Queue other group types in the second priority buffer; as well as The output of the packet from the first priority buffer takes precedence over the output of the packet from the second priority buffer.

3. The method of claim 2, further comprising: Access data, wherein the data defines at least a plurality of memory regions of the device’s local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture physical layer interfaces to access each memory region per time interval; Maintain a count of the number of memory accesses to the defined memory region performed through the memory architecture physical layer interface within the time interval; When the maximum number of memory architecture physical layer interface accesses is exceeded, the read packets are stored in the second priority buffer; as well as In the next time interval, the stored packets are provided from the second priority buffer to the memory access logic.

4. The method of claim 1, wherein prioritizing the transmission of the incoming packet with the atomic request over other packet types of incoming packets comprises queuing incoming packets determined to contain a storage request in a buffer while providing the incoming packet with the atomic request to the memory access logic.

5. The method of claim 3, further comprising: The second priority buffer is allocated to include a plurality of second priority buffers, each of which corresponds to a different defined memory region; as well as Incoming packets, which are determined to be of the type containing a read request based on the address associated with the incoming packet, are stored in a corresponding second priority buffer in a memory region corresponding to the different definitions.

6. An apparatus, said apparatus comprising: One or more processors; Memory access logic, which is operatively coupled to the one or more processors; The local memory within the device is operatively coupled to the memory access logic and configurable as an addressable portion of the memory, which can be addressed as a framework-attached memory and can be addressed by other devices through a memory framework physical layer interface. A physical layer interface operatively coupled to the memory access logic and operable to receive incoming packets from the memory architecture physical layer interface for the local memory configured as architecture-attached memory, the memory architecture physical layer interface having a higher data rate than the physical layer interface, at least some of the packets comprising different instruction types. A controller, operatively coupled to the physical layer interface and configured to: Determine the packet type of the incoming packet received from the memory architecture physical layer interface through the physical layer interface; as well as When the determined incoming packet type is a type containing an atomic request, the incoming packet containing the atomic request is queued in the first priority buffer, and the incoming packet containing other types of requests is queued in one or more priority buffers with lower priority than the first priority buffer, and the incoming packet with the atomic request is transmitted to the memory access logic with priority over other packet types of incoming packets.

7. The device of claim 6, wherein the device comprises: The first priority buffer and the second priority buffer having a lower priority than the first priority buffer; and The controller is further configured to: The incoming packet containing the atomic request is prioritized over other packet types by being queued in the first priority buffer. Queue other group types in the second priority buffer; as well as The output of the packet from the first priority buffer takes precedence over the output of the packet from the second priority buffer.

8. The device of claim 6, further comprising: A buffer, which is operatively coupled to the controller; and The controller is further configured to prioritize the incoming packet with the atomic request over other packet types of incoming packets by queuing incoming packets identified as containing a storage request in the buffer while providing the incoming packet with the atomic request to the memory access logic.

9. The device of claim 7, further comprising: A configuration register is configured to include data that defines at least a plurality of memory regions of the device’s local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture physical layer interface accesses per time interval; and The controller is further configured to: Maintain a count of the number of memory accesses to the defined memory region performed through the memory architecture physical layer interface within the time interval; When the maximum number of memory architecture physical layer interface accesses is exceeded, the read packets are stored in the second priority buffer; as well as In the next time interval, the stored packets are provided from the second priority buffer to the memory access logic.

10. The device as claimed in claim 9, wherein: The second priority buffer includes a plurality of second priority buffers, each of which corresponds to a different defined memory region; and The controller is also configured to store incoming packets, which are determined to contain a read request type, in a corresponding second priority buffer in a memory region corresponding to the different definitions, based on the address associated with the incoming packet.

11. The device of claim 6, further comprising a memory architecture bridge circuit operatively coupled to the controller and the memory architecture physical layer interface, and operable to transfer packets between the physical layer interface and the memory architecture physical layer interface.

12. An apparatus, the apparatus comprising: Local memory, which is operatively coupled to memory access logic and configurable as an addressable portion of memory addressable via a memory architecture physical layer interface; A physical layer interface that operates to receive incoming packets from a memory architecture physical layer interface, the memory architecture physical layer interface having a higher data rate than the physical layer interface, and at least some of the packets comprising different instruction types; An incoming packet buffer structure, the incoming packet buffer structure including a hierarchically ordered priority buffer structure, the priority buffer structure including at least a first priority buffer and a second priority buffer having a priority lower than the first priority buffer; A controller, which is operatively coupled to the physical layer interface and the incoming packet buffer structure; The controller is configured as follows: Determine the packet type of the incoming packet from the physical layer interface of the memory architecture; as well as When the determined packet type indicates that an atomic request exists in the incoming packet, the incoming packet is stored in the first priority buffer. When the determined packet type indicates that a load instruction exists in the incoming packet, the incoming packet is stored in the second priority buffer; as well as The incoming packets of the storage are provided to the memory access logic in hierarchical order according to the priority buffer order.

13. The apparatus of claim 12, further comprising: A buffer, which is operatively coupled to the controller; and The controller is further configured to prioritize the incoming packet with the atomic request over other packet types of incoming packets by queuing incoming packets identified as containing a storage request in a storage buffer while providing the incoming packet with the atomic request to the memory access logic.

14. The apparatus of claim 13, further comprising: A configuration register is configured to include data that defines at least a plurality of memory regions of the device’s local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture physical layer interface accesses per time interval. and The controller is further configured to: Maintain a count of the number of memory accesses to the defined memory region performed through the memory architecture physical layer interface within the time interval; When the maximum number of memory architecture physical layer interface accesses is exceeded, the packet is stored in the second priority buffer; as well as In the next time interval, the stored packets in the second priority buffer are provided from the buffer to the memory access logic.

15. The apparatus of claim 12, wherein: The second priority buffer includes a plurality of second priority buffers, each of which corresponds to a different defined memory region; and The controller is also configured to store incoming packets, which are determined to contain a read request type, in a corresponding second priority buffer in a memory region corresponding to a different definition, based on the address associated with the incoming packet.

16. A system comprising: A memory architecture that operates to interconnect multiple distributed non-volatile memories; A first device, operatively coupled to the memory architecture; as well as A second device, operatively coupled to the memory architecture, wherein the first device and the second device have a physical layer interface to receive memory access requests from each other via the memory architecture, the second device comprising: Local memory, operatively coupled to the memory access logic of the second device, the local memory being configurable as an addressable portion of the distributed nonvolatile memory addressable via the memory architecture; The physical layer interface, operatively coupled to the memory access logic, is configured to receive incoming packets from the first device via the memory architecture addressing the local memory of the second device, the memory architecture having a higher data rate than the physical layer interface, and at least some of the packets comprising different instruction types. A controller, operatively coupled to the physical layer interface and configured to: Determine the packet type of the incoming packet received from the memory architecture; and When the determined incoming packet type is a type containing an atomic request, the incoming packet is queued in a first priority buffer. Incoming packets determined to contain other types of requests are queued in one or more priority buffers with lower priority than the first priority buffer, so that the incoming packet with the atomic request is transmitted to the memory access logic with priority over other packet types of incoming packets.

17. The system of claim 16, wherein: The second device includes a first priority buffer and a second priority buffer having a lower priority than the first priority buffer; and The controller is also configured to: The incoming packet containing the atomic request is prioritized over other packet types by being queued in the first priority buffer. Queue other group types in the second priority buffer; as well as The output of the packet from the first priority buffer takes precedence over the output of the packet from the second priority buffer.

18. The system of claim 16, wherein: The second device includes a buffer operatively coupled to the controller; and the controller is further configured to prioritize the incoming packet with the atomic request over other packet types of incoming packets by queuing incoming packets determined to contain a storage request in the buffer while providing the incoming packet with the atomic request to the memory access logic.

19. The system of claim 17, wherein the second means comprises: A configuration register is configured to include data that defines at least a plurality of memory regions of the device's local memory as priority memory regions, wherein each memory region is allowed an unlimited maximum number of memory architecture accesses per time interval; and The controller is further configured to: Maintain a count of the number of memory accesses to the defined memory region performed by the memory architecture within the time interval; When the maximum number of memory architecture accesses is exceeded, the read packets are stored in the second priority buffer; as well as In the next time interval, the stored packets are provided from the second priority buffer to the memory access logic.

20. The system of claim 19, wherein: The second priority buffer includes a plurality of second priority buffers, each of which corresponds to a different defined memory region; and The controller is also configured to store incoming packets, which are determined to contain a read request type, in a corresponding second priority buffer in a memory region corresponding to a different definition, based on the address associated with the incoming packet.

Citation Information

Patent Citations

  • glow wire candle

    DE620717C

  • Technologies for scalable remotely accessible memory segments

    US20160314073A1

  • Data storage system having acceleration path for congested packet switching network

    US7979588B1