Network-aware memory proxy
By reducing software management through Network Aware Memory Agent (NAMA), the problem of insufficient access performance of the MPI protocol in high-performance computing is solved, achieving seamless integration of hardware and MPI and improving memory access efficiency.
Patent Information
- Application Number
- CN202510572916.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-06
- Publication Date
- 2025-11-14
AI Technical Summary
Existing MPI protocols cannot meet the performance requirements of distributed and shared memory access in modern applications such as machine learning and artificial intelligence in high-performance computing. Software management increases computational load and latency, and the integration of hardware accelerators with MPI is not seamless enough, resulting in excessive network communication overhead.
By employing a Network Aware Memory Agent (NAMA), MPI message handshake and packetization are performed through hardware logic, reducing software management and enabling direct network transmission of memory requests, thus supporting the functionality of distributed and shared memory.
It significantly reduces computational load and memory latency, improves memory access performance, achieves seamless integration of hardware and MPI, and reduces network communication overhead.
Smart Images

Figure CN120950426A_ABST
Abstract
Description
Technical Field
[0001] This document generally relates to memory controllers, and more specifically, to network-aware memory agents for manageable memory and memory controllers. Background Technology
[0002] The Message Passing Interface (MPI) is a messaging standard that has been used for many years in high-performance computing (HPC) distributed memory systems. MPI provides hardware and software manufacturers with a set of application programming interfaces (APIs) supporting high-performance communication for distributed memory applications. Developers of compute-intensive programs frequently use these APIs to distribute computational workloads across clusters using distributed memory systems. MPI messages from the originating process are handled by the kernel-mode MPI message library, which (in the case of "placement") enriches the MPI message with data from local memory and then typically sends the enriched message to other processes via Transmission Control Protocol (TCP) sockets. The receiver-side API retrieves the data embedded in the message and writes it to its local memory. The receiver API may (in the case of "fetch") send data from its local memory back to the originating process (again, typically using TCP sockets), where the originating API copies the data to its local memory.
[0003] Distributed memory describes an architecture where multiple nodes (computers, computing cores, etc.) can access memory over a network. In contrast, shared memory describes an architecture where multiple nodes can access the same memory, for example, via a bus or interconnect. Distributed memory is far more scalable than shared memory, but it introduces networking complexity and overhead. Hybrid memory describes an architecture that incorporates both distributed and shared memory. Traditionally, MPI was defined on top of distributed memory, but starting with MPI-3, the shared memory (SHMEM) extension was introduced, where programmers create shared regions for various processes running on independent CPUs sharing a common memory space. More recently, MPI-4.0 and OpenSHMEM defined the framework for hybrid memory. Due to the widespread availability of well-tested libraries, the primary use of MPI remains on distributed memory.
[0004] However, in its current implementation, MPI alone is insufficient to support the demands of modern HPC. For example, applications such as machine learning (ML) and artificial intelligence (AI) require significantly higher distributed and shared memory access performance than MPI currently offers. The MPI message base (containing TCP sockets and related software stack) imposes a substantial computational load and increases latency.
[0005] Hardware accelerators exist that implement most networking layers (transport layer and below) in hardware. Protocols such as Remote Direct Memory Access (RDMA), Remote Direct Memory Transfer over Converged Ethernet (RoCE), and MPI tag match offload are also being developed. These protocols improve the performance of lower-level data transfers but cannot be seamlessly integrated with MPI and therefore require significant software management. For example, RDMA implements the movement of data from the application area to packetization hardware (and vice versa), but this typically requires software management, which adds overhead. For instance, in an MPI_Put operation (sending data to distributed memory), the software must accept MPI messages, collect the relevant data, move the data to the RDMA data area, update the pointers in the relevant send queue, and then monitor for acknowledgments. Similarly, on the receiving side, the software must read the receive queue, move the data to the actual MPI memory, and send an acknowledgment back to the hardware. The MPI_Get operation (requesting data from distributed memory) requires even more software management.
[0006] Software involvement in networking distributed memory introduces additional overhead, especially considering that much of the software management must be performed by the host operating system in kernel mode, requiring significant mode switching. This increased overhead limits the computational load and latency improvements that these protocols can provide.
[0007] Therefore, there is a need for improved hardware solutions that facilitate networking of distributed memory with reduced software management. Summary of the Invention
[0008] On one hand, this disclosure relates to a network-aware memory agent, comprising: a memory interface configured to communicate with a first system memory including random access memory (RAM); at least one communication interface configured to communicate with: a first central processing unit (CPU); and a packet networking device; logic for identifying a first memory window including the first system memory; logic for communicating with a local CPU via an interconnect; logic for receiving a first memory request; logic for determining that the first memory request addresses a second memory window that does not include the first system memory; logic for identifying a target memory agent associated with the second memory window based on a memory window identifier in the first memory request; logic for generating a first enhanced memory request (EMR) proxy-to-proxy request (a2aReq) message corresponding to the first memory request; logic for identifying packet header information associated with the target memory agent; and logic for packetizing the first EMR2aReq message using the packet header information to generate one or more data packets; and logic for transmitting the one or more data packets via the packet networking device for reception by the target memory agent.
[0009] On the other hand, this disclosure relates to a network-aware memory controller, comprising: a memory controller including: a memory interface configured to communicate with a first system memory including random access memory (RAM); an interconnect interface configured to communicate with an interconnect providing communication with a first central processing unit (CPU); a packet network interface configured to communicate with a packet network; and logic for transmitting and receiving memory requests via the packet network.
[0010] On the other hand, this disclosure relates to a method comprising: identifying a first memory window including a first system memory, the first system memory including random access memory (RAM), via a network-aware memory agent; the network-aware memory agent comprising: a memory interface configured to communicate with the first system memory including the RAM; at least one communication interface configured to communicate with: a first central processing unit (CPU); and packet networking equipment; communicating with a local CPU via the interconnect via the network-aware memory agent; receiving a first memory request via the network-aware memory agent; and determining via the network-aware memory agent that the first memory request address does not include the first system memory. The second memory window; identifying a target memory agent associated with the second memory window through the network-aware memory agent and based on the memory window identifier in the first memory request; generating a first enhanced memory request (EMR) a2aReq message embedded in the first memory request through the network-aware memory agent; identifying packet header information associated with the target memory agent through the network-aware memory agent; packetizing the a2aReq message using the packet header information through the network-aware memory agent to generate one or more data packets; and transmitting the one or more data packets through the network-aware memory agent for reception by the target memory agent via the packet networking equipment. Attached Figure Description
[0011] Figures 1A to 1D This is a block diagram illustrating a network-aware memory agent according to some embodiments.
[0012] Figure 2 This is a block diagram illustrating a system employing multiple network-aware memory agents in a shared memory architecture according to some embodiments.
[0013] Figure 3 This is a block diagram illustrating a system employing multiple network-aware memory agents in a distributed memory architecture according to some embodiments.
[0014] Figure 4 This is a block diagram illustrating a system employing multiple network-aware memory agents in a hybrid memory architecture according to some embodiments.
[0015] Figure 5A This describes an interconnect with a ring topology according to some embodiments.
[0016] Figure 5B This describes an interconnection with a mesh topology according to some embodiments.
[0017] Figures 6 to 8 This is a message flow diagram illustrating a communication model for remote memory operations according to some embodiments.
[0018] Figure 9 This is a flowchart illustrating a method for operating a memory agent according to some embodiments.
[0019] Figure 10 and 11 This is a flowchart illustrating a method for processing memory requests addressed to a remote memory window according to some embodiments.
[0020] Figure 12 This is a flowchart illustrating a method for processing memory requests addressed to a local memory window according to some embodiments.
[0021] Figure 13 This describes a method for managing a queue of memory requests according to some embodiments.
[0022] Figure 14 This describes a method for processing network computing requests according to some embodiments.
[0023] Figure 15 This describes a method for managing fences on a local storage window according to some embodiments.
[0024] Figure 16 This is a block diagram illustrating example components of a computer system according to some embodiments. Detailed Implementation
[0025] Some embodiments provide “Network Aware Memory” (NAM), which addresses many shortcomings of current solutions. Some embodiments employ hardware to perform most MPI interactions, significantly reducing computational load and memory latency. In some responses, various embodiments enable NAM by using an agent (referred to herein as “namAgent” or “NAMA”), which can be implemented in hardware and / or firmware logic. In some cases, NAMA may be implemented with and / or integrated with a memory controller; such implementations are referred to herein as “namController” or “NAMC”. It should be understood that NAMC is a specific implementation of NAMA. In one aspect of some embodiments, NAMA is aware of its location in the distributed memory system and has logic that allows communication with network equipment (e.g., Ethernet components, such as network interface cards (NICs), Ethernet switches, etc.). In some cases, NAMA may include such networking components.
[0026] According to some embodiments, NAMA can perform MPI message handshake, enrich MPI messages with local data (if applicable), handle the packetization of MPI messages, and / or exchange packetized messages with existing networked devices. On the receiving side, some embodiments enable NAMA to receive (e.g., via network sockets), depacketize and / or decode MPI messages, perform some or all memory operations, and / or respond to the originating side (if needed).
[0027] From a global memory perspective, some embodiments can provide SHMEM functionality similar to a distributed memory system (and not just shared memory). In other words, such embodiments can localize remote memory to the compute processor. In some aspects, the memory localized by some embodiments does not need to be cache-consistent, as such embodiments can employ an MPI synchronous API. Various embodiments can provide additional functionality. For example, when a message arrives at a network device (e.g., a switch port or NIC) communicating with a set of memories on a common bus, NAMA according to some embodiments can identify the appropriate specific memory controller to receive the message, for example, using MPI tag matching. Similarly, if the existing memory controller has a fast compute link... TM With a (CXL) interface, NAMA, according to some embodiments, can translate message packets into CXL messages to allow the controller to select request messages and handle memory reads / writes.
[0028] In this document, certain terms are used as follows:
[0029]
[0030]
[0031]
[0032]
[0033]
[0034] Table 1: Vocabulary
[0035] For illustrative purposes, the following description uses several exemplary message types and primitive functions. For ease of reference, these exemplary message types are listed and described in Table 2, and the description of each exemplary message type includes some exemplary fields that can be found in messages of this type. Table 3 lists and describes two exemplary fields in more detail. Exemplary primitive functions are listed and described in Table 4, and several exemplary variables for these exemplary primitives are described in Table 5. It should be noted that the message types and primitive functions (and any fields or variables described therein) are not intended to be limiting, and those skilled in the art will understand that various embodiments may use various message types, fields, and primitive functions, including (but not limited to) those described herein.
[0036]
[0037]
[0038]
[0039]
[0040] Table 2: Message Types
[0041]
[0042]
[0043] Table 3: Exemplary Fields
[0044]
[0045]
[0046]
[0047] Table 4: Exemplary Primitive Functions
[0048]
[0049]
[0050]
[0051] Table 5: Exemplary Variables
[0052] Examples of NAMA functionality in various embodiments
[0053] NAMA, according to various embodiments, can exhibit several novel functionalities and properties through its network-aware nature. In some embodiments, NAMA maintains a list of memory regions (or memory windows) and their locations in the network (including locations in physically connected memory). In some embodiments, memory regions are defined by a pair (e.g., {win_handle, win_addr}), and NAMA can create them using functions (e.g., MPI_Win_create) that have this mapping.
[0054] 1) Conventional memory request
[0055] In some embodiments, NAMA listens for regular memory read / write requests and Enhanced Memory Requests (EMRs) from local processes or from peer memory agents (e.g., NAMA). In some embodiments, NAMA is positioned in the path of all messages to the memory controller (and / or, as described further in detail below, may include the memory controller and / or be integrated with the memory controller). In such embodiments, NAMA may forward regular memory read / write requests to the memory controller and forward read data / write completions back to the originating process / processor (e.g., via the interconnect).
[0056] In some embodiments, NAMA may then ignore (or, if inline, forward unaltered) regular memory read / write requests to local memory and allow the requests to be handled by the memory controller (and / or by an integrated memory controller, such as in the case of NAMC), for example, by returning an unaltered response (containing any read data) to the memory controller via the interconnect.
[0057] 2) Enhanced memory requests
[0058] In some embodiments, an Enhanced Memory Request (EMR) may be one of the following: (a) a message originating from a local process (e.g., p2aReq) or (b) a message originating from a remote processor (e.g., via a remote NAMA on a network) (e.g., a2aReq, a2aCRM). See the preceding definitions section. In some embodiments, an EMR may contain data read from memory, in which case the message may be considered a partially satisfied request (PFR). A PFR request message may be transmitted from NAMA to NAMA as an a2aReq message (or a portion thereof). In one aspect, in addition to a PFR message, an a2a (agent-to-agent) message may also contain parameters including the agentID of the source NAMA and the destination NAMA. In some cases, as further specified in detail below, an a2a message may also contain addressing and / or routing information (e.g., an IP header) to enable routing over packet networks.
[0059] a) Enhanced memory requests from the local CPU
[0060] If the message originates from the local CPU's EMR (which implies the source memory window must be located in locally attached system memory), and the target memory region to which the request is directed is also local (i.e., a memory window managed by NAMA), then NAMA will translate the enhanced request into a set of regular memory read / write requests and forward them to the local memory controller (and / or, if NAMC, the integrated memory controller to execute the request). In some embodiments, NAMC will also increment the ucMPICount variable for the specified memory window. This may be done for fault tolerance, such that subsequent MPI_Win_fence operations are preserved if the operation does not complete for some reason. If the request is a "place" request, then NAMA may read the specified amount of data from the "source" location specified in the EMR and write it to the "target" location specified in the EMR. Conversely, if the request is a "get" request, then NAMA may read the specified amount of data from the "target" location and write it to the "source" location. In either case, if a Completion Response Message (CRM) is enabled, the agent may send an a2pCRM message back to the calling process. This message can be an interrupt, a write to a specific memory region configured by the process for this purpose, etc. Typically, this specific memory region is organized as a ring, allowing the process to freely read such a2pCRM messages at idle times. After performing the requested memory operation, NAMA can decrement the ucMPICount of the specified memory window (e.g., if the counter was incremented at the start of message processing). If ucMPICount becomes 0 and the ucMPIFenceFlag of the target memory window is set, then NAMA can generate and / or send the CRM of the fence and / or clear the corresponding ucMPIFenceFlag.
[0061] On the other hand, if the message originates from an EMR on the local CPU and the target memory region is controlled by another memory controller (e.g., on the same network node), then the NAMA can act as the source NAMA in a shared memory bridge (SMB). For example, in some embodiments, the NAMA can update the request (and / or generate a new request) in such a way that the request is routed to a target NAMA via the local interconnect, which manages the target memory window and sends the message back to the local interconnect as an a2a message. The source NAMA can also increment the ucMPICount of the specified memory window. If the request is a "place" request, then the NAMA can read a specified amount of data from the "source" location, enrich the EMR with the read data, and change the request to a PFR. The source NAMA can generate an a2aReq containing the PFR and send the request to the target NAMA. If the request is a "get" request, then the source NAMA can swap the source and target fields, update the msgSubtype, and / or send the EMR request to the target NAMA. In some embodiments, the target NAMA can treat such modified messages as "place".
[0062] If the request originates from the local CPU and is directed to a non-local memory window and is not on the same network node, then the NAMA can act as the source NAMA and generate a request to be satisfied by the target NAMA over the network (e.g., a packet network). The source NAMA can extract local data (e.g., if the request is a "place" request) and form a PFR. The source NAMA can embed the PFR in an a2aReq message. In some embodiments, the source NAMA packets the message into packets (or multiple packets, if necessary) containing the a2aReq message and a header that allows the packet to be routed by a locally connected network equipment. The source NAMA may then send the packets directly or via an interconnect to the network equipment. In some cases, the source NAMA also increments the ucMPICount of the target memory window. Based on a typical MPI message structure, the message can typically be packed into a single Ethernet packet; if not, the source NAMA can generate multiple packets that will be reassembled by the target NAMA. The generation and reassembly of multiple packets can be performed by the source and target NAMA respectively using any of a variety of such techniques, or may be performed by an intermediate network equipment. In addition to the functions necessary for delivery by the packet network, source and destination NAMA can operate similarly to the SMB behavior described above.
[0063] b) Enhanced memory requests from peer agents
[0064] However, if the request originates from a peer NAMA on the same network node, then the peer NAMA can act as the source NAMA, and the local NAMA can act as the target NAMA for the SMB. From the perspective of the local (target) NAMA, if the request is an a2aReq message with a "place" request (e.g., PFR), then the target NAMA writes the data to local storage and sends the a2aCRM message back to the source NAMA.
[0065] If the request is an a2aReq with a "get" request but converted to a PFR by the source NAMA (as described above), then the target NAMA may write the data to local storage and / or decrement the "win"-specific ucMPICount (if any). If the count becomes 0 and ucFenceFlag is set for said "win", then the target NAMA (if enabled, e.g., via the enableA2pFenceCRM flag) may generate an a2pCRM message for past fence messages (which initially caused the ucFenceFlag setting) and / or clear the ucFenceFlag. If the request is a "get" request with exchange fields, then the target agent may read the data from local storage, convert the request to a PFR, and / or send the message (along with the read data) as an a2aReq message to the source NAMA.
[0066] In any case, whether it's a "place" or "get" request, the target NAMA can increment the target memory window's ucMPICount before the operation begins and decrement it when the operation completes. The use of counters can be beneficial for fault tolerance, even if the target agent can perform its work in local memory in a non-blocking manner. In some cases, for example, for clarity, the target NAMA can implement a different counter (e.g., ucMPICount_t) to distinguish the count of a2a memory requests from the count of similar requests from local processes.
[0067] If the request originates from a peer NAMA on a different network node, the local NAMA acts as the target NAMA, performing any necessary functions to depacketize or otherwise recover the a2aReq message, and otherwise operating in a manner similar to the SMB behavior described above. However, in this case, the target NAMA is aware of the source NAMA's network node and can function to packetize the response message and transmit it to network equipment (e.g., directly or via interconnect) for transmission over a packet network. This may include sending an a2aCRM message (if applicable).
[0068] If the message received from a peer NAMA is an a2aCRM message received from a peer NAMA of a remote network node (e.g., via a packet network) or a local node (e.g., via an interconnect), then the local NAMA may decrement ucMPICount and create an a2pCRM message (if enabled, e.g., with the enableA2pCRM flag) to the originating process. If ucMPICount becomes 0 and the flag ucFenceFlag of the window is set, then the NAMA may generate an a2pCRM message (if enabled) for the fence message and / or clear the ucFenceFlag of the target window.
[0069] c) Fence message
[0070] In some embodiments, if the EMR is a fence message (e.g., MPI_Win_fence), then NAMA can infer that the fence message originated from the local process. In this case, NAMA can set a flag (e.g., ucMPIFenceFlag) for the memory window to which the fence message was directed. If the counter of the memory window (e.g., ucMPICount) is zero, then NAMA can send a CRM message (if enabled) to the local process and / or clear the flag. If the counter is non-zero, then NAMA can allow the flag to remain set until the counter reaches zero due to the execution of other requests; and at that time, NAMA can send a CRM message (if enabled).
[0071] Exemplary architecture and topology
[0072] Figure 1A An exemplary NAMA 100 is described according to a set of embodiments. In the illustrated embodiments, the NAMA 100 is integrated with the memory controller 105 within a NAMC 110. The NAMC 110 includes the memory controller 105 and the NAMA 100 integrated within a single package (e.g., a chip, a system-on-a-chip (SoC), a printed circuit board (PCB), etc.). The NAMA 100 includes a NAMA context memory 115, which can be used to buffer MPI messages, data, etc., during memory operations (including, but not limited to, distributed memory operations). Figure 1AIn the illustrated embodiments, NAMC 110 further includes an interconnect interface 120, a memory interface 125, and a network interface 130. These interfaces advantageously allow NAMC 110 to communicate with interconnect 135, system memory 140, and network device 150. In various embodiments, the nature of these interfaces may depend on the type of connection established; for example, in some cases, memory interface 125 may be configured to interface with memory controller 105 in a manner similar to how a CPU interfaces with memory controller 105. In other embodiments, as described below, for example, the memory interface may be incorporated into memory controller 105 and / or provide some or all of the functionality of the memory controller; in this case, memory interface 125 may interface directly with memory 140, for example, using commands similar to those used by the memory controller. In some embodiments, one or more of interfaces 120 to 130 may resemble typical interfaces on a chip package or PCB for interfacing with components (e.g., interconnect 135, system memory 140, and / or network device 150).
[0073] In some embodiments, these interfaces 120 to 130 enable NAMA 100 (including (but not limited to) NAMC 110) to communicate via interconnects (e.g., with MPI messages), memory (e.g., read / write input / output operations (IO) performed on system memory 140 via memory controller 105, which in certain embodiments are controlled by instructions provided to controller 105 by NAMA 100 through memory interface 125), and / or other NAMAs (e.g., packetized EMR messages transmitted via network interface 130 through network component 150). In some embodiments, network interface 130 may also be incorporated into network device 150 (e.g., NIC or switch). In some embodiments, network interface 130 and / or network device 150 may provide communication with packet networks (e.g., Internet Protocol (IP) networks). As used herein, the term “network” includes (but is not limited to) such packet networks unless the context explicitly indicates otherwise.
[0074] Although Figure 1A This describes one possible arrangement of the NAMA 100, but other embodiments may feature many different arrangements. Figures 1B to 1D Here are a few examples illustrating this type of arrangement. For instance, in... Figure 1B In this configuration, the NAMA 100 is not incorporated or integrated with the memory controller; instead, it communicates with an external memory controller 105 (via memory interface 125). Conversely, Figure 1C The NAMC110 integrates the functionality of the NAMA 100 into the memory controller 105 itself, while Figure 1DRather than interfacing directly with network component 150 through interconnect 135, a combined interconnect / network interface 155 is employed. From these examples, those skilled in the art will appreciate that many different architectural arrangements are possible within the various embodiments, and all such arrangements can possess some or all of the NAMA 100 functionality described herein. Accordingly, no particular exemplary architecture should be considered limiting. For ease of description, many of the following examples describe the NAM-C 110 that includes the integrated NAMA 100 and memory controller 105; however, it should be understood that similar principles and techniques can be applied to NAMA 100s that do not have an integrated memory controller 105, and the NAMA 100 and NAM-C 110 can be used interchangeably in such examples, in which case an external memory controller 105 can be used as necessary.
[0075] Figure 2 Illustrate a system 200 that includes multiple NAMA 100a, 100b communicating with an interconnect 135. In the embodiment Figure 2 illustrated, the NAMA 100a, 100b are exactly NAMCs, each including an integrated memory controller, but the various embodiments can equally employ NAMAs that do not have an integrated memory controller and may communicate with an external memory controller, for example, as Figure 1B illustrated. This system 200 is an example of a shared memory architecture that uses NAMCs 110a, 100b. The first NAMC 110a communicates with a first system memory 140a, while the second NAMC 110b communicates with a second system memory 140b.
[0076] In some embodiments, the NAMC 110 can communicate via interconnect 135 to implement a hybrid memory arrangement, wherein, for example, the CPU 205a can access and use memory 140b by issuing memory I / O requests to the NAM-C 110a, and the NAM-C 110a can, for example, communicate these memory I / O requests to the NAMC 110b via interconnect 135, for example, using the techniques described herein, and the NAMC 110b can perform the requested memory I / O on memory 140b. Similarly, the CPU 205b can issue memory I / O requests to memory 140a via the NAMC 110b, and the NAMC 110b can communicate these requests to the NAMC 110a, and the NAMC 110a can perform the requested I / O on memory 140a. In some embodiments, since each of the NAMCs is a “network-aware memory controller,” or more precisely, memory controller 105 is integrated with NAMA 100, these devices can handle all NAMC communication and memory operations regardless of location, allowing CPU 205 to ignore the location of memory 140 where the requested I / O is to be performed. From the perspective of the CPU (e.g., CPU 205a), it only needs to make regular memory I / O requests to the memory controller that acts as its local memory controller (NAMC 110a), which handles all the complexities of the hybrid memory layout.
[0077] In some embodiments, interconnect 135 may be a shared interconnect providing communication between CPU 205a, NAM-C 110a, CPU 205b, and NAMC 110b. In other cases (not specified), interconnect 135 may be split, with one interconnect providing communication between CPU 205 (e.g., CPU 205a) and its local NAMC 110 (e.g., NAMC 110a), and another interconnect providing communication between another CPU 205 (e.g., CPU 205b) and its local NAM-C 110 (e.g., NAMC 110b). Two NAMCs 110 may communicate via a third interconnect 135 and / or via one of the other two interconnects. As described in detail elsewhere herein, in some embodiments, NAMC 110 may communicate using EMRs (which may be embedded or incorporated into MPI messages or other memory requests) carried through interconnect 135. Specifically, the NAMC 110 can communicate with the CPU using MPI or regular memory I / O instructions, and it can communicate with another NAMC 110 using EMR or any other suitable communication protocol. In a sense, the NAMC 110 can be used as an interface between a local CPU and another NAMC.
[0078] In some embodiments, the NAMC 110 can communicate via a variety of different media. Figure 2 Provides an example of communication via interconnect 135. Figure 3 Another example of a system 300 in which two NAMC 110s communicate via network 305 (e.g., via network equipment 150) is provided. On one hand, Figure 3 This can be viewed as a distributed memory arrangement. In certain embodiments, network 305 may be a packet network, such as an Internet Protocol (IP) network, which can operate across various media (e.g., Ethernet, Fibre Channel, and / or the like). In some embodiments, NAMC 110 may generate and / or packetize MPI communication and / or transmit it through network 305 (e.g., via a reference). Figures 1A to 1D Such communications can be transmitted / received via network equipment 150 through network interface 130 and / or combined interface 155. Except for the nature of the transmission (e.g., network 305 and interconnect 135) and / or any necessary encapsulation (e.g., packetizing messages for transmission over the network), operation between NAMCs 110 is possible. Figure 2 and 3 The process is similar in China.
[0079] In other embodiments, the NAMC 110 may be provided in a hybrid memory arrangement, wherein a plurality of NAMCs 110 communicate via one or more interconnects 135 and / or one or more networks 305. This is merely an example. Figure 4 This describes a system 400 employing a hybrid memory layout, in which multiple CPUs 205 (e.g., 205a, 205b) have a shared memory layout, similar to the above description. Figure 2 Those described, for example, CPU 205a can access CPU 205b's local memory 140b by issuing a memory I / O request to NAMC 110a, and NAMC 110a can relay the request to NAMC 110b, for example, regarding... Figure 2 And as described elsewhere in this document. Furthermore, System 400 provides distributed memory functionality, for example, similar to that described above. Figure 3 Those described in the context. For example only, CPU 205 (e.g., CPU 205a) can access memory 140 (e.g., memory 140c, 140d) across network 305 by issuing a memory I / O request to NAMC 110a. NAMC 110a can communicate with NAMC 110c and / or NAMC 110d as appropriate to serve the I / O request from memory 140c and / or 140d, for example, as described above regarding... Figure 3 And as described elsewhere in this article.
[0080] although Figure 2 and 4 The interconnect 135 is described as a bus, but it should be understood that the topology of the interconnect 135 can vary. For example only, in some embodiments, instances of it are... Figure 5A Note that the interconnect 135 can adopt a ring topology 500, in which all nodes 505 (i.e., entities that use the interconnect to send and receive data from other entities connected to the same interconnect, such as CPU 205 and / or NAMC 110) are connected sequentially. The ring can be unidirectional (as shown by the solid line links between nodes 505) or bidirectional (as shown by both solid and dashed line links between nodes 505). In a unidirectional ring, a node can send data directly to only one other node (left or right). In the bidirectional case, a node can send data directly to two other nodes (one on its left and one on its right). Ring topologies (e.g., 500) provide good traffic control, but have relatively high latency. In a unidirectional ring with M nodes, the maximum latency will be M-1 hops, and the average latency will be M / 2 hops. In a bidirectional ring, the maximum latency will be ~M / 2 hops, and the average latency will be ~M / 4 hops. Conversely, in a mesh topology (e.g., by...) Figure 5B In the illustrated topology (550), service control is more complex than, for example, in a ring topology, but latency is relatively low. Therefore, in a 4-way mesh topology (such as that described above), service control is more complex than in a ring topology, but latency is relatively low. Figure 5B As illustrated by the solid line link between nodes 505 (on the network), the maximum delay will be approximately ~M / 4 hops, and the average delay will be ~M / 8 hops. "Omnidirectional" networks (e.g., such as...) Figure 5B (The solid and dashed lines connecting nodes 505 together illustrate this) connecting each node to all other nodes, providing a maximum latency of 1 hop, but this may be impractical for larger interconnect sizes.
[0081] It should be understood that, by Figures 1A to 1D The architectures and topologies illustrated in sections 2 to 4 and 5A to 5B are exemplary in nature and should not be considered as limiting in any way. These are merely examples, although... Figure 4 While NAMC 110a and NAMC 110b each have a direct connection to network device 150, in some embodiments, each NAMC 110 may have a connection to a separate network device, and / or such connections may be indirect, for example, via interconnect 135. Similarly, as mentioned above, although Figures 2 to 4For simplicity, NAMC 110 is described, but other embodiments can readily employ one or more NAMA 100s (possibly with an external memory controller 105 where appropriate). Furthermore, according to some embodiments, the nature of the remote memory controller and / or agent communicating with an individual NAMA 100 or NAMC 110 (e.g., issuing or fulfilling shared and / or distributed memory requests) can vary with different implementations, as long as such remote memory controllers and / or agents are capable of participating in the communications described elsewhere herein and specifically in the following description. More generally, according to some embodiments, the architecture of the apparatus and systems implementing the communication techniques and memory operations described herein is not critical, as long as these apparatus and / or systems are capable of performing the communication techniques and / or memory operations described herein. Similarly, according to other embodiments, the NAMA 100 or NAMC 110, as described architecturally herein, may employ different communication techniques and / or memory operations than those described herein without departing from the scope of these embodiments.
[0082] Exemplary Message and Communication Model
[0083] Table 6 (below) provides an exemplary non-exhaustive list of messages and some exemplary non-exhaustive fields defined for these messages, which can be exchanged by NAMC 110 according to some embodiments. Table 6 is not intended to provide an exhaustive or limiting list of messages, but rather to provide an overview of possible data structures that NAMC 100 may be able to process according to various embodiments, and to illustrate examples of various implementation-specific details to enable those skilled in the art to understand some principles of a set of embodiments.
[0084]
[0085]
[0086]
[0087]
[0088]
[0089] Table 6: Exemplary MPI Messages
[0090] In some embodiments, the MPI_send and MPI_recv EMR messages are the simplest and, in practice, are typically the most commonly used. These messages may be used, for example, by a source NAMC 110 based on a request from an initiating process to request remote memory read or write I / O from a target NAMC 110, respectively. On one hand, these messages can be considered substantially similar to the MPI_Get and MPI_Put RMA messages, respectively. On the other hand, these can be considered unicast (from one NAMC 110 to another) communication. Some embodiments also provide collective (e.g., multicast and / or broadcast) messages involving all NAMCs in a win group. MPI_Win_fence is an example of a collective message described herein. Those skilled in the art will understand from these examples that some embodiments may support a variety of different unicast and collective messages not described in detail herein.
[0091] Figures 6 to 8 This section describes exemplary communication models that may be used according to some embodiments. For illustrative purposes, most of the following description uses one-way messages as examples, but it should be understood that the principles illustrated by these models can be applied to some or all one-way and two-way messages, including (but not limited to) those described in Table 2 above.
[0092] Depend on Figures 6 to 8 The described message flow describes the behavior of four entities involved. These four entities are: the initiating process (in some embodiments, a software entity executing on a processor, such as, as illustrated, CPU 205), the local agent (in some embodiments, a hardware entity, such as NAMA 100, or as illustrated, NAMC 110a), the remote agent (in some embodiments, a hardware entity, such as NAMA 100, or as illustrated, NAMC 110b), and in some cases, the target process (in some embodiments, a software entity executing on a processor, such as, as illustrated, CPU 205b). In some embodiments, it is assumed that the remote agent resides on a different network node.
[0093] Figure 6 This describes the communication model of the message flow of the MPI_Put() operation, and is described herein to provide those skilled in the art with knowledge of implementing this operation according to various embodiments. Figure 7 and 8 The message flows for the exemplary MPI_Get() and MPI_Win_Fence() operations are illustrated separately. For the sake of brevity, these diagrams are not described in great detail, but those skilled in the art will understand. Figure 6 The principles described can also be applied to Figure 7 and 8 In the context of.
[0094] According to some embodiments, high-performance application code (e.g., artificial intelligence and / or machine learning function libraries) employs MPI_* function calls. Such calls can be compiled into low-level code that uses function identifiers to search a table for pointers to the actual MPI function code and then invokes those pointers. The function pointer table can be populated at load time. In some embodiments, instead of invoking MPI functions, the calls in the application code are replaced with calls to different functions that write data to an identified memory region located in the user memory space. The data written may be in the form of p2aReq and / or may contain an MPI message identifier and appropriate MPI message parameters. In some embodiments, the memory region can be easily created by writing code that declares several static variables of the p2aReq message type and then compiling the code along with the application code. The static variables allow the address of the region to be programmed into the NAMA configuration as a p2aMsgAddr field.
[0095] It should be understood that Figures 6 to 8 This describes the flow of a message from start (by a process) to end (by the process's local namAgent). Typically, many such messages may exist in a pipeline or queue, and therefore many embodiments will provide context for N numbers of messages in such a pipeline or queue. If the total turnaround time for a message is T seconds, and the expected throughput is P msg / sec, then the number of messages N is given by the formula N = P * T. Therefore, if the expected throughput is 50M messages per second, and the turnaround time for a message is 10μs, then context needs to be maintained for 500 messages.
[0096] According to some embodiments, NAMA can address two aspects of concurrency: (a) multiple messages from a single entity and (b) multiple messages from different entities in different win groups (e.g., local processes and / or one or more peer NAMAs). In some embodiments, since the a2aReq message is self-contained, it is not necessary to identify multiple concurrent messages within the same win group (even if they are unordered). However, in some embodiments, it is indeed necessary to identify concurrent requests belonging to multiple win groups (for example, for ucMPCount adjustment), but for this purpose, a group identifier (“win” handle) is present in the message. Some embodiments may include a msgID in the a2a message (e.g., as described above). The msgID gives each message a unique identity (of course, for a limited time, as the ID can sometimes be reused).
[0097] As mentioned above, Figure 6This describes the main steps of an exemplary message flow involving an MPI_Put() message to a source NAMC 110a and a target NAMC 110b. At operation 605, the initiating process (e.g., running on CPU 205a) sends an MPI_Put() message to the locally connected NAMC 110a. In some embodiments, the initiating process performs a “write,” “move,” or “store” of the message word at p2aMsgAddr, for example, using any technique that allows NAMC 110a to access the data. By way of example only, the process may perform a direct write, for example, using a “first-in, first-out” (FIFO) technique to write consecutive messages to the same address (e.g., using DMA, using message rings, etc.). In some cases, such writes may accumulate in the processor’s data cache, where they arrive at the nearest interconnect node in bursts. On the one hand, some embodiments may allow the entire MPI message to be loaded into a single burst; for example, according to some embodiments, a 64-byte burst size can accommodate most MPI messages. This message can be categorized as an Enhanced Memory Request (EMR), or more specifically, a p2aReq message.
[0098] At operation 610, NAMC 110a matches the start address of the write operation from the interconnect with p2aMsgAddr, and if a match is found, NAMC 110a identifies the message as a p2aReq message from the local process. NAMC 110a stores the entire message in local storage (e.g., ...). Figures 1A to 1D The NAMC 110a (as shown above, NAMC 115) decodes the message after storing it and determines that the message is p2aReq (where msgSubtype is 0). NAMC 110a also determines that mpiMsgType implies that the message is MPI_Put. In some embodiments, NAMC 110a increments the ucMPICount of the win group (defined by the "win" handle). On the one hand, if NAMC 110a receives the message during multiple bursts, it may maintain sufficient context to cascade the fields of the message.
[0099] At operation 615, NAMC 110a generates an a2aReq message, which includes the data to be "placed" (i.e., an IO write request) to be transferred to the target NAMC 110b. In some embodiments, NAMC 110a may first use an integrated memory controller (or, in the case of NAMA without an integrated controller, send a read request to an external memory controller to obtain the "placed" data from "origin_addr"). NAMC 110a may read the amount of data indicated by SizeOf(origin_datatype)*origin_count. Based on this data, NAMC 110a creates a PFR message. NAMC 110a may determine its own agentID, for example, based on configuration variables, and the agentID of the destination namAgent (e.g., NAMC 110b) located in the "win" window, which can be used as the target memory handle. Using these identifiers, NAMC 110a may generate the a2aReq message. Figure 6 In the illustrated embodiments, the target NAMC 110b is accessible via a packet network, thus the NAMC 110a can perform packetization of the a2aReq message. In some embodiments, the NAMC 110b stores all necessary header information (UDP port, IP address, etc.) for each potential peer agent, and therefore the NAMC 110a can add such header information associated with the NAMC 110b to the packet. Once a packet including the a2aReq message has been generated, the NAMC 110a sends the packet, for example, via network device 150, as discussed above. If the NAMC 110a is configured to communicate with network device 150 via interconnect 135, for example, as by... Figure 1D As explained, NAMC 110a can add interconnect node address information to packets so that packets can be routed to network equipment 150 via the interconnect.
[0100] Operation 620 describes the journey of a packetized a2aReq message (with any necessary additional routing headers) from NAMC 110a to the destination NAMC 110b. At the destination network node, the network equipment can deliver the a2aReq message (with any necessary additional routing headers) to NAMC 110b directly or via an interconnect. If routed via an interconnect, the network equipment can provide an a2aMsgAddr, allowing the interconnect to write the a2aReq message to NAMC 110b.
[0101] In some embodiments, certain aspects of operation 625 at NAMC 110b may be similar to those of operation 610 at NAMC 110a. In this case, msgSubtype may indicate that it comes from another namAgent (e.g., NAMC 110a), and the message may be a PFR request. In one aspect of some embodiments, NAMC 110b may, for example, increment its local ucMPICount (or ucMPICount_t) variable for a given “win” upon receiving an a2aReq message, and / or decrement the variable (hereinafter) upon completion of operation 635. At operation 630, NAMC 110b writes the payload from the a2aReq message (i.e., the data requested to be written to memory) to an address in the “win” region, for example, calculated using the address described in Table 6 regarding the MPI_Put() message at the target NAMA.
[0102] At operation 635, NAMC 110b may generate a CRM. In some embodiments, for example, the a2aCRM message may take the format and / or fields described in Table 2. NAMC 110b may add routing information to the a2aCRM message, packetize it, and / or send the resulting packets to the nearest switch port or network equipment. These operations may be similar to those described above with respect to operation 610. Operation 640 describes the journey of the a2aCRM message through the network to the source NAMC 110b. Ultimately, this message arrives at the source namAgent.
[0103] At operation 645, NAMC 110a decrements its local ucMPICount (for the given "win"), and at box 650, NAMC 110a generates an a2pCRM message, which in some respects indicates to the source NAMA that the MPI_put operation is complete, thus allowing the source NAMA to decrement its local ucMPICount_t and, if applicable, indicate to the initiating process that the operation is complete. The format and / or fields of the a2pCRM message according to some embodiments are described in Table 2 above.
[0104] In some embodiments, the MPI_Accumulate() message stream can be similar to that regarding Figure 6 The message stream described for MPI_Put(). However, in some embodiments, at operation 635, the target NAMC 110b may read the raw data, perform the requested operation (as indicated by “op”), and then write back the resulting data, rather than replacing the data. Further details about this message type can be found in Table 6.
[0105] Figure 7This illustrates an exemplary communication model for the MPI_Get() message flow 700. Operations 705 to 755 can be similar. Figure 6 The corresponding operations 605 to 655 differ from the following: at operation 705, the initiating process (e.g., running on CPU 205a) sends an EMR message to the locally connected NAMC 110a that embeds an MPI_Get() message instead of an MPI_Put() message.
[0106] At operation 710, NAMC 110a generates the a2aReq message in a manner similar to operation 605. However, since the message is an MPI_Get() message, NAMC 110a does not need to read any data. In one embodiment, NAMC 110a may swap the source and destination fields from the corresponding MPI_Get() type a2aReq message to generate an MPI_Put() type a2aReq message; in some cases, NAMC 110a may modify msgSubtype to indicate that such a swap has been performed. In other embodiments, this swap may be performed at the target NAMC 110b.
[0107] At operation 725, the target NAMC 110b receives and processes the a2aReq message in a manner similar to that of operation 625. However, instead of writing data to memory 140b, it reads the requested data from a specified address in the "win" memory at operation 730 (e.g., as described in Table 6). At operation 735, the target NAMC 110b generates a PFR2aReq message from the message received from the source NAMC 110a, indicated by msgSubtype. Alternatively and / or further, the NAMC 110b may simply convert the incoming a2aReq message into a PFR message, for example, by appropriately changing the subtype. The target NAMC 110b then packets and transmits the message, for example, as described above.
[0108] At process 745, source NAMC 110a processes the incoming a2aReq message and determines from the message the data to be written to the initiating process's memory 140a. At block 750, the data is written to memory 140a, and at block 755, source NAMC 110a sends an a2p message to the initiating process. In this case, if the request from the initiator enables the operation, the message may be an a2pCRM message; otherwise, the message may be omitted if the initiating process does not request confirmation of the MPI_Get() operation.
[0109] Figure 8This describes the main steps of an exemplary message flow for an MPI_Win_fence type message. In various embodiments, many messages may exist in this category. For example only, MPI_Win_start() and MPI_Win_complete() may define access periods; MPI_Win_post() and MPI_Win_wait() may define exposure periods. While those skilled in the art will understand that the affected processes and the order in which fence-related flags are set and cleared vary depending on the message, all fence-related messages can be implemented using the same methods as those described above. Figure 8 The message flow described is similar.
[0110] At operation 805, the initiating process sends a p2aReq message to the source NAMC 110a, for example, as described above regarding operation 605, the message type is MPI_Win_fence. At block 810, NAMC 110a identifies the message as a fence message. For example, based on this identification, NAMC 110a may determine whether the ucMPICount of “win” identified in the message is zero. If not, then in some embodiments, NAMC 110a sets the flag ucMPIFenceFlag (if it has not already been set) and completes the operation. At this point, NAMC 110a may return to process the next message. NAMC 110a may maintain the ucMPIFenceFlag setting until the execution of one or more future MPI messages reduces the ucMPICount to zero.
[0111] At block 810, if ucMPICount is zero (e.g., upon arrival of the MPI_Win_fence message or due to the completion of other pending messages), then NAMC 110a may generate, store, and / or transmit a2pCRM of the fence type. In some embodiments, the message may be sent via a message loop; in other embodiments, it may not be sent (e.g., if the calling process is waiting for the return of an a2pCRM message). In this case, there may be a separate location in memory for NAMC 110a to store a2pCRM messages for different “win” groups (operation 815). In some embodiments, NAMC 110a may be configured to send an interrupt message (operation 820) to indicate fence completion; in this case, a2pCRM messages may not be generated. In some embodiments, operation 815 may involve sending the a2pCRM message directly to the CPU cache.
[0112] Methods performed by various embodiments
[0113] Figures 9 to 15This section describes various methods for managing memory using memory agents (e.g., network-aware memory agents), examples of which include (but are not limited to) the NAMA 100 and NAMC 110 described in detail above. Those skilled in the art will understand that... Figures 9 to 15 The description includes methods for exemplary operations that can be performed partially or entirely by an agent (such as the NAMA 100 or NAMC 110 described in detail above). These operations are exemplary in nature, and the operation of such agents can generally vary depending on configuration variable settings and / or message fields, as described in Tables 2 through 6 above. For example, in some cases, the methods may describe the transmission of a response, or any such description may be omitted. However, in implementations, whether and how the agent responds may depend on the setting of values for, for example, configuration variables (such as enableA2pCRM, enableA2aCRM, enableFenceCRM, ucMPICount, ucMPIFenceFlag, and / or similar).
[0114] Although Figures 9 to 15 For descriptive purposes, it is described as a separate method, but it should be noted that... Figures 9 to 15 Some or all of the methods described herein and / or their various operations may be performed in combination and / or may be considered as part of a single method. Similarly, while the operations of these methods are shown in a specific order, those skilled in the art will understand that, within the scope of various embodiments, the various operations of the different methods described may be combined, reordered, and / or omitted. “Managing” memory (e.g., system memory) may include some or all of such operations, and more generally may include any (or any subset of) the functionality attributable to NAMA herein, including (but not limited to) performing memory operations, controlling and / or providing access to memory and / or memory windows, providing access to system memory and / or communicating with system memory, and / or the like.
[0115] Depend on Figures 9 to 15 The methods described may employ and / or implement various functions, messages, and / or message flows, including (but not limited to) the messages, message types, fields, functions, and / or variables described in Tables 2 to 6 above, as well as the various techniques described in detail above. However, it should be understood that unless explicitly stated in the context of a particular method or its operation, all other interpretations are subject to change. Figures 9 to 15 The methods and operations described are not limited to any particular architecture, method, message format, or data structure.
[0116] Figure 9 This describes a method for operating a memory agent. In some embodiments, the memory agent may be NAMA or NAMC, and non-limiting examples have been described above. Figures 1A to 1D Described in the context of 2 to 4 and 6 to 8.
[0117] At box 905, method 900 includes identifying one or more local memory windows, each of which may include at least a portion of system memory (e.g., in...). Figure 1A In the context of NAMA 100 of D, memory 140, or... Figures 2 to 4 In the context of NAMC 110a, this is memory 140a. In some embodiments, such as those mentioned above, configuration variables can be used to configure an agent to identify one or more local memory windows (“win”), including variables that hold handles to the win windows and / or variables that store the address space in local memory (e.g., 140) corresponding to each of the win windows. These variables can be used by the agent to identify the local memory windows.
[0118] At box 910, method 900 includes identifying one or more remote memory windows, each of which can be managed by a remote memory agent (e.g., in...). Figures 2 to 4 In the context of NAMC 110a, memory 140b is managed by NAMC 110b. In some embodiments, such as those mentioned above, configuration variables can be used to configure an agent to identify one or more remote memory windows (“win”), including variables that, for example, hold handles to the win windows and / or variables that store the address space in the remote memory (e.g., 140b) corresponding to each of the win windows. These variables can be used by the agent to identify the remote memory windows. In some embodiments, the agent may store communication information for each of the remote agents managing the remote memory windows, such as the interconnect address and / or IP address of the remote memory agent; such information can be used to construct message headers and / or packetize request messages to be sent to the remote agents, for example, as further described elsewhere herein.
[0119] At box 915, method 900 includes interconnection with a local CPU (e.g., Figures 2 to 4 The agent may communicate with the local CPU (CPU 205a) in the local CPU, for example, as further described above. For example, in some embodiments, the agent may communicate with the local CPU using a2p messages and / or p2a messages, as further described herein. In some embodiments, communication with the local CPU may include communication with processes running on the local CPU, such as application processes, operating system processes, and / or the like.
[0120] At block 920, method 900 includes communicating with a remote memory agent via an interconnect. As described herein, such communication may include MPI messages, which may take the form of various a2a messages described elsewhere herein. At block 925, method 900 includes communicating with a remote memory agent via a packet network. In some aspects, the operation of communicating with a remote memory agent via a packet network may be similar to that of communicating with a remote memory agent via an interconnect, except that, in some embodiments, communication via a packet network may further include packetizing outgoing messages and / or depacketizing incoming messages, as further described in detail herein.
[0121] At box 930, method 900 includes managing local memory. Managing local memory may include various operations, including (but not limited to) handling memory I / O requests addressed to said local memory (and / or addressed to memory windows including some or all of the local memory) (e.g., such as...). Figure 12 (in the context of this document and elsewhere), a queue for managing memory requests (e.g., such as...) Figure 13 (in the context of and elsewhere in this document), handling network computing requests (e.g., such as...) Figure 14 (in the context of and elsewhere described herein) and / or manage memory fences on local memory (e.g., as described elsewhere) Figure 15 (As described in the context of and elsewhere in this document), to name just a few examples. As used herein, in the context of a memory window, the verbs “address” and “bootstrap”, and their variations, are used, for example, if a memory request seeks to perform an operation on memory (e.g., I / O operation, fence operation, network computing operation, etc.) or otherwise relate to memory within or associated with a memory window in the memory window (e.g., containing a memory address), then the memory request is “addressed” or “addressed to” or “bootstrap” to the memory window.
[0122] At box 935, method 900 includes handling requests for one or more local memory windows. Various operations (e.g., as further described in detail herein) can be used to handle memory requests for local memory windows. (Example only) Figure 12 (Discussed in further detail below) An exemplary method for handling memory requests addressed to a local memory window is described.
[0123] At box 940, method 900 includes handling requests for one or more remote memory windows. Various operations (e.g., as described further in detail herein) can be used to handle memory requests for remote memory windows. (Example only) Figure 10 and 11(Discussed in further detail below) An exemplary method for handling memory requests directed to a remote memory window.
[0124] For example, Figure 10 This describes a method 1000 for handling memory requests addressed to a remote memory window (e.g., a memory window not managed by the agent executing method 1000). In some embodiments, method 1000 may be implemented similar to... Figures 6 to 8 A communication model that describes one or more communication models.
[0125] At box 1005, method 1000 includes receiving a memory request. In some cases, the memory request may be received from the local CPU and / or from a process executing on the local CPU (e.g., as in...). Figure 6 and 7 (Discussed in the context of this document). Memory requests may include DMA requests, MPI messages (e.g., encapsulated as p2a EMR messages and / or similar, non-limiting examples of which are described in the table above). By way of example only, in some embodiments, the request includes an MPI message from the local CPU via an interconnect. In other cases, the memory request may be received from different sources. Memory requests may contain any kind of memory I / O or memory management operation, including (but not limited to) those described herein, such as write I / O (e.g., send, write, or place requests), read I / O (e.g., receive, read, or get requests), network computing requests, fence requests, and / or similar.
[0126] At box 1010, method 1000 includes determining a memory request addressing remote memory window (e.g., a memory window that does not include the first system memory). For example, as described above... Figure 9 As described in the context and elsewhere, a proxy may include a memory that stores configuration variables, and the proxy may perform lookups against one or more such variables to identify the memory window to which the request is addressed (also referred to herein as a bootstrap). In other cases, the request itself may contain information from which the memory window can be identified and / or from which the identity of the memory window can be derived.
[0127] At block 1015, method 1000 includes identifying a target memory agent associated with a memory window based on a memory window identifier in a first memory request. By way of example only, in some cases, as mentioned above, a local agent may store the identity of each remote agent that manages one or more memory windows addressable by any request from the local agent.
[0128] At box 1020, method 1000 includes generating a remote agent request message, such as an EMRa2aReq message corresponding to a memory request. By way of example only, in some cases, the request message may be generated from a request (e.g., if the request is an MPI message from the local CPU), may have a memory request (and / or a portion thereof) embedded therein, may include the request message (and / or a portion thereof), and / or may be derived from the request message. In this context, the term "corresponding to" means that the request message has sufficient information to enable the remote agent to perform the requested memory operation (e.g., read, write, or accumulate data as the object of the request, establish a fence for the memory window corresponding to the request, etc.).
[0129] In some embodiments, method 1000 includes identifying packet header information associated with a target memory agent (block 1025), packetizing a request message (block 1030), and / or, for example, transmitting one or more data packets via a packet networking device for reception by the target memory agent (block 1035). Each of these operations may be performed, for example, as described in further detail above.
[0130] The following text is about Figure 12 As described, some embodiments may employ various techniques, such as using Explicit Congestion Notification (ECN) information, to avoid packet loss when exchanging EMR messages over a packet network. In this case, method 1000 may include obtaining ECN information and / or transmitting such ECN information (block 1040). Those skilled in the art will appreciate that several techniques exist for performing such operations, and any of such techniques may be implemented according to various embodiments. By way of example only, in some cases, NAMA may mark one or more packet headers (e.g., one or more bits of the traffic category field of the header) to indicate ECN capability, which may allow network equipment (e.g., switches, routers) to further mark packets to provide additional ECN information, for example, by modifying bits of the traffic category header to indicate experienced congestion.
[0131] At block 1045, method 1000 may include receiving a packetized response message from a target memory agent via a packet network (e.g., via a packet network apparatus). At block 1050, method 1000 includes generating a response to the original request. In some embodiments, the response may include an EMR message (e.g., an a2P message, such as an a2pCRM message, etc.) that may encapsulate an appropriate MPI message, and / or the response to the request may be based at least in part on the packetized response message. An example of such an EMR message may be an a2pCRM message, as described further in detail herein. At block 1055, method 1000 includes transmitting the MPI response message via an interconnect for reception by a CPU.
[0132] Figure 11 Method 1100 describes a method for managing memory requests addressed to remote memory windows (e.g., memory windows not managed by the agent executing method 1100).
[0133] In some respects, the various operations of method 1100 can be similar to Figure 10 The corresponding operations of method 1000. For example, method 1100 includes receiving a memory request (box 1105), determining the remote memory window addressed by the memory request (box 1110), and identifying the target memory agent associated with the memory window (box 1115).
[0134] At block 1120, method 1100 includes determining a communication route for a target memory agent. In some aspects, the identity of the target memory agent can determine the communication route for the target memory agent. For example, as mentioned above, one or more configuration variables can provide the identity of the target memory agent for a given memory window and / or may include addressing information for the agent. Such addressing information can be used to determine the communication route for the agent. One or more variables (or another data structure) stored in the memory of a local agent can identify a first target agent for a first memory window, indicate that the first target agent is accessible via an interconnect, and provide addressing information for communicating with the first target agent via the interconnect. One or more variables (or another data structure) can identify a second target agent for a second memory window, indicate that the first target agent is accessible via a packet network, and provide addressing information (e.g., IP address) for communicating with the second target agent via a packet network.
[0135] Method 1100 further includes generating a remote agent request message (box 1125). This operation may be similar to the one about... Figure 10 The operation described is 1020.
[0136] If the communication routing indicates that the communication of the target memory agent is routed via the interconnect, then method 1100 includes transmitting a request message via the interconnect (block 1130).
[0137] If the communication routing indicates that the communication of the target memory agent is routed via a packet network, then method 1100 includes packetizing the message using packet header information from the addressing information to generate one or more data packets (box 1135) and / or transmitting the request message as one or more packets via the packet network (box 1140). These operations may be similar to those described above. Figure 10 The described operations are 1030 to 1035.
[0138] Although not by Figure 11 Note that method 1100 may further include similar components as described above. Figure 10 The operations described in boxes 1040 to 1050.
[0139] Figure 12 This describes a method 1200 for handling memory requests addressed to a local memory window. In some embodiments, method 1200 may be provided by a local memory management system (e.g., Figures 1A to 1D The agent (e.g., target NAMA) of NAMA 100 and memory 140 is executed.
[0140] At block 1205, method 1200 includes receiving a memory request from a peer memory agent, for example, via a communication interface. The memory request may be received via an interconnect (e.g., from a local CPU or from a remote agent) or via a packet network (e.g., from a remote agent).
[0141] At box 1210, method 1200 includes determining a memory request addressing a local memory window. In some aspects, the request may be an MPI message, such as an a2a or p2a message, and the request may include a win handle field (e.g., as described in the table and description above). By examining the win handle field, the local agent can identify the win handle and determine, based on configuration variables, that the win handle addresses a local memory window (e.g., a memory window that includes at least a portion of the system memory managed by the agent).
[0142] At block 1215, method 1200 includes performing a memory request on the first system memory based on determining a memory request addressing local memory window. In some cases, method 1200 (or a portion thereof) may be performed by a NAMA located adjacent to (or otherwise communicating with) an external controller, and in this case, performing the memory request may include communicating the first memory request to the memory controller (block 1220). In other cases, method 1200 (or a portion thereof) may be performed by a NAMC having an integrated memory controller; in this case, method 1200 may include causing the included memory controller to perform the first memory request on the first system memory (block 1225).
[0143] Although not in Figure 12 The above description may vary, but method 1200 may be adjusted depending on the nature of the request and / or how it is received. For example, in some cases, the response may include MPI messages, such as a2a or a2p messages, depending on the source of the request. The type of message transmitted may depend on the nature of the request. For example, as described above... Figure 6 As explained, if the request is an MPI_Get() request from a remote proxy, the response message may include an a2aReq message, while if the request is an MPI_Put() message, the response may include an a2aCRM message. Examples of request types and corresponding messages are described in further detail above.
[0144] As another example of this variant, if the request is received via a packet network, the method may further include degrouping the request and packetizing the response for transmission over the packet network, for example, as described in further detail above. Furthermore, as mentioned above, in some cases, messaging according to some embodiments may employ connectionless communication and Universal Datagram Protocol (UDP) packets, without the overhead of connection-oriented protocols such as Terminal Control Protocol (TCP). In this case, various embodiments may employ measures to reduce the risk of communication loss when transferring data for memory I / O.
[0145] By way of example only, in some embodiments, method 1200 may include receiving explicit congestion notification (ECN) information from a separate source and / or the like, for example, as part of a request referenced at box 1210 (box 1230). By way of example only, as mentioned above, one or more request packets may be marked with ECN information (e.g., marked in one or more bits of the service category field in the packet header to indicate ECN capability and / or indicate congestion experienced during the delivery of the packet). At box 1235, method 1200 may include determining the timing of the transmission of one or more response IP packets to reduce the risk of communication loss (e.g., the timing of transmissions referenced by box 1040 discussed below). This determination may be based at least in part on the ECN information. For example, a local agent may periodically receive ECN information and / or use this information to track or estimate congestion trends. The agent may then delay the transmission of response packets until the ECN information indicates that the current congestion is lower than historical values, the congestion is estimated to increase in the future, and / or the congestion is currently at an acceptable level (e.g., below a defined threshold) and / or is estimated to remain at an acceptable level for a defined period of time, in order to allow low-risk transmission of response packets.
[0146] At block 1240, method 1200 includes transmitting a response to be received by the peer memory agent. As mentioned above, the techniques used for transmitting the response may vary depending on the nature of the request and / or implementation-specific details (e.g., whether the embodiment implements ECN, etc.).
[0147] Figure 13 This describes a method 1300 for managing a queue of memory requests. In some embodiments, some or all of method 1300 may be managed by managing local memory (e.g., Figures 1A to 1D The agent execution of NAMA 100 and memory 140 in the memory).
[0148] At block 1305, method 1300 includes maintaining at least one request queue. In some embodiments, the request queue may be stored in the agent's onboard memory (e.g., Figures 1A to 1D(NAMA context memory 115). On one hand, each queue can be used to store requests for a specific memory window (or multiple memory windows) managed by the agent until those requests can be processed. Thus, at block 1310, method 1300 includes storing multiple received requests in at least one queue.
[0149] At block 1315, method 1300 includes maintaining a source context that associates each message in the queue with the source of the message. In some embodiments, the source context identifies the source of the request (e.g., a remote agent, a process running on the local CPU, etc.); this information may include, for example, the value of the message's "agentIDSrc" field (e.g., as described in the contexts of the various message types in Table 2). At block 1320, method 1300 includes maintaining a window context that associates each message in the queue with the memory window to which the request is directed. The window context may include, for example, a window handle, such as the value of the request message's "win" field (as described in the contexts of the various message types in Table 2 above). This context data may be stored along with the stored request and / or may be stored in a location (e.g., in NAMA context memory 115 that addresses, is addressed by, or is otherwise linked to / from the request).
[0150] At block 1325, method 1300 includes processing multiple received requests in a queue based on a source context and a window context. For example, the agent may process each request in a specific order, and in processing the request, the source context and window context may be referenced to determine how to process the request (e.g., the location or window in system memory 140 for processing the request, where and how to send a response to the request, etc.).
[0151] Figure 14 This describes a method 1400 for handling network computing requests. In some embodiments, method 1400 may be managed by a system that manages local memory (e.g., Figures 1A to 1D The agent execution of NAMA 100 and memory 140.
[0152] At box 1405, method 1400 includes receiving a computation request within the network. The computation request may request the execution of any computational operation that the agent can perform and / or is configured to perform, such as arithmetic operations, logical operations, etc., involving one or more values stored in a memory window managed by the agent. Instances may include (but are not limited to) accumulation operations (e.g., MPI_Accumulate(), MPI_Get_Accumulate(), and MPI_Fetch_and_op() requests, as described further in Table 6).
[0153] At block 1410, method 1400 includes performing an in-network computing function in response to receiving an in-network computing request. In some aspects, the agent may include hardware circuitry configured to perform such functions and / or may employ firmware, a processor, or external computing resources to perform such functions. In some embodiments, the computing function may include performing operations on data provided in the request and / or data retrieved from local memory (e.g., as part of the request) (block 1415). At block 1420, method 1400 includes transmitting a response to the in-network computing request, for example, using any of the various techniques described herein, where appropriate for the source of the request.
[0154] Figure 15 This describes a method 1500 for managing fences on a local storage window. In some embodiments, method 1500 may be provided by managing local storage (e.g., ...). Figures 1A to 1D The proxy execution of NAMA 100 and memory 140). In some embodiments, method 1500 can be used to implement the above regarding Figure 8 The communication model described.
[0155] At block 1505, method 1500 includes receiving a memory request addressed to a local memory window. As mentioned above, in some embodiments, a single NAMA can manage multiple memory windows. In such embodiments, each request may be directed to (e.g., directed to) a specific memory window, for example by including the win handle as a field in the memory request, as described above. For the purposes of this example, we will assume that a specific request is directed to a window called Win_A.
[0156] In some cases, memory requests include fence messages from a specific source (e.g., a specific local initiator process, a source NAMA—in this case, method 1500 may be executed by the target NAMA—and / or the like). Examples of such fence messages are described, for example, in Table 6 and elsewhere herein. Thus, at box 1510, method 1500 may include determining whether the request includes a fence message. For the purposes of this example, if the request is not a fence request (e.g., a request for another memory operation, such as MPI_Get(), MPI_Put(), MPI_Accumulate(), etc.), then method 1500 may include determining at box 1515 whether a fence has been set for the window to which the request is directed (e.g., in this example, determining whether a fence flag has been set for Win_A). If no fence has been set for the target window, then the method may include at box 1520 queuing the request to be executed (e.g., for any requested I / O or other memory operation to be executed). As mentioned herein, in some cases, the queue may have the context of the associated window (here, Win_A). Then, at box 1525, method 1500 may include incrementing a pending request counter (e.g., the ucMPICount field) of the target window to indicate that there are requested memory operations directed to the target window that have not yet been performed.
[0157] At this point, method 1500 may proceed to box 1530, where method 1500 may include waiting for new requests and / or processing existing requests, for example, as described in further detail below.
[0158] Briefly returning to box 1510, if the request is a fence message, then method 1500 may include setting a fence flag at box 1535 to establish a memory fence on the target memory window (Win_A in this case). At this point, method 1500 may proceed to box 1570 to determine whether a fence flag has been set (yes in this case) and whether the pending request counter for the memory window is zero. This procedure is described in further detail below.
[0159] Returning to box 1515, if the immediate request is not a fenced request and a fence has been set, then method 1500 may include determining whether the request is directed to the same window on which a fence is set (box 1540). If the request is directed to a different window (e.g., if the request is directed to Win_A and a fence is set on Win_B), then the method proceeds to box 1520, where the request is queued for the target window, and box 1525, where a counter is incremented for the target window. On the other hand, if the request is directed to a window on which a fence has been set (e.g., if the request is directed to Win_A and a fence has been set on Win_A), then method 1500 may include holding the request at box 1545 (e.g., after the pending request counter for Win_A reaches zero and / or no fence is set for Win_A) and / or simply ignoring the request, in which case the source of the request may later retry a new fenced request. At this point, the method returns to box 1530.
[0160] Then, at block 1530, the process may include waiting for new incoming requests and / or processing pending requests. In some embodiments, these operations may be performed in parallel, while in other embodiments, method 1500 may process pending requests only when no new requests are incoming.
[0161] Therefore, at block 1550, if a new memory request (of any type) is available (e.g., if a new memory request has been received from any source), then method 1500 may be repeated from block 1505, as illustrated. Meanwhile, and / or when no new request is received, method 1500 may include executing one or more memory requests, for example, by performing any requested I / O or other memory operations on local memory in any window to which the request is targeted (block 1555). In some embodiments, memory requests in different queues (e.g., for different windows) may be executed in parallel, while in other cases, the requests may be executed serially. After a request has been executed, method 1500 may include decrementing a pending request counter applied to any window / queue (block 1565), for example, the queue of the window to which the request is targeted and / or the queue to which the request was previously placed (in many embodiments, this will be the same queue for the reasons described above).
[0162] In some cases, method 1500 includes checking whether the pending request counter of the window has reached zero. This can occur at any time, and / or specifically, if a fence flag has been set (e.g., at box 1535) and / or the window's counter has decremented (e.g., at box 1565). If no fence flag is set on the window to which the currently executing request is targeted, and / or if the pending request counter is greater than zero (box 1570), then the method may return to box 1530. However, if a fence flag is set on the window, and the window's counter has reached zero, then method 1500 may include unsetting (i.e., clearing) the fence flag (box 1575), and / or transmitting a fence completion response message to the source that initially requested the fence flag (e.g., the initiating process, source NAMA, etc.) (box 1580), for example using various techniques described herein for sending completion messages. In this case, method 1500 may continue from box 1530.
[0163] Exemplary computing environment
[0164] Figure 16 This is a block diagram illustrating an example of device 1600, which can function as described herein, including (but not limited to) acting as a computing node, including a computer system such as NAMA or NAMC according to various embodiments, and / or performing some or all of the operations of the methods described herein. Figure 16 No component shown herein should be considered necessary or essential to every embodiment. For example, many embodiments may not include a processor and / or may be implemented entirely as hardware or firmware circuitry. Similarly, many embodiments may not include input devices, output devices, or network interfaces.
[0165] Based on this, such as Figure 16 As shown, device 1600 may include bus 1605. Bus 1605 may include one or more components that enable wired and / or wireless communication between components of device 1600. Bus 1605 can... Figure 16 Two or more components are coupled together, for example via operative coupling, communicative coupling, electronic coupling, and / or electrical coupling. Such components may include a processor 1610, a non-volatile storage device 1615, a system memory (e.g., random access memory (DRAM)) 1620, and / or a circuit system 1625. In some cases, system 1600 may include a human-machine interface component 1630 and / or a communication interface 1635.
[0166] Although these components are shown as being integrated within device 1600, some components may be located outside device 1600. Therefore, in addition to or besides the components themselves, device 1600 may include facilities for communicating with such external devices, and thus in some embodiments, these facilities may be considered part of device 1600.
[0167] By way of example only, non-volatile storage device 1615 may include hard disk drives (HDDs), solid-state drives (SSDs), and / or any other form of persistent storage device (i.e., storage devices that do not require power to maintain the state of stored data). While such storage devices are typically incorporated into device 1600 itself, they may be located external to device 1600 and may include external HDDs, SSDs, flash drives, etc., as well as networked storage devices (e.g., shared storage devices on a file server), storage area network (SAN) storage devices, cloud-based storage devices, and / or the like. Unless the context otherwise indicates, any such storage device may be considered part of device 1600 according to various embodiments. On the one hand, storage device 1615 may be non-transitory.
[0168] Similarly, the human-machine interface 1630 may include an input component 1640 and / or an output component 1645, which may be located within, outside, and / or in combinations thereof in the device 1600. The input component 1640 enables the device 1600 to receive input, such as user input and / or sensed input. For example, the input component 1640 may include a touchscreen, keyboard, keypad, mouse, button, microphone, switch, sensor, GPS sensor, accelerometer, gyroscope, and / or actuator. In some cases, such components may be located outside the device 1600 and / or communicate with components inside the device 1600 (e.g., input jacks, USB ports, Bluetooth radios, and / or the like). Similarly, the output component 1645 enables the device 1600 to provide output, such as via a display, printer, speaker, and / or the like, any of which may be located inside the device 1600 and / or outside the device but communicate with internal components (e.g., USB ports, Bluetooth radios, video ports, and / or the like). Similarly, unless the context otherwise indicates, any such component may be considered part of the apparatus 1600 according to various embodiments.
[0169] From these examples, it should be understood that various embodiments can support various arrangements of external and / or internal components, all of which can be considered part of the apparatus 1600. In some embodiments, some or all of these components may be virtualized; instances may include virtual machines, containers (e.g., Docker containers), cloud computing environments, Platform as a Service (PaaS) environments, and / or the like.
[0170] On one hand, the non-volatile storage device 1615 can be considered as a non-transitory computer-readable medium. In some embodiments, the non-volatile storage device 1615 can be used to store software and / or data used by the device 1600. Such software / data may include an operating system 1650, data 1655, and / or instructions 1660. The operating system may contain instructions for controlling the basic operation of the device 1600, and depending on the nature of the device 1600, may include various personal computer or server operating systems, embedded operating systems, and / or the like. Data 1655 may contain any of various types of data used or generated by the device 1600 (and / or its operation), such as media content, databases, documents, and / or the like. Instructions 1660 may contain software code (e.g., application programs, object code, assembly, binary, etc.) for programming the processor 1610 to perform operations according to various embodiments. On one hand, in some embodiments, the operating system 1650 may be considered as part of the instructions 1660.
[0171] Processor 1610 may include one or more of the following: a central processing unit, a graphics processing unit, a microprocessor, a controller, a microcontroller, a digital signal processor (DSP), programmable logic (e.g., a field-programmable gate array (FPGA), an erasable programmable logic device (EPLD), etc.), an application-specific integrated circuit (ASIC), a system-on-a-chip (SoC), and / or another type of processing component. Processor 1610 may be implemented in hardware, firmware, or a combination of hardware, firmware, and / or software. In some embodiments, processor 1610 includes one or more processors capable of being programmed to perform one or more operations or processes described elsewhere herein.
[0172] For example, in some embodiments, device 1600 may include logic 1665. Such logic may be any kind of code, instructions, circuitry, etc., that enables device 1600 to operate according to embodiments herein (e.g., to perform some or all of the processes and / or operations described herein). By way of example only, logic 1665 may include instructions 1660, which may be stored on non-volatile storage device 1615 as mentioned above, loaded into working memory 1620, and / or executed by processor 1610 to perform operations and methods according to various embodiments. In one respect, these instructions 1660 may be considered as programming processor 1610 to operate according to such embodiments. Similarly, operating system 1650 (within its scope separate from instructions 1660) may be stored on non-volatile storage device 1615, loaded into working memory 1620, and / or executed by processor 1610.
[0173] Alternatively and / or additionally, the logic may include a circuit system 1625 (e.g., hardware or firmware) that may operate independently of or in cooperation with any processor 1610 that device 1600 may or may not have. (As mentioned above, in some cases, circuit system 1650 itself may be considered processor 1610.) Circuit system 1625 may be embodied by a chip, SoC, ASIC, programmable logic device (FPGA, EPLD, etc.), and / or the like. Thus, some or all of the logic that enables or causes some or all of the operations described herein to be performed may be encoded in and executed directly by a hardware or firmware circuit system (e.g., circuit system 1650), rather than by software instructions 1660 loaded into working memory 1620. (In some cases, this functionality may be embodied by hardware instructions.) Therefore, unless the context otherwise indicates, the embodiments described herein are not limited to any particular combination of hardware, firmware, and / or software.
[0174] Device 1600 may also include a communication interface 1635 that enables device 1600 to communicate with other devices via wired (e.g., electrical and / or optical) and / or wireless (RF) connections. For example, communication interface 1635 may include one or more RF subsystems (e.g., Bluetooth subsystems, such as those described above, e.g., Wi-Fi subsystems, 5G or cellular subsystems, etc.). Some such systems may be implemented in combination as discrete chips, SoCs, and / or the like. Communication interface 1635 may further include a modem, network interface card, and / or antenna. In some cases, communication interface 1630 may include multiple I / O ports, each of which may be any facility providing communication between device 1600 and other devices; in certain embodiments, such ports may be network ports, such as Ethernet ports, fiber optic ports, etc. Other embodiments may include different types of I / O ports, such as serial ports, pin outputs, and / or the like. Depending on the nature of device 1600, communication interface 1635 may include any standard or proprietary components to allow communication as described according to various embodiments.
[0175] In some embodiments, the apparatus 1600 includes a memory controller 1670 that performs memory I / O operations on system memory 1620 and / or NAMA 1675, and that manages one or more memory windows 1680, such as those described in detail above. In some embodiments, NAMA 1675 may be similar to NAMA 100 described above. In some embodiments, the memory controller 1670 may be external to and / or adjacent to NAMA 1675. In some embodiments, NAMA 1675 may be incorporated into, included in, integrated with, packaged with, include the memory controller 1670, and / or be a component of the memory controller 1670, in which case NAMA 1675 may be an NAMC similar to NAMC 110 described above.
[0176] In some embodiments, NAMA 1675 may include logic that enables or configures NAMA to perform various operations, including (but not limited to) the methods, operations, and / or other functionalities described above. In some embodiments, this logic may be similar to logic 1665 described above. For example, NAMA 1675 may include a processor and software or firmware instructions similar to instructions 1660, which can be executed on the internal processor of NAMA and / or the processor 1610 of the device, and / or NAMA 1675 may not include its own processor and / or may include a hardware circuitry similar to circuitry 1625 that enables NAMA 1675 to operate independently of any processor and / or in conjunction with the processor within NAMA 1675 and / or other processor operations outside NAMA 1675, including (but not limited to) the processor 1610 of device 1600.
[0177] In some embodiments, memory window 1675 may encompass the entire system memory 1620, in which case, for example, execution of OS 1650 and / or instructions 1660 may be stored within memory window 1670. In other embodiments, memory window 1670 may be separate from the memory region storing execution of OS 1650 and / or instructions. In some embodiments, multiple memory windows 1675 may comprise different portions of system memory 1620.
[0178] In some embodiments, bus 1605 may include multiple buses. One such bus may be an interconnect and / or front-side bus, such as interconnect 135 described above. In some embodiments, all communication between processor 1610 and memory controller 1605 may flow via NAMA 170. In other embodiments, the processor may have a separate communication route (e.g., directly via bus 1605).
[0179] As mentioned above, the communication interface of device 1635 may include, for example, a NIC, an internal or external switch, and / or the like, and / or be able to communicate with it. Therefore, in some embodiments, communication interface 1635 may be used as a network accessory for NAMA 1675, for example, by providing connectivity to a packet network. As mentioned above, NAMA 1675 may communicate directly with such network accessories and / or may communicate with accessories via bus 1605 and / or a portion thereof (e.g., an interconnect).
[0180] The NAMA 1675, memory controller 1665, and / or system memory 1665 may share a housing with other components of the device 1600, and / or one or more of the components of the device 1600 (including (but not limited to) the NAMA 1675, memory controller 1665, and / or system memory 1665) may be physically separable from other components of the device 1600.
[0181] Additional instances
[0182] Some exemplary embodiments are described below. As those skilled in the art will understand, each of the described embodiments can be implemented individually or in any combination. Therefore, a single embodiment or combination of embodiments should not be considered limiting.
[0183] According to some embodiments, a network-aware memory agent includes a memory interface configured to communicate with a first system memory including random access memory (RAM), and a communication interface configured to communicate with at least one central processing unit (CPU) and packet networking equipment. In some embodiments, the communication interface includes an interconnect interface providing communication with the CPU and a network interface providing communication with the packet networking equipment. In some cases, the interconnect includes a front-side bus of the CPU.
[0184] The network-aware memory agent may further include logic for performing one or more operations. In a particular embodiment, the logic is implemented as a hardware circuit system. In other embodiments, some of the logic may be implemented as firmware or software executable by a processor.
[0185] In some embodiments, the logic includes logic for identifying a first memory window including the first system memory, logic for communicating with a local CPU via an interconnect, and logic for receiving a first memory request. In some embodiments, the logic may further include logic for determining that the first memory request addresses a second memory window that does not include the first system memory; logic for identifying a target memory agent associated with the second memory window based on a memory window identifier in the first memory request; logic for generating a first enhanced memory request (EMR) agent-to-agent request (a2aReq) message embedding corresponding to the first memory request; logic for identifying packet header information associated with the target memory agent; and / or logic for packetizing the first EMRa2aReq message using the packet header information to generate one or more data packets; and logic for transmitting one or more data packets via a packet networking device for reception by the target memory agent.
[0186] In some embodiments, the first memory request is received from the local CPU via an interconnect as a Message Passing Interface (MPI) message; in such embodiments, the network-aware memory agent may further include logic for generating a first EMRaReq message from the MPI message.
[0187] In some embodiments, the network-aware memory agent further includes logic for receiving an in-network computation request; logic for performing an in-network computation function in response to receiving the in-network computation request; and logic for transmitting a response to the in-network computation request. In one aspect of some embodiments, the in-network computation request includes an accumulation request containing a first set of data; in such embodiments, the logic for performing the in-network computation function may include logic for performing an operation using the first set of data.
[0188] In some embodiments, the network-aware memory agent further includes logic for receiving a packetized response message from a target memory agent via a packet network through a packet network apparatus, logic for generating an MPI response message based at least in part on the packetized response message, and / or logic for transmitting the MPI response via an interconnect for reception by the CPU.
[0189] In some embodiments, the network-aware memory agent further includes logic for receiving a second memory request addressed to a first memory window; the second memory request may be a fence message from a request source. In some embodiments, the network-aware memory agent may therefore further include logic for setting a fence flag for the first memory window; logic for receiving a plurality of subsequent memory requests directed to the first memory window from a request source; logic for incrementing a counter upon receiving each of the plurality of subsequent memory requests; logic for executing each of the plurality of subsequent memory requests directed to the first memory request; logic for decrementing a counter upon executing each of the plurality of subsequent memory requests; and / or logic for clearing the fence flag when it is determined that the counter has decremented to zero; and logic for transmitting a completion response message when it is determined that the counter has decremented to zero.
[0190] In some embodiments, the network-aware memory agent further includes logic for receiving a second memory request addressed to a second memory window; logic for identifying a target memory agent associated with the second memory window; logic for determining a communication route for the target memory agent; and / or logic for generating a second a2aReq message encapsulating the second memory request. The network-aware memory agent may also include logic for transmitting the second EMR message via an interconnect if the communication route indicates that the target memory agent's communication is routed via an interconnect, and / or logic for transmitting the second EMR message as one or more packets via a packet network if the communication route indicates that the target memory agent's communication is routed via a packet network.
[0191] In some embodiments, the network-aware memory agent further includes logic for receiving a second memory request from a peer memory agent via at least one communication interface; logic for determining that the second memory request addresses a first memory window; logic for executing the second memory request on a first system memory based on determining that the second memory request addresses the first memory window; and logic for transmitting a response to be received by the peer memory agent.
[0192] In some embodiments, the memory interface is coupled to a memory controller, and the logic for executing the first memory request includes logic for communicating the first memory request to the memory controller. In some embodiments, the network-aware memory agent further includes a memory controller coupled to the first system memory, and the logic for executing the first memory request on the first system memory includes logic for causing the included memory controller to execute the first memory request on the first system memory.
[0193] In some embodiments, the second memory request includes a write command, and / or the response includes a completion message indicating that data has been written to the first system memory. In some embodiments, the second memory request includes a read command, and / or the response includes data read from the first system memory.
[0194] In some embodiments, the second memory request includes a compute fast link (CXL) message received via an interconnect.
[0195] In some embodiments, the second memory request includes one or more request Internet Protocol (IP) packets received via a packet network through a packet network device, and one or more response packets include one or more response IP packets. In some embodiments, the network-aware memory agent further includes logic for receiving explicit congestion notification (ECN) information; and logic for determining the timing of the transmission of one or more IP packets (e.g., one or more response IP packets) based at least in part on the ECN information to reduce the risk of communication loss.
[0196] In some embodiments, the network-aware memory agent further includes logic for maintaining at least one request queue and / or logic for storing a plurality of received requests in at least one queue. In some embodiments, the network-aware memory agent further includes logic for maintaining a source context that associates each message in the queue with the source of the message, and / or logic for maintaining a window context that associates each message in the queue with the memory window to which the request is directed. In some embodiments, the network-aware memory agent further includes logic for processing a plurality of received requests in the queue according to the source context and the window context.
[0197] According to some embodiments, a network-aware memory controller includes a memory controller. The memory controller may include a memory interface configured to communicate with a first system memory including random access memory (RAM). In some embodiments, the network-aware memory controller further includes: an interconnect interface configured to communicate with an interconnect providing communication with a first central processing unit (CPU); and a packet network interface configured to communicate with packet network equipment. In some embodiments, the network-aware memory controller further includes logic for transmitting and receiving memory requests via a packet network.
[0198] According to some embodiments, the method may include performing any of the operations described above. In some embodiments, such a method may be performed by a network-aware memory agent, which may include logic for causing the network-aware memory agent to perform such operations, such as as described above. By way of example only, one such method may include, for example, identifying a first memory window including a first system memory via a network-aware memory agent; the first system memory may include random access memory (RAM). In some embodiments, the network-aware memory agent includes a memory interface configured to communicate with the first system memory including the random access memory (RAM), and at least one communication interface configured to communicate with a first central processing unit (CPU) and packet networking equipment.
[0199] In some embodiments, the method includes, for example, communicating with a local CPU via an interconnect through a network-aware memory agent; for example, receiving a first memory request through a network-aware memory agent; for example, determining, through a network-aware memory agent, that the first memory request addresses a second memory window that does not include the first system memory; and / or, for example, identifying a target memory agent associated with the second memory window through a network-aware memory agent and based on a memory window identifier in the first memory request.
[0200] In some embodiments, the method includes, for example, generating a first enhanced memory request (EMR) a2aReq message embedding a first memory request via a network-aware memory agent; for example, identifying packet header information associated with a target memory agent via a network-aware memory agent; for example, packetizing the a2aReq message using the packet header information via a network-aware memory agent to generate one or more data packets; and / or, for example, transmitting one or more data packets via a network-aware memory agent for reception by the target memory agent via a packet networking device.
[0201] in conclusion
[0202] In the foregoing description, numerous details have been set forth for illustrative purposes to provide a thorough understanding of the described embodiments. However, it will be apparent to those skilled in the art that other embodiments may be practiced without some of these details. In other instances, structures and apparatuses are shown as block diagrams for clarity, but without complete detail. Several embodiments are described herein, and while various features are attributed to different embodiments, it should be understood that features described with respect to one embodiment may also be incorporated into other embodiments. However, for the same reason, no single feature or features of any described embodiment should be considered essential to every embodiment of the invention, as such features may be omitted in other embodiments of the invention.
[0203] Therefore, the foregoing description provides an illustration and description of some features and aspects of various embodiments, but is not intended to be exhaustive or to limit the embodiments as a whole to the precise forms disclosed. Those skilled in the art will recognize that modifications can be made in light of the foregoing disclosure, or modifications can be obtained from the practice of the embodiments, all of which fall within the scope of the various embodiments. For example, as mentioned above, the methods and processes described herein can be implemented using software components, firmware and / or hardware components (including (but not limited to) processors, other hardware circuit systems, custom integrated circuits (ICs), programmable logic, etc.) and / or any combination thereof.
[0204] Furthermore, although the various methods and processes described herein may be described with respect to specific structures and / or functional components for ease of description, the methods provided by the various embodiments are not limited to any particular structure and / or functional architecture, but can be implemented in any suitable hardware configuration. Similarly, although some functionality is attributed to one or more system components, unless the context otherwise indicates, this functionality may be distributed across a variety of other system components according to several embodiments.
[0205] Similarly, although the procedures of the methods and processes described herein are presented in a specific order for ease of description, various procedures may be reordered, added, and / or omitted according to various embodiments unless the context otherwise indicates. Furthermore, procedures described with respect to a method or process may be incorporated into other described methods or processes; similarly, system components described with respect to a particular architecture and / or system may be organized into alternative architectures and / or incorporated into other described systems. Therefore, although various embodiments may or may not be described with certain features for ease of description and illustration of aspects of these embodiments, various components and / or features described herein with respect to particular embodiments may be replaced, added, and / or subtracted from other described embodiments unless the context otherwise indicates.
[0206] As used herein, the term "component" is intended to be broadly interpreted as hardware, firmware, software, or a combination of any of these. It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, and / or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit any embodiment, unless expressly stated in the appended claims. Therefore, when the operation and behavior of systems and / or methods are described herein without reference to specific software code, those skilled in the art will understand that the software and hardware can be used to implement the systems and / or methods based on the description herein.
[0207] In this disclosure, when an element is referred to herein as “connected” or “coupled” to another element, it should be understood that an element may be directly connected to another element, or that an intervening element may exist between the elements. Conversely, when an element is referred to as “directly connected” or “directly coupled” to another element, it should be understood that there is no intervening element in the “direct” connection between the elements. However, the presence of a direct connection does not preclude the existence of other connections where an intervening element may exist. Similarly, although the methods and processes described herein may be described in a specific order for ease of description, it should be understood that, unless the context otherwise indicates, the intervention process may occur before and / or after any part of the described process, and as mentioned above, the described procedures may be reordered, added, and / or omitted according to various embodiments.
[0208] In this application, unless otherwise expressly stated, the use of the singular form includes the plural form, and unless otherwise indicated, the use of the term "and" means "and / or". Furthermore, as used herein, the term "or" is intended to be inclusive when used in a series, and may be used interchangeably with "and / or" unless otherwise expressly stated (e.g., if used in combination with "any" or "only one"). Additionally, the use of the term "including" and other forms (e.g., "includes / included") should be considered non-exclusive. Furthermore, unless otherwise expressly stated, terms such as "element" or "component" cover both elements and components comprising one unit and elements and components comprising more than one unit. As used herein, the phrase "at least one" preceding a series of items (separating "any" from the items with the terms "and" or "or") modifies the entire list, not each member of the list (i.e., each item). The phrase "at least one" does not require the selection of at least one of each of the listed items; rather, the phrase allows for the inclusion of at least one of any one of the items, and / or at least one of any combination of items. For example, the phrases "at least one of A, B, and C" or "at least one of A, B, or C" each refer to only A, only B, or only C; and / or any combination of A, B, and C. In examples where the desired selection is "at least one of each of A, B, and C" or alternatively "at least one of A, at least one of B, and at least one of C," it is explicitly described as such.
[0209] Unless otherwise indicated, all figures used herein to express quantity, size, etc., should be understood to be modified by the term “about” in all instances. As used herein, the article “a / an” is intended to include one or more items and is used interchangeably with “one or more”. Similarly, as used herein, the article “described” is intended to include one or more items referenced in conjunction with the article “described” and is used interchangeably with “one or more”. As used herein, the term “group” is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, and / or similar) and is used interchangeably with “one or more”. Where only one item is desired, the phrase “only one” or similar language is used. As used herein, the terms “has / have / having” are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the phrase “based on” is intended to mean “at least partially based on”. In the foregoing description, depending on the context, satisfying a threshold can refer to a value greater than a threshold, greater than or equal to a threshold, less than a threshold, less than or equal to a threshold, equal to a threshold, and / or similar.
[0210] Although specific combinations of features are stated in the claims and / or disclosed in the description, these combinations are not intended to limit the disclosure of various embodiments. In fact, many of these features can be combined in ways not expressly stated in the claims and / or disclosed in the description. Therefore, while each dependent claim listed below may be directly attached to only one claim, the disclosure of various embodiments includes a combination of each dependent claim with every other claim in the claim set. No element, action, or instruction used herein should be construed as critical or essential unless expressly stated otherwise.
Claims
1. A network-aware memory agent, comprising: A memory interface configured to communicate with a first system memory including random access memory (RAM); At least one communication interface configured to communicate with: First Central Processing Unit (CPU); and Packet network equipment; Logic for identifying the first memory window including the first system memory; Logic used for communication with the local CPU via the interconnect; Logic for receiving the first memory request; Logic used to determine that the second memory window not included in the first system memory is being addressed by the first memory request; Logic for identifying the target memory agent associated with the second memory window based on the memory window identifier in the first memory request; Logic for generating a first enhanced memory request EMR proxy to proxy request a2aReq message corresponding to the first memory request; Logic used to identify packet header information associated with the target memory agent; and Logic for grouping the first EMRa2aReq message using the packet header information to generate one or more data packets; and Logic for transmitting the one or more data packets via the packet network equipment for reception by the target memory agent.
2. The network-aware memory agent according to claim 1, wherein: The first memory request is received from the local CPU via the interconnect as a message passing interface (MPI) message; and The network-aware memory agent further includes: Logic for generating the first EMRa2aReq message from the MPI message.
3. The network-aware memory agent according to claim 2, further comprising: Logic used to receive computation requests within the network; Logic for performing intra-network computing functions in response to receiving the intra-network computing request; and Logic used to transmit responses to computation requests within the network.
4. The network-aware memory agent according to claim 3, wherein: The network computation request includes a cumulative request containing the first set of data; and The logic for performing network-in-network computation functions includes: The logic used to perform operations with the first set of data.
5. The network-aware memory agent according to claim 2, further comprising: Logic for receiving packetized response messages from the target memory agent via the packet network through the packet network via the packet network equipment; Logic for generating an MPI response message based at least in part on the grouped response message; and Logic for transmitting the MPI response message via the interconnect for reception by the local CPU.
6. The network-aware memory agent according to claim 1, wherein the at least one communication interface is a plurality of separate interfaces, the plurality of separate interfaces comprising: Provides an interconnect interface for communication with the local CPU; and Provides a network interface for communication with the packet network equipment.
7. The network-aware memory agent of claim 6, wherein the interconnect includes the front-side bus of the local CPU.
8. The network-aware memory agent according to claim 1, wherein: The network-aware memory agent further includes: Logic for receiving a second memory request addressed to the first memory window, the second memory request including a fence message from the request source; Logic for setting a fence flag for the first memory window; Logic for receiving multiple subsequent memory requests directed from the request source to the first memory window; Logic for incrementing a counter upon receiving each of the plurality of subsequent memory requests; Logic for executing each of the plurality of subsequent memory requests that are directed to the first memory request; Logic for decrementing the counter as each of the plurality of subsequent memory requests is executed; and Logic for clearing the fence flag when it is determined that the counter has decremented to zero; and The logic for transmitting a completion response message (CRM) when it is determined that the counter has decremented to zero.
9. The network-aware memory agent according to claim 1, further comprising: Logic for receiving second memory requests addressed to the second memory window; Logic used to identify the target memory agent associated with the second memory window; Logic used to determine the communication route for the target memory proxy; Logic for generating a second a2aReq message that encapsulates the second memory request; Logic for transmitting the second a2aReq message via the interconnect when the communication route indicates that the communication of the target memory agent is via the interconnect route; and Logic for transmitting the second a2aReq message as one or more packets via the packet network when the communication routing indicates that the communication of the target memory agent is routed via the packet network.
10. The network-aware memory agent according to claim 1, further comprising: Logic for receiving a second memory request from a peer memory agent via the at least one communication interface; Logic used to determine if the second memory requests to address the first memory window; Logic for executing the second memory request on the first system memory based on determining that the second memory request addresses the first memory window; and Logic used to transmit responses to be received by the peer memory agent.
11. The network-aware memory agent according to claim 10, wherein: The second memory request includes one or more request Internet Protocol (IP) packets received via the packet network through the packet network equipment; and The one or more response packets include one or more response IP packets.
12. The network-aware memory agent of claim 11, further comprising: Logic used to receive explicit congestion notification (ECN) information; and Logic for determining the timing of the transmission of one or more response IP packets, at least in part, based on the ECN information, to reduce the risk of communication loss.
13. The network-aware memory agent according to claim 10, wherein: The second memory request includes a compute fast link (CXL) message received via the interconnect.
14. The network-aware memory agent according to claim 10, wherein: The memory interface is coupled to the memory controller, and The logic for executing the first memory request includes: Logic for communicating the first memory request to the memory controller.
15. The network-aware memory agent according to claim 10, wherein: The network-aware memory agent further includes a memory controller; The memory controller is coupled to the first system memory; and The logic for executing the first memory request on the first system memory includes logic for causing the included memory controller to execute the first memory request on the first system memory.
16. The network-aware memory agent according to claim 10, wherein: The second memory request includes a write command; and The response includes a completion message indicating that data has been written to the first system memory.
17. The network-aware memory agent according to claim 10, wherein: The second memory request includes a read command; and The response includes data read from the first system memory.
18. The network-aware memory agent according to claim 1, further comprising: Logic for maintaining at least one request queue; Logic for storing multiple received requests in the at least one queue; Logic for maintaining a source context, which associates each message in the queue with the source of the message; Logic for maintaining a window context, which associates each message in the queue with the memory window to which the request is directed; and Logic for processing the plurality of received requests in the queue based on the source context and the window context.
19. A network-aware memory controller, comprising: The memory controller includes: A memory interface configured to communicate with a first system memory including random access memory (RAM); An interconnect interface configured to communicate with an interconnect that provides communication with a first central processing unit (CPU); A packet network interface configured to communicate with a packet network; and Logic for transmitting and receiving memory requests via the packet network.
20. A method comprising: The network-aware memory agent identifies a first memory window including a first system memory, wherein the first system memory includes random access memory (RAM). A memory interface configured to communicate with a first system memory including random access memory (RAM); At least one communication interface configured to communicate with: First Central Processing Unit (CPU); and Packet network equipment; The network-aware memory agent communicates with the local CPU via the interconnect; The network-aware memory agent receives the first memory request. The network-aware memory agent determines that the second memory window for the first memory request address does not include the first system memory; The target memory agent associated with the second memory window is identified by the network-aware memory agent and based on the memory window identifier in the first memory request. The network-aware memory agent generates a first enhanced memory request EMRa2aReq message that embeds the first memory request. The network-aware memory agent identifies packet header information associated with the target memory agent; The network-aware memory agent uses the packet header information to packetize the a2aReq message to generate one or more data packets; and The one or more data packets are transmitted through the network-aware memory agent for reception by the target memory agent via the packet network equipment.