Adaptive Memory Consistency in Disaggregated Data Centers

By introducing an adaptive consistency controller (ACC) into the FAM system, dynamically manages the memory consistency model, solving the challenges of memory consistency management in the FAM system, achieving efficient data consistency and performance improvements.

CN117120976BActive Publication Date: 2025-05-30ADVANCED MICRO DEVICES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280026190.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-31
Filing Date
2022-03-16
Publication Date
2025-05-30
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

In structurally attached memory (FAM) systems, memory consistency management challenges exist, especially in multiprocessor or multithreaded systems, how to effectively sort memory instructions to ensure data consistency is a difficult problem.

Method used

Adaptive consistency controller (ACC) is used to manage the memory consistency model. By collaborating with the structure manager, the requester's access rights changes to the memory area are monitored, and appropriate memory consistency models are set according to different permissions, and the corresponding fence mechanism is activated to ensure data consistency.

Benefits of technology

It realizes dynamic management of memory consistency in the FAM system, adjusts the consistency model in real time according to the number and permissions of the requester, and improves the system's data consistency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117120976B_ABST
    Figure CN117120976B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processor, which includes a fabric-attached memory (FAM) interface for coupling to a data structure and executing memory access instructions. A requester-side adaptive coherence controller coupled to the FAM interface requests from a fabric manager of the fabric-attached memory a notification of a change of a requester authorized to access a FAM region to which the data processor is authorized to access. If the notification indicates that more than one requester is authorized to access the FAM region, a fence is activated for a selected memory access instruction in a local application program.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Emerging fabric standards such as Compute Express Link (CXL) 3.0, Gen-Z, or Slingshot exemplify methods for data center disaggregation, where a central processing unit (CPU) host can access fabric-attached memory (FAM) modules. Such modules contain memory attached to the data center fabric but have no or little computing power associated with them. With FAM, the host is not constrained by the memory capacity limitations of its local server. Instead, the host gains access to a large pool of memory that does not need to be attached to any particular host. FAM is partitioned among hosts, and the partition can be dedicated to one host or shared among multiple hosts.

[0002] Memory consistency is an important consideration in the process of developing software applications for use with FAM systems. Consistency defines how memory instructions (to different memory locations) in a multiprocessor or multithreaded system will be ordered and is enforced by reordering independent memory operations according to a consistency model.

[0003] Various consistency models have been developed that impose various ordering constraints on independent memory operations in the instruction streams of individual processors involving high-level dependencies. In a simple consistency model called sequential consistency, a processor is not allowed to reorder reads and writes. Another model called "Total Store Ordering" (TSO) allows a store buffer. In this scheme, the store buffer holds store operations that need to be sent to memory until a specified condition is met and can send a group of operations to memory. Loads are allowed to pass the stored contents, but the stored contents are sent to memory in program order. The address of a load operation is checked against the addresses in the store buffer, and if there is an address match, the store buffer is used to satisfy the load operation.

[0004] Other consistency models, called relaxed or weak consistency models, rely on a certain version of fence (or barrier) operations that demarcate regions where reordering of operations is permissible. Release consistency is an example of a weak consistency model where synchronous access is divided into "acquire" and "release", where in "acquire", operations like locking must be completed before all subsequent memory accesses, and in "release", operations like unlocking must be completed with all memory operations before the release is complete. Brief Description of the Drawings

[0005] Figure 1 A fabric-attached memory (FAM) system according to the prior art is shown in block diagram form;

[0006] Figure 2 A block diagram illustrates an Accelerated Processing Unit (APU) according to some embodiments, which is suitable for use as a computing unit in a FAM system such as Figure 1 ;

[0007] Figure 3 A block diagram illustrates certain elements of a FAM memory system including an Adaptive Coherence Controller (ACC) according to some embodiments;

[0008] Figure 4 A diagram illustrates a process of operating Figure 3 the ACC according to some embodiments;

[0009] Figure 5 A flowchart illustrates a process for managing a memory coherence model at an Adaptive Coherence Controller according to an exemplary embodiment;

[0010] Figure 6 A flowchart illustrates a process for implementing a memory coherence model according to some embodiments; and

[0011] Figure 7 A flowchart illustrates another process for implementing a memory coherence model according to some additional embodiments.

[0012] In the following description, the same reference numerals are used in different figures to indicate similar or identical items. Unless otherwise specified, the word "coupled" and its associated verb forms include both direct connection and indirect electrical connection by means known in the art, and unless otherwise specified, any description of direct connection also implies an alternative embodiment using a suitable form of indirect electrical connection. Detailed Description

[0013] A method for use with a fabric-attached memory system that includes a fabric-attached memory and a plurality of requesters coupled to the fabric-attached memory through the fabric. Request a notification from a fabric manager about changes to the requesters authorized to access a fabric-attached memory region. In response to a notification from the fabric manager indicating that more than one requester is authorized to access a fabric-attached memory region, for each such authorized requester, activate a fence for selected memory access instructions in a local application related to the fabric-attached memory region.

[0014] The data processor includes a processing core, a fabric-attached memory interface, and a requester-side adaptive coherence controller. The processing core executes an application program. The fabric-attached memory interface is coupled to the processor core and is adapted to connect to a data structure and execute memory access instructions from the processing core to the fabric-attached memory. The requester-side adaptive coherence controller is coupled to the processing core and the fabric-attached memory interface and requests from the fabric manager of the fabric-attached memory a notification of changes to the requester authorized to access the fabric-attached memory region to which the data processor is authorized access. In response to a notification from the fabric manager indicating that more than one requester is authorized to access the fabric-attached memory region, the requester-side adaptive coherence controller causes a fence to be activated for selected memory access instructions in the local application program.

[0015] The fabric-attached memory system includes a fabric-attached memory, a data structure, a fabric manager, and a plurality of data processors. The data structure is connected to the fabric-attached memory. The fabric manager is connected to the data structure and is operable to authorize and de-authorize requesters to access memory regions of the fabric-attached memory. The plurality of data processors are connected to the data structure, and each data processor includes a processing core that executes an application program, a fabric-attached memory interface, and a requester-side adaptive coherence controller coupled to the processing core and the fabric-attached memory interface. The requester-side adaptive coherence controller requests from the fabric manager a notification of changes to the requester authorized to access the fabric-attached memory region to which the data processor is authorized access. In response to a notification from the fabric manager indicating that more than one requester is authorized to access the fabric-attached memory region, the requester-side adaptive coherence controller causes a fence to be activated for selected memory access instructions in the local application program.

[0016] Figure 1 A fabric-attached memory (FAM) system 100 according to the prior art is shown in block diagram form. The depicted FAM system 100 is merely one example of a data structure topology among many topologies commonly used to disaggregate data centers. The FAM system 100 generally includes a data center fabric 102 and a plurality of device groups 104 called pods.

[0017] Each cluster 104 includes a plurality of compute nodes "C", a plurality of memory nodes "M", and an interconnect network "ICN". The compute nodes C are connected to the ICN via routers "R". The compute nodes C include multiple CPUs (each with multiple cores) or multiple accelerated processing units (APUs) that are part of the same coherence domain. Each compute node C includes a fabric bridge, such as a network interface card (NIC), a CXL interface, or other suitable fabric interfaces that serve as a gateway for the compute node C into the data center fabric 102. The memory nodes M are connected to the ICN via routers R. Each memory node M includes a similar fabric interface and a media controller that satisfies requests for the FAM. The ICN includes switches for interconnecting the respective compute nodes C with the memory nodes M and may include routers in some topologies.

[0018] The depicted topology includes a local data center fabric formed by routers R and the ICN, and a global data center fabric labeled as data center fabric 102. In this embodiment, the local data center fabric is within a rack, and the global data center fabric includes multiple racks. However, various fabric topologies may be implemented within a rack or within a data center and may include compute nodes that remotely access the data center via a network. Note that many topologies have a compute node C that also includes memory as part of a FAM pool. This memory can be mapped as fabric-attached memory and made available for use by other compute nodes according to a resource allocation process known as "composability".

[0019] The data center fabric 102 provides data interconnectivity between clusters 104, including switches and routers that couple data traffic in a protocol such as CXL, Gen-Z, or other suitable memory fabric protocols. Note that multiple protocols may be employed together in the data center fabric. In this exemplary embodiment, CXL is used to interconnect devices within a rack, while Gen-Z is used to interconnect various racks within the data center.

[0020] Figure 2 The APU 200 is shown in block diagram form and is suitable for use as a compute unit C in a FAM system of a FAM system 100 such as Figure 1 The APU 200 is an integrated circuit suitable for use as a processor in a host data processing system and generally includes a central processing unit (CPU) core complex 210, a graphics core 220, a set of display engines 222, a data structure 225, a memory management hub 240, a set of peripheral controllers 260, a set of peripheral bus controllers 270, and a system management unit (SMU) 280, as well as a set of memory interfaces 290.

[0021] The CPU core complex 210 includes CPU cores 212 and 214. In this example, the CPU core complex 210 includes two CPU cores, but in other embodiments, the CPU core complex 210 may include any number of CPU cores. Each of the CPU cores 212 and 214 is bidirectionally connected to a system management network (SMN) (which forms a control structure) and a local data structure 225, and is capable of providing memory access requests to the data structure 225. Each of the CPU cores 212 and 214 may be a monolithic core, or may further be a core complex of two or more monolithic cores that share certain resources such as caches. Each of the CPU cores 212 and 214 includes μCode 216, which runs to execute certain instructions on the CPU, including performing certain functions for memory consistency on the data center fabric as further described below.

[0022] The graphics core 220 is a high-performance graphics processing unit (GPU) that is capable of performing graphics operations such as vertex processing, fragment processing, shading, texture blending, etc. in a highly integrated and parallel manner. The graphics core 220 is bidirectionally connected to the SMN and the data structure 225, and is capable of providing memory access requests to the data structure 225. In this regard, the APU 200 may support a unified memory architecture in which the CPU core complex 210 and the graphics core 220 share the same storage space, or a memory architecture in which the CPU core complex 210 and the graphics core 220 share a portion of the storage space, while the graphics core 220 also uses a private graphics memory that is not accessible to the CPU core complex 210. Memory regions may be allocated from local memory or the data center fabric.

[0023] The display engine 222 renders and rasterizes the objects generated by the graphics core 220 for display on a monitor. The graphics core 220 and the display engine 222 are bidirectionally connected to a common memory management hub 240 for unified translation to the appropriate address in system memory.

[0024] The local data structure 250 includes a crossbar switch for routing memory access requests and memory responses between any memory access agent and the memory management hub 240. The data structure also includes a system memory map defined by the basic input / output system (BIOS) for determining the destination of memory accesses based on the system configuration, and buffers for each virtual connection.

[0025] The peripheral controller 260 includes a Universal Serial Bus (USB) controller 262 and a Serial Advanced Technology Attachment (SATA) interface controller 264, each of which is bi-directionally connected to the system hub 266 and the SMN bus. These two controllers are merely examples of peripheral controllers available for the APU 200.

[0026] The peripheral bus controller 270 includes a system controller or “South Bridge” (SB) 272 and a Peripheral Component Interconnect Express (PCIe) controller 274, each of which is bi-directionally connected to the Input / Output (I / O) hub 276 and the SMN bus. The I / O hub 276 is also bi-directionally connected to the system hub 266 and the data structure 225. Thus, for example, the CPU cores can program the registers in the USB controller 262, the SATA interface controller 264, the SB 272, or the PCIe controller 274 through access routed by the I / O hub 276 through the data structure 225. The software and firmware of the APU 200 are stored in a system data drive or a system BIOS memory (not shown), which can be any of a variety of non-volatile memory types, such as read-only memory (ROM), flash electrically erasable programmable ROM (EEPROM), etc. Generally, the BIOS memory is accessed through the PCIe bus, and the system data drive is accessed through the SATA interface.

[0027] The SMU 280 is a local controller that controls the operation of the resources on the APU 200 and synchronizes the communication between these resources. The SMU 280 manages the power-on sequencing of the various processors on the APU 200 and controls multiple off-chip devices via reset, enable, and other signals. The SMU 280 includes one or more clock sources (not shown), such as a Phase-Locked Loop (PLL), to provide clock signals for each component of the APU 200. The SMU 280 also manages the power of the various processors and other functional blocks, and can receive measured power consumption values from the CPU cores 212 and 214, as well as the graphics core 220, to determine the appropriate power state.

[0028] The Memory Management Hub 240 is connected to the local data structure 250, the graphics core 220, and the display engine 230 to provide direct memory access capabilities to the graphics core 220 and the display engine 230.

[0029] The memory interface 290 includes two memory controllers 291 and 292, DRAM media 293 and 294, and a FAM memory interface 295. Each of the memory controllers 291 and 292 is connected to the local data structure 250 and is connected to a corresponding one of the DRAM media 293 and 294 through a physical layer (PHY) interface. In this embodiment, the DRAM media 293 and 294 include memory modules based on DDR memory such as DDR version five (DDR5). In other embodiments, other types of DRAM memory are used, such as low-power DDR4 (LPDDR4), graphics DDR version five (GDDR5), and high-bandwidth memory (HBM).

[0030] The FAM memory interface 295 includes a fabric bridge 296, an adaptive coherence controller (ACC) 297, and a fabric PHY 298. The fabric bridge 296 is a fabric-attached memory interface connected to the local data structure 250 for receiving and executing memory requests for a FAM system such as the FAM system 100. Such memory requests may come from the CPU core complex 210 or may be direct memory access (DMA) requests from other system components such as the graphics core 220. The fabric bridge 296 is also connected bidirectionally to the fabric PHY 298 to provide a connection of the APU 200 to the data center fabric. The adaptive coherence controller (ACC) 297 is connected bidirectionally to the fabric bridge 296 for providing memory coherence control inputs to the fabric bridge 296 and the CPU core complex 210, as further described below. In operation, the ACC 297 communicates with the CPU cores in the CPU core complex 210 to receive notifications of specified memory access instructions identified by the μCode 216 running on the CPU cores 212 and 214, as further described below. The ACC 297 also provides configuration inputs to the CPU core complex 210 to configure the memory coherence model.

[0031] Figure 3 Certain elements of a FAM system 300 according to some embodiments are shown in block diagram form. The FAM system 300 generally includes a compute node 302, a data center fabric 102, a FAM memory node 310, and a fabric manager 320.

[0032] Compute node 302 is one of many requester compute nodes connected to data center fabric 102 and is typically implemented using an APU such as APU 200. Compute node 302 may implement an Internet server, an application server, a supercomputing node, or another suitable computing node that benefits from access to the FAM. Only the FAM interface components of compute node 302 are depicted to focus on the relevant parts of the system. Compute node 302 includes fabric bridge 296, fabric PHY 298, and ACC 297.

[0033] Fabric bridge 296 is connected to the local data structure as described above in connection with Figure 2 and fabric PHY 298 and ACC 297. ACC 297 in this version includes a microcontroller (μC) 304. μC 304 performs the memory coherence control function described below and is also typically connected to a tangible non-transitory memory for storing firmware to initialize and configure μC 304 to perform its functions.

[0034] μC 304 performs the memory coherence control function described below and is also typically connected to a tangible non-transitory memory for storing firmware to initialize and configure μC 304 to perform its functions.

[0035] Fabric manager 320 is a controller connected to data center fabric 102 for managing the configuration and access to FAM system 300. Fabric manager 320 executes a data structure management application for a particular standard (such as CXL or Gen-Z) employed on data center fabric 102. The data structure management application manages and configures data center fabric functions such as authorizing compute nodes, allocating memory regions, and managing composability by identifying and configuring memory resources among the various nodes in FAM system 300. Note that while one FAM memory node 310 and one compute node 302 are shown, the system includes multiple such nodes, which may appear in many configurations such as the exemplary configuration depicted in Figure 1 In some embodiments, fabric manager 320 has an Adaptive Coherence Controller (ACC) module 322 that is installed to access the data structure management application and report data to each ACC 297 on each of the various compute nodes throughout FAM system 300.

[0036] The FAM memory node 310 includes a media controller 312 and a memory 314. The media controller 312 generally includes a memory controller that is adapted to select any type of memory to be used in the memory 314. For example, if the memory 314 is a DRAM memory, a DRAM memory controller is used. The memory 314 may also include a persistent memory module and a hybrid memory module. In some embodiments, the ACC 297 maintains data in the requester table 306 at the FAM memory node 310 regarding other requesters authorized to access the FAM memory regions allocated to the compute node 302. The requester table 306 tracks updates to the compute nodes authorized to access the same memory region and includes fields for a "timestamp" that reflects the update time, a "region ID" that reflects the identifier of the FAM memory region allocated to the compute node 302, and a "# requesters" that reflects the number of requesters on the FAM system 300 authorized to access the memory region since each update. As described further below, the requester table 306 is updated based on reports from the fabric manager 320. In this embodiment, the FAM memory node 310 includes a buffer that holds the requester table 306 and is accessible by the media controller 312.

[0037] Figure 4 FIG. 400 shows a process for operating the ACC 297 according to some embodiments. FIG. 400 depicts activities at a requester compute node at a data structure and shows the CPU core complex 210, the operating system 402, the ACC 297, and the fabric bridge 296 at the requester compute node. FIG. 400 also shows the fabric manager 320 and the media controller 312, which is one of many media controllers on the data structure.

[0038] When the fabric manager 320 allocates a particular FAM region to a requester for system memory, the ACC 297 makes a callback request to the fabric manager 320 to request notification when there is a change in the number of requesters authorized to use the same memory region as the requester compute node, as shown by the outgoing request labeled "callback". Whenever there is a change in the number of requesters authorized to use the memory region, the fabric manager 320 provides the notification back to the ACC 297, as indicated by the "# users" response on FIG. 400. In this embodiment, the requester table 306 at the FAM memory node 310 is updated whenever the number of users changes ( Figure 3 ) as indicated by the "# users" arrow going to the media controller 312 on FIG. 400. In some embodiments, the ACC 287 maintains the requester table in a buffer local to the ACC 297. In some embodiments, an ACC module 322 running at the fabric manager 320 ( Figure 3)The process of managing monitors for requesters authorized to access a memory region and sending notifications to the ACC 297. Based on the # user update notification, the ACC 297 sets the coherence model to be used with the allocated FAM memory region. Generally, the process includes: setting the coherence model to a first coherence model when the current requester is the only requester authorized to access the memory region, and setting the coherence model to a second coherence model when more than one requester is authorized to access the memory region. Setting the coherence model is shown by the command "Set Model (FAM i ) = (WB,FENCED)", where "FAM i " identifies the structural attached memory region involved, and "WB,FENCED" indicates which coherence model to activate. In some embodiments, the first coherence mode is characterized as being lax relative to the second coherence model. An example of the process is further described below in connection with Figure 5 .

[0039] In FIG. 400, the μCode 216 ( Figure 2 ) running at the CPU core complex 210 helps implement the second coherence model when the second coherence model is active. Specifically, the μCode 216 identifies specified memory access instructions in the application program executed at the CPU core complex 210, which indicate the data fences required for the specified memory instructions regarding the structural attached memory region allocated to the requester node. The μCode 216 has various ways to identify the specified instructions, as further described below in connection with Figure 6 and Figure 7 . When such an instruction is identified, the μCode 216 communicates with the ACC 297 to send a notification identifying the selected instruction, which will be sent to the data structure through the structural bridge 296. The ACC 297 then adds a fence command to the command stream going to the media controller 312, as indicated by the outgoing arrow labeled "Fence". When the first coherence model is active, the μCode 216 will not make such a notification, but will allow the selected command to execute normally with the memory coherence handled by the first coherence model set and the operating system 402. Although in this embodiment, the μCode 216 identifies the specified instruction, in other embodiments, this function is performed by the CPU firmware or a combination of the CPU firmware and the μCode.

[0040] Although this figure shows a fence command going to the media controller 312, in a topology including a local data center structure and a global data center structure, in a scenario where the memory region includes two levels of the access structure topology, the ACC 297 will cause the fence command to be sent to the media controllers on both the local data center architecture and the global data center architecture.

[0041] Figure 5 FIG. 500 is a flow chart of a process for managing a memory coherence model at an adaptive coherence controller in accordance with an exemplary embodiment. The process begins at block 502, where a requester node on a data structure is authorized to access a specified FAM memory region. This authorization is typically provided by a fabric manager, but in some embodiments may be configured by other system components. Based on this authorization, a memory region is established in the addressable system memory space of the requester node. The memory region is typically for an application running at the requester node, which may have dependencies with other requester nodes.

[0042] At block 504, the ACC 297 at the requester node requests from the fabric manager a notification of any change in the number of requesters authorized to access a particular memory region. In one embodiment, this request has the form of a callback request to the fabric manager to track the number of compute requesters accessing the FAM region. The ACC module 322 ( Figure 3 ) may be employed to receive such requests and implement the request at the fabric manager or configure the fabric manager to implement the request. In other embodiments, the fabric manager may have this capability as part of the fabric manager application and no additional module is required.

[0043] At block 506, the ACC 297 at the requester node receives a notification from the fabric manager in response to the request. Based on this notification, the ACC 297 determines the number of requesters currently authorized to use the FAM region and updates the requester table 306 ( Figure 3 ). In some embodiments, the notification includes data fields used in the requester table 306, including a timestamp, a region ID, and # requesters. In other embodiments, the fabric manager may not provide all the data, but only provide data indicating that a requester authorization has been added or removed for the FAM region. In this case, the ACC 297 will update the data in the requester table 306 based on the current notification data and the previous update to the requester table 306.

[0044] At block 508, if more than one requester is authorized for the memory region, the process proceeds to block 510. If this is not the case, the process proceeds to block 512. At block 510, the process causes a fence to be activated for selected memory access instructions in the local application associated with the FAM region. If a transition is made at block 508 from only one requester being authorized to more than one requester, the process includes deactivating a first memory coherence model and activating a second memory coherence model at the requester node, as described above in connection with Figure 4As described. If more than one requester has been authorized at block 508, then the second memory consistency model is already active and no change is needed.

[0045] At block 512, the update notification received at block 506 has resulted in only one requester being authorized for the state of the FAM region, and thus the process activates the first memory consistency model. In some embodiments, the first memory consistency model includes mapping the FAM region as a write-back memory for a single requester. Mapping the FAM region as a write-back memory for the local compute node is preferably accomplished by sending an appropriate message via the ACC 297 to the operating system running on the CPU core complex 210, which then marks one or more memory pages corresponding to the FAM region as write-back in the requester's page table. Mapping the FAM region as a write-back memory ensures that normal local consistency schemes, such as the x86 Total Store Ordering (TSO) consistency scheme, will be applied by the compute node to the FAM region. When only one requester is authorized for the FAM region, no change to the application's functionality is needed to use the FAM instead of local memory. The process of blocks 506, 508, 510, and 512 is repeated whenever the fabric manager sends an update notification regarding the FAM memory region.

[0046] Figure 6 FIG. 600 is a flow chart showing a process for implementing a memory consistency model according to some embodiments. Generally, for applications that were not originally written for a FAM system, if they have any dependencies between compute nodes, they need to be modified so that the FAM system appropriately accounts for the dependencies between compute nodes. Such modifications can be made at compile time or can be applied post hoc by adding instructions to the selected memory commands in the application. The Figure 6 process is performed for applications that have had compile-time instructions added so that the applications are adapted to work with the FAM system. Such applications include applications written and compiled using tools for handling consistency in the FAM system, and applications that have been modified using appropriate tools to be adapted to work with the FAM system.

[0047] At block 602, the process activates a fence for an application that includes compile-time directives in the corresponding application. Block 602 occurs when the ACC 297 changes the coherence model of the compute node to activate a more relaxed second coherence model. Generally, data structure protocols such as CXL 3.0 and Gen-Z employ a relaxed ordering coherence model implemented via fences. When the relaxed coherence model is active, the ACC 297 seeks to insert datacenter structure fences such as CXL / Gen-Z fences at critical locations transparent to the application code. This transparency means that the activities performed by the ACC 297 should not require any adjustment by the applications involved.

[0048] At block 604, the process identifies the compile-time structure interface instructions for the selected memory access instructions. Generally, the selected memory access instructions are part of the code that performs parallel synchronization between compute threads, such as flags, locks, semaphores, control variables, etc., and thus require ordering of the memory accesses. The location of the selected memory access instructions is identified using hints in the application, which can be provided by the software developer to the adaptive controller via several mechanisms.

[0049] Compiler hints (such as C++11 atomic_load / atomic_store coherence constructs or primitives) or runtime hints (such as OpenMP's flush construct) can be integrated with development tools (such as the "CodeAnalyst" tool provided by Advanced Micro Devices, Santa Clara, California) to make coherence hints more accessible to developers. The compiler inserts requester-side structure interface tags, such as a special "FABRIC_ACQUIRE_FENCE" instruction, before the marked control variable and a special "FABRIC_RELEASE_FENCE" instruction after the control variable. These special instructions or tags are converted to no-ops (NOPs) by the CPU's μCode 216 on non-FAM systems or on FAM systems where only one requester node accesses the relevant FAM region, but are recognized by μCode 216 at block 604 when more than one application is authorized to access the relevant FAM region.

[0050] In response to identifying such instructions, at block 606, as Figure 4As indicated by the selected instruction arrow in , CPU μCode 216 notifies ACC 297 that the instruction requires a structural fence. In some embodiments, such notification occurs along with a message on the local data structure from the CPU to ACC 297. The notification includes identification information for ACC 297 to identify the instruction when receiving the instruction through the local data structure, such as the memory address of the variable involved. In other embodiments, the notification can be achieved by adding a predetermined flag or marker to the memory access instruction to be modified when the CPU sends it to the structure bridge 296 through the local data structure. In some embodiments, such a marker includes information necessary to determine whether an acquire fence instruction or a release fence instruction is required.

[0051] At block 608, in response to receiving each notification, ACC 297 will issue acquire / release data center structure fences. As Figure 4 indicated by the fence arrows in , these fences are commands inserted into the command stream that goes to the data structure and continues to the media controller in the relevant FAM region. The media controller then executes these commands to provide fences for the variables.

[0052] Although the depicted process occurs after the application has been modified to include FABRIC_ACQUIRE_FENCE and FABRIC_RELEASE_FENCE instructions, in some embodiments, the process also includes inserting such hints or markers into the application so that μCode 216 can identify the selected memory instructions.

[0053] Figure 7 Flowchart 700 shows another process for implementing a memory consistency model according to some additional embodiments. The depicted process is executed for an application that has not been modified with compile-time instructions to implement structural fences. This process has the advantages that an application for which the developer has not provided a version configured to work with the FAM system can still work with the FAM system without causing the consistency problems that would normally occur. Another advantage is that for a deployment where a separate or more expensive software license is required to obtain a version of the application that works with the FAM system, this process can enable the use of a non-FAM version with the FAM system.

[0054] At block 702, the process starts activating FAM fences for such an application. Identifying the selected memory access instructions that require fence commands is different in this process from Figure 6The process because there are no specific instructions in the application for the structural fence. At block 704, the CPU μCode 216 identifies memory access instructions in the application that require a fence based on the identification of instructions in a predetermined list of instruction types and instruction prefixes associated with synchronization between threads. In some embodiments, a predetermined list is provided to μCode 216 to cover all types of applications for which it can identify the selected memory access instructions. The list includes instructions and instruction prefixes associated with dependencies between applications. For example, the list may include the LOCK instruction prefix for x86, x86's xacquire / xrelease, and the SYNC instruction for PowerPC.

[0055] At block 706, whenever any instruction from the predefined list is called, the CPU μCode 216 will notify the ACC 297, which will then issue a structural fence as shown at block 708.

[0056] The FAM memory interface 295 or any part thereof (such as the ACC 297 or the structural bridge 296) can be described or represented by a computer-accessible data structure in the form of a database or other data structures that can be read by a program and used directly or indirectly in the manufacture of integrated circuits. For example, the data structure can be a behavioral-level description or a register transfer level (RTL) description of the hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool, which can synthesize the description to produce a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represents the functionality of the hardware including the integrated circuit. The netlist can then be placed and routed to produce a dataset that describes the geometry to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce the integrated circuit. Alternatively, the database on the computer-accessible storage medium can be a netlist (with or without a synthesis library) or a dataset (as needed) or a Graphics Data System (GDS) II data.

[0057] Although specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, the internal architectures of the structural bridge 296 and the ACC 297 can vary in different embodiments. The types of FAM and FAM protocols employed can also vary. Additionally, a particular structural architecture can vary from an architecture that provides disaggregation within a data node or a rack of multiple data nodes using a structural protocol and PCIe or CXL-based transports to an architecture that can employ an optical or copper network using a protocol such as Gen-Z to connect devices and racks in a data center. Accordingly, the appended claims are intended to cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.

Claims

1. A method for use with a fabric-attached memory system, the fabric-attached memory system including a fabric-attached memory and a plurality of requesters coupled to the fabric-attached memory via a fabric, the method comprises: requesting from a fabric manager a notification of a change to requesters authorized to access a fabric-attached memory region; and in response to a notification from the fabric manager indicating that more than one requester is authorized to access the fabric-attached memory region, causing a fence to be activated for a selected memory access instruction.

2. The method according to claim 1, further comprises: in response to a notification from the fabric manager indicating that a single requester is authorized to access the fabric-attached memory region, mapping the fabric-attached memory region as a write-back memory region for the single requester.

3. The method according to claim 2, further comprises: in response to a notification from the fabric manager indicating that more than one requester is authorized to access the fabric-attached memory region, for each requester so authorized: if the requester is configured to issue a local fence command, causing a local fence command to be issued for the selected memory access instruction and mapping the fabric-attached memory region as a write-through memory for the requester; and if the requester is not configured to issue a local fence command, mapping the fabric-attached memory region as non-cacheable.

4. The method according to claim 1, wherein causing the fence to be activated comprises: in response to identifying the selected memory access instruction as indicating a dependency between requesters, causing a requester-side adaptive coherence controller to add a fence command for the selected memory access instruction.

5. The method according to claim 4, wherein identifying the selected memory access instruction is performed by processor instruction microcode running at each respective requester authorized to access the fabric-attached memory region.

6. The method according to claim 4, further comprising inserting a requester fabric interface tag into an application during compilation in response to identifying a coherence primitive and a coherence construct.

7. The method according to claim 4, wherein the selected memory access instruction is identified based on an identification of an instruction in a predetermined list of instruction types and instruction prefixes associated with data synchronization between threads.

8. The method according to claim 1, further comprises: at each requester authorized to access the fabric-attached memory region, maintaining a table of requesters authorized to access the fabric-attached memory region and updating the table in response to each of the notifications.

9. The method according to claim 8, wherein the table includes at least a fabric-attached memory region identifier and a count of computational requesters authorized to access the fabric-attached memory region.

10. A data processor, comprises: a processing core that executes an application; A structure-attached memory interface coupled to the processing core and adapted to be coupled to a data structure and execute memory access instructions from the processing core to the structure-attached memory; A requester-side adaptive coherence controller coupled to the processing core and the structure-attached memory interface and operative to: Request from the structure manager of the structure-attached memory a notification of a change of a requester authorized to access a structure-attached memory region to which the data processor is authorized to access; And In response to a notification from the structure manager indicating that more than one requester is authorized to access the structure-attached memory region, cause a fence to be activated for selected memory access instructions in a local application.

11. The data processor according to claim 10, wherein the requester-side adaptive coherence controller is further operative to, in response to a notification from the structure manager indicating that a single requester is authorized to access the structure-attached memory region, cause the structure-attached memory region to be mapped as a write-back memory region.

12. The data processor according to claim 11, wherein the requester-side adaptive coherence controller is further operative to, in response to a notification from the structure manager indicating that more than one requester is authorized to access the structure-attached memory region, perform the following operations: If the data processor is configured to issue a local fence command, cause a local fence command to be issued for selected memory access instructions and map the structure-attached memory region as a write-through memory; and If the data processor is not configured to issue a local fence command, cause the structure-attached memory region to be mapped as non-cacheable.

13. The data processor according to claim 10, further comprising processor instruction microcode stored in a tangible non-transitory memory accessible by and executable by the processing core for identifying selected memory access instructions.

14. The data processor according to claim 13, wherein the processor instruction microcode is executed in response to identifying a selected memory access instruction as indicating a dependency between requesters to notify the requester-side adaptive coherence controller of the identified selected memory access instruction, thereby causing the requester-side adaptive coherence controller to insert a requester-side fence command into a memory command to the structure-attached memory interface.

15. The data processor according to claim 13, wherein identifying the selected memory access instruction includes responding to an identifier identified in the local application, the identifier including one of a coherence primitive and a coherence construct.

16. The data processor according to claim 13, wherein identifying the selected memory access instruction includes identifying a structure interface tag inserted into the application during compilation.

17. The data processor according to claim 10, wherein the requester-side adaptive coherence controller maintains a table of requesters authorized to access the structure-attached memory region, and updates the table in response to the notification.

18. The data processor according to claim 17, wherein the table includes at least a structure-attached memory region identifier and the number of computational requesters authorized to access the structure-attached memory region.

19. A structure-attached memory system comprising: a structure-attached memory; a data structure coupled to the structure-attached memory; a structure manager coupled to the data structure and operable to authorize and de-authorize requesters to access memory regions of the structure-attached memory; a plurality of data processors coupled to the data structure, each data processor including a processing core that executes an application, a structure-attached memory interface, and a requester-side adaptive coherence controller coupled to the processing core and the structure-attached memory interface, and operable to: request from the structure manager a notification of a change in requesters authorized to access the structure-attached memory region to which the data processor is authorized access; and in response to a notification from the structure manager indicating that more than one requester is authorized to access the structure-attached memory region, cause a fence to be activated for selected memory access instructions in the local application.

20. The structure-attached memory system according to claim 19, wherein the requester-side adaptive coherence controller further operates to, in response to a notification from the structure manager indicating that a single requester is authorized to access the structure-attached memory region, cause the structure-attached memory region to be mapped as a write-back memory region.

21. The structure-attached memory system according to claim 20, wherein each respective requester-side adaptive coherence controller further operates to, in response to a notification from the structure manager indicating that more than one requester is authorized to access the structure-attached memory region, perform the following operations: if the data processor of the respective requester-side adaptive coherence controller is configured to issue a local fence command, cause a local fence command to be issued for selected memory access instructions and map the structure-attached memory region as a write-through memory; and if the data processor of the requester-side adaptive coherence controller is not configured to issue a local fence command, cause the structure-attached memory region to be mapped as non-cacheable.

22. The structure-attached memory system according to claim 19, wherein each respective data processor includes processor instruction microcode stored in a tangible non-transitory memory accessible by and executable by the processor core of the respective data processor for identifying selected memory access instructions.

23. The structured attached memory system according to claim 22, wherein the processor instruction microcode is executed in response to identifying a selected memory access instruction as indicating a dependency among requesters to notify the requester-side adaptive coherence controller of the identified selected memory access instruction, thereby causing the requester-side adaptive coherence controller to insert a requester-side fence command into the memory commands going to the structured attached memory interface.

24. The structured attached memory system according to claim 22, wherein identifying the selected memory access instruction includes identifying an indicator in the local application program, the indicator including one of a coherence primitive and a coherence construct.

25. The structured attached memory system according to claim 22, wherein identifying the selected memory access instruction includes identifying a structure interface tag inserted into the application program during compilation.

26. The structured attached memory system according to claim 19, wherein the requester-side adaptive coherence controller maintains a table of the requesters authorized to access the structured attached memory region and updates the table in response to the notification.

27. The structured attached memory system according to claim 26, wherein the table includes at least a structured attached memory region identifier and the number of computational requesters authorized to access the structured attached memory region.

Citation Information

Patent Citations

  • Ensuring consistency between a data cache and a main memory

    CN102388372A

  • Storage device unit restoring device and method and storage device system comprising device

    CN103295648A