Method and apparatus for performing memory reconfiguration without system reboot

By offloading the reconfiguration of system memory to the operating system with the assistance of the platform/firmware, the downtime and performance loss problems caused by server pool reconfiguration in the prior art are solved, and efficient memory subsystem reconfiguration without cold reset is achieved.

CN120723152APending Publication Date: 2025-09-30INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510231330.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-02-28
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the prior art, in cloud service providers, the reconfiguration process of the server pool requires a cold reset, resulting in loss of server uptime and performance loss, and the operating system is unable to dynamically enumerate memory configuration and performance characteristics, resulting in increased downtime.

Method used

By offloading system memory reconfiguration to the operating system with the assistance of the platform/firmware, the operating system enumerates different potential memory configurations and their performance characteristics and performs the memory subsystem reconfiguration without a cold reset.

Benefits of technology

Reduces server downtime, meets the stringent server uptime requirements of cloud service providers, and improves the efficiency and flexibility of storage subsystem configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723152A_ABST
    Figure CN120723152A_ABST
Patent Text Reader

Abstract

The cloud service provider reconfigures the memory subsystem during routine operations while minimizing the amount of time the server is not online. Server downtime is reduced by offloading reconfiguration of system memory to an operating system with the assistance of a platform. The operating system enumerates potential memory configurations and associated performance characteristics of the memory subsystem in an abstract manner and performs reconfiguration of the memory subsystem without a cold reset. When the operating system considers that reconfiguration of the memory subsystem is necessary, the operating system checks the enumerated memory subsystem configuration provided by the system firmware. After selecting the memory subsystem configuration, the operating system initiates a reconfiguration process. The reconfiguration process saves any existing memory context to the secondary device, requests system firmware to perform memory subsystem reconfiguration, and restores the existing memory context from the secondary device after the memory subsystem reconfiguration has completed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cloud computing provides access to servers, storage, databases, and a broad set of application services over the internet. Cloud service providers offer cloud services, such as computing, networking, and storage services, as well as business applications, hosted on servers in one or more data centers that can be accessed by companies or individuals over the internet. Hyperscale cloud service providers typically have hundreds of thousands of servers. Each server in a hyperscale cloud includes storage to store user data, such as for business intelligence, data mining, analytics, social media, and microservices. Cloud service providers generate revenue from companies and individuals (also known as tenants) who use cloud services. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Features of embodiments of the claimed subject matter will become apparent as the following detailed description proceeds, and with reference to the accompanying drawings, in which like reference numerals represent like parts, and in which:

[0003] Figure 1 is a simplified diagram of at least one embodiment of a data center for executing workloads utilizing disaggregated resources;

[0004] Figure 2 is a simplified diagram of at least one embodiment of a pod that may be included in a data center;

[0005] Figure 3 is a simplified block diagram of at least one embodiment of a top side of a node;

[0006] Figure 4 is a simplified block diagram of at least one embodiment of the bottom side of a node;

[0007] Figure 5 is a simplified block diagram of at least one embodiment of a computing node;

[0008] Figure 6 is a simplified block diagram of at least one embodiment of an accelerator node that may be used in a data center;

[0009] Figure 7 is a simplified block diagram of at least one embodiment of a storage node that may be used in a data center;

[0010] Figure 8 is a simplified block diagram of at least one embodiment of a memory node that may be used in a data center;

[0011] Figure 9 depicting a system for executing one or more workloads;

[0012] Figure 10 Depicts an example system;

[0013] Figure 11 An example system is shown;

[0014] Figure 12 It is a combination Figure 9 a block diagram of an embodiment of the described system;

[0015] Figure 13 Shown Figure 12 An embodiment of the ACPI System Resource Affinity Table (SRAT) in the table in;

[0016] Figure 14 Shown Figure 13 An embodiment of a flag field in an ACPI system resource affinity table (SRAT) in an SRAT; and

[0017] Figure 15 is a flow diagram illustrating offloading system memory reconfiguration to the operating system using the ACPI System Resource Affinity Table (SRAT) with assistance from the platform / firmware.

[0018] Although the following detailed description will be made with reference to illustrative embodiments of the claimed subject matter, many alternatives, modifications and variations thereof will be apparent to those skilled in the art. Therefore, the claimed subject matter is intended to be viewed broadly and defined only as set forth in the appended claims. DETAILED DESCRIPTION

[0019] Cloud service providers (CSPs) typically offer different tiers of virtual machines (VMs) as part of their service offerings. For example, servers can be categorized into pools: a general purpose ("GP") pool, a trusted domain ("TD") pool that runs secure VMs, and a pool with advanced memory capacity (a high memory pool). When operating a large data center with different pools, it may be necessary to reclassify some server instances from one pool to another for business reasons in certain scenarios. For example, a TD instance may need to be migrated to a GP instance, or vice versa. Such a migration may require reconfiguring the processor and / or memory devices to match the target pool.

[0020] Typically, nodes are reclassified by manually selecting a storage configuration, triggering a cold reset of the data server, and rebooting the server to the desired target storage configuration. However, this approach is very disruptive to currently active VMs and results in a loss of server uptime (the amount of time a service or server is online) due to the server cold reset. Furthermore, this approach does not meet the stringent server uptime requirements of cloud service providers.

[0021] Cloud service providers often need to reconfigure memory subsystems during routine operations while minimizing any downtime (the amount of time a server is offline). Reconfiguring a server's memory subsystem typically requires a cold reset of the server, which increases downtime and results in performance losses for the cloud service provider. Furthermore, because each server in a data center can uniquely implement its memory subsystem, the cloud service provider operating system (OS) cannot dynamically enumerate all possible memory configurations and the performance characteristics of each memory configuration.

[0022] During a cold reboot, the system is shut down, platform settings (e.g., hard / soft strap configuration, Unified Extensible Firmware Interface (UEFI) firmware setting knobs) are modified for the requested system configuration, and the system is rebooted. This results in reduced server uptime (and thus performance loss) and typically requires manual intervention. UEFI is a specification that defines an architecture for platform firmware used to boot computer hardware and its interfaces to interact with the operating system.

[0023] Instead of performing a cold reboot, a system management interrupt (SMI) can be injected to allow the Unified Extensible Firmware Interface (UEFI) firmware to reprogram the silicon / platform to match the desired target configuration. The use of a system management interrupt (SMI) introduces significant platform complexity to coordinate the activities of various agents across the system. Furthermore, the operating system (OS) / virtual memory manager (VMM) has little control over the timing of such events. For example, if a system management interrupt is triggered when the operating system has multiple jobs scheduled and is operating at its peak load, the performance loss can be significant.

[0024] In an embodiment, server downtime is reduced by offloading system memory reconfiguration to the operating system with the assistance of the platform / firmware. The operating system abstractly enumerates different potential memory configurations of the memory subsystem and associated performance characteristics and performs the memory subsystem reconfiguration without a cold reset.

[0025] When the operating system kernel determines that a memory subsystem reconfiguration is necessary, the operating system kernel examines the enumerated potential memory subsystem configurations provided by the system firmware. Each possible memory subsystem configuration includes performance characteristics (e.g., latency and bandwidth information). The performance characteristics allow the operating system kernel to make an informed choice of memory subsystem configuration.

[0026] After selecting the memory subsystem configuration, the operating system kernel initiates the reconfiguration process.

[0027] The reconfiguration process first saves any existing memory context to the auxiliary device, requests the system firmware to perform the memory subsystem reconfiguration, and restores the existing memory context from the auxiliary device after the memory subsystem reconfiguration has completed.

[0028] Various embodiments and aspects of the present invention will be described with reference to the details discussed below, and the accompanying drawings will illustrate various embodiments. The following description and the accompanying drawings are illustrative of the present invention and should not be construed as limiting the present invention. Many specific details are described to provide a thorough understanding of the various embodiments of the present invention. However, in specific examples, well-known or conventional details are not described in order to provide a concise discussion of embodiments of the present invention.

[0029] Reference in the specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase "in one embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.

[0030] Various embodiments and aspects of the present invention will be described with reference to the details discussed below, and the accompanying drawings will illustrate various embodiments. The following description and the accompanying drawings are illustrative of the present invention and should not be construed as limiting the present invention. Many specific details are described to provide a thorough understanding of the various embodiments of the present invention. However, in specific examples, well-known or conventional details are not described in order to provide a concise discussion of embodiments of the present invention.

[0031] Reference in the specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase "in one embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.

[0032] Figure 1A data center is depicted in which disaggregated resources can collaboratively execute one or more workloads (e.g., applications on behalf of a customer), the data center including a plurality of pods 110, 120, 130, 140, where a pod is or includes one or more rows of racks. Of course, while the data center 100 is shown as having multiple pods, in some embodiments, the data center 100 can be embodied as a single pod. As described in more detail herein, each rack houses a plurality of nodes, some of which can be equipped with one or more types of resources (e.g., memory devices, data storage devices, accelerator devices, general-purpose processors). Resources can be logically coupled to form combined nodes or composite nodes, which can be used, for example, as servers that execute jobs, workloads, or microservices. In an illustrative embodiment, the nodes in each pod 110, 120, 130, 140 are connected to a plurality of pod switches (e.g., switches that route data communications to or from nodes within a pod). The pod switches, in turn, connect to the spine switches 150 that exchange communications between the pods (e.g., pods 110, 120, 130, 140) in the data center 100. In some embodiments, the nodes can communicate with each other using The structure of Omni-Path technology is connected. In other embodiments, the nodes can be connected to other structures (e.g., InfiniBand or Ethernet or PCI Express or direct optical interconnect). As described in more detail herein, the resources within the nodes in the data center 100 can be assigned to a group containing resources from one or more nodes (referred to herein as "managed nodes") to be jointly utilized in the execution of the workload. The workload can be executed as if the resources belonging to the managed nodes are located on the same node. The resources in the managed nodes can belong to nodes belonging to different racks (and even to different pods 110, 120, 130, 140). Therefore, some resources of a single node can be assigned to a managed node, while other resources of the same node can be assigned to different managed nodes (e.g., a processor is assigned to a managed node, and another processor of the same node is assigned to a different managed node).

[0033] Data centers including disaggregated resources, such as data center 100, can be used in a variety of contexts, such as enterprises, governments, cloud service providers, and communications service providers (e.g., Telcos), and come in a variety of sizes, from large cloud service provider data centers consuming over 60,000 square feet to single-rack or multi-rack installations for use in base stations.

[0034] The decomposition of resources into nodes that primarily include a single type of resource (e.g., compute nodes that primarily include compute resources, memory nodes that primarily include memory resources), and the selective allocation and release of the decomposed resources to form managed nodes assigned to execute workloads, improves the operation and resource usage of the data center 100 relative to a typical data center composed of hyperconverged servers that contain compute, memory, storage, and possibly additional resources in a single chassis. For example, because the nodes primarily include a specific type of resource, resources of a given type can be upgraded independently of other resources. Additionally, because different resource types (processors, storage devices, accelerators, etc.) typically have different refresh rates, higher resource utilization and reduced total cost of ownership can be achieved. For example, a data center operator can upgrade processors throughout its facility by swapping out only the compute nodes. In this case, accelerators and storage resources can be allowed to continue operating instead of being upgraded simultaneously until these resources are scheduled for their own refresh. Resource utilization can also be increased. For example, if managed nodes are composed based on the requirements of the workloads that will be running on them, the resources within the nodes are more likely to be fully utilized. Such utilization may allow more managed nodes to be run in a data center with a given set of resources, or allow a data center expected to run a given set of workloads to be built using fewer resources.

[0035] Figure 2pods are depicted. A pod can include a collection of racks 240 arranged in rows 200, 210, 220, 230. Each rack 240 can house a plurality of nodes (e.g., sixteen nodes) and provide power and data connectivity to the housed nodes, as described in more detail herein. In the illustrative embodiment, the racks in each row 200, 210, 220, 230 are connected to a plurality of pod switches 250, 260. The pod switches 250 include a collection of ports 252 to which the nodes of the racks of the pod 110 are connected, and another collection of ports 254 that connects the pod 110 to the spine switch 150 to provide connectivity to other pods in the data center 100. Similarly, pod switch 260 includes a set of ports 262 to which the nodes of the rack of pod 110 are connected, and a set of ports 264 that connects pod 110 to spine switch 150. Thus, the use of a pair of switches 250, 260 provides a degree of redundancy for pod 110. For example, if either switch 250, 260 fails, the nodes in pod 110 can still maintain data communications with the rest of data center 100 (e.g., nodes of other pods) via the other switch 250, 260. Furthermore, in the illustrative embodiment, switches 150, 250, 260 may be embodied as dual-mode optical switches capable of routing both Ethernet protocol communications carrying Internet Protocol (IP) packets and communications according to a second, high-performance link layer protocol (e.g., PCI Express or Compute Express) via the optical signal transmission medium of fiber optics.

[0036] It should be appreciated that each of the other pods 120, 130, 140 (and any additional pods of the data center 100) may be connected to the Figure 2 Shown in and about Figure 2 The depicted pod 110 is similarly constructed and has similar components as pod 110 (e.g., each pod may have rows of racks housing multiple nodes, as described above). Additionally, while two pod switches 250, 260 are shown, it should be understood that in other embodiments, each pod 110, 120, 130, 140 may be connected to a different number of pod switches, thereby providing even more failover capabilities. Of course, in other embodiments, the pods may be connected to multiple pod switches. Figure 1-Figure 2 The configuration of the rows of racks shown in may be arranged differently. For example, a pod may be embodied as multiple sets of racks, wherein each set of racks is arranged radially, ie, the racks are equidistant from a central switch.

[0037] Now refer to Figure 3In the illustrative embodiment, the nodes 400 are configured to be installed in corresponding racks 240 of the data center 100, as discussed above. In some embodiments, each node 400 may be optimized or otherwise configured to perform a specific task, such as a computing task, an acceleration task, a data storage task, etc. For example, the node 400 may be embodied as follows: Figure 5 The computing node 500 discussed below Figure 6 The accelerator node 600 discussed below is Figure 7 The storage node 700 discussed herein may be embodied as a node optimized or otherwise configured to perform other specialized tasks, such as described below with respect to Figure 8 Discussed memory node 800. Each rack 240 may contain one or more nodes of a single or multiple node types (compute, storage, accelerator, memory, etc.).

[0038] As discussed above, the illustrative node 400 includes a circuit board substrate 302 that supports various physical resources (eg, electrical components) mounted thereon.

[0039] As discussed above, the illustrative node 400 includes one or more physical resources 320 mounted to the top side 350 of the circuit board substrate 302. Figure 3 , two physical resources 320 are shown, but it should be appreciated that in other embodiments, node 400 may include one, two, or more physical resources 320. Physical resources 320 may be embodied as any type of processor, controller, or other computing circuitry capable of performing various tasks (e.g., computing functions) and / or controlling the functionality of node 400 (depending, for example, on the type or intended functionality of node 400). For example, as discussed in more detail below, physical resources 320 may be embodied as a high-performance processor in embodiments in which node 400 is embodied as a compute node, as an accelerator coprocessor or circuit in embodiments in which node 400 is embodied as an accelerator node, as a storage controller in embodiments in which node 400 is embodied as a storage node, or as a collection of memory devices in embodiments in which node 400 is embodied as a memory node.

[0040] Node 400 also includes one or more additional physical resources 330 mounted to top side 350 of circuit board substrate 302. In the illustrative embodiment, the additional physical resources include a network interface controller (NIC), as discussed in more detail below. Of course, depending on the type and functionality of node 400, physical resources 330 may include additional or other electrical components, circuits, and / or devices in other embodiments.

[0041] Physical resource 320 may be communicatively coupled to physical resource 330 via input / output (I / O) subsystem 322. I / O subsystem 322 may embody circuitry and / or components for facilitating input / output operations with physical resource 320, physical resource 330, and / or other components of node 400. For example, I / O subsystem 322 may embody or otherwise include a memory controller hub, an input / output control hub, an integrated sensor hub, a firmware device, communication links (e.g., point-to-point links, bus links, wires, cables, waveguides, optical guides, printed circuit board traces, etc.), and / or other components and subsystems for facilitating input / output operations.

[0042] In some embodiments, node 400 may further include a resource-to-resource interconnect 324. Resource-to-resource interconnect 324 may be embodied as any type of communication interconnect capable of facilitating resource-to-resource communication. In an illustrative embodiment, resource-to-resource interconnect 324 is embodied as a high-speed point-to-point interconnect (e.g., faster than I / O subsystem 322). For example, resource-to-resource interconnect 324 may be embodied as QuickPath Interconnect (QPI), UltraPath Interconnect (UPI), PCI Express (PCIe), or other high-speed point-to-point interconnects dedicated for resource-to-resource communication.

[0043] The node 400 also includes a power connector 340 that is configured to mate with a corresponding power connector of the rack 240 when the node 400 is installed in the corresponding rack 240. The node 400 receives power from the power supply of the rack 240 via the power connector 340 to supply power to the various electrical components of the node 400. That is, the node 400 does not include any local power supply (e.g., an on-board power supply) for providing power to the electrical components of the node 400. The exclusion of a local power supply or an on-board power supply facilitates reducing the overall footprint of the circuit board substrate 302, which can increase the thermal cooling characteristics of the various electrical components mounted on the circuit board substrate 302, as discussed above. In some embodiments, the voltage regulator is placed on the circuit board substrate 302 at a location adjacent to the processor 520 (see Figure 5 ) opposite the bottom side 450 (see Figure 4 ), and power is routed from the voltage regulator to the processor 520 through vias extending through the circuit board substrate 302. This configuration provides an increased thermal budget, additional current and / or voltage, and better voltage control relative to a typical printed circuit board (where processor power is delivered in part from the voltage regulator through printed circuit traces).

[0044] In some embodiments, the node 400 may also include mounting features 342 that are configured to cooperate with a mounting arm or other structure of the robot to facilitate the robot's placement of the node 400 in the rack 240. The mounting features 342 may be embodied as any type of physical structure that allows the robot to grasp the node 400 without damaging the circuit board substrate 302 or the electrical components mounted thereon. For example, in some embodiments, the mounting features 342 may be embodied as non-conductive pads attached to the circuit board substrate 302. In other embodiments, the mounting features may be embodied as brackets, mounts, or other similar structures attached to the circuit board substrate 302. The specific number, shape, size, and / or composition of the mounting features 342 may depend on the design of the robot configured to manage the node 400.

[0045] Now refer to Figure 4 In addition to the physical resources 330 mounted on the top side 350 of the circuit board substrate 302, the node 400 also includes one or more memory devices 420 mounted to the bottom side 450 of the circuit board substrate 302. That is, the circuit board substrate 302 can be embodied as a double-sided circuit board. The physical resources 320 can be communicatively coupled to the memory devices 420 via the I / O subsystem 322. For example, the physical resources 320 and the memory devices 420 can be communicatively coupled by one or more through-holes extending through the circuit board substrate 302. In some embodiments, the physical resources 320 can be communicatively coupled to different sets of one or more memory devices 420. Alternatively, in other embodiments, each physical resource 320 can be communicatively coupled to each memory device 420.

[0046] Memory device 420 may be embodied as any type of memory device capable of storing data for physical resources 320 during operation of node 400 , such as any type of volatile memory (eg, dynamic random access memory (DRAM), etc.) or non-volatile memory.

[0047] Volatile memory is memory whose state (and therefore the data stored therein) is indeterminate if power to the device is interrupted. Dynamic volatile memory requires that the data stored in the device be refreshed to maintain the state. An example of dynamic volatile memory includes DRAM (Dynamic Random Access Memory) or some variants, such as Synchronous Dynamic Random Access Memory (SDRAM). The memory subsystem described herein may be compatible with a variety of memory technologies, such as DDR3 (Double Data Rate Version 3) JESD79-3F originally released by JEDEC (Joint Electron Device Engineering Council) in June 2007, DDR4 (DDR Version 4) JESD209-4D originally released in September 2012, DDR5 (DDR Version 5) JESD79-5B originally released in June 2021, DDR6 (DDR Version 6) currently under discussion by JEDEC, LPDDR3 (Low Power DDR Version 3) JESD209-3C originally released in August 2015, and LPDDR4 (LPDDR Version 4) originally released in June 2021. JESD209-4D, LPDDR5 (LPDDR Version 5) JESD209-5B, originally published in June 2021, WIO2 (Wide Input / Output Version 2) JESD229-2, originally published in August 2014, HBM (High Bandwidth Memory) JESD235B, originally published in December 2018, HBM2 (HBM Version 2) JESD235D, originally published in March 2021, HBM3 (HBM Version 3) JESD238A, originally published in January 2023, or HBM4 (HBM Version 4) currently under discussion by JEDEC, or other memory technologies or combinations of memory technologies, and derivatives or extensions based on such specifications. JEDEC standards are available at www.jedec.org.

[0048] In one embodiment, the memory device is a block addressable memory device, such as those based on NAND or NOR technology, such as multi-threshold level NAND flash memory and NOR flash memory. Blocks can have any size, such as, but not limited to, 2KB, 4KB, 5KB, etc. The memory device may also include next generation non-volatile devices, such as Intel Memory or other byte-addressable write-in-place non-volatile memory devices (e.g., memory devices using chalcogenide glass), single-level or multi-level phase change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), antiferroelectric memory, magnetoresistive random access memory (MRAM) memory combined with memristor technology, resistive memory including metal oxide-based, oxygen vacancy-based and conductive bridge random access memory (CB-RAM) or spin transfer torque (STT)-MRAM, devices based on spintronic magnetic junction memory, devices based on magnetic tunneling junction (MTJ), devices based on DW (domain wall) and SOT (spin-orbit transfer), thyristor-based memory devices, or any combination of the above memories, or other memories. The memory device may refer to the die itself and / or to the packaged memory product. In some embodiments, the memory device may include a transistor-free stackable cross-point architecture in which memory cells are located at the intersection of word lines and bit lines and are individually addressable, and in which bit storage is based on changes in bulk resistance.

[0049] Now refer to Figure 5 In some embodiments, node 400 may be embodied as a computing node 500. Computing node 500 may be configured to perform computing tasks. Of course, as discussed above, computing node 500 may rely on other nodes (e.g., acceleration nodes and / or storage nodes) to perform computing tasks.

[0050] In the illustrative computing node 500, the physical resources 320 are embodied as processors 520. Although Figure 5 Only two processors 520 are shown in FIG, but it should be appreciated that in other embodiments, the computing node 500 may include additional processors 520. Illustratively, the processors 520 are embodied as high-performance processors 520 and may be configured to operate at a relatively high power rating.

[0051] In some embodiments, the compute node 500 may also include a processor-to-processor interconnect 542. The processor-to-processor interconnect 542 may be embodied as any type of communication interconnect capable of facilitating communication across the processor-to-processor interconnect 542. In an illustrative embodiment, the processor-to-processor interconnect 542 is embodied as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 322). For example, the processor-to-processor interconnect 542 may be embodied as a QuickPath Interconnect (QPI), an UltraPath Interconnect (UPI), or other high-speed point-to-point interconnect dedicated for processor-to-processor communication (e.g., PCIe or CXL).

[0052] The computing node 500 also includes communication circuitry 530. The illustrative communication circuitry 530 includes a network interface controller (NIC) 532, which may also be referred to as a host fabric interface (HFI). The NIC 532 may be embodied as or otherwise include any type of integrated circuit, discrete circuit, controller chip, chipset, add-in board, daughter card, network interface card, or other device that can be used by the computing node 500 to connect to another computing device (e.g., with another node 400). In some embodiments, the NIC 532 may be embodied as part of a system on a chip (SoC) that includes one or more processors, or included on a multi-chip package that also includes one or more processors. In some embodiments, the NIC 532 may include a local processor (not shown) and / or local memory (not shown), both of which are local to the NIC 532. In such embodiments, the local processor of the NIC 532 may be capable of performing one or more of the functions of the processor 520. Additionally or alternatively, in such embodiments, the local memory of the NIC 532 can be integrated into one or more components of the compute node at the board level, socket level, chip level, and / or other level. In some examples, the network interface comprises a network interface controller or a network interface card. In some examples, the network interface can include one or more of a network interface controller (NIC) 532, a host fabric interface (HFI), a host bus adapter (HBA), a network interface connected to a bus or connection (e.g., PCIe, CXL, DDR, etc.). In some examples, the network interface can be part of a switch or a system on a chip (SoC).

[0053] Communication circuitry 530 is communicatively coupled to optical data connector 534. Optical data connector 534 is configured to mate with a corresponding optical data connector of the rack when compute node 500 is installed in the rack. Illustratively, optical data connector 534 includes a plurality of optical fibers leading from a mating surface of optical data connector 534 to optical transceiver 536. Optical transceiver 536 is configured to convert incoming optical signals from the rack-side optical data connector into electrical signals, and convert electrical signals into outgoing optical signals destined for the rack-side optical data connector. While optical transceiver 536 is shown in the illustrative embodiment as forming part of optical data connector 534, in other embodiments, optical transceiver 536 may form part of communication circuitry 530 or even processor 520.

[0054] In some embodiments, the computing node 500 may further include an expansion connector 540. In such embodiments, the expansion connector 540 is configured to mate with a corresponding connector of an expansion circuit board substrate to provide additional physical resources to the computing node 500. During operation of the computing node 500, the additional physical resources may be used, for example, by the processor 520. The expansion circuit board substrate may be substantially similar to the circuit board substrate 302 discussed above and may include various electrical components mounted thereon. The specific electrical components mounted to the expansion circuit board substrate may depend on the intended function of the expansion circuit board substrate. For example, the expansion circuit board substrate may provide additional computing resources, memory resources, and / or storage resources. Therefore, the additional physical resources of the expansion circuit board substrate may include, but are not limited to, processors, memory devices, storage devices, and / or accelerator circuits, including, for example, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), security coprocessors, graphics processing units (GPUs), machine learning circuits, or other specialized processors, controllers, devices, and / or circuits.

[0055] Now refer to Figure 6 In some embodiments, node 400 may be embodied as an accelerator node 600. Accelerator node 600 is configured to perform specialized computational tasks, such as machine learning, encryption, hashing, or other computationally intensive tasks. In some embodiments, for example, compute node 500 may offload tasks to accelerator node 600 during operation. Accelerator node 600 includes various components similar to those of node 400 and / or compute node 500, which have been described in detail in the prior art. Figure 6 The same reference numerals are used for identification.

[0056] In the illustrative accelerator node 600, the physical resources 320 are embodied as accelerator circuits 620. Although Figure 6 Only two accelerator circuits 620 are shown, but it should be appreciated that in other embodiments, the accelerator node 600 may include additional accelerator circuits 620. The accelerator circuits 620 may be embodied as any type of processor, coprocessor, computational circuit, or other device capable of performing computational or processing operations. For example, the accelerator circuits 620 may be embodied as, for example, a central processing unit, a core, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a programmable control logic (PCL), a security coprocessor, a graphics processing unit (GPU), a neuromorphic processor unit, a quantum computer, a machine learning circuit, or other specialized processors, controllers, devices, and / or circuits.

[0057] In some embodiments, the accelerator node 600 may further include an accelerator-to-accelerator interconnect 642. Similar to the resource-to-resource interconnect 324 of the node 400 discussed above, the accelerator-to-accelerator interconnect 642 may be embodied as any type of communication interconnect capable of facilitating accelerator-to-accelerator communication. In an illustrative embodiment, the accelerator-to-accelerator interconnect 642 is embodied as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 622). For example, the accelerator-to-accelerator interconnect 642 may be embodied as a QuickPath Interconnect (QPI), an UltraPath Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to processor-to-processor communication. In some embodiments, the accelerator circuit 620 may be daisy-chained with a primary accelerator circuit 620 connected to the NIC 532 and the memory 420 via the I / O subsystem 622, and a secondary accelerator circuit 620 connected to the NIC 532 and the memory 420 via the primary accelerator circuit 620.

[0058] Now refer to Figure 7 In some embodiments, the node 400 may be embodied as a storage node 700. The storage node 700 is configured to store data in a data storage device 750 local to the storage node 700. For example, during operation, the compute node 500 or the accelerator node 600 may store and retrieve data from the data storage device 750 of the storage node 700. The storage node 700 includes various components similar to those of the node 400 and / or the compute node 500, which have been described in detail in the accompanying drawings. Figure 7 The same reference numerals are used for identification.

[0059] In the illustrative storage node 700, the physical resource 320 is embodied as a storage controller 720. Although Figure 7 Only two storage controllers 720 are shown in FIG. 5 , but it should be appreciated that in other embodiments, the storage node 700 may include additional storage controllers 720. The storage controllers 720 may be embodied as any type of processor, controller, or control circuit capable of controlling the storage and retrieval of data to and from the data storage device 750 based on requests received via the communication circuit 530. In an illustrative embodiment, the storage controllers 720 are embodied as relatively low-power processors or controllers. For example, in some embodiments, the storage controllers 720 may be configured to operate at a nominal power of approximately 75 watts.

[0060] In some embodiments, storage node 700 may also include a controller-to-controller interconnect 742. Similar to the resource-to-resource interconnect 324 of node 400 discussed above, controller-to-controller interconnect 742 may be embodied as any type of communication interconnect capable of facilitating controller-to-controller communication. In an illustrative embodiment, controller-to-controller interconnect 742 is embodied as a high-speed point-to-point interconnect (e.g., faster than I / O subsystem 622). For example, controller-to-controller interconnect 742 may be embodied as a QuickPath Interconnect (QPI), an UltraPath Interconnect (UPI), or other high-speed point-to-point interconnect dedicated for processor-to-processor communication.

[0061] Now refer to Figure 8 In some embodiments, the node 400 may be embodied as a memory node 800. The memory node 800 is configured to provide other nodes 400 (e.g., compute nodes 500, accelerator nodes 600, etc.) with access to a pool of memory local to the storage node 700 (e.g., in two or more sets 830, 832 of memory devices 420). For example, during operation, the compute node 500 or accelerator node 600 may remotely write to and / or read from one or more of the memory sets 830, 832 of the memory node 800 using a logical address space that maps to physical addresses in the memory sets 830, 832.

[0062] In the illustrative memory node 800, the physical resource 320 is embodied as a memory controller 820. Although Figure 8 830 , 832 . The memory controllers 820 are shown in FIG. 800 , but it should be appreciated that in other embodiments, the memory node 800 may include additional memory controllers 820. The memory controllers 820 may be embodied as any type of processor, controller, or control circuitry capable of controlling the writing and reading of data into the memory sets 830, 832 based on requests received via the communication circuitry 530. In the illustrative embodiment, each memory controller 820 is connected to a corresponding memory set 830, 832 to write to and read from the memory devices 420 within the corresponding memory set 830, 832, and enforce any permissions (e.g., read, write, etc.) associated with a node 400 that has sent a request to the memory node 800 to perform a memory access operation (e.g., read or write).

[0063] In some embodiments, the memory node 800 may also include a controller-to-controller interconnect 842. Similar to the resource-to-resource interconnect 324 of the node 400 discussed above, the controller-to-controller interconnect 842 may be embodied as any type of communication interconnect capable of facilitating controller-to-controller communication. In an illustrative embodiment, the controller-to-controller interconnect 842 is embodied as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 622). For example, the controller-to-controller interconnect 842 may be embodied as a QuickPath Interconnect (QPI), an UltraPath Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to processor-to-processor communication. Thus, in some embodiments, the memory controller 820 may access memory associated with another memory controller 820 within the memory set 832 via the controller-to-controller interconnect 842. In some embodiments, the scalable memory controller is made of multiple smaller memory controllers (referred to herein as "chiplets") on a memory node (e.g., the memory node 800). The chiplets may be interconnected (e.g., using EMIB (Embedded Multi-Die Interconnect Bridge)). The combined chip memory controller can be extended to a relatively large number of memory controllers and I / O ports (e.g., up to 16 memory channels). In some embodiments, memory controller 820 can implement memory interleaving (e.g., one memory address is mapped to memory set 830, the next memory address is mapped to memory set 832, and the third address is mapped to memory set 830, etc.). Interleaving can be managed within memory controller 820 or from a CPU socket (e.g., a CPU socket of compute node 500) across a network link to memory sets 830, 832, and this can improve the latency associated with performing memory access operations compared to accessing consecutive memory addresses from the same memory device.

[0064] In addition, in some embodiments, the memory node 800 can be connected to one or more other nodes 400 (e.g., in the same rack 240 or an adjacent rack 240) via waveguide using a waveguide connector 880. In the illustrative embodiment, the waveguide is a 64 mm waveguide that provides 16 Rx (e.g., receive) channels and 16 Tx (e.g., transmit) channels. In the illustrative embodiment, each channel is 16 GHz or 32 GHz. In other embodiments, the frequency can be different. Using a waveguide can provide high throughput access to a memory pool (e.g., memory collections 830, 832) to another node (e.g., a node 400 in the same rack 240 or an adjacent rack 240 as the memory node 800) without increasing the load on the optical data connector 534.

[0065] Now refer to Figure 9, a system for executing one or more workloads (e.g., applications) can be implemented. In an illustrative embodiment, system 910 includes an orchestrator server 920, which can be embodied as a managed node including a computing device (e.g., processor 520 on compute node 500) executing management software (e.g., a cloud operating environment, such as OpenStack), the managed node being communicatively coupled to a plurality of nodes 400, including a plurality of compute nodes 930 (e.g., each compute node 930 is similar to compute node 500), memory nodes 940 (e.g., each memory node 940 is similar to memory node 800), accelerator nodes 950 (e.g., each accelerator node 950 is similar to accelerator node 600), and storage nodes 960 (e.g., each storage node 960 is similar to storage node 700). One or more of nodes 930 , 940 , 950 , 960 may be grouped into managed nodes 970 , for example, by orchestrator server 920 , to collectively execute a workload (eg, application 932 executed in a virtual machine or container).

[0066] A managed node 970 may be embodied as a combination of physical resources 320 (e.g., processors 520, memory resources 420, accelerator circuits 620, or data storage devices 750) from the same or different nodes 400. Physical resources 320 from the same compute node 500, or the same memory node 800, or the same accelerator node 600, or the same storage node 700 may be assigned to a single managed node 970. Alternatively, physical resources 320 from the same node 400 may be assigned to different managed nodes 970. Furthermore, a managed node may be created, defined, or "spun up" by the orchestrator server 920 when a workload is assigned to the managed node or at any other time, and may exist regardless of whether any workload is currently assigned to the managed node. In an illustrative embodiment, orchestrator server 920 can selectively allocate and / or de-allocate physical resources 320 from nodes 400 and / or add or remove one or more nodes 400 from managed nodes 970 based on quality of service (QoS) targets (e.g., target throughput, target latency, target number of instructions per second, etc.) associated with a service level agreement for a workload (e.g., application 932). In doing so, orchestrator server 920 can receive telemetry data indicating performance conditions (e.g., throughput, latency, instructions per second, etc.) in each node 400 of managed nodes 970 and compare the telemetry data with the quality of service targets to determine whether the quality of service targets are met. Orchestrator server 920 can also determine whether one or more physical resources can be de-allocated from managed node 970 while still meeting the QoS targets, thereby freeing those physical resources for use in another managed node (e.g., to execute a different workload). Alternatively, if the QoS target is not currently being met, orchestrator server 920 can determine to dynamically allocate additional physical resources to assist with the execution of the workload (e.g., application 932) while the workload is executing. Similarly, if orchestrator server 920 determines that de-allocating the physical resources will result in the QoS target still being met, orchestrator server 920 can determine to dynamically de-allocate the physical resources from the managed node.

[0067] Additionally, in some embodiments, orchestrator server 920 can identify trends in resource utilization of a workload (e.g., application 932) by, for example, identifying execution phases of the workload (e.g., application 932) (e.g., time periods during which different operations are executed, each with different resource utilization characteristics) and preemptively identifying available resources in data center 100 and allocating these resources to managed nodes 970 (e.g., within a predefined time period after the associated phase begins). In some embodiments, orchestrator server 920 can model performance based on various latency and distribution schemes to place workloads among compute nodes and other resources (e.g., accelerator nodes, memory nodes, storage nodes) in the data center. For example, orchestrator server 920 can utilize a model that considers the performance of resources on node 400 (e.g., FPGA performance, memory access latency, etc.) as well as the performance of the path to the resource (e.g., FPGA) through the network (e.g., congestion, latency, bandwidth). Thus, orchestrator server 920 can determine which resource(s) should be used with which workloads based on the total latency associated with each potential resource available in data center 100 (e.g., latency associated with the performance of the resource itself, in addition to latency associated with the path through the network between the compute node executing the workload and node 400 where the resource is located).

[0068] In some embodiments, orchestrator server 920 can use telemetry data reported from nodes 400 (e.g., temperature, fan speed, etc.) to generate a thermal map of data center 100 and allocate resources to managed nodes based on the thermal map and predicted heat generation associated with different workloads to maintain target temperatures and thermal distribution in data center 100. Additionally or alternatively, in some embodiments, orchestrator server 920 can organize the received telemetry data into a hierarchical model that indicates relationships among managed nodes (e.g., spatial relationships (e.g., the physical location of resources of managed nodes within data center 100) and / or functional relationships (e.g., grouping managed nodes by the customers they provide services to, the types of functions typically performed by managed nodes, managed nodes that typically share or exchange workloads with each other, etc.). Based on the differences in physical location and resources among managed nodes, a given workload can exhibit different resource utilization among the resources of different managed nodes (e.g., resulting in different internal temperatures, different percentages of processor or memory capacity used). Orchestrator server 920 can determine the difference based on the telemetry data stored in the hierarchical model and factor the difference into the prediction of the workload's future resource utilization if the workload is reassigned from one managed node to another managed node to accurately balance resource utilization in data center 100. In some embodiments, orchestrator server 920 can identify patterns in the workload's resource utilization phases and use the patterns to predict the workload's future resource utilization.

[0069] To reduce the computational load on orchestrator server 920 and the data transmission load on the network, in some embodiments, orchestrator server 920 may send self-test information to nodes 400, so that each node 400 can locally (e.g., on node 400) determine whether the telemetry data generated by node 400 meets one or more conditions (e.g., available capacity meeting a predefined threshold, temperature meeting a predefined threshold, etc.). Each node 400 may then report a simplified result (e.g., yes or no) back to orchestrator server 920, which may use the simplified result to determine the allocation of resources to the managed node.

[0070] At a general level, edge computing refers to the implementation, coordination, and use of computing and resources at a location closer to the "edge" or a collection of "edges" of a network. The goals of this arrangement are to improve total cost of ownership, reduce application and network latency, reduce network backhaul traffic and associated energy consumption, improve service capabilities, and improve compliance with security or data privacy requirements (especially compared to conventional cloud computing). Components that can perform edge computing operations ("edge nodes") can reside anywhere the system architecture or temporary service requires (e.g., in a high-performance computing data center or cloud installation; in a designated edge node server, enterprise server, roadside server, telecommunications central office; or in a local or peer edge device served by the edge service).

[0071] Applications already suitable for edge computing include, but are not limited to, virtualization of traditional network functions (e.g., to operate telecommunications or internet services) and the introduction of next-generation features and services (e.g., to support 5G network services). Use cases expected to widely utilize edge computing include connected autonomous vehicles, surveillance, Internet of Things (IoT) device data analysis, video encoding and analysis, location-aware services, device sensing in smart cities, and many other network- and compute-intensive services.

[0072] In some scenarios, edge computing can provide or host cloud-like distributed services to provide orchestration and management to applications and coordinate service instances across many types of storage and computing resources. Edge computing is also expected to be closely integrated with existing use cases and technologies developed for IoT and fog / distributed networking configurations, as endpoint devices, clients, and gateways attempt to access network resources and applications closer to the edge of the network.

[0073] The following embodiments generally relate to data processing, service management, resource allocation, compute management, network communications, application partitioning, and communication system implementations, and specifically relate to techniques and configurations for adapting various edge computing devices and entities to dynamically support multiple entities (e.g., multiple tenants, users, stakeholders, service instances, applications, etc.) in a distributed edge computing environment.

[0074] In the following description, methods, configurations, and related apparatus are disclosed for making various improvements to the configuration and functional capabilities of edge computing architectures and for implementing edge computing systems. These improvements can benefit a variety of use cases, particularly those involving multiple stakeholders of an edge computing system, whether in the form of multiple users of the system, multiple tenants on the system, multiple devices or user devices interacting with the system, multiple services provided from the system, multiple resources available or managed within the system, multiple forms of network access exposed to the system, multiple operating locations of the system, and the like. Such multi-dimensional aspects and considerations are generally referred to herein as "multi-entity" constraints, with specific discussion of resources managed or orchestrated in multi-tenant and multi-service edge computing configurations.

[0075] With the illustrative edge networking system described below, computing and storage resources are moved closer to the edge of the network (e.g., closer to clients, endpoint devices, or "things"). By moving computing and storage resources closer to the devices that generate or use data, various latency, compliance, and / or monetary or resource cost constraints can be achieved relative to standard networking (e.g., cloud computing) systems. To this end, in some examples, a pool of computing, memory, and / or storage resources can be located in or otherwise equipped with local servers, routers, and / or other network devices. Such local resources facilitate meeting the constraints imposed on the system. For example, local computing and storage resources allow edge systems to perform computations in real time or near real time, which may be a consideration in low-latency use cases (e.g., autonomous driving, video surveillance, and mobile media consumption). Additionally, these resources will benefit from service management in the edge system, which provides the ability to extend and implement local service level agreements (SLAs), manage tiered service requirements, and enable local features and functionality on a temporary or permanent basis.

[0076] The illustrative edge computing system can support a variety of services and / or provide a variety of services to endpoint devices (e.g., client user equipment (UE)), each of which may have different requirements or constraints. For example, certain services may have priority or quality of service (QoS) constraints (e.g., traffic data for an autonomous vehicle may have a higher priority than temperature sensor data), reliability and resiliency (e.g., traffic data may require mission-critical reliability, while temperature data may tolerate some error variation), as well as power, cooling, and form factor constraints. These and other technical constraints can introduce significant complexity and technical challenges when applied in a multi-stakeholder environment.

[0077] However, along with the advantages of edge computing come the following considerations. Devices located at the edge are typically resource constrained, and therefore there is pressure to use edge resources. Typically, this is addressed by pooling memory and storage resources used by multiple users (tenants) and devices. The edge may be power and cooling constrained, and therefore power usage needs to be considered by the applications that consume the most power. There may be an inherent power-performance tradeoff in these pooled memory resources, as many of these memory resources may use emerging memory technologies where more power requires greater memory bandwidth. Similarly, there is also a requirement to improve the security of hardware and root-of-trust trusted functions, as edge locations may be unmanned and may even require permissioned access (for example, when housed in a third-party location). Such issues are magnified in edge clouds in multi-tenant, multi-owner, or multi-access settings where many users request services and applications, especially when network usage fluctuates dynamically and the composition of multiple stakeholders, use cases, and services changes.

[0078] Figure 10 An edge computing system 1000 for providing edge services and applications to multiple stakeholder entities is generally depicted, the edge computing system 1000 being distributed across multiple layers of the network, among one or more client computing nodes 1002, one or more edge gateway nodes 1012, one or more edge aggregation nodes 1022, one or more core data centers 1032, and a global network cloud 1042. Implementations of the edge computing system 1000 can be provided at or on behalf of a telecommunications service provider ("telco" or "TSP"), an IoT service provider, a cloud service provider (CSP), an enterprise entity, or any other number of entities. Various implementations and configurations of the edge computing system 1000 can be provided dynamically (e.g., when orchestrated to meet service objectives).

[0079] For example, the client computing nodes 1002 are located at the endpoint layer, while the edge gateway nodes 1012 are located at the edge device layer (local layer) of the edge computing system 1000. Additionally, the edge aggregation nodes 1022 (and / or fog devices 1024 if arranged in or operating with the fog networking configuration 1026)) are located at the network access layer (intermediate level). Fog computing (or "fogging") generally refers to the extension of cloud computing to the edge of an enterprise network, or to the ability to manage transactions across cloud / edge environments, typically in a coordinated distributed network or multi-node network. Some forms of fog computing provide deployment of compute, storage, and networking services between end devices and cloud computing data centers representing cloud computing locations. In terms of overall transactions, some forms of fog computing also provide the ability to manage workload / workflow level services by pushing certain workloads to the edge or cloud based on the ability to meet overall service level agreements.

[0080] Fog computing in many scenarios provides a decentralized architecture and serves as an extension of cloud computing by collaborating with one or more edge node devices, providing subsequent localized control, configuration, management, and more for end devices. In addition, fog computing provides the ability for edge resources to identify similar resources and collaborate to create an edge-local cloud, which can be used alone or in conjunction with cloud computing to complete computing, storage, or connectivity-related services. Fog computing can also allow cloud-based services to extend their coverage to the edge of the device network to provide local and faster access to edge devices. Therefore, some forms of fog computing provide operations consistent with the edge computing discussed in this article; the edge computing aspects discussed in this article also apply to fog networks, fogging, and fog configurations. In addition, the aspects of the edge computing system discussed in this article can be configured as fog, or the fog aspects can be integrated into the edge computing architecture.

[0081] The core data center 1032 is located at the core network layer (regional or geographic center level), while the global network cloud 1042 is located at the cloud data center layer (national or global layer). The use of "core" is provided as a term for a centralized network location (deeper in the network) that is accessible by multiple edge nodes or components; however, "core" does not necessarily specify the "center" or deepest location of the network. Therefore, the core data center 1032 can be located within, at, or near the edge cloud 1000. Although the illustrative number of client computing nodes 1002, edge gateway nodes 1012, edge aggregation nodes 1022, edge core data centers 1032, global network cloud 1042 are Figure 10, it should be appreciated that the edge computing system 1000 may include additional devices or systems at each layer. Devices at any layer may be configured as peer nodes to each other and thus act in a collaborative manner to meet service goals.

[0082] Consistent with the examples provided herein, client computing node 1002 may be embodied as any type of endpoint component, device, apparatus, or other thing capable of communicating as a producer or consumer of data. Furthermore, the labels "node" or "device" used in edge computing system 1000 do not necessarily imply that such a node or device operates in the role of a client or agent / slave / follower; rather, any of the nodes or devices in edge computing system 1000 refers to an individual entity, node, or subsystem comprising discrete or connected hardware or software configurations to facilitate or use edge cloud 1000.

[0083] Thus, the edge cloud 1000 is formed of network components and functional features that are operated by and within the edge gateway node 1012 and the edge aggregation node 1022. The edge cloud 1000 may be embodied as any type of network that provides edge computing and / or storage resources located close to endpoint devices (e.g., mobile computing devices, IoT devices, smart devices, etc.) with radio access network (RAN) capabilities. Figure 10 1002. In other words, edge cloud 1000 can be envisioned as an “edge” that connects endpoint devices and traditional network access points, which serves as an entry point into a service provider core network (including a mobile carrier network (e.g., a Global System for Mobile Communications (GSM) network, a Long Term Evolution (LTE) network, a 5G / 6G network, etc.)) while also providing storage and / or computing capabilities. Other types and forms of network access (e.g., Wi-Fi, long-range wireless, wired networks including optical networks) can also be used in place of such 3GPP carrier networks or in combination with such 3GPP carrier networks.

[0084] In some examples, the edge cloud 1000 can form part of, or otherwise provide an entry point into or across, a fog networking configuration 1026 (e.g., a network of fog devices 1024, not shown in detail), which can be embodied as a system-level horizontal distributed architecture that distributes resources and services to perform specific functions. For example, a coordinated distributed network of fog devices 1024 can perform compute, storage, control, or networking aspects in the context of an IoT system arrangement. Other networked, aggregated, and distributed functions can exist in the edge cloud 1000 between the core data center 1032 and client endpoints (e.g., client compute nodes 1002). The following sections will discuss some of these in the context of network function or service virtualization, including the use of virtual edges and virtual services orchestrated for multiple stakeholders.

[0085] As discussed in more detail below, the edge gateway nodes 1012 and edge aggregation nodes 1022 collaborate to provide various edge services and security to the client computing nodes 1002. Furthermore, because the client computing nodes 1002 can be stationary or mobile, the corresponding edge gateway nodes 1012 can collaborate with other edge gateway devices to propagate currently provided edge services, related service data, and security as the corresponding client computing nodes 1002 move within a region. To this end, the edge gateway nodes 1012 and / or edge aggregation nodes 1022 can support multi-tenant and multi-stakeholder configurations, in which services from (or hosted for) multiple service providers, owners, and multiple consumers can be supported and coordinated across a single or multiple computing devices.

[0086] Various security approaches can be utilized within the architecture of the edge cloud 1000. In a multi-stakeholder environment, there can be multiple loadable security modules (LSMs) that provide policies that implement the interests of the stakeholders. The enforcement point environment can support multiple LSMs that apply a combination of loaded LSM policies (e.g., where the most constrained effective policy is applied, such as if any of stakeholder A, stakeholder B, or stakeholder C restricts access, then access is restricted). Within the edge cloud 1000, each edge entity can provide an LSM that implements the interests of the edge entity. A cloud entity can provide an LSM that implements the interests of a cloud entity. Similarly, various fog and IoT network entities can provide LSMs that implement the interests of fog entities.

[0087] In these examples, services can be considered from the perspective of transactions performed against a set of contracts or elements, whether considered at the element level or at a human-perceivable level. Thus, a user who has a service agreement with a service provider expects the service to be delivered under the terms of the SLA. Although not discussed in detail, the use of edge computing technologies discussed herein can play a role during the negotiation of agreements and the measurement of agreement performance (to identify what elements the system requires to perform the service, how the system responds to service conditions and changes, etc.).

[0088] "Services" is a broad term often applied in a variety of contexts, but generally refers to a relationship between two entities where one entity provides and performs work for the benefit of another. However, services delivered from one entity to another must be executed using certain guidelines that ensure trust between the entities and govern transactions according to contractual terms and conditions spelled out at the beginning, during, and end of the service.

[0089] The deployment of a multi-stakeholder edge computing system can be arranged and orchestrated to enable the deployment of multiple services and virtual edge instances across multiple edge nodes and subsystems for use by multiple tenants and service providers. In an example system applicable to a cloud service provider (CSP), the deployment of an edge computing system can be provided via an "over-the-top" approach to introduce edge computing nodes as a complementary tool to cloud computing. In a contrasting example system applicable to a telecommunications service provider (TSP), the deployment of an edge computing system can be provided via a "network aggregation" approach to introduce edge computing nodes at locations where network access (from different types of data access networks) is aggregated. Figure 9 and Figure 10 These over-the-top and network aggregation approaches for networking and services in corresponding edge computing systems are contrasted. However, these over-the-top and network aggregation approaches can be implemented together in a hybrid or combined approach or configuration, as suggested in later examples.

[0090] Figure 11An example is shown in which various client endpoints 1110 (in the form of mobile devices, computers, autonomous vehicles, commercial computing equipment, industrial processing equipment) provide requests 1120 for services or data transactions to an edge cloud 1100 (e.g., via a wireless or wired network 1140) and receive responses 1130 for the services or data transactions from the edge cloud 1100. Within the edge cloud 1000, the CSP can deploy various computing and storage resources (e.g., edge content nodes 1150) to provide cached content from a distributed content delivery network. Other available computing and storage resources available on the edge content nodes 1150 can be used to perform other services and complete other workloads. The edge content nodes 1150 and other systems of the edge cloud 1000 are connected to a cloud or data center 1170, which uses a backhaul network 1160 to facilitate higher latency requests from the cloud / data center to websites, applications, database servers, etc.

[0091] Figure 12 It is a combination Figure 9 A block diagram of an embodiment of a system 910 is described. Figure 12 In the illustrated embodiment, the system is a data center 1200. Data center 1200 includes a plurality of communicatively coupled nodes, including three compute nodes (compute node 0 1220, compute node 1 1222, and compute node 2 1224) that communicate via a CXL switch 1212. Data center 1200 also includes an orchestrator 1202 and dynamic capacity devices (dynamic capacity device 0 1214, dynamic capacity device 1 1216, and dynamic capacity device 2 1218), also communicating via CXL switch 1212. In one embodiment, the dynamic capacity devices may be storage nodes with volatile memory. In another embodiment, the dynamic capacity devices may be storage nodes with non-volatile memory.

[0092] CXL switch 1212 is a communication link based on the CXL standard. CXL is an open standard interconnect based on the PCI Express (PCIe) 5.0 physical layer infrastructure that provides high-performance connectivity between one or more host processors and other devices. The CXL standard includes three protocols: CXL.io, CXL.cache, and CXL.mem (CXL.memory). Compute nodes 1220, 1222, 1224 can use CXL memory (CXL.mem devices) or can be coherent (CXL.cache). The CXL.mem protocol allows compute nodes 1220, 1222, 1224 to directly access memory attached to other CXL devices in a cache-coherent manner. The CXL.cache protocol allows connected devices to cache data.

[0093] Each of compute nodes 1220, 1222, and 1224 is coupled to a corresponding memory node. The memory nodes are CXL.mem devices. Compute node 0 1220 is coupled to memory node 1232, memory node 1234, and memory node 1236. Compute node 1 1222 is coupled to memory node 1240, memory node 1242, and memory node 1244. Compute node 2 1224 is coupled to memory node 1250, memory node 1252, and memory node 1254. Each of compute nodes 1220, 1222, and 1224 includes local memory. Compute node 0 1220 includes local memory 1226. Compute node 1 1222 includes local memory 1228. Compute node 2 1224 includes local memory 1230.

[0094] Orchestrator 1202 includes an operating system (OS) 1204, a system basic input / output system (BIOS) 1206, and a baseboard management controller (BMC) 1210. System BIOS 1206 is firmware that provides runtime services for operating system (OS) 1204 and programs and performs hardware initialization during power-on startup of data center 1200. Baseboard management controller 1210 manages the interface between system management software and platform hardware.

[0095] System BIOS 1206 includes table 1208. In an embodiment, table 1208 is an Advanced Configuration and Power Interface (ACPI) table. ACPI is a standard that can be used by operating systems in a system to discover and configure computer hardware components, perform power management, auto-configuration, and monitor status.

[0096] In an embodiment, table 1208 comprises an ACPI table. Table 1208 comprises an ACPI system resource affinity table (SRAT) (also referred to as a memory affinity structure). The SRAT provides memory topology information to the operating system. The topology information includes the association between a memory range and its adjacent domains and information about whether the memory range can be hot-plugged. The system BIOS 1206 identifies a set of potential memory configurations for the mapped memory range (e.g., alternative memory configurations for the mapped memory range) and populates table 1208. Figure 13 Describe SRAT.

[0097] Table 1208 also includes an ACPI System Resource Affinity Table (SRAT) Flags Table that includes definitions of bits in the Flags field in the ACPI System Resource Affinity Table (SRAT). Figure 14 Describes the ACPI System Resource Affinity Table (SRAT) flag table.

[0098] Figure 13 Shown Figure 12

[0066] An embodiment of an Advanced Configuration and Power Interface (ACPI) System Resource Affinity Table (SRAT) 1300 in a table in FIG. The ACPI SRAT 1300 includes a plurality of fields. Each field is one or more bytes. The number of bytes in each field is stored as a byte length.

[0099] A one-byte type field 1302 identifies the ACPI SRAT 1300 as a memory affinity structure.

[0100] The one-byte length field 1304 stores 40, the number of bytes in the ACPI SRAT 1300.

[0101] The four-byte neighborhood field 1306 stores a four-byte integer representing the neighborhood to which the memory range belongs.

[0102] The 4-byte base address low field 1310 stores the lower 32 bits of the base address of the memory range.

[0103] The 4-byte base address high field 1312 stores the upper 32 bits of the base address of the memory range.

[0104] The 4-byte Length Low field 1314 stores the low 32 bits of the length of the memory range.

[0105] The 4-byte Length High field 1316 stores the high 32 bits of the length of the memory range.

[0106] The 4-byte flag field 1320 stores a flag for the memory affinity structure. This flag indicates whether the memory region defined in the system resource affinity table (SRAT) is enabled and can be hot-plugged. Figure 14 Describes the flags field.

[0107] There are three reserved fields, namely, a 2-byte reserved field 1308 , a 4-byte reserved field 1318 , and an 8-byte reserved field 1322 .

[0108] Figure 14 Shown Figure 13 An embodiment of a flags field 1320 in the ACPI system resource affinity table (SRAT) 1300 in the SRAT.

[0109] A 1-bit enable field 1402 allows system firmware to populate the ACPI System Resource Affinity Table (SRAT) 1300 with a static number of structures that can be enabled when needed. If the enable field 1402 is set to logic "0," OSPM ignores the contents of the memory affinity structures in the ACPI SRAT 1300.

[0110] The 1-bit hot-pluggable field 1404 is used in conjunction with the 1-bit enable field 1402. If the enable field 1402 is set to logic "1" and the 1-bit hot-pluggable field 1404 is also set to logic "1," the system hardware supports hot-addition and hot-removal of the memory region defined in the SRAT in the ACPI System Resource Affinity Table (SRAT) 1300. If the enable field 1402 is set to logic "1" and the 1-bit hot-pluggable field 1404 is also cleared to logic "0," the system hardware does not support hot-addition or hot-removal. If the enable field 1402 is cleared to logic "0" and the 1-bit hot-pluggable field 1404 is also cleared to logic "0," OSPM ignores the contents of the SRAT in the ACPI System Resource Affinity Table (SRAT) 1300.

[0111] The state of the 1-bit non-volatile field 1406 indicates the type of system memory. If set to logic "1," the system memory is non-volatile memory.

[0112] The state of the 1-bit reconfigurable field 1408 indicates whether the memory region is supported by the system BIOS. If set to logic "1," the memory region represents an alternate potential memory configuration supported by the system BIOS 1206.

[0113] The reserved field 1410 is a 28-bit field that is unused and is cleared to logic "0."

[0114] Figure 15 is a flow chart illustrating offloading system memory reconfiguration to the operating system 1204 using the ACPI System Resource Affinity Table (SRAT) with the assistance of the system BIOS 1206 .

[0115] At block 1500, upon system power-up, the system BIOS 1206 identifies a set of alternative potential memory configurations for the mapped memory range. The system BIOS 1206 populates the table 1208 with the combination Figure 13 and Figure 14The ACPI SRAT table is discussed to describe the memory configuration to the OS 1204 to allow the OS 1204 to read the mapped memory ranges stored in table 1208. The memory configuration includes the type of memory (e.g., DRAM, CXL, HBM), the number of memory devices, the configuration of the memory devices, and how the memory is reported to the OS 1204.

[0116] In an embodiment, enumeration of alternative potential memory configurations is performed during each system power-up. The mapped memory range can be a single level of memory (1LM) using DRAM (e.g., DDR5 DRAM) or CXL volatile memory, or it can be two levels of memory (2LM) or CXL persistent memory (PMEM), or a combination of DRAM and PMEM.

[0117] The system BIOS 1206 creates additional memory affinity structures in the System Resource Affinity Table (SRAT). The additional memory affinity structures are programmed by the system BIOS 1206 so that they can be read by the OS 1204. To create the additional memory affinity structures, the system BIOS 1206 first identifies possible alternative memory configurations and then creates an (enhanced) SRAT to describe the alternative memory configurations.

[0118] The alternate memory configuration is identified via the Flags field 1320 in the ACPI SRAT 1300. The Enable field 1402 (bit 0) in the Flags field 1320 is cleared (cleared to logic "0"), and the Reconfigurable field 1408 (bit 3) in the Flags field is set (set to logic "1"). Operating system-directed power management (OSPM) 1260 in the OS 1204 not only ignores the memory affinity structure based solely on the state of the Enable bit field. The system BIOS 1206 also sets the Reconfigurable field 1408 for the enabled range corresponding to the alternate decoding configuration. The system BIOS 1206 creates additional structures in the Heterogeneous Memory Attribute Table (HMAT) in table 1208 to describe the alternate memory proximity domains listed in the SRAT. The system BIOS 1206 creates additional entries in the System Local Information Table (SLIT) in table 1208 to describe the alternate memory proximity domains listed in the SRAT. The additional structures include performance information, such as memory bandwidth and latency, for the alternate proximity domains. Processing continues to block 1502 .

[0119] At block 1502, during runtime, the OS determines if / when one or more of the alternative memory configurations described in ACPI are advantageous to the kernel 1262. The OS 1204 prepares a list of memory contiguous domain identifiers (IDs) for these configurations. These configurations are identified by contiguous domains.

[0120] The OS kernel is the agent responsible for scheduling jobs on the system and can determine when the load on the system is relatively low. The OS kernel can use this knowledge to choose to trigger memory reconfiguration during periods of low utilization, thereby minimizing any performance overhead of the operation compared to other possible performance overhead. After determining to trigger memory reconfiguration, processing continues to block 1504.

[0121] At block 1504, OS 1204 prepares to migrate all activity / data out of the affected memory range. The affected memory range is the memory range that is subject to data loss due to the reconfiguration process. To ensure that data stored in the affected memory range is not lost during the reconfiguration, the data is saved to another memory range before the reconfiguration process begins. OS 1204 identifies the operations to be performed to safely reconfigure memory without losing any data from the ACPI data provided by the BIOS in table 1208. OS 1204 suspends all memory accesses to the memory range being reconfigured that is associated with the affected memory range. OS 1204 relocates all required data from the affected memory range in one or more memory nodes (e.g., memory nodes 1232, 1234, 1236) to an auxiliary device, such as a CXL dynamic capacity device (CXL DCD) (e.g., dynamic capacity device 1216). Processing continues to block 1506.

[0122] At block 1506, the OS 1204 initiates a system BIOS reconfiguration request to the memory configuration identified by the memory contiguous domain ID list selected by the OS 1204. The system BIOS 1206 uses a security agent to calculate and program the memory decode registers to enable the memory configuration identified by the selected memory contiguous domain ID. The security agent is used because the memory decoder is locked before the OS boots, so that malicious Ring 0 software (software with the highest privileges) cannot reprogram the memory decode registers. Additionally, the security agent ensures that the memory decoder is programmed to avoid security issues (e.g., memory aliasing).

[0123] In one embodiment, the system management mode (SMM) in compute node 0 1220, compute node 1 1222, and compute node 2 1224 is used to perform the memory reconfiguration. In another embodiment, the baseboard management controller (BMC) 1210 is used to perform the memory reconfiguration. When the OS 1204 requests the system BIOS 1206 to perform a remapping, a_DSM (ACPI device specific method) is called in the system BIOS 1206. For example, the remapping can be switching from a single level of memory (flat memory) to a two-level memory (hierarchical memory, where the first level of memory serves as a cache for the second level of memory). Processing continues at block 1508.

[0124] At block 1508, the system BIOS 1206 communicates a request to perform a remapping to the baseboard management controller 1210 via the DSM to reprogram the system address map in compute node 0 1220, compute node 1 1222, and compute node 2 1224. In response to receiving the request to perform a remapping, the baseboard management controller (BMC) firmware 1264 performs the memory remapping by reprogramming the system address map accordingly. After the system address map has been reprogrammed by the BMC firmware 1264, control returns to the OS 1204. Processing continues at block 1510.

[0125] At block 1510, the OS restores all relocated data from the auxiliary devices (e.g., CXL DCDs (e.g., dynamic capacity device 1216)) to the newly configured DDR memory ranges (e.g., the relocated data stored in dynamic capacity device 1216 may be restored to the newly configured DDR memory ranges in memory node 1250, memory node 1252, and memory node 1254). The OS 1204 then restarts all suspended activities and continues with the new memory configuration.

[0126] Use combination Figure 15 The described ACPI System Resource Affinity Table (SRAT) offloads system memory reconfiguration to the operating system with assistance from the platform / firmware, enabling the OS 1204 to adapt the platform memory configuration to dynamically changing workload conditions in the data center in an architecture-agnostic manner. For example, if the nature of the incoming workload changes over time, the kernel can reconfigure memory to extract better performance from existing resources. Avoiding reboots improves CSP server uptime requirements and reduces operator intervention.

[0127] Embodiments for reconfiguration of system memory in a data server have been described.In other embodiments, input / output (I / O) devices or accelerator devices may be reconfigured.

[0128] It is contemplated that aspects of the embodiments herein may be implemented in various types of computing and networking equipment, such as switches, routers, and blade servers, such as those employed in data center and / or server farm environments. Servers used in data centers and server farms include arrayed server configurations, such as rack-based servers or blade servers. These servers are communicatively interconnected via various network configurations, such as by partitioning groups of servers into local area networks (LANs) with appropriate switching and routing facilities between the LANs to form a private intranet. For example, a cloud hosting facility may typically employ a large data center with a large number of servers.

[0129] Each blade comprises a separate computing platform configured to perform server-type functions, i.e., a "server on a card." Thus, each blade includes components common to conventional servers, including a main printed circuit board (motherboard) that provides internal wiring (i.e., buses) for coupling appropriate integrated circuits (ICs) and other components mounted on the board. These components may include the previously described Figure 1 Components discussed.

[0130] The flow charts shown herein provide examples of sequences of various processing actions.

[0131] The flowchart may indicate operations to be performed by software or firmware routines as well as physical operations. In one embodiment, the flowchart may illustrate the states of a finite state machine (FSM) that may be implemented in hardware and / or software. Although shown in a particular sequence or order, the order of the actions may be modified unless otherwise stated. Therefore, the illustrated embodiments should be understood as examples only, and the process may be performed in a different order, and some actions may be performed in parallel. Additionally, one or more actions may be omitted in various embodiments; therefore, not all actions are required in every embodiment. Other processing flows are also possible.

[0132] With respect to the various operations or functions described herein, these operations or functions may be described or defined as software code, instructions, configurations and / or data. The content may be directly executable (in the form of an "object" or "executable file"), source code, or differential code (incremental or "patch" code). The software content of the embodiments described herein may be provided via an article having content stored thereon, or via a method of operating a communication interface to send data via a communication interface. A non-transitory machine-readable storage medium (computer-readable storage medium) may enable a machine to perform the functions or operations described, and includes any mechanism for storing information in a form accessible to a machine (e.g., a computing device, an electronic system, etc.), such as recordable / non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), disk storage media, optical storage media, flash memory devices, etc.). A communication interface includes any mechanism that interfaces with any medium, such as a hardwired medium, a wireless medium, an optical medium, etc., to communicate with another device, such as a memory bus interface, a processor bus interface, an Internet connection, a disk controller, etc. The communication interface may be configured by providing configuration parameters and / or sending signals to prepare the communication interface to provide data signals describing the software content. The communication interface may be accessed via one or more commands or signals sent to the communication interface.

[0133] The various components described herein may be units for performing the described operations or functions. Each component described herein includes software, hardware, or a combination of these software and hardware. Components may be implemented as software modules, hardware modules, dedicated hardware (e.g., dedicated hardware, application specific integrated circuits (ASICs), digital signal processors (DSPs), etc.), embedded controllers, hard-wired circuit systems, etc.

[0134] Besides what is described herein, various modifications may be made to the disclosed embodiments and implementations without departing from the scope of these embodiments and implementations.

[0135] The specification and examples herein are, therefore, to be construed in an illustrative rather than a restrictive sense.The scope of the present invention should be measured solely by reference to the appended claims.

[0136] Example

[0137] Illustrative examples of the technology disclosed herein are provided below. Embodiments of the technology may include any one or more of the examples described below, and any combination of the examples described below.

[0138] Example 1 is an orchestrator comprising circuitry configured to: store an alternative memory configuration for a mapped memory range; upon determining that memory reconfiguration is triggered, migrate data from the mapped memory range to an auxiliary device; perform memory remapping for the alternative memory configuration; and restore the data from the auxiliary device to the mapped memory range.

[0139] Example 2 includes the organizer of Example 1, optionally wherein the alternate memory configuration is stored in a table in a system basic input / output system (BIOS).

[0140] Example 3 includes the orchestrator of Example 1, optionally wherein the memory remapping is performed by baseboard management controller (BMC) firmware.

[0141] Example 4 includes the coordinator of Example 1, optionally wherein the auxiliary device is a dynamic capacity device.

[0142] Example 5 includes the orchestrator of Example 4, optionally wherein the dynamic capacity device is a volatile memory.

[0143] Example 6 includes the orchestrator of Example 1, optionally wherein the mapped memory range is a region of memory defined in a system resource affinity table.

[0144] Example 7 includes the orchestrator of Example 4, optionally wherein the dynamic capacity device is a CXL.mem device.

[0145] Example 8 is a system comprising an auxiliary device and circuitry. The circuitry is configured to: store an alternative memory configuration for a mapped memory range; upon determining that memory reconfiguration is triggered, migrate data from the mapped memory range to the auxiliary device; perform memory remapping for the alternative memory configuration; and restore the data from the auxiliary device to the mapped memory range.

[0146] Example 9 includes the system of Example 8, optionally wherein the alternate memory configuration is stored in a table in a system basic input / output system (BIOS).

[0147] Example 10 includes the system of Example 8, optionally wherein the memory remapping is performed by baseboard management controller (BMC) firmware.

[0148] Example 11 includes the system of Example 8, optionally wherein the auxiliary device is a dynamic capacity device.

[0149] Example 12 includes the system of Example 11, optionally wherein the dynamic capacity device is a volatile memory.

[0150] Example 13 includes the system of Example 8, optionally wherein the mapped memory range is a region of memory defined in a system resource affinity table.

[0151] Example 14 includes the system of Example 11, optionally wherein the dynamic capacity device is a CXL.mem device.

[0152] Example 15 is a method comprising: storing an alternative memory configuration for a mapped memory range. The method further comprises: migrating data from the mapped memory range to a secondary device upon determining that memory reconfiguration is triggered. The method further comprises: performing memory remapping for the alternative memory configuration. The method further comprises: restoring the data from the secondary device to the mapped memory range.

[0153] Example 16 includes the method of Example 15, optionally wherein the alternate memory configuration is stored in a table in a system basic input / output system (BIOS).

[0154] Example 17 includes the method of Example 15, optionally wherein the memory remapping is performed by baseboard management controller (BMC) firmware.

[0155] Example 18 includes the method of Example 15, optionally wherein the auxiliary device is a dynamic capacity device.

[0156] Example 19 includes the method of Example 15, optionally wherein the mapped memory range is a region of memory defined in a system resource affinity table.

[0157] Example 20 includes the method of Example 18, optionally wherein the dynamic capacity device is a CXL.mem device.

[0158] Example 21 is at least one machine-readable medium comprising a plurality of instructions that, in response to being executed by a system, may cause the system to perform the method according to any one of Examples 15 to 20.

[0159] Example 22 is an apparatus comprising means for performing the method according to any one of Examples 15 to 20.

Claims

1. A node in a server system, comprising: network interface circuitry for coupling to a system storage device and to an auxiliary memory device; as well as A processor device configured to execute the orchestrator to perform the following operations: storing an alternative memory configuration for the mapped memory range of the system memory device; upon determining that memory reconfiguration is triggered, migrating data from the mapped memory range to the auxiliary memory device; performing a memory remapping of the system memory device for the alternate memory configuration; as well as The data is restored from the auxiliary storage device to the mapped memory range of the system storage device.

2. The node according to claim 1, wherein: The alternate memory configuration is stored in a table in the system basic input / output system (BIOS).

3. The node according to claim 1 or 2, wherein: The memory remapping is performed by baseboard management controller (BMC) firmware.

4. The node according to any one of claims 1 to 3, wherein: The secondary storage device is a dynamic capacity device.

5. The node according to claim 4, wherein: The dynamic capacity device is a volatile memory.

6. The node according to claim 4, wherein: The dynamic capacity device is a CXL.mem device.

7. The node according to any one of claims 1 to 6, wherein: The mapped memory range is the area of ​​memory defined in the system resource affinity table.

8. A server system comprising: System storage devices; Auxiliary storage devices; as well as An orchestrator circuit system configured to: storing an alternative memory configuration for the mapped memory range of the system memory device; upon determining that memory reconfiguration is triggered, migrating data from the mapped memory range to the auxiliary memory device; performing a memory remapping of the system memory device for the alternate memory configuration; as well as The data is restored from the auxiliary storage device to the mapped memory range of the system storage device.

9. The server system according to claim 8, wherein: The alternate memory configuration is stored in a table in the system basic input / output system (BIOS).

10. The server system according to claim 8 or 9, wherein: The memory remapping is performed by baseboard management controller (BMC) firmware.

11. The server system according to any one of claims 8 to 10, wherein: The secondary storage device is a dynamic capacity device.

12. The server system according to claim 11, wherein: The dynamic capacity device is a volatile memory.

13. The server system according to claim 11, wherein: The dynamic capacity device is a CXL.mem device.

14. The server system according to any one of claims 8 to 13, wherein: The mapped memory range is the area of ​​memory defined in the system resource affinity table.

15. A method for memory reconfiguration, comprising: storing an alternative memory configuration for the mapped memory range of the system memory device; upon determining that memory reconfiguration is triggered, migrating data from the mapped memory range to the auxiliary memory device; performing a memory remapping of the system memory device for the alternate memory configuration; as well as The data is restored from the auxiliary storage device to the mapped memory range of the system storage device.

16. The method according to claim 15, wherein The alternate memory configuration is stored in a table in the system basic input / output system (BIOS).

17. The method according to claim 15 or 16, wherein The memory remapping is performed by baseboard management controller (BMC) firmware.

18. The method according to any one of claims 15 to 17, wherein: The secondary storage device is a dynamic capacity device.

19. The method according to any one of claims 15 to 18, wherein The mapped memory range is the area of ​​memory defined in the system resource affinity table.

20. The method according to claim 18, wherein The dynamic capacity device is a CXL.mem device.