Memory Allocation and Memory Write Redirection in a Cloud Computing System Based on the Temperature of a Memory Module

By maintaining temperature profiles and redirecting write requests to cooler memory modules, the method addresses temperature-related failures in cloud computing systems, enhancing reliability and uptime.

JP7704382B2Active Publication Date: 2025-07-08MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022578980
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-29
Filing Date
2021-04-20
Publication Date
2025-07-08
Estimated Expiration
2041-04-20

AI Technical Summary

Technical Problem

There is a need for improving the reliability of memory modules in cloud computing systems by addressing temperature-related malfunctions and failures due to heat and overuse.

Method used

Implementing a method and system that maintains temperature profiles of memory chips using sensors and redirects write requests to memory modules that do not exceed a temperature threshold, and migrates computing entities to cooler memory modules when necessary, utilizing hypervisors to manage memory allocation and temperature profiles.

Benefits of technology

Enhances the reliability and uptime of memory modules by preventing overheating and ensuring efficient memory usage, thereby improving the overall performance and stability of cloud computing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704382000001
    Figure 0007704382000001
  • Figure 0007704382000002
    Figure 0007704382000002
  • Figure 0007704382000003
    Figure 0007704382000003
Patent Text Reader

Abstract

A system and method for allocating memory and redirecting data writes based on temperatures of memory modules in a cloud computing system are described. The method includes maintaining a temperature profile of a plurality of first memory modules and a plurality of second memory modules. The method includes automatically redirecting a first write request to memory from a first computing entity running on a first processor to a selected one of a plurality of first memory chips included in at least the plurality of first memory modules, the temperature of which does not meet or exceed a temperature threshold, and automatically redirecting a first write request to memory from a second computing entity running on a second processor to a selected one of a plurality of second memory chips included in at least the plurality of second memory modules, the temperature of which does not meet or exceed a temperature threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Computing, storage, and network resources are increasingly being accessed via public clouds, private clouds, or two hybrids. A public cloud includes a global network of servers that perform various functions, including delivery of content or services such as data storage and management, application execution, video streaming, email, office productivity software, or social media. Servers and other components can be located in data centers around the world. A public cloud generally provides services via the Internet, but businesses can use private clouds or hybrid clouds. Both private and hybrid clouds also include a network of servers housed in a data center. A cloud service provider provides access to these resources by providing cloud computing and storage resources to customers.

[0002] There is a need for methods and systems for improving the reliability of memory modules used in cloud computing systems.

Summary of the Invention

[0003] One aspect of the present disclosure relates to a method in a cloud computing system including a host server. Here, the host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor. The method may include maintaining a first temperature profile based on information received from temperature sensors associated with respective ones of a plurality of first memory chips included in at least one of the plurality of first memory modules. The method may further include maintaining a second temperature profile based on information received from temperature sensors associated with respective ones of a plurality of second memory chips included in at least the plurality of second memory modules. The method may further include automatically redirecting a first write request to memory from a first computing entity being executed by the first processor to a selected one of a plurality of first memory chips included in at least the plurality of first memory modules, the temperature of which does not meet or exceed a temperature threshold, based on at least the first temperature profile. The method may further include automatically redirecting a second write request to memory from a second computing entity being executed by the second processor to a selected one of a plurality of second memory chips included in at least the plurality of second memory modules, the temperature of which does not meet or exceed a temperature threshold, based on at least the second temperature profile.

[0004] In yet another aspect, the present disclosure relates to a system that includes a host server. The host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor. The system further includes a hypervisor associated with the host server. The hypervisor is configured to: (1) maintain a first temperature profile based on information received from temperature sensors associated with respective ones of a plurality of first memory chips included in at least one of the plurality of first memory modules; (2) maintain a second temperature profile based on information received from temperature sensors associated with respective ones of a plurality of second memory chips included in at least the plurality of second memory modules; (3) automatically redirect a first write request to memory from a first computing entity being executed by the first processor to a selected one of a plurality of first memory chips included in at least the plurality of first memory modules, where the temperature of the selected first memory chip does not meet or exceed a temperature threshold; and (4) automatically redirect a second write request to memory from a second computing entity being executed by the second processor to a selected one of the plurality of second memory chips included in at least the plurality of second memory modules, where the temperature of the selected second memory chip does not meet or exceed a temperature threshold, based at least on the second temperature profile.

[0005] In another aspect, the present disclosure relates to a method in a cloud computing system including a first host server and a second host server. Here, the first host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor, and the second host server includes at least a plurality of third memory modules coupled to a third processor and at least a plurality of fourth memory modules coupled to a fourth processor. Here, the first host server includes a first hypervisor that manages a plurality of first computing entities for execution by the first processor or the second processor, and the second host server includes a second hypervisor that manages a plurality of second computing entities for execution by the third processor or the fourth processor. The method may include maintaining a first temperature profile based on information received from temperature sensors associated with each of a plurality of first memory chips included in at least the plurality of first memory modules and at least the plurality of second memory modules. The method may further include maintaining a second temperature profile based on information received from temperature sensors associated with each of a plurality of second memory chips included in at least the plurality of third memory modules and at least the plurality of fourth memory modules. The method may further include automatically redirecting a first write request to memory from a first computing entity being executed by the first processor to a selected one of a plurality of first memory chips included in at least the plurality of first memory modules and at least the plurality of second memory modules, the temperature of which does not meet or exceed a temperature threshold, based on at least the first temperature profile. The method may further include automatically redirecting a second write request to memory from a second computing entity being executed by the second processor to a selected one of a plurality of second memory chips included in at least the plurality of third memory modules and at least the plurality of fourth memory modules, the temperature of which does not meet or exceed a temperature threshold, based on at least the second temperature profile.This method may further include automatically migrating at least a subset of the first computing entities from the first host server to the second host server when it is determined that the temperatures of at least N of the plurality of first memory chips, where N is a positive integer, satisfy or exceed a temperature threshold and the temperature of at least one of the plurality of second memory chips does not satisfy or exceed the temperature threshold.

[0006] This summary is provided to introduce a selection of concepts in a simplified form that are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

Brief Description of the Drawings

[0007] The present disclosure is illustrated by way of example and is not limited by the accompanying drawings. In the drawings, like references indicate like elements. The elements in the drawings are illustrated for simplicity and clarity and are not necessarily drawn to scale.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

[0008] The embodiments described in this disclosure relate to allocating memory based on the temperature of a memory module of a cloud computing system and redirecting data writes. The memory module may be included in a host server. A plurality of host servers may be included in a server rack or stack of servers. The host server may be any server within a cloud computing environment configured to provide services to a tenant or other subscriber of cloud computing services. Examples of memory technologies include, but are not limited to, volatile memory technologies, non-volatile memory technologies, and quasi-volatile memory technologies. Examples of memory types include dynamic random access memory (DRAM), flash memory (e.g., NAND flash), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), phase change memory (PCM), and resistive random access memory (RRAM). Generally speaking, the present disclosure relates to improving the reliability and uptime of any server having memory based on technologies that are susceptible to malfunction or failure due to heat and overuse.

[0009] In certain examples, the methods and systems described herein can be deployed in a cloud computing environment. Cloud computing can be defined as a model that enables on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with little administrative effort or interaction with a service provider, and can then be scaled accordingly. The cloud computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model can be used to expose various service models such as, for example, hardware as a service ("HaaS"), software as a service ("SaaS"), platform as a service ("PaaS"), infrastructure as a service ("IaaS"). The cloud computing model can also be deployed using various deployment models such as, for example, private cloud, community cloud, public cloud, hybrid cloud, etc.

[0010] FIG. 1 shows a system 100 for controlling memory allocation and data writing in a cloud computing system according to one example. In this example, the system 100 may correspond to a cloud computing stack in a data center. The system 100 may be implemented as a rack of servers. In this example, the system 100 may include host servers 110, 120, and 130. Each host server may include one or more processors configured to provide computing functionality in at least some form. For example, host server 110 may include CPU-0 112 and CPU-1 114, host server 120 may include CPU-0 122 and CPU-1 124, and host server 130 may include CPU-0 132 and CPU-1 134. Host server 110 may further include memory modules 116 and 118. Host server 120 may further include memory modules 126 and 128. Host server 130 may further include memory modules 136 and 138.

[0011] Continuing to refer to FIG. 1, host server 110 may be configured to execute instructions corresponding to hypervisor 140. Hypervisor 140 may further be configured to interface with virtual machines (VMs) (e.g., VM142, VM144, and VM146). Instructions corresponding to the VMs may be executed using either CPU-0 112 or CPU-1 114 associated with host server 110. Hypervisor 150 may further be configured to interface with virtual machines (VMs) (e.g., VM152, VM154, and VM156). Instructions corresponding to the VMs may be executed using either CPU-0 122 or CPU-1 124 associated with host server 120. Hypervisor 160 may further be configured to interface with virtual machines (VMs) (e.g., VM162, VM164, and VM166). Instructions corresponding to the VMs may be executed using either CPU-0 132 or CPU-1 134 associated with host server 130.

[0012] The hypervisor 140 can share control information with the hypervisor 150 via a control path. The control path may correspond to a path implemented using a bus system (e.g., a server rack bus system or another type of bus system). The hypervisor 150 can share control information with the hypervisor 160 via another control path. The control path may correspond to a path implemented using a bus system (e.g., a server rack bus system or other types of bus systems). Each of the hypervisor 140, hypervisor 150, and hypervisor 160 can be a kernel-based virtual machine (KVM) hypervisor, a Hyper-V hypervisor, or another type of hypervisor. FIG. 1 shows the system 100 as including a predetermined number of components arranged and coupled in a predetermined manner, but may include fewer or additional components arranged and coupled in different ways. As one example, although not shown in FIG. 1, each host server may include an operating system for managing a predetermined aspect of the host server. As another example, the system 100 may include any number of host servers coupled as part of a rack or stack. As another example, each host server may include any number of CPUs, GPUs, memory modules, or other components, as needed, to provide cloud computing, storage, and / or network functions. Further, the functions associated with the system 100 can be distributed or combined as needed. Further, although FIG. 1 illustrates access to the memory of a host server by a VM, other types of computing entities, such as containers, microVMs, microservices, and unikernels of serverless functions, can access memory in a similar manner.As used herein, the term "compute entity" includes, without limitation, functionality, applications, services, microservices, containers, unikernels for serverless computing, or any executable code (in the form of hardware, firmware, software, or any combination of the foregoing) that implements a portion of the foregoing.

[0013] FIG. 2 shows a block diagram of a server 200 implementing a host server (e.g., any of host server 110, host server 120, or host server 130) according to one example. Server 200 may include server board 210 and server board 250. Server board 210 and server board 250 may be coupled via an interconnect, a high-speed cable, a bus system associated with a rack, or another structure housing server board 210 and server board 250. Server board 210 may include CPU-0 212, dual in-line memory modules (DIMMs) (e.g., the DIMMs shown in FIG. 2 as part of server board 210), and solid state drives (SSDs) 245, 247, and 249. In this example, CPU-0 212 may include 24 cores (identified as blocks having a capital C in FIG. 2). Each DIMM may be installed in a DIMM slot / connector. In this example, six DIMMs may be arranged on one side of CPU-0 212, and another six DIMMs may be arranged on the opposite side of CPU-0 212. In this example, DIMM0 222, DIMM1 224, DIMM2 226, DIMM3 228, DIMM4 230, and DIMM5 232 may be arranged on the left side of CPU-0 212. DIMM6 234, DIMM7 236, DIMM8 238, DIMM9 240, DIMM10 242, and DIMM11 244 may be arranged on the right side of CPU-0 212.

[0014] Continuing to refer to FIG. 2, server board 250 may include CPU-1 252, dual in-line memory modules (DIMMs) (e.g., the DIMMs shown in FIG. 2 as part of server board 250), and solid state drives (SSDs) 285, 287, and 289. In this example, CPU-1 252 may include 24 cores (identified as blocks having capital letter C in FIG. 2). Each DIMM may be installed in a DIMM slot / connector. In this example, six DIMMs may be arranged on one side of CPU-1 252, and another six DIMMs may be arranged on the opposite side of CPU-1 252. In this example, DIMM0 262, DIMM1 264, DIMM2 266, DIMM3 268, DIMM4 270, and DIMM5 272 may be arranged on the left side of CPU-1 252. DIMM6 274, DIMM7 276, DIMM8 278, DIMM9 280, DIMM10 282, and DIMM11 284 may be arranged on the right side of CPU-1 252.

[0015] Still referring to FIG. 2, in this example, the server 200 can be cooled using air. As one example, inlets 292, 294, and 296 can be used to supply cooling air to various parts of server board 210 and server board 250. In FIG. 2, cooling is performed using cooling air, but the server 200 can also be cooled using a liquid or other forms of substances. Regardless of the cooling method used, the various components incorporated in the server 200 can have different temperatures. There can be several reasons why the temperatures of the components are not uniform. As one example, a given component mounted on either server board 210 or 250 may generate more heat than other components. As another example, the cooling air received by some downstream components may be preheated by upstream components (e.g., the SSD shown in FIG. 2). In addition, due to the number and arrangement of the inlets, the temperature inside the server rack or another structure used to house server boards 210 and 250 may become non-uniform. The DIMMs may also have non-uniform temperatures. FIG. 2 shows the server 200 as including a given number of components arranged and coupled in a given manner, but it may include fewer or additional components arranged and coupled in different ways. As one example, the server 200 can include any number of server boards arranged inside a rack and other structures. As another example, each server board can include any number of CPUs, GPUs, memory modules, or other components as needed to provide computing, storage, and / or networking functions. In addition, although server boards 210 and 250 are described as having DIMMs, other types of memory modules may be included as part of server boards 210 and 250. As one example, such memory modules can be Single-Inline Memory Modules (SIMMs).

[0016] Figure 3 shows a host server 300 including a memory module 350 according to one example. The host server 300 may include a portion 310. In this example, the portion 310 may correspond to a part of the host server 300 that includes a central processing function and a bus / memory controller. As one example, the portion 310 may include a CPU 312, a cache 314, a Peripheral Component Interconnect Express (PCIe) controller 316, and a memory controller (MC) 318. The CPU 312 is coupled to the cache 314 via a bus 313 and can enable fast access to cached instructions or data. The CPU 312 may also be connected to the PCIe controller 316 via a bus 315. The CPU 312 may be connected to the memory controller 318 via a bus 317. The cache 314 may be connected to the memory controller 318 via a bus 319. In one example, the CPU 312, the cache 314, and the PCIe controller 316 can be incorporated within a single module (e.g., a CPU module).

[0017] The PCIe controller 316 may be connected to a PCIe bridge 320 via a PCIe bus 321. The PCIe bridge 320 may include a peer-to-peer (P2P) controller 322. The PCIe bridge 320 may also provide functions related to the PCIe controller and other functionality as needed and can enable an interface with various storage and network resources. In this example, the P2P controller 322 may be coupled to a P2P port including P2P 328, P2P 330, and P2P 332 via a bus 334. In this example, P2P 328 may be coupled to an SSD 340, P2P 330 may be coupled to an SSD 342, and P2P 332 may be coupled to an SSD 344.

[0018] Continuing to refer to FIG. 3, the memory controller 318 can be coupled to the memory module 350 via buses 352 and 354. In this example, the coupling to the memory module can be performed via advanced memory buffers (e.g., AMBs 362, 372, and 382). Additionally, in this example, bus 352 can transfer data / control / status signals from the memory controller 318 to the memory module 350, and bus 354 can transfer data / control / status signals from the memory module 350 to the memory controller 318. Additionally, a clock source 346 can be used to synchronize signals as needed. The clock source 346 can be implemented as a phase-locked loop (PLL) circuit or another type of clock circuit.

[0019] Memory modules 360, 370, and 380 may each be DIMMs, as described above. Memory module 360 may include memory chips 363, 364, 365, 366, 367, and 368. Memory module 360 may further include a memory module controller (MMC) 361. Memory module 370 may include memory chips 373, 374, 375, 376, 377, and 378. Memory module 370 may further include MMC 371. Memory module 380 may include memory chips 383, 384, 385, 386, 387, and 388. Memory module 380 may further include MMC 381. Each memory chip may include a temperature sensor (not shown) for continuously monitoring and tracking the temperature within the memory chip. Such temperature sensors can be implemented using semiconductor manufacturing techniques during the manufacture of the memory chip. In this example, each of MMCs 361, 371, and 381 may be coupled to memory controller 318 via bus 351. Each MMC is responsible, among other things, for collecting the temperature sensor values from each of its respective memory chips. Memory controller 318 can obtain temperature-related information from each of the respective MMCs corresponding to memory modules 360, 370, and 380. Alternatively, each of MMCs 361, 371, and 381 can periodically provide temperature-related information to memory controller 318 or another controller, which can then store the information in a manner accessible from a hypervisor associated with the host server.

[0020] Still referring to FIG. 3, in one example, the collected temperature values may be stored in a control / status register or other type of memory structure accessible to the CPU 312. In this example, the memory controller may maintain temperature profiles for the memory modules 360, 370, and 380. An example of a temperature profile may include information regarding the most recent measured temperature of each memory chip associated with each memory module. In one example, the hypervisor can control the scanning of the temperature profiles so that updated information can be accessed by the hypervisor periodically. The temperature profile may also include the relative difference in temperature compared to a baseline. Thus, in this example, a memory chip may have a temperature that is lower or higher than the baseline temperature. The relative temperature difference between memory chips may be 10 degrees Celsius or more. The CPU 312 may have access to temperature measurements associated with the memory chips and more detailed data as needed. The hypervisor associated with the host server can access the temperature profile for each memory module as part of a memory allocation decision or as part of redirecting writes to other physical memory locations.

[0021] Regarding access to memory (e.g., DIMM) associated with a host server, at a high level, there can be two ways for a computing entity (e.g., a virtual machine (VM)) to access the memory of the host server. In these instances, when the VM is accessing the physical memory associated with the CPU on which it is running, then the load or store access can be converted to a bus transaction by the hardware associated with the system. However, when the VM is provided access to physical memory associated with another CPU, then, in one example, the hypervisor can manage this by using a hardware exception that occurs due to an attempt to access an unmapped page. Each hypervisor can be permitted access to the host-side page table, or other memory mapping tables. Access to an unmapped page may cause a hardware exception such as a page fault. After moving the page to local memory associated with another host server, the hypervisor can access the host memory and install the page table mapping.

[0022] In one example, prior to such memory operations (or I / O operations) being performed, control information can be exchanged between host servers that are part of a server stack or group. The exchange of information can occur between hypervisors (e.g., the hypervisor shown in FIG. 1). To enable live migration, each host server can reserve a portion of the entire host memory to enable restart of the VMs within the host server. The host server may also be required to hold a certain amount of memory reserved for other purposes, including the overhead of the stack infrastructure and a resiliency reserve. This may be related to the memory reserved to enable migration of the VMs in case of server failure or other such problems that result in insufficient availability of another host server. Thus, at least a portion of the control information can be associated with each host server that designates a memory space that can be accessed by virtual machines being run by other host servers to enable live migration of the VMs. In one example, prior to initiating live migration, the hypervisor can determine that the temperature of at least a predetermined number of memory chips meets or exceeds a temperature threshold. If so determined, the hypervisor can automatically migrate at least a subset of the computing entities from that host server to another host server if the temperature of at least the memory associated with the host server does not meet or exceed the temperature threshold. As part of this process, apart from ensuring live migration to cooler DIMMs, the hypervisor can also ensure that sufficient physical memory exists within another host server to effectuate live migration.

[0023] Still referring to FIG. 3, in one example, as part of host server 300, loads and stores can be performed using Remote Direct Memory Access (RDMA). RDMA can enable direct copying of data from the memory of one system (e.g., host server 110 of FIG. 1) to the memory of another system (e.g., host server 120 of FIG. 1) without involvement of the operating system of either system. In this way, a host server supporting RDMA can achieve the zero-copy advantage by directly transferring data to or from the memory space of a process, which can eliminate the extra data copy between application memory and data buffers in the operating system. In other words, in this example, by using address translation / mapping across various software / hardware layers, only one copy of the data can be maintained in the memory (or I / O device) associated with the host server.

[0024] Continuing to refer to FIG. 3, the temperature profile of the SSD coupled via the PCIe bus can also be monitored as needed, and data corresponding to the VM can be stored on a cooler SSD. As described above, bus 321 can correspond to a PCIe bus that can function according to the PCIe specification and includes support for non-transparent bridging as needed. PCIe transactions can be routed using address routing, ID-based routing (e.g., using bus number, device number, and function number), or implicit routing using messages. Transaction types can include memory read / write, I / O read / write, configuration read / write, and transactions associated with messaging operations. The endpoints of the PCIe system can be configured using base address registers (BARs). The type of BAR can be configured as a BAR for memory operations or I / O operations. Other setup and configuration can also be performed as needed. The hardware associated with the PCIe system (e.g., any root complexes and ports) can further provide functionality that enables the performance of memory read / write operations and I / O operations. As one example, the address translation logic associated with the PCIe system can be used for address translation for packet processing, including packet transfer or packet drop.

[0025] In one example, the hypervisor running on host server 300 can map the memory area associated with the SSD associated with host server 300 to the guest address space of the virtual machine executed using CPU 312. When data loading is required by the VM, the load can be directly converted into a PCIe transaction. In the case of a store operation, PCIe controller 316 sends the PCIe packet to P2P controller 322, and the P2P controller can then send it to any of P2P ports 328, 330, or 332. In this way, data can be stored in an I / O device (e.g., SSD, HD, or other I / O device) associated with host server 300. The transfer may also include address translation by the PCIe system. FIG. 3 shows host server 300 as including a given number of components arranged and coupled in a given way, but the host server may include fewer or additional components arranged and coupled differently. Additionally, functions related to host server 300 can be distributed or combined as needed. As one example, FIG. 3 shows P2P ports to enable the performance of I / O operations, but other types of interconnects can also be used to enable such functionality. Alternatively, and / or additionally, any access operation to the SSD associated with the virtual machine executed by CPU 312 can be enabled using Remote Direct Memory Access (RDMA).

[0026] Figure 4 shows a system environment 400 for implementing a system and method according to one example. In this example, the system environment 400 may correspond to a part of a data center. As one example, the data center may include several clusters of racks containing platform hardware such as server nodes, storage nodes, network nodes, or other types of nodes. Server nodes may be connected to switches to form a network. The network can enable connections between each possible combination of switches. The system environment 400 may include Server 1 410 and Server N 430. The system environment 400 may further include data center-related functions 460 including deployment / monitoring 470, directory / identity services 472, load balancing 474, data center controller 476 (e.g., software-defined network (SDN) controller, and other controllers), router / switch 478. Server 1 410 may include a host processor 411, a host hypervisor 412, memory 413, a storage interface controller (SIC) 414, cooling 415 (e.g., cooling fan or other cooling device), a network interface controller (NIC) 416, and storage disks 417 and 418. Server N 430 may include a host processor 431, a host hypervisor 432, memory 433, a storage interface controller (SIC) 434, cooling 435 (e.g., cooling fan or other cooling device), a network interface controller (NIC) 436, and storage disks 437 and 438. Server 1 410 may be configured to support virtual machines including VM1 419, VM2 420, VMN 421. The virtual machines may further be configured to support applications such as APP1 422, APP2 423, APPN 424. Server N 430 may be configured to support virtual machines including VM1 439, VM2 440, and VMN 441.The virtual machine can be further configured to support applications such as APP1 442, APP2 443, APPN 444.

[0027] Continuing to refer to FIG. 4, in one example, a virtual extensible local area network (VXLAN) framework can be used to enable the system environment 400 for multiple tenants. Each virtual machine (VM) can be permitted to communicate with VMs within the same VXLAN segment. Each VXLAN segment can be identified by a VXLAN network identifier (VNI). FIG. 4 shows the system environment 400 as including a given number of components arranged and coupled in a given manner, but the system environment can include fewer or additional components arranged and coupled differently. Additionally, the functionality associated with the system environment 400 can be distributed or combined as needed. Further, although FIG. 4 depicts access by VMs to unused resources, other types of computing entities such as containers, micro-VMs, microservices, unikernels for serverless functions can access unused resources associated with the host server in a similar manner.

[0028] FIG. 5 shows a block diagram of a computing platform 500 (e.g., for implementing certain aspects related to the methods and algorithms associated with the present disclosure). The computing platform 500 may include a processor 502, I / O components 504, memory 506, a presentation component 508, sensors 510, a database 512, a network interface 514, and I / O ports. These may be interconnected via a bus 520. The processor 502 can execute instructions stored in the memory 506. The I / O components 504 may include user interface devices such as a keyboard, a mouse, a speech recognition processor, or a touch screen. The memory 506 can be any combination of non-volatile storage or volatile storage (e.g., flash memory, DRAM, SRAM, or other types of memory). The presentation component 508 can be any type of display such as an LCD, an LED, or other types of displays. The sensors 510 may include telemetry configured to detect and / or receive information (e.g., conditions related to the device), or other types of sensors. The sensors 510 may include sensors configured to sense conditions related to a CPU, memory or other storage components, FPGA, motherboard, baseboard management controller, etc. The sensors 510 may also include sensors configured to sense conditions related to a rack, chassis, fan, power supply unit (PSU), etc. The sensors 510 may also include sensors configured to sense conditions related to a network interface controller (NIC), top-of-rack (TOR) switch, middle-of-rack (MOR) switch, router, power distribution unit (PDU), rack-level uninterruptible power supply (UPS) system, etc.

[0029] Continuing to refer to FIG. 5, sensor 510 can be implemented in hardware, software, or a combination of hardware and software. Some sensors 510 can be implemented using a sensor API that enables the sensor 510 to receive information via the sensor API. Software configured to detect or listen for a given condition or event can communicate, via the sensor API, any conditions associated with a device that is part of a data center or other similar system. Remote sensors or other telemetry devices can be incorporated within the data center and can sense the states associated with components installed therein. Remote sensors or other telemetry can also be used to monitor other adverse signals within the data center. As an example, if a fan cooling a rack stops operating, it can be read by a sensor and reported to the deployment and monitoring functions. This type of monitoring can ensure that any impact on temperature profile-based memory write redirects is detected, recorded, and corrected as necessary.

[0030] Referring further to FIG. 5, database 512 can be used to store records related to temperature profiles for memory write redirects and VM migrations, and includes policy records that establish which host servers can implement such functionality. Additionally, database 512 can also store data used to generate reports related to memory write redirects and VM migrations based on the temperature profile.

[0031] The network interface 514 may include a communication interface such as Ethernet (registered trademark), cellular radio, Bluetooth (registered trademark) radio, UWB radio, or other types of wireless or wired communication interfaces. The I / O port may include an Ethernet port, an InfiniBand port, an optical fiber port, or other types of ports. FIG. 5 shows a computing platform 500 including a predetermined number of components arranged and coupled in a predetermined manner, but may include fewer or additional components arranged and coupled differently. Additionally, functions related to the computing platform 500 may be distributed as needed.

[0032] FIG. 6 shows a flowchart 600 of a method according to one example. In this example, the method may be executed in a cloud computing system including a host server. Here, the host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor. As one example, the method may be executed as part of the host server 300 of FIG. 3 as part of the system 100 of FIG. 1. Step 610 may include maintaining a first temperature profile based on information received from temperature sensors associated with each of a plurality of first memory chips included in at least the plurality of first memory modules. As an example, the first temperature profile may correspond to temperature data associated with a memory chip included as part of the aforementioned memory module (e.g., one of the memory modules 360, 370, and 380 of FIG. 3). In one example, a hypervisor associated with the host server may manage the first temperature profile.

[0033] Step 620 may include maintaining a second temperature profile based on information received from temperature sensors associated with each of a plurality of second memory chips included in at least a plurality of second memory modules. As an example, the second temperature profile may correspond to temperature data associated with a memory chip included as part of one of the other memory modules described above (e.g., one of memory modules 360, 370, and 380 of FIG. 3). In one example, a hypervisor associated with the host server may manage the second temperature profile. Additionally, the hypervisor may also periodically initiate a temperature scan to update at least one of the first temperature profile or the second temperature profile. As described above, either a push or pull technique (or a combination of both) may be used to update the temperature profile. Instructions corresponding to the hypervisor and associated modules may be stored in a memory including memory 506 associated with computing platform 500, as needed.

[0034] Step 630 may include automatically redirecting a first write request to a memory from a first computing entity being executed by a first processor to a selected one of a plurality of first memory chips included in at least a plurality of first memory modules, where the temperature of the selected first memory chip does not meet or exceed a temperature threshold, based at least on the first temperature profile. In one example, a hypervisor associated with the host server may assist in automatically redirecting the memory write operation. The CPU that initiates the write operation can write to physical memory with the help of a memory controller (e.g., the memory controller previously described with respect to FIG. 3) based on a memory mapping table maintained by the hypervisor for managing the memory of the host server. Instructions corresponding to the hypervisor and associated modules may be stored in a memory including memory 506 associated with computing platform 500, as needed.

[0035] Step 640 may include automatically redirecting a second write request to the memory from a second computing entity being executed by a second processor to a selected one of a plurality of second memory chips included in at least a plurality of second memory modules, where the temperature of the second memory chips does not meet or exceed a temperature threshold. In one example, a hypervisor associated with a host server may help automatically redirect a memory write operation. The CPU that initiates the write operation can write to physical memory with the help of a memory controller (e.g., the memory controller previously described with respect to FIG. 3) based on a memory mapping table maintained by the hypervisor to manage the memory of the host server. Instructions corresponding to the hypervisor and related modules may be stored in a memory including the memory 506 associated with the computing platform 500 as needed. FIG. 6 illustrates a flowchart 600 as including a predetermined number of steps executed in a predetermined order, but the method may include additional or fewer steps executed in a different order. As one example, the hypervisor may periodically initiate a temperature scan to update at least one of the first temperature profile or the second temperature profile. Further, based on an analysis of the first temperature profile and the second temperature profile as to whether a memory module includes K memory chips, where K is a positive integer, that have a temperature exceeding the temperature threshold for all of a default time frame or for a selected number of times during all of a default time frame, the hypervisor may quarantine a memory module (e.g., any memory module including the aforementioned DIMM) selected from at least one of the first memory module or the second memory module. Finally, the hypervisor can track metrics related to the use of each of the plurality of first memory modules and second memory modules by a computing entity to prevent overuse of a particular memory module by other memory modules.As an example, metrics related to the use of memory modules can be related to the number of times different memory modules are accessed within a given time frame. A histogram can also be used to bin memory modules that are being overused.

[0036] FIG. 7 shows another flowchart 700 of a method according to one example. In this example, the method can be executed in a cloud computing system that includes a first host server and a second host server. Here, the first host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor, where the second host server includes at least a plurality of third memory modules coupled to a third processor and at least a plurality of fourth memory modules coupled to a fourth processor, where the first host server includes a first hypervisor that manages a plurality of first computing entities for execution by the first processor or the second processor, and the second host server includes a second hypervisor that manages a plurality of second computing entities for execution by the third processor or the fourth processor. In this example, the method can be executed on the host server 300 of FIG. 3 as part of the system 100 of FIG. 1.

[0037] Step 710 can include maintaining a first temperature profile based on information received from temperature sensors associated with each of a plurality of first memory chips included in at least a plurality of first memory modules and at least a plurality of second memory modules. As an example, the first temperature profile can correspond to temperature data associated with memory chips included as part of the memory modules described above (e.g., memory modules 360, 370, and 380 of FIG. 3). In one example, a hypervisor associated with the host server can manage the first temperature profile.

[0038] Step 720 may include maintaining a second temperature profile based on information received from temperature sensors associated with respective ones of a plurality of second memory chips included in at least a plurality of third memory modules and at least a plurality of fourth memory modules. As one example, the second temperature profile may correspond to temperature data associated with a memory chip included as part of one of the other memory modules described above (e.g., one of memory modules 360, 370, and 380 in FIG. 3). In one example, a hypervisor associated with a host server may manage the second temperature profile. Additionally, the hypervisor may also periodically initiate a temperature scan to update at least one of the first temperature profile or the second temperature profile. As described above, either a push or pull technique (or a combination of both) may be used to update the temperature profile.

[0039] Step 730 may include automatically redirecting a first write request to memory from a first computing entity being executed by a first processor to a selected one of a plurality of first memory chips included in at least a plurality of first memory modules and at least a plurality of second memory modules, where the temperature of the selected first memory chip does not meet or exceed a temperature threshold, based at least on the first temperature profile. In one example, a hypervisor associated with a host server may assist in automatically redirecting a memory write operation. The CPU that initiates the write operation can write to physical memory with the help of a memory controller (e.g., the memory controller previously described with respect to FIG. 3) based on a memory mapping table maintained by the hypervisor for managing the memory of the host server.

[0040] Step 740 may include automatically redirecting a second write request to the memory from a second computing entity executed by a second processor to a selected one of a plurality of second memory chips included in at least a plurality of third memory modules and at least a plurality of fourth memory modules, where the temperature does not meet or exceed a temperature threshold. In one example, a hypervisor associated with a host server may help to automatically redirect a memory write operation. The CPU that initiates the write operation can write to physical memory with the help of a memory controller (e.g., the memory controller previously described with respect to FIG. 3) based on a memory mapping table maintained by the hypervisor to manage the memory of the host server.

[0041] Step 750 may include automatically migrating at least a subset of the first computing entities from the first host server to the second host server when it is determined that the temperatures of at least N of the plurality of first memory chips, where N is a positive integer, meet or exceed a temperature threshold and the temperature of at least one of the plurality of second memory chips does not meet or exceed the temperature threshold. As described above, with respect to FIG. 3, live migration of a computing entity (e.g., a VM) can be performed by coordination between hypervisors associated with two host servers (e.g., the host server from which the VM is migrated and the host server to which the VM is migrated). In FIG. 7, flowchart 700 is described as including a predetermined number of steps executed in a predetermined order, but the method of collecting resources may include additional steps executed in a different order. As one example, the hypervisor may automatically migrate at least a subset of the first computing entities from the second host server to the first host server when it is determined that the temperatures of at least O of the plurality of second memory chips, where O is a positive integer, meet or exceed a temperature threshold and the temperature of at least one of the plurality of first memory chips does not meet or exceed the temperature threshold. In other words, when the temperature profile changes, the VM can be migrated from one host server to another and then back to the same host server.

[0042] In conclusion, the present disclosure relates to a method in a cloud computing system including a host server, the host server including at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor. The method may include maintaining a first temperature profile based on information received from temperature sensors associated with respective ones of a plurality of first memory chips included in at least one of the plurality of first memory modules. The method may further include maintaining a second temperature profile based on information received from temperature sensors associated with respective ones of a plurality of second memory chips included in at least the plurality of second memory modules. The method may further include automatically redirecting a first write request to memory from a first computing entity being executed by the first processor to a selected one of a plurality of first memory chips included in at least the plurality of first memory modules, the temperature of which does not meet or exceed a temperature threshold, based at least on the first temperature profile. The method may further include automatically redirecting a second write request to memory from a second computing entity being executed by the second processor to a selected one of a plurality of second memory chips included in at least the plurality of second memory modules, the temperature of which does not meet or exceed a temperature threshold, based at least on the second temperature profile.

[0043] The host server includes a hypervisor for managing a plurality of computing entities for execution by the first or second processor, and the hypervisor may be configured to maintain both the first temperature profile and the second temperature profile. The method may further include the hypervisor periodically initiating a temperature scan to update at least one of the first temperature profile or the second temperature profile.

[0044] The method further includes isolating a memory module selected from at least one of the first memory module or the second memory module based on an analysis of a first temperature profile and a second temperature profile, the method including a hypervisor that has K memory chips, where K is a positive integer, having a temperature exceeding a temperature threshold for all of a predefined time frame or for a selected number of times during all of the predefined time frame. The method further includes a hypervisor that tracks metrics associated with the use of each of the plurality of first memory modules and the second memory modules by a computing entity to prevent overuse of a particular memory module relative to other memory modules.

[0045] The method further includes managing a mapping between virtual memory and physical memory allocated to a computing entity. Each of the first computing entity and the second computing entity includes at least one of a virtual machine (VM) for serverless functions, a microVM, a microservice, or a unikernel.

[0046] In yet another aspect, the present disclosure relates to a system that includes a host server. The host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor. The system further includes a hypervisor associated with the host server. The hypervisor is configured to: (1) maintain a first temperature profile based on information received from temperature sensors associated with respective ones of a plurality of first memory chips included in at least one of the plurality of first memory modules; (2) maintain a second temperature profile based on information received from temperature sensors associated with respective ones of a plurality of second memory chips included in at least the plurality of second memory modules; (3) automatically redirect a first write request to memory from a first computing entity being executed by the first processor to a selected one of a plurality of first memory chips included in at least the plurality of first memory modules that do not meet or exceed a temperature threshold; and (4) automatically redirect a second write request to memory from a second computing entity being executed by the second processor to a selected one of the plurality of second memory chips included in at least the plurality of second memory modules that do not meet or exceed a temperature threshold, based at least on the second temperature profile.

[0047] The hypervisor is further configured to periodically initiate a temperature scan to update at least one of the first temperature profile or the second temperature profile. The hypervisor is further configured to isolate a memory module selected from at least one of the first memory modules or the second memory modules based on an analysis of the first temperature profile and the second temperature profile as to whether, for all or a selected number of times during a given time frame, there are K memory chips, where K is a positive integer, having a temperature that exceeds a temperature threshold.

[0048] The hypervisor is further configured to track metrics associated with the use of each of a plurality of first memory modules and second memory modules by a computing entity in order to prevent overuse of a particular memory module with respect to other memory modules. The hypervisor is further configured to manage a mapping between virtual memory and physical memory assigned to a computing entity. Each of the first computing entity and the second computing entity includes at least one of a virtual machine (VM) for serverless functions, a micro VM, a microservice, or a unikernel.

[0049] In another aspect, the present disclosure relates to a method in a cloud computing system including a first host server and a second host server. Here, the first host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor, and the second host server includes at least a plurality of third memory modules coupled to a third processor and at least a plurality of fourth memory modules coupled to a fourth processor. Here, the first host server includes a first hypervisor that manages a plurality of first computing entities for execution by the first processor or the second processor, and the second host server includes a second hypervisor that manages a plurality of second computing entities for execution by the third processor or the fourth processor. The method may include maintaining a first temperature profile based on information received from temperature sensors associated with respective ones of the plurality of first memory chips included in at least the plurality of first memory modules and at least the plurality of second memory modules. The method may further include maintaining a second temperature profile based on information received from temperature sensors associated with respective ones of the plurality of second memory chips included in at least the plurality of third memory modules and at least the plurality of fourth memory modules. The method may further include automatically redirecting a first write request to memory from a first computing entity being executed by the first processor to a selected one of the plurality of first memory chips included in at least the plurality of first memory modules and at least the plurality of second memory modules, where the temperature does not meet or exceed a temperature threshold. The method may further include automatically redirecting a second write request to memory from a second computing entity being executed by the second processor to a selected one of the plurality of second memory chips included in at least the plurality of third memory modules and at least the plurality of fourth memory modules, where the temperature does not meet or exceed a temperature threshold, based at least on the second temperature profile.The method may further include automatically migrating at least a subset of the first computing entities from the first host server to the second host server when it is determined that the temperatures of at least N of the plurality of first memory chips, where N is a positive integer, meet or exceed a temperature threshold, and the temperature of at least one of the plurality of second memory chips does not meet or exceed the temperature threshold.

[0050] The method may further include automatically migrating at least a subset of the first computing entities from the second host server to the first host server when it is determined that the temperatures of at least O of the plurality of second memory chips, where O is a positive integer, meet or exceed a temperature threshold, and the temperature of at least one of the plurality of first memory chips does not meet or exceed the temperature threshold. The method may further include the first hypervisor periodically initiating a temperature scan to update the first temperature profile and the second hypervisor periodically initiating a temperature scan to update the second temperature profile.

[0051] The method may further include the first hypervisor isolating a memory module selected from at least one of the first memory module or the second memory module by analyzing the first temperature profile to determine whether K memory chips, where K is a positive integer, having temperatures exceeding the temperature threshold are included during the entire period of a predefined time frame or for a selected number of times during the entire period of the predefined time frame. The method may further include the second hypervisor isolating a memory module selected from at least one of the third memory module or the fourth memory module by analyzing the second temperature profile to determine whether K memory chips, where K is a positive integer, having temperatures exceeding the temperature threshold are included during the entire period of a predefined time frame or for a selected number of times during the entire period of the predefined time frame.

[0052] The method may further include a first hypervisor tracking metrics associated with the use of each of a plurality of first memory modules and second memory modules by a computing entity to prevent overuse of a particular memory module relative to other memory modules. The method may further include a second hypervisor tracking metrics associated with the use of each of a plurality of third memory modules and fourth memory modules by a computing entity to prevent overuse of a particular memory module relative to other memory modules.

[0053] It should be understood that the methods, modules, and components shown herein are merely exemplary. Alternatively, or in addition, the functions described herein may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like. Although abstract, yet in a clear sense, any configuration of components for achieving the same functionality is effectively "associated" so as to achieve the desired functionality. Thus, two components combined herein to achieve a particular functionality can be considered to be "associated with" each other, regardless of architecture or intermediate components, so as to achieve the desired functionality. Similarly, two components so associated can also be considered to be "operably connected" or "coupled" to each other to achieve the desired functionality.

[0054] The functionality associated with some of the examples described in this disclosure can also include instructions stored on non-transitory media. As used herein, "non-transitory media" refers to any media that stores data and / or instructions for operating a machine in a specified manner. Representative non-transitory media include non-volatile media and / or volatile media. Non-volatile media includes, for example, hard disks, solid state drives, magnetic disks or tapes, optical disks or tapes, flash memory, EPROM, NVRAM, PRAM, or other such media, or network versions of such media. Volatile media includes, for example, dynamic memories such as DRAM, SRAM, cache, or other such media. Non-transitory media is different from, but can be used in combination with, transmission media. Transmission media is used to transfer data and / or instructions to or from a machine. Representative transmission media include coaxial cables, fiber optic cables, copper wire, and wireless media such as radio waves.

[0055] Furthermore, those skilled in the art will recognize that the boundaries between the functionality of the operations described above are merely exemplary. The functionality of multiple operations may be combined into one operation and / or the functionality of one operation may be distributed in additional operations. Additionally, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be changed in various other embodiments.

[0056] Although specific embodiments are shown, various modifications and changes can be made without departing from the scope of the disclosure as defined in the following claims. Accordingly, the specification and drawings are to be considered in an illustrative rather than a limiting sense, and all such changes are intended to be included within the scope of the disclosure. Benefits, advantages, or solutions to problems described herein with respect to specific examples are not intended to be construed as critical, required, or essential features or elements of any or all the claims.

[0057] Furthermore, as used herein, the term “a” is defined as one or more. Also, the use of introductory phrases such as “at least one” and “one or more” in the claims should not be understood to limit any particular claim containing such an introductory phrase to inventions containing only one of the elements introduced by the indefinite article “a,” even if the same claim includes both such introductory phrases and the indefinite article “a.” The same holds true for the use of definite articles.

[0058] Unless specifically stated otherwise, terms such as “first” and “second” are used to arbitrarily distinguish between elements represented by such terms. Accordingly, these terms are not necessarily intended to indicate a temporal or other ranking of such elements.

Claims

1. A method in a cloud computing system including a first host server and a second host server, wherein the first host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor, and the second host server includes at least a plurality of third memory modules coupled to a third processor and at least a plurality of fourth memory modules coupled to a fourth processor, wherein the first host server includes a first hypervisor for managing a plurality of first computing entities executed by the first processor or the second processor, and the second host server includes a second hypervisor for managing a plurality of second computing entities executed by the third processor or the fourth processor, and the method includes: maintaining a first temperature profile by the first hypervisor based on information received from first temperature sensors respectively associated with a plurality of first memory chips included in at least the plurality of first memory modules; maintaining a second temperature profile by the first hypervisor based on information received from second temperature sensors respectively associated with a plurality of second memory chips included in at least the plurality of second memory modules; automatically redirecting, by the first hypervisor based at least on the first temperature profile, a first write request to memory from a first computing entity executed by the first processor to a selected one of the plurality of first memory chips included in at least the plurality of first memory modules, the temperature of which does not meet a temperature threshold; automatically redirecting, by the first hypervisor based at least on the second temperature profile, a second write request to memory from a second computing entity executed by the second processor to a selected one of the plurality of second memory chips included in at least the plurality of second memory modules, the temperature of which does not meet a temperature threshold. When it is determined that the temperatures of at least N of the plurality of the first memory chips and the plurality of the second memory chips, where N is a positive integer, satisfy the temperature threshold, the first hypervisor (1) the temperature of at least one memory chip among the plurality of third memory modules or the plurality of fourth memory modules does not satisfy the temperature threshold, and (2) when, based on the exchange of information between the first hypervisor and the second hypervisor, the second host server has sufficient memory to enable live migration, automatically migrates at least a subset of the first computing entities from the first host server to the second host server. Method. **Claim 2** The method further includes a statement in which the first hypervisor and the second hypervisor exchange control information, a step of determining whether the plurality of third memory modules or the plurality of fourth memory modules have sufficient physical memory to enable live migration. The method according to claim 1, including this. **Claim 3** The method further includes a step in which the first hypervisor periodically starts a temperature scan to update at least one of the first temperature profile or the second temperature profile. The method according to claim 2, including this. **Claim 4** The method further includes based on the analysis of the first temperature profile and the second temperature profile, whether K memory chips, where K is a positive integer, having temperatures exceeding the temperature threshold are included during the entire period of a predetermined time frame or for a selected number of times during the entire period of the predetermined time frame, the first hypervisor isolates a memory module selected from at least one of the plurality of first memory modules or the plurality of second memory modules. The method according to claim 3, including this. **Claim 5** The method further includes a step in which the first hypervisor tracks metrics related to the use of each of the plurality of first memory modules and the plurality of second memory modules by the plurality of first computing entities to prevent overuse of a specific memory module with respect to other memory modules. The method according to claim 2, including this. **Claim 6** The method further includes Managing the mapping between virtual memory and physical memory assigned to a computing entity The method according to claim 2, comprising:

7. Each of the first computing entity and the second computing entity includes at least one of a virtual machine (VM) for serverless functions, a micro VM, a microservice, or a unikernel The method according to claim 1

8. A system, comprising: A first host server and a second host server, the first host server including at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor, and the second host server including at least a plurality of third memory modules coupled to a third processor and at least a plurality of fourth memory modules coupled to a fourth processor, the first host server and the second host server; A first hypervisor associated with the first host server; Managing a plurality of first computing entities executed by the first processor or the second processor; Maintaining a first temperature profile based on information received from a first temperature sensor associated with each of a plurality of first memory chips included in at least the plurality of first memory modules, and Maintaining a second temperature profile based on information received from a second temperature sensor associated with each of a plurality of second memory chips included in at least the plurality of second memory modules, a first hypervisor configured as such; A second hypervisor associated with the second host server, configured to manage a plurality of second computing entities executed by the third processor or the fourth processor, including the second hypervisor; The first hypervisor further includes: Based on at least the first temperature profile, automatically redirecting a first write request to memory from a first computing entity executed by the first processor to a selected one of the plurality of first memory chips included in at least the plurality of first memory modules, the temperature of which does not satisfy a temperature threshold; Based on at least the second temperature profile, a second write request to the memory is automatically redirected by the second computing entity executed by the second processor from the second computing entity to a selected one of the plurality of second memory chips included in at least a plurality of the second memory modules, the temperature of which does not satisfy the temperature threshold. When it is determined that the temperatures of at least N of the plurality of first memory chips and the plurality of second memory chips, where N is a positive integer, satisfy the temperature threshold. (1) The temperature of at least one memory chip among the plurality of third memory modules or the plurality of fourth memory modules does not satisfy the temperature threshold, and (2) Based on the exchange of information between the first hypervisor and the second hypervisor, when the second host server has sufficient memory to enable live migration. At least a subset of the first computing entity is automatically migrated from the first host server to the second host server. Is configured as follows. System.

9. The first hypervisor further Periodically initiates a temperature scan to update at least one of the first temperature profile or the second temperature profile. The system according to claim 8, which is configured as follows.

10. The first hypervisor further Based on the analysis of the first temperature profile and the second temperature profile, whether or not to include K memory chips, where K is a positive integer, having a temperature exceeding the temperature threshold during the entire period of a predetermined time frame or for a selected number of times during the entire period of the predetermined time frame. Isolate a selected memory module from at least one of the first memory module or the second memory module. The system according to claim 9, which is configured as follows.

11. The first hypervisor further Tracks metrics related to the use of each of the plurality of first memory modules and the second memory modules by a computing entity to prevent overuse of a particular memory module by other memory modules. The system according to claim 8, which is configured as follows.

12. The first hypervisor further Managing the mapping between virtual memory and physical memory assigned to a computing entity, The system according to claim 8, configured as such.

13. Each of the first computing entity and the second computing entity includes at least one of a virtual machine (VM) for serverless functions, a micro VM, a microservice, or a unikernel. The system according to claim 9.

14. A method in a cloud computing system including a first host server and a second host server, The first host server includes at least a plurality of first memory modules coupled to a first processor and at least a plurality of second memory modules coupled to a second processor. The second host server includes at least a plurality of third memory modules coupled to a third processor and at least a plurality of fourth memory modules coupled to a fourth processor. The first host server includes a first hypervisor that manages a plurality of first computing entities for execution by the first processor or the second processor, and the second host server includes a second hypervisor that manages a plurality of second computing entities for execution by the third processor or the fourth processor. The method includes: Maintaining a first temperature profile based on information received from first temperature sensors associated with respective first memory chips included in at least a plurality of the first memory modules and at least a plurality of the second memory modules by the first hypervisor. The first temperature profile includes information regarding a first temperature value of at least one of the plurality of first memory chips. Maintaining a second temperature profile based on information received from second temperature sensors associated with respective second memory chips included in at least a plurality of the third memory modules and at least a plurality of the fourth memory modules by the second hypervisor. The second temperature profile includes information regarding a second temperature value of at least one of the plurality of second memory chips. Based on at least the first temperature profile, the first hypervisor automatically redirects a first write request to memory from a first computing entity being executed by the first processor to a selected one of a plurality of the first memory chips included in at least the plurality of the first memory modules and at least the plurality of the second memory modules, where the temperature does not meet the temperature threshold. Based on at least the second temperature profile, the second hypervisor automatically redirects a second write request to memory from a second computing entity being executed by the third processor to a selected one of a plurality of the second memory chips included in at least the plurality of the third memory modules and at least the plurality of the fourth memory modules, where the temperature does not meet the temperature threshold. The method includes the above steps. When it is determined that at least N, where N is a positive integer, of the plurality of the first memory chips have a temperature that meets the temperature threshold, the first hypervisor (1) When the temperature of at least one memory chip of the plurality of the third memory modules or the plurality of the fourth memory modules does not meet the temperature threshold and (2) Based on the exchange of information between the first hypervisor and the second hypervisor, when the second host server has sufficient memory to enable live migration to occur, automatically migrates at least a subset of the first computing entity from the first host server to the second host server. Method.

15. The method further includes When it is determined that at least O, where O is a positive integer, of the plurality of the second memory chips have a temperature that meets or exceeds the temperature threshold, and when the temperature of at least one of the plurality of the first memory chips does not meet or exceed the temperature threshold, automatically migrating at least a subset of the first computing entity from the second host server to the first host server. The method according to claim 14, including the above step.

16. The method further includes the step of the first hypervisor periodically initiating a temperature scan to update the first temperature profile. The step in which the second hypervisor periodically starts a temperature scan to update the second temperature profile The method according to claim 14, comprising the above.

17. The method further comprises The step in which the first hypervisor isolates a memory module selected from at least one of the first memory module or the second memory module by analyzing the first temperature profile including the step of determining whether the memory module includes at least K memory chips where K is a positive integer The temperature exceeds a temperature threshold during the entire predefined time frame or for a selected number of times during the entire predefined time frame The method according to claim 16.

18. The method further comprises The step in which the second hypervisor isolates a memory module selected from at least one of the third memory module or the fourth memory module by analyzing the second temperature profile including the step of determining whether the memory module includes at least K memory chips where K is a positive integer The temperature exceeds a temperature threshold during the entire predefined time frame or for a selected number of times during the entire predefined time frame The method according to claim 16.

19. The method further comprises The step in which the first hypervisor tracks metrics related to the use of each of a plurality of the first memory modules and a plurality of the second memory modules by a computer entity including the step of preventing overuse of a specific memory module with respect to other memory modules The method according to claim 14.

20. The method further comprises The step in which the second hypervisor tracks metrics related to the use of each of a plurality of the third memory modules and a plurality of the fourth memory modules by a computer entity including the step of preventing overuse of a specific memory module with respect to other memory modules The method according to claim 14.

Citation Information

Patent Citations

  • Information processor, information processing method, and program

    JP2006018758A

  • Semiconductor device

    JP2016212745A

  • Logical server management program, logical server management device, and logical server management method

    JP2019061560A