Dynamic Platform Feature Tuning Based on Virtual Machine Runtime Requirements
Through the virtual machine dynamic tuner system, the platform characteristics are dynamically adjusted based on the virtual machine runtime requirements, and the performance matching problem between virtual machines on cloud servers is solved, and efficient feature management and performance optimization are achieved.
Patent Information
- Application Number
- CN201810990688.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-28
- Filing Date
- 2018-08-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2038-08-28
AI Technical Summary
On public cloud servers, it is difficult to efficiently adjust platform characteristics for each virtual machine to match its runtime requirements, especially due to frequent context switching and workload differences between different virtual machines, which can lead to performance penalties in existing methods.
Through the virtual machine dynamic tuner system, the platform characteristics are dynamically adjusted based on the virtual machine runtime requirements, and the feature administrator and virtual machine monitor are used to independently configure the feature settings of the logical core at each virtual machine level, avoiding the writing of specific registers of the platform feature model per thread, and achieving efficient feature management.
It realizes the optimization of the runtime performance of each virtual machine without affecting the performance of other virtual machines, improves the overall performance and efficiency of the cloud server platform, and adapts to the unique needs of different virtual machines.
Smart Images

Figure CN109582434B_ABST
Abstract
Description
Technical Field
[0001] The embodiments generally relate to cloud server management. More specifically, the embodiments relate to techniques for dynamically tuning platform features based on the runtime requirements of virtual machines. Background Art
[0002] In a native server, an administrator can optimize workload performance by modifying (e.g., by enabling, disabling, or changing) platform features of the entire platform (e.g., prefetcher, C state, P state, etc.). For example, in some instances, enabling the prefetcher for the entire platform can improve the performance of high-performance computing (HPC) applications by several orders of magnitude. On a public cloud server, such modifications to platform features (e.g., by enabling, disabling, or changing platform features) can be applied across the entire platform (e.g., to every and each virtual machine running on this platform). Brief Description of the Drawings
[0003] The various advantages of the embodiments will become apparent to those skilled in the art by reading the following specification and the appended claims and by referring to the following drawings, in which:
[0004] Figure 1 is a block diagram of an example of a cloud server platform system according to an embodiment;
[0005] Figure 2 is a sequence diagram of an example of an operation performed in a cloud server platform according to an embodiment;
[0006] Figure 3 is a screenshot of an example of a graphical user interface according to an embodiment;
[0007] Figure 4 is a graph of an example of experimental results according to an embodiment;
[0008] Figure 5 is a block diagram of an example of a logical architecture according to an embodiment;
[0009] Figure 6 is a block diagram of an example of a processor according to an embodiment; and
[0010] Figure 7 is a block diagram of an example of a computing system according to an embodiment. Detailed Description
[0011] As already discussed, in conventional approaches, enabling, disabling, and / or changing platform features on a public cloud server can apply to every single virtual machine running on that platform. For example, a workload can use the same platform features (e.g., prefetchers, C-states, P-states, etc.) across the entire cloud server platform. Thus, there may not be an efficient way to adjust these features for each individual virtual machine. Another additional difficulty may be that virtual machines are typically scheduled across multiple cores and are also regularly rescheduled by a virtual machine monitor (VMM / Orchestrator).
[0012] Accordingly, individual virtual machines running on a public cloud server platform can benefit from using unique feature / control settings at runtime. One option to address this need could be to write per-thread platform feature model specific registers (MSRs) at each context switch to adjust settings that match the runtime requirements of the workload running on an individual core. However, due to the time involved in reading / writing these MSRs at each context switch, such an approach may be impractical from a performance perspective. A modern "2P" server running a heavy workload (e.g., two CPU sockets in a single motherboard) can have over a million context switches per second. In fact, context switches may already incur a measurable performance penalty. Thus, adding more overhead to context switches may not be an appropriate solution.
[0013] As will be described in more detail, an enhanced approach can tune platform features differently for different virtual machines operating on the same platform, since each workload running in a virtual machine can require different platform features at runtime. For example, such platform setting requirements can typically be workload-specific. For example, a software-defined networking (SDN) load balancer / multiplexer application (e.g., which can be latency-sensitive) may require disabling C-states, while other applications may not necessarily need to do so. More generally, basic input / output system (BIOS) optimization guides often involve different settings for different kinds of workloads.
[0014] Now turning to Figure 1 , a virtual machine dynamic tuner system 100 for managing virtual machines is shown. System 100 can dynamically tune platform features based on virtual machine runtime requirements.
[0015] In the illustrated example, system 100 may include virtual resources 110 and a cloud server platform 112. The virtual resources 110 may include multiple virtual machines 114. Each of the multiple virtual machines 114 may be used as a self - contained unit, running its own operating system (OS) and / or application software. For example, the multiple virtual machines 114 may include a first virtual machine 116 running a first application ("App") 118, and an active second virtual machine 120 running an active second application ("App") 122. Similarly, the Nth virtual machine 124 may run the Nth application ("App") 126.
[0016] For example, the active second application 122 associated with the active second virtual machine 120 may be associated with application runtime requirements. Such application runtime requirements may include an indication of latency sensitivity or an indication of a data reference pattern, and / or combinations thereof. The data reference pattern may include an indication of whether the active second application 122 accesses memory in a random or sequential manner, such as, for example, the active second application 122 includes a streaming application.
[0017] Each of the first virtual machine 116 and the second virtual machine 120 in the system 100 may be associated with one or more of a portion of the associated hardware cores 131 ("HC0" to "HC N ") of the cloud server platform 112, one or more logical cores 130 ("LC0" to "LC 2N+1 "). As illustrated, each of the hardware cores 131 may include two or more logical cores 130. As will be described in more detail below, the techniques described herein may be used to independently configure each of the logical cores 130 (e.g., without using writes to per - thread platform feature model specific registers / MSRs). The first virtual machine 116 may have a first configuration for efficiently supporting a first feature set ("Feature") 132 on the associated logical cores 130. The second virtual machine 120 may have a different second configuration for efficiently supporting a different second feature set ("Feature") 134 on a different portion of the associated logical cores 130. For example, the first feature set 132 may be at least partially based on the application runtime requirements associated with the first application 118, while the second feature set 134 may be at least partially based on the application runtime requirements associated with the second application 122. Similarly, the Nth virtual machine 124 may have a still different configuration for efficiently supporting a still different feature set ("Feature") 136 associated with the Nth application 126.
[0018] The virtual resource 110 may include a feature administrator 140. As used herein, the term "feature administrator" generally may refer to an automated administrator and / or a user administrator that is configured to control platform settings for each virtual machine on a cloud server platform at least in part based on application runtime requirements. Such an automated agent-type feature administrator 140 may be an intelligent agent that can identify appropriate feature settings to be set on the virtual machine control 160. For example, such an automated agent-type feature administrator 140 may infer feature settings by collecting historical performance data that will point to the runtime requirements of a particular virtual machine workload. In the illustrated example, the feature administrator 140 may determine active second feature settings 134 that are specific to the active second application 122 at least in part based on application runtime requirements. As discussed above, the active second feature settings may include an idle power state setting, an execution power state setting, a prefetch setting, etc., and / or a combination thereof. The prefetch settings may include a core prefetch setting, a non-core prefetch setting, etc., and / or a combination thereof. The power state settings may include a C-state latency tolerance, a C-state cache dump purge tolerance, a hardware control performance (HWP) minimum, an HWP maximum, an HWP energy / performance preference (EPP), etc., and / or a combination thereof.
[0019] The virtual resource 110 may include a virtual machine monitor 150. As used herein, the term "virtual machine monitor" generally may refer to, for example, a virtual machine monitor (VMM), a hypervisor, a virtual machine coordinator, etc., that is configured to manage the execution of a guest operating system in a virtual environment on a virtual machine. In the illustrated example, the virtual machine monitor 150 may independently assign the active second feature settings 134 to the active second virtual machine 120 on a per-virtual machine basis, while independently assigning the first feature settings 132 to the first virtual machine 116. Additionally, the virtual machine monitor 150 may iteratively identify rescheduling events and then, in response to the identified rescheduling events, iteratively schedule the active second virtual machine 120 to run on a particular set of logical cores 130.
[0020] The virtual resource 110 may include a virtual machine control 160. As used herein, the term "virtual machine control" may generally refer to a virtual machine control system (VMCS), virtual machine control (VMC), or any other similar data structure in memory for managing virtual machines, including data structures in memory that can be assigned per individual virtual machine when managed by a virtual machine monitor. In the illustrated example, the virtual machine control 160 may include a control field 162 for storing a bitmask of the active second feature settings 134. A particular set of logical cores 130 associated with the active second virtual machine 120 may pass the bitmask of the active second feature settings 134 specific to the active second application 122 from the control field 162 of the virtual machine control 160 and implement the active second feature settings 134 specific to the active second application 122 by adapting the operation of the particular set of logical cores 130 associated with the active second virtual machine 120 to support the active second virtual machine 120 and the active second application 122.
[0021] In some examples, the virtual machine control 160 can be a virtual machine control system (VMCS). In such examples, each VMCS may include a per-virtual machine feature management bitmask. The bitmask may allow privileged software to manipulate the feature settings of each of the individual virtual machines 114 running on the cloud server platform 112. The proposed feature settings are virtual machine-specific requirements (e.g., HWP level), which are abstractions of the underlying microarchitecture set of platform features (e.g., core frequency). The VMCS region may include multiple different zones. The control field 162 may store the feature bitmask in the previously reserved bits (e.g., undefined / unused bits) of the VMCS. For example, the software interface may specify certain bits as reserved so that when the need arises to extend the interface to add new capabilities, these reserved bits can be used. Thus, when new capabilities for the feature bitmask are added, the previously reserved bits can then be redefined as the feature bitmask. The logical cores 130 may implement the bitmask to replace the current model-specific register (MSR) settings (e.g., IA32_MISC_ENABLE (IA32_MISC enabled), etc.). That is, when scheduling one of the virtual machines 114 on a particular set of logical cores 130, the particular set of logical cores 130 will reset the platform features according to the bitmask field in the VMCS.
[0022] The cloud server platform 112 may include a power source 170. The power source 170 may supply power to the system 100. It should be understood that for clarity, the system 100 and the cloud server platform 112 may include other components not illustrated herein.
[0023] In operation, the virtual machine monitor 150 can control / manage workload runtime requirements by setting a bitmask on the virtual machine control 160. This bitmask can define the active feature settings 134 required for the active second virtual machine 120. The logical cores 130 associated with the active second virtual machine 120 can be changed to read this bitmask and implement the active feature settings 134 requested for the active second virtual machine 120. Regardless of which core / slot or system the virtual machine 114 is scheduled on, these settings can remain configured for each of the virtual machines 114. For example, the virtual machine 114 can migrate across physical systems. In such a migration, the virtual machine 114 can retain all application and state information.
[0024] Figure 2 Illustrates a method 300 for dynamically tuning a virtual machine. Method 300 can generally be implemented via one or more of the components in the system 100 that has been discussed ( Figure 1 ).
[0025] Continuing to refer to Figure 1 and Figure 2 , at operation 302 (e.g., "pass application runtime requirements"), application runtime requirements can be passed. For example, application runtime requirements can be passed between the active application 122 ("application") and the feature administrator 140.
[0026] At operation 304 (e.g., "determine feature settings based on application runtime requirements"), feature settings can be determined at least in part based on application runtime requirements. For example, the active feature settings 134 ("features") specific to the active application 122 associated with the active second virtual machine 120 can be determined by the feature administrator 140 at least in part based on application runtime requirements.
[0027] In some implementations, application runtime requirements can include one or more of the following: an indication of latency sensitivity or an indication of a data reference pattern, and / or a combination thereof. The data reference pattern can include an indication of whether the active application 122 accesses memory in a random or sequential manner. For example, operation 304 can determine whether the active application 122 is a streaming application. The active feature settings 134 can include one or more of the following: an idle power state setting, an execution power state setting, or a prefetch setting, and / or a combination thereof.
[0028] At operation 306 (e.g., "assign the determined feature settings to the associated virtual machine"), the determined feature settings can be assigned to the associated virtual machine. For example, the active feature settings 134 can be independently assigned to the active second virtual machine 120 via the virtual machine monitor 150. Similarly, different active feature settings (not shown here) can be independently assigned to other virtual machines (not shown here) on a per-virtual machine basis.
[0029] At operation 308 (e.g., "store the assigned feature settings as a bit map"), the assigned feature settings may be stored as a bit mask. For example, the active feature settings 134 may be stored as a bit mask in the control field 162 of the virtual machine control 160.
[0030] At operation 310 (e.g., "identify a rescheduling event"), a rescheduling event may be identified. For example, the rescheduling event may be iteratively identified by the virtual machine monitor 150.
[0031] At operation 312 (e.g., "schedule the active virtual machine to run on a particular set of logical cores"), the active virtual machine may be scheduled to run on a particular set of logical cores. For example, in response to the identified rescheduling event, the virtual machine monitor 150 may iteratively schedule the active second virtual machine 120 to run on a particular set of the logical cores 130.
[0032] At operation 314 (e.g., "pass the bit mask"), the bit mask may be passed. For example, the bit mask of the active feature settings 134 that is specific to the active application 122 may be passed from the control field 162 of the virtual machine control 160 to a particular set of the logical cores 130.
[0033] At operation 316 (e.g., "implement the feature settings"), the feature settings may be implemented. For example, the active feature settings 134 that are specific to the active application 122 may be implemented on a particular set of the logical cores 130 by adapting the operations of the particular set of the logical cores 130 to support the active second virtual machine 120 and the active application 122. In one implementation, the first virtual machine 116 and the second virtual machine 120 of the cloud server platform 112 may each be associated with one or more of the logical cores 130. The first virtual machine 116 may have a first configuration for efficiently supporting a first feature settings arrangement on the associated logical cores 130, where this configuration may be implemented in an operation similar to operation 316. The second virtual machine 120 may have a different second configuration for efficiently supporting a different second feature settings arrangement on different associated logical cores 130, where this configuration may be implemented at operation 316. As described above, the feature settings specific to any particular application associated with a corresponding virtual machine may be determined based on the application runtime requirements.
[0034] In one implementation, method 300 is operable such that a feature administrator 140 can assign the active feature settings 134 in a virtual basic input / output system (BIOS) to an associated active second virtual machine 120 via, for example, a virtual machine monitor 150. The virtual machine monitor 150 can assign the active feature settings 134 to an associated active second virtual machine 120 (e.g., to turn on a high performance computing / HPC style prefetcher). The virtual machine monitor 150 can write the feature settings as a bitmask of all the feature settings 134 required by the active second virtual machine 120 into a control field 162 of a virtual machine control 160. The active second virtual machine 120 can be scheduled to run on a particular set of logical cores 130. The particular set of logical cores 130 accesses the bitmask of the active feature settings 134 required by the active second virtual machine 120 from the control field 162 of the virtual machine control 160. The particular set of logical cores 130 implements platform settings required by the active second virtual machine 120 at least in part based on the active feature settings 134.
[0035] Embodiments of method 300 can be implemented in a system, apparatus, processor, reconfigurable device, etc. (e.g., such as those described herein). More specifically, a hardware implementation of method 300 can include configurable logic (such as, for example, a PLA, FPGA, CPLD), or fixed function logic hardware using circuit technologies (such as, for example, ASIC, CMOS or TTL technologies), or any combination thereof. Alternatively or additionally, method 300 can be implemented in one or more modules as a set of logical instructions stored in a machine or computer readable storage medium (such as, RAM, ROM, PROM, firmware, flash memory, etc.) to be executed by a processor or computing device. For example, computer program code for implementing component operations can be written in any combination of one or more OS applicable / appropriate programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.
[0036] For example, embodiments or portions of method 300 can be implemented in an application running on an OS (e.g., via an application programming interface / API) or driver software. Additionally, the logical instructions can include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, status - setting data, configuration data for an integrated circuit, status information that personalizes an electronic circuit and / or other structural components made from hardware (e.g., a host processor, a central processing unit / CPU, a microcontroller, etc.).
[0037] Figure 3An illustration of interface 400 is shown. Interface 400 is an example of a feature manager that specifies application runtime requirements for determining feature settings. As described above, application runtime requirements may include one or more of the following: an indication of latency sensitivity or an indication of a data reference pattern, and / or a combination thereof. The data reference pattern may include an indication of whether the active application is a streaming application. In addition to providing a user experience (UX) to information technology (IT) operators / administrators in plain English via interface 400, the techniques described herein may obtain these application runtime requirements and convert them into active feature settings (e.g., via underlying per-thread platform feature model specific registers / MSR settings) to be programmed as bitmasks for feature management of each individual virtual machine. Thus, the resulting active feature settings may include one or more of the following: idle power state settings, execution power state settings, or prefetch settings, and / or a combination thereof.
[0038] Some examples of high-level use cases that may be optimized using the techniques described herein are as follows. In one example, an HPC virtual machine may need to enable the prefetcher, while a web virtual machine may not need the same prefetcher, and the two virtual machines operate on the same platform. In another example, a latency-sensitive virtual machine may need to disable the C state during a specific time of day. In yet another example, a virtual machine that requires deterministic performance may set the HWP minimum to 100, while other virtual machines may require different levels of HWP, and all virtual machines operate on the same platform.
[0039] Additionally, in one example implementation, an application responsible for monitoring a health diagnostic and life support system may be extremely latency-sensitive and not have a specific data access pattern. In such an implementation, interface 400 may also expose profile options (e.g., "HPC profile") at the hypervisor level immediately with preset definitions for all application requirements. Such life support examples may need to disable aggressive power management features to reduce latency and most likely have no preference for prefetcher settings, which is beneficial for streaming applications or applications with high spatial locality.
[0040] Figure 4 A chart 500 illustrating experimental results is shown. In the laboratory, the impact of platform features was prototyped by running different workloads simultaneously and changing one of these features at a time (e.g., baseline 510, power management / PM 520, Turbo on 530, prefetcher off / "Pre-f" 540, etc.). As shown in chart 500, platform features have a significant impact on workload performance.
[0041] Figure 4The results illustrated therein show that for high-performance computing (HPC) workloads, performance is negatively affected when the prefetcher is turned off, while it does not affect the performance of web transaction workloads. Additionally, the results show that switching P-states and C-states have no effect on the performance of certain workloads (e.g., HPC-type applications), while it does have a negative impact on the performance of other workloads (e.g., web transaction workloads). Finally, the results show that turning on the turbo has a positive impact on web transaction workloads and a minimal positive impact on the performance of HPC workloads.
[0042] Table 1 shows another experiment conducted, which runs web workloads (e.g., each workload running within a separate virtual machine) against another workload that requires the prefetcher to be on (e.g., an HPC workload in a separate virtual machine). The following table shows the results of the different runs.
[0043]
[0044] The first column in Table 1 shows the throughput and response time of the workload running alone on the platform with the prefetcher on. The second column shows the same numbers when a noisy neighbor that requires the prefetcher to be on is added to the mix. The last column shows the throughput and response time that return to the same level as the initial run by turning off the prefetcher.
[0045] The results illustrated in Table 1 show that web workloads do not necessarily require the prefetcher to be on (e.g., when the prefetcher is off, the response time actually improves by 4% even when an HPC workload is running).
[0046] Furthermore, the techniques described herein can efficiently implement such optimizations illustrated in Table 1. For example, by experimenting with the techniques described herein, the virtual machine monitor will be able to turn off the prefetcher (e.g., by setting the configuration on the virtual machine control to the web prefetcher setting). As described above, this can be triggered by a user administrator and / or an automated administrator.
[0047] Finally, when the prefetcher is off, the performance of high-performance computing (HPC) workloads is negatively affected. The techniques described herein can be used to significantly reduce this impact, as the virtual machine monitor can turn on these prefetchers in the virtual machine control for this workload (e.g., by setting the configuration on the virtual machine control to the high-performance computing (HPC) prefetcher setting). This can mean that even all workloads - regardless of their runtime requirements - can benefit from the techniques described herein.
[0048] The above examples illustrate how the techniques described herein can be used to optimize the runtime performance of different workloads by a virtual machine monitor. This optimization occurs without compromising the operation of other workloads; however, without an appropriate virtual machine control configuration, the performance of multiple workloads can be optimized simultaneously.
[0049] As illustrated herein, it is difficult to always predict the runtime platform requirements of all workloads in a public cloud. Accordingly, the procedures described herein can allow a facility to modify platform characteristics when it deems appropriate.
[0050] Figure 5 Illustrated is a virtual machine dynamic tuner device 600 (e.g., a semiconductor package, chip, die). The device 600 can implement one or more aspects of the method 300 ( Figure 2 ). The device 600 can readily replace some or all of the virtual machine dynamic tuner system 100 ( Figure 1 ) that has been discussed. The illustrated device 600 includes one or more substrates 602 (e.g., silicon, sapphire, gallium arsenide) and logic 604 (e.g., a transistor array and other integrated circuit / IC components) coupled to the substrate(s) 602. The logic 604 can be at least partially implemented in configurable logic or fixed function logic hardware.
[0051] In addition, the logic 604 can configure one or more first logical cores associated with a first virtual machine in a cloud server platform, wherein the configuration of the one or more first logical cores is at least partially based on one or more first feature settings. The logic 604 can also configure one or more active logical cores associated with an active virtual machine in the cloud server platform, wherein the configuration of the one or more logical cores is based on one or more active feature settings, and wherein the active feature settings are different from the first feature settings.
[0052] Figure 6 Illustrated is a processor core 700 according to one embodiment. The processor core 700 can be a core of any type of processor, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, or other device for executing code. Although Figure 6 only one processor core 700 is illustrated, the processing element can alternatively include more than one Figure 6 processor core 700 as illustrated. The processor core 700 can be a single-threaded core, or for at least one embodiment, the processor core 700 can be multi-threaded in that each core can include more than one hardware thread context (or “logical processor”).
[0053] Figure 6Also illustrated is a memory 770 coupled to a processor core 700. The memory 770 can be any of a variety of memories (including the various levels of the memory hierarchy) known to or otherwise available to those of skill in the art. The memory 770 may include one or more code 713 instructions to be executed by the processor core 700, where the code 713 may implement one or more aspects of the method 300( Figure 2 ) discussed above.
[0054] The processor core 700 follows a program sequence of instructions indicated by the code 713. Each instruction may enter the front-end portion 710 and be processed by one or more decoders 720. The decoder 720 may generate micro-operations (such as, fixed-width micro-operations in a predefined format) as its output, or may generate other instructions, micro-instructions, or control signals that reflect the original code instruction. The illustrated front-end portion 710 also includes register renaming logic 725 and scheduling logic 730, which generally allocate resources and queue the operations corresponding to the translated instructions for execution.
[0055] The processor core 700 is shown as including execution logic 750 having a set of execution units 755-1 through 755-N. Some embodiments may include multiple execution units dedicated to a particular function or set of functions. Other embodiments may include only one execution unit or only one execution unit that can perform a particular function. The illustrated execution logic 750 performs the operations specified by the code instructions.
[0056] After completing the execution of the operations specified by the code instructions, the backend logic 760 retires the instructions of the code 713. In one embodiment, the processor core 700 allows out-of-order execution but requires in-order retirement of instructions. The retirement logic 765 may take various forms known to those of skill in the art (e.g., reorder buffer, etc.). In this manner, the processor core 700 is transformed during the execution of the code 713, at least in terms of the output generated by the decoder, the hardware registers and tables utilized by the register renaming logic 725, and any registers (not shown) modified by the execution logic 750.
[0057] Although not illustrated in Figure 6 , the processing element may include other elements on the chip with the processor core 700. For example, the processing element may include memory control logic with the processor core 700. The processing element may include I / O control logic and / or may include I / O control logic integrated with the memory control logic. The processing element may also include one or more caches.
[0058] Now referring to Figure 7 , shown is a block diagram of an embodiment of a computing system 1000 according to one embodiment. Figure 7Shown in [the figure] is a multi-processor system 1000, which includes a first processing element 1070 and a second processing element 1080. Although two processing elements 1070 and 1080 are shown, it is to be understood that embodiments of the system 1000 may also include only one such processing element.
[0059] The system 1000 is illustrated as a point-to-point interconnect system, in which the first processing element 1070 and the second processing element 1080 are coupled via a point-to-point interconnect 1050. It should be understood that Figure 7 any or all of the interconnects illustrated in [the figure] may be implemented as a multi-drop bus rather than a point-to-point interconnect.
[0060] As Figure 7 shown in [the figure], each of the processing elements 1070 and 1080 may be a multi-core processor including a first and a second processor core (i.e., processor cores 1074a and 1074b, and processor cores 1084a and 1084b). Such cores 1074a, 1074b, 1084a, 1084b may be configured to execute instruction codes in a manner similar to that discussed above in connection with Figure 6 [the relevant content].
[0061] Each processing element 1070, 1080 may include at least one shared cache 1896a, 1896b. The shared caches 1896a, 1896b may store data (e.g., instructions) utilized by one or more components of the processor, such as cores 1074a, 1074b, and 1084a, 1084b. For example, the shared caches 1896a, 1896b may locally cache the data stored in the memories 1032, 1034 for faster access by the components of the processor. In one or more embodiments, the shared caches 1896a, 1896b may include one or more mid-level caches (such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of caches), last-level caches (LLC), and / or combinations thereof.
[0062] Although shown as having only two processing elements 1070, 1080, it should be understood that the scope of each embodiment is not limited thereto. In other embodiments, one or more additional processing elements may be present in a given processor. Alternatively, one or more of the processing elements 1070, 1080 may be elements other than a processor, such as an accelerator or a field programmable gate array. For example, the additional processing element(s) may include additional processors identical to the first processor 1070, additional processors heterogeneous or asymmetric to the first processor 1070, accelerators (such as, by way of example, a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array, or any other processing element. There may be various differences between the processing elements 1070, 1080 in terms of a series of quality metrics including architecture, microarchitecture, thermal, power consumption characteristics, etc. These differences themselves may effectively manifest as asymmetry and heterogeneity among the processing elements 1070, 1080. For at least one embodiment, the various processing elements 1070, 1080 may reside in the same die package.
[0063] The first processing element 1070 may further include memory controller logic (MC) 1072 and point-to-point (P-P) interfaces 1076 and 1078. Similarly, the second processing element 1080 may include MC 1082 and P-P interfaces 1086 and 1088. As Figure 7 shown, MCs 1072 and 1082 couple the processors to respective memories, namely memories 1032 and 1034, which may be portions of the main memory locally attached to the respective processors. Although MCs 1072 and 1082 are illustrated as integrated into the processing elements 1070, 1080, for alternative embodiments, the MC logic may be discrete logic external to the processing elements 1070, 1080 rather than being integrated therein.
[0064] The first processing element 1070 and the second processing element 1080 may be coupled to the I / O subsystem 1090 via P-P interconnections 1076, 1086 respectively. As Figure 7 shown, the I / O subsystem 1090 includes P-P interfaces 1094 and 1098. In addition, the I / O subsystem 1090 includes an interface 1092 that couples the I / O subsystem 1090 to the high performance graphics engine 1038. In one embodiment, the graphics engine 1038 may be coupled to the I / O subsystem 1090 using a bus 1049. Alternatively, a point-to-point interconnect may couple these components.
[0065] Further, the I / O subsystem 1090 can be coupled to the first bus 1016 via the interface 1096. In one embodiment, the first bus 1016 can be a Peripheral Component Interconnect (PCI) bus, or a bus such as a high-speed PCI bus or another third-generation I / O interconnect bus, although the scope of the embodiments is not limited thereto.
[0066] As Figure 7 shown, various I / O devices 1014 (e.g., speakers, cameras, sensors) can be coupled to the first bus 1016 together with a bus bridge 1018, which can couple the first bus 1016 to the second bus 1020. In one embodiment, the second bus 1020 can be a Low Pin Count (LPC) bus. In one embodiment, various devices can be coupled to the second bus 1020, including, for example, a keyboard / mouse 1012, a communication device(s) 1026, and a data storage unit 1019 such as a disk drive or other mass storage device that can include code 1030.
[0067] The illustrated code 1030, which can be similar to code 713 ( Figure 6 ), can implement one or more aspects of the method 300 ( Figure 2 ) that has been discussed. Further, the audio I / O 1024 can be coupled to the second bus 1020, and the battery port 1010 can supply power to the computing system 1000.
[0068] Note that other embodiments are contemplated. For example, the system can implement a multi-branched bus or another such communication topology instead of Figure 7 the point-to-point architecture. Additionally, more or fewer integrated chips than those Figure 7 shown can be used alternatively to partition the Figure 7 elements.
[0069] Additional notes and examples:
[0070] Example 1 can include a system for managing virtual machines, including a cloud server platform that includes a substrate and logic coupled to the substrate, where the logic is for: configuring one or more first logical cores associated with a first virtual machine of the cloud server platform, where the configuration of the one or more first logical cores is at least partially based on one or more first feature settings; determining, at least in part based on application runtime requirements, one or more active feature settings specific to an active application associated with an active virtual machine, where the active feature settings are different from the first feature settings; configuring one or more active logical cores associated with the active virtual machine of the cloud server platform, where the configuration of the one or more active logical cores is at least partially based on the one or more active feature settings.
[0071] Example 2 may include the system of Example 1, where the application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, and the data reference pattern includes an indication of whether the active application is a streaming application.
[0072] Example 3 may include the system of Example 1, where the active feature settings include an idle power state setting, an execution power state setting, or a prefetch setting.
[0073] Example 4 may include the system of Example 1, where the logic is to: on a per-virtual machine basis, independently assign active feature settings to active virtual machines, and independently assign first feature settings to a first virtual machine.
[0074] Example 5 may include the system of Example 4, where the logic is to: iteratively identify rescheduling events, and in response to the identified rescheduling events, iteratively schedule the active virtual machines to run on a specific set of logical cores that includes one or more active logical cores.
[0075] Example 6 may include the system of Example 1, where the logic is to: store a bitmask of the active feature settings in a control field controlled by the virtual machine associated with the active virtual machine.
[0076] Example 7 may include the system of Example 6, where the logic is to: pass a bitmask of the active feature settings specific to the active application from the control field controlled by the virtual machine, and implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
[0077] Example 8 may include a method for managing virtual machines, including configuring one or more first logical cores associated with a first virtual machine of a cloud server platform, where configuring the one or more first logical cores is at least partially based on one or more first feature settings; determining one or more active feature settings specific to an active application associated with an active virtual machine at least partially based on application runtime requirements, where the active feature settings are different from the first feature settings; and configuring one or more active logical cores associated with the active virtual machine of the cloud server platform, where configuring the one or more active logical cores is at least partially based on the one or more active feature settings.
[0078] Example 9 may include the method of Example 8, where the application runtime requirements include one or more of an indication of latency sensitivity and an indication of a data reference pattern, where the data reference pattern includes an indication of whether the active application is a streaming application; and where the active feature settings include one or more of the following: an idle power state setting, an execution power state setting, or a prefetch setting.
[0079] Example 10 may include the method of Example 8, further comprising: on a per-virtual machine basis, independently assigning an active feature setting to an active virtual machine and independently assigning a first feature setting to a first virtual machine; iteratively identifying rescheduling events; and in response to an identified rescheduling event, iteratively scheduling the active virtual machine to run on a particular set of logical cores that includes one or more active logical cores.
[0080] Example 11 may include the method of Example 8, further comprising: storing a bitmask of an active feature setting in a control field controlled by a virtual machine associated with the active virtual machine, passing the bitmask of the active feature setting specific to the active application from the control field controlled by the virtual machine, and implementing the active feature setting specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
[0081] Example 12 may include at least one computer-readable storage medium including a set of instructions that, when executed by a computing system, cause the computing system to: configure one or more first logical cores associated with a first virtual machine of a cloud server platform, wherein the configuration of the one or more first logical cores is at least partially based on one or more first feature settings; determine, at least in part based on application runtime requirements, one or more active feature settings specific to an active application associated with the active virtual machine, wherein the active feature settings are different from the first feature settings; and configure one or more active logical cores associated with the active virtual machine of the cloud server platform, wherein the configuration of the one or more active logical cores is at least partially based on the one or more active feature settings.
[0082] Example 13 may include the at least one computer-readable storage medium of Example 12, wherein the application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, wherein the data reference pattern includes an indication of whether the active application is a streaming application; and wherein the active feature settings include: an idle power state setting, an execution power state setting, or a prefetch setting.
[0083] Example 14 may include the at least one computer-readable storage medium of Example 12, wherein the instructions, when executed, cause the computing system to: on a per-virtual machine basis, independently assign an active feature setting to the active virtual machine and independently assign a first feature setting to the first virtual machine; iteratively identify rescheduling events; and in response to an identified rescheduling event, iteratively schedule the active virtual machine to run on a particular set of logical cores that includes one or more active logical cores.
[0084] Example 15 may include at least one computer-readable storage medium of Example 12, wherein when the instructions are executed, cause the computing system to: store a bitmask of the active feature settings in a control field controlled by the virtual machine associated with the active virtual machine, transfer the bitmask of the active feature settings specific to the active application from the control field controlled by the virtual machine, and implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
[0085] Example 16 may include an apparatus for managing virtual machines, including: a first virtual machine in a cloud server platform, the first virtual machine including one or more first logical cores, the first virtual machine having a first configuration for efficiently supporting a first feature settings arrangement on the one or more first logical cores; a feature manager for determining, at least in part based on application runtime requirements, an active feature settings arrangement specific to an active application, wherein the active feature settings arrangement is different from the first feature settings arrangement; and an active virtual machine of the cloud server platform, the active virtual machine associated with the active application, the active virtual machine including one or more active logical cores, the active virtual machine having an active configuration for efficiently supporting an active feature settings configuration on the one or more active logical cores.
[0086] Example 17 may include the apparatus of Example 16, further including an active application associated with the active virtual machine, wherein the active application is associated with application runtime requirements, wherein the application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, wherein the data reference pattern includes an indication of whether the active application is a streaming application.
[0087] Example 18 may include the apparatus of Example 16, wherein the active feature settings include an idle power state setting, an execution power state setting, or a prefetch setting.
[0088] Example 19 may include the apparatus of Example 16, further including a virtual machine monitor for: on a per-virtual machine basis, independently assign active feature settings to the active virtual machine and independently assign the first feature settings to the first virtual machine; iteratively identify rescheduling events; and in response to the identified rescheduling events, iteratively schedule the active virtual machine to run on a specific set of logical cores including one or more active logical cores.
[0089] Example 20 may include the apparatus of Example 16, further including virtual machine control, wherein the virtual machine control includes a control field for storing a bitmask of active feature settings, the virtual machine control being associated with an active virtual machine; wherein one or more active logical cores associated with the active virtual machine are configured to: transfer a bitmask of active feature settings specific to an active application from the control field of the virtual machine control, and implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
[0090] Example 21 may include the apparatus of Example 16, further including an active application associated with the active virtual machine, wherein the active application is associated with application runtime requirements, wherein the application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, wherein the data reference pattern includes an indication of whether the active application is a streaming application; wherein the active feature settings include an idle power state setting, an execution power state setting, or a prefetch setting; a virtual machine monitor configured to: on a per-virtual machine basis, independently assign the active feature settings to the active virtual machine, and independently assign a first feature setting to a first virtual machine, iteratively identify rescheduling events, and in response to the identified rescheduling events, iteratively schedule the active virtual machine to run on a specific set of logical cores including one or more active logical cores; virtual machine control, wherein the virtual machine control includes a control field for storing a bitmask of active feature settings, the virtual machine control being associated with the active virtual machine; wherein one or more active logical cores associated with the active virtual machine are configured to: transfer a bitmask of active feature settings specific to an active application from the control field of the virtual machine control, and implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
[0091] Example 22 may include a device including means for performing the method described in any of the foregoing examples.
[0092] Example 23 may include a machine-readable storage including machine-readable instructions that, when executed, perform the method described in any of the foregoing examples or implement the apparatus described in any of the foregoing examples.
[0093] Each embodiment is applicable for use with various types of semiconductor integrated circuit (“IC”) chips. Examples of such IC chips include, but are not limited to, processors, controllers, chipset components, programmable logic arrays (PLAs), memory chips, network chips, system-on-chip (SoC), SSD / NAND controller ASICs, and the like. Additionally, in some of the figures, signal conductors are represented by lines. Some lines may be different to represent more constitutive signal paths, having numeric labels to denote the number of constitutive signal paths, and / or having arrows at one or more ends to denote the primary information flow direction. However, this should not be construed in a limiting manner. Instead, this added detail may be used in conjunction with one or more exemplary embodiments to more easily understand the circuitry. Any represented signal line, whether or not having additional information, may in fact include one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, such as digital or analog lines implemented using differential pairs, fiber optic lines, and / or single-ended lines.
[0094] Example dimensions / models / values / ranges may have been given, but the embodiments are not limited thereto. As manufacturing technology (e.g., lithography) matures over time, it is expected that devices of smaller dimensions can be manufactured. Additionally, for simplicity of illustration and explanation, well-known power / ground connections and other components of the IC chips may or may not be shown in the figures, and doing so is also to avoid obscuring certain aspects of the embodiments. Furthermore, various configurations may be shown in block diagram form to avoid obscuring the embodiments, and in view of the fact that the specific details of the implementation relative to these block diagram configurations largely depend on the platform on which the embodiments are implemented, i.e., these specific details should be within the purview of those skilled in the art. In cases where specific details (e.g., circuitry) are set forth to describe exemplary embodiments, it is apparent that those skilled in the art can implement the embodiments without or with variations to these specific details. The specification is thus to be regarded as illustrative rather than restrictive.
[0095] The term “coupled” is used herein to denote any type of direct or indirect relationship between the components being discussed and can apply to electrical, mechanical, fluidic, optical, electromagnetic, electromechanical, or other connections. Additionally, the terms “first,” “second,” etc. are used herein solely for ease of discussion and have no temporal or chronological significance unless otherwise stated.
[0096] As used in this application and the claims, a list of items joined by the term “one or more” may mean any combination of the listed items. For example, the phrase “one or more of A, B, or C” may mean A; B; C; A and B; A and C; B and C; or A, B, and C.
[0097] Those skilled in the art will understand from the foregoing description that the broad techniques of the embodiments can be implemented in many forms. Accordingly, while the embodiments have been described in connection with their specific examples, the true scope of the embodiments is not so limited, since other modifications will become readily apparent to those skilled in the art after study of the drawings, specification, and following claims.
Claims
1. A system for managing virtual machines, comprising: A cloud server platform, the cloud server platform including a substrate and logic coupled to the substrate, wherein the logic is operative to: Configure one or more first logical cores associated with a first virtual machine of the cloud server platform, wherein the configuration of the one or more first logical cores is at least partially based on one or more first feature settings; Determine, at least in part based on application runtime requirements, one or more active feature settings specific to an active application associated with an active virtual machine, wherein the active feature settings are different from the first feature settings; Configure one or more active logical cores associated with the active virtual machine of the cloud server platform, wherein the configuration of the one or more active logical cores is at least partially based on the one or more active feature settings; Store a bitmask of the active feature settings in a control field of a virtual machine control associated with the active virtual machine; And A power source for providing power to the cloud server platform.
2. The system according to claim 1, wherein The application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, wherein the data reference pattern includes an indication of whether the active application is a streaming application.
3. The system according to claim 1, characterized in that, The active feature settings include an idle power state setting, an execution power state setting, or a prefetch setting.
4. The system according to claim 1, wherein The logic is operative to: On a virtual machine-by-virtual machine basis, independently assign the active feature settings to the active virtual machine and independently assign the first feature settings to the first virtual machine.
5. The system according to claim 4, characterized in that, The logic is operative to: Iteratively identify rescheduling events; and In response to the identified rescheduling events, iteratively schedule the active virtual machine to run on a specific set of logical cores including the one or more active logical cores.
6. The system according to claim 1, characterized in that, The logic is operative to: Transfer the bitmask of the active feature settings specific to the active application from the control field of the virtual machine control; And Implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and active application.
7. A method for managing virtual machines, comprising: Configuring one or more first logical cores associated with a first virtual machine of a cloud server platform, wherein configuring the one or more first logical cores is at least partially based on one or more first feature settings; Determining, at least in part based on application runtime requirements, one or more active feature settings specific to an active application associated with an active virtual machine, wherein the active feature settings are different from the first feature settings; Configuring one or more active logical cores associated with the active virtual machine of the cloud server platform, wherein configuring the one or more active logical cores is at least partially based on the one or more active feature settings; And Storing a bitmask of the active feature settings in a control field of a virtual machine control associated with the active virtual machine.
8. The method according to claim 7, wherein The application runtime requirements include one or more of an indication of latency sensitivity and an indication of a data reference pattern, where the data reference pattern includes an indication of whether the active application is a streaming application, and where the active feature settings include one or more of the following: an idle power state setting, an execution power state setting, or a prefetch setting.
9. The method according to claim 7, wherein Further comprising: On a per-virtual machine basis, independently assign the active feature settings to the active virtual machine and independently assign the first feature settings to the first virtual machine; Iteratively identify rescheduling events; And In response to the identified rescheduling events, iteratively schedule the active virtual machine to run on a specific set of logical cores including the one or more active logical cores.
10. The method according to claim 7, wherein Further comprising: Transfer the bitmask of the active feature settings specific to the active application from the control field controlled by the virtual machine; And Implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
11. An apparatus for managing virtual machines, comprising: A first virtual machine of a cloud server platform, the first virtual machine including one or more first logical cores, the first virtual machine having a first configuration for efficiently supporting a first feature settings arrangement on the one or more first logical cores; A feature manager for determining, at least in part based on application runtime requirements, an active feature settings arrangement specific to an active application, where the active feature settings arrangement is different from the first feature settings arrangement; The active virtual machine of the cloud server platform, the active virtual machine being associated with the active application, the active virtual machine including one or more active logical cores, the active virtual machine having an active configuration for efficiently supporting the active feature settings configuration on the one or more active logical cores; And A virtual machine control, where the virtual machine control includes a control field for storing a bitmask of the active feature settings, the virtual machine control being associated with the active virtual machine.
12. The device according to claim 11, wherein Further comprising: An active application associated with the active virtual machine, where the active application is associated with application runtime requirements, where the application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, where the data reference pattern includes an indication of whether the active application is a streaming application.
13. The device according to claim 11, characterized in that, The active feature settings include an idle power state setting, an execution power state setting, or a prefetch setting.
14. The device according to claim 11, wherein, Further comprising: A virtual machine monitor for: On a per-virtual machine basis, independently assign the active feature settings to the active virtual machine and independently assign the first feature settings to the first virtual machine, Iteratively identify rescheduling events, and In response to the identified rescheduling event, iteratively schedule the active virtual machine to run on a specific set of logical cores that includes the one or more active logical cores.
15. The device according to claim 11, characterized in that, One or more active logical cores associated with the active virtual machine are configured to: Transfer the bitmask for the active feature settings specific to the active application from the control field controlled by the virtual machine, and Implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
16. The device according to claim 11, wherein Further comprising: An active application associated with the active virtual machine, wherein the active application is associated with application runtime requirements, and wherein the application runtime requirements include at least one of an indication of latency sensitivity and an indication of a data reference pattern, and wherein the data reference pattern includes an indication of whether the active application is a streaming application; Wherein the active feature settings include an idle power state setting, an execution power state setting, or a prefetch setting; A virtual machine monitor configured to: On a virtual machine-by-virtual machine basis, independently assign the active feature settings to the active virtual machine and independently assign the first feature settings to the first virtual machine, Iteratively identify rescheduling events, In response to the identified rescheduling events, iteratively schedule the active virtual machine to run on a specific set of logical cores that includes the one or more active logical cores; Wherein one or more active logical cores associated with the active virtual machine are configured to: Transfer the bitmask for the active feature settings specific to the active application from the control field controlled by the virtual machine, and Implement the active feature settings specific to the active application by adapting the operation of the active logical cores to support the active virtual machine and the active application.
17. A device for managing virtual machines, comprising means for performing the method according to any one of claims 7 - 10.
18. A machine-readable storage including machine-readable instructions that, when executed, perform the method according to any one of claims 7 - 10 or implement the device according to any one of claims 11 - 16.
Citation Information
Patent Citations
Virtual machine and / or multi-level scheduling support on systems with asymmetric processor cores
CN102402458A