Memory bandwidth balancing method, host machine, electronic device, and storage medium
By using different memory buses for memory access in virtualization technology, and balancing the virtual machine to the processing unit to match the memory bandwidth requirements, the memory bandwidth competition among virtual machines is solved, and the allocation reliability of memory bandwidth and the operating performance of virtual machines are improved.
Patent Information
- Application Number
- PCT/IB2024/063269
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2024-12-30
- Publication Date
- 2025-08-07
AI Technical Summary
In virtualization technology, the competition between virtual machines leads to poor memory bandwidth allocation reliability, poor virtual machine operation performance and tenant experience.
By determining the memory bandwidth requirements of multiple virtual machines for processing units, using different memory buses for memory access, and balancing each virtual machine to the corresponding processing unit, so that its memory bandwidth occupies is matched with the requirements, and weighted computing and resource-oriented technology are used to adjust the memory bandwidth upper limit.
Improves the allocation reliability of memory bandwidth, improves the operating performance and tenant experience of virtual machines.
Smart Images

Figure IB2024063269_07082025_PF_FP_ABST
Abstract
Description
[0001] Memory Bandwidth Balancing Method, Host Machine, Electronic Device, and Storage Medium TECHNICAL FIELD Embodiments of the present invention relate to the field of computer technology, and more particularly to a memory bandwidth balancing method, host machine, electronic device, and storage medium. Background In virtualization technology, a host machine, as a physical machine in the operating environment of a hypervisor, can run multiple virtual machine instances. The host machine provides infrastructure such as computing resources, storage, and networking to support the operation of virtual machines. The host machine is responsible for allocating and managing physical resources and providing an execution environment for virtual machines, including the allocation and scheduling of resources such as CPU, memory, storage, and networking. The host machine typically has higher computing power and hardware configuration to meet the needs of multiple virtual machine instances. When each virtual machine is running, if they do not exclusively occupy the entire memory bandwidth of the entire host machine or the entire memory bandwidth of a separate memory bus, contention for memory bandwidth may occur between different virtual machines. This results in poor memory bandwidth allocation reliability, poor virtual machine operating performance, and a poor user experience for virtual machine tenants. SUMMARY OF THE INVENTION In view of this, embodiments of the present invention provide a memory bandwidth balancing method, host machine, electronic device, and storage medium to at least partially address the aforementioned issues. According to a first aspect of an embodiment of the present invention, a memory bandwidth balancing method is provided, comprising: determining the memory bandwidth requirements of multiple virtual machines for at least one processing unit, wherein different processing units access memory through different memory buses; and evenly scheduling each virtual machine to a corresponding processing unit within the at least one processing unit, such that the memory bandwidth occupied by the virtual machine for the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine. In another implementation of the present invention, determining the memory bandwidth requirements of the multiple virtual machines for the at least one processing unit comprises: obtaining the historical memory bandwidth occupied by the multiple virtual machines in the invoked processing unit; performing a weighted calculation on the historical memory bandwidth occupied by each of the multiple virtual machines to obtain scheduling weight information for each of the multiple virtual machines; and determining the memory bandwidth requirements of the multiple virtual machines for the at least one processing unit based on the scheduling weight information for each of the multiple virtual machines. In another implementation of the present invention, determining the memory bandwidth requirements of the multiple virtual machines for the at least one processing unit based on the scheduling weight information for each of the multiple virtual machines comprises: determining a memory bandwidth percentage that matches the scheduling weight information for each virtual machine; and determining the memory bandwidth requirement of each virtual machine based on the memory bandwidth percentage of each virtual machine and the total memory bandwidth of the at least one processing unit.In another implementation of the present invention, balancing scheduling each virtual machine to a corresponding processing unit in the at least one processing unit includes: when the host machine on which the multiple virtual machines are deployed is in a first resource state, dynamically scheduling different virtual machines with consistent memory bandwidth requirements to different processing units in the at least one processing unit, wherein the computing resource occupancy in the first resource state is less than a first preset occupancy. In another implementation of the present invention, balancing scheduling each virtual machine to a corresponding processing unit in the at least one processing unit further includes: when the host machine is in a second resource state, scheduling at least two virtual machines with consistent memory bandwidth requirements to the same processing unit in the at least one processing unit, wherein the computing resource occupancy in the second resource state is greater than a second preset occupancy, and the second preset occupancy is greater than the first preset occupancy. In another implementation of the present invention, the at least two virtual machines include a first virtual machine and a second virtual machine, and the method further includes: determining a change in a bandwidth occupancy ratio of a memory bus of the same processing unit by the first virtual machine and the second virtual machine; and setting a memory bandwidth upper limit for at least one of the first virtual machine and the second virtual machine on the same processing unit based on the change in the bandwidth occupancy ratio. In another implementation of the present invention, determining a change in bandwidth usage ratio of the first virtual machine and the second virtual machine on the same processing unit's memory bus includes: monitoring a first change in the memory bandwidth usage ratio of the first virtual machine and a second change in the memory bandwidth usage ratio of the second virtual machine; and setting a memory bandwidth upper limit for at least one of the first and second virtual machines on the same processing unit based on the change in the bandwidth usage ratio, including: if the first change in the bandwidth usage ratio and the second change in the bandwidth usage ratio are consistent, setting the memory bandwidth upper limit for the first and second virtual machines to evenly divide the memory bandwidth resources of the same processing unit. In another implementation of the present invention, setting a memory bandwidth upper limit for at least one of the first and second virtual machines on the same processing unit based on the change in the bandwidth usage ratio also includes: if the first change in the bandwidth usage ratio and the second change in the bandwidth usage ratio are inconsistent, setting the memory bandwidth upper limit for at least one of the first and second virtual machines on the same processing unit to less than half the memory bandwidth resources of the processing unit during a period adjacent to a peak moment of the first change in the bandwidth usage ratio.According to a second aspect of an embodiment of the present invention, an application deployment method is provided, comprising: determining memory bandwidth requirements of multiple microservices of an application, wherein the memory bandwidth requirement of each microservice indicates the bandwidth requirement of the microservice for a memory bus in the host machine where it resides; and deploying the multiple microservices to multiple host machines based on the memory bandwidth requirements of the multiple microservices, such that the total memory bandwidth requirement of the microservices deployed in each host machine matches the memory bandwidth upper limit of the host machine. According to a third aspect of an embodiment of the present invention, a host machine is provided, comprising: determining memory bandwidth requirements of multiple virtual machines for at least one processing unit, wherein different processing units access memory through different memory buses; and a scheduling module that evenly schedules each virtual machine to a corresponding processing unit of the at least one processing unit, such that the memory bandwidth occupied by each virtual machine for the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine. According to a fourth aspect of an embodiment of the present invention, an application deployment apparatus is provided, comprising: determining memory bandwidth requirements of multiple microservices of an application, wherein the memory bandwidth requirement of each microservice indicates the bandwidth requirement of the microservice for a memory bus in the host machine in which it resides; and a deployment module, based on the memory bandwidth requirements of the multiple microservices, deploying the multiple microservices to multiple host machines, such that the total memory bandwidth requirement of the microservices deployed in each host machine matches the memory bandwidth upper limit of the host machine. According to a fifth aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is configured to store at least one executable instruction, wherein the executable instruction causes the processor to perform operations corresponding to the method described in the first or second aspect. According to a sixth aspect of an embodiment of the present invention, a computer storage medium is provided, storing a computer program, which, when executed by a processor, implements the method described in the first or second aspect.In the solutions of the embodiments of the present invention, since different processing units access memory through different memory buses, the memory bandwidth demands of multiple virtual machines for at least one processing unit reflect the degree of contention among the multiple virtual machines for the memory bandwidth of the memory bus of the at least one processing unit to be scheduled. Consequently, when each virtual machine is evenly scheduled to a corresponding processing unit within the at least one processing unit, the memory bandwidth occupied by the virtual machine for the memory bus of the corresponding processing unit matches the memory bandwidth demand of the virtual machine. This ensures that virtual machines with greater memory bandwidth demands are allocated greater memory bandwidth, thereby improving the reliability of memory bandwidth allocation, enhancing the operating performance of the virtual machines, and enhancing the user experience of virtual machine tenants. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following briefly introduces the figures required for use in the embodiments or the prior art description. Obviously, the figures described below are only some of the embodiments described in the embodiments of the present invention. Those skilled in the art can also derive other figures based on these figures. FIG. 1 is a schematic block diagram of a host machine in a cloud service system according to some embodiments of the present invention. FIG. 2 is a flow chart of the steps of a memory bandwidth balancing method according to some embodiments of the present invention. FIG. 3 is a flow chart of the steps of a memory bandwidth balancing method according to some other embodiments of the present invention. Figure 4 is a flowchart of a specific example of the memory bandwidth balancing method of the embodiment of Figure 2 . Figure 5 is a block diagram of the host machine structure of other embodiments of the present invention. Figure 6 is a flow chart of the application deployment method of other embodiments of the present invention. Figure 7 is a flow chart of the application deployment apparatus of other embodiments of the present invention. Figure 8 is a schematic diagram of the electronic device structure of other embodiments of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS To help those skilled in the art better understand the technical solutions of the embodiments of the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a portion of the embodiments of the present invention, not all of them. All other embodiments derived by those skilled in the art based on the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention. The specific implementation of the embodiments of the present invention will be further described below in conjunction with the accompanying drawings of the embodiments of the present invention. Figure 1 is a schematic block diagram of a host machine in a cloud service system of some embodiments of the present invention. The cloud service system includes a control node such as a cloud management system (CMS), a configuration node such as a network controller, and computing nodes such as server hosts.As shown in Figure 1, a host machine, acting as a server host, can be locally configured with virtual machines #1, #2, ..., and #N for tenant use. These virtual machines include, but are not limited to, Java virtual machines and container objects such as PODs. The host machine can also be configured with a virtual machine management module, such as a hypervisor, configured to manage each virtual machine. Configuration nodes, such as network controllers, facilitate communication between virtual machines through the virtual machine management module. Control nodes can manage virtual machines by creating, deploying, or orchestrating virtual machines, enabling elastic computing or flexible application deployment. Furthermore, in a microservices architecture, applications can be split into microservices as relatively independent application modules, with inter-process communication implemented between the microservices. Furthermore, when a virtual machine requires computing resources from the host machine to execute a program process from the virtual machine, the virtual machine is dispatched to a processing unit, such as processing unit #1 or processing unit #2, via the virtual machine management module. This allows data access between the processing unit and memory via the corresponding bus to execute the application in the virtual machine. For example, processing unit #1 accesses data from memory via memory bus #1, and processing unit #2 accesses data from memory via memory bus #2. It should be understood that the number of processing units capable of independently utilizing buses to exchange data with memory is illustrative only; without loss of generality, data access to memory may be performed using only one or more buses. The number of processing units herein may correspond to the number of buses capable of independently accessing memory. For the X87 processor architecture, processing units can be implemented as computing resources corresponding to processing unit sockets, which can access data from memory via independent buses. Processing units comprise a processor such as a CPU. Physically, each processing unit may include at least one processing core, and logically, each virtual machine may be configured with at least one virtual machine processing core. Furthermore, in a cloud service system deploying multiple host machines, virtual machines can connect to an external physical network (e.g., a switch) via ports using corresponding network adapters to forward data packets within the physical network. The physical network connects between the host machines to facilitate data forwarding between virtual machines in different host machines. Specifically, each virtual machine can occupy computing resources in a local host machine, and sometimes can also occupy computing resources of a remote host machine. Virtual machine #1, virtual machine #2, ..., virtual machine #N can occupy computing resources of a local host machine.When virtual machine #1 monopolizes the computing resources of processing unit #1, virtual machine #2 and others are scheduled to processing unit #2 or a processing unit on a remote host. Because processing unit #1 can independently access memory data via memory bus #1, there is no competition with other virtual machines, such as virtual machine #2, for the memory bandwidth of memory bus #1. However, when virtual machine #1 and virtual machine #2 are scheduled to processing unit #1, the memory bandwidth of memory bus #1 is occupied by both virtual machines #1 and #2. In this case, if both virtual machines #1 and #2 have high memory bandwidth requirements, contention for memory bus #1 memory bandwidth resources will occur, resulting in poor memory bandwidth management reliability, poor virtual machine performance, and a poor user experience for virtual machine tenants. Figure 2 is a flowchart of the steps of a memory bandwidth balancing method according to some embodiments of the present invention. Specifically, the memory bandwidth balancing method includes:
[0002] S210: Determine the memory bandwidth requirements of multiple virtual machines for at least one processing unit, with different processing units accessing memory via different memory buses. It should be understood that memory bandwidth in this context refers to the access rate at which a processor, such as a processing unit, accesses data stored in memory (e.g., RAM) by a virtual machine program, for example, the number of bits or bytes that can be transferred per second. It should also be understood that, in some examples, the memory bandwidth requirement may indicate the relative size of each virtual machine's memory bandwidth requirement compared to the memory bandwidth requirements of other virtual machines. For example, the requirement may be represented by the ratio of each virtual machine's memory bandwidth requirement to the total memory bandwidth requirement of all virtual machines, the urgency of the memory bandwidth requirement, the sensitivity of the memory bandwidth requirement, or a combination of these indicators. Specifically, a virtual machine with a more urgent memory bandwidth requirement may have a greater memory bandwidth requirement, while a virtual machine with a less urgent memory bandwidth requirement may have a smaller memory bandwidth requirement. A virtual machine with a high sensitivity to memory bandwidth requirement may have a larger memory bandwidth requirement, while a virtual machine with a low sensitivity to memory bandwidth requirement may have a smaller memory bandwidth requirement. The following example uses the ratio between each VM's memory bandwidth requirement and the total memory bandwidth requirement of all VMs as an example. Specifically, assuming that the memory bandwidth sensitivity of each VM is similar, or ignoring the memory bandwidth sensitivity of each VM, and that there are 10 VMs on a local host, and the total memory bandwidth requirement is 20 GB / s, the average memory bandwidth requirement per VM is 2 GB / s. Therefore, if a VM's memory bandwidth requirement is less than 1 GB / s, it indicates a low memory bandwidth requirement. If a VM's memory bandwidth requirement is greater than 5 GB / s and more than 2 GB / s, it indicates a high memory bandwidth requirement.
[0003] S220: Each virtual machine is evenly scheduled to a corresponding processing unit in at least one processing unit, such that the memory bandwidth occupied by the virtual machine on the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine. It should be understood that the memory bandwidth occupied by the virtual machine on the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine. In other words, a virtual machine with a greater memory bandwidth requirement occupies a greater amount of memory bandwidth on the memory bus of the corresponding processing unit. In other words, because different processing units access memory through different memory buses, two virtual machines scheduled to a processing unit corresponding to the same memory bus may experience memory bandwidth contention, while two virtual machines scheduled to processing units corresponding to different memory buses will not experience memory bandwidth contention. By adjusting factors that may or may not cause memory bandwidth contention, multiple virtual machines are scheduled to at least one processing unit, thereby meeting the memory bandwidth requirements of each of the multiple virtual machines. It should also be understood that the number of virtual machines in a processing unit can indicate the memory bandwidth occupied by the processing unit's memory bus. That is, the greater the number of virtual machines in a processing unit, the smaller the proportion of bandwidth occupied by the virtual machines on the processing unit's memory bus, and the smaller the memory bandwidth occupied by the virtual machines on the memory bus of the scheduled processing unit. For example, when the memory bandwidth occupied by a virtual machine on the corresponding processing unit's memory bus matches the virtual machine's memory bandwidth requirement, the virtual machine's memory bandwidth requirement matches the number of virtual machines in the processing unit where the virtual machine resides. In the solution of this embodiment of the present invention, since different processing units access memory through different memory buses, the memory bandwidth requirements of multiple virtual machines on at least one processing unit reflect the degree of contention among the multiple virtual machines for the memory bandwidth of the memory bus of the at least one processing unit to be scheduled. Consequently, when each virtual machine is evenly scheduled to a corresponding processing unit within the at least one processing unit, the memory bandwidth occupied by the virtual machine on the corresponding processing unit's memory bus matches the virtual machine's memory bandwidth requirement. This ensures that virtual machines with greater memory bandwidth requirements are allocated larger memory bandwidths, thereby improving the reliability of memory bandwidth allocation, enhancing the operating performance of the virtual machines, and enhancing the user experience of virtual machine tenants. In other embodiments, as an example of determining the memory bandwidth requirements of multiple virtual machines, historical memory bandwidth data of each of the multiple virtual machines may be obtained; weighted calculations may be performed on the historical memory bandwidth data of each of the multiple virtual machines to obtain scheduling weight information of each of the multiple virtual machines; and memory bandwidth requirements of the multiple virtual machines for at least one processing unit may be determined based on the scheduling weight information of each of the multiple virtual machines.Furthermore, a user profile of a virtual machine tenant can be generated based on the historical memory bandwidth data of the virtual machines. Dimensions of the user profile include, but are not limited to, the ratio of each virtual machine to the total memory bandwidth demand of all virtual machines, a representation of the urgency of the memory bandwidth demand, a representation of the sensitivity of the memory bandwidth demand, the service priority of the virtual machine, and the tenant priority. Furthermore, a weighted calculation can be performed on the historical memory bandwidth data of each of the multiple virtual machines based on a memory bandwidth demand index to obtain scheduling weight information for each of the multiple virtual machines. For example, the memory bandwidth demand index can be a preset weight index for each of the aforementioned dimensions. Accordingly, calculations can be performed on each dimension of the historical memory bandwidth data of each virtual machine based on the preset weight index for each dimension to obtain scheduling weight information for each of the multiple virtual machines. In other words, a virtual machine with a higher scheduling weight has a greater memory bandwidth demand. Alternatively, or optionally, when determining the memory bandwidth demands of the multiple virtual machines, the memory bandwidth demands of the multiple virtual machines can be determined based on the current service traffic of the applications deployed to the multiple virtual machines. For example, the greater the current service traffic of the deployed applications, the greater the memory bandwidth demand of the virtual machine. Specifically, in high-concurrency access scenarios, the current business traffic of an application is high, while in low-concurrency access scenarios, the current business traffic is low. This leads to different memory bandwidth requirements for virtual machines deploying different applications or different microservices of the same application. It should be understood that the current business traffic of the application in the virtual machine can be used as one weighting factor for the scheduling weight information, and the historical memory bandwidth occupied by the virtual machine can be used as another weighting factor to determine the scheduling weight information for each virtual machine. For example, based on the memory bandwidth requirement indicator, a weighted calculation is performed on the historical memory bandwidth occupied data and the current business traffic of multiple virtual machines to obtain the scheduling weight information for each of the multiple virtual machines. Furthermore, the weights between the different weighting factors can be set to be equal or different. Furthermore, to determine the memory bandwidth requirements of the multiple virtual machines for at least one processing unit based on the scheduling weight information of each of the multiple virtual machines, a memory bandwidth share that matches the scheduling weight information of each virtual machine can be determined. Then, based on the memory bandwidth share of each virtual machine and the total memory bandwidth of the at least one processing unit, the memory bandwidth requirement of each virtual machine is determined. For example, the memory bandwidth requirement of each virtual machine is obtained by multiplying the memory bandwidth ratio of each virtual machine by the total memory bandwidth of at least one processing unit.It should also be understood that the scheduling weight information of multiple virtual machines can be implemented as a scheduling weight interval, that is, each virtual machine can correspond to a scheduling weight interval. For example, the higher the weight value in the scheduling weight interval, the greater the weight of the virtual machine. In other embodiments, evenly scheduling each virtual machine to a corresponding processing unit in the at least one processing unit includes: dynamically scheduling different virtual machines with consistent memory bandwidth requirements to different processing units when the host machine deploying the multiple virtual machines is in a first resource state, wherein the calculated occupancy rate of the first resource state is less than a first preset occupancy rate. In other words, the first preset occupancy rate can indicate the resource shortage level of the host machine. When the host machine's resource shortage level is not high, different virtual machines with consistent memory bandwidth requirements can be dynamically scheduled to different processing units. For example, as an example of different virtual machines having consistent memory bandwidth requirements, if the weight values of the scheduling weight information of at least two virtual machines are the same or similar, then the memory bandwidth requirements of the at least two virtual machines are consistent. Furthermore, if the at least two virtual machines belong to the same scheduling weight interval, then the memory bandwidth requirements of the at least two virtual machines are consistent. For another example, different virtual machines with scheduling weights higher than a preset weight can be dynamically scheduled to different processing units, while virtual machines with scheduling weights lower than the preset weight can be randomly scheduled. This allows virtual machines with relatively high memory bandwidth requirements to be allocated to larger memory bandwidth allocations when the host machine's resource status is not limited. It should be understood that resource status includes two dimensions: storage resource status and computing resource status. The storage resource status occupancy and the computing resource status occupancy can be weighted to obtain the computing resource occupancy. More specifically, as an example of dynamically scheduling different virtual machines with consistent memory bandwidth requirements to different processing units, for example, a non-running to-be-scheduled virtual machine can be directly scheduled to a different processing unit from the scheduled virtual machines. Alternatively, if at least two scheduled virtual machines are running, some of the scheduled virtual machines can be migrated so that different scheduled virtual machines are placed in different processing units. In other embodiments, if the processing module of at least one processing unit uses a processor architecture such as AMD, for example, the processing module includes two processing units (i.e., each processing unit corresponds to a socket), each processing unit includes 128 virtual processing cores, and each processing unit includes 8 processing sub-units, then the processing module accordingly includes 16 processing sub-units. In other words, each processing sub-unit includes 16 virtual processing cores.If a virtual machine has 128 virtual processing cores, each processing subunit can be allocated 8 virtual processing cores, meaning that the virtual machine's memory bandwidth can be expanded by up to 100%. Without loss of generality, different virtual machines with consistent memory bandwidth requirements can be dynamically scheduled to multiple processing subunits to ensure balanced memory bandwidth usage by each processing subunit. As another example of evenly scheduling each virtual machine to a corresponding processing unit for processing, when multiple virtual machines are in a second resource state of the host machine, at least two virtual machines with consistent memory bandwidth requirements can be scheduled to the same processing unit, where the computing resource usage in the second resource state is greater than a second preset usage ratio. In other examples, the second preset usage ratio is greater than the first preset usage ratio, thereby achieving more reliable memory bandwidth allocation based on the resource availability of the host machine. For example, the at least two virtual machines include a first virtual machine and a second virtual machine. In the memory bandwidth balancing method, the fluctuation in the bandwidth usage ratio of the first virtual machine and the second virtual machine on the memory bus of the same processing unit can also be determined; and based on the fluctuation in the bandwidth usage ratio, an upper limit on the memory bandwidth usage of at least one of the first virtual machine and the second virtual machine on the same processing unit can be set. For example, Resource Director Technology (RDT) can be used to adjust the memory bandwidth cap. It should be understood that RDT can help system administrators achieve fine-grained control over memory bandwidth, providing priority management of memory bandwidth usage for different applications and tasks. By using RDT, administrators can reserve sufficient memory bandwidth for critical tasks to ensure their normal operation without being impacted by other tasks. At the same time, RDT can also limit the memory bandwidth usage of low-priority tasks to prevent them from consuming excessive resources and impacting the performance of other tasks. In other words, the memory bandwidth usage change reflects the more fine-grained memory bandwidth demand within the same processing unit. This solution further improves the granularity of memory bandwidth allocation and enhances the computing performance of each virtual machine. For another example, determining the change in the bandwidth usage ratio of the first virtual machine and the second virtual machine on the memory bus of the same processing unit includes monitoring a first change in the memory bandwidth usage of the first virtual machine and a second change in the memory bandwidth usage of the second virtual machine.Setting the upper limit of the memory bandwidth occupied by at least one of the first virtual machine and the second virtual machine on the same processing unit based on the change state of the bandwidth occupancy ratio includes: if the first change state (e.g., the fluctuation state of the occupancy ratio) and the second change state (e.g., the fluctuation state of the occupancy ratio) are consistent, setting the upper limit of the memory bandwidth of the first virtual machine and the second virtual machine to evenly divide the memory bandwidth resources of the same processing unit. In other words, Resource Director Technology (RDT) can be employed to evenly divide the memory bandwidth of the same processing unit. For example, each virtual machine's occupied memory bandwidth can be set to 50% of the processing unit's, i.e., the total memory bandwidth of the memory bus is evenly distributed, thereby further improving the granularity of memory bandwidth allocation and further enhancing the computing performance of each virtual machine. It should be understood that as an example of monitoring the first change state of the memory bandwidth occupied by the first virtual machine and the second change state of the memory bandwidth occupied by the second virtual machine, the change state of the first virtual machine after being scheduled to the processing unit can be monitored, and the first change state after the current moment can be predicted based on the previous change state monitored from the time of scheduling to the current moment. Similarly, the changing state of the second virtual machine after being scheduled to the processing unit can be monitored, and a second changing state after the current moment can be predicted based on the previously monitored changing state from the time of scheduling to the current moment. If the changing states of the first and second virtual machines immediately after scheduling do not show complete peaks or troughs, the total memory bandwidth of the processing unit can be evenly distributed between the first and second virtual machines. After the changing states of both the first and second virtual machines show complete peaks or troughs, dynamic allocation of occupied memory bandwidth can be performed based on whether the first and second changing states are consistent. In other embodiments, setting a memory bandwidth upper limit for at least one of the first and second virtual machines on the same processing unit based on the changing state of bandwidth usage ratios further includes: if the first and second changing states are inconsistent, setting the memory bandwidth upper limit for the second virtual machine to less than half the memory bandwidth resources of the processing unit during a period adjacent to the peak of the first changing state. For example, if the peak and / or trough times of the memory bandwidth usage of two virtual machines are inconsistent, that is, staggered, the staggered peaks and / or troughs can be used to dynamically adjust the memory bandwidth of the two virtual machines. For example, when the first variable state approaches its peak, the upper limit of the memory bandwidth usage of the second virtual machine can be limited to increase the memory bandwidth usage of the first virtual machine.Conversely, when approaching the peak time of the second change state, limiting the upper limit of the memory bandwidth occupied by the first virtual machine increases the memory bandwidth occupied by the second virtual machine, thereby further improving the granularity of memory bandwidth allocation, further enhancing the computing performance of each virtual machine and the virtual machine tenant experience. In summary, coarse-grained memory bandwidth balancing can be performed on virtual machines based on their memory bandwidth requirements, and fine-grained memory bandwidth balancing can be performed based on the change state of their memory bandwidth occupation. Figure 3 is a flowchart of the steps of a memory bandwidth balancing method according to other embodiments. The memory bandwidth balancing method includes:
[0004] S310: Determine memory bandwidth requirements of multiple virtual machines for at least one processing unit, where different processing units access memory through different memory buses.
[0005] S320: When multiple virtual machines are in a first resource state of the host machine, dynamically schedule different virtual machines with consistent memory bandwidth requirements to different processing units, wherein the computing occupancy of the first resource state is less than a first preset occupancy.
[0006] S330: When multiple virtual machines are in the second resource state of the host machine, at least two virtual machines with consistent memory bandwidth requirements are scheduled to the same processing unit, wherein the computing resource occupancy rate in the second resource state is greater than a second preset occupancy rate. It should be understood that the execution process of the steps in the memory bandwidth balancing method of this embodiment similar to the embodiment of FIG. 2 will not be repeated here. In various embodiments of the present invention, various embodiments and variations can be obtained by combining various steps. FIG. 4 is a flowchart of a specific example of the memory bandwidth balancing method of the embodiment of FIG. 2 . The memory bandwidth balancing method of FIG. 4 includes: Step S410: Starting memory bandwidth allocation and proceeding to step S415. Step S415: Obtaining historical memory bandwidth data for each of the multiple virtual machines and proceeding to step S420. For example, user profiles of virtual machine tenants can be generated based on the historical memory bandwidth data of the virtual machines. Dimensions of the user profiles include, but are not limited to, the ratio of each virtual machine to the total memory bandwidth requirement of all virtual machines, a representation of the urgency of the memory bandwidth requirement, a representation of the sensitivity of the memory bandwidth requirement, the service priority of the virtual machine, and the tenant priority. Step S420: Calculate scheduling weights for the multiple virtual machines based on their respective historical memory bandwidth data, and proceed to step S425. For example, each dimension of each virtual machine's historical memory bandwidth data can be calculated based on preset weight indicators for each dimension to obtain scheduling weight information for each of the multiple virtual machines. Step S425: Determine whether the host machine's computing resources are scarce. If so, proceed to step S445; if not, proceed to step S430. For example, a first preset occupancy rate and a second preset occupancy rate of the virtual machine's resource state can be used as criteria for determining whether computing resources are scarce. Step S430: Schedule virtual machines with the same scheduling weight to the same processing unit, and proceed to step S435. For example, when multiple virtual machines are in a first resource state of the host machine, different virtual machines with the same memory bandwidth requirements are dynamically scheduled to different processing units, where the computing occupancy rate of the first resource state is less than the first preset occupancy rate. Step S435: Determine whether the first change state of the first virtual machine is consistent with the second change state of the second virtual machine. If so, proceed to step S465; if not, proceed to step S440. Step S440: Set the memory bandwidth upper limit of the second virtual machine to less than half the memory bandwidth resources of the processing unit during a period adjacent to the peak moment of the first change state.For example, Resource Director Technology (RDT) can be used to adjust the memory bandwidth cap. Step S445: Determine whether each processing unit of the host machine contains multiple processing sub-units. If so, execute step S450. For example, when multiple virtual machines are in the second resource state of the host machine, schedule at least two virtual machines with consistent memory bandwidth requirements to the same processing unit, where the computing resource occupancy rate in the second resource state is greater than a second preset occupancy rate. If not, proceed to steps S455 and S60. Step S450: Dynamically schedule the same virtual machine to multiple processing sub-units to ensure balanced memory bandwidth usage by each processing sub-unit. Step S455: If the target virtual machine is not running, directly schedule the target virtual machine to a different processing unit from the already scheduled virtual machines. Step S460: If at least two already scheduled virtual machines are running, migrate some of the already scheduled virtual machines so that different already scheduled virtual machines are in different processing units. Step S465: Set the memory bandwidth caps for the first and second virtual machines to evenly divide the memory bandwidth resources of the same processing unit. For example, set the occupied memory bandwidth of each virtual machine to 50% of the processing unit, i.e., evenly distribute the total memory bandwidth of the memory bus. It should be understood that steps S440 and S465 are parallel processes, and at least one of them can be performed simultaneously. Figure 5 is a block diagram of a host machine according to other embodiments of the present invention. The host machine includes: a determination module 510, which determines the memory bandwidth requirements of multiple virtual machines for at least one processing unit, where different processing units access memory through different memory buses; and a scheduling module 520, which evenly schedules each virtual machine to a corresponding processing unit within the at least one processing unit, so that the memory bandwidth occupied by each virtual machine for the corresponding processing unit's memory bus matches the virtual machine's memory bandwidth requirement.In the solutions of the embodiments of the present invention, since different processing units access memory through different memory buses, the memory bandwidth demands of multiple virtual machines for at least one processing unit reflect the degree of contention among the multiple virtual machines for the memory bandwidth of the memory bus of the at least one processing unit to be scheduled. Consequently, when each virtual machine is evenly scheduled to a corresponding processing unit within the at least one processing unit, the memory bandwidth occupied by each virtual machine for the memory bus of the corresponding processing unit matches the memory bandwidth demand of the virtual machine. This ensures that virtual machines with greater memory bandwidth demands are allocated greater memory bandwidth, thereby improving the reliability of memory bandwidth allocation, enhancing virtual machine operating performance, and enhancing the user experience of virtual machine tenants. In other embodiments, the determination module is specifically configured to: obtain historical memory bandwidth occupied by multiple virtual machines in the invoked processing unit; perform weighted calculation on the historical memory bandwidth occupied by each of the multiple virtual machines to obtain scheduling weight information for each of the multiple virtual machines; and determine the memory bandwidth demands of the multiple virtual machines for the at least one processing unit based on the scheduling weight information for each of the multiple virtual machines. In other embodiments, the determination module is specifically configured to: determine a memory bandwidth share that matches the scheduling weight information of each virtual machine; and determine the memory bandwidth requirement of the virtual machine based on the memory bandwidth share of each virtual machine and the total memory bandwidth of the at least one processing unit. In other embodiments, the scheduling module is specifically configured to: dynamically schedule different virtual machines with consistent memory bandwidth requirements to different processing units in the at least one processing unit when the host machine deploying the multiple virtual machines is in a first resource state, wherein the computing resource occupancy in the first resource state is less than a first preset occupancy. In other embodiments, the scheduling module is further configured to: schedule at least two virtual machines with consistent memory bandwidth requirements to the same processing unit in the at least one processing unit when the host machine is in a second resource state, wherein the computing resource occupancy in the second resource state is greater than a second preset occupancy, and the second preset occupancy is greater than the first preset occupancy. In some other embodiments, the at least two virtual machines include a first virtual machine and a second virtual machine, and the scheduling module is further configured to: determine a change state of a bandwidth occupancy ratio of the first virtual machine and the second virtual machine on a memory bus of the same processing unit; and set an upper limit on the memory bandwidth occupied by at least one of the first virtual machine and the second virtual machine on the same processing unit based on the change state of the bandwidth occupancy ratio.In other embodiments, the scheduling module is specifically configured to: monitor a first change state of the memory bandwidth occupied by the first virtual machine and a second change state of the memory bandwidth occupied by the second virtual machine; if the first change state and the second change state are consistent, set the memory bandwidth upper limits of the first and second virtual machines to evenly divide the memory bandwidth resources of the same processing unit. In other embodiments, the scheduling module is further configured to: if the first change state and the second change state are inconsistent, set the memory bandwidth upper limit of the second virtual machine to less than half the memory bandwidth resources of the processing unit during a period adjacent to the peak moment of the first change state. The specific implementation of each module in the host machine can be found in the corresponding descriptions of the corresponding steps in the above-mentioned method embodiments, and corresponding beneficial effects are achieved, and are not further described here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the modules described above can be found in the corresponding process descriptions in the above-mentioned method embodiments, and are not further described here. Figure 6 is a flowchart of the steps of an application deployment method according to other embodiments of the present invention. The application deployment method in Figure 6 can be executed by a control node, such as a cloud management system (CMS), which is configured to perform the deployment of the application's microservices. Specifically, the application deployment method includes:
[0007] S610: Determine the memory bandwidth requirements of multiple microservices of the application. The memory bandwidth requirement of each microservice indicates the bandwidth requirement of the memory bus in the host machine where the microservice resides. It should be understood that the application includes core microservices such as the order module, recommendation module, and payment module that have high real-time memory bandwidth requirements, as well as non-core microservices such as the calendar module and message module that have lower real-time memory bandwidth requirements. The memory bandwidth requirements of core microservices are higher than those of non-core microservices.
[0008] S620: Based on the memory bandwidth requirements of the multiple microservices, deploy the multiple microservices to multiple host machines, such that the total memory bandwidth requirement of each microservice deployed on each host machine matches the memory bandwidth upper limit of the host machine. It should be understood that the computing resources or storage resources of the multiple host machines may vary. That is, when a host machine has more computing resources or storage resources, the memory bandwidth upper limit of the host machine is higher; when a host machine has fewer computing resources or storage resources, the memory bandwidth upper limit of the host machine is lower. It should also be understood that the memory bandwidth balancing method of the embodiment of the present invention can balance the memory bandwidth occupied by the host machines. Communication between different host machines requires the use of a network card or switch. The memory bandwidth requirement of each microservice deployed on each host machine is consistent with the memory bandwidth upper limit of the host machine, which can further optimize the local memory bandwidth balancing of the host machines. In this embodiment, since the memory bandwidth requirements of each microservice deployed on each host machine are consistent with the memory bandwidth upper limit of that host machine, memory bandwidth requirements are balanced across different host machines. Specifically, memory bandwidth management is efficiently performed at the granularity of management between hosts, from the perspective of the application. Figure 7 is a flowchart of the steps of an application deployment apparatus according to other embodiments of the present invention. The application deployment apparatus in Figure 7 corresponds to an application deployment method and includes: a determination module 710 for determining the memory bandwidth requirements of multiple microservices of an application, where the memory bandwidth requirement of each microservice indicates the bandwidth requirement of the memory bus of the host machine in which it resides. A deployment module 720, based on the memory bandwidth requirements of the multiple microservices, deploys the multiple microservices to multiple host machines, such that the total memory bandwidth requirement of the microservices deployed on each host machine matches the memory bandwidth upper limit of the host machine. The specific implementation of each module in the application deployment apparatus can be found in the corresponding descriptions of the corresponding steps in the above-mentioned method embodiments, and corresponding beneficial effects are achieved, so this description is not repeated here. Those skilled in the art will clearly understand that, for ease of description and brevity, the specific operating processes of the modules described above can be referred to the corresponding process descriptions in the aforementioned method embodiments and will not be repeated here. Referring to FIG8 , a schematic structural diagram of an electronic device according to another embodiment of the present invention is shown. The specific embodiments of the present invention do not limit the specific implementation of the electronic device. As shown in FIG8 , the electronic device may include: a processor 802 configured to execute a program 810; a communications interface 804; a memory 806; and a communications bus.
[0009] 808. The processor, communication interface, and memory communicate with each other via a communication bus. The communication interface is configured to communicate with other electronic devices or servers. The processor is configured to execute a program, specifically, to perform the relevant steps in the above-described method embodiments. Specifically, the program may include program code, which includes computer operating instructions. The processor may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs. The memory is configured to store programs. The memory may include high-speed RAM memory or non-volatile memory, such as at least one disk drive. The program may include multiple computer instructions. Specifically, the program may cause a processor to execute the following steps through the multiple computer instructions: determining the memory bandwidth requirements of multiple virtual machines for at least one processing unit, with different processing units accessing memory via different memory buses; balancing the scheduling of each virtual machine to a corresponding processing unit within the at least one processing unit, such that the memory bandwidth occupied by the virtual machine for the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine; or determining the memory bandwidth requirements of multiple microservices of an application, wherein the memory bandwidth requirement of each microservice indicates the bandwidth requirement of the memory bus of the host machine in which the microservice resides; and deploying the multiple microservices to multiple host machines based on the memory bandwidth requirements of the multiple microservices, such that the total memory bandwidth requirement of the microservices deployed in each host machine matches the memory bandwidth upper limit of the host machine. The specific implementation of each step in the program can be found in the corresponding descriptions of the corresponding steps and units in the above-mentioned method embodiments, and corresponding beneficial effects are achieved, and are not further described here. Those skilled in the art will clearly understand that, for ease of description and brevity, the specific operating procedures of the devices and modules described above can refer to the corresponding process descriptions in the above-mentioned method embodiments, and are not further described here. An embodiment of the present invention further provides a computer storage medium storing a computer program, which, when executed by a processor, implements the method described in any one of the aforementioned method embodiments.The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk. Embodiments of the present invention also provide a computer program product comprising computer instructions that instruct a computing device to perform operations corresponding to any of the memory bandwidth balancing methods described in the aforementioned method embodiments. Furthermore, it should be noted that the user-related information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, sample data used for model training, data used for analysis, stored data, and displayed data, etc.) involved in the embodiments of the present invention are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with relevant regulations and standards, and corresponding operation portals are provided for the user to choose to authorize or deny. It should be noted that, depending on implementation needs, the various components / steps described in the embodiments of the present invention may be split into more components / steps, or two or more components / steps or partial operations of a component / step may be combined into a new component / step to achieve the objectives of the embodiments of the present invention. The methods according to the embodiments of the present invention described above can be implemented in hardware or firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored in a remote recording medium or non-transitory machine-readable medium downloaded via a network and then stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)). It will be understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods described herein are implemented.Furthermore, when a general-purpose computer accesses the code for implementing the methods described herein, the execution of the code transforms the general-purpose computer into a specialized computer for executing the methods described herein. Those skilled in the art will appreciate that the units and method steps described in the various examples in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals skilled in the art may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present invention. The above embodiments are intended only to illustrate the embodiments of the present invention and are not intended to limit them. Persons skilled in the relevant technical fields may make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions are also within the scope of the embodiments of the present invention, and the scope of patent protection for the embodiments of the present invention shall be defined by the claims.
Claims
Claims 1. A memory bandwidth balancing method, wherein: include: Determining memory bandwidth requirements of multiple virtual machines for at least one processing unit, where different processing units access memory through different memory buses; Each virtual machine is evenly scheduled to a corresponding processing unit among the at least one processing unit, so that the memory bandwidth occupied by the virtual machine on the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine.
2. The method according to claim 1, wherein determining memory bandwidth requirements of multiple virtual machines for at least one processing unit comprises: Obtaining historical memory bandwidth occupied by multiple virtual machines in the called processing unit; Performing weighted calculation on the historical memory bandwidths occupied by each of the multiple virtual machines to obtain scheduling weight information of each of the multiple virtual machines; Memory bandwidth requirements of the multiple virtual machines for at least one processing unit are determined based on the respective scheduling weight information of the multiple virtual machines.
3. The method according to claim 2, determining the memory bandwidth requirements of the multiple virtual machines for at least one processing unit based on the scheduling weight information of each of the multiple virtual machines, comprising: Determine the memory bandwidth ratio that matches the scheduling weight information of each virtual machine; Determine a memory bandwidth requirement of the virtual machine based on the memory bandwidth ratio of each virtual machine and the total memory bandwidth of the at least one processing unit.
4. The method according to claim 1, wherein the balancing scheduling of each virtual machine to a corresponding processing unit in the at least one processing unit comprises: When the host machine deploying the multiple virtual machines is in a first resource state, different virtual machines with consistent memory bandwidth requirements are dynamically scheduled to different processing units in the at least one processing unit, wherein the computing occupancy of the first resource state is less than a first preset occupancy.
5. The method according to claim 4, wherein each virtual machine is evenly scheduled to a corresponding processing unit in the at least one processing unit, further comprising: When the host machine is in a second resource state, at least two virtual machines with consistent memory bandwidth requirements are scheduled to the same processing unit of the at least one processing unit, wherein a computing resource occupancy rate in the second resource state is greater than a second preset occupancy rate, and the second preset occupancy rate is greater than the first preset occupancy rate.
6. The method according to claim 5, wherein the at least two virtual machines include a first virtual machine and a second virtual machine, and the method further comprises: Determining a change in bandwidth occupancy ratio of the first virtual machine and the second virtual machine on a memory bus of a same processing unit; Based on the change state of the bandwidth occupancy ratio, an upper limit of memory bandwidth occupied by at least one of the first virtual machine and the second virtual machine on the same processing unit is set.
7. The method according to claim 6, wherein determining a change in bandwidth usage ratio of the first virtual machine and the second virtual machine on a memory bus of a same processing unit comprises: Monitoring a first change state of memory bandwidth occupied by the first virtual machine and a second change state of memory bandwidth occupied by the second virtual machine; Setting a memory bandwidth upper limit for at least one of the first virtual machine and the second virtual machine in the same processing unit based on a change state of the bandwidth occupancy ratio, including: if the first change state and the second change state are consistent, setting the memory bandwidth upper limits of the first virtual machine and the second virtual machine to evenly divide the memory bandwidth resources of the same processing unit.
8. The method according to claim 7, further comprising setting a memory bandwidth upper limit for at least one of the first virtual machine and the second virtual machine on the same processing unit based on the change state of the bandwidth occupancy ratio, and further comprising: if the first change state and the second change state are inconsistent, setting the memory bandwidth upper limit of the second virtual machine to be less than half of the memory bandwidth resources of the processing unit in an adjacent time period of a peak moment of the first change state.
9. An application deployment method, wherein: include: Determine the memory bandwidth requirements of multiple microservices of the application, where the memory bandwidth requirement of each microservice indicates the bandwidth requirement of the microservice for the memory bus in the host machine where it is located; Based on the memory bandwidth requirements of the multiple microservices, the multiple microservices are deployed to multiple host machines, so that the total memory bandwidth requirement of each microservice deployed in each host machine matches the memory bandwidth upper limit of the host machine.
10. A host machine, wherein: include: a determination module, determining memory bandwidth requirements of multiple virtual machines for at least one processing unit, where different processing units access memory through different memory buses; The scheduling module evenly schedules each virtual machine to a corresponding processing unit in the at least one processing unit so that the memory bandwidth occupied by the virtual machine for the memory bus of the corresponding processing unit matches the memory bandwidth requirement of the virtual machine.
11. An application deployment device, wherein: include: A determination module determines memory bandwidth requirements of multiple microservices of an application, where the memory bandwidth requirement of each microservice indicates the bandwidth requirement of the microservice for a memory bus in a host machine. The deployment module deploys the multiple microservices to multiple host machines based on the memory bandwidth requirements of the multiple microservices, so that the total memory bandwidth requirement of each microservice deployed in each host machine matches the memory bandwidth upper limit of the host machine.
12. An electronic device, wherein: include: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is configured to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method according to any one of claims 1 to 9.
13. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method and system for virtual environment
CN108694068A
Virtual machine-oriented hardware resource allocation method, apparatus and device, and storage medium
CN115390983A
Regulating memory bandwidth via CPU scheduling
US8826270B1