Resource allocation method, apparatus and computing device cluster
By exclusively using resources that cause performance issues for high-priority virtual instances among multiple resource types, the problem of reduced resource utilization is solved, achieving a balance between performance assurance and resource utilization for high-priority virtual instances.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-08-27
- Publication Date
- 2026-07-30
AI Technical Summary
While ensuring the performance of high-priority virtual machines, existing technologies can lead to reduced resource utilization.
Among various resource types, when a high-priority virtual instance experiences performance issues, the portion of the resource experiencing performance problems will be allocated to the high-priority virtual instance for exclusive use, reducing resource interference and thus ensuring the performance of the high-priority virtual instance while maintaining resource utilization.
By exclusively using resources, the performance of high-priority virtual instances is improved, while maintaining the overall utilization of resources.
Smart Images

Figure CN2025117318_30072026_PF_FP_ABST
Abstract
Description
Resource allocation methods, devices, and computing equipment clusters
[0001] This application claims priority to Chinese Patent Application No. 202510112726.2, filed on January 23, 2025, entitled “Resource Allocation Method, Apparatus and Computing Device Cluster”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates to the field of virtualization technology, specifically to a resource allocation method, apparatus, and computing device cluster. Background Technology
[0003] A physical machine typically runs multiple virtual machines (VMs) simultaneously. VMs have priorities, which indicate the importance of the services they provide. For high-priority VMs, the physical machine can isolate resources to prevent them from sharing resources with other VMs.
[0004] However, the above solution reduces resource utilization while ensuring the performance of high-priority VMs. Summary of the Invention
[0005] This invention provides a resource allocation method, apparatus, and computing device cluster. In scenarios where multiple types of resources are shared, when any type of resource experiences a performance problem for a high-priority virtual instance, at least a portion of the resource with the performance problem is allocated to the high-priority virtual instance for exclusive use. This reduces the interference of a certain type of resource on the high-priority virtual instance, thereby ensuring the performance of the high-priority virtual instance while maintaining resource utilization.
[0006] In a first aspect, embodiments of the present invention provide a resource allocation method applied to a first device, the first device including multiple virtual instances and multiple types of resources, the multiple virtual instances sharing the multiple types of resources, the method including:
[0007] First, a first virtual instance is determined; wherein, the first virtual instance is any one of multiple virtual instances, and the first virtual instance has a higher priority than the other virtual instances among the multiple virtual instances; after determining the high-priority first virtual instance, then, during the operation of the first virtual instance, the first performance information of the first type of resources is obtained; wherein, the first type of resources is any one of multiple types of resources, and the first performance information is used to indicate the performance of the first type of resources for the first virtual instance; after obtaining the performance of the first type of resources for the first virtual instance, if it is determined based on the first performance information that the first type of resources has a performance abnormality, it indicates that the first virtual instance has an abnormality when sharing the first type of resources with other virtual instances. In order to reduce the impact of sharing the first type of resources on the high-priority first virtual instance, the first resource in the first type of resources is allocated to the first virtual instance for exclusive use.
[0008] In this solution, when multiple types of resources are shared, if any type of resource experiences a performance issue for a high-priority virtual instance, at least a portion of the resource of the type experiencing the performance issue will be allocated to the high-priority virtual instance for exclusive use. This reduces the interference of a particular type of resource on the high-priority virtual instance, thereby ensuring the performance of the high-priority virtual instance while maintaining resource utilization.
[0009] In one possible implementation, the method further includes: firstly, analyzing whether the resources are sufficient based on the amount of resources exclusively used by the first virtual instance for the first type of resources and the total amount of resources of the first type of resources, in order to determine resource insufficiency information; wherein, the resource insufficiency information is used to indicate that the first type of resources are insufficient for the first virtual instance; after determining that the resources are insufficient, the resource insufficiency information can be sent to the scheduling device, thereby facilitating the calling device to determine the virtual instances that need to be migrated in the first device based on the resource insufficiency information, reducing the impact on the high-priority first virtual instance.
[0010] In this solution, when a certain type of resource experiences performance anomalies for a virtual instance, it analyzes whether there is insufficient resources for that type of resource for high-priority virtual instances. If resources are insufficient, it notifies the scheduling device, which then adjusts the virtual instances in the first device to ensure the performance of high-priority virtual instances.
[0011] In one optional example of this implementation, the method further includes: receiving a first migration instruction sent by a scheduling device; wherein the first migration instruction is used to instruct the migration of a second virtual instance to a second device, the second virtual instance being any virtual instance other than the first virtual instance among a plurality of virtual instances; then, in response to the first migration instruction, migrating the second virtual instance to the second device; and determining that the first virtual instance in the first device and other virtual instances other than the first virtual instance in the first device share a first resource.
[0012] In this solution, if a certain type of resource experiences performance anomalies for a virtual instance, and there is insufficient resources for high-priority virtual instances of that type, the number of virtual instances in the first device is adjusted to reduce the number of virtual instances sharing resources, thereby ensuring the performance of high-priority virtual instances.
[0013] In one optional example of this implementation, the method further includes: receiving a second migration instruction sent by a scheduling device; wherein the second migration instruction is used to instruct the migration of a first virtual instance to a third device, wherein the performance of a first type of resource in the third device is greater than the performance of a first type of resource in the first device; and then, in response to the second migration instruction, migrating the first virtual machine instance to the third device.
[0014] In this solution, if a certain type of resource experiences performance anomalies for a virtual instance, and there is insufficient resources for high-priority virtual instances, the high-priority virtual instances will be migrated to a more powerful device to ensure their performance.
[0015] In one possible implementation, the method further includes: during the operation of the first virtual instance, obtaining second performance information of a second type of resource, wherein the second type of resource is any type of resource other than the first type of resource among multiple types of resources, and the second performance information is used to indicate the performance of the second type of resource for the first virtual instance;
[0016] Correspondingly, determining that the first type of resource has a performance abnormality based on the first performance information includes: if, based on the first performance information and the second performance information, it is determined that the second type of resource affects the performance of the first type of resource, it indicates that other types of resources will interfere with the performance of the first type of resource for the first virtual instance, then it can be determined that the first type of resource has a performance abnormality.
[0017] In this solution, when analyzing a certain type of resource, we can refer to the performance of other types of resources for the high-priority first virtual instance to analyze whether other types of resources affect the performance of the first type of resource. If they do, it means that other types of resources will interfere with the performance of the first virtual instance using the first type of resource, and the performance of the first type of resource is abnormal.
[0018] In one possible implementation, the method further includes: sending first performance information to a scheduling device, which determines an analysis result based on the first performance information and fourth performance information. The analysis result is used to indicate whether the first type of resource has experienced performance anomalies, and the fourth performance information is used to indicate the performance of the first type of resource of the fourth device for the third virtual instance. The third virtual instance and the first virtual instance deploy the same application software to provide services that meet user needs. Then, the analysis result is sent to the first device. Subsequently, the first device can determine whether to allow the high-priority first virtual instance to exclusively use the first type of resource based on the analysis result.
[0019] In this solution, the scheduling device can obtain the performance of multiple virtual instances that provide the same service to meet user needs for a certain type of resource. If the performance of a certain virtual instance for a certain type of resource differs significantly from the performance of other virtual instances for the same type of resource, it indicates that the performance of the virtual instance for that type of resource is abnormal.
[0020] In one possible implementation, the first performance information is the performance index value of the first type of resource for the first virtual instance;
[0021] Determining that a first type of resource has a performance anomaly based on the first performance information includes: obtaining a set of indicator values corresponding to the performance indicators; wherein the set of indicator values is used to indicate multiple indicator values of the performance indicators for the first virtual instance under normal operation of the first virtual instance; obtaining a similarity threshold; wherein the similarity threshold is used to indicate the minimum value of the similarity with the set of indicator values; obtaining the similarity between the indicator values of the performance indicators for the first virtual instance and the set of indicator values; if the similarity is less than or equal to the similarity threshold, it indicates that the performance of the first type of resource differs significantly from that under normal operation of the first virtual instance, and the first type of resource is determined to have a performance anomaly.
[0022] In this solution, the similarity of the performance index values of the first type of resource to the high-priority first virtual instance is compared with that of the first virtual instance under normal operation. The analysis is conducted to determine whether there is a significant difference between the current performance of the first type of resource and the performance of the first virtual instance under normal operation. If the difference is significant, it indicates that the first type of resource is experiencing performance abnormalities.
[0023] In one possible implementation, the first performance information is the performance index value of the first type of resource for the first virtual instance;
[0024] Determining that a first type of resource has a performance anomaly based on the first performance information includes: obtaining the value range of the performance indicator corresponding to the performance indicator; wherein, the value range is used to indicate the range of values of the performance indicator for the first virtual instance under normal operation of the first virtual instance; and determining that a first type of resource has a performance anomaly when the value of the performance indicator for the first virtual instance is outside the value range.
[0025] In this solution, by comparing the performance index range with that of the first virtual instance under normal operation, the solution analyzes whether the current performance of the first type of resources for the high-priority first virtual instance is normal, thereby analyzing whether the first type of resources has experienced performance abnormalities.
[0026] In one possible implementation, the method further includes: during the operation of the first virtual instance, obtaining third performance information of the first resource, the third performance information being used to indicate the performance of the first resource for the first virtual instance; if it is determined based on the third performance information that the first resource has a performance abnormality, allocating the second resource in the first type of resources to the first virtual instance for exclusive use, so that the first virtual instance has exclusive use of the first resource and the second resource.
[0027] In this solution, when a certain type of resource experiences performance anomalies for a virtual instance, a resource allocation scheme based on a feedback mechanism is adopted. This scheme can continuously allocate that type of resource to high-priority virtual instances until the performance of that type of resource for the high-priority virtual instance returns to normal, thereby ensuring the performance of the high-priority virtual instance.
[0028] In one possible implementation, the first type of resource is memory bandwidth or a cache shared by the CPU cores of a central processing unit. Obtaining the exclusive resource amount of the first virtual instance for the first type of resource includes: obtaining a first number of virtual CPU cores of the first virtual instance; obtaining a second number of CPU cores in the first device; determining the ratio of the first number and the second number; determining the exclusive resource amount of the first virtual instance for the first type of resource based on the ratio and the total resource amount of the first type of resource; and determining a first resource in the first type of resource that is adapted to the exclusive resource amount.
[0029] In this solution, exclusive resources are allocated by the ratio of the number of virtual CPU cores to the number of physical CPU cores, ensuring that the allocated resources match the computing power of the required CPU cores and guaranteeing the performance of high-priority virtual instances.
[0030] In one possible implementation, the first type of resource also includes a third resource, which is shared by the first virtual instance and other virtual instances among the multiple virtual instances.
[0031] In this solution, high-priority virtual instances can not only exclusively use resources but also share resources with other virtual instances, thereby further ensuring the performance of high-priority virtual instances.
[0032] In one possible implementation, the multiple resources include at least two of the following: central processing unit (CPU) cores, memory bandwidth, and CPU core shared cache.
[0033] In one possible implementation, the first type of resource is a CPU core, and the first performance information is the performance metric of the CPU core; or, the first type of resource is memory bandwidth, and the first performance information is the performance metric of the memory bandwidth; or, the first type of resource is a cache shared by the CPU core, and the first performance information is the performance metric of the cache shared by the CPU core.
[0034] Secondly, embodiments of the present invention provide a resource allocation device, which includes several modules. Each module is used to execute various steps in the resource allocation method provided in the first aspect of the present invention. The division of modules is not limited here. For the specific functions performed by each module of this resource allocation device and the beneficial effects achieved, please refer to the functions of each step in the resource allocation method provided in the first aspect of the present invention; further details will not be repeated here.
[0035] Optionally, the resource allocation device includes multiple virtual instances and multiple types of resources, with the multiple virtual instances sharing the multiple types of resources; the resource allocation device includes:
[0036] The identification module is used to determine the first virtual instance; wherein the first virtual instance is any virtual instance among multiple virtual instances, and the first virtual instance has a higher priority than other virtual instances among multiple virtual instances;
[0037] The acquisition module is used to acquire first performance information of a first type of resource during the operation of the first virtual instance; wherein, the first type of resource is any one of multiple types of resources, and the first performance information is used to indicate the performance of the first type of resource for the first virtual instance;
[0038] The resource allocation module is used to allocate the first resource in the first type of resources to the first virtual instance for exclusive use when it is determined that the first type of resources have experienced performance abnormalities based on the first performance information.
[0039] In one possible implementation, the resource allocation module is used to determine resource insufficiency information based on the amount of resources exclusively used by the first virtual instance for the first type of resources and the total amount of resources of the first type of resources; wherein, the resource insufficiency information is used to indicate that the first type of resources are insufficient for the first virtual instance; the resource insufficiency information is sent to the scheduling device, and the device is invoked to determine the virtual instance that needs to be migrated in the first device based on the resource insufficiency information.
[0040] In one alternative example of this implementation, the apparatus further includes:
[0041] The first migration module is configured to receive a first migration instruction sent by the scheduling device; wherein the first migration instruction is configured to instruct the second virtual instance to be migrated to the second device, and the second virtual instance is any virtual instance other than the first virtual instance among a plurality of virtual instances; in response to the first migration instruction, the second virtual instance is migrated to the second device;
[0042] The resource allocation module is used to determine whether the first virtual instance in the first device and other virtual instances in the first device share the first resource.
[0043] In one alternative example of this implementation, the apparatus further includes:
[0044] The second migration module is used to receive a second migration instruction sent by the scheduling device; wherein the second migration instruction is used to instruct the migration of the first virtual instance to the third device, and the performance of the first type of resources in the third device is greater than the performance of the first type of resources in the first device; and to migrate the first virtual machine instance to the third device.
[0045] In one possible implementation, the acquisition module is used to acquire second performance information of a second type of resource during the operation of the first virtual instance. The second type of resource is any type of resource other than the first type of resource among multiple types of resources. The second performance information is used to indicate the performance of the second type of resource for the first virtual instance.
[0046] The resource allocation module is used to determine that the first type of resource has a performance abnormality when it is determined that the second type of resource affects the performance of the first type of resource based on the first performance information and the second performance information.
[0047] In one possible implementation, a resource allocation module sends first performance information to a scheduling device. The scheduling device, based on the first and fourth performance information, determines an analysis result. The analysis result indicates whether a first type of resource has experienced performance anomalies. The fourth performance information indicates the performance of the first type of resource on the fourth device relative to a third virtual instance. The third virtual instance and the first virtual instance deploy the same application software to provide services that meet user needs. The module then sends the analysis result back to the first device. In another possible implementation, the first performance information is the performance index value of the first type of resource relative to the first virtual instance.
[0048] The resource allocation module is used to obtain the set of indicator values corresponding to the performance indicators; wherein, the set of indicator values is used to indicate multiple indicator values of the performance indicators for the first virtual instance under normal operation of the first virtual instance; obtain a similarity threshold; wherein, the similarity threshold is used to indicate the minimum value of the similarity with the set of indicator values; obtain the similarity between the indicator values of the performance indicators for the first virtual instance and the set of indicator values; and determine that the first type of resource has a performance abnormality if the similarity is less than or equal to the similarity threshold.
[0049] In one possible implementation, the first performance information is the performance index value of the first type of resource for the first virtual instance;
[0050] The resource allocation module is used to obtain the value range of the performance indicators; wherein, the value range of the performance indicators is used to indicate the range of values of the performance indicators for the first virtual instance under normal operation of the first virtual instance; if the value of the performance indicator for the first virtual instance is outside the value range, it is determined that the first type of resource has a performance abnormality.
[0051] In one possible implementation, the resource allocation module is used to obtain third performance information of the first resource during the operation of the first virtual instance. The third performance information is used to indicate the performance of the first resource for the first virtual instance. If it is determined based on the third performance information that the first resource has a performance abnormality, the second resource in the first type of resources is allocated to the first virtual instance for exclusive use, so that the first virtual instance can exclusively use the first resource and the second resource.
[0052] In one possible implementation, the first type of resource is memory bandwidth or a cache shared by the CPU cores.
[0053] The resource allocation module is used to obtain the first number of virtual CPU cores of the first virtual instance; obtain the second number of CPU cores in the first device; determine the ratio of the first number and the second number; determine the amount of exclusive resources of the first virtual instance for the first type of resources based on the ratio and the total amount of resources of the first type of resources; and determine the first resource in the first type of resources that is adapted to the amount of exclusive resources.
[0054] In one possible implementation, the first type of resource also includes a third resource, which is shared by the first virtual instance and other virtual instances among the multiple virtual instances.
[0055] In one possible implementation, the multiple resource types include at least two of the following:
[0056] Central processing unit (CPU) cores, memory bandwidth, and CPU core shared cache.
[0057] In one possible implementation, the first type of resource is a CPU core, and the first performance information is the performance metric of the CPU core; or, the first type of resource is memory bandwidth, and the first performance information is the performance metric of the memory bandwidth; or, the first type of resource is a cache shared by the CPU core, and the first performance information is the performance metric of the cache shared by the CPU core.
[0058] Thirdly, embodiments of the present invention provide a resource allocation apparatus, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method provided in the first aspect.
[0059] Fourthly, embodiments of the present invention provide a resource allocation apparatus that executes computer program instructions to perform the method provided in the first aspect. Exemplarily, the apparatus may be a chip or a processor.
[0060] In one example, the device may include a processor that can be coupled to memory, read instructions from the memory, and execute the methods provided in the first aspect according to those instructions. The memory may be integrated into the chip or processor, or it may be independent of the chip or processor.
[0061] Fifthly, embodiments of the present invention provide a first device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method provided in the first aspect.
[0062] In a sixth aspect, embodiments of the present invention provide a computing device cluster, characterized in that it includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the method provided in the first aspect.
[0063] In a seventh aspect, embodiments of the present invention provide a computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method provided in the first aspect.
[0064] Eighthly, embodiments of the present invention provide a computer program product containing instructions that, when executed on a computer, cause the computer to perform the method provided in the first aspect. Attached Figure Description
[0065] Figure 1 is a schematic diagram of the structure of a computing node provided in an embodiment of the present invention;
[0066] Figure 2a is a schematic diagram of a CPU provided in an embodiment of the present invention;
[0067] Figure 2b is a schematic diagram of another CPU structure provided in an embodiment of the present invention;
[0068] Figure 2c is a schematic diagram of n virtual CPU cores sharing a common CPU core provided in an embodiment of the present invention;
[0069] Figure 2d is a schematic diagram of n virtual CPU cores that can be used exclusively and shared by an embodiment of the present invention;
[0070] Figure 2e is a schematic diagram of another n virtual CPU cores used exclusively and shared by an embodiment of the present invention;
[0071] Figure 3a is a schematic diagram of the architecture of the calling system provided in an embodiment of the present invention;
[0072] Figure 3b is a schematic diagram of the architecture of the calling system deployed in the cloud according to an embodiment of the present invention;
[0073] Figure 4a is a schematic diagram of the principle of the resource allocation method provided in the embodiment of the present invention;
[0074] Figure 4b is a schematic diagram of the first resource allocation scenario provided by an embodiment of the present invention;
[0075] Figure 4c is a schematic diagram of the second resource allocation scenario provided by an embodiment of the present invention;
[0076] Figure 4d is a flowchart illustrating a resource allocation method provided in an embodiment of the present invention;
[0077] Figure 5a is a schematic diagram of the scenario of the performance index values provided in the embodiment of the present invention;
[0078] Figure 5b is a schematic diagram of the first scenario for determining that the i-th type of resource has a performance abnormality for virtual instance A, provided by an embodiment of the present invention;
[0079] Figure 5c is a schematic diagram of the second scenario for determining that the i-th type of resource has a performance abnormality for virtual instance A, provided by an embodiment of the present invention;
[0080] Figure 5d is a schematic diagram of the third scenario for determining that the i-th type of resource has a performance abnormality for virtual instance A, provided by an embodiment of the present invention;
[0081] Figure 5e is a schematic diagram of the fourth scenario for determining whether the i-th type of resource has abnormal performance for virtual instance A, provided by an embodiment of the present invention.
[0082] Figure 5f is a flowchart illustrating the third resource allocation scenario provided in an embodiment of the present invention;
[0083] Figure 6a is a schematic diagram of the fourth resource allocation scenario provided by an embodiment of the present invention;
[0084] Figure 6b is a flowchart illustrating another resource allocation method in the scenario of Figure 6a;
[0085] Figure 6c is a schematic diagram of a CPU core allocation scenario provided in an embodiment of the present invention;
[0086] Figure 6d is a schematic diagram of an LLC allocation scenario provided in an embodiment of the present invention;
[0087] Figure 7a is a schematic diagram of the fifth resource allocation scenario provided by an embodiment of the present invention;
[0088] Figure 7b is a flowchart illustrating another resource allocation method in the scenario shown in Figure 6a;
[0089] Figure 8a is a flowchart illustrating a resource allocation method based on Figure 7b;
[0090] Figure 8b is a schematic diagram of the sixth resource allocation scenario provided by an embodiment of the present invention;
[0091] Figure 9a is a flowchart illustrating another resource allocation method based on Figure 7b;
[0092] Figure 9b is a schematic diagram of the seventh resource allocation scenario provided by an embodiment of the present invention;
[0093] Figure 10 is a schematic diagram of a resource allocation device provided in an embodiment of the present invention;
[0094] Figure 11 is a schematic diagram of the structure of the computing device provided in an embodiment of the present invention;
[0095] Figure 12 is a schematic diagram of the structure of the computing device cluster provided in an embodiment of the present invention;
[0096] Figure 13 is a schematic diagram of computing devices in a computer cluster connected via a network according to an embodiment of the present invention. Detailed Implementation
[0097] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be described below with reference to the accompanying drawings.
[0098] In the description of the embodiments of the present invention, the words "exemplary," "for example," or "for instance" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary," "for example," or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the words "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a specific manner.
[0099] In the description of the embodiments of this invention, the term "and / or" is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, B existing alone, and A and B existing simultaneously. Furthermore, unless otherwise stated, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.
[0100] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0101] The following explanations cover some of the terms used in this embodiment. It should be noted that these explanations are for the convenience of those skilled in the art and are not intended to limit the scope of protection claimed by this invention.
[0102] A virtual machine (VM) is a complete computer system simulated by software, possessing full hardware system functionality and running in a completely isolated environment. Any task that can be performed on a server can also be performed in a VM. When creating a VM on a server, a portion of the physical machine's hard drive and memory capacity is used as the VM's hard drive and memory capacity. Each VM has its own independent hard drive and operating system, and VM users can operate it just like they would on a server. A VM creates an environment between a computer platform and the end user, allowing the end user to operate other software within that virtual machine environment. From the application's perspective, a program running on a VM is identical to one running on its corresponding physical computer.
[0103] An ESC (Elastic Compute Service) instance is a virtual computing environment consisting of CPU, memory, system disk, and a running operating system. As the core concept of a cloud server, other resources, such as disks, images, and snapshots, only have practical use when combined with an ESC instance.
[0104] Virtual Machine Manager (VMM): Also known as a virtual machine monitor, it is a special type of software. A VMM can manage and externally monitor virtual machines. Additionally, a VMM is also called a hypervisor.
[0105] Central Processing Unit (CPU): As the core of a computer system for computation and control, it is the final execution unit for information processing and program execution.
[0106] Level 1 Cache: The fastest and smallest in size, it runs close to the CPU core and is typically tens to hundreds of KB. It stores the CPU's most frequently used instructions and data.
[0107] Secondary cache (L2 cache): Slightly slower, but larger in capacity, typically ranging from several hundred KB to several MB, storing hot data not covered by L1.
[0108] Level 3 Cache: Largest capacity, slowest speed, located outside the CPU core, typically several MB to tens of MB, supported by multiple cores, primarily expanding cache capacity and optimizing multi-threaded applications. The Level 3 cache is designed to cache data that is missed after accessing the Level 2 cache. In CPUs with a Level 3 cache, only about 5% of data needs to be retrieved from main memory, further improving CPU efficiency. The main function of the Level 3 cache is to store data that is missed in the Level 2 cache, providing larger storage space and faster access speeds. Especially in multi-core processors, the Level 3 cache can coordinate data access and reduce the frequency of access to main memory.
[0109] Last Level Cache (LLC): This refers to the highest level of cache in a computer architecture, typically L3 Cache. It resides between processor cores, stores large blocks of data, and is shared by multiple cores. The main function of LLC is to reduce access to main memory, thereby improving overall system performance.
[0110] Central Processing Unit (CPU): As the core of a computer system for computation and control, it is the final execution unit for information processing and program execution.
[0111] CPU core: It interprets computer instructions and processes software data. It is the core of a computer's computing and control, responsible for reading instructions, translating and encoding them, and executing them.
[0112] Cluster analysis is a method of dividing data into different groups or clusters. It groups data objects by similarity measurement, so that the data within the same group has high similarity and the data between different groups has low similarity.
[0113] A way is an organization method in CPU cache used to improve cache access efficiency and reduce access latency. In CPU cache, a way means that the storage units in the cache are divided into multiple independent storage areas, each of which can independently store data. When the CPU accesses the cache, it selects the appropriate way to access the data based on the data's address information. In this way, the number of accesses to main memory can be reduced, thus improving access speed.
[0114] Memory bandwidth refers to the amount of data a memory module can transfer per unit of time, usually measured in GB / s (gigabytes per second). It determines how quickly the memory can provide data to or receive data from the processor.
[0115] VM hybrid deployment: refers to a single physical machine where VMs from different businesses or users are deployed and running. The applications of the business or users run inside the VMs.
[0116] Isolation: The resources of a physical machine are divided so that each isolated VM occupies a portion of non-overlapping resources.
[0117] On-demand isolation: Based on the interference experienced by VMs on different resource dimensions on the machine, resources with a high degree of interference are isolated, while resources in other dimensions are shared.
[0118] Complete isolation: Isolate VMs on the physical machine in all available resource dimensions, including CPU cores, LLC, and memory bandwidth.
[0119] Core binding: Isolate the CPU cores of the VM, with each VM having its own exclusive access to a number of CPU cores.
[0120] Cloud: A software platform that uses application virtualization technology, integrating multiple functions such as software search, download, use, management, and backup.
[0121] Virtual instances: Virtual servers, virtual machines, containers, etc., created through virtualization technology. These instances can run independently on the cloud platform, with a complete operating system and application runtime environment, just like running on a real physical server.
[0122] Performance refers to the overall performance of a virtual instance, such as a VM, in terms of computing power, response speed, stability, and resource utilization during operation.
[0123] Performance metrics: These are metrics that describe performance usage, such as the computing power, response speed, and resource utilization of virtual instances (VMs) during operation.
[0124] Performance anomalies refer to problems that affect the performance of virtual instances, such as VMs, during operation, such as reduced computing power, reduced response speed, and reduced resource utilization.
[0125] Physical node: A fundamental concept in computer networks, referring to the physical device within the network. Physical nodes are responsible for processing and forwarding data, and provide the network's computing and storage capabilities.
[0126] Load: This refers to the number of tasks running on a physical node and the measure of resource utilization. The number of tasks refers to the number of tasks running on a physical node. Too many tasks may lead to performance degradation or even crashes. Resource utilization refers to the proportion of resources used by the tasks running on a physical node, including CPU, memory, and disk. If resource utilization is too high, it may cause tasks to execute slowly or fail.
[0127] Business: Services that meet user needs.
[0128] Services: refers to the various functions performed by a computer system.
[0129] Application software, as opposed to system software, refers to a collection of various programming languages and applications written in those languages that users can use. It is divided into application packages and user programs. Application packages are collections of programs designed to solve specific types of problems using a computer, and are primarily intended for user use.
[0130] Clustering: The process of dividing a collection of physical or abstract objects into multiple classes composed of similar objects is called clustering. A cluster generated by clustering is a set of data objects that are similar to objects in the same cluster and different from objects in other clusters.
[0131] First, the physical nodes to which the method provided in the embodiments of the present invention may be applied will be described. Figure 1 is a schematic diagram of the architecture of a physical node provided in an embodiment of the present invention.
[0132] In this embodiment of the invention, a physical node is shared with multiple tenants at the virtual machine level, enabling tenants to conveniently and flexibly use physical hardware resources under the premise of secure isolation, and greatly improving the utilization rate of physical hardware resources.
[0133] As shown in Figure 1, physical node 100 includes physical hardware resources 110, a virtual machine manager 120, and n (a positive integer greater than or equal to 2) virtual instances 130. It should be noted that n represents the total number of virtual instances 130. For example, physical node 100 may include three virtual instances 130, denoted as virtual instance 1, virtual instance 2, and virtual instance 3. Virtual instances 130 can be virtual machines, or in other scenarios, containers, bare metal servers, etc. In a cloud scenario, virtual instances 130 can be ECS instances purchased by users. Virtual instances 130 are used to provide services (various services required by users) to tenants. Different virtual instances 130 provide different services; for example, as shown in Figure 1, virtual instances 1, 2, ..., n run application software 1, application software 2, ..., application software n respectively, and application software 1, application software 2, ..., application software n provide service 1, service 2, ..., service n respectively.
[0134] Physical hardware resources 110 may include at least computing resources, storage resources, and network resources. For example, computing resources may include a processor and the memory shown in Figure 1.
[0135] The processor may include a CPU, and may also include any one or more of the following: a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). In this embodiment, as shown in Figures 2a and 2b, the CPU may include CPU cores, L1 cache, L2 cache, and LLC. The number of CPU cores may be N (a positive integer greater than or equal to 2). The L1 cache may be integrated into the CPU cores. The L2 cache may be integrated into the CPU cores as shown in Figure 2a, or it may be located outside the N CPU cores and shared by multiple CPU cores as shown in Figure 2b. The LLC may be located outside the N CPU cores and shared by the N CPU cores.
[0136] In the n virtual instances 130, the resource status of each virtual instance 130 for the CPU core can be shared and / or exclusive.
[0137] Optionally, in some implementations, the resource status of each virtual instance 130 among the n virtual instances 130 for the CPU core can be shared; as shown in Figure 2c, each CPU core among the N CPU cores can run any one of the n virtual instances 130.
[0138] In some optional implementations, some of the n virtual instances 130 may have exclusive access to CPU core resources, while others may have shared access to CPU core resources. As shown in Figure 2d, virtual instance 1 exclusively uses CPU cores 1 and 2, while virtual instances 2, ..., n share CPU cores 3, ..., N. In other words, each CPU core in CPU cores 1 and 2 runs virtual instance 1, and each CPU core in CPU cores 3, ..., N runs virtual instance 2, ..., n.
[0139] In some optional implementations, among the n virtual instances 130, some virtual instances 130 can have exclusive and shared CPU core resource states, while other virtual instances 130 can have shared CPU core resource states; as shown in Figure 2e, virtual instance 1 exclusively uses CPU core 1 and CPU core 2, while virtual instances 1, virtual instance 2, ..., virtual instance n share CPU core 3, ..., CPU core N. In other words, each CPU core in CPU core 1 and CPU core 2 runs virtual instance 1, and each CPU core in CPU core 3, ..., CPU core N runs virtual instance 1, virtual instance 2, ..., virtual instance n.
[0140] Additionally, the resource status of each virtual instance 130 in the n virtual instances 130 for the LLC can be shared and / or exclusive. For details, please refer to the description of the resource status of each virtual instance 130 in the n virtual instances 130 for the CPU core, and the description of Figures 2c to 2e above, the difference being that the CPU core is replaced by LLC.
[0141] In this context, the memory is volatile memory, such as random access memory (RAM). A memory bus connects the CPU and the memory, and this memory bus has memory bandwidth. The resource state of the memory bandwidth for each of the n virtual instances 130 can be shared and / or exclusive. For details, refer to the description of the resource state of each virtual instance 130 for the CPU core, and the descriptions of Figures 2c to 2e above, the difference being that the CPU core is replaced by memory bandwidth.
[0142] The storage resources may include non-volatile memory, such as a disk. For example, the disk may be a hard disk drive (HDD) or a solid state drive (SSD). The disk is an example of non-volatile memory and does not constitute a specific limitation.
[0143] The network resources may include the network interface card (NIC) shown in Figure 1. The NIC is merely an example and does not constitute a specific limitation.
[0144] The virtual machine manager 120 is responsible for managing the physical hardware resources 110, which are then allocated to virtual machines 130 as needed. Specifically, the virtualization management system 120 determines how much physical hardware resource 110 is actually allocated to each virtual machine 130.
[0145] It should be noted that multiple virtual instances 130 are created on a single physical node 100, and these virtual instances 130 share the physical hardware resources 110 of the entire physical node 100. For any given virtual instance 130, apart from the virtual resources it requests, it cannot see the actual physical hardware resources 110 of the physical node 100, nor can it see other virtual instances 130 on the same physical node 100. The virtual resources of a virtual instance 130 may include virtual CPU (vCPU), virtual memory, virtual hard disk, and virtual network interface card. From the perspective of the virtualization management system 120, virtual instance 130 appears as a single task, but from the client's perspective, it appears as a "host." This "host" has its own physical resources and its own operating system, and users can run any application on this "host."
[0146] The virtual machine manager 120 is used to implement the functions of compute virtualization, network virtualization, and storage virtualization (the latter two can be summarized as IO virtualization) as well as cloud management platform client.
[0147] Computational virtualization provides computing resources such as CPU and memory of physical node 100 to virtual instance 130.
[0148] In CPU virtualization, physical node 100 can be configured with multiple CPU cores, and virtual machine manager 120 needs to associate the vCPUs of virtual instance 130 with several CPU cores. It should be noted that virtual instance 130 cannot perceive the physical CPU cores; the virtual resources of virtual instance 130 are presented through vCPUs, and the virtual machine can only perceive the vCPUs presented to the virtual machine by virtual machine manager 120.
[0149] The purpose of memory virtualization is to provide virtual instance 130 with a contiguous physical memory space starting from address 0, effectively isolating and scheduling memory resources among virtual instances 130. In this embodiment, the memory bandwidth is shared by n virtual instances 130.
[0150] When virtual instance 130 is running, when a user or program issues a resource request to obtain physical hardware resources 110, including the required number of vCPUs, memory capacity, network bandwidth, etc., the virtualization management system 120 establishes an association between the virtual resources of virtual machine 130 and the physical hardware resources 110 based on the resource request. For example, each vCPU of virtual instance 130 corresponds to a CPU core, and virtual memory corresponds to the actual memory address, thereby allocating physical resources to virtual machine 130. Subsequently, running virtual machine 130 can access physical resources. For any virtual instance 130, the number of CPU cores and the number of vCPUs can be the same or different. For different virtual instances 130, the CPU cores associated with vCPUs can be the same or different.
[0151] I / O virtualization is a device virtualization technology that aims to abstract physical I / O devices, such as hard drives and network cards, into multiple virtual devices, allowing multiple virtual instances to share the same physical I / O device while ensuring the isolation and security of each virtual instance and improving overall system performance.
[0152] The following section describes the scheduling systems that may be used for physical nodes.
[0153] Figure 3a is a schematic diagram of the architecture of a scheduling system provided in an embodiment of the present invention.
[0154] As shown in Figure 3a, the scheduling system includes multiple physical nodes 100 and a scheduling device 300. The physical nodes 100 and the scheduling device 300 can communicate via a network. This network can be a wired network or a wireless network. It is understood that the network can use any known network communication protocol to achieve communication between different client layers and gateways. This network communication protocol can be various wired or wireless communication protocols, such as Ethernet, Universal Serial Bus (USB), or any combination thereof.
[0155] In this context, the multiple physical nodes 100 can be physical node 1, physical node 2, physical node 3, ... as shown in Figure 3a. Each physical node 100 deploys multiple virtual instances 130 and m (a positive integer greater than or equal to 2) types of resources. The m types of resources are resources that affect the computing performance of the n virtual instances. They can include several types of CPU-related resources, such as CPU cores and LLC, and several types of memory-related resources, such as memory bandwidth. For example, the m types of resources can include CPU cores, LLC, memory bandwidth, ..., the m-th type of resource. As shown in Figure 3a, n (greater than or equal to 2) virtual instances 130 can be deployed on physical node 1. It should be noted that the number of virtual instances 130 deployed on different physical nodes 100 can be different or the same, depending on the actual needs.
[0156] It should be noted that CPU cores, LLC, and memory bandwidth are merely examples and do not constitute specific limitations. In practical applications, m-type resources can be flexibly designed based on the types of resources that the CPU and memory can provide. In addition to CPU and memory, if high network quality requirements are needed, network bandwidth and other network-affecting resources can also be considered. It is worth noting that m-type resources are resources that can be provided by physical hardware resources; that is, physical hardware resources include m-type resources.
[0157] In this embodiment, the scheduling device 300 is used to record the correspondence between virtual instances 130 and physical nodes 100, such as which virtual instances 130 are running on physical node 100; and the scheduling device 100 is also used to migrate virtual instances 130, for example, migrating virtual instance 1 from physical node 1 to physical node 2, thereby removing virtual instance 1 from physical node 1 and adding virtual instance 1 to physical node 2. Furthermore, the structure of the scheduling device 300 can be referenced from that of physical node 100; for example, the scheduling device 300 includes physical hardware resources 110 in physical node 100.
[0158] Additionally, as shown in Figure 3b, the deployment environment of the scheduling system can be the cloud, which includes a cloud management platform and a data center. The data center can include a physical node cluster consisting of multiple physical servers. The physical node cluster provides infrastructure, which can include databases, services, physical servers, and virtual servers such as virtual machines and containers. Tenants can access the cloud management platform, purchase virtual instance 130, and configure the application software that needs to run on virtual instance 130 and the priority of the application software. The cloud management platform can create virtual instance 130 on physical node 100 through scheduling device 300 based on the virtual instance 130 purchased by the user, and run the application software configured by the tenant on the created virtual instance 130. The created virtual instance 130 is tagged, and the tag is used to indicate the priority of the application software configured by the tenant. Tenants can access the cloud management platform through a terminal, which can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. Exemplary embodiments of the terminals involved in this solution include, but are not limited to, electronic devices running iOS, Android, Windows, Harmony OS, or other operating systems. This embodiment of the invention does not specifically limit the type of terminal 110.
[0159] It should be noted that the cloud is only one possible scenario and does not constitute a specific limitation. The deployment environment of the calling system can be determined based on actual needs.
[0160] Next, in conjunction with the physical node 100 and scheduling system provided above, a resource allocation method provided by an embodiment of the present invention will be described in detail.
[0161] Virtual instance 130 has a priority, which indicates the importance of the services provided by virtual instance 130. In related technologies, physical node 100 can isolate high-priority virtual instance 130 from other virtual instances 130 by not sharing resources. However, this solution reduces resource utilization while ensuring the performance of high-priority virtual instance 130.
[0162] Based on this, embodiments of the present invention provide the following resource allocation method:
[0163] Multiple virtual instances 130 in physical node 100 share multiple types of resources. For a high-priority virtual instance 130 among the multiple virtual instances 130, the performance of each type of resource for that virtual instance 130 is analyzed. When any type of resource has a performance problem, at least a portion of the resource with the performance problem is allocated to that virtual instance 130 for exclusive use, reducing the interference of a certain type of resource on the high-priority virtual instance, thereby ensuring the performance of the high-priority virtual instance while ensuring resource utilization.
[0164] This embodiment can be applied to any physical node 100 in the calling system, such as physical node 1. Physical node 1 includes n virtual instances 130 and m types of resources. The n virtual instances 130 share the m types of resources, which are resources that affect the computing performance of the n virtual instances. The sharing of m types of resources by the n virtual instances can be understood as each of the n virtual instances being able to use any one of the m types of resources. It is worth noting that for other devices besides physical node 1, the difference lies in the number of virtual instances 130, which may be different.
[0165] Figure 4a illustrates a scenario diagram of the resource allocation method provided in this embodiment of the invention. As shown in Figure 4a, in the resource allocation method provided in this embodiment of the invention, physical node 1 identifies high-priority virtual instances 130, determining high-priority virtual instances 1, 2, ... For high-priority virtual instance 1, its performance is determined, and then anomaly detection is performed on m types of resources. These m types of resources include, but are not limited to, CPU cores, LLC, and memory bandwidth. For each type of resource that is abnormal, it is allocated to virtual instance 1 for exclusive use, achieving on-demand isolation. The same applies to other high-priority virtual instances, which will not be elaborated further. This reduces interference to virtual instances in shared environments and increases the probability of high-priority virtual instances operating normally. It is worth noting that since user metrics (e.g., response speed) for virtual instance 130 are difficult to obtain, and considering the real-time nature of virtual instance 130 resource usage, a resource dimension is used to analyze whether the performance of virtual instance 130 is abnormal. For example, if a certain type of resource is abnormal, it indicates that virtual instance 130 is abnormal.
[0166] It should be noted that the CPU cores, LLC, and memory bandwidth mentioned above are merely examples and do not constitute specific limitations. In practical applications, more resource types that affect the computing performance of n virtual instances can be flexibly designed according to actual needs. For example, when the CPU cores share the L2 cache, the m-type resources can also include the L2 cache shared by the CPU cores; furthermore, if high network quality is required, the m-type resources can also include several types of resources that affect the network, such as network bandwidth.
[0167] In an exemplary scenario, as shown in Figure 4b, physical node 1 includes n virtual instances 130: virtual instance 1, virtual instance 2, ..., virtual instance n, and m types of resources: CPU core, LLC, memory bandwidth, ..., the mth type of resource. Virtual instances 1, virtual instance 2, ..., virtual instance n share the m types of resources. Virtual instance 1 has high priority. When virtual instance 1 is running, the performance of CPU core, LLC, and memory bandwidth is collected to obtain performance information 1, performance information 2, ..., performance information m. Then, performance analysis is performed based on performance information 1. If the CPU core is found to be abnormal, virtual instance 1 is isolated on demand for the CPU core, and the resources in the CPU core are allocated to virtual instance 1 for exclusive use. Virtual instances 2, ..., virtual instance n share other resources in the CPU core. In this scenario, the resource status of virtual instance 1 for the CPU core changes from shared to exclusive.
[0168] In one exemplary scenario, as shown in Figure 4c, the difference between this scenario and the one in Figure 4b is that virtual instances 2, ..., n share other resources in the CPU core. In this scenario, the resource status of virtual instance 1 on the CPU core changes from shared to shared and exclusive. Subsequently, less important processing of virtual instance 1 can be assigned to the shared CPU core, and more important processing can be assigned to the exclusive CPU core.
[0169] This is merely an illustrative description; for details, please refer to Figure 4d and its description.
[0170] In this embodiment, the processing of each high-priority virtual instance 130 is the same. For ease of description, the first virtual instance is used as an example. The first virtual instance is any high-priority virtual instance 130 in physical node 1. The processing of each type of resource in m types of resources is the same. For ease of description and understanding, the technical solution is explained using the i-th type of resource as an example. The i-th type of resource is any type of resource in m types of resources.
[0171] Figure 4d is a flowchart illustrating the resource allocation method provided in an embodiment of the present invention. As shown in Figure 4d, the resource allocation method provided in an embodiment of the present invention includes at least the following steps:
[0172] Step 401: Physical node 1 determines virtual instance A, where virtual instance A is any virtual instance among n virtual instances 130, and virtual instance A has a higher priority than the other virtual instances 130 among the n virtual instances 130.
[0173] In one optional implementation of this embodiment, physical node 1 can obtain the tag of each virtual instance 130 among n virtual instances 130. The tag is used to indicate the priority of the virtual instance 130. Based on the tag of each virtual instance 130 among n virtual instances 130, the highest priority virtual instances 130 are determined as virtual instances A respectively.
[0174] In one optional implementation of this embodiment, physical node 1 can obtain user input tags for each application software running in n virtual instances 130. The user input tags are used to indicate the priority of the application software entered by the user. Based on the user input tags for each application software running in n virtual instances 130, a number of application software with the highest priority are determined, and each virtual instance 130 running the application software with the highest priority is respectively designated as virtual instance A.
[0175] Step 402: When physical node 1 is running virtual instance A, it obtains the performance information i1 of the i-th type of resource. The performance information i1 is used to indicate the performance of the i-th type of resource for virtual instance A.
[0176] In some optional implementations of this embodiment, when physical node 1 is running virtual instance A, it can determine the resources in the i-th type of resources that are running virtual instance A, collect the running status of these resources, perform performance analysis based on the running status, and obtain performance information i1. The i-th type of resource can be any one of the following: CPU cores, CPU LLC, or memory bandwidth. CPU cores, LLC, and memory bandwidth are merely examples and do not constitute specific limitations. In practical applications, more resource types affecting the computing performance of n virtual instances can be flexibly designed according to actual needs. For example, when CPU cores share a L2 cache, the i-th type of resource can be the CPU cores' shared L2 cache; or, if high network quality is required, the i-th type of resource can be a resource affecting the network, such as network bandwidth.
[0177] Performance information i1 is used to indicate the performance metrics of the i-th type of resource.
[0178] Optionally, in one example, the i-th type of resource is a CPU core, and the performance metrics may include, but are not limited to, one or more of the following metrics:
[0179] Instructions: The number of instructions executed;
[0180] mem_load_retired.l1_hit: Represents the number of micro-operations loaded from memory when the L1 cache is hit;
[0181] mem_load_retired.l1_miss: indicates the number of instructions that load data from memory and retire it when the L1 cache is missed;
[0182] mem_load_retired.l2_hit: Indicates the number of micro-operations loaded from memory when the L2 cache is hit;
[0183] mem_load_retired.l2_miss: indicates the number of instructions that load data from memory and retire it when the L2 cache is missed;
[0184] branch-misses: refers to the number of times a branch prediction failed;
[0185] branch-instructions: branch instructions.
[0186] Optionally, in one example, the i-th type of resource is an LLC, and the performance metrics may include, but are not limited to, any of the following:
[0187] mem_load_retired.l3_hit: Indicates the number of micro-operations loaded from memory when the L3 cache is hit;
[0188] mem_load_retired.l3_miss: This indicates the number of instructions that load data from memory and retire when the L1 cache is missed.
[0189] Optionally, in one example, the i-th type of resource is memory bandwidth, and the performance metrics may include, but are not limited to, any of the following:
[0190] mem_inst_retired.all_stores: refers to the number of all memory instructions retired;
[0191] mem_inst_retired.all_loads: refers to the number of retired instructions in the processor;
[0192] cache-misses: refers to the number of cache misses that occur in the cache hierarchy.
[0193] It should be noted that the above performance indicators are merely examples and do not constitute specific limitations. In practical applications, performance indicators can be flexibly designed according to actual needs.
[0194] Step 403: Based on performance information i1, physical node 1 determines whether the i-th type of resource has a performance anomaly. If so, proceed to step 404.
[0195] In one optional implementation of this embodiment, performance information i1 is the target value of the performance index of the i-th type of resource for virtual instance A (for ease of description and distinction, it can be called the target index value), as shown in Figure 5a. The performance index can be multiple indices, and the index value is the index value of each indice. As shown in Figure 5b, physical node 1 can obtain the index value range corresponding to the performance index of the i-th type of resource. The index value range indicates the range of values for the performance index of the i-th type of resource for virtual instance A under normal operating conditions. The index value range is a set consisting of all real numbers between the maximum and minimum values. Optionally, physical node 1 can determine the target value based on the historical performance information of the i-th type of resource for virtual instance A. The target value range and historical performance information can be the performance indicators of the i-th type of resource collected by physical node 1 under low load for virtual instance A. The collection time of each indicator value is different. Then, the multiple indicator values are clustered to obtain the indicator value set. The indicator value range is determined based on the maximum and minimum values of the performance indicators in the indicator value set. Here, the clustering can use one category or two categories, and one of the clusters can be manually specified as the indicator value set. Subsequently, if the target indicator value of the i-th type of resource is outside the indicator value range, it means that the performance of the i-th type of resource is out of the normal range, and it can be determined that the i-th type of resource has a performance abnormality for virtual instance A.
[0196] Optionally, as shown in Figure 5b, the performance metrics can be multiple metrics, and the metric value range is the range corresponding to each of the multiple metrics. The target metric value of the performance metric for the i-th type of resource is the metric value of each of the performance metrics. Then, if the metric value of each of the performance metrics exceeds the corresponding range, it indicates that the i-th type of resource has a performance anomaly for virtual instance A. For example, the i-th type of resource is LLC, and the performance metrics may include mem_load_retired.l3_hit and em_load_retired.l3_miss. The metric value range is range 1 corresponding to mem_load_retired.l3_hit and range 2 corresponding to em_load_retired.l3_miss. The performance information i1 is the metric value a of mem_load_retired.l3_hit and the metric value b of em_load_retired.l3_miss. If a is outside range 1 and b is outside range 2, it is determined that the LLC has a performance anomaly for virtual instance A.
[0197] In one optional implementation of this embodiment, performance information i1 is the target indicator value of the performance index of the i-th type of resource for virtual instance A. As shown in Figure 5c, physical node 1 can obtain the set of indicator values corresponding to the performance index of the i-th type of resource. The set of indicator values indicates multiple indicator values of the performance index of the i-th type of resource for virtual instance A under normal operating conditions. For example, physical node 1 can determine the set of indicator values based on historical performance information of the i-th type of resource for virtual instance A. The historical performance information can be multiple indicator values of the performance index of the i-th type of resource for virtual instance A collected by physical node 1 under low load conditions. Since the data collection time varies, multiple indicator values can be clustered to obtain an indicator value set. Clustering can use one category or two categories, and one cluster can be manually designated as the indicator value set. Next, a similarity threshold is obtained, which indicates the minimum similarity with the indicator value set. The similarity threshold can be designed according to the actual situation. Then, the similarity between the target indicator value and the indicator value set is obtained. If the similarity between the target indicator value and the indicator value set is less than or equal to the similarity threshold, it indicates that the performance of the i-th type of resource is significantly different from that of the virtual instance A when it is running normally, and it is determined that the i-th type of resource has a performance abnormality for the virtual instance A.
[0198] It should be noted that, considering the performance metric values are vectors formed by the metric values of each metric, obtaining the similarity between the target metric value and the set of metric values can include: calculating the distance between the target metric value and the set of metric values, and obtaining the similarity based on the distance; the closer the distance, the higher the similarity, for example, using the reciprocal of the distance as the similarity. The distance can be calculated using various methods, such as Euclidean distance, Manhattan distance, or cosine distance, etc., and the specific method can be designed according to actual needs. Each metric value in the set of metric values can be considered as a point; therefore, the set of metric values can include a center point, which can be the mean of all points in the set of metric values. The target indicator value can be a virtual point or a point in the set of indicator values, as shown in Figure 5c. The set of indicator values also has a boundary, which is used to define the range of the set of indicator values. For example, it can be a circle, an ellipse, etc. Optionally, the method of calculating the distance between the target indicator value and the set of indicator values can include: calculating the distance between the target indicator value and the center point in the set of indicator values. Optionally, the method of calculating the distance between the target indicator value and the set of indicator values can include: calculating the minimum distance between the target indicator value and the boundary of the set of indicator values. For example, calculating the distance between the target indicator value and the boundary point (the point on the boundary of the set of indicator values) that is closest to the target indicator value.
[0199] For example, if the i-th type of resource is LLC, performance metrics may include mem_load_retired.l3_hit and em_load_retired.l3_miss.
[0200] The set of indicator values is indicator value 1, indicator value 2, indicator value 3, indicator value 4, indicator value 5. Indicator value 1 is the indicator value of mem_load_retired.l3_hit and the indicator value of mem_load_retired.l3_miss. Indicator values 2, 3, 4, and 5 are similar, differing only in their magnitude. Assuming indicator value 3 is the center point, and performance information i1 is the target indicator value ab: mem_load_retired.l3_hit If the reciprocal of the distance between the target metric value a and the metric value b of em_load_retired.l3_miss is less than or equal to the similarity threshold, it is determined that LLC is experiencing a performance anomaly for virtual instance A; or, if the boundary of the metric value set is determined, which includes metric value 1, metric value 2, metric value 3, metric value 4, metric value 5, and the minimum distance between the target metric value a and the boundary is determined, and the reciprocal of the minimum distance is less than or equal to the similarity threshold, it is determined that LLC is experiencing a performance anomaly for virtual instance A.In one optional implementation of this embodiment, performance information i1 is the target performance index value of the i-th type of resource for virtual instance A. Physical node 1 can send performance information i1 to scheduling device 300. Scheduling device 300 can determine the index value range or index value set based on the performance information of virtual instance A in physical node 1 and the i-th type of resource for related virtual instances in other physical nodes (for ease of description and distinction, this can be called historical related performance information). The historical related performance information can be multiple index values of the i-th type of resource for virtual instance A collected by physical node 1 under low load, and multiple index values of the i-th type of resource for related virtual instance A collected by other physical nodes under low load. Then, clustering the historical related performance information yields the index value set, and subsequently, the index value range. Related virtual instances can be understood as those related to virtual instances. Instance A provides one or more virtual instances of the same service. For example, virtual instance A runs application software A, as shown in Figure 5e. Application software A can be distributed across related virtual instances on physical node 1 and other physical nodes 100 (one or more). For instance, virtual instances on physical node 1 and other physical nodes 100 run the same application software A, or virtual instances on physical node 1 and other physical nodes 100 each run a portion of application software A. Then, if the target metric value is outside the metric value range, or if the similarity between the target metric value and the metric value set is less than or equal to the similarity threshold, the analysis result is determined to be that the i-th type of resource on physical node 1 exhibits performance anomalies for virtual instance A; otherwise, the analysis result is that the i-th type of resource on physical node 1 does not exhibit performance anomalies for virtual instance A. The detailed calculation of the similarity between the target metric value and the metric value set is described above and will not be repeated here. It should be noted that "other physical nodes" is merely one possible name, for example; other physical nodes can also be referred to as a fourth device.
[0201] In one optional implementation of this embodiment, physical node 1 can input performance information i1 into the first classification model. The first classification model can output analysis results, which are used to indicate whether the i-th type of resource has performance anomalies for virtual instance A. For example, 0 indicates no anomalies, and 1 indicates anomalies. This embodiment does not intend to limit the structure of the first classification model; it can be designed according to actual needs. For example, the first classification model can be a neural network model.
[0202] Optionally, the training method for the first classification model may include: acquiring multiple training samples of the i-th type of resource for virtual instance A, each training sample having a label indicating "normal" or "abnormal". It should be noted that the training samples are used to illustrate the performance of the i-th type of resource in physical node 1 for virtual instance A. The training samples can be real-collected performance information or artificially constructed performance information, such as the performance information of the i-th type of resource for virtual instance A under simulated performance abnormalities. Subsequently, physical node 1 iteratively trains the first classification model based on the multiple training samples until the training completion condition is met. The training completion condition refers to the conditions for stopping model training, including but not limited to reaching the maximum number of iterations, the model loss information reaching a preset threshold, and the model parameters no longer changing.
[0203] For example, physical node 1 can iteratively train the first classification model based on multiple training samples, which may include: for each training sample, inputting the training sample into the first classification model, determining the output of the first classification model, obtaining the error corresponding to the training sample based on the error between the output of the first classification model and the label; updating the model parameters of the first classification model based on the errors corresponding to multiple training samples, and then replacing the previous first classification model with the updated first classification model for the next iteration.
[0204] In one optional implementation of this embodiment, as shown in Figure 5d, physical node 1 can obtain performance information (referred to as reference performance information) of other resource classes besides the i-th resource class among the m resource classes for virtual instance A. These other resource classes can be one or more resource classes other than the i-th resource class among the m resource classes. Subsequently, physical node 1 analyzes whether the other resource classes affect the performance of the i-th resource class based on the performance information i1 of the i-th resource class and the reference performance information of the other resource classes for virtual instance A. If it is determined that the other resource classes affect the performance of the i-th resource class, it is determined that the i-th resource class has a performance anomaly. It should be noted that "reference performance information" is merely one possible name; for example, it can also be called "second performance information," and "other resource classes" is also merely one possible name; for example, it can also be called "second resource."
[0205] Optionally, physical node 1 can input performance information i1 and reference performance information of other resource types for virtual instance A into the second classification model. The second classification model can output analysis results, which are used to indicate whether the i-th resource type has performance anomalies for virtual instance A. For example, 0 indicates no anomaly, and 1 indicates an anomaly. This embodiment does not intend to limit the structure of the second classification model. It can be designed according to actual needs. For example, the second classification model can be a neural network model.
[0206] The training method for the second classification model is the same as that for the first training model. The difference lies in the training samples. The second classification model uses multiple training samples. Each training sample describes the performance of each resource in the m categories for virtual instance A. Each training sample has a label, which indicates whether each resource in the m categories is abnormal. It should be noted that the training samples can be real performance information or artificially constructed performance information, such as the performance information of each resource in the m categories for virtual instance A when several resources in the m categories experience performance abnormalities.
[0207] In one optional implementation of this embodiment, as shown in Figure 5e, physical node 1 can send performance information i1 to scheduling device 300. Scheduling device 300 can analyze whether there is a significant difference between the performance of the i-th type of resource for virtual instance A and the performance of the i-th type of resource for the relevant virtual instance, based on performance information i1 and performance information of the i-th type of resource in other physical nodes for the relevant virtual instance (for ease of description and distinction, this can be referred to as relevant performance information). If there is a significant difference, the analysis result is determined to be that the i-th type of resource of physical node 1 has a performance anomaly for virtual instance A; otherwise, the analysis result is determined to be that the i-th type of resource of physical node 1 has no performance anomaly for virtual instance A, and the analysis result is sent to physical node 1. It should be noted that "relevant performance information" is merely one possible name, for example; it can also be referred to as "fourth performance information." "Other physical nodes" is also merely one possible name, for example; it can also be referred to as "fourth device."
[0208] Optionally, in one example, the scheduling device 300 determines that the performance of the i-th type of resource in physical node 1 for virtual instance A is abnormal based on the performance information i1 and the relevant performance information of the i-th type of resource in other physical nodes for the relevant virtual instance. If the difference between the index value of the performance index of the i-th type of resource in physical node 1 for virtual instance A and the index value of the performance index of the i-th type of resource in other physical nodes for the relevant virtual instance is greater than or equal to a preset difference threshold, the scheduling device 300 determines that the performance of the i-th type of resource in physical node 1 for virtual instance A is abnormal.
[0209] Optionally, in one example, physical node 1 inputs performance information i1 and relevant performance information of the i-th type of resources from other physical nodes for the relevant virtual instance into the third classification model. The training method of the third classification model is the same as that of the first training model, except that the training samples are different. The training samples are used to describe the performance of the i-th type of resources for virtual instance A, and the performance of the i-th type of resources for the relevant virtual instance. This embodiment does not intend to limit the structure of the third classification model; it can be designed according to actual needs. For example, the third classification model can be a neural network model.
[0210] In one optional implementation of this embodiment, physical node 1 can send performance information i1 and reference performance information of other types of resources to scheduling device 300. Scheduling device 300 can combine performance information i1, reference performance information of other types of resources, and performance information of the i-th type of resources and other types of resources in other physical nodes for the relevant virtual instance (for ease of description and distinction, this can be referred to as relevant performance information) to determine whether the i-th type of resources of physical node 1 has an abnormal performance for virtual instance A, and send the analysis result to physical node 1.
[0211] Optionally, in one example, physical node 1 analyzes whether there is a significant difference between the performance of the i-th type of resource for virtual instance A and the performance of the i-th type of resource for the relevant virtual instance, based on performance information i1, reference performance information of other types of resources, and relevant performance information of the i-th type of resource and other types of resources in other physical nodes for the relevant virtual instance, and whether other types of resources affect the performance of the i-th type of resource. If there is a significant difference and / or affects the performance of the i-th type of resource, the analysis result is determined to be that the i-th type of resource of physical node 1 has a performance abnormality for virtual instance A; otherwise, the analysis result is determined to be that the i-th type of resource of physical node 1 has no performance abnormality for virtual instance A.
[0212] Optionally, in one example, physical node 1 inputs performance information i1, reference performance information, and relevant performance information of the i-th type of resources and other types of resources from other physical nodes for the relevant virtual instance into the fourth classification model. The training method of the fourth classification model is the same as that of the first training model, except that the training samples are different. The training samples are used to describe the performance of the i-th type of resources and other types of resources for virtual instance A, and the performance of the i-th type of resources and other types of resources for the relevant virtual instance. This embodiment does not intend to limit the structure of the fourth classification model. It can be designed according to actual needs. For example, the fourth classification model can be a neural network model.
[0213] Step 404: Physical node 1 allocates resource i1 from the i-th type of resource to virtual instance A for exclusive use.
[0214] In this embodiment, when physical node 1 determines that the i-th type of resource has a performance anomaly, it performs exclusive isolation of the i-th type of resource for virtual instance A, allocating resource i1 (also referred to as the first resource) from the i-th type of resource to virtual instance A for exclusive use. This reduces interference to virtual instance A under shared conditions and increases the probability of the high-priority virtual instance A operating normally. Here, resource AI can be at least a portion of the shared resources of virtual instance A, or it can be unused resources from the i-th type of resource.
[0215] It should be noted that before allocating virtual instance A, physical node 1 can determine that the resource status of virtual instance A for the i-th type of resource includes exclusive. For example, the resource status can be exclusive; for example, the resource status can be both exclusive and shared.
[0216] Optionally, if virtual instance A specifies a resource requirement for the i-th type of resource, and the resource status is exclusive and shared, then the amount of the i-th type of resource exclusively used and shared by virtual instance A is equal to the resource requirement. For example, the i-th type of resource can be CPU cores, the total number of CPU cores is 64, the number of CPU cores specified by virtual instance A is 16, and the number of CPU cores shared by virtual instance A and exclusively used by virtual instance A and other virtual instances 130 is 16. If the resource status is exclusive, then the amount of the i-th type of resource exclusively used by virtual instance A is less than or equal to the resource requirement. For example, the i-th type of resource can be CPU cores, the total number of CPU cores is 64, the number of CPU cores specified by virtual instance A is 16, and the number of CPU cores exclusively used by virtual instance A can be less than 16, such as 4.
[0217] Optionally, if virtual instance A does not specify resource requirements for the i-th type of resource, and the resource status is exclusive and shared, then the amount of the i-th type of resource exclusively used and shared by virtual instance A is less than or equal to the total amount of the i-th type of resource. For example, the i-th type of resource can be an LLC, and the total amount of resources of the LLC is 12 channels. The number of channels shared and exclusively used by virtual instance A is less than or equal to 12 channels. If the resource status is exclusive, then the amount of the i-th type of resource exclusively used by virtual instance A is less than or equal to the total amount of resources of the i-th type of resource. For example, the i-th type of resource can be an LLC, and the total amount of resources of the LLC is 12 channels. The number of channels exclusively used by virtual instance A is less than or equal to 12 channels.
[0218] In one optional implementation of this embodiment, physical node 1 can obtain the exclusive resource amount of virtual instance A for the i-th type of resource; and allocate resource i1 (also called the first resource) in the i-th type of resource that is adapted to the exclusive resource amount.
[0219] Optionally, virtual instance A specifies a resource requirement for the i-th type of resource; physical node 1 can obtain the exclusive resource quantity of virtual instance A for the i-th type of resource by: obtaining the resource requirement of virtual instance A for the i-th type of resource, and obtaining the exclusive resource quantity based on the resource requirement.
[0220] Alternatively, in one example, the amount of exclusive resources can be equal to the amount of resources required.
[0221] Alternatively, in another example, the amount of exclusive resources can be less than the resource requirement, thereby reserving more resources for other virtual instances 130 and reducing the impact on other virtual instances 130. For example, the amount of exclusive resources can be determined based on the proportion of the resource requirement of the i-th type of resource. This proportion can be pre-set or the ratio of the resource requirement of the i-th type of resource to the total resource quantity of the i-th type of resource. For instance, if the i-th type of resource can be CPU cores, and the total number of CPU cores is 64, and virtual instance A specifies 16 CPU cores, then the proportion of resource requirement is (16 / 64), and the number of CPU cores allocated to it is 16*(16 / 64) = 4.
[0222] Optionally, virtual instance A does not specify resource requirements for the i-th type of resource. Virtual instance A can fully utilize the i-th type of resource provided by physical node 1. To reduce the impact on other virtual instances 130, the exclusive resource amount of the i-th type of resource is less than the total resource amount of the i-th type of resource. For example, the i-th type of resource can be memory bandwidth or LLC.
[0223] Optionally, in one example, the amount of exclusive resources can be determined based on the proportion of the total resources of the i-th type of resources. This proportion can be the ratio of the resource requirement of the j-th type of resources specified by virtual instance A to the total resources of the j-th type of resources. The j-th type of resources can be CPU cores. By using the ratio of the number of virtual CPU cores to the number of physical CPU cores, exclusive resources of other types of resources are allocated, ensuring that the allocated resources match the computing power of the required CPU cores and guaranteeing the performance of high-priority virtual instances. For example, the j-th type of resources can be CPU cores, and the i-th type of resources can be LLCs. For instance, if the total number of CPU cores is 64 and the LLC is 12-way, and virtual instance A specifies 16 CPU cores, then the exclusive resource amount allocated to its LLCs is 12 * (16 / 64) = 3-way.
[0224] In one optional scenario of this embodiment, the i-th type of resource can be a CPU core. Optionally, physical node 1 allocates a certain number of CPU cores to virtual instance A for exclusive use, based on the number of CPU cores required by virtual instance A. Optionally, physical node 1 supports dynamically adjusting the number of CPU cores, determining the number of exclusive CPU cores required by virtual instance A, such as one, and allocating a certain number of exclusive CPU cores to virtual instance A for exclusive use.
[0225] In one optional scenario of this embodiment, the i-th type of resource can be an LLC of the CPU. Optionally, physical node 1 determines the exclusive resource amount of the LLC based on the proportion of the number of CPU cores required by virtual instance A to the total number of CPU cores. For example, if the total number of CPU cores is 64, the LLC is 12-way, and the number of CPU cores specified by virtual instance A is 16, then the exclusive resource amount of its allocated LLC is 12*(16 / 64) = 3-way.
[0226] In one optional scenario of this embodiment, the i-th type of resource can be memory bandwidth. Optionally, physical node 1 determines the amount of dedicated memory bandwidth resources based on the proportion of the number of CPU cores required by virtual instance A to the total number of CPU cores.
[0227] In this solution, when multiple types of resources are shared, if any type of resource experiences a performance issue for a high-priority virtual instance, at least a portion of the resource of the type experiencing the performance issue will be allocated to the high-priority virtual instance for exclusive use. This reduces the interference of a particular type of resource on the high-priority virtual instance, thereby ensuring the performance of the high-priority virtual instance while maintaining resource utilization.
[0228] In some possible scenarios, as shown in Figure 5f, physical node 1 includes virtual instance 1, virtual instance 2, virtual instance 3, ..., virtual instance n, as well as three types of resources: CPU cores, LLC and memory bandwidth. Virtual instance 1, virtual instance 2, and virtual instance 3 are high-priority virtual instances 1, 3, and 4 respectively, and the priority of other virtual instances is lower than that of virtual instance 1, virtual instance 2, and virtual instance 3.
[0229] The CPU cores are used separately for the performance anomalies of virtual instance 1, virtual instance 2, and virtual instance 3. Virtual instance 1, virtual instance 2, and virtual instance 3 each exclusively use a portion of the CPU cores, while other virtual instances share the remaining CPU cores.
[0230] In response to the performance anomaly of virtual instance 2, virtual instance 2 exclusively uses a portion of the LLC, while other virtual instances share the remaining LLC.
[0231] Regarding the performance anomaly of virtual instance 3, virtual instance 3 exclusively uses a portion of the LLC, while other virtual instances share the remaining memory bandwidth.
[0232] When physical node 1 determines that the i-th type of resource is abnormal, it needs to isolate the i-th type of resource of virtual instance A on demand. In the process of on-demand isolation, feedback-based cybernetics regulation can be adopted. For example, as shown in Figure 6a, if the performance information i1 of the i-th type of resource is abnormal during the operation of virtual instance A, the resource i1 in the i-th type of resource is allocated to virtual instance A for exclusive use. Then, during the operation of virtual instance A, the performance information i2 of resource i1 is obtained. Based on the performance information i2, it is analyzed whether resource i1 has a performance abnormality. If so, the resource i2 in the i-th type of resource is allocated to virtual instance A for exclusive use. Then, during the operation of virtual instance A, the performance information i3 of the resources (resource i1 and resource i2) exclusively used by virtual instance A is obtained. Based on the performance information i3, it is analyzed whether the resources (resource i1 and resource i2) exclusively used by virtual instance A have a performance abnormality. If so, the resource i3 in the i-th type of resource is allocated to virtual instance A for exclusive use. This cycle is repeated until the performance of the resources exclusively used by virtual instance A is normal, thereby reducing the interference to virtual instances in the shared situation and increasing the probability of high-priority virtual instances running normally.
[0233] This description is merely illustrative; for details, please refer to Figure 6b and its description.
[0234] Figure 6b shows a flowchart of another resource allocation method provided by an embodiment of the present invention.
[0235] As shown in Figure 6b, based on steps 401 to 404 shown in Figure 2, this embodiment of the invention further includes at least the following steps:
[0236] Step 405: When physical node 1 is running virtual instance A, it obtains the performance information i2 of resource i1. The performance information i2 is used to indicate the performance of virtual instance A for resource i1.
[0237] Step 406: Based on performance information i2, physical node 1 determines whether resource i1 has a performance abnormality. If so, proceed to step 407.
[0238] For details, please refer to the description of step 303, which will not be repeated here. The difference is that the performance information i1 is replaced with the performance information i2.
[0239] Step 407: Physical node 1 allocates resource i2 from the i-th type of resource to virtual instance A for exclusive use.
[0240] In this embodiment, if resource AI malfunctions, indicating insufficient resource i1, resource i2 (also known as the second resource) from the i-th resource category is allocated to virtual instance A for exclusive use. Steps 401 to 402 are repeated, continuously monitoring the performance information of all resources of the i-th resource category used by virtual instance A until the performance of the i-th resource category used by virtual instance A returns to normal.
[0241] Optionally, physical node 1 can determine the exclusive resource increment for virtual instance A for the i-th type of resource, and determine resource i2 (also called the second resource) within the i-th type of resource that adapts to the exclusive resource increment. The exclusive resource increment can be pre-set; for example, if the i-th type of resource is a CPU core, the exclusive resource increment can be 1 CPU core. Alternatively, if the i-th type of resource is an LLC, the exclusive resource increment can be 1 way.
[0242] In one optional case of this embodiment, i=1, the first type of resource can be a CPU core, and physical node 1 supports dynamic adjustment of the number of CPU cores.
[0243] Optionally, based on a pre-set initial number of dedicated CPU cores, such as one, the number of CPU cores initially allocated to virtual instance A is determined for exclusive use. Performance information 12 of the CPU cores exclusively used by virtual instance A is collected. If performance information 12 is abnormal, one CPU core is added, and this process is repeated to gradually increase the number of CPU cores exclusively used by virtual instance A until the performance of the CPU cores exclusively used by virtual instance A is normal.
[0244] Optionally, as shown in Figure 6c, physical node 1 determines the number of dedicated CPU cores for virtual instance A based on the number of CPU cores required by virtual instance A and the proportion of the required CPU cores to the total number of CPU cores. For example, if the total number of CPU cores is 64 and virtual instance A specifies 16 CPU cores, then the number of dedicated CPU cores allocated to it is 16 * (16 / 64) = 4. These dedicated CPU cores are then allocated exclusively to virtual instance A. Performance information 12 of the dedicated CPU cores used by virtual instance A is collected. If performance information 12 shows an anomaly, one CPU core is added. This process is repeated, gradually increasing the number of dedicated CPU cores used by virtual instance A until the performance of the dedicated CPU cores used by virtual instance A is normal.
[0245] In one optional case of this embodiment, i = 2, and the second type of resource can be LLC. Optionally, as shown in Figure 6d, physical node 1 determines the exclusive resource amount of LLC based on the proportion of the number of CPU cores required by virtual instance A to the total number of CPU cores. For example, if the total number of CPU cores is 64, the LLC is 12-way, and the number of CPU cores specified by virtual instance A is 16, then the exclusive resource amount of its allocated LLC is 12*(16 / 64) = 3-way. Performance information 22 of the LLC exclusively used by virtual instance A is collected. If performance information 22 is abnormal, an LLC with one way is added. This process is repeated, gradually increasing the number of ways of the LLC exclusively used by virtual instance A until the performance of the LLC exclusively used by virtual instance A is normal.
[0246] In one optional case of this embodiment, i = 3, and the third type of resource can be memory bandwidth. Optionally, similar to Figure 6d, physical node 1 determines the amount of dedicated memory bandwidth based on the proportion of the number of CPU cores required by virtual instance A to the total number of CPU cores. Performance information 32 of the dedicated memory bandwidth used by virtual instance A is collected. If performance information 32 is abnormal, the memory bandwidth is increased. This process is repeated, gradually increasing the dedicated memory bandwidth used by virtual instance A until the performance of the dedicated memory bandwidth used by virtual instance A is normal.
[0247] In this solution, when a certain type of resource experiences performance anomalies for a virtual instance, a resource allocation scheme based on a feedback mechanism is adopted. This scheme can continuously allocate that type of resource to high-priority virtual instances until the performance of that type of resource for the high-priority virtual instance returns to normal, thereby ensuring the performance of the high-priority virtual instance.
[0248] In an alternative embodiment, the i-th type of resource may further include shared resources (also referred to as third resources), which are resources shared by virtual instance A and the other virtual instances 130 among the n virtual instances 130.
[0249] In this solution, high-priority virtual instances can not only exclusively use resources but also share resources with other virtual instances, thereby further ensuring the performance of high-priority virtual instances.
[0250] As shown in Figure 7a, when physical node 1 determines that the i-th type of resource is abnormal, it needs to isolate the i-th type of resource of virtual instance A on demand, and can exclusively use at least a portion of the i-th type of resource; for example, in the process of on-demand isolation, feedback-based cybernetics regulation can be adopted (see Figure 6a for details). If the i-th type of resource exclusively used by virtual instance A continues to have performance abnormalities, it is necessary to continuously increase the amount of exclusive resources of virtual instance A for the i-th type of resource.
[0251] Considering that the i-th type of resources of physical node 1 is limited, as shown in Figure 7a, physical node 1 needs to determine whether the i-th type of resources is insufficient based on the amount of exclusive resources of virtual instance A for the i-th type of resources and the total amount of resources of the i-th type of resources. For example, if the amount of exclusive resources of virtual instance A for the i-th type of resources reaches or is close to the total amount of resources of the i-th type of resources, it is determined that the i-th type of resources is insufficient. If the i-th type of resources is determined to be insufficient, resource insufficiency information is determined, which is used to indicate that there is a shortage of resources for virtual instance A for the i-th type of resources. Subsequently, in order to ensure the normal operation of virtual instance A, physical node 1 needs to send the resource insufficiency information to the scheduling device 300. Based on the resource insufficiency information, the scheduling device 300 determines the virtual instance 130 that needs to be migrated in physical node 1.
[0252] This description is merely illustrative; for details, please refer to Figure 7b and its description.
[0253] Figure 7b shows a flowchart of another resource allocation method provided by an embodiment of the present invention. As shown in Figure 7b, based on steps 401 to 404 shown in Figure 4 and steps 405 to 407 shown in Figure 6, this embodiment of the present invention further includes at least the following steps:
[0254] Step 701: Physical node 1 determines resource insufficiency information based on the amount of resources exclusively used by virtual instance A for the i-th type of resource and the total amount of resources for the i-th type of resource; wherein, the resource insufficiency information is used to indicate that there is a shortage of resources for the i-th type of resource for virtual instance A.
[0255] The resource shortage information can include the identifier of physical node 1, such as its number, and the identifier of the i-th type of resource.
[0256] In one optional implementation of this embodiment, if the amount of resources exclusively used by virtual instance A for the i-th type of resource is the same as the total amount of resources for the i-th type of resource, or if the proportion of the amount of resources exclusively used by virtual instance A for the i-th type of resource to the total amount of resources for the i-th type of resource is less than or equal to a preset threshold, then physical node 1 is considered to have insufficient resources for the i-th type of resource.
[0257] Optionally, for the amount of resources exclusively used by virtual instance A for the i-th type of resource, the performance of the exclusive resources used by virtual instance A for the i-th type of resource is abnormally normal.
[0258] Step 702: Physical node 1 sends a resource shortage message to scheduling device 300.
[0259] In this solution, when a certain type of resource experiences performance anomalies for a virtual instance, it analyzes whether there is insufficient resources for that type of resource for high-priority virtual instances. If resources are insufficient, it notifies the scheduling device, which then adjusts the virtual instances in physical node 1 to ensure the performance of high-priority virtual instances.
[0260] In the scenario shown in Figure 7a, the calling device 300 determines that the virtual instance 130 that needs to be migrated in physical node 1 can be a virtual instance B other than virtual instance A, and determines that the device to which virtual instance B needs to be migrated can be physical node 2. Physical node 2 is any other physical node in the calling system other than physical node 1. By adjusting the number of virtual instances in physical node 1, the performance of high-priority virtual instances can be guaranteed.
[0261] Figure 8a shows a flowchart of another resource allocation method provided by an embodiment of the present invention.
[0262] As shown in Figure 8a, based on steps 401 to 404 shown in Figure 4 and steps 405 to 407 shown in Figure 6, this embodiment of the invention further includes at least the following steps:
[0263] Step 801: Based on the resource shortage information, device 300 determines the first migration instruction. The first migration instruction is used to instruct the migration of virtual instance B to physical node 2. Virtual instance B is virtual instance 130 other than virtual instance A in physical node 1.
[0264] The first migration instruction may include the identifier of virtual instance B and the identifier of physical node 2. Furthermore, there may be one or more virtual instances B.
[0265] Here, physical node 2 can be a device with a low load, such as running a smaller number of virtual instances 130, thereby reducing the probability of failure after virtual instance 2 is migrated to physical node 2.
[0266] Step 802: Invoke device 300 to send the first migration command to physical node 1.
[0267] Step 803: Physical node 1 responds to the first migration command and migrates virtual instance B to physical node 2.
[0268] Step 804: Physical node 1 determines that virtual instance A and other virtual instances 130 share the i-th type of resource.
[0269] After physical node 1 removes virtual instance 2, the available resources in resource i increase. At this time, virtual instance 1's exclusive access to resource i can be revoked, and virtual instance 1 and virtual instance 3 can share resource i, thereby improving resource utilization.
[0270] It should be noted that step 804 is an optional step, not a mandatory step. Step 804 can be added or deleted depending on the actual situation.
[0271] In one optional scenario of this embodiment, as shown in Figure 8b, physical node 1 determines that there is a shortage of resources of type i based on the exclusive resource amount of virtual instance 1 for type i and the total resource amount of type i. The resource shortage information is used to indicate that there is a shortage of resources of type i for virtual instance A.
[0272] Next, to ensure the normal operation of virtual instance A, physical node 1 needs to send resource shortage information to scheduling device 300. Based on the resource shortage information, scheduling device 300 determines that the virtual instance 130 that needs to be migrated from physical node 1 can be a virtual instance other than virtual instance 1, such as virtual instance 2, and determines that the device to which virtual instance 2 needs to be migrated can be physical node 2. Physical node 2 can be any other physical node in the calling system besides physical node 1, thus obtaining a first migration instruction. The first migration instruction is used to instruct virtual instance 2 to be migrated to physical node 2. Here, physical node 2 can be a device with a lower load, such as running a smaller number of virtual instances 130, thereby reducing the probability of failure after virtual instance 2 is migrated to physical node 2.
[0273] Next, physical node 1 responds to the first migration command and migrates virtual instance 2 to physical node 2.
[0274] Subsequently, after physical node 1 removes virtual instance 2, the available resources increase. At this time, the exclusive right of virtual instance 1 to the i-th type of resource can be revoked, and virtual instance 1 and virtual instance 3 can share the i-th type of resource, thereby improving resource utilization.
[0275] In this solution, if a certain type of resource experiences performance anomalies for a virtual instance, and there is insufficient resources for that type of resource for high-priority virtual instances, the number of virtual instances in physical node 1 will be adjusted to ensure the performance of high-priority virtual instances.
[0276] In the scenario shown in Figure 7a, the calling device 300 determines that the virtual instance 130 that needs to be migrated in physical node 1 can be virtual instance A, and determines that the device to which virtual instance A needs to be migrated can be physical node 3. Physical node 3 is the physical node in the calling system whose performance of the i-th type of resource is higher than any other physical node besides physical node 1. Thus, by migrating the high-priority virtual instance to the physical node with stronger performance, the performance of the high-priority virtual instance is guaranteed.
[0277] As shown in Figure 9a, based on steps 401 to 404 shown in Figure 4 and steps 405 to 407 shown in Figure 6, this embodiment of the invention further includes at least the following steps:
[0278] Step 901: Based on the resource shortage information, device 300 determines a second migration instruction. The second migration instruction is used to instruct the virtual instance A to migrate to physical node 3, where the performance of the i-th type of resource in physical node 3 is greater than that of the i-th type of resource in physical node 1.
[0279] For example, the i-th type of resource can be a CPU core, and the CPU core in physical node 3 has stronger computing power.
[0280] For example, the i-th type of resource can be an LLC, and the LLC in physical node 3 has a larger storage space.
[0281] For example, the i-th type of resource can be memory bandwidth, with physical node 3 having a larger memory bandwidth.
[0282] Step 902: Invoke device 300 to send a second migration command to physical node 1.
[0283] Step 903: Physical node 1 responds to the second migration command and migrates virtual instance A to physical node 3.
[0284] In one optional scenario of this embodiment, as shown in Figure 9b, physical node 1 determines that there is a shortage of resources of type i based on the exclusive resource amount of virtual instance 1 for type i and the total resource amount of type i. The resource shortage information is used to indicate that there is a shortage of resources of type i for virtual instance A.
[0285] Next, in order to ensure the normal operation of virtual instance A, physical node 1 needs to send the resource shortage information to scheduling device 300. Based on the resource shortage information, scheduling device 300 determines that the virtual instance 130 that needs to be migrated in physical node 1 can be virtual instance 1, and determines that the device to which virtual instance 1 needs to be migrated can be physical node 3. Physical node 3 can be any other physical node in the calling system whose i-th type of resource is stronger than physical node 1, thereby obtaining a second migration instruction. The second migration instruction is used to instruct virtual instance 1 to be migrated to physical node 2.
[0286] Subsequently, physical node 1 responds to the second migration command and migrates virtual instance 1 to physical node 3.
[0287] In this solution, if a certain type of resource experiences performance anomalies for a virtual instance, and there is insufficient resources for high-priority virtual instances, the high-priority virtual instances will be migrated to a physical node with stronger performance, thereby ensuring the performance of high-priority virtual instances.
[0288] It should be noted that, for Figures 4a to 9b above, Physical Node 1, Physical Node 2, Physical Node 3, Other Physical Nodes, Resource i1, Resource i2, Performance Information i1, Performance Information i2, Reference Performance Information, Related Performance Information, Type i Resource, Other Type Resources, Virtual Instance A, Virtual Instance B, and Related Virtual Instance are one possible naming convention. Optionally, in one example, Physical Node 1, Physical Node 2, Physical Node 3, Other Physical Nodes, Resource i1, Resource i2, Performance Information i1, Performance Information i2, Reference Performance Information, Related Performance Information, Type i Resource, Other Type Resources, Virtual Instance A, Virtual Instance B, and Related Virtual Instance can also be named First Device, Second Device, Third Device, Fourth Device, First Resource, Second Resource, First Performance Information, Third Performance Information, Second Performance Information, Fourth Performance Information, First Type Resource, Second Type Resource, First Virtual Instance, Second Virtual Instance, and Third Virtual Instance.
[0289] The present invention also provides a resource allocation device, which includes multiple virtual instances and multiple types of resources, wherein the multiple virtual instances share the multiple types of resources; as shown in Figure 10, the resource allocation device includes:
[0290] The identification module is used to determine the first virtual instance; wherein the first virtual instance is any virtual instance among multiple virtual instances, and the first virtual instance has a higher priority than other virtual instances among multiple virtual instances;
[0291] The acquisition module is used to acquire first performance information of a first type of resource during the operation of the first virtual instance; wherein, the first type of resource is any one of multiple types of resources, and the first performance information is used to indicate the performance of the first type of resource for the first virtual instance;
[0292] The resource allocation module is used to allocate the first resource in the first type of resources to the first virtual instance for exclusive use when it is determined that the first type of resources have experienced performance abnormalities based on the first performance information.
[0293] The identification module, data acquisition module, and resource allocation module can all be implemented in software or hardware. For example, the implementation of the identification module will be described below. Similarly, the implementation methods of the data acquisition module and the resource allocation module can refer to the implementation method of the identification module.
[0294] As an example of a software functional unit, a module can include code running on a computing instance. A computing instance can include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, the aforementioned computing instance can be one or more. For example, a module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region can include multiple AZs.
[0295] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0296] As an example of a hardware functional unit, an identification module may include at least one computing device, such as a server. Alternatively, an identification module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The aforementioned PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0297] The identification module includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the identification module includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the identification module includes multiple computing devices that can be distributed within the same Virtual Private Cloud (VPC) or multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0298] It should be noted that, in other embodiments, the identification module can be used to execute any step in the above-described resource allocation method, such as any step in the method shown in Figures 4d, 6b, 7b, 8a, or 9a; the acquisition module can be used to execute any step in the above-described resource allocation method, such as any step in the method shown in Figures 4d, 6b, 7b, 8a, or 9a; and the resource allocation module can be used to execute any step in the above-described resource allocation method, such as any step in the method shown in Figures 4d, 6b, 7b, 8a, or 9a. The steps implemented by the identification module, acquisition module, and resource allocation module can be specified as needed. By implementing different steps in the above-described resource allocation method through the identification module, acquisition module, and resource allocation module, all functions of the resource allocation device can be realized.
[0299] The present invention also provides a computing device 1100. As shown in FIG11, the computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via the bus 1102. The computing device 1100 may be a server or a terminal device. It should be understood that the present invention does not limit the number of processors 1104 and memories in the computing device 1100.
[0300] Bus 1102 can be a Peripheral Component Interconnect (PCI) bus 1102 or an Extended Industry Standard Architecture (EISA) bus 1102, etc. Bus 1102 can be divided into address bus 1102, data bus 1102, control bus 1102, etc. For ease of illustration, only one line is used to represent it in Figure 11, but this does not mean that there is only one bus 1102 or a bus 1102 of one type of resource. Bus 1102 can include a path for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, communication interface 1108).
[0301] The processor 1104 may include any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0302] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0303] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the functions of the aforementioned identification module, acquisition module, and resource allocation module, thereby implementing the aforementioned resource allocation method, such as the method shown in FIG4d, FIG6b, FIG7b, FIG8a, or FIG9a. That is, the memory 1106 stores instructions for executing the aforementioned resource allocation method, such as the instructions shown in FIG4d, FIG6b, FIG7b, FIG8a, or FIG9a.
[0304] The communication interface 1108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1100 and other devices or communication networks.
[0305] This invention also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0306] As shown in Figure 12, the computing device cluster includes at least one computing device 1100. The memory 1106 of one or more computing devices 1100 in the computing device cluster may store the same instructions for performing the resource allocation method described above, such as the instructions for the method shown in Figures 4d, 6b, 7b, 8a, or 9a.
[0307] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the resource allocation method described above, such as partial instructions of the method shown in FIG4d, FIG6b, FIG7b, FIG8a, or FIG9a. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions for executing the resource allocation method described above, such as the instructions of the method shown in FIG4d, FIG6b, FIG7b, FIG8a, or FIG9a.
[0308] It should be noted that the memory 1106 in different computing devices 1100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the resource allocation device. That is, the instructions stored in the memory 1106 of different computing devices 1100 can implement the functions of one or more modules among the identification module, acquisition module, and resource allocation module.
[0309] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 13 illustrates one possible implementation. As shown in Figure 13, two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1106 in computing device 1100A stores instructions for performing the functions of the identification module and the acquisition module. Simultaneously, the memory 1106 in computing device 1100AB stores instructions for performing the functions of the resource allocation module.
[0310] The connection method between the computing device clusters shown in Figure 13 can be considered as follows: taking into account that the resource allocation method provided by the present invention needs to determine high-priority virtual instances and collect performance information of each type of resource for high-priority virtual instances, the function implemented by the resource allocation module is considered to be executed by the computing device 1100B.
[0311] It should be understood that the functions of computing device 1100A shown in Figure 13 can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.
[0312] This invention also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any available medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to perform the resource allocation method described above, such as the method shown in FIG4d, FIG6b, FIG7b, FIG8a, or FIG9a.
[0313] This invention also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the resource allocation method described above, such as the method shown in Figures 4d, 6b, 7b, 8a, or 9a.
[0314] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0315] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0316] The basic principles of the present invention have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in the present invention are merely examples and not limitations, and should not be considered as essential features of the various embodiments of the present disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of the present disclosure to the necessity of employing the specific details described above.
[0317] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0318] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0319] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
[0320] It is understood that the various numerical designations used in the embodiments of the present invention are merely for descriptive convenience and are not intended to limit the scope of the embodiments of the present invention.
Claims
1. A resource allocation method characterized by, Applied to a first device, the first device comprising multiple virtual instances and multiple types of resources, wherein the multiple virtual instances share the multiple types of resources, the method includes: A first virtual instance is determined; wherein the first virtual instance is any virtual instance among the plurality of virtual instances, and the first virtual instance has a higher priority than the other virtual instances among the plurality of virtual instances; During the operation of the first virtual instance, first performance information of a first type of resource is obtained; wherein, the first type of resource is any one of the multiple types of resources, and the first performance information is used to indicate the performance of the first type of resource for the first virtual instance; If it is determined based on the first performance information that the first type of resource has a performance abnormality, the first resource in the first type of resource will be allocated to the first virtual instance for exclusive use.
2. The method of claim 1, wherein, The method further includes: Based on the amount of resources exclusively used by the first virtual instance for the first type of resources and the total amount of resources of the first type of resources, resource insufficiency information is determined; wherein, the resource insufficiency information is used to indicate that the first type of resources are insufficient for the first virtual instance; The resource shortage information is sent to the scheduling device, and the calling device is used to determine the virtual instances in the first device that need to be migrated based on the resource shortage information.
3. The method of claim 2, wherein, The method further includes: The system receives a first migration instruction sent by the scheduling device; wherein the first migration instruction is used to instruct the second virtual instance to be migrated to the second device, and the second virtual instance is any virtual instance other than the first virtual instance among the plurality of virtual instances; In response to the first migration command, the second virtual instance is migrated to the second device; It is determined that the first virtual instance in the first device and other virtual instances in the first device besides the first virtual instance share the first resource.
4. The method according to claim 2, characterized in that, The method further includes: The system receives a second migration instruction sent by the scheduling device; wherein the second migration instruction is used to instruct the first virtual instance to be migrated to a third device, and the performance of the first type of resource in the third device is greater than the performance of the first type of resource in the first device; In response to the second migration instruction, the first virtual machine instance is migrated to the third device.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: During the operation of the first virtual instance, second performance information of the second type of resource is obtained. The second type of resource is any type of resource other than the first type of resource among the multiple types of resources. The second performance information is used to indicate the performance of the second type of resource for the first virtual instance. The step of determining that the first type of resource has a performance anomaly based on the first performance information includes: If, based on the first performance information and the second performance information, it is determined that the second type of resource affects the performance of the first type of resource, then it is determined that the first type of resource has experienced a performance anomaly.
6. The method according to any one of claims 1 to 4, characterized in that, The first performance information is the performance index value of the first type of resource for the first virtual instance; The step of determining that the first type of resource has a performance anomaly based on the first performance information includes: Obtain the set of indicator values corresponding to the performance indicator; wherein, the set of indicator values is used to indicate multiple indicator values of the performance indicator for the first virtual instance under normal operation of the first virtual instance; Obtain a similarity threshold; wherein the similarity threshold is used to indicate the minimum value of similarity with the set of index values; Obtain the similarity between the performance metric value for the first virtual instance and the set of metric values; If the similarity is less than or equal to the similarity threshold, it is determined that the first type of resource has a performance abnormality.
7. The method according to any one of claims 1 to 4, characterized in that, The first performance information is the performance index value of the first type of resource for the first virtual instance; The step of determining that the first type of resource has a performance anomaly based on the first performance information includes: Obtain the value range of the performance metric corresponding to the performance metric; wherein the value range is used to indicate the range of values of the performance metric for the first virtual instance under normal operation of the first virtual instance; If the performance metric value for the first virtual instance is outside the specified range, it is determined that the first type of resource has experienced a performance anomaly.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: During the operation of the first virtual instance, third performance information of the first resource is obtained, the third performance information being used to indicate the performance of the first resource for the first virtual instance; If it is determined that the first resource has a performance abnormality based on the third performance information, the second resource in the first type of resource is allocated to the first virtual instance for exclusive use, so that the first virtual instance can exclusively use the first resource and the second resource.
9. The method according to any one of claims 1 to 8, characterized in that, The first type of resource is memory bandwidth or a cache shared by the CPU cores of a central processing unit. Obtaining the exclusive resource amount of the first virtual instance for the first type of resource includes: Obtain the first number of virtual CPU cores of the first virtual instance; Obtain the second number of CPU cores in the first device; Determine the ratio of the first number to the second number; Based on the ratio and the total amount of resources of the first type of resources, the exclusive resource amount of the first virtual instance for the first type of resources is determined; Determine the first resource in the first type of resources that is compatible with the amount of exclusive resources.
10. The method according to any one of claims 1 to 9, characterized in that, The multiple types of resources include at least two of the following: Central processing unit (CPU) cores, memory bandwidth, and cache shared by the CPU cores.
11. A resource allocation device, characterized in that, The resource allocation device includes multiple virtual instances and multiple types of resources, wherein the multiple virtual instances share the multiple types of resources, and the device includes: An identification module is used to determine a first virtual instance; wherein the first virtual instance is any virtual instance among the plurality of virtual instances, and the first virtual instance has a higher priority than the other virtual instances among the plurality of virtual instances; The acquisition module is used to acquire first performance information of a first type of resource during the operation of the first virtual instance; wherein, the first type of resource is any one of the multiple types of resources, and the first performance information is used to indicate the performance of the first type of resource for the first virtual instance; The resource allocation module is used to allocate the first resource in the first type of resources to the first virtual instance for exclusive use when it is determined that the first type of resources have experienced performance abnormalities based on the first performance information.
12. The apparatus according to claim 11, characterized in that, The resource allocation module is used to determine resource shortage information based on the amount of resources exclusively used by the first virtual instance for the first type of resources and the total amount of resources of the first type of resources; wherein, the resource shortage information is used to indicate that the first type of resources are insufficient for the first virtual instance; and to send the resource shortage information to the scheduling device, wherein the calling device is used to determine the virtual instance that needs to be migrated in the first device based on the resource shortage information.
13. The apparatus according to claim 12, characterized in that, The device further includes: A first migration module is configured to receive a first migration instruction sent by the scheduling device; wherein the first migration instruction is configured to instruct the second virtual instance to be migrated to the second device, and the second virtual instance is any virtual instance other than the first virtual instance among the plurality of virtual instances; in response to the first migration instruction, the second virtual instance is migrated to the second device; The resource allocation module is used to determine that the first virtual instance in the first device and other virtual instances in the first device besides the first virtual instance share the first resource.
14. The apparatus according to claim 12, characterized in that, The device further includes: The second migration module is configured to receive a second migration instruction sent by the scheduling device; wherein the second migration instruction is configured to instruct the first virtual instance to be migrated to a third device, wherein the performance of the first type of resource in the third device is greater than the performance of the first type of resource in the first device; in response to the second migration instruction, the first virtual machine instance is migrated to the third device.
15. The apparatus according to any one of claims 11 to 14, characterized in that, The acquisition module is used to acquire second performance information of a second type of resource during the operation of the first virtual instance. The second type of resource is any resource other than the first type of resource among the multiple types of resources. The second performance information is used to indicate the performance of the second type of resource for the first virtual instance. The resource allocation module is used to determine that the first type of resource has a performance abnormality when it is determined, based on the first performance information and the second performance information, that the second type of resource affects the performance of the first type of resource.
16. The apparatus according to any one of claims 11 to 14, characterized in that, The first performance information is the performance index value of the first type of resource for the first virtual instance; The resource allocation module is configured to: obtain a set of indicator values corresponding to the performance indicator; wherein the set of indicator values indicates multiple indicator values of the performance indicator for the first virtual instance under normal operation of the first virtual instance; obtain a similarity threshold; wherein the similarity threshold indicates the minimum value of similarity with the set of indicator values; obtain the similarity between the indicator value of the performance indicator for the first virtual instance and the set of indicator values; and determine that the first type of resource has a performance abnormality if the similarity is less than or equal to the similarity threshold.
17. The apparatus according to any one of claims 11 to 14, characterized in that, The first performance information is the performance index value of the first type of resource for the first virtual instance; The resource allocation module is used to obtain the index value range corresponding to the performance index; wherein, the index value range is used to indicate the value range of the performance index for the first virtual instance under normal operation of the first virtual instance; when the index value of the performance index for the first virtual instance is outside the index value range, it is determined that the first type of resource has a performance abnormality.
18. The apparatus according to any one of claims 11 to 17, characterized in that, The resource allocation module is used to obtain third performance information of the first resource during the operation of the first virtual instance, and the third performance information is used to indicate the performance of the first resource for the first virtual instance. If it is determined that the first resource has a performance abnormality based on the third performance information, the second resource in the first type of resource is allocated to the first virtual instance for exclusive use, so that the first virtual instance can exclusively use the first resource and the second resource.
19. The apparatus according to any one of claims 11 to 18, characterized in that, The first type of resource is memory bandwidth or cache shared by the CPU cores of the central processing unit; The resource allocation module is used to obtain a first number of virtual CPU cores of the first virtual instance; obtain a second number of CPU cores in the first device; and determine the ratio of the first number to the second number. Based on the ratio and the total amount of resources of the first type of resources, the exclusive resource amount of the first virtual instance for the first type of resources is determined; Determine the first resource in the first type of resources that is compatible with the amount of exclusive resources.
20. The apparatus according to any one of claims 11 to 19, characterized in that, The multiple types of resources include at least two of the following: Central processing unit (CPU) cores, memory bandwidth, and cache shared by the CPU cores.
21. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 10.
22. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 10.
23. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 10.