Resource scheduling method based on cloud management platform and cloud management platform
By using the resource pooling scheduling method of the cloud management platform, the problem of resource fragmentation in cloud data centers is solved, and efficient utilization and flexible scheduling of resources are achieved.
Patent Information
- Application Number
- PCT/CN2025/086901
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-30
AI Technical Summary
In cloud data centers, the resources of different physical servers cannot be used in an efficient manner, resulting in low resource utilization and resource fragmentation.
Multiple resource pools are managed through a cloud management platform to achieve pooled scheduling of resources. Resources in the same or adjacent racks are connected by a high-speed bus, and resources are dynamically allocated according to tenant needs.
It improves resource utilization, avoids resource fragmentation, simplifies the creation, expansion, and migration of cloud instances, and enhances resource scheduling efficiency.
Smart Images

Figure CN2025086901_30102025_PF_FP_ABST
Abstract
Description
A resource scheduling method based on a cloud management platform and the cloud management platform.
[0001] This application claims priority to Chinese Patent Application No. 202410494340.8, filed with the State Intellectual Property Office of China on April 23, 2024, entitled "A Resource Scheduling Method and Cloud Management Platform Based on a Cloud Management Platform", and to Chinese Patent Application No. 202410703571.5, filed with the State Intellectual Property Office of China on May 31, 2024, entitled "A Resource Scheduling Method and Cloud Management Platform Based on a Cloud Management Platform", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud technology, and in particular to a resource scheduling method based on a cloud management platform and a cloud management platform. Background Technology
[0003] With the rapid development of cloud technology, more and more tenants are choosing to purchase cloud instances provided by cloud vendors to complete their business operations. Generally, tenant cloud instances can be deployed on physical servers in cloud data centers. Cloud instances can use various types of resources on physical servers to complete tenant business operations and thus meet tenant business needs.
[0004] In related technologies, a cloud data center can contain multiple physical servers, each with a fixed amount of computing, storage, and network resources. Based on these resources, multiple cloud instances can be deployed on each physical server. Because the resources of different physical servers are relatively independent—meaning they typically cannot communicate at high speed—their resources cannot usually be combined for collaborative use.
[0005] In the aforementioned technologies, resource management and scheduling are highly limited. When a cloud instance requires a large amount of resources, there will often be insufficient resources on a certain physical server, while the resources of other physical servers remain idle, resulting in low resource utilization of the cloud data center. Summary of the Invention
[0006] This application provides a resource scheduling method based on a cloud management platform and a cloud management platform, which can improve resource utilization to a certain extent and avoid resource fragmentation.
[0007] The first aspect of this application provides a resource scheduling method based on a cloud management platform. The cloud management platform implementing this method can manage infrastructure providing cloud services. This infrastructure includes multiple resource pools of different types, each resource pool containing multiple resources located in the same or adjacent racks, and each resource pool containing multiple resources of the same type. The method includes:
[0008] When a tenant needs to create a cloud instance, the cloud management platform can provide a management interface to the tenant. The tenant can then input a cloud instance creation request into the management interface through their client. This request specifies the first type and first specification of resources required for the tenant's first cloud instance. In this way, the cloud management platform can receive the cloud instance creation request sent by the tenant through the management interface.
[0009] Upon receiving a cloud instance creation request, the cloud management platform can select a first resource pool that meets the first type from multiple resource pools, and then select a first resource that meets the first specification from the first resource pool. It should be noted that since the first resource pool contains multiple resources, the first resource selected by the cloud management platform is usually one or more of these resources.
[0010] After determining the primary resource, the cloud management platform can create the primary cloud instance and allocate the primary resource to the primary cloud instance for its use, thereby meeting the tenant's cloud instance creation needs.
[0011] As can be seen from the above method, the first resource pool selected by the cloud management platform contains multiple resources located in the same rack or adjacent racks. This is equivalent to presenting multiple resources of the same type in a pooled manner. Since the tenant's first cloud instance and these multiple resources in the first resource pool are connected through a high-speed bus, when the tenant's first cloud instance needs one or more of these resources, the cloud management platform can schedule resources in the first resource pool according to the specifications of the resources required, and allocate a sufficient number of first resources to the tenant's first cloud instance for use. Therefore, during the cloud instance creation process, resource utilization can be improved to a certain extent and resource fragmentation can be avoided.
[0012] In one possible implementation, the cloud instance creation request is further used to instruct the tenant to set a Service Level Agreement (SLA) threshold. The cloud management platform determines, from a first resource pool, first resources that meet the first specification by: the cloud management platform determining multiple first resource groups that meet the first specification from multiple resources in the first resource pool, each first resource group containing at least one resource; the cloud management platform determining first SLA metrics for the multiple first resource groups, where the first SLA metric of any first resource group includes at least one of the following: the latency required for a first cloud instance to access the first resource group, and the bandwidth required for a first cloud instance to access the first resource group; and the cloud management platform selecting first resource groups whose first SLA metric is less than the SLA threshold as first resources. In the aforementioned implementation, after determining the first resource pool, the cloud management platform can determine multiple first resource groups that meet the first specification from multiple resources in the first resource pool, each first resource group containing at least one resource from the first resource pool. After identifying multiple primary resource groups, the cloud management platform can evaluate these groups to obtain primary SLA metrics. These primary SLA metrics include one or more of the following: latency required for the primary cloud instance to access the multiple primary resource groups, and bandwidth required for the primary cloud instance to access the multiple primary resource groups. Once the primary SLA metrics are obtained, the cloud management platform can select the primary resource group whose primary SLA metric is less than the SLA threshold as the primary resource available for the primary cloud instance.
[0013] In one possible implementation, the method further includes: a cloud management platform receiving a cloud instance expansion request from a tenant, the cloud instance expansion request indicating a first type and a second specification of resources required to be added to the tenant's first cloud instance; based on the cloud instance expansion request, the cloud management platform determining a first resource pool that meets the first type from multiple resource pools, and determining a second resource that meets the second specification from the first resource pool, wherein the first resource and the second resource are different resources within the first resource pool; and the cloud management platform allocating the second resource to the first cloud instance. In the aforementioned implementation, after the first cloud instance is created, if the tenant needs to expand the first cloud instance, the cloud management platform can provide a management interface to the tenant. The tenant can then input a cloud instance expansion request specified by the tenant into the management interface, the cloud instance expansion request indicating the first type and the second specification of resources required to be added to the tenant's first cloud instance. In this way, the cloud management platform can receive the cloud instance expansion request sent by the tenant through the management interface. After receiving the cloud instance expansion request, the cloud management platform can select a first resource pool that meets the first type from multiple resource pools based on the cloud instance expansion request, and select a second resource that meets the second specification from the first resource pool. It should be noted that the second resource and the first resource are typically different resources among multiple resources in the first resource pool. After determining the second resource, the cloud management platform can allocate the second resource to the first cloud instance for its use, thereby meeting the tenant's cloud instance expansion needs. Therefore, when a tenant needs to expand the first cloud instance, the cloud management platform selects the first resource pool for it. Since the tenant's first cloud instance and the other resources in the first resource pool (excluding the first resource) are connected via a high-speed bus, when the tenant's first cloud instance needs one or more of the remaining resources, the cloud management platform can schedule resources in the first resource pool according to the required resource specifications, allocating a sufficient number of second resources to the tenant's first cloud instance. Thus, during the cloud instance expansion process, resource utilization can be improved to a certain extent, avoiding resource fragmentation.
[0014] In one possible implementation, the cloud management platform determines second resources that meet the second specification from the first resource pool by: identifying multiple second resource groups that meet the second specification from the remaining resources in the first resource pool (excluding the first resources); determining second SLA metrics for the multiple second resource groups; and selecting second resource groups whose second SLA metrics are less than the SLA threshold as second resources. In the aforementioned implementation, after determining the first resource pool, the cloud management platform can determine multiple second resource groups that meet the second specification from the remaining resources in the first resource pool (excluding the first resources), each second resource group containing at least one resource from the first resource pool. After determining the multiple second resource groups, the cloud management platform can evaluate the multiple second resource groups to obtain second SLA metrics for the multiple second resource groups. The second SLA metrics for the multiple second resource groups include one or more of the following information: the latency required for the first cloud instance to access the multiple second resource groups, and the bandwidth required for the first cloud instance to access the multiple second resource groups. After obtaining the second SLA metrics for multiple second resource groups, the cloud management platform can select a second resource group whose second SLA metrics are less than the SLA threshold as a second resource available for use by the first cloud instance.
[0015] In one possible implementation, the cloud instance creation request is further used to indicate a second type and a third specification of resources required by the tenant's first cloud instance. The method further includes: the cloud management platform determining a second resource pool that meets the second type from multiple resource pools, and determining a third resource that meets the third specification from the second resource pool, wherein the third resource is at least one of multiple resources in the second resource pool; and the cloud management platform allocating the third resource to the first cloud instance.
[0016] In one possible implementation, the method further includes: a cloud management platform receiving a cloud instance migration request from a tenant for a first cloud instance; the cloud management platform determining a third SLA metric and a fourth SLA metric for the first resource based on the cloud instance migration request; if the third SLA metric and the fourth SLA metric are less than an SLA threshold, the cloud management platform creating a second cloud instance, allocating the first resource and the third resource to the second cloud instance, and releasing the first cloud instance; if the third SLA metric is greater than or equal to the SLA threshold and the fourth SLA metric is less than the SLA threshold, the cloud management platform determining a fourth resource matching the first resource in the first resource pool, creating a second cloud instance, allocating the fourth resource and the third resource to the second cloud instance, and releasing the first cloud instance; if the third SLA metric and the fourth SLA metric are greater than or equal to the SLA threshold, the cloud management platform determining a fourth resource matching the first resource in the first resource pool, determining a fifth resource matching the third resource in the second resource pool, creating a second cloud instance, allocating the fourth resource and the fifth resource to the second cloud instance, and releasing the first cloud instance. In the aforementioned implementation, when a tenant determines that the first cloud instance needs to be migrated during its use, the cloud management platform can provide a management interface to the tenant. The tenant can then input a cloud instance migration request for the first cloud instance into the management interface, allowing the cloud management platform to receive the migration request sent by the tenant through the client. Upon receiving the migration request, the cloud management platform can evaluate the first and third resources used by the first cloud instance, thereby obtaining a third SLA metric and a fourth SLA metric for the first resource. The third SLA metric includes one or more information such as the latency required for the second cloud instance to access the first resource and the bandwidth required for the second cloud instance to access the first resource. The fourth SLA metric includes one or more information such as the latency required for the second cloud instance to access the third resource and the bandwidth required for the second instance to access the third resource. After obtaining the third and fourth SLA metrics, the cloud management platform can determine the relationship between the third and fourth SLA metrics and the SLA threshold set by the tenant. Based on this relationship, the cloud management platform can perform different migration operations for the first cloud instance, such as a full migration or a partial migration. Therefore, when a tenant needs to migrate the first cloud instance, the cloud management platform can determine whether the resources originally used by the first cloud instance need to be migrated. Only the resources that need to be migrated will be migrated, while the resources that do not need to be migrated will be associated with the second cloud instance. In this way, the data migration process can be simplified to a certain extent during cloud instance migration, thereby reducing migration time and improving the migration success rate.
[0017] In one possible implementation, the first resource pool includes any one of the following: a computing resource pool, a storage resource pool, and a network resource pool, wherein the computing resource pool includes multiple computing resources located in the same rack or adjacent racks, the storage resource pool includes multiple storage resources located in the same rack or adjacent racks, and the network resource pool includes multiple network resources located in the same rack or adjacent racks.
[0018] In one possible implementation, the first cloud instance includes any of the following: physical servers, virtual machines, containers, microvirtual machines, and bare metal servers.
[0019] A second aspect of this application provides a cloud management platform for managing infrastructure providing cloud services. The infrastructure includes multiple resource pools of different types, each resource pool containing multiple resources located in the same or adjacent racks, and each resource pool containing multiple resources of the same type. The cloud management platform includes: a first receiving module for receiving a cloud instance creation request from a tenant, the cloud instance creation request indicating a first type and a first specification of resources required by the tenant's first cloud instance; a first determining module for determining, based on the cloud instance creation request, a first resource pool satisfying the first type from the multiple resource pools, and a first resource satisfying the first specification from the first resource pool, the first resource being at least one of the multiple resources in the first resource pool; and a first allocation module for creating a first cloud instance and allocating the first resource to the first cloud instance.
[0020] In one possible implementation, the cloud instance creation request is further used to indicate the SLA threshold set by the tenant. The first determining module is used to: determine multiple first resource groups that meet the first specification from multiple resources in the first resource pool, each first resource group containing at least one resource; determine the first SLA metric of the multiple first resource groups, wherein the first SLA metric of any first resource group includes at least one of the following: the latency required for the first cloud instance to access the first resource group, and the bandwidth required for the first cloud instance to access the first resource group; and from the multiple first resource groups, the first resource group whose first SLA metric is less than the SLA threshold is selected as the first resource.
[0021] In one possible implementation, the cloud management platform further includes: a second receiving module for receiving a cloud instance expansion request from a tenant, the cloud instance expansion request indicating a first type and a second specification of resources required to be added to the tenant's first cloud instance; a second determining module for determining, based on the cloud instance expansion request, a first resource pool satisfying the first type from multiple resource pools, and a second resource satisfying the second specification from the first resource pool, the first resource and the second resource being different resources in the first resource pool; and a second allocation module for allocating the second resource to the first cloud instance.
[0022] In one possible implementation, the second determining module is configured to: determine a plurality of second resource groups that meet the second specification from the remaining resources in the first resource pool other than the first resource; determine the second SLA index of the plurality of second resource groups; and from the plurality of second resource groups, select the second resource group whose second SLA index is less than the SLA threshold as the second resource.
[0023] In one possible implementation, the cloud instance creation request is also used to indicate the second type and third specification of the resources required by the tenant's first cloud instance. The cloud management platform further includes: a third determination module, used to determine a second resource pool that meets the second type from multiple resource pools, and to determine a third resource that meets the third specification from the second resource pool, wherein the third resource is at least one of the multiple resources in the second resource pool; and a third allocation module, used to allocate the third resource to the first cloud instance.
[0024] In one possible implementation, the cloud management platform further includes: a third receiving module for receiving a cloud instance migration request from a tenant for a first cloud instance; a fourth determining module for determining a third SLA metric and a fourth SLA metric for the first resource based on the cloud instance migration request; and a fourth allocation module for: if the third SLA metric and the fourth SLA metric are less than the SLA threshold, creating a second cloud instance, allocating the first resource and the third resource to the second cloud instance, and releasing the first cloud instance; if the third SLA metric is greater than or equal to the SLA threshold and the fourth SLA metric is less than the SLA threshold, determining a fourth resource matching the first resource in the first resource pool, creating a second cloud instance, allocating the fourth resource and the third resource to the second cloud instance, and releasing the first cloud instance; if the third SLA metric and the fourth SLA metric are greater than or equal to the SLA threshold, determining a fourth resource matching the first resource in the first resource pool, determining a fifth resource matching the third resource in the second resource pool, creating a second cloud instance, allocating the fourth resource and the fifth resource to the second cloud instance, and releasing the first cloud instance.
[0025] In one possible implementation, the first resource pool includes any one of the following: a computing resource pool, a storage resource pool, and a network resource pool, wherein the computing resource pool includes multiple computing resources located in the same rack or adjacent racks, the storage resource pool includes multiple storage resources located in the same rack or adjacent racks, and the network resource pool includes multiple network resources located in the same rack or adjacent racks.
[0026] In one possible implementation, the first cloud instance includes any of the following: physical servers, virtual machines, containers, microvirtual machines, and bare metal servers.
[0027] A third aspect of this application provides a cloud service system. The cloud service system includes infrastructure for providing cloud services and a cloud management platform for managing the infrastructure. The infrastructure includes multiple resource pools of different types. Each resource pool includes multiple resources set in the same rack or adjacent racks. Each resource pool includes multiple resources of the same type. The cloud management platform is used to implement the steps performed by the cloud management platform in the method described in the first aspect or any possible implementation of the first aspect.
[0028] A fourth aspect of this application provides a computing device cluster, the computing device cluster including at least one computing device, each computing device including a processor and a memory: the memory is used to store instructions; the processor is used to cause the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect according to the instructions.
[0029] A fifth aspect of this application provides a computer storage medium storing one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method described in the first aspect or any possible implementation of the first aspect.
[0030] A sixth aspect of this application provides a computer program product storing instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect or any possible implementation of the first aspect.
[0031] In this embodiment, after receiving a cloud instance creation request from a tenant, the cloud management platform can determine the first type and first specification of resources required by the tenant's first cloud instance based on the request. Therefore, the cloud management platform can identify a first resource pool that meets the first type and select first resources that meet the first specification from the first resource pool. After creating the tenant's first cloud instance, the cloud management platform allocates the first resources to the tenant's first cloud instance for its use. Thus, the first resource pool selected by the cloud management platform contains multiple resources located in the same or adjacent racks, which is equivalent to these multiple resources of the same type being presented in a pooled manner. Since the tenant's first cloud instance and these multiple resources in the first resource pool are connected via a high-speed bus, when the tenant's first cloud instance needs one or more of these resources, the cloud management platform can schedule resources in the first resource pool according to the required resource specifications, allocating a sufficient number of first resources to the tenant's first cloud instance. Therefore, during the cloud instance creation process, resource utilization can be improved to a certain extent, and resource fragmentation can be avoided. Attached Figure Description
[0032] Figure 1 is a schematic diagram of a cloud service system provided in an embodiment of this application;
[0033] Figure 2 is another structural schematic diagram of the cloud service system provided in the embodiment of this application;
[0034] Figure 3 is another structural schematic diagram of the cloud service system provided in the embodiment of this application;
[0035] Figure 4 is a flowchart illustrating a resource scheduling method based on a cloud management platform provided in an embodiment of this application;
[0036] Figure 5 is another structural schematic diagram of the cloud service system provided in the embodiment of this application;
[0037] Figure 6 is another flowchart illustrating the resource scheduling method based on a cloud management platform provided in an embodiment of this application;
[0038] Figure 7 is another structural schematic diagram of the cloud service system provided in the embodiment of this application;
[0039] Figure 8 is another flowchart illustrating the resource scheduling method based on a cloud management platform provided in an embodiment of this application;
[0040] Figure 9 is another structural schematic diagram of the cloud service system provided in the embodiment of this application;
[0041] Figure 10 is another structural schematic diagram of the cloud service system provided in the embodiment of this application;
[0042] Figure 11 is a schematic diagram of the structure of a cloud management platform provided in an embodiment of this application;
[0043] Figure 12 is a schematic diagram of a computing device provided in an embodiment of this application;
[0044] Figure 13 is a schematic diagram of a computing device cluster provided in an embodiment of this application;
[0045] Figure 14 is a schematic diagram of computer devices in a computer cluster connected via a network according to an embodiment of this application. Detailed Implementation
[0046] This application provides a resource scheduling method based on a cloud management platform and a cloud management platform, which can improve resource utilization to a certain extent and avoid resource fragmentation.
[0047] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0048] With the rapid development of cloud technology, more and more tenants are choosing to purchase cloud instances provided by cloud vendors to complete their business operations. Generally, tenant cloud instances can be deployed on physical servers in cloud data centers. Cloud instances can use various types of resources on physical servers to complete tenant business operations and thus meet tenant business needs.
[0049] In related technologies, a cloud data center can contain multiple physical servers, each containing a fixed number of computing, storage, and network resources. Based on these resources, multiple cloud instances can be deployed on each physical server. Because the resources of different physical servers are relatively independent—meaning they typically cannot communicate at high speed via a bus—the resources of different physical servers usually cannot be combined for collaborative use.
[0050] In the aforementioned technologies, resource management and scheduling are highly limited. When a cloud instance requires a large amount of resources, there will often be insufficient resources on a certain physical server, while the resources of other physical servers remain idle, resulting in low resource utilization and high resource fragmentation in the cloud data center.
[0051] Furthermore, when migrating a cloud instance, all resources on a physical server used by the virtual machine need to be converted to resources on another physical server, which can lead to problems such as excessively long migration time and low migration success rate.
[0052] To address the aforementioned problems, this application provides a resource scheduling method based on a cloud management platform. This method can be implemented through a cloud service system (e.g., a public cloud system). Figure 1 is a schematic diagram of a cloud service system provided in this application embodiment. As shown in Figure 1, the cloud service system includes infrastructure that can provide cloud services and a cloud management platform that manages this infrastructure. The cloud management platform and the infrastructure are described in detail below:
[0053] A cloud management platform provides comprehensive management of the infrastructure within the entire cloud service system. (For example, within the infrastructure, it creates cloud instances for tenants according to their instructions, allocates various resources to these instances to enable them to run the applications specified by the tenants, and provides corresponding data services.) The cloud management platform can also be accessible to tenants outside the cloud service system and respond to their requests. For instance, it can provide various interfaces, such as login and management interfaces, for tenant clients (e.g., the terminal devices used by the tenants or the browsers on those devices) to access. Specifically, the cloud management platform can authenticate tenant clients through the login interface, allowing them to log in after successful authentication. For example, the cloud management platform can also allow tenant clients to send cloud instance creation requests to the platform via a management interface. Based on these requests, the platform can determine the type of resource pool selected by the tenant and select resources of a specific specification from these pools. After creating the tenant's cloud instance, these resources can be allocated to the instance to provide remote services. Similarly, the platform can also allow tenant clients to send cloud instance expansion requests to the platform via a management interface. Based on these expansion requests, the platform can determine the type of resource pool selected by the tenant and select resources of a specific specification from these pools. These resources can then be allocated to the tenant's cloud instance for additional use, thus expanding its capacity. For example, the cloud management platform can also allow tenant clients to send cloud instance migration requests to the cloud management platform through the management interface. Therefore, based on the cloud instance migration request, the cloud management platform can choose to migrate the resources used by the tenant's cloud instance to the new cloud instance, or partially or completely, depending on the actual situation, thereby flexibly and quickly completing the cloud instance migration.
[0054] Infrastructure comprises multiple pools of physical resources of different types serving tenants. These pools can include compute resource pools, storage resource pools, and network resource pools, among others. Compute resource pools can contain processing resource pools and control resource pools. Processing resource pools can contain multiple processing resources (also called processors, such as central processing units (CPUs)) located in the same or adjacent racks. Control resource pools can contain multiple control resources (also called controllers, such as service processing units (SPUs)) located in the same or adjacent racks. Storage resource pools can contain multiple storage resources (e.g., memory, hard drives) located in the same or adjacent racks. Network resources can contain multiple network resources (e.g., network interface cards (NICs)) located in the same or adjacent racks.
[0055] The following description, in conjunction with Figure 2, further illustrates the various resource types described above. As shown in Figure 2 (another structural schematic diagram of the cloud service system provided in this application embodiment), multiple resource pools of different types can be connected via a high-speed bus, and the cloud management platform can connect to multiple resource pools. Within these resource pools, the cloud management platform can create tenant cloud instances on processing resources. It can also create virtual control resources (e.g., control nodes) on control resources to enable the creation, expansion, and migration of cloud instances. Furthermore, the cloud management platform can create virtual storage resources (e.g., virtual memory, virtual hard disks, etc.) on storage resources and virtual network resources (e.g., virtual network cards, etc.) on network resources. When the cloud management platform performs cloud instance creation, expansion, and migration, it involves the allocation of virtual storage resources and virtual network resources, which is equivalent to the allocation of storage and network resources, but this will not be elaborated upon here.
[0056] Furthermore, the cloud management platform may include a separate resource management module, which may include a separate resource monitoring submodule and a separate resource scheduling submodule. As shown in Figure 3 (Figure 3 is another structural schematic diagram of the cloud service system provided in this application embodiment), the separate resource management module can be deployed remotely in the infrastructure, but it can also be regarded as part of the cloud management platform. When the cloud management platform receives cloud instance creation requests, cloud instance expansion requests, and cloud instance migration requests from tenants, it can send these requests to the separate resource management module, so that the separate resource monitoring submodule and the separate resource scheduling submodule in this module can call virtual control resources to perform a series of processes based on these requests, thereby completing the creation, expansion, and migration of cloud instances. This will not be elaborated here. It should be noted that the separate resource management module may also not be deployed remotely, but integrated into the cloud management platform; this application embodiment does not impose any restrictions.
[0057] Furthermore, for tenant cloud instances, these cloud instances can be presented in various ways. For example, these cloud instances can be physical servers (i.e., the processors of physical servers) selected by the cloud management platform; they can be bare metal servers (i.e., the processors of bare metal servers) selected by the cloud management platform; they can be virtual machines (VMs) created on physical servers by the cloud management platform using virtualization technology; they can also be containers created on physical servers by the cloud management platform using virtualization technology; they can also be micro VMs created on physical servers by the cloud management platform using virtualization technology, and so on.
[0058] Furthermore, for multiple resource pools of different types, these resource pools can be deployed in the same site or different sites. The site can be presented in various forms. For example, the site can be a region in the infrastructure, an availability zone in the infrastructure, a data center (DC) in the infrastructure, a room in the infrastructure, and so on.
[0059] Based on the aforementioned cloud service system, after receiving a cloud instance creation request from a tenant, the cloud management platform can determine the type and specifications of the resources required by the tenant's cloud instance. Therefore, the cloud management platform can identify resource pools that meet the required type and select at least one resource from these pools that meets the specified specifications. After creating the tenant's cloud instance, this at least one resource is then allocated to the tenant's cloud instance for its use. Thus, the resource pool selected by the cloud management platform contains multiple resources located in the same or adjacent racks, effectively presenting these resources of the same type in a pooled manner. Since the tenant's cloud instance and these resources in the resource pool communicate via a high-speed bus, when the tenant's cloud instance needs one or more of these resources, the cloud management platform can schedule resources within the resource pool according to the required specifications, allocating a sufficient number of resources to the tenant's cloud instance. This can improve resource utilization to a certain extent and avoid resource fragmentation. To further understand the aforementioned process, the following description refers to Figure 4. Figure 4 is a flowchart illustrating a resource scheduling method based on a cloud management platform provided in an embodiment of this application. As shown in Figure 4, this method can be implemented through a cloud service system as shown in Figure 1. The cloud service system includes infrastructure that can provide cloud services to tenants and a cloud management platform that manages this infrastructure. This infrastructure may include multiple resource pools, each resource pool containing multiple resources located in the same rack or adjacent racks, and each resource pool containing multiple resources of the same type. The method includes:
[0060] 401. The cloud management platform receives a cloud instance creation request from a tenant. The cloud instance creation request is used to indicate the first type and first specification of resources required by the tenant's first cloud instance.
[0061] In this embodiment, when a tenant needs to create a cloud instance, the cloud management platform can provide a management interface to the tenant's client (e.g., the cloud instance management section of the tenant's interface). The tenant can then input a cloud instance creation request specified by the tenant into the management interface through their client. This request indicates the first type of resources required for the first cloud instance the tenant wants to create, as well as the first specifications of those resources. In this way, the cloud management platform can receive the cloud instance creation request sent by the tenant through the client via the management interface.
[0062] Specifically, a cloud instance creation request may include the following information: (1) the first type of resources required by the first cloud instance to be created by the tenant (e.g., storage type resources). (2) the first specification of the resources required by the first cloud instance to be created by the tenant. (3) the second type of resources required by the first cloud instance to be created by the tenant (e.g., network type resources). It should be noted that the first type and the second type are different types. (4) the third specification of the resources required by the first cloud instance to be created by the tenant. It should be noted that the first specification and the third specification are the same or different specifications. (5) the service-level agreement (SLA) thresholds set by the tenant, such as the latency threshold and bandwidth threshold set by the tenant.
[0063] For example, as shown in Figure 5 (Figure 5 is another structural schematic diagram of the cloud service system provided in the embodiment of this application), when a tenant needs to create virtual machine 1, the tenant can send a creation request for virtual machine 1 to the cloud management platform. The creation request can be used to indicate the types of resources required by virtual machine 1, including memory and network card. The memory specification is 16G and the network card specification is 500Mbps. The creation request can also be used for the SLA thresholds set by the tenant, including the latency threshold and bandwidth threshold set by the tenant.
[0064] 402. Based on the cloud instance creation request, the cloud management platform determines a first resource pool that meets the first type from multiple resource pools, and determines a first resource that meets the first specification from the first resource pool. The first resource is at least one of the multiple resources in the first resource pool.
[0065] Upon receiving a cloud instance creation request, the cloud management platform can determine, based on the request, the first cloud instance that the tenant needs to create, as well as the first type and first specification of the resources required for the first cloud instance. Therefore, the cloud management platform can select a first resource pool from multiple resource pools that meets the first type, and from the first resource pool, select a first resource that meets the first specification. It should be noted that since the first resource pool contains multiple resources, the first resource selected by the cloud management platform is usually one or more of these resources.
[0066] Specifically, since the cloud management platform can also determine the second type and third specification of resources required by the tenant to create the first cloud instance based on the cloud instance creation request, the cloud management platform determines a second resource pool that meets the second type from multiple resource pools, and then determines a third resource that meets the third specification from the second resource pool. It should be noted that since the second resource pool contains multiple resources, the third resource selected by the cloud management platform is one or more of the multiple resources in the second resource pool.
[0067] More specifically, the cloud management platform can identify the primary and tertiary resources in the following ways:
[0068] (1) After determining the first resource pool, the cloud management platform can determine multiple first resource groups that meet the first specification from the multiple resources in the first resource pool. Each first resource group contains at least one resource in the first resource pool. It should be noted that for any one of the multiple first resource groups, the first specification means that the total specification of all resources in the first resource group is equal to the first specification.
[0069] (2) After identifying multiple first resource groups, the cloud management platform can evaluate the multiple first resource groups to obtain the first SLA indicators of the multiple first resource groups. It should be noted that for any one of the multiple first resource groups, the first SLA indicator of the first resource group includes one or more of the following information: the latency required for the first cloud instance to access the first resource group and the bandwidth required for the first cloud instance to access the first resource group.
[0070] (3) After obtaining the first SLA index of multiple first resource groups, the cloud management platform can select a first resource group whose first SLA index is less than the SLA threshold from the multiple first resource groups as the first resource available for the first cloud instance.
[0071] (4) Similarly, after determining the second resource pool, the cloud management platform can also determine the third resource from the second resource pool by following operations similar to (1) to (3), which will not be elaborated here.
[0072] Continuing with the example above, after receiving the creation request, the cloud management platform can send it to the split resource management module. The split resource monitoring submodule within this module can then determine the memory pool (storage resource pool) and the network interface card (NIC) pool (network resource pool) from multiple resource pools based on this creation request. Since the memory pool contains multiple virtual memory pools such as Virtual Memory 1, Virtual Memory 2, Virtual Memory 3, and Virtual Memory 4, the split resource monitoring submodule can determine multiple virtual memory groups with a total size of 16GB from these virtual memory pools. Virtual Memory Group 1: Virtual Memory 1 + Virtual Memory 2; Virtual Memory Group 2: Virtual Memory 3 + Virtual Memory 4; Virtual Memory Group 3: Virtual Memory 5 + Virtual Memory 6 + Virtual Memory 7, etc. Next, the separate resource monitoring submodule determines the SLA metrics for each virtual memory group (including latency and bandwidth of virtual machine 1 using each virtual memory group, etc.), and then compares the SLA metrics of each virtual memory group with the SLA threshold. Since the SLA metrics of virtual memory group 1 are less than the SLA threshold, the separate resource monitoring submodule can determine that virtual memory group 1 is a virtual memory that cloud instance 1 can use. Similarly, the separate resource monitoring submodule can also determine virtual network interface card 1 from the network interface card pool in a similar way as a virtual network interface card that cloud instance 1 can use.
[0073] 403. The cloud management platform creates the first cloud instance and allocates the first resource to the first cloud instance.
[0074] After determining the primary resource, the cloud management platform can create the primary cloud instance and allocate the primary resource to the primary cloud instance for its use, thereby meeting the tenant's cloud instance creation needs.
[0075] Specifically, since the cloud management platform also selected a third resource for the first cloud instance, the cloud management platform can allocate the third resource to the first cloud instance after it is created, so that the first cloud instance can use it.
[0076] More specifically, the cloud management platform can allocate first and third resources to the first cloud instance in the following ways:
[0077] After determining the primary resource, the cloud management platform can identify a third resource pool from multiple resource pools. This third resource pool is the control resource pool, and the platform selects one resource from among the multiple resources in the third resource pool as the primary resource (i.e., a control resource). Next, the cloud management platform sends the address of the primary resource to the third resource, enabling the third resource to create the primary cloud instance. Based on the address of the primary resource, a communication connection is established between the primary cloud instance and the primary resource, thus successfully allocating the primary resource to the primary cloud instance.
[0078] Similarly, after the third resource is determined, the cloud management platform can allocate the third resource to the first cloud instance in the same way as described above, which will not be repeated here.
[0079] As in the example above, the separate resource monitoring submodule can provide the addresses of virtual memory group 1 (including virtual memory 1 and virtual memory 2) and virtual network interface card 1 to the separate resource scheduling submodule. The separate resource scheduling submodule can select control node 1 from multiple control nodes in the controller pool to serve virtual machine 1, and provide the addresses of virtual memory group 1 and virtual network interface card 1 to control node 1. This allows control node 1 to notify the hypervisor 1 to start virtual machine 1 and provide the addresses of virtual memory group 1 and virtual network interface card 1 to the hypervisor 1, so that the hypervisor 1 can allocate virtual memory group 1 and virtual network interface card 1 to virtual machine 1 for use based on the addresses of virtual memory group 1 and virtual network interface card 1.
[0080] Figure 6 is another flowchart illustrating the resource scheduling method based on a cloud management platform provided in this application embodiment. As shown in Figure 6, this method can be implemented through the cloud service system shown in Figure 1. The cloud service system includes infrastructure that can provide cloud services to tenants and a cloud management platform that manages this infrastructure. This infrastructure may include multiple resource pools, each resource pool containing multiple resources located in the same rack or adjacent racks, and each resource pool containing multiple resources of the same type. The method includes:
[0081] 601. The cloud management platform receives a cloud instance expansion request from a tenant. The cloud instance expansion request is used to indicate the first type and second specification of the resources required to be added to the tenant's first cloud instance.
[0082] In this embodiment, after the first cloud instance is created, it can use the first resources to complete the tenant's business. During this process, if the tenant finds that the first cloud instance's resources are insufficient, it can expand the first cloud instance. The cloud management platform can then provide a management interface to the tenant's client (e.g., the cloud instance management section of the tenant's interface). Next, the tenant can input a cloud instance expansion request through its client into the management interface. This request indicates the first type of resources required to be added to the tenant's first cloud instance, and the second specification of the resources needed. In this way, the cloud management platform can receive the cloud instance expansion request sent by the tenant through its client via the management interface.
[0083] For example, as shown in Figure 7 (Figure 7 is another structural schematic diagram of the cloud service system provided in the embodiment of this application), when a tenant needs to expand the capacity of virtual machine 1, the tenant can send an expansion request for virtual machine 1 to the cloud management platform. The expansion request can be used to indicate the type of resources that virtual machine 1 needs to add, including memory, and the required memory specification is 4G.
[0084] 602. Based on the cloud instance expansion request, the cloud management platform determines a first resource pool that meets the first type from multiple resource pools, and determines a second resource that meets the second specification from the first resource pool. The first resource and the second resource are different resources in the first resource pool.
[0085] Upon receiving a cloud instance expansion request, the cloud management platform can determine, based on the request, the tenant's need to expand the first cloud instance, and the first type and second specification of the resources required to be added to the first cloud instance. Therefore, the cloud management platform can select a first resource pool from multiple resource pools that meets the first type, and select a second resource from the first resource pool that meets the second specification. It should be noted that the second resource and the first resource are typically different resources among the multiple resources in the first resource pool.
[0086] More specifically, the cloud management platform can identify the second resource in the following ways:
[0087] (1) After determining the first resource pool, the cloud management platform can determine multiple second resource groups that meet the second specification from the remaining resources in the first resource pool other than the first resource. Each second resource group contains at least one resource from the first resource pool. It should be noted that for any one of the multiple second resource groups, the second specification means that the total specification of all resources in the second resource group is equal to the second specification.
[0088] (2) After identifying multiple second resource groups, the cloud management platform can evaluate the multiple second resource groups to obtain the second SLA indicators for the multiple second resource groups. It should be noted that for any one of the multiple second resource groups, the second SLA indicator of the second resource group includes one or more of the following information: the latency required for the first cloud instance to access the second resource group and the bandwidth required for the first cloud instance to access the second resource group.
[0089] (3) After obtaining the second SLA index of multiple second resource groups, the cloud management platform can select a second resource group whose second SLA index is less than the SLA threshold from the multiple second resource groups as a second resource available for use by the first cloud instance.
[0090] Continuing with the example above, after receiving the expansion request, the cloud management platform can send it to the split resource management module. The split resource monitoring submodule within this module can then determine the memory pool (storage resource pool) from multiple resource pools based on this expansion request. Since virtual memory 1 and virtual memory 2 have already been allocated to virtual machine 1, the split resource monitoring submodule can determine multiple virtual memory groups with a total specification of 4GB from virtual memory 3, virtual memory 4, and other virtual memory pools. Examples include virtual memory group 1: virtual memory 3, virtual memory group 2: virtual memory 5, virtual memory group 3: virtual memory 8, etc. Next, the split resource monitoring submodule can determine the SLA metrics for each virtual memory group (including latency and bandwidth for virtual machine 1 using each virtual memory group), and then compare the SLA metrics of each virtual memory group with the SLA threshold. Since the SLA metrics of virtual memory group 1 are less than the SLA threshold, the split resource monitoring submodule can determine virtual memory group 1 as the additional virtual memory that cloud instance 1 can use.
[0091] 603. The cloud management platform allocates the second resource to the first cloud instance.
[0092] After the second resource is determined, the cloud management platform can allocate the second resource to the first cloud instance for the first cloud instance to use, thereby meeting the tenant's cloud instance expansion needs.
[0093] Specifically, the cloud management platform can allocate first and third resources to the first cloud instance in the following ways:
[0094] After determining the second resource, the cloud management platform can identify a third resource pool from multiple resource pools. This third resource pool is the control resource pool, and the platform selects one resource from among the multiple resources in the third resource pool as the third resource (i.e., a control resource). Next, the cloud management platform can send the address of the second resource to the third resource, enabling the third resource to establish a communication connection between the first cloud instance and the second resource based on that address, thus successfully allocating the second resource to the first cloud instance.
[0095] Continuing with the example above, the separate resource monitoring submodule can provide the address of virtual memory group 1 (including virtual memory 3) to the separate resource scheduling submodule. The separate resource scheduling submodule can then select control node 1 from among the multiple control nodes in the controller pool to serve virtual machine 1, and provide the address of virtual memory group 1 to control node 1. This allows control node 1 to notify the virtual machine supervisor 1 to allocate virtual memory group 1 to virtual machine 1 based on its address. At this point, virtual machine 1 possesses virtual memory resources such as virtual memory 1, virtual memory 2, and virtual memory 3, successfully achieving expansion.
[0096] Figure 8 is another flowchart illustrating the resource scheduling method based on a cloud management platform provided in this application embodiment. As shown in Figure 8, the method can be implemented through the cloud service system shown in Figure 1. The cloud service system includes infrastructure that can provide cloud services to tenants and a cloud management platform that manages this infrastructure. This infrastructure may include multiple resource pools, each resource pool containing multiple resources located in the same rack or adjacent racks, and each resource pool containing multiple resources of the same type. The method includes:
[0097] 801. The cloud management platform receives a cloud instance migration request from a tenant for the first cloud instance.
[0098] In this embodiment, when a tenant determines that the first cloud instance needs to be migrated during its use (e.g., due to a failed expansion of the first cloud instance), the cloud management platform can provide a management interface to the tenant's client (e.g., the cloud instance management section of the tenant's interface). Then, the tenant can input a cloud instance migration request for the first cloud instance into the management interface through their client. Therefore, the cloud management platform can receive the cloud instance migration request sent by the tenant through the client via the management interface.
[0099] For example, as shown in Figure 9 (Figure 9 is another structural schematic diagram of the cloud service system provided in the embodiment of this application), when a tenant needs to migrate virtual machine 1, the tenant can send a migration request for virtual machine 1 to the cloud management platform.
[0100] 802. Based on the cloud instance migration request, the cloud management platform determines the third SLA metric of the first resource and the fourth SLA metric of the third resource.
[0101] Upon receiving a cloud instance migration request, the cloud management platform can determine that the tenant needs to migrate the first cloud instance. Therefore, the cloud management platform can evaluate the first and third resources used by the first cloud instance to obtain the third SLA metric and the fourth SLA metric for the first resource. The third SLA metric includes one or more of the following: the latency required for the second cloud instance (i.e., the cloud instance to be created corresponding to the first cloud instance) to access the first resource; and the bandwidth required for the second cloud instance to access the first resource. The fourth SLA metric includes one or more of the following: the latency required for the second cloud instance to access the third resource; and the bandwidth required for the second instance to access the third resource.
[0102] 803. If the third and fourth SLA metrics are less than the SLA threshold, the cloud management platform creates a second cloud instance, allocates the first and third resources to the second cloud instance, and releases the first cloud instance.
[0103] 804. If the third SLA indicator is greater than or equal to the SLA threshold and the fourth SLA indicator is less than the SLA threshold, the cloud management platform determines the fourth resource that matches the first resource in the first resource pool, creates a second cloud instance, allocates the fourth resource and the third resource to the second cloud instance, and releases the first cloud instance.
[0104] 805. If the third SLA indicator is greater than or equal to the SLA threshold and the fourth SLA indicator is greater than or equal to the SLA threshold, the cloud management platform determines the fourth resource that matches the first resource in the first resource pool, determines the fifth resource that matches the third resource in the second resource pool, creates a second cloud instance, allocates the fourth and fifth resources to the second cloud instance, and releases the first cloud instance.
[0105] After obtaining the third and fourth SLA metrics, the cloud management platform can determine the relationship between these metrics and the SLA thresholds set by the tenant. Based on this relationship, the cloud management platform can perform different migration operations for the first cloud instance.
[0106] (1) If both the third and fourth SLA metrics are less than the SLA threshold, it means that in this migration scenario, the first and third resources originally used by the first cloud instance can still be used by the second cloud instance. Therefore, the cloud management platform creates the second cloud instance, allocates the first and third resources to the second cloud instance, and releases the first cloud instance, i.e., shuts it down. In this way, the cloud management platform successfully migrates the first cloud instance, thereby meeting the tenant's cloud instance migration needs.
[0107] (2) If the third SLA metric is greater than or equal to the SLA threshold and the fourth SLA metric is less than the SLA threshold, it means that in this migration scenario, the third resource originally used by the first cloud instance can still be used by the second cloud instance. However, if the second cloud instance uses the first resource originally used by the first cloud instance, it will cause fluctuations and instability in the SLA metric. Therefore, the second cloud instance cannot use the first resource. Thus, the cloud management platform can determine a fourth resource matching the first resource from the first resource pool (e.g., the fourth resource's specifications are greater than the first resource's specifications), create the second cloud instance, allocate the fourth and third resources to the second cloud instance, and release the first cloud instance, i.e., shut it down. In this way, the cloud management platform successfully migrates the first cloud instance, thereby meeting the tenant's cloud instance migration needs.
[0108] (3) If both the third and fourth SLA metrics are greater than or equal to the SLA threshold, it indicates that in this migration scenario, the use of the first and third resources originally used by the first cloud instance by the second cloud instance would cause fluctuations and instability in the SLA metrics. Therefore, the second cloud instance cannot use the first and third resources. Thus, the cloud management platform determines a fourth resource matching the first resource in the first resource pool and a fifth resource matching the third resource in the second resource pool (e.g., the fifth resource's specifications are greater than the third resource's third specifications). Then, it creates the second cloud instance, allocates the fourth and fifth resources to the second cloud instance, and releases the first cloud instance, effectively shutting it down. In this way, the cloud management platform successfully migrates the first cloud instance, thus meeting the tenant's cloud instance migration requirements.
[0109] It is worth noting that in this embodiment, the method for migrating cloud instances is usually hot migration.
[0110] It should be noted that the cloud management platform's operations for determining the fourth resource in the first resource pool and the fifth resource in the second resource pool can be found in the aforementioned explanation of the cloud management platform's operations for determining the first resource in the first resource pool, and will not be repeated here.
[0111] It should also be noted that the cloud management platform will create a second cloud instance and allocate the first, third, fourth, and fifth resources to the second cloud instance. The operation can be referred to the operation of the cloud management platform creating a first cloud instance and allocating the first and third resources to the first cloud instance, which will not be repeated here.
[0112] Continuing with the example above, upon receiving the migration request, the cloud management platform can send it to the split resource management module. The split resource monitoring submodule within this module can then instruct the split resource scheduling submodule to determine the SLA metrics for virtual machine 2 (the migration target of virtual machine 1) accessing virtual memory 1, virtual memory 2, and virtual network interface 1 based on the migration request. These SLA metrics are then compared to SLA thresholds. Since the SLA metrics for virtual machine 3 accessing virtual memory 1 and virtual memory 2 are greater than the SLA thresholds, while the SLA metrics for virtual machine 3 accessing virtual network interface 1 are less than the SLA thresholds, it indicates that migrating virtual machine 1 only involves migrating its virtual memory while retaining its virtual network interface. This will not cause significant fluctuations in SLA metrics such as latency and bandwidth. Therefore, the split resource scheduling submodule can select virtual memory 4 with a specification larger than virtual memory 1 and virtual memory 2 from the memory pool. The split resource scheduling submodule can then select control node 1 from among the multiple control nodes in the controller pool to serve virtual machine 1 and instruct control node 1 to initiate the migration. Control node 1 can provide the addresses of virtual memory 1, virtual memory 2, virtual memory 4, and virtual network interface card 1 to hypervisor 1 for migration. Specifically, hypervisor 1 migrates data from virtual memory 1 and virtual memory 2 to virtual memory 4, while keeping virtual network interface card 1 unchanged. Control node 1 can then use hypervisor 2 to launch virtual machine 3, allocate virtual memory 4 and virtual network interface card 1 to virtual machine 3 based on their addresses, and then shut down virtual machine 1. Thus, the migration is successfully completed.
[0113] It is worth noting that, as shown in Figure 10 (Figure 10 is another structural schematic diagram of the cloud service system provided in this application embodiment), when the cloud management platform needs to create a new control node to realize its management and control business, the cloud management platform can send a control node creation request to the separate resource management module. Based on the control node creation request, the separate resource monitoring submodule will query the controller pool. When the remaining resources in the controller pool are insufficient, but the virtual machine pool (containing multiple processors for deploying virtual machine supervisors and virtual machines) has idle resources available, the separate resource monitoring submodule can find a control node in the controller pool through the separate resource scheduling submodule, for example, control node 2. Then, the separate resource monitoring submodule sends a control node creation request to control node 2, so that control node 2 can start virtual machine 4 through virtual machine supervisor 2 and make virtual machine 4 run the relevant business of the control node for use as a control node.
[0114] In this embodiment, after receiving a cloud instance creation request from a tenant, the cloud management platform can determine the first type and first specification of resources required by the tenant's first cloud instance based on the request. Therefore, the cloud management platform can identify a first resource pool that meets the first type and select first resources that meet the first specification from the first resource pool. After creating the tenant's first cloud instance, the cloud management platform allocates the first resources to the tenant's first cloud instance for its use. Thus, the first resource pool selected by the cloud management platform contains multiple resources located in the same or adjacent racks, which is equivalent to these multiple resources of the same type being presented in a pooled manner. Since the tenant's first cloud instance and these multiple resources in the first resource pool are connected via a high-speed bus, when the tenant's first cloud instance needs one or more of these resources, the cloud management platform can schedule resources in the first resource pool according to the required resource specifications, allocating a sufficient number of first resources to the tenant's first cloud instance. Therefore, during the cloud instance creation process, resource utilization can be improved to a certain extent, and resource fragmentation can be avoided.
[0115] Furthermore, in this embodiment, when a tenant needs to expand the capacity of the first cloud instance, the cloud management platform selects a first resource pool for it. Since the tenant's first cloud instance and the other resources in the first resource pool (excluding the first resource) are connected via a high-speed bus, when the tenant's first cloud instance needs one or more of the other resources, the cloud management platform can schedule resources in the first resource pool according to the specifications of the required resources, and allocate a sufficient number of second resources to the tenant's first cloud instance. Therefore, during the cloud instance expansion process, resource utilization can be improved to a certain extent, and resource fragmentation can be avoided.
[0116] Furthermore, in this embodiment, when a tenant needs to migrate the first cloud instance, the cloud management platform can determine whether the resources originally used by the first cloud instance (i.e., the first resource and the third resource) need to be migrated. Only the resources that need to be migrated (e.g., the first resource) are migrated, and the resources that do not need to be migrated (e.g., the third resource) are associated with the second cloud instance. In this way, the data migration process can be simplified to a certain extent during the cloud instance migration process, thereby reducing migration time and improving the migration success rate.
[0117] The above describes the resource scheduling method based on a cloud management platform provided in this application embodiment. The cloud management platform will be described below. Figure 11 is a structural diagram of the cloud management platform provided in this application embodiment. As shown in Figure 11, the cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes multiple resource pools of different types. Each resource pool contains multiple resources located in the same rack or adjacent racks. Each resource pool contains multiple resources of the same type. The cloud management platform includes:
[0118] The first receiving module 1101 is used to receive a cloud instance creation request from a tenant. The cloud instance creation request is used to indicate the first type and first specification of the resources required by the tenant's first cloud instance. For example, the first receiving module 1101 can be used to implement step 401 in the embodiment shown in FIG4.
[0119] The first determining module 1102 is used to determine a first resource pool that meets a first type from multiple resource pools based on a cloud instance creation request, and to determine a first resource that meets a first specification from the first resource pool. The first resource is at least one of the multiple resources in the first resource pool. For example, the first determining module 1102 can be used to implement step 402 in the embodiment shown in FIG4.
[0120] The first allocation module 1103 is used to create a first cloud instance and allocate first resources to the first cloud instance. For example, the first allocation module 1103 can be used to implement step 403 in the embodiment shown in FIG4.
[0121] In one possible implementation, the cloud instance creation request is further used to indicate the SLA threshold set by the tenant. The first determining module is used to: determine multiple first resource groups that meet the first specification from multiple resources in the first resource pool, each first resource group containing at least one resource; determine the first SLA metric of the multiple first resource groups, wherein the first SLA metric of any first resource group includes at least one of the following: the latency required for the first cloud instance to access the first resource group, and the bandwidth required for the first cloud instance to access the first resource group; and from the multiple first resource groups, the first resource group whose first SLA metric is less than the SLA threshold is selected as the first resource.
[0122] In one possible implementation, the cloud management platform further includes: a second receiving module for receiving a cloud instance expansion request from a tenant, the cloud instance expansion request indicating a first type and a second specification of resources required to be added to the tenant's first cloud instance; a second determining module for determining, based on the cloud instance expansion request, a first resource pool satisfying the first type from multiple resource pools, and a second resource satisfying the second specification from the first resource pool, the first resource and the second resource being different resources in the first resource pool; and a second allocation module for allocating the second resource to the first cloud instance.
[0123] In one possible implementation, the second determining module is configured to: determine a plurality of second resource groups that meet the second specification from the remaining resources in the first resource pool other than the first resource; determine the second SLA index of the plurality of second resource groups; and from the plurality of second resource groups, select the second resource group whose second SLA index is less than the SLA threshold as the second resource.
[0124] In one possible implementation, the cloud instance creation request is also used to indicate the second type and third specification of the resources required by the tenant's first cloud instance. The cloud management platform further includes: a third determination module, used to determine a second resource pool that meets the second type from multiple resource pools, and to determine a third resource that meets the third specification from the second resource pool, wherein the third resource is at least one of the multiple resources in the second resource pool; and a third allocation module, used to allocate the third resource to the first cloud instance.
[0125] In one possible implementation, the cloud management platform further includes: a third receiving module for receiving a cloud instance migration request from a tenant for a first cloud instance; a fourth determining module for determining a third SLA metric and a fourth SLA metric for the first resource based on the cloud instance migration request; and a fourth allocation module for: if the third SLA metric and the fourth SLA metric are less than the SLA threshold, creating a second cloud instance, allocating the first resource and the third resource to the second cloud instance, and releasing the first cloud instance; if the third SLA metric is greater than or equal to the SLA threshold and the fourth SLA metric is less than the SLA threshold, determining a fourth resource matching the first resource in the first resource pool, creating a second cloud instance, allocating the fourth resource and the third resource to the second cloud instance, and releasing the first cloud instance; if the third SLA metric and the fourth SLA metric are greater than or equal to the SLA threshold, determining a fourth resource matching the first resource in the first resource pool, determining a fifth resource matching the third resource in the second resource pool, creating a second cloud instance, allocating the fourth resource and the fifth resource to the second cloud instance, and releasing the first cloud instance.
[0126] In one possible implementation, the first resource pool includes any one of the following: a computing resource pool, a storage resource pool, and a network resource pool, wherein the computing resource pool includes multiple computing resources located in the same rack or adjacent racks, the storage resource pool includes multiple storage resources located in the same rack or adjacent racks, and the network resource pool includes multiple network resources located in the same rack or adjacent racks.
[0127] In one possible implementation, the first cloud instance includes any of the following: physical servers, virtual machines, containers, microvirtual machines, and bare metal servers.
[0128] It should be noted that the information interaction and implementation process between the modules / units of the above-mentioned device are based on the same concept as the method embodiments of this application, and the resulting technical effects are the same as those of the method embodiments of this application. For details, please refer to the description in the method embodiments shown above in the embodiments of this application, and will not be repeated here.
[0129] Please refer to Figure 12, which is a schematic diagram of a computing device provided in an embodiment of this application. As shown in Figure 12, the computing device 1200 (which can be used to present the aforementioned cloud management platform) includes: a processor 1201, a memory 1202, a communication interface 1203, and a bus 1204. The processor 1201, the memory 1202, and the communication interface 1203 are coupled through the bus (not labeled in the figure). The memory 1202 stores instructions. When the execution instructions in the memory 1202 are executed, the computing device 1200 executes the method performed by the cloud management platform in the above-described method embodiment.
[0130] The computing device 1200 may be one or more integrated circuits configured to implement the methods described above, such as: one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these forms of integrated circuits. Furthermore, when the units in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these units may be integrated together and implemented as a system-on-a-chip (SOC).
[0131] Processor 1201 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.
[0132] The memory 1202 can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0133] The memory 1202 stores executable program code, and the processor 1201 executes the executable program code to implement the functions of the aforementioned first receiving module, first determining module, and first allocation module, thereby realizing the resource scheduling method based on the cloud management platform. That is, the memory 1202 stores instructions for executing the resource scheduling method based on the cloud management platform.
[0134] The communication interface 1203 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1200 and other devices or communication networks.
[0135] In addition to the data bus, the 1204 bus can also include a power bus, a control bus, and a status signal bus. The bus can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. The bus can be divided into address bus, data bus, and control bus.
[0136] Please refer to Figure 13, which is a schematic diagram of a computing device cluster provided in an embodiment of this application. As shown in Figure 13, the computing device cluster 1300 includes at least one computing device 1200.
[0137] As shown in Figure 13, the computing device cluster 1300 includes at least one computing device 1200. The memory 1202 of one or more computing devices 1200 in the computing device cluster 1300 may store the same instructions for executing the resource scheduling method based on the cloud management platform described above.
[0138] In some possible implementations, the memory 1202 of one or more computing devices 1200 in the computing device cluster 1300 may also store partial instructions for executing the resource scheduling method based on the cloud management platform described above. In other words, a combination of one or more computing devices 1200 can jointly execute the resource scheduling method based on the cloud management platform described above.
[0139] It should be noted that the memory 1202 in different computing devices 1200 within the computing device cluster 1300 can store different instructions, each used to execute a portion of the functions of the aforementioned cloud management platform. That is, the memory 1202 in different computing devices 1200 stores the functions of one or more modules, such as the first receiving module, the first determining module, and the first allocating module.
[0140] In some possible implementations, one or more computing devices 1200 in the computing device cluster 1300 can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.
[0141] Please refer to Figure 14, which is a schematic diagram of computer devices in a computer cluster 1400 provided in this embodiment of the application being connected via a network. As shown in Figure 14, two computing devices 1200A and 1200B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0142] In one possible implementation, the memory in computing device 1200A stores instructions for performing the functions of modules such as the first receiving module. Meanwhile, the memory in computing device 1200B stores instructions for performing the functions of modules such as the first determining module and the first allocating module.
[0143] It should be understood that the functions of computing device 1200A shown in Figure 14 can also be performed by multiple computing devices. Similarly, the functions of computing device 1200B can also be performed by multiple computing devices.
[0144] This application also relates to a computer storage medium storing a program for signal processing, which, when run on a computer, causes the computer to perform the steps performed by the cloud management platform in the embodiments shown in FIG4, FIG6 or FIG8.
[0145] This application also relates to a computer program product that stores instructions that, when executed by a computer, cause the computer to perform the steps performed by the cloud management platform in the embodiments shown in FIG4, FIG6 or FIG8.
[0146] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0147] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0149] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0150] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A resource scheduling method based on a cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes multiple resource pools of different types. Each resource pool contains multiple resources located in the same or adjacent racks. Each resource pool contains multiple resources of the same type. The method includes: The cloud management platform receives a cloud instance creation request from a tenant, the cloud instance creation request being used to indicate the first type and first specification of resources required by the tenant's first cloud instance; Based on the cloud instance creation request, the cloud management platform determines a first resource pool that meets the first type from the plurality of resource pools, and determines a first resource that meets the first specification from the first resource pool, wherein the first resource is at least one of the plurality of resources in the first resource pool; The cloud management platform creates the first cloud instance and allocates the first resource to the first cloud instance.
2. The method according to claim 1, characterized in that, The cloud instance creation request is also used to indicate the Service Level Agreement (SLA) threshold set by the tenant, and the cloud management platform determines from the first resource pool that the first resources meeting the first specification include: The cloud management platform determines multiple first resource groups that meet the first specification from multiple resources in the first resource pool, and each first resource group contains at least one resource; The cloud management platform determines the first SLA indicators of the plurality of first resource groups. The first SLA indicator of any first resource group includes at least one of the following: the latency required for the first cloud instance to access the first resource group, and the bandwidth required for the first cloud instance to access the first resource group. The cloud management platform selects the resource group whose first SLA index is less than the SLA threshold from the plurality of first resource groups as the first resource.
3. The method according to claim 1 or 2, characterized in that, The method further includes: The cloud management platform receives a cloud instance expansion request from the tenant, the cloud instance expansion request being used to indicate the first type and second specification of the resources required to be added to the tenant's first cloud instance; Based on the cloud instance expansion request, the cloud management platform determines the first resource pool that meets the first type from the multiple resource pools, and determines the second resource that meets the second specification from the first resource pool. The first resource and the second resource are different resources in the first resource pool. The cloud management platform allocates the second resource to the first cloud instance.
4. The method according to claim 3, characterized in that, The cloud management platform determines from the first resource pool the second resources that meet the second specification, including: The cloud management platform determines multiple second resource groups that meet the second specification from the remaining resources in the first resource pool, excluding the first resource. The cloud management platform determines the second SLA metrics for the plurality of second resource groups; The cloud management platform selects the second resource group whose second SLA index is less than the SLA threshold from the plurality of second resource groups as the second resource.
5. The method according to any one of claims 1 to 4, characterized in that, The cloud instance creation request is also used to indicate the second type and third specification of the resources required by the tenant's first cloud instance, and the method further includes: The cloud management platform determines a second resource pool that meets the second type from the plurality of resource pools, and determines a third resource that meets the third specification from the second resource pool, wherein the third resource is at least one of the plurality of resources in the second resource pool; The cloud management platform allocates the third resource to the first cloud instance.
6. The method according to claim 5, characterized in that, The method further includes: The cloud management platform receives a cloud instance migration request from the tenant for the first cloud instance; Based on the cloud instance migration request, the cloud management platform determines the third SLA metric of the first resource and the fourth SLA metric of the third resource. If the third SLA metric and the fourth SLA metric are less than the SLA threshold, the cloud management platform creates a second cloud instance, allocates the first resource and the third resource to the second cloud instance, and releases the first cloud instance. If the third SLA metric is greater than or equal to the SLA threshold and the fourth SLA metric is less than the SLA threshold, the cloud management platform determines a fourth resource in the first resource pool that matches the first resource, creates a second cloud instance, allocates the fourth resource and the third resource to the second cloud instance, and releases the first cloud instance. If the third SLA metric and the fourth SLA metric are greater than or equal to the SLA threshold, the cloud management platform determines a fourth resource matching the first resource in the first resource pool, determines a fifth resource matching the third resource in the second resource pool, creates a second cloud instance, allocates the fourth resource and the fifth resource to the second cloud instance, and releases the first cloud instance.
7. The method according to any one of claims 1 to 6, characterized in that, The first resource pool includes any one of the following: a computing resource pool, a storage resource pool, and a network resource pool. The computing resource pool includes multiple computing resources located in the same rack or adjacent racks. The storage resource pool includes multiple storage resources located in the same rack or adjacent racks. The network resource pool includes multiple network resources located in the same rack or adjacent racks.
8. The method according to any one of claims 1 to 7, characterized in that, The first cloud instance includes any of the following: physical server, virtual machine, container, micro virtual machine, and bare metal server.
9. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes multiple resource pools of different types. Each resource pool contains multiple resources located in the same or adjacent racks. Each resource pool contains multiple resources of the same type. The cloud management platform includes: The first receiving module is configured to receive a cloud instance creation request from a tenant, wherein the cloud instance creation request is used to indicate the first type and first specification of the resources required by the tenant's first cloud instance; The first determining module is configured to, based on the cloud instance creation request, determine a first resource pool that meets the first type from the plurality of resource pools, and determine a first resource that meets the first specification from the first resource pool, wherein the first resource is at least one of the plurality of resources in the first resource pool; The first allocation module is used to create the first cloud instance and allocate the first resource to the first cloud instance.
10. The cloud management platform according to claim 9, characterized in that, The cloud instance creation request is also used to indicate the SLA threshold set by the tenant, and the first determining module is used to: From the multiple resources in the first resource pool, determine multiple first resource groups that meet the first specification, each first resource group containing at least one resource; Determine the first SLA metrics for the plurality of first resource groups, wherein the first SLA metrics for any one of the first resource groups include at least one of the following: the latency required for the first cloud instance to access the first resource group, and the bandwidth required for the first cloud instance to access the first resource group. From the plurality of first resource groups, the first resource group whose first SLA index is less than the SLA threshold is selected as the first resource.
11. The cloud management platform according to claim 9 or 10, characterized in that, The cloud management platform also includes: The second receiving module is configured to receive a cloud instance expansion request from the tenant, wherein the cloud instance expansion request is configured to indicate the first type and second specification of the resources required to be added to the tenant's first cloud instance; The second determining module is used to determine, based on the cloud instance expansion request, a first resource pool that meets the first type from the plurality of resource pools, and a second resource that meets the second specification from the first resource pool, wherein the first resource and the second resource are different resources in the first resource pool; The second allocation module is used to allocate the second resource to the first cloud instance.
12. The cloud management platform according to claim 11, characterized in that, The second determining module is used for: From the remaining resources in the first resource pool, excluding the first resource, determine a plurality of second resource groups that meet the second specification; Determine the second SLA metrics for the plurality of second resource groups; From the plurality of second resource groups, the second resource group whose second SLA index is less than the SLA threshold is selected as the second resource.
13. The cloud management platform according to any one of claims 9 to 12, characterized in that, The cloud instance creation request is also used to indicate the second type and third specification of the resources required by the tenant's first cloud instance, and the cloud management platform also includes: The third determining module is used to determine a second resource pool that satisfies the second type from the plurality of resource pools, and to determine a third resource that satisfies the third specification from the second resource pool, wherein the third resource is at least one of the plurality of resources in the second resource pool; The third allocation module is used to allocate the third resource to the first cloud instance.
14. The cloud management platform according to claim 13, characterized in that, The cloud management platform also includes: The third receiving module is used to receive a cloud instance migration request for the first cloud instance from the tenant; The fourth determining module is used to determine the third SLA metric of the first resource and the fourth SLA metric of the third resource based on the cloud instance migration request. The fourth allocation module is used for: If the third SLA metric and the fourth SLA metric are less than the SLA threshold, create a second cloud instance, allocate the first resource and the third resource to the second cloud instance, and release the first cloud instance. If the third SLA metric is greater than or equal to the SLA threshold and the fourth SLA metric is less than the SLA threshold, a fourth resource matching the first resource is determined in the first resource pool, a second cloud instance is created, the fourth resource and the third resource are allocated to the second cloud instance, and the first cloud instance is released. If the third SLA metric and the fourth SLA metric are greater than or equal to the SLA threshold, a fourth resource matching the first resource is determined in the first resource pool, a fifth resource matching the third resource is determined in the second resource pool, a second cloud instance is created, the fourth resource and the fifth resource are allocated to the second cloud instance, and the first cloud instance is released.
15. The cloud management platform according to any one of claims 9 to 14, characterized in that, The first resource pool includes any one of the following: a computing resource pool, a storage resource pool, and a network resource pool. The computing resource pool includes multiple computing resources located in the same rack or adjacent racks. The storage resource pool includes multiple storage resources located in the same rack or adjacent racks. The network resource pool includes multiple network resources located in the same rack or adjacent racks.
16. The cloud management platform according to any one of claims 9 to 15, characterized in that, The first cloud instance includes any of the following: physical server, virtual machine, container, micro virtual machine, and bare metal server.
17. A cloud service system, characterized in that, The cloud service system includes infrastructure that provides cloud services and a cloud management platform that manages the infrastructure. The infrastructure includes multiple resource pools of different types. Each resource pool contains multiple resources set in the same rack or adjacent racks. Each resource pool contains multiple resources of the same type. The cloud management platform is used to implement the steps performed by the cloud management platform in any one of the methods described in claims 1 to 8.
18. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, each computing device including a processor and memory: The memory is used to store instructions; The processor is configured to, according to the instructions, cause the computing device cluster to perform the method of any one of claims 1 to 8.
19. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 8.
20. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Cloud platform, cloud platform management method and device, electronic equipment and storage medium
CN112286632A
Graphic program online development method and system based on cloud technology and related equipment
CN116185366A
Distributed heterogeneous resource pool scheduling method and device, server and storage medium
CN116360994A
Resource management method and device, equipment and storage medium
CN117608823A
Computing power management method and device, computing power scheduling equipment and storage medium
CN117632509A