Cloud host change method based on public cloud technology, and cloud management platform
By classifying cloud servers into server categories and optimizing change batches based on importance and regional distribution, the problem of abnormal cloud server changes in the public cloud was solved, achieving secure and efficient canary deployments and minimizing business impact.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-07-30
AI Technical Summary
In public clouds, due to the differences between regions, anomalies can easily occur during the cloud host change process, leading to service failures and damage to core businesses, thus affecting the competitiveness of cloud services.
By dividing cloud servers into multiple server categories and selecting target cloud servers within each category for modification, and using operational data and importance to divide modification batches, verification is prioritized in representative areas to ensure that the modification does not cause failures, and then gradually expanded to other servers.
This reduces the probability of abnormal changes to cloud servers, ensures the safe and efficient canary rollout of cloud servers, minimizes the impact on tenant businesses, and improves user experience.
Smart Images

Figure CN2026075044_30072026_PF_FP_ABST
Abstract
Description
Cloud Host Modification Method and Cloud Management Platform Based on Public Cloud Technology
[0001] This application claims priority to Chinese Patent Application No. 202510127456.2, filed on January 27, 2025, entitled "Method for Changing Cloud Hosts and Cloud Management Platform Based on Public Cloud Technology", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud service technology, and in particular to a cloud host modification method and cloud management platform based on public cloud technology. Background Technology
[0003] With the continuous operation and growth of public clouds, the number of cloud hosts (hereinafter referred to as "hosts" or "cloud servers") in public clouds can reach millions or more. The infrastructure environment is becoming increasingly complex, and there may be differences in configuration between different cloud resource deployment regions / availability zones (AZs), such as significant differences in hardware, software, resources, and customers. Therefore, based on certain requirements, cloud management platforms need to make changes to the software services, operating systems, and physical devices on cloud hosts, which will be collectively referred to as cloud host changes. If an anomaly occurs during the change process, a group failure of services or cloud hosts may occur, or even damage the customer's core business. The impact of the failure is uncontrollable and directly affects the product competitiveness of the cloud business.
[0004] Currently, cloud vendors typically determine the order in which to modify cloud hosts in different regions based on the importance of those regions, in order to control the risk of service launch after modifications to cloud hosts.
[0005] However, due to the differences between regions, making changes to cloud servers in the current way can easily lead to changes that cause abnormalities. Summary of the Invention
[0006] This application provides a cloud server modification method and cloud management platform based on public cloud technology. This application can reduce the probability of modification anomalies during cloud server modifications, and helps to ensure the safe and efficient canary deployment of cloud servers. The technical solution provided by this application is as follows:
[0007] Firstly, this application provides a method for modifying cloud hosts based on public cloud technology. This method is applied to a cloud management platform. The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes cloud hosts. Cloud hosts are used to deploy instances that implement tenant services. The method includes: based on the operation and maintenance data of M cloud hosts managed by the cloud management platform, dividing the M cloud hosts into N host classes, each host class including multiple cloud hosts, where the operation and maintenance data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts, and M and N are both positive integers greater than 1; selecting multiple target cloud hosts from the multiple cloud hosts included in each of the N host classes, wherein the multiple target cloud hosts in a first host class are a subset of the multiple cloud hosts included in the first host class, and the first host class is any one of the N host classes; in a first batch of modifications to the M cloud hosts, modifying the target cloud hosts selected from the N host classes; in subsequent batches of modifications to the M cloud hosts, modifying the other cloud hosts among the M cloud hosts besides the target cloud hosts.
[0008] In this method, since the changes to multiple cloud hosts within the same host class are essentially the same, the cloud management platform selects a subset of cloud hosts from each of the N host classes. In the first batch of changes involving M cloud hosts, these selected hosts are modified. This is equivalent to first modifying a subset of cloud hosts for any given host class to verify the effect of the changes on that host class. This allows for the continuation of modifications to the remaining cloud hosts in the host class when the change results indicate that the changes will not cause failures, and timely termination of the changes when the results indicate that the changes will cause failures. This reduces the probability of changes causing anomalies and helps to ensure the safe and efficient canary rollout of cloud hosts.
[0009] In one possible implementation, selecting a target cloud host from the cloud hosts included in each of the N host classes includes: obtaining the importance of each cloud host among the P cloud hosts in the first host class to the tenant's business, where the first host class is any one of the N host classes and P is a positive integer; dividing the P cloud hosts into multiple change batches to be executed sequentially based on their importance, wherein the cloud hosts in the later change batches have a greater importance than the cloud hosts in the earlier change batches; and determining the cloud hosts in the first change batch among the multiple change batches as the target cloud hosts selected from the cloud hosts included in the first host class.
[0010] By dividing P cloud hosts into multiple batches of changes to be executed sequentially based on their importance to the tenant's business, and ensuring that the cloud hosts in the later batches of changes are more important than those in the earlier batches, changes to cloud hosts with lower importance can be made first, thereby reducing the impact of change anomalies on the tenant's business and improving the user experience.
[0011] In one possible implementation, the importance of a cloud server to a tenant's business is determined based on one or more of the following: the resource utilization of the cloud server, the number and level of tenants served by the cloud server, or the number and level of business applications deployed on the cloud server.
[0012] In one possible implementation, M cloud hosts are deployed in multiple cloud resource deployment regions. Selecting a target cloud host from the cloud hosts included in each of the N host classes includes: determining the distribution of the N host classes in the multiple cloud resource deployment regions; based on the distribution, determining the minimum coverage of the N host classes in the multiple cloud resource deployment regions, wherein the minimum coverage includes the N host classes; and obtaining the target cloud host selected from the N host classes based on the cloud hosts included in the minimum coverage.
[0013] By analyzing the distribution of N host classes across multiple cloud resource deployment regions, determining the minimum coverage area for each host class, and then making changes to the cloud hosts included in the minimum coverage area in the first change batch, this approach prioritizes changes to cloud hosts within the representative minimum coverage area. This allows for the continuation of changes to the remaining cloud hosts in the same host class if the change results indicate that the change will not cause a failure, and timely termination of changes if the change results indicate that the change will cause a failure. This reduces the probability of change anomalies caused by changes to cloud hosts and helps to ensure the safe and efficient canary rollout of cloud hosts.
[0014] Furthermore, selecting the target cloud host from the cloud hosts included in each of the N host classes further includes: determining the extension range of the minimum coverage area in multiple cloud resource deployment regions based on the distribution situation, and the union of the cloud resource deployment regions where the cloud hosts included in the minimum coverage area and the extension range are located constitutes multiple cloud resource deployment regions. Accordingly, based on the cloud hosts included in the minimum coverage area, the target cloud host selected from the N host classes is obtained, including: obtaining the target cloud host selected from the N host classes based on the cloud hosts included in the minimum coverage area and the extension range.
[0015] When the union of the cloud resource deployment regions where the cloud hosts included in the minimum coverage area and the extended coverage area are located constitutes multiple cloud resource deployment regions, and the cloud hosts included in the minimum coverage area and the extended coverage area are selected as the target cloud hosts, it is equivalent to selecting the target cloud hosts in all the multiple cloud resource deployment regions where the cloud hosts to be changed are located. When the target cloud hosts are changed in the first change batch, it is equivalent to changing them across the entire range of multiple cloud resource deployment regions where multiple cloud hosts to be changed are located. This ensures that the representative range of the cloud hosts to be changed first is included, achieving coverage and immersion verification of the entire deployment range. It is also more helpful to discover problems in advance and control the impact of changes within a safe range.
[0016] In one possible implementation, based on the distribution, determining the extension range of the minimum coverage area across multiple cloud resource deployment areas includes: determining a first combination of cloud resource deployment areas where the cloud hosts included in the minimum coverage area are located, the first combination of cloud resource deployment areas including one or more target cloud resource deployment areas; determining a preceding extension range of the minimum coverage area in cloud resource deployment areas whose importance is lower than the lowest importance among the one or more target cloud resource deployment areas, the hosts included in the preceding extension range having a change order prior to the hosts included in the minimum coverage area; determining a subsequent extension range of the minimum coverage area in cloud resource deployment areas whose importance is higher than the highest importance among the one or more target cloud resource deployment areas, the hosts included in the subsequent extension range having a change order after the hosts included in the minimum coverage area; and, if the first combination of cloud resource deployment areas includes multiple target cloud resource deployment areas, for any two target cloud resource deployment areas with adjacent importance, determining an intermediate extension range of the minimum coverage area in cloud resource deployment areas whose importance is between that of the two target cloud resource deployment areas, the hosts included in the intermediate extension range having a change order between that of the hosts included in the two target cloud resource deployment areas.
[0017] In one possible implementation, for any target coverage area among the minimum coverage area and the extended coverage area, the target coverage area is determined in multiple cloud resource deployment areas based on the distribution, including: in the candidate cloud resource deployment areas used to determine the target coverage area, a second cloud resource deployment area combination is determined based on the combination of candidate cloud resource deployment areas including the most host classes; and the target coverage area is determined based on the second cloud resource deployment area combination.
[0018] Optionally, the second cloud resource deployment region combination also meets one or more of the following conditions:
[0019] 1) Among multiple combinations of candidate cloud resource deployment regions that include the most host types, the combination with the smallest total number of cloud resource deployment regions is selected. When a second cloud resource deployment region combination is selected based on this condition, and a target cloud host is chosen based on the selected second cloud resource deployment region combination, this is equivalent to selecting the target cloud host from cloud hosts that include all host types and whose deployment scope covers the fewest cloud resource deployment regions. In this way, if an anomaly occurs, the cloud resource deployment regions affected by the anomaly are minimized, thus controlling the risk of changes within the fewest possible cloud resource deployment regions.
[0020] 2) Among multiple combinations of candidate cloud resource deployment regions that include the most host types, the cloud host with the lowest importance is deployed. When a second cloud resource deployment region combination is selected from multiple combinations based on this condition, and the target cloud host is selected based on the selected second cloud resource deployment region combination, it is equivalent to selecting the target cloud host from a group of cloud hosts that include all host types and have the lowest importance. In this way, if an anomaly occurs, the anomaly will only affect the cloud host with the lowest importance to the tenant's business, thus reducing the impact of the anomaly on the tenant's business.
[0021] 3) Among multiple combinations of candidate cloud resource deployment regions that include the most host types, the minimum total number of cloud hosts is deployed. When selecting a second cloud resource deployment region combination based on this condition, and then selecting the target cloud host based on the selected second cloud resource deployment region combination, it is equivalent to selecting the target cloud host from the cloud resource deployment region combination that includes all host types and has the minimum number of deployed cloud hosts. In this way, if an anomaly occurs, the impact of the anomaly can be controlled within the minimum number of cloud hosts, reducing the scope of the impact of the anomaly.
[0022] In one possible implementation, each cloud resource deployment area includes multiple availability zones. Based on the second cloud resource deployment area combination, determining any target coverage area includes: in the second cloud resource deployment area combination, determining a target availability zone combination based on the combination of availability zones that include the most host classes; and determining the target availability zone combination as any target coverage area.
[0023] In one possible implementation, the target availability zone composition also satisfies one or more of the following conditions:
[0024] 1) Among multiple availability zone combinations that include the most host types in the second cloud resource deployment region combination, the one with the smallest total number of availability zones is selected. When selecting a target availability zone combination from multiple combinations based on this condition, and then selecting a target cloud host based on the selected target availability zone combination, it is equivalent to selecting a target cloud host from cloud hosts that include the most host types and whose deployment scope covers the fewest availability zones. In this way, if a change anomaly occurs, the availability zones affected by the change anomaly are minimized, and the change risk can be controlled within the fewest availability zones.
[0025] 2) In the second cloud resource deployment region combination, among multiple availability zone combinations that include the most host types, the cloud host with the lowest importance is deployed. When selecting a target availability zone combination from multiple combinations based on this condition, and then selecting the target cloud host based on the selected target availability zone combination, it is equivalent to selecting the target cloud host from the group of cloud hosts that include the most host types and have the lowest importance. In this way, if an anomaly occurs, the anomaly will only affect the cloud host with the lowest importance to the tenant's business, thus reducing the impact of the anomaly on the tenant's business.
[0026] 3) Alternatively, among multiple availability zone combinations that include the most host types in the second cloud resource deployment region combination, the minimum total number of cloud hosts is deployed. When selecting a target availability zone combination from multiple combinations based on this condition, and then selecting the target cloud host based on the selected target availability zone combination, it is equivalent to selecting the target cloud host from the cloud resource deployment region combination that includes the most host types and has the fewest deployed cloud hosts. In this way, if a change anomaly occurs, the impact of the change anomaly can be controlled within the minimum total number of cloud hosts, reducing the scope of the impact of the change anomaly.
[0027] Furthermore, when determining the multiple sub-batches that are executed sequentially within the second change batch, further consideration can be given to the first sub-batchens that are executed first, in order to further control the impact of the change within a safe range. In one possible implementation, M cloud hosts are modified in multiple change batches, and the second change batch of the multiple change batches includes multiple sub-batches executed sequentially, wherein the second change batch is any one of the multiple change batches. The first sub-batch executed first among multiple sub-batches includes one or more of the following cloud servers: idle cloud servers, and the total number of idle cloud servers in the first sub-batch does not exceed the first proportion of the total number of cloud servers in the second change batch; cloud servers belonging to a specified type of tenant, and the total number of cloud servers belonging to a specified type of tenant in the first sub-batch does not exceed the corresponding second proportion of the total number of cloud servers in the second change batch, the second proportion corresponding to the specified type of tenant is negatively correlated with the importance of the specified type of tenant; or, cloud servers used for business testing, and the total number of cloud servers used for business testing in the first sub-batch does not exceed the third proportion of the total number of cloud servers in the second change batch; the priority of changing idle cloud servers in the first sub-batch, the priority of changing cloud servers belonging to a specified type of tenant in the first sub-batch, and the priority of changing cloud servers used for business testing in the first sub-batch decreases in that order, and the priority of changing cloud servers belonging to a specified type of tenant in the first sub-batch is negatively correlated with the importance of the specified type of tenant.
[0028] In one possible implementation, the cloud hosts included in the first sub-batch also satisfy the following:
[0029] The first approach involves multiple cloud hosts with the same priority being changed in the first sub-batch. This first sub-batch includes cloud hosts serving fewer tenants and / or cloud hosts with fewer instances deployed. When multiple cloud hosts have the same priority in the first sub-batch, and a cloud host serving fewer tenants is selected, if an anomaly occurs during the change, it ensures that fewer tenants are affected, thus controlling the impact of the anomaly from the perspective of tenant quantity. Similarly, when multiple cloud hosts have the same priority in the first sub-batch, and a cloud host with fewer instances is selected, if an anomaly occurs during the change, it ensures that fewer instances are affected, thus controlling the impact of the anomaly from the perspective of instance quantity. Compared to cloud hosts with fewer instances, prioritizing cloud hosts serving fewer tenants ensures fewer instances are affected if an anomaly occurs during the change, reducing the impact of the anomaly on business operations and controlling the impact of the anomaly from the perspective of minimizing business impact.
[0030] The second approach involves multiple cloud hosts with services deployed under a specific tenant. The total number of cloud hosts with services deployed under a specific tenant in the first sub-batch is less than a first quantity threshold. By controlling the total number of cloud hosts with services deployed under a specific tenant in the first sub-batch, if an anomaly occurs during the change, the total number of cloud hosts with services deployed under a specific tenant affected by the anomaly can be controlled, thus keeping the impact of the anomaly on the services of the same tenant within a manageable range.
[0031] In one possible implementation, before selecting the target cloud host from the cloud hosts included in each of the N host classes, the method further includes merging the second host class and the third host class if the second host class and the third host class meet one or more of the following merging conditions. The merging conditions include: the feature similarity between the second host class and the third host class is less than a first similarity threshold; or, the number of cloud hosts included in both the second host class and the third host class is less than a first quantity threshold; and both the second host class and the third host class are one of the N host classes.
[0032] When the feature similarity between the second and third host classes is less than the first similarity threshold, the deployment strategies and risk controls for changes to the second and third host classes are basically the same. Therefore, the cloud management platform can merge the second and third host classes. When the number of cloud hosts included in both the second and third host classes is less than the first quantity threshold, since the number of cloud hosts included in the second and third host classes accounts for a small proportion of the total number of cloud hosts in the N host classes, even if the second host class is merged with the third host class, the impact of any change anomalies will be small. Therefore, the cloud management platform can merge the second host class with the third host class.
[0033] In one possible implementation, the method further includes one or more of the following operations: determining the start time and / or start conditions for each change batch in the first host class; determining the start time and / or start conditions for changes in the first host class; determining the start dependency order and / or start dependency conditions between different change batches in the first host class; determining the start dependency order and / or start dependency conditions between changes in different host classes; or, determining the sub-batch in the change batch in the first host class and the number of cloud hosts it includes. By determining the above factors and making changes to cloud hosts based on these factors, the effectiveness of making changes to cloud hosts can be further improved.
[0034] In one possible implementation, based on the operational data of M cloud hosts managed by the cloud management platform, the M cloud hosts are divided into N host classes. This includes: determining the operational characteristics of each cloud host based on the operational data of the M cloud hosts; and dividing the M cloud hosts into N host classes based on the operational characteristics of each cloud host. The operational data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. The cloud management platform can obtain operational characteristics reflecting the operating characteristics of the cloud hosts based on the operational data. The cloud management platform can classify the M cloud hosts based on their operational characteristics, and can group cloud hosts with similar characteristics into the same host class.
[0035] Secondly, this application provides a cloud management platform. This cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes cloud servers. The cloud servers are used to deploy instances that implement tenant services. The cloud management platform includes:
[0036] The classification module is used to divide the M cloud hosts managed by the cloud management platform into N host classes based on their operation and maintenance data. Each host class includes multiple cloud hosts. The operation and maintenance data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. M and N are both positive integers greater than 1.
[0037] The selection module is used to select multiple target cloud hosts from multiple cloud hosts included in each of the N host classes. The multiple target cloud hosts in the first host class are a part of the multiple cloud hosts included in the first host class, and the first host class is any one of the N host classes.
[0038] The change module is used to make changes to the target cloud hosts selected from N host classes in the first change batch that makes changes to M cloud hosts;
[0039] The change module is also used to make changes to other cloud hosts among the M cloud hosts, excluding the target cloud host, in subsequent change batches that make changes to M cloud hosts.
[0040] In one possible implementation, the selection module is specifically used to: obtain the importance of each cloud host among P cloud hosts in the first host class to the tenant's business, where the first host class is any one of N host classes and P is a positive integer; divide the P cloud hosts into multiple change batches to be executed sequentially based on their importance, wherein the cloud hosts in the later change batches are more important than the cloud hosts in the earlier change batches; and determine the cloud hosts in the first change batch among the multiple change batches as the target cloud hosts selected from the cloud hosts included in the first host class.
[0041] In one possible implementation, the importance of a cloud server to a tenant's business is determined based on one or more of the following: the resource utilization of the cloud server, the number and level of tenants served by the cloud server, or the number and level of business applications deployed on the cloud server.
[0042] In one possible implementation, M cloud hosts are deployed in multiple cloud resource deployment areas. The selection module is specifically used to: determine the distribution of N host classes in the multiple cloud resource deployment areas; based on the distribution, determine the minimum coverage of the N host classes in the multiple cloud resource deployment areas, wherein the minimum coverage includes the N host classes; and based on the cloud hosts included in the minimum coverage, obtain the target cloud host selected from the N host classes.
[0043] In one possible implementation, the selection module selects the target cloud host from the cloud hosts included in each of the N host classes, and further includes: determining the extension range of the minimum coverage area in multiple cloud resource deployment regions based on the distribution, and the union of the cloud resource deployment regions where the cloud hosts included in the minimum coverage area and the extension range are located is the multiple cloud resource deployment regions.
[0044] Accordingly, the selection module, based on the cloud hosts included in the minimum coverage area, obtains the target cloud hosts selected from N host classes, including: based on the cloud hosts included in the minimum coverage area and the extended coverage area, the target cloud hosts selected from N host classes.
[0045] In one possible implementation, the selection module determines the extension range of the minimum coverage area across multiple cloud resource deployment regions based on the distribution, including: determining a first combination of cloud resource deployment regions where the cloud hosts included in the minimum coverage area are located, the first combination of cloud resource deployment regions including one or more target cloud resource deployment regions; determining a preceding extension range of the minimum coverage area in cloud resource deployment regions with an importance lower than the lowest importance among the one or more target cloud resource deployment regions, wherein the change order of the hosts included in the preceding extension range precedes the change order of the hosts included in the minimum coverage area; and determining the highest importance among the cloud resource deployment regions with an importance higher than the highest importance among the one or more target cloud resource deployment regions. In a cloud resource deployment area of a certain importance, a subsequent extension range of minimum coverage is determined, wherein the change order of hosts included in the subsequent extension range is after the change order of hosts included in the minimum coverage area; in the case where the first cloud resource deployment area combination includes multiple target cloud resource deployment areas, for any two target cloud resource deployment areas with adjacent importance, an intermediate extension range of minimum coverage is determined in the cloud resource deployment area with an importance between the importance of any two target cloud resource deployment areas, wherein the change order of hosts included in the intermediate extension range is between the change order of hosts included in any two target cloud resource deployment areas.
[0046] In one possible implementation, for any target coverage area among the minimum coverage area and the extended coverage area, the selection module determines any target coverage area among multiple cloud resource deployment areas based on the distribution, including: determining a second cloud resource deployment area combination based on the combination of candidate cloud resource deployment areas including the most host classes among the candidate cloud resource deployment areas used to determine any target coverage area; and determining any target coverage area based on the second cloud resource deployment area combination.
[0047] In one possible implementation, the second cloud resource deployment area combination also satisfies one or more of the following conditions: among multiple combinations of candidate cloud resource deployment areas including the most host classes, it has the minimum total number of cloud resource deployment areas; among multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys cloud hosts of the lowest importance level; or, among multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys the minimum total number of cloud hosts.
[0048] In one possible implementation, each cloud resource deployment area includes multiple availability zones. The selection module determines any target coverage area based on the second cloud resource deployment area combination, including: determining a target availability zone combination based on the combination of availability zones that include the most host classes in the second cloud resource deployment area combination; and determining the target availability zone combination as any target coverage area.
[0049] In one possible implementation, the target availability zone combination also satisfies one or more of the following conditions: among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it has the minimum total number of availability zones; among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it deploys cloud hosts of the lowest importance; or, among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it deploys the minimum total number of cloud hosts.
[0050] In one possible implementation, M cloud hosts are modified in multiple change batches, and the second change batch of the multiple change batches includes multiple sub-batches executed sequentially, and the second change batch is any one of the multiple change batches. The first sub-batch executed first among multiple sub-batches includes one or more of the following cloud servers: idle cloud servers, and the total number of idle cloud servers in the first sub-batch does not exceed the first proportion of the total number of cloud servers in the second change batch; cloud servers belonging to a specified type of tenant, and the total number of cloud servers belonging to a specified type of tenant in the first sub-batch does not exceed the corresponding second proportion of the total number of cloud servers in the second change batch, the second proportion corresponding to the specified type of tenant is negatively correlated with the importance of the specified type of tenant; or, cloud servers used for business testing, and the total number of cloud servers used for business testing in the first sub-batch does not exceed the third proportion of the total number of cloud servers in the second change batch; the priority of changing idle cloud servers in the first sub-batch, the priority of changing cloud servers belonging to a specified type of tenant in the first sub-batch, and the priority of changing cloud servers used for business testing in the first sub-batch decreases in that order, and the priority of changing cloud servers belonging to a specified type of tenant in the first sub-batch is negatively correlated with the importance of the specified type of tenant.
[0051] In one possible implementation, the cloud hosts included in the first sub-batch also satisfy the following: for multiple cloud hosts with the same priority that change in the first sub-batch, the first sub-batch includes: cloud hosts that provide services to fewer tenants, and / or cloud hosts with fewer instances deployed; and / or, for multiple cloud hosts that deploy services to a specified tenant, the total number of cloud hosts that deploy services to a specified tenant included in the first sub-batch is less than a first quantity threshold.
[0052] In one possible implementation, the cloud management platform further includes a merging module, used to merge the second host class and the third host class when the second host class and the third host class meet one or more of the following merging conditions, the merging conditions including: the feature similarity between the second host class and the third host class is less than a first similarity threshold, or the number of cloud hosts included in the second host class and the third host class is less than a first quantity threshold, and the second host class and the third host class are both one of N host classes.
[0053] In one possible implementation, the selection module is also configured to perform one or more of the following operations: determine the start time and / or start conditions for each change batch in the first host class; determine the start time and / or start conditions for changes in the first host class; determine the start dependency order and / or start dependency conditions between different change batches in the first host class; determine the start dependency order and / or start dependency conditions between changes in different host classes; or, determine the sub-batch in the change batch in the first host class and the number of cloud hosts it includes.
[0054] In one possible implementation, the classification module is specifically used to: determine the operation and maintenance characteristics of each cloud host among the multiple cloud hosts based on the operation and maintenance data of M cloud hosts; and classify the M cloud hosts into N host classes based on the operation and maintenance characteristics of each cloud host among the multiple cloud hosts.
[0055] Thirdly, this application provides a computing device including a memory and a processor, wherein the memory stores program instructions and the processor executes the program instructions to implement the methods provided in the first aspect of this application and any of its possible implementations.
[0056] Fourthly, this application provides a computing device cluster, including multiple computing devices, each computing device including multiple processors and multiple memories, the multiple memories storing program instructions, and the multiple processors executing the program instructions, so that the computing device cluster implements the method provided in the first aspect of this application and any possible implementation thereof.
[0057] Fifthly, this application provides a computer-readable storage medium that is a non-volatile computer-readable storage medium, which includes program instructions that, when executed on a computing device cluster, cause the computing device cluster to implement the methods provided in the first aspect of this application and any of its possible implementations.
[0058] Sixthly, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the methods provided in the first aspect of this application and any possible implementation thereof. Attached Figure Description
[0059] Figure 1 is a structural diagram of an implementation scenario involving a cloud host modification method based on public cloud technology provided in an embodiment of this application;
[0060] Figure 2 is a schematic diagram of the deployment of basic resources in a data center according to an embodiment of this application;
[0061] Figure 3 is a flowchart illustrating how to modify a cloud host according to its host class, as provided in an embodiment of this application.
[0062] Figure 4 is a flowchart illustrating how to select multiple target cloud hosts from multiple cloud hosts included in each of N host classes, according to an embodiment of this application.
[0063] Figure 5 is a schematic diagram of multiple change batches provided in an embodiment of this application;
[0064] Figure 6 is a flowchart illustrating another method for selecting multiple target cloud hosts among multiple cloud hosts included in each of N host classes, as provided in an embodiment of this application.
[0065] Figure 7 is a flowchart illustrating another method for selecting multiple target cloud hosts among multiple cloud hosts included in each of N host classes, as provided in an embodiment of this application.
[0066] Figure 8 is a flowchart illustrating how to determine the minimum coverage of N host classes in multiple cloud resource deployment regions according to an embodiment of this application;
[0067] Figure 9 is a flowchart illustrating how to determine the minimum coverage area based on a combination of second cloud resource deployment areas according to an embodiment of this application.
[0068] Figure 10 is a flowchart illustrating the process of determining the minimum coverage extension range in multiple cloud resource deployment areas according to an embodiment of this application.
[0069] Figure 11 is a schematic diagram of a modified pipeline provided in an embodiment of this application;
[0070] Figure 12 is a flowchart of a method for classifying a cloud host to be modified into multiple host classes according to an embodiment of this application;
[0071] Figure 13 is a flowchart of a cloud management platform according to an embodiment of this application, which divides the cloud host to be changed into multiple host classes;
[0072] Figure 14 is a flowchart of another method for classifying cloud hosts to be modified into multiple host classes, provided in an embodiment of this application.
[0073] Figure 15 is a schematic diagram of the functional modules for implementing a cloud host modification method based on public cloud technology, provided in an embodiment of this application;
[0074] Figure 16 is a schematic diagram of the structure of a cloud management platform provided in an embodiment of this application;
[0075] Figure 17 is a schematic diagram of another cloud management platform provided in an embodiment of this application;
[0076] Figure 18 is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0077] Figure 19 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0078] Figure 20 is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0080] To facilitate understanding, the technologies and background involved in the embodiments of this application will be introduced below.
[0081] Cloud computing is a type of distributed computing that refers to a network that centrally manages and schedules a large number of computing and storage resources to provide on-demand services to users. These computing and storage resources are provided through clusters of computing devices located in data centers. Furthermore, cloud computing can provide users with various types of services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Infrastructure as a Service provides virtual machines or other resources as a service to tenants. Platform as a Service provides a development platform as a service to tenants. Software as a Service provides applications (Apps) as a service to customers.
[0082] An Internet Data Center (IDC) is a facility and related service system that provides operation and maintenance for equipment that centrally collects, stores, processes, and transmits data, based on the Internet. Conceptually, it can be understood as a public, commercial Internet "server room," and it is also a professional IT service and a crucial infrastructure for the IT industry. IDC is not only a service concept but also a network concept; it constitutes part of the network infrastructure resources, like backbone networks and access networks, providing high-end data delivery and high-speed access services. Generally, a tenant's on-premises IDC can be understood as their physical server room, where the tenant utilizes existing Internet communication lines and bandwidth resources to establish a standardized, telecommunications-grade server room environment to provide comprehensive services such as server hosting, leasing, and related value-added services. A cloud data center is an Internet data center deployed using the infrastructure resources owned by cloud vendors.
[0083] A resource pool is a collection of various hardware and software resources involved in a cloud data center. Typically, resources in a resource pool can be categorized by type, such as computing resources, storage resources, and network resources.
[0084] A physical machine (PM) is the physical resource used to host virtualization technology. It is also called a physical server. Typically, a physical machine is used to deploy virtual instances. A physical machine has multiple physical devices. For example, a physical server has physical devices such as processors and memory. Multiple virtual instances can be deployed on a single physical machine, sharing the machine's physical resources. Depending on the use case, multiple virtual instances deployed on a single physical machine can belong to the same tenant or to different tenants.
[0085] Virtualization is a resource management technology. Virtualization abstracts and transforms various physical resources of a host, such as computing, network, and storage resources, breaking down the indivisible barriers between the host's physical structures. This allows tenants to utilize these resources in a better way than the original configuration. Resources obtained through virtualization are called virtualized resources, and virtualized resources are not limited by the existing physical resource deployment methods, geographical location, or physical configuration.
[0086] Virtualized resources are typically provided to tenants in the form of virtual instances. Virtual instances utilize the host's hardware resources and run on the host's operating system (OS). Applications run within the virtual instance to implement the tenant's business logic. The host's hardware resources can be allocated to one or more tenants at the virtual instance level. Different virtual instances are isolated from each other, allowing tenants to use physical resources conveniently and flexibly while maintaining security and isolation, and significantly improving the utilization of physical resources. Typically, virtual instances can be virtual machines, containers, or independent processes (such as functions). Virtual instances can also be called Elastic Compute Service (ECS) or Elastic Instances (different cloud service providers may use different names).
[0087] A virtual machine (VM) is a complete computer system with full hardware system functionality, simulated using virtualization technology and running in a completely isolated environment. A subset of the instructions in a VM can be processed on the host machine, while other instructions can be executed in a simulated manner. A VM is also called a virtual server. A VM can be viewed as a collection of virtual devices, which possess full hardware system functionality and run in a completely isolated environment. Virtual devices are created by virtualizing physical devices that can share resources. For example, a virtual processor, created by virtualizing a processor, is a virtual device. Similarly, a training card, created by virtualizing a field-programmable gate array (FPGA), is also a virtual device. For instance, the VM in this application can be a kernel-based virtual machine (KVM). Any task that can be performed on a server can also be performed in a VM. When creating a virtual machine on a server, a portion of the physical machine's hard drive and memory capacity is used as the virtual machine's hard drive and memory capacity. Each virtual machine has its own independent hard drive and operating system, and virtual machine tenants can operate the virtual machine as if it were a server. The runtime environments (such as virtual machine applications, operating systems, and virtual hardware) in different virtual machines are completely isolated, and communication between different virtual machines requires the virtual machine manager to forward network packets.
[0088] Containers utilize the namespace and cgroup technologies supported by the Linux kernel to isolate application processes and their dependencies (the runtime environment's bins / libs, specifically all files required to run the application) within an independent runtime environment. Containers provide a lightweight virtual runtime environment. Containers are created by packaging all the code, libraries, and dependencies of a tenant's application into an image. When the image is executed, it runs in a virtual runtime environment. At this point, the container is a runtime instance of the image, similar to a lightweight sandbox, which can be started, stopped, and deleted. The infrastructure for containers can be server hardware or virtual machines in the cloud (i.e., containers can also be deployed within virtual machines). The operating system uses the Linux kernel and supports namespaces and cgroups. Namespaces are used to isolate processes, while cgroups are used to allocate process resources, specifically virtual processors and memory allocated to the process. The container engine, similar to a virtual machine manager, runs within the operating system and is used to manage containers. Compared to virtual machines, which come with their own operating system, containers do not have an operating system. Instead, containers run as processes within the host machine's operating system. As a result, containers start up faster than virtual machines, making them particularly suitable for lightweight applications. Furthermore, a single host machine can run thousands of containers (processes) simultaneously.
[0089] Resource pooling refers to integrating various computing and storage resources into a unified resource pool for unified dynamic allocation and management. Resource pooling enables high resource sharing, improves resource utilization, simplifies resource management, and provides users with flexible on-demand allocation services.
[0090] In the field of computer science, orchestration refers to the automated arrangement, coordination, and management of complex computer systems, middleware, and business processes. Orchestration typically involves three aspects: 1) resource orchestration, responsible for resource allocation; 2) workload orchestration, responsible for sharing workloads among resources and managing their lifecycles; and 3) service orchestration, responsible for service discovery and high availability, etc.
[0091] Network interface card (NIC): also known as network interface controller, network adapter, or local area network receiver, is a type of computer hardware designed to allow hosts or computing devices to communicate over a network.
[0092] Memory (RAM): Also known as internal memory or main memory, its function is to temporarily store the data processed by the CPU, as well as the data exchanged with external storage devices such as hard drives.
[0093] Checkpoint restore (CR) technology is a fault recovery technique for compute-intensive applications, widely used in high-performance computing (HPC), scientific computing, and artificial intelligence (AI) training. Its basic principle is to periodically or periodically save the cluster's running state during task execution, persisting completed results and task progress states to storage. When a fault occurs, the cluster reads the most recently saved checkpoint and re-executes the task from that checkpoint. These saved running states are called checkpoints (CKPT). This allows the cluster to directly restore the computation progress to the checkpoint after a fault, ensuring that subsequent recovery only loses the progress between the fault time and the previous checkpoint, avoiding the time overhead of re-execution. CR technology can only handle transient faults, i.e., faults that can be recovered by rerunning the program or restarting the server. For other types of faults, CR technology must be combined with faulty node replacement capabilities to achieve fault recovery and ensure high cluster availability.
[0094] A computing cluster is a group of computers connected through various hardware and software technologies. These computers work closely together to complete computational tasks that are difficult for a single computer to perform. Because computing clusters have powerful overall computing capabilities, and because the failure of any computer during computation will cause the cluster's tasks to be interrupted, they are also called tightly coupled, heavily coupled computing clusters. The computers in a computing cluster are also called computing nodes (or simply nodes).
[0095] A logical cluster is a logical computing cluster presented to the user, hosted on top of a physical cluster. A logical cluster may be the entire physical cluster or a part of it. Physical computing clusters often contain thousands or even tens of thousands of computers, while practical applications only require tens or hundreds of computers. Therefore, the physical cluster needs to be sliced, i.e., logical clusters. Network isolation between logical clusters is achieved by default. In public cloud technology, the user is the tenant.
[0096] As public clouds continue to operate and grow in scale, the number of cloud servers in public clouds can reach millions or more. The infrastructure environment is becoming increasingly complex, and there may be significant differences in configuration between different cloud resource deployment regions / availability zones (AZs), including substantial variations in hardware, software, resources, and customers. Therefore, based on certain requirements, cloud management platforms need to modify the software services, operating systems, and physical devices on cloud servers; this is collectively referred to as cloud server modification. Examples include service deployment, configuration modification, service restart, and server restart. If an anomaly occurs during the modification process, a group failure of services or cloud servers may occur, potentially even damaging the customer's core business. The impact of such failures is uncontrollable and directly affects the product competitiveness of cloud services.
[0097] Currently, cloud vendors typically determine the order in which to modify cloud servers in different regions based on the region's importance, in order to control the risk of service deployment after changes. However, due to the differences between regions, modifying cloud servers in this current manner can easily lead to changes causing anomalies.
[0098] For example, the order in which changes are made to cloud hosts in different regions is typically determined by the site reliability engineer (SRE) based on the importance of the regions. This order can be represented using a ring system. A ring system consists of multiple sequential stages, each requiring changes to cloud hosts deployed in a specific Availability Zone (AZ) within a specific region. Each stage can be called a Ring. SREs usually require strict adherence to the Ring sequence for canary deployments. On one hand, the pipelined change system creates a deployment pipeline in accordance with Ring rules: first, the regions in the earlier Rings are deployed, followed by the later Rings, with a certain soaking time between different Rings. Within the same Ring, changes can be implemented across different Regions but not simultaneously. Changes between different AZs within the same Region must be implemented sequentially. An Availability Zone (AZ) phase is set up within each AZ, and after the changes in the AZ phase are completed, a period is retained to verify whether the new functionality will cause any anomalies. The gray-scale testing phase involves a certain soaking time between different Rings. After modifying the previous Ring, the modification to the next Ring is not immediately executed. Instead, the modification waits for the previous Ring to run for a certain period. If the previous Ring's operation indicates that the modification did not cause any anomalies, then the modification to the next Ring is executed. This process from the start of modifying one Ring to the start of modifying the next Ring is called the gray-scale soaking phase of a Ring. Furthermore, based on the business requirements of the services provided by the cloud server, more granular gray-scale policy rules can be customized to collaboratively control deployment risks.
[0099] However, due to differences between regions, under the current SRE Ring rules, each Ring cannot cover all functional scenarios of the service. Although the number of issues decreases with each subsequent Ring, there is still a significant probability of new issues arising. The gray-scale phase is distributed throughout the entire deployment process, resulting in low efficiency in issue discovery and reduced deployment efficiency. For example, if a gray-scale phase includes four Rings, and each Ring undergoes a one-week gray-scale phase, the overall gray-scale phase can last up to one month. When new events occur in the production network, the SRE will add new constraints to the Ring rules, increasing the operational rules for the service and making change operations more difficult, thus worsening the deployment experience. This requires more cautious deployment planning and the deployment of more operations engineers (OPS) to ensure a smooth deployment, increasing deployment costs. While issues during the deployment process can be resolved by releasing patches, patch versions must also be deployed according to the Ring rules. If new issues continue to emerge in multiple Rings, the service requires continuous R&D / testing to release patches. The more patches released for a service, the lower the development efficiency and the higher the deployment cost. Furthermore, because development / testing cannot obtain the full runtime picture of the service across the entire network in advance, testing and evaluation during the development / testing / pre-production phases can only be conducted from a feature perspective or a single runtime perspective. It is impossible to simulate and test from all runtime dimensions of the production environment in advance to identify problems, thus affecting the quality of service version releases. At the same time, SREs cannot effectively and completely identify the impact of changes on the core service, upstream and downstream services, and customer businesses. Therefore, currently, efficient and secure canary rollouts are not feasible.
[0100] In view of this, this application provides a method for modifying cloud hosts based on public cloud technology. This method is applied to a cloud management platform. The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes cloud hosts. Cloud hosts are used to deploy instances that implement tenant services. In this method, the cloud management platform, based on the operation and maintenance data of M cloud hosts managed by the cloud management platform, divides the M cloud hosts into N host classes. Then, within each of the N host classes, a subset of cloud hosts is selected from among the multiple cloud hosts included. The selected cloud hosts are the target cloud hosts. Then, in the first batch of modifications to the M cloud hosts, the target cloud hosts selected from the N host classes are modified. In subsequent batches of modifications to the M cloud hosts, the other cloud hosts among the M cloud hosts, excluding the target cloud hosts, are modified. Each host class includes multiple cloud hosts. The operation and maintenance data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. M and N are both positive integers greater than 1.
[0101] In this method, since the changes to multiple cloud hosts within the same host class are essentially the same, the cloud management platform selects a subset of cloud hosts from each of the N host classes. In the first batch of changes involving M cloud hosts, these selected hosts are modified. This is equivalent to first modifying a subset of cloud hosts for any given host class to verify the effect of the changes on that host class. This allows for the continuation of modifications to the remaining cloud hosts in the host class when the change results indicate that the changes will not cause failures, and timely termination of the changes when the results indicate that the changes will cause failures. This reduces the probability of changes causing anomalies and helps to ensure the safe and efficient canary rollout of cloud hosts.
[0102] This article provides a detailed introduction to the technical solution of this application from multiple perspectives, including implementation scenarios, methods and processes, hardware devices, and software devices.
[0103] The following are examples illustrating the implementation scenarios of the embodiments of this application.
[0104] Figure 1 is a structural diagram of an implementation scenario involving a cloud host modification method based on public cloud technology provided in this application embodiment. As shown in Figure 1, the implementation scenario includes: cloud system 1 and client 2. Cloud system 1 and client 2 can establish a communication connection through a network. Optionally, the network can be the Internet or other networks, which is not limited in this application embodiment. Tenants can interact with cloud system 1 through client 2. For example, tenants can send cloud service requests and other information to cloud system 1 through client 2. Cloud system 1 responds based on the information sent by client 2.
[0105] As shown in Figure 1, cloud system 1 includes a cloud management platform and infrastructure. The cloud management platform and infrastructure are connected via an internal network of the cloud system. In another implementation, the cloud management platform may optionally be located within the infrastructure. The cloud management platform is used to manage the infrastructure. The infrastructure is used to provide public cloud services. The infrastructure includes at least one data center (DC). The cloud management platform can connect to this at least one data center. In this case, the cloud management platform is used to manage this at least one data center. The data center deploys a large number of cloud resources owned by the cloud service provider, such as computing resources, storage resources, and network resources. Computing resources can be computing devices (such as servers) capable of providing computing power. For example, as shown in Figure 1, multiple servers are deployed in the data center. Cloud services may optionally be deployed on the servers. Cloud services are implemented by running virtual instances, hence also referred to as virtual instances deployed on servers to implement tenant services. Tenants can send cloud service requests and related information to the server through their client 2. The server can process the cloud service requests and related information and provide cloud services to the tenant based on the processed cloud service requests and related information.
[0106] The cloud management platform can be logically divided into: tenant console, compute management service, network management service, storage management service, authentication service, and image management service. The tenant console provides a user interface or application programming interface (API) for interaction with tenants. The compute management service manages servers running virtual instances and bare metal servers. The network management service manages network services (such as gateways and firewalls). The storage management service manages storage services (such as data bucket services). The authentication service manages tenant accounts and passwords. The image management service manages virtual instance images.
[0107] In the implementation scenario shown in Figure 1, a data center contains multiple servers. The servers consist of a hardware layer and a software layer. The hardware layer comprises the standard server configuration, including hardware devices such as processors, memory, network interface cards (NICs), disks, and buses. The software layer includes the operating system installed and running on the server. The operating system relative to the virtual machine can be called the host operating system. The host operating system runs a virtual machine manager (also known as a hypervisor). The virtual machine manager's role is to implement compute virtualization, network virtualization, and storage virtualization for the virtual machines, and to manage the virtual machines.
[0108] The virtual machine manager runs a cloud management platform client. This client receives control plane commands from the cloud management platform, creates virtual instances on the server based on these commands, and manages the virtual instances throughout their lifecycle. For example, the client can monitor the hardware resource usage of the server in real time and report it to the cloud management platform. When the cloud management platform confirms that a virtual instance needs to be created on a specific server, it sends a virtual instance creation command to the client on that server. Upon receiving the command, the client creates the virtual instance on that server. In this way, tenants can create, manage, log in to, and operate virtual instances through the cloud management platform.
[0109] Servers can run virtual machines of different specifications. Virtual machine specifications are categorized as: general-purpose computing, memory-optimized, ultra-large memory, etc., with specific specifications under each type. After a tenant selects a virtual machine specification, the cloud management platform selects a server in the data center that supports that specification and ensures sufficient idle hardware resources on that server. Then, it creates and configures the virtual machine with that specification on that server. Configuring servers through the cloud management platform allows for the analysis and planning of server hardware resources. Based on the server's hardware performance, it plans the corresponding computing products for the physical hardware, such as planning virtual machines of different specifications, to meet the diverse needs of different tenants. Furthermore, differentiated pricing strategies can be implemented based on the performance differences of virtual machines of different specifications. For example, high-performance virtual instances can be sold at a higher price, while ordinary performance virtual instances can be sold at a lower price, allowing tenants to purchase virtual instances as needed.
[0110] In one implementation, as shown in Figure 2, the location of infrastructure can be described by cloud resource deployment regions (regions) and availability zones (AZs). Tenants can choose to deploy cloud services based on resources within specific regions and AZs. Regions are defined based on geographical location and network latency. Using the same resource pool within the same region can be understood as sharing public services such as elastic computing, block storage, object storage, virtual private cloud (VPC) networks, elastic internet protocol (EIP) addresses, and images. Regions are divided into general-purpose regions and dedicated regions. General-purpose regions provide general cloud services to public tenants. Dedicated regions are dedicated regions that host the same type of business or provide business services to specific tenants. A region typically includes multiple AZs. Multiple AZs within a region are connected via high-speed fiber optic cables to meet the needs of tenants building high-availability systems across AZs. Computing, network, and storage resources within an AZ are logically divided into multiple clusters.
[0111] Tenants can send instructions to the cloud management platform through their client 2 to create, manage, log in to, and operate virtual instances on the server, and use the cloud services provided by these virtual instances. For example, the cloud management platform can provide an access interface. This interface can be provided either as a user interface or an API. Tenants can operate their client to remotely access the access interface to register a cloud account and password on the cloud management platform, and then log in using these accounts and passwords. The cloud management platform can also authenticate the cloud account and password. After successful authentication, the tenant can further select and purchase a virtual instance with specific specifications (processor, memory, disk) on the cloud management platform. After the tenant successfully purchases the virtual instance, the cloud management platform provides the tenant with a remote login account and password for the purchased virtual instance. The tenant can use the remote login account and password to remotely log in to the virtual instance on their client, install and run their application within the virtual instance, and use the application to implement their business operations.
[0112] Client 2 can be selected from computers, personal computers, laptops, mobile phones, smartphones, tablets, cloud servers, portable mobile terminals, multimedia players, e-book readers, wearable devices, smart home appliances, artificial intelligence devices, smart wearable devices, smart in-vehicle devices, or Internet of Things devices, etc.
[0113] In one implementation, the cloud host modification method based on public cloud technology provided in this application embodiment can be implemented by running an executable program on a computing device in cloud system 1. When the cloud host modification method based on public cloud technology provided in this application embodiment is applied to a cloud management platform, the server used to implement the cloud management platform can implement the cloud host modification method based on public cloud technology provided in this application embodiment by running the executable program of the cloud host modification method based on public cloud technology provided in this application embodiment. Furthermore, the executable program implementing the cloud host modification method based on public cloud technology can optionally be presented in the form of an application installation package. After the server installs the application installation package, it can implement the cloud host modification method based on public cloud technology provided in this application embodiment by running the executable program therein.
[0114] It should be understood that the above content is an exemplary description of the implementation scenarios of the cloud host modification method based on public cloud technology provided in the embodiments of this application, and does not constitute a limitation on the implementation scenarios of the cloud host modification method based on public cloud technology. Those skilled in the art will know that as business needs change, the implementation scenarios can be adjusted according to application requirements, and the embodiments of this application do not specifically limit them. Furthermore, when the cloud host modification method based on public cloud technology provided in the embodiments of this application is applied to other scenarios, the executable program of the method can also be presented in the form of an application installation package or in other ways, and the embodiments of this application do not list them all.
[0115] The following describes the implementation process of a cloud host modification method based on public cloud technology, as provided in this application embodiment, executed by a cloud management platform. The cloud management platform manages the infrastructure providing cloud services. The infrastructure includes cloud hosts. Cloud hosts are used to deploy instances that implement tenant services. The cloud host modification method based on public cloud technology provided in this application embodiment mainly includes two parts. One part is the process of classifying the cloud host to be modified into multiple host classes. The other part is the process of modifying the cloud host according to its host class. The process of modifying the cloud host according to its host class will be described first, followed by the process of classifying the cloud host to be modified into multiple host classes.
[0116] Figure 3 is a flowchart illustrating how to modify a cloud host according to its host class, as provided in an embodiment of this application. As shown in Figure 3, the implementation process includes the following steps.
[0117] Step 301: Obtain N host classes, each host class including multiple cloud hosts to be changed.
[0118] The cloud management platform can use relevant measures to classify the M cloud hosts to be changed into N host classes. There are several ways to implement this, and examples of these implementations are provided below for ease of understanding.
[0119] Step 302: Select multiple target cloud hosts from the multiple cloud hosts included in each of the N host classes. The multiple target cloud hosts in the first host class are a part of the multiple cloud hosts included in the first host class, and the first host class is any one of the N host classes.
[0120] There are several ways to implement step 302. This application embodiment will illustrate it using the following two examples.
[0121] In the first implementation of step 302, as shown in Figure 4, the implementation process of step 302 includes steps 302a1 to 302a3.
[0122] Step 302a1: Obtain the importance of each cloud host in the P cloud hosts in the first host class to the tenant's business. The first host class is any one of the N host classes, and P is a positive integer.
[0123] In one possible implementation, the importance of a cloud server to a tenant's business is determined based on one or more of the following influencing factors: the cloud server's resource utilization, the number and level of tenants served by the cloud server, or the number and level of business applications deployed on the cloud server. Optionally, the importance of a cloud server to a tenant's business is positively correlated with the cloud server's resource utilization, the number of tenants served by the cloud server, the level of tenants served by the cloud server, the number of business applications deployed on the cloud server, and the level of business applications deployed on the cloud server. Furthermore, the degree of influence of the above factors on the importance of the cloud server to the tenant's business can vary. For example, the influence of the level of tenants served by the cloud server and the level of business applications deployed on the cloud server on the importance of the cloud server to the tenant's business is greater than the influence of the number of tenants served by the cloud server and the number of business applications deployed on the cloud server on the importance of the cloud server to the tenant's business. The influence of the number of tenants served by the cloud server and the number of business applications deployed on the cloud server on the importance of the cloud server to the tenant's business is greater than the influence of the cloud server's resource utilization on the importance of the cloud server to the tenant's business. The level of the business applications deployed on a cloud server has a greater impact on the importance of the cloud server to the tenant's business than the level of the tenants served by the cloud server. Similarly, the number of business applications deployed on a cloud server has a greater impact on the importance of the cloud server to the tenant's business than the number of tenants served by the cloud server. For example, when determining the importance of a cloud server to a tenant's business, one can first determine the weight of each influencing factor in the tenant's business importance, including the cloud server's resource utilization, the number and level of tenants served by the cloud server, and the number and level of business applications deployed on the cloud server. Then, obtain the statistical values for each influencing factor, including the cloud server's resource utilization, the number and level of tenants served by the cloud server, and the number and level of business applications deployed on the cloud server. Finally, calculate the weighted average of the weights of multiple influencing factors and the statistical values; this weighted average is the importance of the cloud server to the tenant's business.
[0124] Furthermore, considering that cloud servers are typically deployed according to regions and Availability Zones (AZs), and that multiple cloud servers deployed within the same region or AZ often share similar characteristics, host classes can also be divided by region or AZ. For example, when host classes are divided by region, a host class can include all cloud servers in one or more regions. Similarly, when host classes are divided by AZ, a host class can include all cloud servers in one or more AZs. In this case, the importance of cloud servers to the tenant's business can also be determined on a region or AZ basis. For example, when host classes are divided by region, if a host class includes cloud servers in multiple regions, the importance of each region to the tenant's business can be determined separately. Likewise, when host classes are divided by AZ, if a host class includes cloud servers in multiple AZs, the importance of each AZ to the tenant's business can be determined separately. Similarly, the impact of cloud servers in multiple other units included in a host class on the tenant's business can also be determined according to other units. When determining the impact of cloud hosts on tenant services within a unit based on regions, AZs, or other units, statistical values of one or more of the above-mentioned influencing factors can be obtained for each unit. Then, based on these statistical values, the impact of cloud hosts on tenant services within the corresponding unit can be determined. Please refer to the corresponding content above for the implementation method. Furthermore, when determining the impact of cloud hosts on tenant services within a unit, the influencing factors can also include the total number of cloud hosts in the unit, the total number of idle cloud hosts in the unit, and the total number of cloud hosts with a specified purpose in the unit. For example, a specified purpose can be a purpose that has a significant impact on tenant services. For instance, a cloud host with a specified purpose might be a business testing host. The importance of cloud hosts to tenant services is positively correlated with the total number of cloud hosts in the unit, the total number of cloud hosts with a specified purpose in the unit, and the total number of cloud hosts with a specified purpose in the unit.
[0125] It should be noted that the above-mentioned influencing factors are just one example of this application and are not intended to limit the factors affecting the importance of cloud servers to tenant businesses. They can also be determined based on other influencing factors, and this application does not provide examples of each of them.
[0126] Step 302a2: Divide the P cloud hosts into multiple change batches to be executed sequentially based on their importance, wherein the cloud hosts in the later change batches are more important than the cloud hosts in the earlier change batches.
[0127] After obtaining the importance of each of the P cloud hosts in the first host class to the tenant's business, the cloud management platform can divide the P hosts into multiple change batches to be executed sequentially based on this importance, so that changes can be made to the cloud hosts in the first host class in multiple change batches. Determining multiple change batches to be executed sequentially for the first host class is equivalent to determining a change pipeline for changing the P cloud hosts in the first host class; one change batch is equivalent to a ring in the change pipeline. In this application, there are several ways to implement dividing the P cloud hosts into multiple change batches based on importance. This application will illustrate several implementation methods as examples.
[0128] In the first implementation, if the importance level is determined for each cloud host in step 302a1, multiple cloud hosts with the same importance level or multiple cloud hosts with the same importance level that differ within a specified threshold can be divided into a change batch.
[0129] In the second implementation, if the importance level is determined for the unit where the cloud host is located in step 302a1, then cloud hosts in multiple units with the same importance level, or cloud hosts in multiple units with the same importance level but within a specified threshold, can be divided into a change batch.
[0130] For both the first and second implementation methods, the value of the specified threshold can be determined based on application requirements. For example, when the application has a higher tolerance for change anomalies, the specified threshold can be set higher; when the application has a lower tolerance for change anomalies, the specified threshold can be set lower. Furthermore, different change batches can use the same or different specified thresholds. For example, since earlier change batches are more likely to expose change anomalies, the specified threshold used for earlier change batches can be lower than the specified threshold used for later change batches.
[0131] For example, as shown in Figure 5, assume that the cloud management platform divides the multiple cloud hosts to be changed into the following 6 host classes based on the Availability Zone (AZ):
[0132] Host Class 1 includes: all hosts deployed in AZ1 of region 1, all hosts deployed in AZ1 of region 2, all hosts deployed in AZ1 of region 3, all hosts deployed in AZ1 of region 4, all hosts deployed in AZ1 of region 5, all hosts deployed in AZ1 of region 6, all hosts deployed in AZ1 of region 7, and all hosts deployed in AZ1 of region 8. Among these, the hosts deployed in AZ1 of region 1 and AZ1 of region 2 have the same level of importance to the tenant's business; the hosts deployed in AZ1 of region 3 and AZ1 of region 4 have the same level of importance to the tenant's business; the hosts deployed in AZ1 of region 5 and AZ1 of region 6 have the same level of importance to the tenant's business; and the hosts deployed in AZ1 of region 7 and AZ1 of region 8 have the same level of importance to the tenant's business. The importance of AZ1 of region 1, AZ1 of region 3, AZ1 of region 5, and AZ1 of region 7 increases sequentially to the tenant's business.
[0133] Host Class 2 includes: all hosts deployed in AZ2 of region 1, all hosts deployed in AZ2 of region 2, all hosts deployed in AZ2 of region 5, all hosts deployed in AZ2 of region 6, all hosts deployed in AZ2 of region 7, and all hosts deployed in AZ2 of region 8. Among these, the hosts deployed in AZ2 of region 1 and AZ2 of region 2 are of equal importance to the tenant's business; the hosts deployed in AZ2 of region 5 and AZ2 of region 6 are of equal importance to the tenant's business; and the hosts deployed in AZ2 of region 7 and AZ2 of region 8 are of equal importance to the tenant's business. The importance of AZ2 of region 1, AZ2 of region 5, and AZ2 of region 7 increases sequentially to the tenant's business.
[0134] Host class 3 includes: all hosts deployed in AZ2 of region 3, all hosts deployed in AZ2 of region 4, all hosts deployed in AZ3 of region 7, and all hosts deployed in AZ3 of region 8. Among these, all hosts deployed in AZ2 of region 3 and all hosts deployed in AZ2 of region 4 are of equal importance to the tenant's business. Similarly, all hosts deployed in AZ3 of region 7 and all hosts deployed in AZ3 of region 8 are of equal importance to the tenant's business. The importance of AZ2 of region 3 and AZ3 of region 7 increases sequentially to the tenant's business.
[0135] Host class 4 includes: all hosts deployed in AZ3 of region 3, all hosts deployed in AZ3 of region 4, all hosts deployed in AZ4 of region 7, and all hosts deployed in AZ4 of region 8. Among these, all hosts deployed in AZ3 of region 3 and all hosts deployed in AZ3 of region 4 are of equal importance to the tenant's business. Similarly, all hosts deployed in AZ4 of region 7 and all hosts deployed in AZ4 of region 8 are of equal importance to the tenant's business. The importance of AZ3 of region 3 and AZ4 of region 7 increases sequentially to the tenant's business.
[0136] Host class 5 includes all hosts deployed in AZ3 of region 5.
[0137] Host class 6 includes all hosts deployed in AZ4 of region 5.
[0138] Based on the importance of cloud servers to tenant businesses, the cloud management platform divides the cloud servers in each cloud server class into multiple change batches to be executed sequentially, as shown in Figure 5. The change batches for the six cloud server classes are as follows:
[0139] The cloud hosts in Host Class 1 are divided into four change batches to be executed sequentially. The first change batch modifies all hosts deployed in AZ1 of region 1 and all hosts deployed in AZ1 of region 2. The second change batch modifies all hosts deployed in AZ1 of region 3 and all hosts deployed in AZ1 of region 4. The third change batch modifies all hosts deployed in AZ1 of region 5 and all hosts deployed in AZ1 of region 6. The fourth change batch modifies all hosts deployed in AZ1 of region 7 and all hosts deployed in AZ1 of region 8.
[0140] The cloud hosts in Host Class 2 are divided into three batches of changes to be executed sequentially. The first batch of changes applies to all hosts deployed in AZ2 of region 1 and all hosts deployed in AZ2 of region 2. The second batch of changes applies to all hosts deployed in AZ2 of region 5 and all hosts deployed in AZ2 of region 6. The third batch of changes applies to all hosts deployed in AZ2 of region 7 and all hosts deployed in AZ2 of region 8.
[0141] The cloud hosts in host class 3 are divided into two batches of changes to be executed sequentially. The first batch of changes applies to all hosts deployed in AZ2 of region 3 and all hosts deployed in AZ2 of region 4. The second batch of changes applies to all hosts deployed in AZ3 of region 7 and all hosts deployed in AZ3 of region 8.
[0142] The cloud hosts in host class 4 are divided into two batches of changes to be executed sequentially. The first batch of changes applies to all hosts deployed in AZ3 of region 3 and all hosts deployed in AZ3 of region 4. The second batch of changes applies to all hosts deployed in AZ4 of region 7 and all hosts deployed in AZ4 of region 8.
[0143] Cloud hosts in host category 5 are classified into a single change batch. This change batch involves changes to all hosts deployed in AZ3 within region 5.
[0144] Cloud hosts in host category 6 are classified into a single change batch. This change batch involves changes to all hosts deployed in AZ4 of region 5.
[0145] Optionally, after dividing the cloud hosts in a host class into multiple change batches to be executed sequentially, the cloud management platform can also configure the execution strategy for the change batches in that host class. This is also known as configuring the execution strategy of the change pipeline. For example, the cloud management platform may also optionally perform one or more of the following operations: determine the start time and / or start conditions for each change batch in the first host class; determine the start time and / or start conditions for changes in the first host class; determine the start dependency order and / or start dependency conditions between different change batches in the first host class; determine the start dependency order and / or start dependency conditions between changes in different host classes; or, determine the sub-batch in the change batches of the first host class and the quantity relationship of the cloud hosts they include.
[0146] The start time of the change batch in the first host class is the moment when the change batch in the first host class begins execution. The start condition of the change batch in the first host class is the condition that needs to be met to start the execution of the change batch in the first host class. In one possible implementation, the start time and / or start condition of the change batch in the first host class can optionally be determined based on the operating load of the change batch in the first host class. For example, the start time and start condition of the change batch in the first host class need to ensure that the change process of the change batch avoids its peak business period. For example, the cloud management platform can determine the resource utilization rate of the cloud hosts in the change batch in multiple time periods in history, determine the operating load of multiple time periods based on the resource utilization rate of multiple time periods, and then determine the start time of the change batch based on the low business period indicated by the operating load of multiple time periods. As another example, the start condition of the change batch is that the business implemented by the cloud hosts included in the change batch is in a low business period.
[0147] The start time of the change to the first host class is the moment when the change to the first host class begins to be executed. The start condition of the change to the first host class is the condition that must be met to begin executing the change to the first host class. In one possible implementation, the start time and / or start condition of the change to the first host class can be determined based on the operating load of the cloud hosts included in the first host class. For example, the start time and start condition of the change to the first host class need to ensure that the change process of the first host class avoids its peak business periods. For example, the cloud management platform can determine the resource utilization rate of the cloud hosts included in the first host class in multiple time periods in history, determine the operating load of multiple time periods based on the resource utilization rate of multiple time periods, and then determine the start time of the change to the first host class based on the low business periods indicated by the operating load of multiple time periods. For another example, the start condition of the change to the first host class is that the business implemented by the cloud hosts included in the first host class is in a low business period. In another possible implementation, the start time of the change to the first host class is the start time of the first batch of changes executed in the first host class, and the implementation method is described in the corresponding description in the previous paragraph. The initiation condition for the change of the first host class is the initiation condition for the first batch of changes executed first in the first host class. Please refer to the corresponding description in the previous paragraph for its implementation.
[0148] The startup dependency order between the second change batch and the subsequent third change batch in the first host class refers to the dependency relationship between the execution status of the second change batch and the start time of the third change batch. The startup dependency condition between the second change batch and the subsequent third change batch in the first host class refers to the condition that the dependency relationship between the execution status of the second change batch and the start time of the third change batch must satisfy. In one possible implementation, the startup dependency order between the second change batch and the subsequent third change batch in the first host class can be determined based on the relationship between the services implemented by the cloud hosts in the second change batch and the services implemented by the cloud hosts in the third change batch. For example, if the relationship between the services implemented by the two indicates that the cloud hosts in the second change batch and the cloud hosts in the third change batch cannot perform changes simultaneously, then the startup dependency order and / or startup dependency condition of the second and third change batches need to ensure that the change times of the second and third change batches do not overlap. For example, if the relationship between the services implemented by the two cloud hosts indicates a very high degree of overlap between the services implemented by the cloud hosts in the second and third change batches, then the startup dependency order and / or startup dependency conditions of the second and third change batches need to ensure that the change process of the third change batch begins after the changes in the second change batch have been executed to a specified extent. For instance, the change process of the third change batch might begin execution after 10% of the changes in the second change batch have been executed. The second change batch can be any one of multiple change batches in the first host class. The third change batch is a change batch in the first host class that has a startup dependency order and / or startup dependency conditions with the second change batch. For example, the third change batch can be a change batch in the first host class that is adjacent to the second change batch, or a change batch that is not adjacent to the second change batch. When the third change batch is a change batch in the first host class that is adjacent to the second change batch, the startup dependency order between the second and third change batches can be represented by the change interval time between the second and third change batches, and the startup dependency conditions between the second and third change batches can be represented by the change interval conditions between the second and third change batches.
[0149] The startup dependency order between the first host class and the second host class refers to the dependency relationship between the change execution processes of the first host class and the second host class. The startup dependency condition between the first host class and the second host class refers to the condition that this dependency relationship must satisfy. The first host class and the second host class are any one of multiple host classes. In one possible implementation, the startup dependency order between the first host class and the second host class can be determined based on the relationship between the services implemented by the cloud hosts in the first host class and the services implemented by the cloud hosts in the second host class. For example, if the relationship between their implemented services indicates that the cloud hosts in the first host class and the cloud hosts in the second host class cannot perform changes simultaneously, then the startup dependency order and / or startup dependency condition of the first host class and the second host class must ensure that the change times of the first host class and the second host class do not overlap. As another example, if the relationship between their implemented services indicates that the services implemented by the cloud hosts in the first host class and the cloud hosts in the second host class are highly closely related, then the startup dependency order and / or startup dependency condition of the first host class and the second host class must ensure that the change process of the second host class begins after the change process of the first host class has reached a specified stage. For example, the change process for the second host class will begin after 10% of the changes for the first host class have been completed.
[0150] For any given change batch, the multiple hosts within that batch may begin implementing the changes all at once, or they may be implemented sequentially in multiple sub-batches. When multiple hosts in a change batch implement the changes in multiple sub-batches, the cloud management platform also needs to confirm the number of sub-batches and their constituent cloud hosts within each change batch of the first host class. In one possible implementation, the cloud management platform can divide all hosts included in the second change batch into multiple parts, with each part containing the cloud hosts included in a sub-batch, thereby obtaining the multiple sub-batches and their constituent cloud hosts within the second change batch. Here, the second change batch is any one of the multiple change batches within the first host class.
[0151] In another implementation, when determining the multiple sub-batches included in the second change batch and the number of cloud hosts they contain, the main consideration is that if a cloud host undergoing change first experiences a change anomaly, the anomaly will be detected earlier. The more cloud hosts with change anomalies, the greater the impact on the business. Therefore, the cloud management platform can optionally specify that the number of cloud hosts included in later sub-batches within the second change batch increases relative to the number of cloud hosts included in earlier sub-batches, in order to minimize the impact of potential change anomalies on the business. The method by which the number of cloud hosts included in later sub-batches increases relative to the number of cloud hosts included in earlier sub-batches can be determined according to application requirements. For example, the first sub-batchens in the second change batch may include one cloud host, the number of cloud hosts in the remaining sub-batches may increase by a factor of 2, and the number of cloud hosts in a single sub-batchens may not exceed 100. Since the first sub-batchens is the first batch to undergo change in the second change batch, the cloud hosts included in the first sub-batchens are also called OneBox cloud hosts.
[0152] Furthermore, when determining the multiple sub-batches that will be executed sequentially within the second change batch, further consideration can be given to the first sub-batch that is executed first, in order to further control the impact of the changes within a safe range. For example, the first sub-batch may optionally include one or more of the following cloud hosts:
[0153] 1) Idle cloud servers, and the total number of idle cloud servers in the first sub-batch accounts for no more than the proportion of the total number of cloud servers in the second change batch. Idle cloud servers are cloud servers that have not deployed instances for implementing tenant services.
[0154] 2) Cloud servers belonging to a specified tenant type, where the total number of cloud servers belonging to the specified tenant type in the first sub-batch accounts for no more than the corresponding second percentage in the total number of cloud servers in the second change batch. The second percentage for the specified tenant type is negatively correlated with the importance of the specified tenant type. Cloud service providers have various tenant types. For example, based on the importance of the implemented business, cloud service providers may have ordinary tenants and important tenants; in this case, the specified tenant type is either an ordinary tenant or an important tenant. If the importance of an ordinary tenant is less than that of an important tenant, then the second percentage for ordinary tenants is greater than that for important tenants.
[0155] 3) Cloud servers used for business testing (also known as business test cloud servers), and the total number of cloud servers used for business testing in the first sub-batch shall not exceed the proportion of the total number of cloud servers in the second change batch in the third sub-batch. Business test cloud servers are cloud servers that run test instances for certain testing purposes.
[0156] In this process, the priority for changing idle cloud servers in the first sub-batch decreases sequentially: the priority for changing cloud servers belonging to a specific tenant type in the first sub-batch, and the priority for changing cloud servers used for business testing. Furthermore, the priority for changing cloud servers belonging to a specific tenant type in the first sub-batch is negatively correlated with the importance of that tenant type. For example, assuming the second change batch includes idle cloud servers, cloud servers belonging to ordinary tenants, cloud servers belonging to important tenants, and business testing cloud servers, then the priority for changing idle cloud servers, cloud servers belonging to ordinary tenants, cloud servers belonging to important tenants, and business testing cloud servers in the first sub-batch decreases sequentially.
[0157] Furthermore, the values of the first, second, and third percentages can be determined based on application requirements. For example, the values of the first, second, and third percentages are determined based on the application's tolerance for changes or anomalies to the corresponding type of cloud server. For instance, the first percentage might be 5%, the second percentage for ordinary tenants might be 3%, the second percentage for important tenants might be 1%, and the third percentage might be 10%.
[0158] Optionally, in addition to this, the cloud servers included in the first sub-batch also meet one or more of the following criteria:
[0159] 1) For multiple cloud hosts with the same priority that are changed in the first sub-batch, the first sub-batch includes: cloud hosts serving fewer tenants, and / or cloud hosts with fewer instances deployed, with priority given to cloud hosts serving fewer tenants. This condition is equivalent to requiring that, among multiple cloud hosts with the same priority that are changed in the first sub-batch, cloud hosts serving fewer tenants and / or cloud hosts with fewer instances deployed be selected. When multiple cloud hosts have the same priority for changes in the first sub-batch, and cloud hosts serving fewer tenants are selected, if an anomaly occurs in the change, it can be guaranteed that fewer tenants will be affected, and the impact of the anomaly can be controlled from the perspective of the number of tenants. When multiple cloud hosts have the same priority for changes in the first sub-batch, and cloud hosts with fewer instances are selected, if an anomaly occurs in the change, it can be guaranteed that fewer instances will be affected, and the impact of the anomaly can be controlled from the perspective of the number of instances. Compared to cloud servers with fewer instances deployed, prioritizing cloud servers that serve fewer tenants ensures that fewer instances are affected if changes occur, thus minimizing the impact of changes on business operations and controlling the scope of impact from the perspective of reducing business disruptions.
[0160] 2) For multiple cloud hosts with services deployed under a specified tenant, the total number of cloud hosts with services deployed under the specified tenant in the first sub-batch is less than a first quantity threshold. By controlling the total number of cloud hosts with services deployed under the specified tenant in the first sub-batch, if an anomaly occurs during a change, the total number of cloud hosts with services deployed under the specified tenant affected by the anomaly can be controlled, thus keeping the impact of the anomaly on the services of the same tenant within a controllable range. The value of the first quantity threshold can be determined according to application requirements. For example, the value of the first quantity threshold is determined based on the application's tolerance for anomalies during changes to the corresponding type of cloud host. For example, the value of the first quantity threshold is 2.
[0161] Similarly, in other sub-batches within the second change batch, the cloud management platform can also control the total number of cloud hosts with services deployed on a specified tenant included in these other sub-batches to be less than a second quantity threshold. Other sub-batches are those sub-batches other than the first sub-batchens of the second change batch. By controlling the total number of cloud hosts with services deployed on a specified tenant included in these other sub-batches, if an anomaly occurs in the change, the total number of cloud hosts with services deployed on a specified tenant affected by the anomaly can be controlled, keeping the impact of the anomaly on the services of the same tenant within a manageable range. The value of the second quantity threshold can be determined according to application requirements. For example, the value of the second quantity threshold is determined based on the application's tolerance for anomalies in the corresponding type of cloud hosts. For instance, the second quantity threshold might be half the total number of cloud hosts with services deployed on a specified tenant.
[0162] For example, as shown in Figure 5, for host class 2, the cloud management platform makes changes to the cloud hosts in Regions 7 and 8 in Ring 3. This change process includes four sub-batches executed sequentially. The first sub-batch, executed first, changes the OneBox cloud hosts deployed in AZ1 of Region 7 (as shown in Region 7 AZ1-Box in Figure 5) and the OneBox cloud hosts deployed in AZ1 of Region 8 (as shown in Region 8 AZ1 in Figure 5). The second sub-batch, executed next, changes the remaining cloud hosts deployed in AZ1 of Region 7 (as shown in Region 7 AZ1 in Figure 5) and the remaining cloud hosts deployed in AZ1 of Region 8 (as shown in Region 8 AZ1 in Figure 5). The third sub-batch, executed next, changes the OneBox cloud hosts deployed in AZ2 of Region 7 (as shown in Region 7 AZ2-Box in Figure 5) and the OneBox cloud hosts deployed in AZ2 of Region 8 (as shown in Region 8 AZ2-Box in Figure 5). The fourth sub-batch executed next will make changes to the remaining cloud hosts deployed in AZ2 of Region 7 (as shown in Figure 5, Region 7 AZ2) and the remaining cloud hosts deployed in AZ2 of Region 8 (as shown in Figure 5, Region 8 AZ2).
[0163] It should be noted that the above descriptions of startup time, startup conditions, startup dependency order, startup dependency conditions, sub-batch in each change batch in the first host class and the number of cloud hosts included therein, and the conditions that the cloud hosts included in the first sub-batch need to meet are all examples of possible implementation methods and are not intended to limit the implementation method. There may be other implementation methods, but this application embodiment does not provide examples of them one by one.
[0164] Step 302a3: Determine the cloud host in the first change batch among multiple change batches as the target cloud host selected from the cloud hosts included in the first host class.
[0165] The cloud management platform divides the P cloud hosts in the first host class into multiple change batches to be executed sequentially based on the importance of the cloud hosts to the tenant's business. Then, the cloud host in the first change batch among the multiple change batches can be identified as the target cloud host selected from the cloud hosts included in the first host class.
[0166] In the second implementation of step 302, when the M cloud hosts are deployed in multiple cloud resource deployment regions, as shown in Figure 6, the implementation process of step 302 includes steps 302b1 to 302b3. Alternatively, as shown in Figure 7, the implementation process of step 302 includes steps 302b1 to 302b5.
[0167] Step 302b1: Determine the distribution of N host classes across multiple cloud resource deployment regions.
[0168] The configuration information of the cloud management platform contains information about the infrastructure it manages. This information includes indications of the infrastructure's deployment status. For example, this information might include: the data centers included in the infrastructure, the cloud resource deployment regions included in the data centers, the availability zones included in the cloud resource deployment regions, and the distribution of cloud hosts within the availability zones. Based on this configuration information, the cloud management platform can determine the distribution of all cloud hosts managed by the platform across the cloud resource deployment regions. By statistically analyzing the cloud hosts included in each host class and the distribution of all cloud hosts across the cloud resource deployment regions, the cloud management platform can obtain the distribution of cloud hosts for each host class across multiple cloud resource deployment regions, thus obtaining the distribution of N host classes across multiple cloud resource deployment regions.
[0169] Table 1
[0170] For example, Table 1 shows the distribution of N host classes across multiple cloud resource deployment regions, calculated by the cloud management platform based on configuration information. As shown in Table 1, taking cloud hosts deployed in region 1 as an example, 20% of the cloud hosts are host class 1 hosts deployed in AZ1 of region 1, 30% of the cloud hosts are host class 2 hosts deployed in AZ1 of region 1, 25% of the cloud hosts are host class 1 hosts deployed in AZ2 of region 1, and 25% of the cloud hosts are host class 2 hosts deployed in AZ2 of region 1.
[0171] Step 302b2: Based on the distribution, determine the minimum coverage of N host classes in multiple cloud resource deployment areas. The minimum coverage includes N host classes.
[0172] As shown in Figure 8, the implementation process of step 302b2 includes:
[0173] Step b21: In the candidate cloud resource deployment areas used to determine the minimum coverage, determine the second cloud resource deployment area combination based on the combination of candidate cloud resource deployment areas including all host classes.
[0174] The minimum coverage area includes N host classes, meaning all host classes deployed within the minimum coverage area. After obtaining the distribution of the N host classes across multiple cloud resource deployment regions, the cloud management platform can obtain a combination of cloud resource deployment regions that deploy all host classes from the N host classes, resulting in a second cloud resource deployment region combination. This second cloud resource deployment region combination may consist of one or multiple cloud resource deployment regions; collectively, it is referred to here as the second cloud resource deployment region combination. Since the minimum coverage area is determined within multiple cloud resource deployment regions deploying M cloud hosts, the candidate cloud resource deployment regions are essentially multiple cloud resource deployment regions deploying M cloud hosts. Therefore, the combination of candidate cloud resource deployment regions including the most host classes is essentially a combination of candidate cloud resource deployment regions including N host classes.
[0175] It should be noted that there may be one or more combinations of cloud resource deployment regions that deploy all host classes across N host classes. When there is only one combination of cloud resource deployment regions that deploy all host classes across N host classes, the cloud management platform determines this combination as the second cloud resource deployment region combination. When there are multiple combinations of cloud resource deployment regions that deploy all host classes across N host classes, the cloud management platform can select one of them as the second cloud resource deployment region combination based on application requirements. For example, according to the deployment situation shown in Table 1, regions 1 / 2 / 3 / 4 / 5 / 6 / 7 / 8 all deploy cloud hosts from host class 1, regions 1 / 2 / 5 / 6 / 7 / 8 all deploy cloud hosts from host class 2, regions 3 / 4 / 7 / 8 all deploy cloud hosts from host class 3, regions 3 / 4 / 7 / 8 all deploy cloud hosts from host class 4, Region 5 deploys cloud hosts from host class 5, and Region 5 deploys cloud hosts from host class 6. Therefore, the combination of cloud resource deployment regions deploying 6 host classes is the combination of the candidate cloud resource deployment regions listed above, each deploying 6 host classes. At this point, there are multiple combinations of cloud resource deployment regions that deploy 6 host classes.
[0176] Optionally, the cloud management platform may select one of multiple combinations of cloud resource deployment regions that deploy all host classes of N host classes as the second cloud resource deployment region combination, based on one or more of the following conditions: That is, the second cloud resource deployment region combination also satisfies one or more of the following three conditions:
[0177] 1) Among multiple combinations of candidate cloud resource deployment regions including all host classes, the one with the smallest total number of cloud resource deployment regions. When selecting a second cloud resource deployment region combination based on this condition, and then selecting a target cloud host based on the selected second cloud resource deployment region combination, this is equivalent to selecting a target cloud host from cloud hosts that include all host classes and whose deployment scope covers the fewest cloud resource deployment regions. In this way, if an anomaly occurs, the affected cloud resource deployment regions are minimized, thus controlling the risk of changes to the fewest possible cloud resource deployment regions.
[0178] 2) Among multiple combinations of candidate cloud resource deployment regions encompassing all host classes, a cloud host with the lowest importance is deployed. The cloud host with the lowest importance can be the cloud host with the lowest overall importance in the combination, which is determined based on the importance of all cloud hosts deployed within the combination. Alternatively, the cloud host with the lowest importance can be a single cloud host with the lowest importance. When a second cloud resource deployment region combination is selected from multiple combinations based on this condition, and a target cloud host is selected based on the selected second cloud resource deployment region combination, this is equivalent to selecting the target cloud host from a group of cloud hosts with the lowest importance across all host classes. In this way, if an anomaly occurs, the anomaly only affects the cloud host with the lowest importance to the tenant's business, thus reducing the impact of the anomaly on the tenant's business. The method for determining the importance of a cloud host to the tenant's business is described in step 302a1, and will not be repeated here.
[0179] 3) Among multiple combinations of candidate cloud resource deployment regions that include all host types, the minimum total number of cloud hosts is deployed. When selecting a second cloud resource deployment region combination based on this condition, and then selecting the target cloud host based on the selected second cloud resource deployment region combination, it is equivalent to selecting the target cloud host from the cloud resource deployment region combination that includes all host types and has the minimum number of deployed cloud hosts. In this way, if an anomaly occurs, the impact of the anomaly can be controlled within the minimum number of cloud hosts, reducing the scope of the impact of the anomaly.
[0180] In one possible implementation, when selecting a second cloud resource deployment region combination from multiple candidate cloud resource deployment region combinations based on at least two of the above conditions, the cloud management platform may first obtain the score for each condition corresponding to the combination of multiple candidate cloud resource deployment regions, as well as the weight of each condition. Then, for any candidate cloud resource deployment region combination, based on the score and the weight of the condition, it obtains the weighted score of that candidate cloud resource deployment region combination. The candidate cloud resource deployment region combination with the highest weighted score is then determined as the second cloud resource deployment region combination. Specifically, for any candidate cloud resource deployment region combination, the total number of cloud resource deployment regions it has, the importance of the deployed cloud hosts to the tenant's business, and the total number of deployed cloud hosts are negatively correlated with the score of that candidate cloud resource deployment region combination. The weights of the above three conditions can be determined according to application requirements; for example, they can be determined based on the impact of changes or anomalies on the corresponding conditions. For example, the weights of the above three conditions are 0.2, 0.5, and 0.3, respectively.
[0181] Step b22: Determine the minimum coverage area based on the combination of deployment regions for the second cloud resources.
[0182] Once the cloud management platform determines the combination of deployment regions for the second cloud resources, it can define this combination as the minimum coverage area.
[0183] Optionally, when each cloud resource deployment region includes multiple availability zones, the granularity of the minimum coverage can be refined to the availability zone level when determining the minimum coverage. In one possible implementation, as shown in Figure 9, the minimum coverage is determined based on the second cloud resource deployment region combination, including:
[0184] Step b221: In the second cloud resource deployment area combination, determine the target availability zone combination based on the combination of availability zones that include the most host types.
[0185] When each cloud resource deployment region includes multiple availability zones, after determining the second cloud resource deployment region combination, the cloud management platform can optionally further determine the combination of availability zones including the most host classes based on the deployment status of cloud hosts in the multiple availability zones of the second cloud resource deployment region combination. Then, based on the combination of availability zones including the most host classes, the target availability zone combination is determined. According to the configuration information, the cloud management platform can obtain the distribution of all cloud hosts managed by the cloud management platform in the availability zones of the cloud resource deployment region. Then, based on this distribution, it obtains the combination of availability zones including the most host classes among all host classes deployed in the cloud resource deployment region, and determines the combination of availability zones including the most host classes as the target availability zone combination.
[0186] It should be noted that there may be one or more combinations of Availability Zones (AZs) that include the most host class. When there is only one combination of AZs that includes the most host class, the cloud management platform will select that combination as the target Availability Zone combination. When there are multiple combinations of AZs that include the most host class, the cloud management platform can select one of them as the target Availability Zone combination based on application requirements.
[0187] Optionally, the cloud management platform may select one of multiple combinations of availability zones, including the one with the most host classes, as the target availability zone combination based on one or more of the following conditions: That is, the target availability zone combination also satisfies one or more of the following three conditions:
[0188] 1) Among multiple availability zone combinations that include the most host types in the second cloud resource deployment region combination, the one with the smallest total number of availability zones is selected. When selecting a target availability zone combination from multiple combinations based on this condition, and then selecting a target cloud host based on the selected target availability zone combination, it is equivalent to selecting a target cloud host from cloud hosts that include the most host types and whose deployment scope covers the fewest availability zones. In this way, if a change anomaly occurs, the availability zones affected by the change anomaly are minimized, and the change risk can be controlled within the fewest availability zones.
[0189] 2) In the second cloud resource deployment area combination, among multiple availability zone combinations that include the most host types, the cloud host with the lowest importance is deployed. When selecting a target availability zone combination from multiple combinations based on this condition, and then selecting the target cloud host based on the selected target availability zone combination, it is equivalent to selecting the target cloud host from the group of cloud hosts that include the most host types and have the lowest importance. In this way, if a change anomaly occurs, the change anomaly will only affect the cloud host with the lowest importance to the tenant's business, thus reducing the impact of the change anomaly on the tenant's business. The method for determining the importance of the cloud host to the tenant's business is described in step 302a1, and will not be repeated here.
[0190] 3) Among the multiple availability zone combinations that include the most host types in the second cloud resource deployment region combination, the minimum total number of cloud hosts is deployed. When selecting a target availability zone combination from multiple combinations based on this condition, and then selecting the target cloud host based on the selected target availability zone combination, it is equivalent to selecting the target cloud host from the cloud resource deployment region combination that includes the most host types and has the fewest deployed cloud hosts. In this way, if an anomaly occurs, the impact of the anomaly can be controlled within the minimum total number of cloud hosts, reducing the scope of the impact of the anomaly.
[0191] In one possible implementation, when selecting a target availability zone combination from multiple candidate availability zone combinations based on at least two of the above conditions, the cloud management platform may first obtain the score for each condition corresponding to the multiple candidate availability zone combinations, as well as the weight of each condition. Then, for any candidate availability zone combination, based on its score and the weight of the condition, it obtains a weighted score for that candidate availability zone combination. The candidate availability zone combination with the highest weighted score is then determined as the target availability zone combination. For any candidate availability zone combination, the total number of availability zones it has, the importance of the deployed cloud hosts to the tenant's business, and the total number of deployed cloud hosts are negatively correlated with the score of that candidate availability zone combination. The weights of the above three conditions can be determined according to application requirements; for example, they can be determined based on the impact of changes or anomalies on the corresponding conditions. For example, the weights of the above three conditions are 0.2, 0.5, and 0.3, respectively.
[0192] Step b222: Combine the target availability zones to determine the minimum coverage area.
[0193] Once the target availability zone combination is determined, the cloud management platform can define the target availability zone combination as the minimum coverage area.
[0194] It should be noted that the above examples of cloud servers being deployed according to cloud resource deployment regions or availability zones within those regions, while underestimating the minimum coverage, are merely illustrative and not intended to limit the implementation method for determining the minimum coverage in this application. For instance, when cloud servers are deployed at other granularities, the minimum coverage can also be determined by referring to the method of deploying cloud servers at those other granularities; however, this application does not provide specific examples of such methods.
[0195] Step 302b3: Based on the cloud hosts included in the minimum coverage area, obtain the target cloud hosts selected from N host classes.
[0196] Once the cloud management platform determines the minimum coverage of N host classes, it can select target cloud hosts from these N host classes based on the cloud hosts included in that minimum coverage. For example, the cloud management platform selects the cloud hosts included in the minimum coverage as the target cloud hosts from the N host classes. For instance, as shown in Table 1, if the target availability zone combination determined by the cloud management platform is Region3 AZ2 + Region5 AZ1, then Region3 AZ2 + Region5 AZ1 can be determined as the minimum coverage of 6 host classes, and the cloud hosts deployed in AZ2 of Region3 and AZ1 of Region5 can be selected as the target cloud hosts from the 6 host classes.
[0197] Optionally, to ensure coverage of the target cloud hosts, the target cloud hosts can also be selected from the N host classes based on other factors. In one possible implementation, as shown in Figure 7, selecting the target cloud host from the cloud hosts included in each of the N host classes further includes:
[0198] Step 302b4: Based on the distribution, determine the extension range of the minimum coverage area in multiple cloud resource deployment areas. The union of the cloud resource deployment areas where the cloud hosts are located, including the minimum coverage area and the extension range, constitutes multiple cloud resource deployment areas.
[0199] When the union of the cloud resource deployment regions where the cloud hosts included in the minimum coverage area and the extended coverage area are located constitutes multiple cloud resource deployment regions, and the cloud hosts included in the minimum coverage area and the extended coverage area are selected as the target cloud hosts, it is equivalent to selecting the target cloud hosts in all the multiple cloud resource deployment regions where the cloud hosts to be changed are located. When the target cloud hosts are changed in the first change batch, it is equivalent to changing them across the entire range of multiple cloud resource deployment regions where multiple cloud hosts to be changed are located. This ensures that the representative range of the cloud hosts to be changed first is included, achieving coverage and immersion verification of the entire deployment range. It is also more helpful to discover problems in advance and control the impact of changes within a safe range.
[0200] In one possible implementation, as shown in Figure 10, the implementation process of step 302b4 includes:
[0201] Step b41: Determine the first cloud resource deployment area combination where the cloud hosts included in the minimum coverage area are located. The first cloud resource deployment area combination includes one or more target cloud resource deployment areas.
[0202] As described above, the first cloud resource deployment region here is the same as the aforementioned combination of second cloud resource deployment regions. The cloud resource deployment regions included in the combination of second cloud resource deployment regions are the target cloud resource deployment regions. For example, assuming the minimum deployment range is Region3 AZ2 + Region5 AZ1, then the multiple target cloud resource deployment regions are Region3 and Region5.
[0203] Step b42: In cloud resource deployment regions with an importance level lower than the lowest importance level in one or more target cloud resource deployment regions, determine the preceding extension range of the minimum coverage range, wherein the change order of hosts included in the preceding extension range is prior to the change order of hosts included in the minimum coverage range.
[0204] Cloud resource deployment regions whose importance is lower than the lowest importance among one or more target cloud resource deployment regions are equivalent to candidate cloud resource deployment regions for determining the preceding extension range. Step b42 is equivalent to determining the preceding extension range among the candidate cloud resource deployment regions. Its implementation process includes: among the candidate cloud resource deployment regions used to determine the preceding extension range, a second cloud resource deployment region combination is determined based on the combination of candidate cloud resource deployment regions including the most host types; then, the preceding extension range is determined based on the second cloud resource deployment region combination. Here, the second cloud resource deployment region combination is the preceding extension range. The importance of a cloud resource deployment region can be obtained based on the importance of the cloud hosts deployed in the cloud resource deployment region to the tenant's business. For example, the importance of a cloud resource deployment region can be the average, maximum, or minimum importance of all cloud hosts deployed in the cloud resource deployment region, etc., which is not specifically limited in this embodiment. For the method of determining the importance of cloud hosts to the tenant's business, please refer to the relevant description in step 302a1, which will not be repeated here.
[0205] As described in step b21 above, after obtaining the distribution of N host classes across multiple cloud resource deployment areas, the cloud management platform can, based on this distribution, obtain a combination of candidate cloud resource deployment areas that includes the most host classes from the candidate cloud resource deployment areas within the preceding extended range, thus obtaining a second cloud resource deployment area combination. Furthermore, the second cloud resource deployment area combination may optionally satisfy one or more of the following conditions: among the multiple combinations of candidate cloud resource deployment areas including the most host classes, it has the minimum total number of cloud resource deployment areas; among the multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys cloud hosts of the lowest importance; or, among the multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys the minimum total number of cloud hosts. For the implementation method of determining the preceding extended range at the cloud resource deployment area granularity, please refer to the previous description of determining the minimum coverage range; it will not be elaborated upon here.
[0206] When each cloud resource deployment region includes multiple availability zones, the granularity of the preceding extension range can be refined to the availability zone level when determining the preceding extension range. In one possible implementation, step b42 further includes: determining a target availability zone combination based on the combination of availability zones with the most host classes in the second cloud resource deployment region combination; and defining the target availability zone combination as the preceding extension range. Optionally, the target availability zone combination may also satisfy one or more of the following conditions: having the minimum total number of availability zones among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination; having cloud hosts of the lowest importance level deployed among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination; or having the minimum total number of cloud hosts deployed among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination. For the implementation method of determining the preceding extension range at the availability zone level, please refer to the relevant description of determining the minimum coverage range above; it will not be elaborated here.
[0207] For example, assume the minimum deployment range is Region3AZ2 + Region5AZ1, multiple target cloud resource deployment regions are Region3 and Region5, and the cloud resource deployment region with the lowest importance among the multiple target cloud resource deployment regions is Region3. Among all cloud resource deployment regions with M cloud hosts, the cloud resource deployment regions with an importance lower than Region3 are Region1 and Region2. Based on the cloud host deployment information recorded in Table 1, assuming the preceding extension range is determined at the availability zone granularity according to the condition of having the minimum total number of cloud hosts deployed, the preceding extension range can be obtained as Region2AZ1.
[0208] Step b43: In the cloud resource deployment region with an importance higher than the highest importance of one or more target cloud resource deployment regions, determine the subsequent extension range of the minimum coverage, wherein the change order of the hosts included in the subsequent extension range is after the change order of the hosts included in the minimum coverage.
[0209] The cloud resource deployment region whose importance is higher than the highest importance among one or more target cloud resource deployment regions is equivalent to a candidate cloud resource deployment region for determining the subsequent extension range. Step b43 is equivalent to determining the subsequent extension range among the candidate cloud resource deployment regions. Its implementation process includes: among the candidate cloud resource deployment regions used to determine the subsequent extension range, a second cloud resource deployment region combination is determined based on the combination of candidate cloud resource deployment regions including the most host types, and then the subsequent extension range is determined based on the second cloud resource deployment region combination. Here, the second cloud resource deployment region combination is the subsequent extension range. The importance of the cloud resource deployment region can be obtained based on the importance of the cloud hosts deployed in the cloud resource deployment region to the tenant's business. For example, the importance of the cloud resource deployment region is the average, maximum, or minimum value of the importance of all cloud hosts deployed in the cloud resource deployment region, etc., which is not specifically limited in this embodiment. For the method of determining the importance of cloud hosts to the tenant's business, please refer to the relevant description in step 302a1, which will not be repeated here.
[0210] As described in step b21 above, after obtaining the distribution of N host classes across multiple cloud resource deployment areas, the cloud management platform can, based on this distribution, determine the combination of candidate cloud resource deployment areas that includes the most host classes within the subsequent extended range, thus obtaining a second cloud resource deployment area combination. Furthermore, the second cloud resource deployment area combination may optionally satisfy one or more of the following conditions: among the multiple combinations of candidate cloud resource deployment areas including the most host classes, it has the minimum total number of cloud resource deployment areas; among the multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys cloud hosts of the lowest importance; or, among the multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys the minimum total number of cloud hosts. For the implementation method of determining the subsequent extended range at the cloud resource deployment area granularity, please refer to the previous description of determining the minimum coverage range; it will not be elaborated upon here.
[0211] When each cloud resource deployment region includes multiple availability zones, the granularity of the subsequent extension range can be refined to the availability zone level when determining the subsequent extension range. In one possible implementation, step b43 further includes: determining a target availability zone combination based on the combination of availability zones with the most host classes in the second cloud resource deployment region combination; and determining the target availability zone combination as the subsequent extension range. Optionally, the target availability zone combination may also satisfy one or more of the following conditions: having the minimum total number of availability zones among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination; deploying cloud hosts of the lowest importance level among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination; or deploying the minimum total number of cloud hosts among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination. For the implementation method of determining the subsequent extension range at the availability zone level, please refer to the relevant description of determining the minimum coverage range above; it will not be elaborated here.
[0212] For example, assume the minimum deployment range is Region3AZ2 + Region5AZ1, multiple target cloud resource deployment regions are Region3 and Region5, and the cloud resource deployment region with the highest importance among these multiple target cloud resource deployment regions is Region5. Among all cloud resource deployment regions with M cloud hosts, the cloud resource deployment regions with a higher importance than Region5 are Region7 and Region8. Based on the cloud host deployment information recorded in Table 1, assuming the subsequent extension range is determined at the availability zone level according to the condition of having the minimum total number of cloud hosts deployed, the subsequent extension range can be obtained as Region7AZ1.
[0213] Step b44: When the first cloud resource deployment area combination includes multiple target cloud resource deployment areas, for any two target cloud resource deployment areas with adjacent importance, and among cloud resource deployment areas with importance between the importance of any two target cloud resource deployment areas, determine the intermediate extension range with the minimum coverage. The change order of hosts included in the intermediate extension range is between the change order of hosts included in any two target cloud resource deployment areas.
[0214] A cloud resource deployment area whose importance falls between any two adjacent target cloud resource deployment areas is equivalent to a candidate cloud resource deployment area for determining an intermediate extension range for those two target cloud resource deployment areas. Step b44 is equivalent to determining an intermediate extension range among the candidate cloud resource deployment areas. Its implementation process includes: among the candidate cloud resource deployment areas used to determine the intermediate extension range, a second cloud resource deployment area combination is determined based on the combination of candidate cloud resource deployment areas including the most host types; then, the intermediate extension range is determined based on the second cloud resource deployment area combination. Here, the second cloud resource deployment area combination is the intermediate extension range. The importance of a cloud resource deployment area can be obtained based on the importance of the cloud hosts deployed in the cloud resource deployment area to the tenant's business. For example, the importance of a cloud resource deployment area can be the average, maximum, or minimum importance of all cloud hosts deployed in the cloud resource deployment area, etc., which is not specifically limited in this embodiment. For the method of determining the importance of cloud hosts to the tenant's business, please refer to the relevant description in step 302a1, which will not be repeated here.
[0215] As described in step b21 above, after obtaining the distribution of N host classes across multiple cloud resource deployment areas, the cloud management platform can, based on this distribution, determine the combination of candidate cloud resource deployment areas containing the most host classes within the intermediate extension range, thus obtaining a second cloud resource deployment area combination. Furthermore, the second cloud resource deployment area combination may optionally satisfy one or more of the following conditions: among the multiple combinations of candidate cloud resource deployment areas containing the most host classes, it has the minimum total number of cloud resource deployment areas; among the multiple combinations of candidate cloud resource deployment areas containing the most host classes, it deploys cloud hosts of the lowest importance; or, among the multiple combinations of candidate cloud resource deployment areas containing the most host classes, it deploys the minimum total number of cloud hosts. For the implementation method of determining the intermediate extension range at the cloud resource deployment area granularity, please refer to the previous description of determining the minimum coverage range; it will not be elaborated upon here.
[0216] When each cloud resource deployment region includes multiple availability zones, the granularity of the intermediate extension range can be refined to the availability zone level when determining the intermediate extension range. In one possible implementation, step b44 further includes: determining a target availability zone combination based on the combination of availability zones with the most host classes in the second cloud resource deployment region combination; and defining the target availability zone combination as the intermediate extension range. Optionally, the target availability zone combination may also satisfy one or more of the following conditions: having the minimum total number of availability zones among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination; deploying cloud hosts of the lowest importance level among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination; or deploying the minimum total number of cloud hosts among the multiple availability zone combinations with the most host classes in the second cloud resource deployment region combination. For the implementation method of determining the intermediate extension range at the availability zone level, please refer to the previous description of determining the minimum coverage range; it will not be elaborated here.
[0217] For example, suppose the minimum deployment range is Region3AZ2 + Region5AZ1, and multiple target cloud resources are deployed in Region3 and Region5, with Region3 having a lower importance than Region5. In all cloud resource deployment regions with M cloud hosts, there is no region with an importance between Region3 and Region5; therefore, there is no intermediate extension of the minimum coverage range Region3AZ2 + Region5AZ1.
[0218] Step 302b5: Based on the cloud hosts included in the minimum coverage and extended coverage, obtain the target cloud hosts selected from N host classes.
[0219] After obtaining the preceding, following, and intermediate extensions of the minimum coverage area, the cloud management platform can determine the shortest change path for modifying cloud hosts across all cloud resource deployment regions with M cloud hosts. The shortest change path is the path obtained by concatenating the minimum coverage area, its preceding, following, and intermediate extensions in the order of modification. For example, in the example above, the minimum coverage area for the six host classes is Region3 AZ2 + Region5 AZ1, its preceding extension is Region2 AZ1, its following extension is Region7 AZ1, and there are no intermediate extensions. Therefore, the shortest change path is: Region2 AZ1 + Region3 AZ2 + Region5 AZ1 + Region7 AZ1. Then, based on the cloud hosts included in the shortest change path, the cloud management platform can select the target cloud host from N host classes. For example, the cloud management platform selects the cloud hosts included in the shortest change path as the target cloud host from the N host classes. It should be noted that the processes of determining the scope of coverage and identifying the target cloud host based on the scope of coverage are optional. When these processes are not executed, the minimum coverage area is the shortest change path.
[0220] Furthermore, in the example above for determining the preceding and subsequent extension ranges, the preceding extension range has only two candidate cloud resource deployment regions, and these two regions have the same level of importance. Therefore, the resulting preceding extension range includes one cloud resource deployment region. However, when the multiple candidate cloud resource deployment regions in the preceding extension range have different levels of importance, it is necessary to iterate through these regions according to the above implementation process to obtain a preceding extension range that includes multiple cloud resource deployment regions, each with a different level of importance. Similarly, the subsequent extension range has only two candidate cloud resource deployment regions, and these regions have the same level of importance. Therefore, the resulting subsequent extension range includes one cloud resource deployment region. At this point, when the multiple candidate cloud resource deployment regions in the subsequent extension range have different levels of importance, it is necessary to iterate through these regions again according to the above implementation process to obtain a subsequent extension range that includes multiple cloud resource deployment regions, each with a different level of importance. Similarly, when the importance of multiple candidate cloud resource deployment areas within the intermediate extension range differs, it is necessary to iterate through these multiple candidate cloud resources according to the above implementation process to obtain an intermediate extension range that includes multiple cloud resource deployment areas, each with varying levels of importance. The principle of the iteration process can be found in the description of the implementation process described above. A simple example of the implementation process, including the iteration process, is provided below for ease of understanding.
[0221] 1. The cloud management platform determines the distribution of N host classes across multiple cloud resource deployment areas. This is equivalent to calculating the distribution of the N host profiles (HP) across each Region / AZ.
[0222] 2. The cloud management platform determines the shortest change path for N host classes based on their distribution across multiple cloud resource deployment regions. This is equivalent to calculating the shortest change path covering the entire host profile based on its distribution, also known as the shortest path of ring (SP). Determining the shortest change path includes defining the minimum coverage area and its preceding, succeeding, and intermediate extensions, following rules that cover each level of importance. The calculation method is as follows:
[0223] 1) The cloud management platform calculates the minimum AZ set S that covers the entire profile. That is, it determines the minimum coverage area that covers N host classes. It satisfies the following conditions:
[0224] a) Select the region combination with the minimum number of regions that covers the entire image, based on the granularity of the region;
[0225] b) When the number of Regions is the same, select the Region combination with the cloud hosts deployed at the lowest importance level;
[0226] c) When deploying cloud servers of equal importance, select the Region with fewer cloud servers;
[0227] d) When the regions are the same, select the AZ with fewer cloud servers;
[0228] For example, referring to the example in Table 1 above, the full profile is: profile 1 / 2 / 3 / 4 / 5 / 6, and the selected minimum AZ set S is: Region3AZ2+Region5AZ1.
[0229] 2) The cloud management platform calculates the shortest path SP1 of the intermediate Ring ring (including the boundary Ring) of the minimum AZ set S. This determines the intermediate extension range of the minimum coverage area. The implementation process includes:
[0230] Step 1: Following the implementation process of step b44 above, determine the minimum AZ set S1 of the intermediate Ring.
[0231] Step 2: Following the implementation process of step b44 above, iteratively calculate the shortest path SP11-SP11 of the middle Ring ring of the smallest AZ set S1.
[0232] Step 3: Following the implementation process of step b42 above, iteratively calculate the shortest path SP11-SP12 of the preceding Ring ring of the minimum AZ set S1.
[0233] Step 4: Following the implementation process of step b43 above, iteratively calculate the shortest path SP11-SP13 of the rear Ring ring of the smallest AZ set S1.
[0234] Step 5: The shortest path SP1 of the middle Ring ring of the smallest AZ set S obtained by splicing according to the change order is: (SP11-SP12)+(SP11-SP11)+(SP11-SP13).
[0235] For example, referring to the example in Table 1 above, when the minimum AZ set S is Region3AZ2 + Region5AZ1, since there is no intermediate Ring in the minimum AZ set S, there is no Ring with an importance between Region3 and Region5.
[0236] 3) The cloud management platform calculates the shortest path of the preceding Ring ring of the minimum AZ set S. That is, it determines the preceding extension range of the minimum coverage area. The implementation process includes:
[0237] Step 1: Following the implementation process of step b42 above, determine the minimum AZ set S2 of the preceding Ring.
[0238] Step 2: Following the implementation process of step b44 above, iteratively calculate the shortest path SP22-SP21 of the middle Ring ring of the smallest AZ set S2.
[0239] Step 3: Following the implementation process of step b42 above, iteratively calculate the shortest path SP22-SP22 of the preceding Ring ring of the minimum AZ set S2.
[0240] Step 4: Following the implementation process of step b43 above, iteratively calculate the shortest path SP22-SP23 of the rear Ring ring of the smallest AZ set S2.
[0241] Step 5: The shortest path SP2 of the preceding Ring ring obtained by splicing according to the change order is: (SP22-SP22)+(SP22-SP21)+(SP22-SP23).
[0242] For example, referring to the example in Table 1 above, when the minimum AZ set S is Region3AZ2 + Region5AZ1, the importance of Region3 is less than that of Region5, and the cloud resource deployment regions with an importance lower than that of Region3 are Region1 and Region2. Region1 and Region2 both cover Image 1 / 2. Based on the condition of having the minimum total number of cloud hosts deployed, the minimum AZ set S2 of the front ring is selected as: Region2AZ1, and the shortest path SP2 is: Region2AZ1.
[0243] 4) The cloud management platform calculates the shortest path of the subsequent Ring ring of the minimum AZ set S. That is, it determines the subsequent extension range of the minimum coverage area. The implementation process includes:
[0244] Step 1: Following the implementation process of step b43 above, determine the minimum AZ set S3 of the subsequent Ring.
[0245] Step 2: Following the implementation process of step b44 above, iteratively calculate the shortest path SP33-SP31 of the middle Ring ring of the minimum AZ set S3;
[0246] Step 3: Following the implementation process of step b44 above, iteratively calculate the shortest path SP33-SP32 of the preceding Ring cycle of the minimum AZ set S3;
[0247] Step 4: Following the implementation process of step b44 above, iteratively calculate the shortest path SP33-SP33 of the rear Ring cycle of the minimum AZ set S3;
[0248] Step 5: The shortest path SP3 of the rear Ring ring of the smallest AZ set S obtained by splicing according to the change order is: (SP33-SP32)+(SP33-SP31)+(SP33-SP33).
[0249] For example, referring to the example in Table 1 above, when the minimum AZ set S is Region3AZ2 + Region5AZ1, the importance of Region3 is less than that of Region5, and the cloud resource deployment regions with higher importance than Region5 are Region7 and Region8. Region7 and Region8 cover the images 1 / 2 / 3 / 4. Based on the condition of having the minimum total number of cloud hosts deployed, the minimum AZ set S3 of the selected rear ring is: Region7AZ1, and the shortest path SP3 is: Region7AZ1.
[0250] After the above process, we can obtain the shortest path SP for the full image: SP1+S+SP3. Referring to the example in Table 1 above, the shortest path SP is: Region2 AZ1+Region3 AZ2+Region5 AZ1+Region7 AZ1.
[0251] Optionally, after dividing the cloud hosts in a host class into multiple change batches to be executed sequentially, the cloud management platform can also configure the execution strategy for the change batches in that host class. This is also known as configuring the execution strategy of the change pipeline. For example, the cloud management platform may also optionally perform one or more of the following operations: determine the start time and / or start conditions for each change batch in the first host class; determine the start time and / or start conditions for changes in the first host class; determine the start dependency order and / or start dependency conditions between different change batches in the first host class; determine the start dependency order and / or start dependency conditions between changes in different host classes; or, determine the sub-batch in the change batch of the first host class and the quantity relationship of the cloud hosts it includes. Please refer to the relevant descriptions above for the implementation details, which will not be elaborated here.
[0252] For any given change batch, the multiple hosts within that batch may begin implementing the changes all at once, or they may be implemented sequentially in multiple sub-batches. When multiple hosts in a change batch implement the changes in multiple sub-batches, the cloud management platform also needs to confirm the number of sub-batches and their constituent cloud hosts within each change batch of the first host class. In one possible implementation, the cloud management platform can divide all hosts included in the second change batch into multiple parts, with each part containing the cloud hosts included in a sub-batch, thereby obtaining the multiple sub-batches and their constituent cloud hosts within the second change batch. Here, the second change batch is any one of the multiple change batches within the first host class.
[0253] In another implementation, when determining the multiple sub-batches included in the second change batch and the number of cloud hosts they contain, the main consideration is that if a cloud host undergoing change first experiences a change anomaly, the anomaly will be detected earlier. The more cloud hosts with change anomalies, the greater the impact on the business. Therefore, the cloud management platform can optionally specify that the number of cloud hosts included in later sub-batches within the second change batch increases relative to the number of cloud hosts included in earlier sub-batches, in order to minimize the impact of potential change anomalies on the business. The method by which the number of cloud hosts included in later sub-batches increases relative to the number of cloud hosts included in earlier sub-batches can be determined according to application requirements. For example, the first sub-batchens in the second change batch may include one cloud host, the number of cloud hosts in the remaining sub-batches may increase by a factor of 2, and the number of cloud hosts in a single sub-batchens may not exceed 100. Since the first sub-batchens is the first batch to undergo change in the second change batch, the cloud hosts included in the first sub-batchens are also called OneBox cloud hosts.
[0254] Furthermore, when determining the multiple sub-batches that will be executed sequentially within the second change batch, further consideration can be given to the first sub-batch that is executed first, in order to further control the impact of the changes within a safe range. For example, the first sub-batch may optionally include one or more of the following cloud hosts:
[0255] 1) Idle cloud servers, and the total number of idle cloud servers in the first sub-batch accounts for no more than the proportion of the total number of cloud servers in the second change batch. Idle cloud servers are cloud servers that have not deployed instances for implementing tenant services.
[0256] 2) Cloud servers belonging to a specified tenant type, where the total number of cloud servers belonging to the specified tenant type in the first sub-batch accounts for no more than the corresponding second percentage in the total number of cloud servers in the second change batch. The second percentage for the specified tenant type is negatively correlated with the importance of the specified tenant type. Cloud service providers have various tenant types. For example, based on the importance of the implemented business, cloud service providers may have ordinary tenants and important tenants; in this case, the specified tenant type is either an ordinary tenant or an important tenant. If the importance of an ordinary tenant is less than that of an important tenant, then the second percentage for ordinary tenants is greater than that for important tenants.
[0257] 3) Cloud servers used for business testing (also known as business test cloud servers), and the total number of cloud servers used for business testing in the first sub-batch shall not exceed the proportion of the total number of cloud servers in the second change batch in the third sub-batch. Business test cloud servers are cloud servers that run test instances for certain testing purposes.
[0258] In this process, the priority for changing idle cloud servers in the first sub-batch decreases sequentially: the priority for changing cloud servers belonging to a specific tenant type in the first sub-batch, and the priority for changing cloud servers used for business testing. Furthermore, the priority for changing cloud servers belonging to a specific tenant type in the first sub-batch is negatively correlated with the importance of that tenant type. For example, assuming the second change batch includes idle cloud servers, cloud servers belonging to ordinary tenants, cloud servers belonging to important tenants, and business testing cloud servers, then the priority for changing idle cloud servers, cloud servers belonging to ordinary tenants, cloud servers belonging to important tenants, and business testing cloud servers in the first sub-batch decreases sequentially.
[0259] Furthermore, the values of the first, second, and third percentages can be determined based on application requirements. For example, the values of the first, second, and third percentages are determined based on the application's tolerance for changes or anomalies to the corresponding type of cloud server. For instance, the first percentage might be 5%, the second percentage for ordinary tenants might be 3%, the second percentage for important tenants might be 1%, and the third percentage might be 10%.
[0260] Optionally, in addition to this, the cloud servers included in the first sub-batch also meet one or more of the following criteria:
[0261] 1) For multiple cloud hosts with the same priority that are changed in the first sub-batch, the first sub-batch includes: cloud hosts serving fewer tenants, and / or cloud hosts with fewer instances deployed, with priority given to cloud hosts serving fewer tenants. This condition is equivalent to requiring that, among multiple cloud hosts with the same priority that are changed in the first sub-batch, cloud hosts serving fewer tenants and / or cloud hosts with fewer instances deployed be selected. When multiple cloud hosts have the same priority for changes in the first sub-batch, and cloud hosts serving fewer tenants are selected, if an anomaly occurs in the change, it can be guaranteed that fewer tenants will be affected, and the impact of the anomaly can be controlled from the perspective of the number of tenants. When multiple cloud hosts have the same priority for changes in the first sub-batch, and cloud hosts with fewer instances are selected, if an anomaly occurs in the change, it can be guaranteed that fewer instances will be affected, and the impact of the anomaly can be controlled from the perspective of the number of instances. Compared to cloud servers with fewer instances deployed, prioritizing cloud servers that serve fewer tenants ensures that fewer instances are affected if changes occur, thus minimizing the impact of changes on business operations and controlling the scope of impact from the perspective of reducing business disruptions.
[0262] 2) For multiple cloud hosts with services deployed under a specified tenant, the total number of cloud hosts with services deployed under the specified tenant in the first sub-batch is less than a first quantity threshold. By controlling the total number of cloud hosts with services deployed under the specified tenant in the first sub-batch, if an anomaly occurs during a change, the total number of cloud hosts with services deployed under the specified tenant affected by the anomaly can be controlled, thus keeping the impact of the anomaly on the services of the same tenant within a controllable range. The value of the first quantity threshold can be determined according to application requirements. For example, the value of the first quantity threshold is determined based on the application's tolerance for anomalies during changes to the corresponding type of cloud host. For example, the value of the first quantity threshold is 2.
[0263] Similarly, in other sub-batches within the second change batch, the cloud management platform can also control the total number of cloud hosts with services deployed on a specified tenant included in these other sub-batches to be less than a second quantity threshold. Other sub-batches are those sub-batches other than the first sub-batchens of the second change batch. By controlling the total number of cloud hosts with services deployed on a specified tenant included in these other sub-batches, if an anomaly occurs in the change, the total number of cloud hosts with services deployed on a specified tenant affected by the anomaly can be controlled, keeping the impact of the anomaly on the services of the same tenant within a manageable range. The value of the second quantity threshold can be determined according to application requirements. For example, the value of the second quantity threshold is determined based on the application's tolerance for anomalies in the corresponding type of cloud hosts. For instance, the second quantity threshold might be half the total number of cloud hosts with services deployed on a specified tenant.
[0264] It should be noted that the above descriptions of startup time, startup conditions, startup dependency order, startup dependency conditions, sub-batch in each change batch in the first host class and the number of cloud hosts included therein, and the conditions that the cloud hosts included in the first sub-batch need to meet are all examples of possible implementation methods and are not intended to limit the implementation method. There may be other implementation methods, but this application embodiment does not provide examples of them one by one.
[0265] Step 303: In the first batch of changes to M cloud hosts, change the target cloud host selected from N host classes. In subsequent batches of changes to M cloud hosts, change the other cloud hosts among the M cloud hosts except for the target cloud host.
[0266] After selecting multiple target cloud hosts from among the numerous cloud hosts included in each of the N host classes, the cloud management platform can then modify the selected target cloud hosts in the first batch of changes affecting M cloud hosts. Following the completion of the first batch, subsequent batches of changes affecting the M cloud hosts will then modify the other cloud hosts (excluding the target hosts). The first batch of changes serves as a profiling soak ring, used to soak and verify the full profile of the cloud hosts before deployment. This allows for rapid changes, achieving full profile coverage and soak verification, early problem detection, and keeping the impact of changes within a safe range. The order in which the other cloud hosts (excluding the target hosts) are modified can be determined by the cloud management platform based on the importance of the cloud host and its cloud resource deployment region or availability zone. For example, cloud hosts in less important cloud resource deployment regions may be modified first, followed by those in more important cloud resource deployment regions.
[0267] For example, Figure 11 is a schematic diagram of a change pipeline provided in an embodiment of this application. This change pipeline is obtained according to the second implementation of step 402. As shown in Figure 11, in the first change batch, the cloud management platform changes the cloud hosts deployed in Region2 AZ1+Region3 AZ2+Region5 AZ1+Region7 AZ1; in the second change batch, it changes the remaining cloud hosts in Region1 and Region2; in the third change batch, it changes the remaining cloud hosts in Region3 and Region4; in the fourth change batch, it changes the remaining cloud hosts in Region5 and Region6; and in the fifth change batch, it changes the remaining cloud hosts in Region7 and Region8.
[0268] As shown above, in the cloud host modification method based on public cloud technology, since the modification results of multiple cloud hosts within the same host class are basically the same, the cloud management platform selects a portion of cloud hosts from each of the N host classes, and modifies these selected cloud hosts in the first batch of modifications to M cloud hosts. This is equivalent to modifying a portion of cloud hosts for any host class first to verify the effect of the modification on that host class. This allows for the continuation of modifications to the remaining cloud hosts in the host class when the modification results indicate that the modification will not cause a failure, and timely termination of the modification when the modification results indicate that the modification will cause a failure. This reduces the probability of modification anomalies caused by cloud host modifications, keeps the impact of the modification within a safe range, and helps to safely and efficiently deploy cloud hosts in a phased rollout.
[0269] The following describes the process by which the cloud management platform classifies the cloud host to be modified into multiple host classes. Figure 12 is a flowchart of a method for classifying a cloud host to be modified into multiple host classes according to an embodiment of this application. As shown in Figure 12, the implementation process includes the following steps.
[0270] Step 1201: Obtain the operation and maintenance data of M cloud hosts.
[0271] Cloud server operation and maintenance data refers to the data required for the operation and maintenance of cloud servers. Based on this data, the cloud management platform can obtain operational characteristics reflecting the operational features of the cloud servers. This data is typically recorded during the operation and maintenance process. Optionally, this system can retrieve not only the operation and maintenance data of the M cloud servers to be modified, but also the operation and maintenance data of all cloud servers managed by the cloud management platform, providing more reference data for making changes to the cloud servers.
[0272] In one possible implementation, the operation and maintenance data includes the following categories of information: infrastructure, configuration, resources, alarms, monitoring, changes, and events. Infrastructure data includes, for example, host type, CPU, network interface card (NIC), disk, and memory. Configuration data includes, for example, product architecture, operating system (OS) version, hardware drivers, service deployment, service version, service configuration, and service characteristics. Resource data includes, for example, resource pools, virtual machines, virtual machine specifications, cloud services, and cloud tenants. Alarm data includes, for example, cloud service alarms and hardware device alarms. Monitoring data includes, for example, CPU utilization, memory utilization, input / output (IO) utilization, network traffic, service processes, and service business metrics. Change data includes, for example, a change list and change results. Event data includes, for example, a list of abnormal events. Table 2 shows an example of the operation and maintenance data used in the embodiments of this application.
[0273] Table 2
[0274] Step 1202: Based on the operation and maintenance data of M cloud hosts, determine the operation and maintenance characteristics of each cloud host among the multiple cloud hosts.
[0275] After the cloud management platform obtains the operation and maintenance data of M cloud hosts, it can optionally use feature extraction technology to obtain the feature set of the cloud hosts from the operation and maintenance data.
[0276] Step 1203: Based on the operation and maintenance characteristics of each cloud host in the multiple cloud hosts, divide the M cloud hosts into N host classes.
[0277] After obtaining the operational characteristics of each cloud host from multiple cloud hosts, the cloud management platform can divide the M cloud hosts into N host classes based on these characteristics. The division principle can be selected as follows: multiple cloud hosts with the same operational characteristics or whose similarity to operational characteristics is less than a specified threshold are grouped into the same host class.
[0278] The following is an example of how a cloud management platform can classify cloud hosts to be modified into multiple host classes. As shown in Figure 13, the implementation process includes:
[0279] Step 1: The cloud management platform integrates the operation and maintenance data into the data lake, generating the operation and maintenance dataset DS.
[0280] Step 2: The cloud management platform obtains all features of all cloud hosts based on the operation and maintenance data in the operation and maintenance dataset DS, resulting in a feature set FS that includes all features of all cloud hosts. For example, the cloud management platform learns all features of all cloud hosts from the operation and maintenance dataset DS using methods such as unsupervised learning. Optionally, after obtaining the feature set FS, the cloud management platform can also label the features with weights and add associated event lists to the features. For example, generally, the default weight of a feature is 1. For features that require special attention due to application requirements, the cloud management platform can increase the weight of the features, such as increasing the weight of the features mentioned in Table 2 to 5. The cloud management platform can also determine the association between features and operation and maintenance events, and increase the importance of the features according to the importance of the operation and maintenance events associated with the features. For example, when the operation and maintenance event associated with a feature is a general event, the importance of the feature is increased by 1; when the operation and maintenance event associated with a feature is a minor event, the importance of the feature is increased by 3; and when the operation and maintenance event associated with a feature is a critical event, the importance of the feature is increased by 5. The features are obtained based on the operation and maintenance data of the operation and maintenance events with which they are associated.
[0281] Step 3: The cloud management platform identifies the features needed for classifying cloud hosts from the feature set FS, obtaining the host profile feature set HP-FS. For example, based on the needs of classifying and modifying cloud hosts, the cloud management platform identifies the features needed for classifying cloud hosts from the feature set FS using methods such as expert experience, event feedback, and feature weights, obtaining the host profile feature set HP-FS. Identifying features through expert experience means determining the features needed for classifying and modifying cloud hosts based on expert experience. Identifying features through event feedback means determining the features needed for classifying and modifying cloud hosts based on the importance of the associated operational events and the operational events that occurred during historical changes. Identifying features through feature weights means identifying features with larger weights (e.g., weight greater than 2) as the features needed for classifying and modifying cloud hosts. Furthermore, since the cloud management platform needs to classify cloud hosts based on features, the features in the host profile feature set HP-FS also need to support attribute definitions, such as being able to be labeled with feature weights (FW) and being able to measure feature distances (FD) between features.
[0282] Step 4: The cloud management platform cleans and preprocesses the operations and maintenance dataset DS based on the host profile feature set HP-FS, generating the training dataset HP-DS. In other words, the cloud management platform cleans and preprocesses the features in the operations and maintenance dataset DS based on the selected features from the host profile feature set HP-FS, obtaining the features needed for classifying and modifying cloud hosts. These cleaned and preprocessed features are then stored in the training dataset HP-DS to optimize the features. The feature cleaning and preprocessing includes deleting some features or deleting data from certain dimensions of existing features.
[0283] Step 5: The cloud management platform collects the feature value set of each cloud host from the training dataset HP-DS and performs vectorization processing to form the feature vector FV.
[0284] Step 6: The cloud management platform calculates the feature distance FD between feature values based on the feature vectors FV of all cloud hosts using clustering methods such as hierarchical clustering.
[0285] Step 7: Based on the business requirements of the cloud services implemented by the cloud host, the cloud management platform selects cloud service-sensitive features from the host profile feature set HP-FS to obtain the host profile feature set HP-FS2.
[0286] Step 8: The cloud management platform adjusts the features in the host profile feature set HP-FS2 according to the business requirements of the cloud services implemented by the cloud hosts. This adjustment includes adjusting the feature weights (FW) and / or feature distances (FD) to improve the resolution of the generated host profiles. The adjusted host profile feature set HP-FS2 is the host profile template HP-FT. For example, the cloud management platform adjusts the feature distance based on the sensitivity of cloud services to features. For instance, assuming the cloud hosts to be changed are all used to implement network-related services, which are highly sensitive to network interface card (NIC) features, the cloud management platform fine-tunes the feature distances (FD) of NIC-related features, making the feature distances between cloud hosts with similar features smaller and the feature distances between cloud hosts with different features larger. By adjusting the feature weights (FW) and / or feature distances (FD), the cloud management platform makes the relationships between the features related to the services implemented by the cloud hosts more significant, providing a reference template for classifying cloud hosts. The host profile template HP-FT is essentially the template referenced when classifying cloud hosts, hence it is called the host profile template HP-FT.
[0287] Step 9: The cloud management platform uses a clustering algorithm to train the training dataset HP-DS and the host profiling template HP-FT to generate host profiles (HP). Each host profile represents a host class, and each host class includes multiple cloud hosts. For example, the cloud management platform first selects initial cluster centers, and then uses the K-Modes clustering algorithm to train the training dataset HP-DS and the host profiling template HP-FT, grouping cloud hosts according to the distance between the cloud host's feature vector (FV) and the cluster center. During the clustering process, the cluster centers are iterated according to the clustering progress until the cluster centers no longer change or the clustering cost no longer decreases. The clustering results, after fine-tuning, can be represented as host profiles, and their central features are the profile labels. Here, the host profile is the host equivalence class; the success rate of changes made to cloud hosts within a profile earlier is equivalent to or close to the success rate of changes made to cloud hosts later. Furthermore, R&D / testing can also use host profiles to build test environments and test cases to improve service version quality and R&D efficiency.
[0288] Optionally, as shown in Figure 14, before selecting the target cloud host from the cloud hosts included in each of the N host classes, the method further includes:
[0289] Step 1204: When the second host class and the third host class meet the merging conditions, merge the second host class and the third host class. Both the second host class and the third host class are one of the N host classes.
[0290] After classifying multiple cloud hosts into N host classes, the cloud management platform can further fine-tune the classification results based on application requirements. In one possible implementation, the cloud management platform merges the second and third host classes if one or more of the following merging conditions are met. These merging conditions include: the feature similarity between the second and third host classes is less than a first similarity threshold, or the number of cloud hosts included in both the second and third host classes is less than a first quantity threshold. For example, as shown in Figure 5, host classes 5 and 6 meet the merging conditions, therefore the cloud management platform merges them. When the feature similarity between the second and third host classes is less than the first similarity threshold, the deployment strategies and risk controls for changes to the second and third host classes are essentially the same, thus the cloud management platform can merge the second and third host classes. When the number of cloud hosts included in both the second and third host classes is less than the first threshold, since the proportion of cloud hosts included in the second and third host classes in the total number of cloud hosts in the N host classes is small, even if the second host class is merged with the third host class, the impact of any change anomaly will be small. Therefore, the cloud management platform can merge the second host class with the third host class. It should be noted that there are other ways to implement the merging conditions, which will not be exemplified in this application embodiment. In addition, when step 1204 is executed, the operation performed on the N host classes mentioned above is actually performed on the merged N host classes.
[0291] In one possible implementation, the cloud host modification method based on public cloud technology provided in this application can be implemented through multiple functional modules. For example, Figure 15 is a structural diagram of a functional module for implementing the method provided in an embodiment of this application. As shown in Figure 15, the operation and maintenance module is used to run operation and maintenance tools to execute the operation and maintenance process, obtain operation and maintenance data based on the operation and maintenance process, and unify the operation and maintenance data into the data lake to obtain a dataset. The host profiling module is responsible for managing profile features and profile templates, and using algorithms to train and generate host profiles from the data lake to classify cloud hosts. Among them, the host profiling module needs to design host profile features and host profile templates in advance on the host profiling system before the service goes online, and publish host profile data applications. The host profiling module includes: a profile feature management submodule, a service profile template submodule, a host profile submodule, and a gray-scale batching submodule. The profile feature management submodule is used to obtain a profile feature set based on the dataset, and to perform data cleaning and preprocessing on the dataset based on the profile feature set to obtain a training dataset. The service profile template submodule is used to obtain a host profile template based on the profile feature set. The host profile submodule is used to obtain a host profile based on the training dataset and the host profile module. The gray-scale batching submodule is used to divide multiple cloud hosts within the same host class into multiple sub-batches based on host profiles and host characteristics. The pipeline change module includes a pipeline management submodule and a pipeline implementation submodule. The pipeline management submodule is used to create profile pipelines using host profiles and determine the change strategy for each profile pipeline. The pipeline implementation submodule is used to implement changes to cloud hosts according to the profile pipelines. Cloud hosts deployed in the cloud platform infrastructure are the objects of change implementation. The specific working processes of each of the above modules and submodules can be found in the corresponding content of the aforementioned method embodiments, and will not be repeated here.
[0292] In summary, this application provides a method for modifying cloud hosts based on public cloud technology. This method is applied to a cloud management platform. The cloud management platform manages the infrastructure providing cloud services. The infrastructure includes cloud hosts. Cloud hosts are used to deploy instances that implement tenant services. In this method, the cloud management platform, based on the operational data of M cloud hosts managed by the platform, divides the M cloud hosts into N host classes. Then, within each of the N host classes, a subset of cloud hosts is selected from among the multiple cloud hosts included. The selected cloud hosts are designated as target cloud hosts. In the first batch of modifications to the M cloud hosts, the target cloud hosts selected from the N host classes are modified. In subsequent batches of modifications to the M cloud hosts, the other cloud hosts besides the target cloud hosts are modified. Each host class includes multiple cloud hosts. The operational data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. M and N are both positive integers greater than 1.
[0293] In this method, since the changes to multiple cloud hosts within the same host class are essentially the same, the cloud management platform selects a subset of cloud hosts from each of the N host classes. In the first batch of changes involving M cloud hosts, these selected hosts are modified. This is equivalent to first modifying a subset of cloud hosts for any given host class to verify the effect of the changes on that host class. This allows for the continuation of modifications to the remaining cloud hosts in the host class when the change results indicate that the changes will not cause failures, and timely termination of the changes when the results indicate that the changes will cause failures. This reduces the probability of changes causing anomalies and helps to ensure the safe and efficient canary rollout of cloud hosts.
[0294] Based on the above, this application can solve at least the following problems: 1) Low quality of service release versions. It advances the problem discovery phase to the R&D / testing side, reducing the number of patches deployed and improving R&D efficiency; 2) Long service deployment soaking period and low deployment efficiency. It can discover as many problems as possible at the beginning of deployment, improving deployment efficiency; 3) Risk control issues during service deployment. It simplifies canary deployment rules, controlling change risks for each batch, improving deployment experience and efficiency. For example, this application can build a test environment and test cases based on the entire network host profile, achieving 100% test coverage, improving version quality and R&D efficiency. Furthermore, this application can quickly generate the optimal soaking path for deployment, achieving 100% soaking coverage. The deployment soaking period can be shortened to one week, reducing the soaking period by more than 50%. During the deployment soaking phase, the problem discovery rate can reach over 95%. It can reduce the number of patches by at least 50%, improving R&D efficiency. It simplifies deployment strategies, requiring only host profiles and tenant business batches, making risk control simpler. Assist SREs in identifying risks in the live network and improving the quality of live network operations.
[0295] Furthermore, the order of steps in the cloud host modification method based on public cloud technology provided in this application embodiment can be appropriately adjusted, and steps can also be added or removed as needed. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.
[0296] The following describes an example of a virtual device in an embodiment of this application.
[0297] The above describes a cloud host modification method based on public cloud technology according to embodiments of this application. Corresponding to the above method, embodiments of this application also provide a cloud management platform for executing the method. Figure 16 is a schematic diagram of the structure of a cloud management platform provided in an embodiment of this application. Based on the following components shown in Figure 16, the cloud management platform shown in Figure 16 can perform all or part of the operations shown in Figure 3 above. It should be understood that the cloud management platform may include more additional components than those shown or omit some of the components shown, and embodiments of this application do not impose any limitations on this. The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes cloud hosts. Cloud hosts are used to deploy instances that implement tenant services. As shown in Figure 16, the cloud management platform 160 includes:
[0298] The classification module 1601 is used to classify the M cloud hosts managed by the cloud management platform into N host classes based on the operation and maintenance data of the cloud hosts. Each host class includes multiple cloud hosts. The operation and maintenance data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. M and N are both positive integers greater than 1.
[0299] Selection module 1602 is used to select multiple target cloud hosts from multiple cloud hosts included in each of the N host classes, wherein the multiple target cloud hosts in the first host class are a part of the multiple cloud hosts included in the first host class, and the first host class is any one of the N host classes.
[0300] The change module 1603 is used to make changes to the target cloud host selected from N host classes in the first change batch that makes changes to M cloud hosts.
[0301] The change module 1603 is also used to change other cloud hosts among the M cloud hosts, excluding the target cloud host, in subsequent change batches that change the M cloud hosts.
[0302] In one possible implementation, the selection module 1602 is specifically used to: obtain the importance of each cloud host in the P cloud hosts in the first host class to the tenant's business, where the first host class is any one of the N host classes and P is a positive integer; divide the P cloud hosts into multiple change batches to be executed sequentially based on their importance, wherein the cloud hosts in the later change batches are more important than the cloud hosts in the earlier change batches; and determine the cloud hosts in the first change batch among the multiple change batches as the target cloud hosts selected from the cloud hosts included in the first host class.
[0303] In one possible implementation, the importance of a cloud server to a tenant's business is determined based on one or more of the following: the resource utilization of the cloud server, the number and level of tenants served by the cloud server, or the number and level of business applications deployed on the cloud server.
[0304] In one possible implementation, M cloud hosts are deployed in multiple cloud resource deployment areas. The selection module 1602 is specifically used to: determine the distribution of N host classes in the multiple cloud resource deployment areas; based on the distribution, determine the minimum coverage of the N host classes in the multiple cloud resource deployment areas, wherein the minimum coverage includes the N host classes; and based on the cloud hosts included in the minimum coverage, obtain the target cloud host selected from the N host classes.
[0305] In one possible implementation, the selection module 1602 selects the target cloud host from the cloud hosts included in each of the N host classes, and further includes: determining the extension range of the minimum coverage area in multiple cloud resource deployment areas based on the distribution, and the union of the cloud resource deployment areas where the cloud hosts included in the minimum coverage area and the extension range are located is the multiple cloud resource deployment areas.
[0306] Accordingly, the selection module 1602, based on the cloud hosts included in the minimum coverage area, obtains the target cloud hosts selected from N host classes, including: based on the cloud hosts included in the minimum coverage area and the extended coverage area, the target cloud hosts selected from N host classes.
[0307] In one possible implementation, the selection module 1602 determines the extension range of the minimum coverage area across multiple cloud resource deployment areas based on the distribution, including: determining a first cloud resource deployment area combination where the cloud hosts included in the minimum coverage area are located, the first cloud resource deployment area combination including one or more target cloud resource deployment areas; in cloud resource deployment areas with an importance lower than the lowest importance among the one or more target cloud resource deployment areas, determining a preceding extension range of the minimum coverage area, the hosts included in the preceding extension range having a change order prior to the hosts included in the minimum coverage area; and in cloud resource deployment areas with an importance higher than the lowest importance among the one or more target cloud resource deployment areas, determining a preceding extension range of the minimum coverage area where the change order of the hosts is prior to the change order of the hosts included in the minimum coverage area. In high-importance cloud resource deployment areas, a subsequent extension range of minimum coverage is determined, wherein the change order of hosts included in the subsequent extension range follows the change order of hosts included in the minimum coverage area; when the first cloud resource deployment area combination includes multiple target cloud resource deployment areas, for any two target cloud resource deployment areas with adjacent importance, an intermediate extension range of minimum coverage is determined in cloud resource deployment areas with importance between the importance of any two target cloud resource deployment areas, wherein the change order of hosts included in the intermediate extension range is between the change order of hosts included in any two target cloud resource deployment areas.
[0308] In one possible implementation, for any target coverage area among the minimum coverage area and the extended coverage area, the selection module 1602 determines any target coverage area among multiple cloud resource deployment areas based on the distribution, including: determining a second cloud resource deployment area combination based on the combination of candidate cloud resource deployment areas including the most host classes among the candidate cloud resource deployment areas used to determine any target coverage area; and determining any target coverage area based on the second cloud resource deployment area combination.
[0309] In one possible implementation, the second cloud resource deployment area combination also satisfies one or more of the following conditions: among multiple combinations of candidate cloud resource deployment areas including the most host classes, it has the minimum total number of cloud resource deployment areas; among multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys cloud hosts of the lowest importance level; or, among multiple combinations of candidate cloud resource deployment areas including the most host classes, it deploys the minimum total number of cloud hosts.
[0310] In one possible implementation, each cloud resource deployment area includes multiple availability zones. The selection module 1602 determines any target coverage area based on the second cloud resource deployment area combination, including: determining a target availability zone combination based on the combination of availability zones that include the most host classes in the second cloud resource deployment area combination; and determining the target availability zone combination as any target coverage area.
[0311] In one possible implementation, the target availability zone combination also satisfies one or more of the following conditions: among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it has the minimum total number of availability zones; among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it deploys cloud hosts of the lowest importance; or, among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it deploys the minimum total number of cloud hosts.
[0312] In one possible implementation, M cloud hosts are modified in multiple change batches, and the second change batch of the multiple change batches includes multiple sub-batches executed sequentially, and the second change batch is any one of the multiple change batches. The first sub-batch executed first among multiple sub-batches includes one or more of the following cloud servers: idle cloud servers, and the total number of idle cloud servers in the first sub-batch does not exceed the first proportion of the total number of cloud servers in the second change batch; cloud servers belonging to a specified type of tenant, and the total number of cloud servers belonging to a specified type of tenant in the first sub-batch does not exceed the corresponding second proportion of the total number of cloud servers in the second change batch, the second proportion corresponding to the specified type of tenant is negatively correlated with the importance of the specified type of tenant; or, cloud servers used for business testing, and the total number of cloud servers used for business testing in the first sub-batch does not exceed the third proportion of the total number of cloud servers in the second change batch; the priority of changing idle cloud servers in the first sub-batch, the priority of changing cloud servers belonging to a specified type of tenant in the first sub-batch, and the priority of changing cloud servers used for business testing in the first sub-batch decreases in that order, and the priority of changing cloud servers belonging to a specified type of tenant in the first sub-batch is negatively correlated with the importance of the specified type of tenant.
[0313] In one possible implementation, the cloud hosts included in the first sub-batch also satisfy the following: for multiple cloud hosts with the same priority that change in the first sub-batch, the first sub-batch includes: cloud hosts that provide services to fewer tenants, and / or cloud hosts with fewer instances deployed; and / or, for multiple cloud hosts that deploy services to a specified tenant, the total number of cloud hosts that deploy services to a specified tenant included in the first sub-batch is less than a first quantity threshold.
[0314] In one possible implementation, as shown in Figure 17, the cloud management platform 160 further includes a merging module 1604, used to merge the second host class and the third host class when the second host class and the third host class meet one or more of the following merging conditions, the merging conditions including: the feature similarity between the second host class and the third host class is less than a first similarity threshold, or the number of cloud hosts included in the second host class and the third host class is less than a first quantity threshold, and the second host class and the third host class are both one of N host classes.
[0315] In one possible implementation, the selection module 1602 is further configured to perform one or more of the following operations: determine the start time and / or start conditions of each change batch in the first host class; determine the start time and / or start conditions of changes in the first host class; determine the start dependency order and / or start dependency conditions between different change batches in the first host class; determine the start dependency order and / or start dependency conditions between changes in different host classes; or, determine the sub-batch in the change batch in the first host class and the number of cloud hosts it includes.
[0316] In one possible implementation, the classification module 1601 is specifically used to: determine the operation and maintenance characteristics of each cloud host among the multiple cloud hosts based on the operation and maintenance data of M cloud hosts; and classify the M cloud hosts into N host classes based on the operation and maintenance characteristics of each cloud host among the multiple cloud hosts.
[0317] The classification module 1601, selection module 1602, modification module 1603, and merging module 1604 can all be implemented in software or hardware. For example, the implementation of the classification module 1601 will be described below. Similarly, the implementation of the selection module 1602, modification module 1603, and merging module 1604 can refer to the implementation of the classification module 1601.
[0318] As an example of a software functional unit, classification module 1601 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, classification module 1601 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one cloud data center or multiple geographically proximate cloud data centers. Typically, a region may include multiple AZs.
[0319] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0320] As an example of a hardware functional unit, the classification module 1601 may include at least one computing device, such as a server. Alternatively, the classification module 1601 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0321] The multiple computing devices included in the classification module 1601 can be distributed within the same region or in different regions. Similarly, the multiple computing devices included in the classification module 1601 can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the classification module 1601 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0322] It should be noted that, in other embodiments, any one of the classification module 1601, selection module 1602, change module 1603, and merging module 1604 can be used to execute any step in the cloud host change method based on public cloud technology. The steps implemented by the classification module 1601, selection module 1602, change module 1603, and merging module 1604 can be specified as needed. By implementing different steps in the cloud host change method based on public cloud technology through the classification module 1601, selection module 1602, change module 1603, and merging module 1604, all functions of the cloud resource management device based on public cloud technology can be realized.
[0323] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of each component described above can be referred to the corresponding content in the foregoing method embodiments, and will not be repeated here.
[0324] The following provides examples illustrating the basic hardware structures involved in the embodiments of this application.
[0325] This application also provides a computing device 1800. As shown in FIG18, the computing device 1800 includes: a bus 1802, a processor 1804, a memory 1806, and a communication interface 1808. The processor 1804, the memory 1806, and the communication interface 1808 communicate with each other via the bus 1802. The computing device 1800 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1800.
[0326] Bus 1802 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 18, but this does not imply that there is only one bus or one type of bus. Bus 1802 can include pathways for transmitting information between various components of computing device 1800 (e.g., memory 1806, processor 1804, communication interface 1808).
[0327] Processor 1804 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0328] The memory 1806 may include volatile memory, such as random access memory (RAM). The processor 1804 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0329] The memory 1806 stores executable program code, which the processor 1804 executes to implement the functions of the aforementioned classification module 1601, selection module 1602, change module 1603, and merging module 1604, thereby realizing the cloud host change method based on public cloud technology. In other words, the memory 1806 stores instructions for executing the cloud host change method based on public cloud technology.
[0330] The communication interface 1808 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 1800 and other devices or communication networks.
[0331] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0332] As shown in Figure 19, the computing device cluster includes at least one computing device 1800. The memory 1806 of one or more computing devices 1800 in the computing device cluster may store the same instructions for executing cloud host modification methods based on public cloud technology.
[0333] In some possible implementations, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also store partial instructions for executing the cloud host change method based on public cloud technology. In other words, a combination of one or more computing devices 1800 can jointly execute instructions for executing the cloud host change method based on public cloud technology.
[0334] It should be noted that the memory 1806 in different computing devices 1800 within the computing device cluster can store different instructions, which are used to execute certain functions of the cloud resource management device based on public cloud technology. That is, the instructions stored in the memory 1806 of different computing devices 1800 can implement the functions of one or more modules among the classification module 1601, selection module 1602, change module 1603, and merging module 1604.
[0335] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 20 illustrates one possible implementation. As shown in Figure 20, two computing devices 1800A and 1800B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1806 in computing device 1800A stores instructions for executing the functions of the classification module 1601, the selection module 1602, and the merging module 1604. Simultaneously, the memory 1806 in computing device 1800B stores instructions for executing the function of the change module 1603.
[0336] The connection method between the computing device clusters shown in Figure 20 can be considered as follows: taking into account that the cloud host change method based on public cloud technology provided in this application requires a large amount of data storage, the function implemented by the change module 1603 is to be executed by the computing device 1800B.
[0337] It should be understood that the functions of computing device 1800A shown in Figure 20 can also be performed by multiple computing devices 1800. Similarly, the functions of computing device 1800B can also be performed by multiple computing devices 1800.
[0338] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device clusters in Figures 19 and 20. The difference is that the memory 1806 in one or more computing devices 1800 in this computing device cluster can store the same instructions for executing the cloud host change method based on public cloud technology.
[0339] In some possible implementations, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also store partial instructions for executing the cloud host change method based on public cloud technology. In other words, a combination of one or more computing devices 1800 can jointly execute instructions for executing the cloud host change method based on public cloud technology.
[0340] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to perform a cloud host modification method based on public cloud technology.
[0341] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a cloud host modification method based on public cloud technology, or instruct the computing device to perform a cloud host modification method based on public cloud technology.
[0342] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0343] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.
[0344] In the embodiments of this application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "multiple" refers to two or more, unless otherwise expressly defined.
[0345] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0346] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the concept and principles of this application should be included within the protection scope of this application.
Claims
1. A method for modifying a cloud host based on public cloud technology, characterized in that, The method is applied to a cloud management platform, which manages infrastructure providing cloud services. The infrastructure includes cloud servers, which are used to deploy instances that implement tenant services. The method includes: Based on the operation and maintenance data of the M cloud hosts managed by the cloud management platform, the M cloud hosts are divided into N host classes, each host class including multiple cloud hosts. The operation and maintenance data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. M and N are both positive integers greater than 1. Multiple target cloud hosts are selected from the multiple cloud hosts included in each of the N host classes. The multiple target cloud hosts in the first host class are a part of the multiple cloud hosts included in the first host class. The first host class is any one of the N host classes. In the first batch of changes to the M cloud hosts, changes are made to the target cloud hosts selected from the N host classes; In subsequent batches of changes to the M cloud hosts, changes will be made to the other cloud hosts among the M cloud hosts, excluding the target cloud host.
2. The method according to claim 1, characterized in that, The step of selecting the target cloud host in each of the N host classes includes: Obtain the importance of each cloud host to the tenant's business among the P cloud hosts in the first host class, where the first host class is any one of the N host classes, and P is a positive integer; Based on the importance level, the P cloud hosts are divided into multiple change batches to be executed sequentially, wherein the cloud hosts in the later change batches are more important than the cloud hosts in the earlier change batches. The cloud host in the first change batch among the multiple change batches is determined as the target cloud host selected from the cloud hosts included in the first host class.
3. The method according to claim 2, characterized in that, The importance of the cloud server to the tenant's business is determined based on one or more of the following: the resource utilization rate of the cloud server, the number and level of tenants served by the cloud server, or the number and level of business applications deployed in the cloud server.
4. The method according to claim 1, characterized in that, The M cloud hosts are deployed in multiple cloud resource deployment regions, and the selection of the target cloud host in each of the N host classes includes: Determine the distribution of the N host classes in the multiple cloud resource deployment regions; Based on the distribution, the minimum coverage range of the N host classes is determined in the multiple cloud resource deployment areas, and the minimum coverage range includes the N host classes; Based on the cloud hosts included in the minimum coverage area, the target cloud hosts selected from the N host classes are obtained.
5. The method according to claim 4, characterized in that, The step of selecting the target cloud host from the cloud hosts included in each of the N host classes further includes: Based on the distribution, the extension range of the minimum coverage area is determined in the multiple cloud resource deployment areas, and the union of the cloud resource deployment areas where the cloud hosts are located, including the minimum coverage area and the extension range, is the multiple cloud resource deployment areas. The process of obtaining the target cloud host selected from the N host classes based on the cloud hosts included in the minimum coverage area includes: Based on the cloud hosts included in the minimum coverage area and the extended coverage area, the target cloud host selected from the N host classes is obtained.
6. The method according to claim 5, characterized in that, The step of determining the extension range of the minimum coverage area based on the distribution includes: The minimum coverage area includes a first cloud resource deployment area combination where the cloud hosts are located, and the first cloud resource deployment area combination includes one or more target cloud resource deployment areas; In cloud resource deployment regions with an importance lower than the lowest importance among the one or more target cloud resource deployment regions, a preceding extension range of the minimum coverage is determined, wherein the change order of hosts included in the preceding extension range precedes the change order of hosts included in the minimum coverage range; In a cloud resource deployment region that is more important than the highest importance among the one or more target cloud resource deployment regions, a subsequent extension range of the minimum coverage is determined, wherein the change order of hosts included in the subsequent extension range follows the change order of hosts included in the minimum coverage. In the case where the first cloud resource deployment area combination includes multiple target cloud resource deployment areas, for any two target cloud resource deployment areas whose importance is adjacent, and in the cloud resource deployment areas whose importance is between the importance of the two target cloud resource deployment areas, an intermediate extension range of the minimum coverage range is determined, and the change order of the hosts included in the intermediate extension range is between the change order of the hosts included in the two target cloud resource deployment areas.
7. The method according to any one of claims 4 to 6, characterized in that, For any target coverage area among the minimum coverage area and the extended coverage area, based on the distribution, determining the target coverage area in the plurality of cloud resource deployment regions includes: In the candidate cloud resource deployment regions used to determine the coverage of any of the targets, a second cloud resource deployment region combination is determined based on the combination of candidate cloud resource deployment regions that includes the most host classes; Based on the second cloud resource deployment area combination, determine the coverage area of any one of the targets.
8. The method according to claim 7, characterized in that, The second cloud resource deployment region combination also meets one or more of the following conditions: Among multiple combinations of candidate cloud resource deployment regions that include the most host classes, the one with the smallest total number of cloud resource deployment regions is the one with the smallest total number of cloud resource deployment regions. Among multiple combinations of alternative cloud resource deployment regions that include the most host classes, cloud hosts of the lowest importance are deployed. Alternatively, among multiple combinations of alternative cloud resource deployment regions that include the most host classes, the minimum total number of cloud hosts are deployed.
9. The method according to claim 7 or 8, characterized in that, Each cloud resource deployment region includes multiple availability zones. Determining the coverage area of any target region based on the second cloud resource deployment region combination includes: In the second cloud resource deployment area combination, the target availability zone combination is determined based on the combination of availability zones that include the most host types; The target availability zones are combined to determine the coverage area of any target.
10. The method according to claim 9, characterized in that, The target availability zone combination also meets one or more of the following conditions: Among the multiple availability zone combinations that include the most host classes in the second cloud resource deployment area combination, it has the minimum total number of availability zones; In the second cloud resource deployment area combination, which includes multiple availability zone combinations with the most host types, cloud hosts of the lowest importance are deployed. Alternatively, in the second cloud resource deployment area combination, which includes the combination of multiple availability zones with the most host classes, the minimum total number of cloud hosts is deployed.
11. The method according to any one of claims 1 to 10, characterized in that, The M cloud hosts are modified in multiple change batches. The second change batch of the multiple change batches includes multiple sub-batches executed sequentially. The second change batch is any one of the multiple change batches. The first sub-batch executed first among the multiple sub-batches includes one or more of the following cloud hosts: Idle cloud servers, and the total number of idle cloud servers in the first sub-batch accounts for no more than the first proportion of the total number of cloud servers in the second change batch; The cloud servers belonging to the specified tenant type, and the total number of cloud servers belonging to the specified tenant type in the first sub-batch accounts for no more than the corresponding second proportion in the total number of cloud servers in the second change batch, and the second proportion corresponding to the specified tenant type is negatively correlated with the importance of the specified tenant type; Alternatively, cloud hosts used for business testing, wherein the total number of cloud hosts used for business testing in the first sub-batch accounts for no more than the third proportion of the total number of cloud hosts in the second change batch. The priority of changing idle cloud hosts in the first sub-batch, the priority of changing cloud hosts belonging to a specified type of tenant in the first sub-batch, and the priority of changing cloud hosts used for business testing in the first sub-batch decrease in that order. Furthermore, the priority of changing cloud hosts belonging to a specified type of tenant in the first sub-batch is negatively correlated with the importance of the specified type of tenant.
12. The method according to claim 11, characterized in that, The cloud servers included in the first sub-batch also meet the following requirements: For multiple cloud hosts with the same priority that change in the first sub-batch, the first sub-batch includes: cloud hosts that serve fewer tenants, and / or cloud hosts with fewer instances deployed. And / or, for multiple cloud hosts that have services deployed with a specified tenant, the total number of cloud hosts in the first sub-batch that have services deployed with the specified tenant is less than a first quantity threshold.
13. The method according to any one of claims 1 to 12, characterized in that, Before selecting the target cloud host from the cloud hosts included in each of the N host classes, the method further includes: The second host class and the third host class shall be merged if one or more of the following merging conditions are met: the feature similarity between the second host class and the third host class is less than a first similarity threshold; or the number of cloud hosts included in both the second host class and the third host class is less than a first quantity threshold; and both the second host class and the third host class are one of the N host classes.
14. The method according to any one of claims 1 to 13, characterized in that, The method further includes one or more of the following operations: Determine the start time and / or start conditions for each change batch in the first host class; Determine the start time and / or start conditions for the change of the first host class; Determine the startup dependency order and / or startup dependency conditions among different change batches in the first host class; Determine the startup dependency order and / or startup dependency conditions between changes of different host classes; Alternatively, determine the sub-batches in the change batch within the first host class and the number of cloud hosts they include.
15. The method according to any one of claims 1 to 14, characterized in that, The operation and maintenance data of the M cloud hosts managed by the cloud management platform are used to divide the M cloud hosts into N host classes, including: Based on the operation and maintenance data of the M cloud hosts, determine the operation and maintenance characteristics of each cloud host among the multiple cloud hosts; Based on the operational characteristics of each cloud host among the multiple cloud hosts, the M cloud hosts are divided into N host classes.
16. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes cloud servers, which are used to deploy instances that implement tenant services. The cloud management platform includes: The classification module is used to divide the M cloud hosts managed by the cloud management platform into N host classes based on the operation and maintenance data of the M cloud hosts. Each host class includes multiple cloud hosts. The operation and maintenance data of the cloud hosts is the data required for the operation and maintenance of the cloud hosts. M and N are both positive integers greater than 1. The selection module is used to select multiple target cloud hosts from the multiple cloud hosts included in each of the N host classes, wherein the multiple target cloud hosts in the first host class are a part of the multiple cloud hosts included in the first host class, and the first host class is any one of the N host classes; The change module is used to change the target cloud host selected from the N host classes in the first change batch that changes the M cloud hosts; The change module is also used to change other cloud hosts among the M cloud hosts, excluding the target cloud host, in subsequent change batches that change the M cloud hosts.
17. The cloud management platform according to claim 16, characterized in that, The selection module is specifically used for: Obtain the importance of each cloud host to the tenant's business among the P cloud hosts in the first host class, where the first host class is any one of the N host classes, and P is a positive integer; Based on the importance level, the P cloud hosts are divided into multiple change batches to be executed sequentially, wherein the cloud hosts in the later change batches are more important than the cloud hosts in the earlier change batches. The cloud host in the first change batch among the multiple change batches is determined as the target cloud host selected from the cloud hosts included in the first host class.
18. The cloud management platform according to claim 16, characterized in that, The M cloud hosts are deployed in multiple cloud resource deployment regions, and the selection module is specifically used for: Determine the distribution of the N host classes in the multiple cloud resource deployment regions; Based on the distribution, the minimum coverage range of the N host classes is determined in the multiple cloud resource deployment areas, and the minimum coverage range includes the N host classes; Based on the cloud hosts included in the minimum coverage area, the target cloud hosts selected from the N host classes are obtained.
19. A computing device cluster, characterized in that, The system includes multiple computing devices, each comprising multiple processors and multiple memories, the multiple memories storing program instructions, and the multiple processors executing the program instructions to enable the cluster of computing devices to implement the method as described in any one of claims 1 to 15.
20. A computer-readable storage medium, characterized in that, Includes program instructions that, when executed on a computing device, cause the computing device to perform the method as described in any one of claims 1 to 15.
21. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 15.