A multi-cluster application fault migration method and system supporting multi-tenancy

By supporting multi-tenants to deploy applications and configure failure migration policies in a kubernetes multi-cluster environment, an automated failure migration solution is generated, which solves the problems of insufficient flexibility and low resource utilization of a single-cluster failure migration solution, and achieves efficient and flexible application failure migration.

CN119088623BActive Publication Date: 2025-06-20CHINA ACADEMY OF RAILWAY SCI CORP LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411132256.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-06-20
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

In the case of a single kubernetes cluster failure, the application failure migration solution has problems such as insufficient flexibility, failure to support migration by tenant, complex configuration, and low resource utilization.

Method used

By creating a multi-cluster environment, using kubernetes application software to manage and manage clusters and work clusters, support multi-tenants to deploy applications and configure failure migration policies, generate resource quota information tables and automatically generate failure migration solutions, and realize flexible migration of tenants, applications and components.

Benefits of technology

Improves the flexibility and resource utilization of application failure migration, reduces configuration complexity, supports migration and reuse by tenant, and minimizes service stop time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119088623B_ABST
    Figure CN119088623B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-cluster application fault migration method and system supporting multi-tenants, which relates to the technical field of cluster application fault migration. The method includes: creating a multi-cluster environment by using kubernetes application software; for each working cluster in the multi-cluster environment, creating tenants within the working cluster; for each tenant, deploying applications within the tenant and configuring fault migration policies associated with the applications; generating a resource quota information table according to the fault migration policies and the status information of the working clusters; generating a fault migration plan according to the resource quota information table, which is used to provide technical guidance after a certain working cluster fails, so as to migrate the tenants, applications and components that can be migrated in the failed working cluster to other working clusters that have not failed and redeploy them. The present application not only supports migration and reuse by tenant, but also improves the flexibility of adjustment, reduces the configuration complexity, and improves the resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cluster application fault migration, and particularly to a multi-cluster application fault migration method and system supporting multi-tenants. Background Art

[0002] With the increasing popularity of containerized applications, most enterprises use kubernetes (a container orchestration engine open-sourced by Google) as a container orchestration tool to host the operation of containers. Although the high availability of the kubernetes cluster can, to a certain extent, ensure the reliability of applications on the cluster, it still cannot avoid single-cluster failures caused by network failures or host failures. Therefore, in order to enable enterprise applications to stably provide services externally under any circumstances, it is necessary to solve the problem of application fault migration on the premise of a single kubernetes cluster failure and improve the reliability of applications.

[0003] Currently, it is usually to form a multi-cluster with multiple kubernetes clusters. When a certain cluster fails, the applications running on the faulty cluster can be migrated to other healthy clusters. However, this method has the following problems:

[0004] (1) Point-to-point migration. It only supports migration from Cluster A to Cluster B and cannot be flexibly adjusted. Each adjustment requires reconfiguring the policy.

[0005] (2) Tenant-based migration is not supported. In the current cluster fault migration solutions, they are all for the overall migration of the cluster or the fault migration of a single application, without involving the concept of tenants and not facilitating the metering of tenant resources.

[0006] (3) Case-by-case discussion, no reuse supported. After the cluster fault migration configuration occurs when the cluster fails, with the change of the cluster environment, the policy needs to be re-adjusted to adapt to the new cluster environment, and the flexibility is poor.

[0007] (4) Complex configuration. In order to control the fault migration process, users need to configure more parameters, increasing the difficulty and threshold for users to use and being unfavorable to the implementation of the overall strategy.

[0008] (5) More resource waste. In enterprises, many existing kubernetes clusters are running at a relatively full load. Each cluster has scattered remaining resources but cannot be effectively utilized; and in order to meet the overall migration of application faults, new backup clusters need to be built to accommodate the applications to be migrated, resulting in a large amount of resource waste and increasing enterprise costs.

[0009] In summary, how to provide a multi-cluster application fault migration method that can be flexibly adjusted, supports migration and reuse by tenant, has simple configuration, and high resource utilization rate is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0010] The purpose of this application is to provide a multi-cluster application fault migration method and system that support multi-tenants, which not only support migration and reuse by tenant, but also improve the flexibility of adjustment, reduce the configuration complexity, and enhance the resource utilization rate.

[0011] To achieve the above purpose, this application provides the following solutions:

[0012] In the first aspect, this application provides a multi-cluster application fault migration method that supports multi-tenants. The multi-cluster application fault migration method that supports multi-tenants includes:

[0013] Using the kubernetes application software, create a multi-cluster environment; the multi-cluster environment includes a management cluster and several working clusters, the management cluster is used to manage each of the working clusters, and the working clusters are used to deploy tenants.

[0014] For each working cluster in the multi-cluster environment, create a tenant within the working cluster. Each working cluster includes several of the tenants, and the tenants are used for resource isolation in the single working cluster scenario and cross-cluster resource isolation in the multi-working cluster scenario.

[0015] For each tenant, deploy an application within the tenant and configure a fault migration strategy associated with the application; the application includes multiple components, the fault migration strategy includes tenant information and application information, the application information includes application weight and application topology, the application weight is used to represent the priority of the application during fault migration, and the application topology includes the maximum number of replicas and the minimum number of replicas of each component in the application after fault migration.

[0016] Generate a resource quota information table according to the fault migration strategy and the status information of the working cluster; the status information of the working cluster includes the application deployment information, application resource occupancy information of each tenant in the working cluster, and the overall resource occupancy information of the working cluster. The resource quota information table includes the total resource information of each working cluster, the resource occupancy information of each tenant, the maximum resource usage value, the average resource occupancy value, and the working cluster status.

[0017] Generate a fault migration plan according to the resource quota information table; the fault migration plan is used to provide technical guidance after a failure occurs in one of the working clusters, so as to migrate the tenants, applications, and components that can be migrated in the failed working cluster to other working clusters that have not failed and redeploy them.

[0018] Optionally, a multi-cluster status collection module, a fault migration scheduling module, and a tenant management module are set up in the management cluster.

[0019] Among them, the multi-cluster status collection module is used to configure access authorization certificates for each working cluster in the management cluster, and obtain the status information and resource information of each working cluster through the APIServer of each working cluster.

[0020] The fault migration scheduling module is used to obtain the status information and resource information of each working cluster from the multi-cluster status collection module, and generate the fault migration plan in real time according to the fault migration strategy, and execute the fault migration of the application.

[0021] The tenant management module is used to manage tenant information in each working cluster.

[0022] Optionally, for each working cluster in the multi-cluster environment, creating a tenant in the working cluster specifically includes:

[0023] Create corresponding tenants in each working cluster according to tenant configuration information; where the tenant configuration information includes tenant name, total resource quota, and allocation amount, the total resource quota represents the total CPU quota, total memory quota, and total storage quota occupied by the tenant in the entire multi-cluster environment, and the allocation amount represents the CPU quota, memory quota, and storage quota occupied by the tenant in each working cluster respectively.

[0024] Generate an identification code for each tenant, and the identification code is used to identify the identity information of the tenant.

[0025] Optionally, the identification code is a UUID code.

[0026] Optionally, for each tenant, deploy an application in the tenant and configure a fault migration strategy associated with the application, specifically including:

[0027] Use an automatic deployment tool to deploy applications for each tenant respectively, and obtain tenant information and application information corresponding to each tenant.

[0028] Determine the fault migration strategy according to the tenant information and application information corresponding to each tenant.

[0029] Optionally, the automatic deployment tool is Argocd.

[0030] Optionally, according to the resource quota information table, a fault migration plan is generated, specifically including:

[0031] According to the resource quota information table, a preliminary fault migration plan is generated.

[0032] Optimize the preliminary fault migration plan to obtain an optimized fault migration plan; the optimized fault migration plan is used as the finally generated fault migration plan.

[0033] Optionally, optimizing the preliminary fault migration plan to obtain an optimized fault migration plan specifically includes:

[0034] According to the application weights corresponding to each application, determine the application with the smallest application weight from all the applications that have met the migration conditions of the tenant.

[0035] For each component in the application with the smallest application weight, when the maximum number of replicas of a certain component is greater than 1, then reduce the number of replicas of this component by 1 from the preliminary fault migration plan. In this way of reducing the number of replicas, reduce the reducible number of replicas of all components in all applications until the components that do not meet the migration conditions are included in the preliminary fault migration plan.

[0036] When all the reducible numbers of replicas have been reduced and there are still applications and their components that do not meet the migration conditions, mark the applications and their components that do not meet the migration conditions as non-migratable types, and delete the other components that have met the migration conditions of this application from the preliminary fault migration plan to obtain the optimized fault migration plan.

[0037] Optionally, after the step of optimizing the preliminary fault migration plan to obtain an optimized fault migration plan, the multi-tenant multi-cluster application fault migration method further includes:

[0038] After a certain working cluster fails, according to the optimized fault migration plan, perform fault migration on the tenants, applications, and components that can be migrated in the failed working cluster, and generate a fault migration report; the fault migration report includes the information of the migrated working cluster and the migrated working cluster of each successfully migrated application in each tenant in the failed working cluster, as well as the names and components of the applications that cannot be migrated.

[0039] Second aspect, the present application provides a multi-cluster application fault migration system supporting multi-tenancy. The multi-cluster application fault migration system supporting multi-tenancy includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the multi-cluster application fault migration method according to the first aspect.

[0040] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:

[0041] The present application provides a multi-cluster application fault migration method and system supporting multi-tenancy. By creating a multi-cluster environment, deploying applications for tenants, and configuring fault migration strategies associated with the applications, a resource quota information table can be generated according to the fault migration strategies and the status information of the working clusters. Further, a fault migration plan can be automatically generated according to the resource quota information table. When a certain working cluster fails, according to this fault migration plan, the tenants, applications, and components that can be migrated in the failed working cluster can be migrated and redeployed to other normal working clusters. It not only supports migration and reuse by tenant, but also can flexibly adjust the resource quota information table and the fault migration plan according to the fault migration strategies and the status information of the working clusters, effectively improving the flexibility of adjustment and reducing the configuration complexity. Moreover, it can enable tenants, applications, and their components to automatically implement fault migration in a multi-cluster environment, minimizing the service stop time to the greatest extent, and thus effectively improving the utilization rate of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a schematic flowchart of a multi-cluster application fault migration method supporting multi-tenancy in an embodiment of the present application.

[0044] Figure 2 It is a multi-cluster architecture diagram provided in an embodiment of the present application.

[0045] Figure 3 It is a relationship diagram of tenants and multi-clusters provided in an embodiment of the present application.

[0046] Figure 4 It is a configuration parameter diagram of application migration strategies in a tenant provided in an embodiment of the present application.

[0047] Figure 5The diagram showing the relationship between tenants and the total amount of application resources provided by an embodiment of this application.

[0048] Figure 6 The schematic diagram of overall tenant failure migration provided by an embodiment of this application.

[0049] Figure 7 The diagram showing the relationship between applications and the total amount of component resources provided by an embodiment of this application.

[0050] Figure 8 The schematic diagram of overall application migration within a tenant provided by an embodiment of this application.

[0051] Figure 9 The schematic diagram of completed migration of application sub-components provided by an embodiment of this application.

[0052] Figure 10 The schematic structural diagram of a multi-cluster application failure migration system supporting multi-tenants provided by an embodiment of this application. Detailed implementation manners

[0053] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0054] To make the above objects, features, and advantages of this application more obvious and understandable, the following further describes this application in detail in conjunction with the accompanying drawings and specific implementation manners.

[0055] As Figure 1 shown, this embodiment proposes a multi-cluster application failure migration method supporting multi-tenants. The multi-cluster application failure migration method supporting multi-tenants includes the following steps:

[0056] Step S1: Use the kubernetes application software to create a multi-cluster environment.

[0057] In this embodiment, the multi-cluster environment includes a management cluster and several working clusters. The management cluster is used to manage each of the working clusters, etc., and the working clusters are used to deploy tenants, etc.

[0058] In this embodiment, a multi-cluster status collection module, a failure migration scheduling module, and a tenant management module are set in the management cluster.

[0059] Among them, the multi-cluster status collection module is mainly used to configure access authorization certificates for each of the working clusters in the management cluster, and obtain status information and resource information of each of the working clusters through the APIServers of each of the working clusters, etc.

[0060] The fault migration scheduling module is mainly used to obtain status information and resource information of each of the working clusters from the multi-cluster status collection module, generate the fault migration plan in real time according to the fault migration strategy, and perform fault migration of applications, etc.

[0061] The tenant management module is mainly used to manage tenant information in each of the working clusters, etc.

[0062] Step S2: For each of the working clusters in the multi-cluster environment, create a tenant within the working cluster.

[0063] In this embodiment, each of the working clusters includes several tenants, and the tenants are used for resource isolation in a single working cluster scenario and cross-cluster resource isolation in a multi-working cluster scenario.

[0064] In this embodiment, step S2 for each of the working clusters in the multi-cluster environment to create a tenant within the working cluster specifically includes:

[0065] Step S21: Create the corresponding tenant within each of the working clusters according to the tenant configuration information; among them, the tenant configuration information includes tenant name, total resource quota and allocation amount. The total resource quota represents the total CPU quota, total memory quota and total storage quota occupied by the tenant in the entire multi-cluster environment, and the allocation amount represents the CPU quota, memory quota and storage quota respectively occupied by the tenant in each of the working clusters.

[0066] After constructing the tenant, a unique identifier can also be generated for each tenant, and the unique identifier is used to identify the identity information of the tenant. Among them, the unique identifier can be a UUID (Universally Unique Identifier), or other types of unique identifiers.

[0067] Step S3: For each tenant, deploy an application within the tenant and configure a fault migration strategy associated with the application.

[0068] In this embodiment, the application includes multiple components, and the fault migration policy includes tenant information and application information. The application information includes application weights and an application topology. The application weights are used to represent the priorities of the application during fault migration, and the application topology includes the maximum and minimum number of replicas of each component in the application after fault migration.

[0069] In this embodiment, step S3, for each tenant, deploys an application within the tenant and configures a fault migration policy associated with the application, specifically including:

[0070] Step S31: Use an automatic deployment tool to deploy the application for each tenant respectively, obtaining the tenant information and application information corresponding to each tenant. Among them, the automatic deployment tool can be Argocd or other automatic deployment tools.

[0071] Step S32: Determine the fault migration policy according to the tenant information and application information corresponding to each tenant.

[0072] Step S4: Generate a resource quota information table according to the fault migration policy and the status information of the working cluster.

[0073] In this embodiment, the status information of the working cluster includes the application deployment information, application resource occupancy information of each tenant in the working cluster, and the overall resource occupancy information of the working cluster. The resource quota information table includes the total resource information of each working cluster, the resource occupancy information of each tenant, the maximum resource usage value, the average resource occupancy value, and the status of the working cluster.

[0074] Step S5: Generate a fault migration plan according to the resource quota information table.

[0075] In this embodiment, the fault migration plan can be used to provide technical guidance after a certain working cluster fails, so as to migrate the tenants, applications, and components that can be migrated in the failed working cluster to other working clusters that have not failed and redeploy them.

[0076] In this embodiment, step S5 generates a fault migration plan according to the resource quota information table, specifically including:

[0077] Step S51: Generate a preliminary fault migration plan according to the resource quota information table.

[0078] Step S52. Due to the limitation of the tenant resource quota, it may not be possible to meet the fault migration of all applications and components in the initially created fault migration plan. Therefore, in this embodiment, it is considered necessary to optimize the initial fault migration plan so that it can meet the fault migration of as many applications or components as possible, thereby maximizing the resource utilization rate. Specifically, the initial fault migration plan is optimized to obtain an optimized fault migration plan. The optimized fault migration plan is used as the finally generated fault migration plan.

[0079] In this embodiment, step S52 optimizes the initial fault migration plan to obtain an optimized fault migration plan, which specifically includes:

[0080] Step S521. According to the application weights corresponding to each application, determine the application with the smallest application weight from all the applications that the tenant has met the migration conditions.

[0081] Step S522. For each component in the application with the smallest application weight, when the maximum number of replicas of a certain component is greater than 1, reduce the number of replicas of this component by 1 from the initial fault migration plan. In this way of reducing the number of replicas, reduce the reducible number of replicas of all components in all applications until the components that do not meet the migration conditions are included in the initial fault migration plan.

[0082] Step S523. When all reducible numbers of replicas have been reduced and there are still applications and their components that do not meet the migration conditions, mark the applications and their components that do not meet the migration conditions as non-migratable types, and delete the other components that have met the migration conditions of this application from the initial fault migration plan to obtain the optimized fault migration plan.

[0083] In this embodiment, after the step of step S52 optimizes the initial fault migration plan to obtain an optimized fault migration plan, the following steps are further included:

[0084] Step S53. After a certain working cluster fails, according to the optimized fault migration plan, perform fault migration on the tenants, applications, and components that can be migrated in the failed working cluster, and generate a fault migration report. The fault migration report includes the information of the migrated working cluster and the migrated working cluster of the successfully migrated applications in each tenant in the failed working cluster, as well as the names and components of the applications that cannot be migrated.

[0085] A multi-cluster application fault migration method supporting multiple tenants proposed in this embodiment, in actual application, the specific operation process is as follows:

[0086] Step (1): Create a multi-cluster environment.

[0087] In this embodiment, first, in a multi-cluster environment composed of multiple Kubernetes environments, the multi-cluster environment refers to an environment of multiple clusters. Here, the clusters are divided into two categories: management clusters and working clusters. Among them, there is usually one management cluster, and there are several working clusters, as Figure 2 shown. Select one of the clusters (3 master nodes and 1 - 2 node nodes are sufficient) as the management cluster, and the other clusters as working clusters. Among them, Kubernetes is abbreviated as k8s. k8s is an abbreviation formed by using the number "8" to replace the 8 characters "ubernete" in the middle of the name. It is an open-source tool used to manage containerized applications on multiple hosts in a cloud platform. The goal of Kubernetes is to make the deployment of containerized applications simple and efficient. Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance.

[0088] In this embodiment, the following modules need to be deployed in the management cluster: multi-cluster status collection module, fault migration scheduling module, and tenant management module. The above three modules together constitute the multi-cluster management plane. The multi-cluster architecture is as Figure 2 shown, and the main functions of each module are as follows:

[0089] 1) Multi-cluster status collection module: It is mainly used to obtain the working cluster status and resource information through each working cluster by configuring access authorization certificates for each working cluster in the management cluster. This module runs as a containerized application in the management cluster. Among them, APIServer is the central API coordinator in the k8s cluster, acting as the brain and entry point of the cluster. It is responsible for receiving and processing API requests from users, controllers, and other components, and executing these requests in the cluster.

[0090] 2) Fault migration scheduling module: It is mainly used to obtain the status information and resource information of each working cluster from the multi-cluster status collection module, generate a scheduling strategy in real time according to the fault migration strategy configured by the user, and execute the fault migration of the application. This module runs as a containerized application in the management cluster.

[0091] 3) Tenant management module: It is mainly used to manage the tenant information of each working cluster, etc.

[0092] In this embodiment, the following two operations are required to incorporate the working clusters into the management cluster: Import the kubeconfig files of each working cluster into the specified file directory of the management cluster as the identity basis for the multi-cluster status collection module to read the information of the working clusters. And configure the IP addresses and port numbers of the APIServers of each working cluster in the multi-cluster status collection module of the management cluster. The multi-cluster status collection module obtains the information of the working clusters through the IP addresses and ports.

[0093] Step (2): Create a cross-cluster tenant in the multi-cluster environment.

[0094] In this embodiment, in a single-cluster scenario, a tenant is a division of the resources of the working cluster for resource isolation within the cluster. In a multi-cluster scenario, a tenant is used for cross-cluster resource isolation, that is, a tenant can occupy resources on multiple clusters, and the total resources it occupies on each working cluster are the overall resource quota of the tenant. Several users can be created in the tenant, and the user with the ability to manage tenant resources is called the tenant administrator. In the management cluster, the kubernetes administrator submits the tenant configuration information in the multi-cluster environment to the tenant management module, and the tenant configuration information is used to create a tenant. The tenant configuration information includes the tenant name, total resource quota, allocation amount, etc. Among them, the total resource quota represents the total CPU quota, total memory quota, and total storage quota occupied by the tenant in the entire multi-cluster environment. The allocation amount represents the CPU quota, memory quota, and storage quota respectively occupied by the tenant on each working cluster (which can be all working clusters or some working clusters). The sum of the allocation amounts on each working cluster should be less than or equal to the total quota. The relationship between the tenant and the multi-cluster is as Figure 3 shown.

[0095] In this embodiment, after the tenant is created, the tenant management module generates a UUID to represent the tenant, and creates a cluster user in all the working clusters involved in the tenant and associates the UUID, representing the tenant administrator in the working cluster. Among them, UUID is an industry standard used in computer systems to ensure the high uniqueness of information. Usually, the platform and programming language provide corresponding functions or methods to generate UUIDs. The purpose is to enable all elements in the distributed system to have unique identification information without the need to specify the identification information through a central control end.

[0096] Step (3): Deploy applications within the tenant.

[0097] In this embodiment, the application is a logical integration of multiple components, and each component is a container image that can be deployed on Kubernetes. The containerized application consists of one or more components, and each component is a containerized image. The application can be deployed in any cluster of the working cluster, and the deployment method and script are the same as those for deploying services in Kubernetes. It is necessary to write Kubernetes deployment, service, and configmap scripts for each application, and the scripts must include the resource size required by the application components. The tenant administrator of the working cluster created in step (2) above uses an automatic deployment tool (such as Argocd) to deploy the application in a single working cluster. After successful deployment, the application will occupy a certain amount of resource quota, and at this time, the remaining available resource quota of the tenant should be equal to the current available resource quota minus the resource quota occupied by the application.

[0098] Step (4): Configure a failover policy associated with the application within the tenant.

[0099] In this embodiment, after the application is deployed, the Kubernetes administrator of the management cluster submits the basic failover policies of all applications under each tenant to the failover module. The failover policy mainly includes tenant information and application information, and the policy configuration parameter file is as Figure 4 shown. Among them, the tenant information includes the tenant name and tenant ID, etc. The application information includes the application weight and application topology, etc. Among them, the application weight is the priority of the application during failover, and applications not listed are configured with the lowest priority of 1 by default. The application topology is the maximum and minimum number of replicas of each component in the application after failover, and components not listed are configured with the maximum and minimum number of replicas of 1 by default.

[0100] It should be noted that the Kubernetes administrator and the tenant administrator are different. The Kubernetes administrator, that is, the system administrator, refers to the user with the highest permissions in the Kubernetes cluster and is responsible for the management of the Kubernetes cluster itself; while the tenant administrator refers to a user with management permissions in the tenant and is responsible for the management of resources within the tenant.

[0101] Step (5): The management cluster generates and maintains a resource quota information table in real time based on the status information of the current working cluster and the failover policy.

[0102] In this embodiment, there can be multiple tenants in the multi-cluster environment, and the resources of each tenant can span multiple clusters. The multi-cluster status collection module in the multi-cluster management plane will refresh the status of the working clusters regularly. For example, every 15 seconds, this multi-cluster status collection module reads the status information of the working cluster through the working cluster APIServer interface. The status information includes the deployment status of the applications under each tenant in the current working cluster, the application resource occupancy information, and the overall occupancy status of the current cluster resources. Therefore, the multi-cluster status collection module will generate a resource quota information table based on the status of each working cluster and tenant resource quotas and other information. As shown in Table 1, Table 1 takes two working clusters, namely Working Cluster 1 and Working Cluster 2, as examples for illustration. Actually, it can include more working clusters.

[0103] Table 1 Resource Quota Information Table

[0104]

[0105] Step (6): Once the management cluster senses a failure of the working cluster, it generates a fault migration plan for the application according to the resource quota information table.

[0106] In this embodiment, when the working cluster fails to communicate with the multi-cluster management plane for two consecutive refresh cycles, the status of this working cluster is marked as "abnormal". When a certain working cluster is marked as the "abnormal" state, the multi-cluster management plane will initiate the fault migration of the applications in this working cluster. In this embodiment, the refresh cycle can be set to 30 seconds.

[0107] In this embodiment, when the multi-cluster management plane of the management cluster senses that a certain working cluster is in an abnormal state, such as abnormal states like connection interruption or inability to obtain the working cluster information, the fault migration scheduling module will automatically create a fault migration plan according to the current "resource quota information table" (i.e., Table 1) and the deployment information of the applications in each working cluster. The creation process of the fault migration plan includes the following content:

[0108] 1) The working cluster that has failed can be simply referred to as the faulty cluster. Regarding all tenants in the faulty cluster as independent entities, the total actual resources occupied by tenant i in the faulty cluster, R(i), is equal to the sum of the resource usage quotas of the applications within the tenant, as Figure 5 shown.

[0109] 2) Among all the working clusters with normal working status, find the working cluster that meets the following conditions: The available resources RA(i) of tenant i in the working cluster exceed the total resources R(i) already occupied by tenant i in the faulty cluster; RA(i) - R(i) is the largest among all the working clusters that meet the conditions; this working cluster serves as the target cluster for the overall migration of all applications within tenant i of the faulty cluster, that is, the migration ultimately needs to be towards this working cluster, which can also be called the incoming working cluster. Correspondingly, the faulty working cluster is called the outgoing working cluster. The overall failure migration of the tenant is as Figure 6 shown.

[0110] 3) If no working cluster can be found to meet RA(i) ≥ R(i), it indicates that all working clusters are not sufficient to accommodate all the applications of tenant i as a whole; then a separate failure migration plan is formulated for each application in tenant i of the faulty cluster.

[0111] 4) Sort the applications of tenant i in descending order of application priority; for those with the same priority configuration, sort them according to the principle of the smallest total resources required by the application first. The application with less resource consumption has a higher priority, and the application with more resource consumption has a lower priority. The comparison relationship of resource consumption is compared in the order of storage, CPU, and memory. For the same storage requirement, the application with less CPU requirement has a smaller total resource; for the same storage and CPU requirements, the application with less memory occupancy has a smaller total resource.

[0112] 5) From the sorted queue, for each application j in turn, calculate its total resource occupancy Rapp(i, j), as Figure 7 shown. Among all the working clusters with normal working status, find the working cluster that meets the following conditions: The available resources RA(i) of tenant i in the working cluster exceed the total resources Rapp(i, j) occupied by application j of tenant i in the faulty cluster. RA(i) - Rapp(i, j) is the largest among all the clusters that meet the conditions; this working cluster serves as the target cluster for the migration of application j of tenant i in the faulty cluster, and application j is deleted from the queue. The overall migration of the application within the tenant is as Figure 8 shown.

[0113] 6) If all the components of application j of tenant i in the faulty cluster cannot be fully accommodated within the space of tenant i in all the working clusters with normal working status, then the failure migration of application j is subdivided into migrating different components of application j to different working clusters.

[0114] 7) For component k of application j in faulty cluster tenant i, calculate its total resource occupancy RappComp(i, j, k). Search for clusters in all working clusters with normal working status that meet the following conditions: The available resource RA(i) of tenant i within the working cluster exceeds the resource occupancy RappComp(i, j, k) of component k of application j of faulty cluster tenant i; RA(i) - RappComp(i, j, k) is the largest among all working clusters that meet the conditions; this working cluster serves as the target cluster for migrating component k of application j of faulty cluster tenant i.

[0115] 8) If all components in application j of tenant i meet the migration conditions, then delete application j from the queue; select the next application from the queue and jump to step 5). The migration of applications by component is as Figure 9 shown.

[0116] Step (7): The management cluster optimizes the initially generated faulty migration plan.

[0117] In this embodiment, the principles for the faulty migration scheduling module to optimize and adjust the initially generated faulty migration plan are as follows:

[0118] 1) When all components of all applications within a tenant have met the migration conditions (the sorting queue is emptied), there is no need for plan optimization, and directly go to step (8).

[0119] 2) When there are still applications or application components in the sorting queue that do not meet the migration conditions, select the application j with the lowest priority from the applications that have met the migration conditions of the current tenant. For its component k, if the maxReplicas (maximum number of replicas) of component k is greater than 1, then reduce its number of replicas by 1 from the initially generated faulty migration plan to save some resources, but the value after reduction is not less than minReplicas (minimum number of replicas), so as to reduce the resource occupancy after the migration of this component, and jump to step 6) of step (6). Among them, the initially generated faulty migration plan is a faulty migration plan created according to the default configuration policy. For example, after an application component is migrated to a new working cluster, it maintains 3 replicas, etc. maxReplicas and minReplicas represent the maximum number of replicas and the minimum number of replicas respectively.

[0120] 3) Continuously try to reduce the reducible number of replicas of all applications until the components that do not meet the migration conditions are included in the faulty migration plan.

[0121] 4) If all reducible numbers of replicas have been reduced, and there are still application components that do not meet the conditions in the queue, mark that this application cannot be faultily migrated (to provide information for generating a migration report), and delete the other components of this application that have met the migration conditions from the plan (either all components of an application are migrated or none are migrated).

[0122] In this embodiment, if some components of the applications in the preliminary fault migration plan cannot be migrated, it means that the resources are insufficient. Then, the number of replicas of the application components with relatively low priorities is reduced from the preliminary fault migration plan to save some resources. Then, it is checked whether the remaining resources can include the components that could not be migrated due to insufficient resources into the migration process. If the resources are still insufficient, some replicas are deleted in turn, and so on. If all the reducible replicas have been reduced and the saved resources still cannot accommodate all the components, then these components cannot achieve fault migration.

[0123] Step (8): The management cluster implements fault migration according to the final fault migration plan.

[0124] In this embodiment, the fault migration scheduling module implements fault migration according to the finally optimized fault migration plan generated in the above steps, redeploys all the migratable applications and components in the faulty cluster in the target cluster, thereby completing the fault migration.

[0125] Step (9): After completing the fault migration, the management cluster generates a fault migration report.

[0126] In this embodiment, when the migration of tenants, applications, and components is completed according to the fault migration plan, the fault migration scheduling module will generate a fault migration report. Taking tenant A as an example, the form and specific content of the fault migration report are as follows:

[0127] In tenant A:

[0128] Successfully migrated applications: which working cluster the application is located in after the fault migration (that is, which working cluster it has migrated into), or which working cluster each component is located in, etc.

[0129] Applications that cannot be migrated: the names and components of the applications that cannot be migrated.

[0130] In this embodiment, in a multi-cluster environment composed of multiple kubernetes clusters, the management cluster manages other working clusters and real-time perceives the health status of the working clusters; the tenants created by the management cluster each occupy a part of the resource quota in each working cluster. When the management cluster perceives that a certain working cluster fails, it will automatically generate a fault migration plan and implement the migration according to the tenant application migration configuration information and the status information of each current working cluster, so that in a multi-cluster and multi-tenant environment, the application can minimize the service stop time to the greatest extent and effectively improve the resource utilization rate.

[0131] This embodiment proposes a multi-cluster application fault migration method and system supporting multi-tenants, and constructs a multi-cluster environment composed of one Kubernetes management cluster and multiple Kubernetes working clusters. In this multi-cluster environment, user applications belong to tenants of the working clusters; when any working cluster fails, the management cluster automatically generates a fault migration plan according to the status of the working clusters, and migrates the applications from the failed working cluster to other non-failed working clusters in terms of tenants, applications, and components based on the remaining status of the tenants' resources, thus achieving high reliability of the applications and improving resource utilization. This system is a software system including a multi-cluster status collection module, a fault migration scheduling module, and a tenant management module. It not only supports migrating applications from one failed cluster to multiple normal clusters, but also supports fault migration of applications according to tenants, thus facilitating resource measurement for fault migration. Moreover, there is no need to specify the target cluster to be migrated in advance, and the fault migration plan is automatically generated based on brief configuration and cluster resource information, with high flexibility; at the same time, this method also supports the idempotency principle. This method and system support both the situation where one working cluster fails and the situation where multiple working clusters fail simultaneously or successively, and the processing solutions are the same without special treatment. In addition, it also supports migrating application components to different working clusters respectively, making full use of the fragmented resources in each working cluster, thus avoiding the increase in resource application volume due to overall application disaster recovery.

[0132] In an exemplary embodiment, a multi-cluster application fault migration system supporting multi-tenants is provided. This system can be a computer device, which can be a server or a terminal, and its internal structure diagram can be as Figure 10 shown. This computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store data related to multi-cluster application fault migration supporting multi-tenants. The input / output interface of this computer device is used to exchange information between the processor and external devices. The communication interface of this computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a multi-cluster application fault migration method supporting multi-tenants.

[0133] Those skilled in the art can understand, Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0134] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAMs), magnetoresistive random access memories (MRAMs), ferroelectric random access memories (FRAMs), phase change memories (PCMs), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0135] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0136] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on this application.

Claims

1. A multi-cluster application failure migration method supporting multi-tenants, characterized in that: The multi-cluster application failure migration method supporting multi-tenants includes: Using kubernetes application software, a multi-cluster environment is created; the multi-cluster environment includes a management cluster and several working clusters, the management cluster is used to manage each of the working clusters, and the working clusters are used to deploy tenants; For each of the working clusters in the multi-cluster environment, a tenant is created in the working cluster; each of the working clusters includes a plurality of tenants, and the tenants are used for resource isolation in a single working cluster scenario and cross-cluster resource isolation in a multi-working cluster scenario; For each of the tenants, deploy an application in the tenant and configure a fault migration policy associated with the application; the application includes multiple components, the fault migration policy includes tenant information and application information, the application information includes application weight and application topology, the application weight is used to indicate the priority of the application during fault migration, and the application topology includes the maximum number of replicas and the minimum number of replicas of each component in the application after fault migration; Generate a resource quota information table according to the fault migration strategy and the status information of the working cluster; the status information of the working cluster includes the application deployment information, application resource occupancy information and the overall occupancy information of the working cluster resources of each tenant in the working cluster, and the resource quota information table includes the total resource information of each working cluster, the resource occupancy information of each tenant, the maximum resource usage, the average resource occupancy value and the working cluster status; A fault migration plan is generated according to the resource quota information table; the fault migration plan is used to provide technical guidance after a certain working cluster fails, so as to migrate the tenants, applications and components that can be migrated in the failed working cluster to other working clusters that have not failed and redeploy them.

2. The multi-cluster application failure migration method supporting multi-tenants according to claim 1, characterized in that: The management cluster is provided with a multi-cluster status collection module, a fault migration scheduling module and a tenant management module; The multi-cluster status collection module is used to configure the access authorization certificate of each working cluster in the management cluster, and obtain the status information and resource information of each working cluster through the API Server of each working cluster; The fault migration scheduling module is used to obtain the status information and resource information of each working cluster from the multi-cluster status collection module, and generate the fault migration plan in real time according to the fault migration strategy, and execute the fault migration of the application; The tenant management module is used to manage tenant information in each of the working clusters.

3. The multi-cluster application failure migration method supporting multi-tenants according to claim 1, characterized in that: For each of the working clusters in the multi-cluster environment, creating a tenant in the working cluster specifically includes: According to the tenant configuration information, the corresponding tenant is created in each of the working clusters; wherein the tenant configuration information includes the tenant name, the total resource quota and the allocation quota, the total resource quota represents the total CPU quota, total memory quota and total storage quota occupied by the tenant in the entire multi-cluster environment, and the allocation quota represents the CPU quota, memory quota and storage quota occupied by the tenant in each of the working clusters; An identification code is generated for each of the tenants, where the identification code is used to identify the identity information of the tenant.

4. The multi-cluster application failure migration method supporting multi-tenants according to claim 3 is characterized in that: The identification code is a UUID code.

5. The multi-cluster application failure migration method supporting multi-tenants according to claim 1, characterized in that: For each of the tenants, deploy an application in the tenant and configure a fault migration strategy associated with the application, specifically including: Using an automatic deployment tool, deploy applications to each tenant respectively to obtain tenant information and application information corresponding to each tenant; The fault migration strategy is determined according to the tenant information and application information corresponding to each tenant.

6. The multi-cluster application failure migration method supporting multi-tenants according to claim 5, characterized in that: The automatic deployment tool is Argocd.

7. The multi-cluster application failure migration method supporting multi-tenants according to claim 1, characterized in that: Generate a fault migration plan based on the resource quota information table, specifically including: Generate a preliminary fault migration plan according to the resource quota information table; The preliminary fault migration plan is optimized to obtain an optimized fault migration plan; the optimized fault migration plan is used as the finally generated fault migration plan.

8. The multi-cluster application failure migration method supporting multi-tenants according to claim 7, characterized in that: The preliminary fault migration plan is optimized to obtain an optimized fault migration plan, which specifically includes: According to the application weights corresponding to the respective applications, the application with the smallest application weight is determined from all the applications of the tenant that have met the migration conditions; wherein the migration condition refers to that the sorting queue has been cleared, and the sorting queue is a queue formed after sorting the tenant's applications. When sorting the tenant's applications, the applications are sorted from high to low according to the application priority; for applications with the same priority configuration, they are sorted according to the principle of priority of the minimum total amount of resources required by the application, the application with less resource consumption has a higher priority, and the application with more resource consumption has a lower priority, and the resource consumption size comparison relationship is compared in the order of storage, CPU, and memory. For applications with the same storage requirements, the less CPU requirements, the smaller the total amount of resources; for applications with the same storage and CPU requirements, the less memory usage, the smaller the total amount of resources; For each component in the application with the smallest application weight, when the maximum number of copies of a component is greater than 1, the number of copies of the component in the preliminary fault migration plan is reduced by 1, and the number of reducible copies of all components in all applications is reduced in this manner until the components that do not meet the migration conditions are included in the preliminary fault migration plan; When all reducible copies have been reduced and there are still applications and their components that do not meet the migration conditions, the applications and their components that do not meet the migration conditions are marked as types that cannot be migrated, and other components of the application that already meet the migration conditions are deleted from the preliminary fault migration plan to obtain the optimized fault migration plan.

9. The multi-cluster application failure migration method supporting multi-tenants according to claim 7, characterized in that: After the step of optimizing the preliminary fault migration scheme to obtain an optimized fault migration scheme, the multi-cluster application fault migration method supporting multiple tenants further includes: After a failure occurs in a certain working cluster, the tenants, applications and components that can be migrated in the failed working cluster are failed over according to the optimized failure migration plan, and a failure migration report is generated; the failure migration report includes the migration-in working cluster information and the migration-out working cluster information of the successfully migrated applications in each tenant in the failed working cluster, as well as the names and components of the applications that cannot be migrated.

10. A multi-cluster application fault migration system supporting multiple tenants, the multi-cluster application fault migration system supporting multiple tenants comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the multi-cluster application failure migration method supporting multiple tenants according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data migration method and device for distributed cache system

    CN115292293A

  • Full-link implementation method for cross-cluster application migration

    CN116866159A