One - click Deployment and Lifecycle Management Method and System for Kubernetes Cluster

By generating node hardware and container mirroring behavior fingerprints, allocating global optimal nodes to the Kubernetes cluster, and performing parallel grayscale testing, the dynamic behavior and implicit dependency problems when nodes match container instances in the prior art are solved, and the adaptability and stability of the cluster are improved.

CN120029637BActive Publication Date: 2025-07-25NANJING TORTOISE & HARE RACE SOFTWARE RES INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510513092.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing Kubernetes cluster deployment technology ignores the dynamic behavior and implicit dependence of the container instance when the node matches the container instance, resulting in a high deployment failure rate, lack of dynamic adaptability of resource configuration, and grayscale testing cannot fully evaluate the impact of configuration changes, increasing the risk of business exceptions.

Method used

By generating node hardware fingerprints and container mirroring behavior fingerprints, the global optimal nodes are allocated to container instances based on these fingerprints, and parallel grayscale tests are performed. The split resource configuration strategy is to independently configure actions, identify and avoid implicit dependencies between configuration actions, and realize full deployment or state rollback.

Benefits of technology

It reduces the deployment failure rate of container instances, improves the accuracy of resource utilization and grayscale testing, ensures the stability and reliability of the deployment process, and reduces the risk of business abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029637B_ABST
    Figure CN120029637B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of application program control, and discloses a method and system for one-click deployment and life cycle management of a Kubernetes cluster. The method includes: collecting node parameters; generating a node hardware fingerprint based on the node parameters; parsing the metadata of a container image; generating a behavior fingerprint of the container image based on the metadata; allocating a globally optimal node for each container instance based on the node hardware fingerprint and the behavior fingerprint of the container image; and executing the deployment of the container instance on the corresponding globally optimal node. This application reduces the deployment failure rate, improves the intelligent level of resource configuration, and enhances the adaptability and stability of the Kubernetes cluster in complex business scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of application program control, and particularly to a method and system for one-click deployment and life cycle management of a Kubernetes cluster. Background Art

[0002] At present, certain progress has been made in the one-click deployment technology of Kubernetes clusters. Some tools and platforms can achieve basic cluster deployment functions to help users quickly set up a Kubernetes environment. However, there are still some deficiencies in the existing technologies.

[0003] In terms of the matching between nodes and container instances, most existing methods are based on resource requirements and allocation strategies. For example, static hardware dependencies such as the number of CPU cores and the amount of memory are considered, while the dynamic behaviors and implicit dependencies during the runtime of container instances are ignored. As a result, in actual deployment, it may occur that the image runs normally in local testing but crashes in the cluster environment, increasing the failure rate of deployment and the operation and maintenance costs.

[0004] The resource configuration and update of traditional container instances mostly rely on preset rules. For example, setting thresholds for CPU request values, etc., lacking the ability of automatic adaptation to the dynamic characteristics of business scenarios. The business load may change at any time, and static configuration rules cannot respond in a timely manner, resulting in problems of resource waste or resource shortage. When the business peak period of the Kubernetes cluster arrives, the pre-set resources may not meet the requirements, leading to a decline in application performance; while in the business trough period, excessive resource allocation will cause resource idleness, reducing the resource utilization rate.

[0005] In the testing link during the deployment and update processes, there are also deficiencies in the existing technologies. The current gray-box testing method lacks effective analysis and processing of the policy-level dependencies between configuration actions. Factors such as the collaborative operations and fault correlations of different configuration actions in historical operation data have not been fully considered, making the gray-box testing unable to comprehensively evaluate the impact of configuration changes on the entire system, increasing the risk of business anomalies after full-scale deployment.

[0006] The patent application with the publication number CN116860386A discloses a method for deploying a Kubernetes cluster, which includes the following steps: dividing the deployment process of the Kubernetes cluster into several independent tasks; further abstracting the divided tasks and defining a unified structure, so that each actual task becomes an implementation instance of the structure; based on the definition of the tasks, configuring appropriate characteristics for each task; organizing and orchestrating all tasks in sequence based on the Kubernetes cluster deployment process; when all tasks are executed, calling an interface to write the results into the management platform database, facilitating users to view the deployment results on the platform. This application solves the problems in the one-time deployment method where one error affects the whole, one node failure affects the overall process, and it has to start from scratch after failure, but there are still problems raised in the background technology of this application: the lack of the ability to automatically adapt to the dynamic characteristics of business scenarios.

[0007] The information disclosed in this background technology section is only intended to enhance the overall understanding of the present application and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention

[0008] The technical problem to be solved by this application is to overcome the defects of the prior art, provide a one-click deployment and life cycle management method and system for a Kubernetes cluster, and enhance the adaptability and stability of the Kubernetes cluster in complex business scenarios.

[0009] To solve the above technical problems, this application provides the following technical solutions:

[0010] On the one hand, this application provides a one-click deployment and life cycle management method for a Kubernetes cluster, which includes the following steps:

[0011] Collect node parameters; generate a node hardware fingerprint based on the node parameters; parse the metadata of the container image; generate a behavior fingerprint of the container image based on the metadata;

[0012] Based on the node hardware fingerprint and the behavior fingerprint of the container image, assign a globally optimal node to each container instance;

[0013] Execute the deployment of the container instance on the corresponding globally optimal node, specifically including:

[0014] Match the business characteristics of the container instance with a predefined resource configuration template to obtain a resource configuration policy;

[0015] Split the resource configuration policy into independent configuration actions;

[0016] Perform parallel gray-box testing on the configuration actions; based on the results of the parallel gray-box testing, perform full-scale deployment or state rollback of container instances on the corresponding globally optimal nodes.

[0017] As a preferred solution of the Kubernetes cluster one-click deployment and life cycle management method described in this application, wherein: the node parameters include hardware topology parameters, system environment parameters, runtime status parameters, and edge device parameters; the metadata includes explicit hardware dependencies, system environment dependencies, runtime behavior constraints, and edge computing requirements; any one of the node parameters corresponds to one piece of metadata.

[0018] The node hardware fingerprint is structured data containing all node parameters; the behavior fingerprint of the container image is structured data containing all metadata; the methods for generating the node hardware fingerprint and the behavior fingerprint of the container image are as follows: Structurally encode the node parameters and the metadata of the container image respectively using the same encoding method to obtain the node hardware fingerprint and the behavior fingerprint of the container image.

[0019] As a preferred solution of the Kubernetes cluster one-click deployment and life cycle management method described in this application, wherein: based on the node hardware fingerprint and the behavior fingerprint of the container image, assign globally optimal nodes to each container instance, specifically including:

[0020] Obtain the behavior fingerprint of the container image corresponding to each container instance as the behavior fingerprint of the corresponding container instance;

[0021] Based on the node hardware fingerprint and the behavior fingerprint of the container instance, calculate the matching weight between any container instance and any node;

[0022] Construct a weight matrix; any row in the weight matrix corresponds to a container instance, any column corresponds to a node, and the element value of any element is the matching weight between the corresponding container instance and the node;

[0023] Use an optimization algorithm to process the weight matrix to obtain a final matching solution; the final matching solution includes a set of element values selected in the weight matrix; wherein, there is exactly one element value selected in each row of the weight matrix, and any selected element value is not 0;

[0024] Assign globally optimal nodes to each container instance based on the final matching solution.

[0025] As a preferred solution of the Kubernetes cluster one-click deployment and life cycle management method described in this application, wherein: the specific calculation of the matching weight between any container instance and any node includes:

[0026] Obtain the weight scoring model for the container instance and the node; input the node hardware fingerprint and the behavior fingerprint of the container instance into the weight scoring model, and the weight scoring model calculates and outputs the matching weight between the corresponding container instance and the node;

[0027] The weight scoring model has built-in weight scoring rules; according to any one of the weight scoring rules, the weight scoring model calculates the single-item matching degree based on one node parameter in the node hardware fingerprint and the corresponding item metadata of the behavior fingerprint; the weight scoring model multiplies all the single-item matching degrees calculated based on each weight scoring rule to obtain the matching weight between the corresponding container instance and the node.

[0028] As a preferred solution of the Kubernetes cluster one-click deployment and life cycle management method described in this application, wherein: the business characteristics of the container instance include at least one of the hardware type required by the container instance, the runtime performance metrics of the container instance, and the SLA definition of the container instance; among them, the hardware type required by the container instance is extracted based on the explicit hardware dependencies of the corresponding container image; the runtime performance metrics of the container instance are extracted based on the runtime behavior constraints of the corresponding container image; the SLA definition of the container instance is extracted based on the system environment dependencies and edge computing requirements of the corresponding container image;

[0029] The resource configuration policy is a Kubernetes object description file, including at least one of a resource allocation policy, a scheduling constraint policy, and an elastic scaling policy.

[0030] As a preferred solution of the Kubernetes cluster one-click deployment and life cycle management method described in this application, wherein: the resource configuration template is a structured configuration file, including a resource allocation rule, a scheduling constraint rule, and an elastic scaling rule;

[0031] The resource allocation rule includes: generating a resource allocation policy based on the hardware type required by the container instance;

[0032] The scheduling constraint rule includes: generating a scheduling constraint policy based on the runtime performance metrics of the container instance;

[0033] The elastic scaling rule includes: generating an elastic scaling policy based on the SLA definition of the container instance.

[0034] As a preferred solution of the Kubernetes cluster one-click deployment and life cycle management method described in this application, wherein: splitting the resource configuration policy into independent configuration actions, specifically including:

[0035] Based on the declarative syntax of Kubernetes resource objects, parsing the resource configuration policy into atomic operation units;

[0036] Perform statelessness verification on each atomization operation unit; the method for performing statelessness verification includes explicit dependency detection; the explicit dependency detection includes: identifying explicit dependencies between atomization operation units by statically analyzing resource references in the atomization operation units;

[0037] Encapsulate the atomization operation units that pass the statelessness verification into configuration actions; any configuration action is a set of configuration instructions.

[0038] As a preferred solution of the method for one-click deployment and life cycle management of the Kubernetes cluster described in this application, wherein: perform parallel gray-box testing on the configuration actions, specifically including:

[0039] Parse the fields of the Kubernetes resource definitions in each configuration action to identify hardware resource conflicts between configuration actions;

[0040] Obtain the historical running logs of container instances; identify policy-level dependencies between configuration actions based on the historical running logs;

[0041] Select configuration actions to form a test group; there are no hardware resource conflicts and policy-level dependencies between any two configuration actions in any test group;

[0042] Perform parallel gray-box testing on the test group, specifically including: select container instances through a label selector to form a gray-box instance group; deploy the configuration actions in the test group to the gray-box instance group; route production traffic to the gray-box instance group, and block the write operations of the container instances in the gray-box instance group to the production data storage; collect the performance metrics of the gray-box instance group, and determine whether the test group passes the test based on the performance metrics.

[0043] As a preferred solution of the method for one-click deployment and life cycle management of the Kubernetes cluster described in this application, wherein: the policy-level dependencies include: the frequency of collaborative operations of any two configuration actions in historical running data exceeds a preset frequency threshold; the collaborative operation is the synchronous execution of configuration instructions of different configuration actions; the policy-level dependencies also include: there is a fault correlation between different configuration actions in historical running data; the fault correlation includes the correlation of fault triggering and the correlation of fault recovery.

[0044] As a preferred solution of the method for one-click deployment and life cycle management of the Kubernetes cluster described in this application, wherein: the test results of the parallel gray-box testing include passing the test and failing the test;

[0045] The full-scale deployment or state rollback of container instances specifically includes:

[0046] If the test group passes the test, deploy the configuration actions in the test group to all container instances;

[0047] If the test group fails the test, identify the conflicting actions, where the conflicting actions are the configuration actions that cause the test group to fail the test, and roll back the deployment status of the conflicting actions in the gray-scale instance group.

[0048] In a second aspect, the present application provides a one-click deployment and life cycle management system for a Kubernetes cluster, including a fingerprint generation module, a node allocation module, a resource configuration module, a gray-scale testing module, and a deployment management module; where:

[0049] The fingerprint generation module is used to collect the node parameters and the metadata of the container image, generate the node hardware fingerprint of each node based on the node parameters, and generate the behavior fingerprint of each container image based on the metadata;

[0050] The node allocation module allocates the globally optimal node for each container instance based on the node hardware fingerprint and the behavior fingerprint of the container image;

[0051] The resource configuration module generates a resource configuration policy based on the business characteristics of the container instance, and splits the resource configuration policy into independent configuration actions;

[0052] The gray-scale testing module is used to perform parallel gray-scale testing on the configuration actions, and feedback the results of the parallel gray-scale testing to the deployment management module;

[0053] The deployment management module performs full-scale deployment or status rollback of the container instances according to the results of the parallel gray-scale testing.

[0054] Compared with the prior art, the beneficial effects achieved by the present application are as follows:

[0055] The present application comprehensively considers the static data and the dynamic behavior and implicit dependencies during the operation of the container instance, constructs the node hardware fingerprint and the container image behavior fingerprint, and allocates the globally optimal node for the container instance through fingerprint matching and verification, solving the problem that the application runs normally locally but crashes in the cluster.

[0056] By constructing a resource configuration template driven by business semantics, automatically maps the business characteristics of the application to the Kubernetes resource configuration policy, and decomposes complex cluster-level changes into multiple independently verifiable atomic operation units, reducing the dependence on manual experience, improving the accuracy and efficiency of gray-scale testing, reducing the business exception rate caused by configuration changes, and ensuring the stability and reliability of the deployment process. Description of the Drawings

[0057] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings. Among them:

[0058] Figure 1 It is a flowchart of the one-click deployment and life cycle management method for the Kubernetes cluster provided by the present application;

[0059] Figure 2 It is a schematic structural diagram of the one-click deployment and life cycle management system for the Kubernetes cluster provided by the present application. Specific implementation manners

[0060] The following will detail the technical solutions of the present application through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0061] Embodiment 1

[0062] This embodiment introduces a one-click deployment and life cycle management method for a Kubernetes cluster. Referring to Figure 1 , the method includes the following steps:

[0063] Collect node parameters; generate a node hardware fingerprint based on the node parameters; parse the metadata of the container image; generate a behavior fingerprint of the container image based on the metadata;

[0064] The node parameters include hardware topology parameters, system environment parameters, runtime status parameters, and edge device parameters; the metadata includes explicit hardware dependencies, system environment dependencies, runtime behavior constraints, and edge computing requirements; any one node parameter corresponds to one item of metadata.

[0065] The following are some preferred node parameters and metadata in this embodiment:

[0066] The hardware topology parameters in the node parameters include GPU architecture, number of physical cores, and storage medium type; the explicit hardware dependencies in the metadata include GPU architecture requirements, number of physical cores requirements, and storage medium type requirements; the corresponding relationship between the node parameters and the metadata includes: GPU architecture corresponds to GPU architecture requirements, number of physical cores corresponds to number of physical cores requirements, and storage medium type corresponds to storage medium type requirements. The system environment parameters in the node parameters include kernel version and system call set; the system environment dependencies in the metadata include kernel version requirements and system call whitelist; the corresponding relationship between the node parameters and the metadata also includes: kernel version corresponds to kernel version requirements, and system call set corresponds to system call whitelist. The runtime state parameters in the node parameters include interrupt affinity configuration, clock synchronization protocol, and NUMA node architecture; the runtime behavior constraints in the metadata include interrupt request binding requirements, clock synchronization accuracy requirements, and NUMA architecture dependencies; the edge device parameters in the node parameters include hot plug interface and encryption acceleration hardware; the edge computing requirements in the metadata include hot plug device support and encryption acceleration requirements; the corresponding relationship between the node parameters and the metadata also includes: hot plug interface corresponds to hot plug device support; encryption acceleration hardware corresponds to encryption acceleration requirements.

[0067] The node hardware fingerprint is structured data containing all node parameters; the behavior fingerprint of the container image is structured data containing all metadata; the method for generating the node hardware fingerprint and the behavior fingerprint of the container image is as follows: Structured encoding is performed on the node parameters and the metadata of the container image respectively using the same encoding method to obtain the node hardware fingerprint and the behavior fingerprint of the container image. In this embodiment, the JSON or Protobuf format is preferably used as the structured encoding format for the node hardware fingerprint and the behavior fingerprint, which can retain the readability of the original data and support the matching and compatibility verification between fingerprints. By constructing the node hardware fingerprint and the behavior fingerprint of the container image, this application not only considers static data (such as explicit hardware dependencies), but also considers the dynamic behavior and implicit dependencies during the runtime of the container instance (such as interrupt request binding requirements, clock synchronization accuracy requirements, etc.). Through the matching and verification between the node hardware fingerprint and the behavior fingerprint for the container instance and the node, the deployment failure rate of the container instance can be reduced, which is more suitable for the heterogeneous hardware environment of the Kubernetes cluster and can solve problems such as the container running normally locally but crashing in the cluster.

[0068] Based on the node hardware fingerprint and the behavior fingerprint of the container image, a globally optimal node is assigned to each container instance; specifically including:

[0069] Obtain the behavior fingerprint of the container image corresponding to each container instance as the behavior fingerprint of the corresponding container instance;

[0070] Calculate the matching weight between any container instance and any node based on the node hardware fingerprint and the behavior fingerprint of the container instance;

[0071] Construct a weight matrix; any row in the weight matrix corresponds to a container instance, any column corresponds to a node, and the element value of any element is the matching weight between the corresponding container instance and the node;

[0072] Process the weight matrix using an optimization algorithm to obtain a final matching scheme; the final matching scheme includes a set of element values selected in the weight matrix; among them, there is exactly one element value selected in each row of the weight matrix, and any selected element value is not 0;

[0073] Selecting an element value from each row in the weight matrix will form a matching scheme; since only one element value in each row is selected in any matching scheme, and each row in the weight matrix corresponds to a container instance, therefore, any matching scheme assigns a node to each container instance. In this embodiment, the KM algorithm is preferably used as the optimization algorithm, and the matching scheme that maximizes the cumulative weight (i.e., the sum of the selected set of element values) is found through the KM algorithm as the final matching scheme, so that the node assigned to each container is its globally optimal node. During the process of the KM algorithm exploring the final matching scheme, the CPU utilization rate of the node, etc. are used as constraints to prevent the processing capacity of the node from being exceeded when multiple container instances are assigned to one node.

[0074] Allocate a globally optimal node for each container instance based on the final matching scheme. Each element value selected in the final matching scheme corresponds to a container instance and its corresponding globally optimal node, and the globally optimal node corresponding to each container instance can be determined through the corresponding relationship of each element value.

[0075] The specific calculation of the matching weight between any container instance and any node includes:

[0076] Obtain the weight scoring model of the container instance and the node; input the node hardware fingerprint and the behavior fingerprint of the container instance into the weight scoring model, and the weight scoring model calculates and outputs the matching weight between the corresponding container instance and the node;

[0077] The weight scoring model has built-in weight scoring rules; according to any weight scoring rule, the weight scoring model calculates the single-item matching degree based on one node parameter in the node hardware fingerprint and the corresponding item metadata of the behavior fingerprint; the weight scoring model multiplies all the single-item matching degrees calculated based on each weight scoring rule to obtain the matching weight between the corresponding container instance and the node.

[0078] The preferred partial weight scoring rules in this embodiment are as follows: If the GPU architecture exactly matches the GPU architecture requirements, the corresponding single-item matching degree is 1; otherwise, the corresponding single-item matching degree is 0. For example, if both the GPU architecture and the GPU architecture requirements are Ampere architecture, they are exactly matched. If the kernel version meets the kernel version requirements, the corresponding single-item matching degree is 1; otherwise, the corresponding single-item matching degree is 0. For example, if the kernel version is 5.1.1 and the kernel version requirement is greater than or equal to 5.1.0, then the kernel version meets the kernel version requirements.

[0079] If the system call set covers all system calls in the system call whitelist, the corresponding single-item matching degree is 1; otherwise, the corresponding single-item matching degree is 0. If the number of physical cores does not meet the physical core requirements, the corresponding single-item matching degree is 0; otherwise, the single-item matching degree is assigned based on the number of physical cores. The more physical cores, the higher the single-item matching degree, and the value range of the single-item matching degree is ; for example, the physical core requirement is not less than 8 cores, and the upper limit of the number of physical cores is set to , if the number of physical cores is not less than , then the single-item matching degree is 1; if the number of physical cores is less than 8, then the single-item matching degree is 0; if the number of physical cores is not less than 8 and less than , then the value range is . Since the matching weight between the container instance and the node is the product of all single-item matching degrees multiplied together, when any single-item matching degree is 0, the matching weight is also 0, and the node with a matching weight of 0 will not be used as the globally optimal node for the corresponding container instance alternative. Thus, it realizes filtering out unmatched nodes through hard conditions.

[0080] Deploy the container instance on the corresponding globally optimal node, which specifically includes:

[0081] Match the business characteristics of the container instance with the predefined resource configuration template to obtain a resource configuration policy;

[0082] The business characteristics of the container instance include at least one of the hardware type required by the container instance, the runtime performance metrics of the container instance, and the SLA definition (i.e., service level agreement definition) of the container instance. Among them, the hardware type required by the container instance is extracted based on the explicit hardware dependencies of the corresponding container image, including GPU architecture, storage medium type, encryption acceleration hardware requirements, etc.; the runtime performance metrics of the container instance are extracted based on the runtime behavior constraints of the corresponding container image, including NUMA node architecture dependencies, interrupt request binding requirements, clock synchronization accuracy requirements, etc.; the SLA definition of the container instance is extracted based on the system environment dependencies and edge computing requirements of the corresponding container image, including the minimum number of replicas, maximum latency threshold, etc.

[0083] The resource configuration policy is an executable Kubernetes object description file, including at least one of a resource allocation policy, a scheduling constraint policy, and an elastic scaling policy.

[0084] The resource configuration template is a structured configuration file, including a resource allocation rule, a scheduling constraint rule, and an elastic scaling rule; in this embodiment, YAML or JSON is preferably used as the specific format of the resource configuration template, and the resource configuration template is associated with the metadata of the container image through a version identifier to extract the business characteristics of the container instance from the corresponding metadata.

[0085] The resource allocation rule includes: generating a resource allocation policy based on the hardware type required by the container instance; the resource allocation policy includes device plugin binding declarations and limit fields of Kubernetes resources, etc.

[0086] The scheduling constraint rule includes: generating a scheduling constraint policy based on the runtime performance metrics of the container instance; the scheduling constraint policy includes node affinity configuration, NUMA core binding policy, interrupt affinity declaration, etc.

[0087] The elastic scaling rule includes: generating an elastic scaling policy based on the SLA definition of the container instance; the elastic scaling policy includes horizontal scaling policy, replica count configuration, resource utilization scaling threshold, etc.

[0088] Split the resource configuration policy into independent configuration actions; specifically including:

[0089] Based on the declarative syntax of Kubernetes resource objects, parse the resource configuration policy into atomic operation units; the atomic operation units include resource allocation declarations, scheduling constraint declarations, runtime optimization declarations, network policy declarations, etc.

[0090] Perform statelessness verification on each atomic operation unit;

[0091] The methods for performing statelessness verification include explicit dependency detection. Specifically, by statically analyzing the resource references in the atomic operation unit, identify the explicit dependencies between atomic operation units, such as service selection dependencies and config map name dependencies. Through dependency detection, it can be ensured that the execution of any atomic operation unit does not depend on the output of other operation units. The methods for statelessness verification also include resource isolation verification to ensure that each atomic operation unit only operates on independent Kubernetes resources, and when performing a state rollback, only the corresponding Kubernetes resource objects need to be deleted. If the atomic operation unit fails the statelessness verification, perform manual reconstruction of the atomic operation unit, such as manually merging the atomic operation units with explicit dependencies.

[0092] Encapsulate the stateless-verified atomic operation unit as a configuration action; any configuration action is an independently executable configuration instruction set. The configuration instructions included in the configuration action include Kubernetes native API call commands, CRD operation commands, custom controller coordination instructions, etc. In this embodiment, the configuration action is stored in the form of a YAML manifest file or a Helm Chart template, and each configuration action is associated with a unique identifier to support the status tracking of subsequent gray-box testing.

[0093] Perform parallel gray-box testing on the configuration action; based on the results of the parallel gray-box testing, perform full-scale deployment or state rollback of container instances on the corresponding globally optimal nodes.

[0094] Performing parallel gray-box testing on the configuration action specifically includes:

[0095] Parse the fields of the Kubernetes resource definitions in each configuration action to identify hardware resource conflicts between configuration actions; identify hardware resource conflicts based on the field correlation of the Kubernetes resource definitions. The hardware resource conflicts include mutually exclusive access declarations of different configuration actions to the same hardware device (such as GPU, storage volume), over-allocation of resource quotas (such as CPU, memory quota) of different configuration actions, and competition for underlying physical resources (such as LLC cache, memory bandwidth) of different configuration actions.

[0096] Obtain the historical running logs of container instances; identify policy-level dependencies between configuration actions based on the historical running logs; perform logical segmentation on the historical running logs of the deployed container instances to obtain atomic operation units consistent with the current configuration action, so as to extract the historical running data of the configuration action; perform statistical analysis on the co-changes between configuration actions based on the historical running data of the configuration action to obtain policy-level dependencies between configuration actions; the policy-level dependencies include: the frequency of co-operations of any two configuration actions in the historical running data exceeds a preset frequency threshold; the co-operation is the synchronous execution of configuration instructions of different configuration actions; the policy-level dependencies also include: there is a fault correlation between different configuration actions in the historical running data; the fault correlation includes the correlation of fault triggering, that is, different configuration actions have associated faults. For example, if configuration action A is the configuration of a database connection pool and configuration action B is the declaration of a database service Endpoint, then when configuration action B fails due to network errors or other reasons, configuration action A will trigger a fault because it cannot connect to the database. The fault correlation also includes the correlation of fault recovery, that is, there is a logical dependency between different configuration actions in fault recovery. For example, if configuration action C is the mounting of a storage volume and configuration action D is the declaration of a storage class, during fault recovery, configuration action C needs to be executed after configuration action D is successfully rolled back, otherwise data corruption may occur due to the inability to unmount the storage volume.

[0097] Select atomic operation units to form a test group; there is no hardware resource conflict and policy-level dependency between any two atomic operation units in any test group;

[0098] Conduct parallel canary testing on the test group;

[0099] The steps of the parallel canary testing are as follows:

[0100] Select container instances through a label selector to form a canary instance group; for example, select 10% of the container instances as the canary instance group and mark the value of the canary label of each selected container instance as true;

[0101] Deploy the configuration actions in the test group to the canary instance group;

[0102] Route a specified proportion of the production traffic to the canary instance group, and at the same time block the write operations of the container instances in the canary instance group to the production data storage; for example, route 10% of the production traffic to the canary instance group through the service mesh, and then block the write operations of the canary instance group to the production data storage through database shadow tables, mirrored topics of message queues, cache replicas isolation, etc. to achieve isolation between test data and production data;

[0103] Collect the performance metrics of the canary instance group; for example, collect performance metrics such as latency, error rate, resource utilization, etc.;

[0104] Based on a preset performance metric threshold, determine whether the test group passes the test. For example, if the error rate is less than 0.1%, the test is passed.

[0105] Record the test results of the parallel canary testing, and split the test group that has completed the parallel canary testing into independent atomic operation units. Split the test group into independent configuration actions to provide independent operation objects for subsequent full-scale deployment or state rollback.

[0106] In the embodiments of the present application, first, statelessness verification is used to ensure that there is no explicit dependency between configuration actions, and then implicit dependencies (hardware resource conflicts, policy-level dependencies) between configuration actions are identified, so as to achieve more refined dependency management and optimize the accuracy and efficiency of parallel canary testing.

[0107] The test results of the parallel canary testing include passing the test and failing the test;

[0108] Based on the results of the parallel canary testing, perform full-scale deployment or state rollback of container instances on the corresponding globally optimal nodes, specifically including:

[0109] If the test group passes the test, deploy the configuration actions in the test group to all container instances; for example, gradually synchronize the configuration actions in the gray instance group to all container instances through progressive traffic switching.

[0110] If the test group fails the test, identify the conflicting actions, where the conflicting actions are the configuration actions that cause the test group to fail the test, and roll back the deployment status of the conflicting actions in the gray instance group.

[0111] Based on real-time gray test data, conflicting actions can be identified in the test group. For example, monitor logs of parallel gray tests are recorded in real time, and error events in the monitor logs can be used to identify conflicting actions. If the test group fails the parallel gray test, roll back the deployment status of the conflicting actions, including deleting the Kubernetes resource definitions of the conflicting actions (such as Deployment, ConfigMap, etc.), deleting the temporary data written by the conflicting units during the test (such as database shadow tables), etc. Other actions in the test group except the conflicting actions do not need to be rolled back, and their deployment status in the gray instance group is retained without full deployment. Subsequently, they form a new test group with other configuration actions. When the new test group passes the parallel gray test, deploy the configuration actions in the new test group to all container instances.

[0112] The resource configuration and update of traditional container instances mostly rely on static rules preset by operation and maintenance personnel (such as manually setting the CPU request value threshold), lacking the ability of automatic adaptation to the dynamic characteristics of business scenarios. This application realizes intelligent optimization through the following technical paths: First, construct a resource configuration template driven by business semantics, automatically map the business characteristics of the application program to the Kubernetes resource configuration policy, and then use declarative resource change splitting to decouple complex cluster-level changes into multiple independently verifiable atomic operation units; finally, through dynamic traffic coloring and isolation technology, implement parallel gray verification on multi-dimensional resource configuration policies to achieve progressive deployment decisions. The technical path of this application realizes the optimization effect of improving resource utilization while reducing the dependence on manual experience, and reduces the business exception rate caused by configuration changes.

[0113] The one-click deployment and life cycle management method of the Kubernetes cluster described in the embodiments of this application is not only applicable to the initial installation stage of container instances, but also applicable to the update and upgrade stage of application programs. In the application upgrade scenario, smooth rolling upgrade and gray release are realized by dynamically adapting to the version differences before and after the upgrade. Specifically, when the application is upgraded, the following key steps need to be adaptively adjusted:

[0114] When generating the behavior fingerprint of a container image, it is necessary to parse the metadata of the upgraded image and generate a differential behavior fingerprint based on the incremental changes of the version upgrade. For example, if the upgraded container image newly adds runtime behavior constraints on the real-time process scheduling policy (such as requiring the CPU core isolation mechanism to be disabled), it is necessary to expand the corresponding scheduler configuration requirement field in the behavior fingerprint and expand the kernel real-time patch status in the node hardware fingerprint to perform incremental matching with the new runtime behavior constraints. Preferably, a version compatibility verification rule is introduced into the weight scoring model. For example, when the intersection of the upgraded system call whitelist and the node's system call set covers the old version whitelist, it is determined as a compatibility match, thus avoiding node mismatch problems caused by version upgrades.

[0115] Furthermore, when constructing the weight matrix, the matching weight is dynamically corrected in combination with the rolling upgrade strategy. For example, taking the node allocation status of the old version container instance as a constraint condition to ensure that the new version container instance is preferentially scheduled to idle nodes or low-load nodes, thereby reducing resource contention during the upgrade process. In addition, when using an optimization algorithm to process the weight matrix, a version affinity constraint is introduced. For example, taking the version distribution uniformity of multiple container instances of the same service as the optimization goal to prevent the risk of single point of failure caused by centralized version deployment.

[0116] Furthermore, in the stage of generating the resource configuration strategy, a version transition strategy is generated based on the business feature differences before and after the upgrade. For example, if the upgraded container instance needs to perform data migration, the data migration rule is expanded in the resource configuration template to generate a composite resource configuration strategy including storage volume declaration conversion, database Schema version management, and cache data preheating strategy. Preferably, in the stateless verification of the atomic operation unit, the isolation verification of cross-version resource references is increased. For example, detecting the name conflict of the configuration mappings of the new and old versions to ensure that the state can be restored by deleting a single resource object during version rollback.

[0117] Furthermore, in the parallel gray box testing stage, the traffic coloring and isolation mechanism is optimized for the upgrade scenario. For example, in the gray box instance group, the mixed deployment status of the new and old version container instances is tested in parallel. The production requests are simultaneously routed to the new and old version container instances through the traffic mirroring function of the service mesh, and the performance index differences between the two are compared to verify the upgrade compatibility. If the test fails, quickly roll back to the behavior fingerprint and resource configuration template of the old version based on the version identifier, and at the same time retain the incrementally generated differential fields as the upgrade failure log for subsequent version iteration analysis.

[0118] Through the above-mentioned adaptive adjustments, the embodiments of the present application can effectively identify the implicit dependency changes caused by version differences in the application upgrade scenario, realize the dynamic adaptation of the resource allocation strategy and the precise control of gray release, improve the success rate and efficiency of cluster upgrade while ensuring service continuity, and thus achieve the life cycle management of the Kubernetes cluster.

[0119] Embodiment 2

[0120] This embodiment is the second embodiment of the present application; based on the same inventive concept as Embodiment 1, referring to Figure 2 , this embodiment introduces a one-click deployment and life cycle management system for Kubernetes clusters, including a fingerprint generation module, a node allocation module, a resource configuration module, a gray box testing module, and a deployment management module; wherein:

[0121] The fingerprint generation module is used to collect the node parameters and the metadata of the container images, generate the node hardware fingerprints for each node based on the node parameters, and generate the behavior fingerprints for each container image based on the metadata;

[0122] The node allocation module allocates the globally optimal nodes for each container instance based on the node hardware fingerprints and the behavior fingerprints of the container images; this module calculates the matching weights between the container instances and the nodes based on the node hardware fingerprints and the behavior fingerprints of the container instances, constructs a weight matrix and processes the weight matrix using an optimization algorithm to obtain the final matching solution, so as to allocate the globally optimal nodes for each container instance.

[0123] The resource configuration module generates a resource configuration strategy based on the business characteristics of the container instances and splits the resource configuration strategy into independent configuration actions; this module is configured with a resource configuration template, maps the business characteristics of the container instances into a resource configuration strategy based on the resource configuration template, and splits the resource configuration strategy into independent configuration actions through stateless verification.

[0124] The gray box testing module is used to perform parallel gray box testing on the configuration actions and feedback the results of the parallel gray box testing to the deployment management module; this module generates test groups by identifying and avoiding hardware resource conflicts and policy-level dependencies between the configuration actions, and performs parallel gray box testing on the test groups to improve the test efficiency and accuracy.

[0125] The deployment management module performs full-scale deployment or status rollback of the container instances according to the results of the parallel gray box testing. If the test group passes the test, the deployment management module deploys the configuration actions in the test group to all container instances; if the test group fails the test, the deployment management module rolls back the deployment status of the conflicting actions in the gray box instance group, and at the same time retains the deployment status of other actions in the gray box instance group except the conflicting actions in the test group for subsequent composition of new test groups with other configuration actions.

[0126] For the specific function implementation of each of the above modules, refer to the relevant content in the method for one-click deployment and life cycle management of the Kubernetes cluster described in Embodiment 1, which will not be elaborated here.

[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose and scope protected by the present application, and these all fall within the protection scope of the present application.

Claims

1. A method for one-click deployment and lifecycle management of a Kubernetes cluster, characterized in that It includes the following steps: Collect node parameters; generate a node hardware fingerprint based on the node parameters; parse the metadata of the container image; generate a behavior fingerprint of the container image based on the metadata; The node parameters include hardware topology parameters, system environment parameters, runtime status parameters, and edge device parameters; the metadata includes explicit hardware dependencies, system environment dependencies, runtime behavior constraints, and edge computing requirements; any one node parameter corresponds to one metadata; the node hardware fingerprint is structured data containing all node parameters; the behavior fingerprint of the container image is structured data containing all metadata; Based on the node hardware fingerprint and the behavior fingerprint of the container image, assign a globally optimal node to each container instance; specifically including: based on the node hardware fingerprint and the behavior fingerprint of the container image, calculate the matching weight between the container instance corresponding to the container image and the node, and assign a globally optimal node to each container instance based on the matching weight; Execute the deployment of the container instance on the corresponding globally optimal node, specifically including: Match the business characteristics of the container instance with a predefined resource configuration template to obtain a resource configuration policy; Split the resource configuration policy into independent configuration actions; Conduct parallel gray-box testing on the configuration actions; specifically including: Parse the fields of the Kubernetes resource definitions in each configuration action to identify hardware resource conflicts between configuration actions; Obtain the historical running logs of the container instance; identify policy-level dependencies between configuration actions based on the historical running logs; Select configuration actions to form a test group; there are no hardware resource conflicts and policy-level dependencies between any two configuration actions in any test group; conduct parallel gray-box testing on the test group; The policy-level dependencies include: the frequency of collaborative operations between any two configuration actions in the historical running data exceeds a preset frequency threshold; the collaborative operation is the synchronous execution of configuration instructions of different configuration actions; the policy-level dependencies also include: there is a fault correlation between different configuration actions in the historical running data; the fault correlation includes the correlation of fault triggering and the correlation of fault recovery; Based on the results of the parallel gray-box testing, perform full-scale deployment or state rollback of the container instance on the corresponding globally optimal node.

2. The one-click deployment and lifecycle management method for a Kubernetes cluster according to claim 1, characterized in that: The method for generating the node hardware fingerprint and the behavior fingerprint of the container image is as follows: use the same encoding method to perform structured encoding on the node parameters and the metadata of the container image respectively to obtain the node hardware fingerprint and the behavior fingerprint of the container image.

3. The Kubernetes cluster one-key deployment and lifecycle management method according to claim 2, characterized in that: Based on the node hardware fingerprint and the behavior fingerprint of the container image, assign a globally optimal node to each container instance, specifically including: Obtain the behavior fingerprint of the container image corresponding to each container instance as the behavior fingerprint of the corresponding container instance; Calculate the matching weight between any container instance and any node based on the node hardware fingerprint and the behavior fingerprint of the container instance; Construct a weight matrix; any row in the weight matrix corresponds to a container instance, any column corresponds to a node, and the element value of any element is the matching weight between the corresponding container instance and the node; Process the weight matrix using an optimization algorithm to obtain a final matching scheme; the final matching scheme includes a set of element values selected in the weight matrix; wherein, for each row in the weight matrix, exactly one element value is selected, and any selected element value is not 0; Allocate a globally optimal node for each container instance based on the final matching scheme.

4. The method for one-click deployment and life cycle management of a Kubernetes cluster according to claim 3, wherein: The specific calculation of the matching weight between any container instance and any node includes: Obtain a weight scoring model for the container instance and the node; input the node hardware fingerprint and the behavior fingerprint of the container instance into the weight scoring model, and the weight scoring model calculates and outputs the matching weight between the corresponding container instance and the node; The weight scoring model has built-in weight scoring rules; according to any weight scoring rule, the weight scoring model calculates a single-item matching degree based on one node parameter in the node hardware fingerprint and the corresponding item metadata of the behavior fingerprint; the weight scoring model multiplies all the single-item matching degrees calculated based on each weight scoring rule to obtain the matching weight between the corresponding container instance and the node.

5. The Kubernetes cluster one-key deployment and lifecycle management method according to claim 4, characterized in that: The service characteristics of the container instance include at least one of the hardware type required by the container instance, the runtime performance metrics of the container instance, and the SLA definition of the container instance; wherein, the hardware type required by the container instance is extracted based on the explicit hardware dependencies of the corresponding container image; the runtime performance metrics of the container instance are extracted based on the runtime behavior constraints of the corresponding container image; the SLA definition of the container instance is extracted based on the system environment dependencies and edge computing requirements of the corresponding container image; The resource configuration policy is a Kubernetes object description file, including at least one of a resource allocation policy, a scheduling constraint policy, and an elastic scaling policy.

6. The method for one-click deployment and life cycle management of a Kubernetes cluster according to claim 5, characterized in that: The resource configuration template is a structured configuration file, including a resource allocation rule, a scheduling constraint rule, and an elastic scaling rule; The resource allocation rule includes: generating a resource allocation policy based on the hardware type required by the container instance; The scheduling constraint rule includes: generating a scheduling constraint policy based on the runtime performance metrics of the container instance; The elastic scaling rule includes: generating an elastic scaling policy based on the SLA definition of the container instance.

7. The method for one-click deployment and life cycle management of a Kubernetes cluster according to claim 6, wherein: Splitting the resource configuration policy into independent configuration actions specifically includes: Based on the declarative syntax of Kubernetes resource objects, parsing the resource configuration policy into atomic operation units; Performing a statelessness check on each atomic operation unit; the method of performing the statelessness check includes explicit dependency detection; the explicit dependency detection includes: identifying the explicit dependencies between atomic operation units by statically analyzing the resource references in the atomic operation units; Encapsulating the atomic operation units that pass the statelessness check into configuration actions; any configuration action is a set of configuration instructions.

8. The Kubernetes cluster one-click deployment and lifecycle management method according to claim 7, characterized in that: The parallel gray-box testing of the test group specifically includes: selecting container instances through a label selector to form a gray-box instance group; deploying the configuration actions in the test group to the gray-box instance group; routing production traffic to the gray-box instance group and blocking the write operations of the container instances in the gray-box instance group to the production data storage; collecting the performance metrics of the gray-box instance group and determining whether the test group passes the test based on the performance metrics.

9. The Kubernetes cluster one-key deployment and lifecycle management method according to claim 8, wherein: The test results of the parallel gray-box testing include passing the test and failing the test; The full-scale deployment or status rollback of the container instances specifically includes: If the test group passes the test, deploy the configuration actions in the test group to all container instances; If the test group fails the test, identify the conflicting actions, where the conflicting actions are the configuration actions that cause the test group to fail the test, and roll back the deployment status of the conflicting actions in the gray-box instance group.

10. A one-click deployment and lifecycle management system for Kubernetes clusters, which is used to implement the one-click deployment and lifecycle management method for Kubernetes clusters described in any one of claims 1-9, characterized in that, It includes a fingerprint generation module, a node allocation module, a resource configuration module, a gray-box testing module, and a deployment management module; among them: The fingerprint generation module is used to collect the node parameters and the metadata of the container images, generate the node hardware fingerprints for each node based on the node parameters, and generate the behavior fingerprints for each container image based on the metadata; The node allocation module allocates the globally optimal nodes for each container instance based on the node hardware fingerprints and the behavior fingerprints of the container images; The resource configuration module generates a resource configuration policy based on the business characteristics of the container instances and splits the resource configuration policy into independent configuration actions; The gray-box testing module is used to perform parallel gray-box testing on the configuration actions and feedback the results of the parallel gray-box testing to the deployment management module; The deployment management module performs full-scale deployment or status rollback of the container instances according to the results of the parallel gray-box testing.

Citation Information

Patent Citations

  • Deployment method of Kubernetes cluster

    CN116860386A

  • Managing deployment of workloads

    CN114610477A