Kubernetes cluster one-key deployment and life cycle management method and system
By generating node hardware fingerprints and container mirroring behavior fingerprints, global optimal nodes are assigned to container instances, and automated configuration and parallel grayscale testing are performed through business semantic-driven resource configuration templates, the adaptability and stability of Kubernetes clusters in complex business scenarios is solved, and efficient deployment and operation and maintenance are achieved.
Patent Information
- Application Number
- CN202510513092.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing technology lacks the ability to adapt the dynamic characteristics of business scenarios in the deployment and life cycle management of Kubernetes clusters, resulting in unreasonable resource allocation, increasing the deployment failure rate and operation and maintenance costs, and grayscale testing cannot comprehensively evaluate the impact of configuration changes on the system.
By collecting node parameters and container mirror metadata, the node hardware fingerprint and container mirror behavior fingerprint are generated. Based on these fingerprints, the global optimal node is assigned to the container instance, and the business characteristics are automatically mapped to the Kubernetes resource configuration policy through the business semantic-driven resource configuration template, split into independent configuration actions for parallel grayscale testing, and finally, full deployment or state rollback is performed on the global optimal node.
It improves the adaptability and stability of Kubernetes clusters in complex business scenarios, reduces deployment failure rate and operation and maintenance costs, and improves the accuracy and efficiency of resource utilization rate and grayscale testing.
Smart Images

Figure CN120029637A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of application control, and in particular to a one-click deployment and lifecycle management method and system for a Kubernetes cluster. Background Art
[0002] At present, the one-click deployment technology of Kubernetes clusters has made certain progress. Some tools and platforms can realize basic cluster deployment functions and help users quickly build a Kubernetes environment. However, the existing technology still has some shortcomings.
[0003] In terms of matching nodes with container instances, existing methods are mostly based on resource requirements and allocation strategies. For example, they consider static hardware dependencies such as the number of CPU cores and memory size, but ignore the dynamic behavior and implicit dependencies of container instances at runtime. As a result, in actual deployment, the image may run normally in local tests but crash in the cluster environment, increasing the deployment failure rate and operation and maintenance costs.
[0004] Traditional container instance resource configuration and updates rely on preset rules, such as setting CPU request value thresholds, which lack the ability to automatically adapt to the dynamic characteristics of business scenarios. Business loads may change at any time, and static configuration rules cannot respond in time, resulting in resource waste or insufficient resources. When the business peak of the Kubernetes cluster arrives, the pre-set resources may not meet the demand, resulting in a decline in application performance; and during the business trough, excessive resource allocation will cause idle resources and reduce resource utilization.
[0005] Existing technologies also have deficiencies in the testing phase of the deployment and update process. The current grayscale testing method lacks effective analysis and processing of policy-level dependencies between configuration actions. Factors such as the coordinated operation of different configuration actions in historical operation data and fault correlation have not been fully considered, making it impossible for grayscale testing to fully evaluate the impact of configuration changes on the entire system, increasing the risk of business anomalies after full deployment.
[0006] For example, the patent application with publication number CN116860386A discloses a method for deploying a Kubernetes cluster, including the following steps: dividing the deployment process of the Kubernetes cluster into several independent tasks; further abstracting the divided tasks and defining a unified structure so that each actual task becomes an implementation instance of the structure; configuring appropriate features for each task based on the definition of the task; organizing and choreographing all tasks in sequence based on the Kubernetes cluster deployment process; at the end of all task executions, calling an interface to write the results to the management platform database, so that users can view the deployment results on the platform. This application solves the problem that an error in a one-time deployment method affects the global situation, a node failure affects the overall process, and the process must be started from scratch after failure, but the problem raised by this application and this background technology still exists: lack of automatic adaptation capabilities for dynamic characteristics of business scenarios.
[0007] The information disclosed in this background technology section is only intended to enhance the understanding of the overall background of the application and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to ordinary technicians in the field. Summary of the invention
[0008] The technical problem to be solved by this application is to overcome the defects of the existing technology, provide a one-click deployment and lifecycle management method and system for Kubernetes clusters, and enhance the adaptability and stability of Kubernetes clusters in complex business scenarios.
[0009] In order to solve the above technical problems, this application provides the following technical solutions:
[0010] On the one hand, this application provides a one-click deployment and lifecycle management method for a Kubernetes cluster, including the following steps:
[0011] Collect node parameters; generate node hardware fingerprints based on the node parameters; parse metadata of the container image; generate behavior fingerprints of the container image based on the metadata;
[0012] Based on the node hardware fingerprint and the behavior fingerprint of the container image, a global optimal node is allocated to each container instance;
[0013] Execute the deployment of container instances on the corresponding global optimal node, including:
[0014] Match the business characteristics of the container instance with the predefined resource configuration template to obtain the resource configuration policy;
[0015] Splitting the resource configuration strategy into independent configuration actions;
[0016] Performing a parallel grayscale test on the configuration action; and performing a full deployment or state rollback of the container instance on the corresponding global optimal node based on the result of the parallel grayscale test.
[0017] As an optimal solution for the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, wherein: the node parameters include hardware topology parameters, system environment parameters, runtime state parameters, and edge device parameters; the metadata includes explicit hardware dependencies, system environment dependencies, runtime behavior constraints, and edge computing requirements; any node parameter corresponds to a metadata.
[0018] The node hardware fingerprint is structured data containing all node parameters; the behavior fingerprint of the container image is structured data containing all metadata; the method of generating the node hardware fingerprint and the behavior fingerprint of the container image is as follows: the node parameters and the metadata of the container image are structuredly encoded respectively using the same encoding method to obtain the node hardware fingerprint and the behavior fingerprint of the container image.
[0019] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, the globally optimal node is allocated to each container instance based on the node hardware fingerprint and the behavior fingerprint of the container image, specifically including:
[0020] Obtain the behavior fingerprint of the container image corresponding to each container instance as the behavior fingerprint of the corresponding container instance;
[0021] Based on the node hardware fingerprint and the behavior fingerprint of the container instance, the matching weight between any container instance and any node is calculated;
[0022] Construct a weight matrix; each row in the weight matrix corresponds to a container instance, each column corresponds to a node, and the element value of each element is the matching weight between the corresponding container instance and the node;
[0023] The weight matrix is processed by an optimization algorithm to obtain a final matching solution; the final matching solution includes a set of element values selected in the weight matrix; wherein each row in the weight matrix has one and only one element value selected, and any selected element value is not 0;
[0024] A global optimal node is allocated to each container instance based on the final matching solution.
[0025] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, the calculation of the matching weight between any container instance and any node specifically includes:
[0026] Obtain a weight scoring model for the container instance and the node; input the node hardware fingerprint and the behavior fingerprint of the container instance into the weight scoring model, and the weight scoring model calculates and outputs the matching weight between the corresponding container instance and the node;
[0027] The weight scoring model has built-in weight scoring rules; the weight scoring model calculates a single matching degree based on a node parameter in the node hardware fingerprint and the corresponding metadata of the behavior fingerprint according to any weight scoring rule; the weight scoring model multiplies all single matching degrees calculated based on each weight scoring rule to obtain a matching weight between the corresponding container instance and the node.
[0028] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in the present application, wherein: the business characteristics of the container instance include at least one of the hardware type required for the container instance, the runtime performance indicators of the container instance, and the SLA definition of the container instance; wherein the hardware type required for the container instance is extracted based on the explicit hardware dependency of the corresponding container image; the runtime performance indicators of the container instance are extracted based on the runtime behavior constraints of the corresponding container image; the SLA definition of the container instance is extracted based on the system environment dependency and edge computing requirements of the corresponding container image;
[0029] The resource configuration strategy is a Kubernetes object description file, including at least one of a resource allocation strategy, a scheduling constraint strategy, and an elastic scaling strategy.
[0030] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, wherein: the resource configuration template is a structured configuration file, including resource allocation rules, scheduling constraint rules, and elastic expansion rules;
[0031] The resource allocation rules include: generating a resource allocation policy based on the hardware type required by the container instance;
[0032] The scheduling constraint rules include: generating a scheduling constraint strategy based on runtime performance indicators of the container instance;
[0033] The elastic scaling rule includes: generating an elastic scaling strategy based on the SLA definition of the container instance.
[0034] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, the resource configuration strategy is split into independent configuration actions, specifically including:
[0035] Based on the declarative syntax of Kubernetes resource objects, the resource configuration policy is parsed into atomic operation units;
[0036] Performing a stateless check on each atomic operation unit; the stateless check includes explicit dependency detection; the explicit dependency detection includes: identifying explicit dependencies between atomic operation units by statically analyzing resource references in the atomic operation units;
[0037] The atomic operation units that pass the stateless check are encapsulated as configuration actions; any configuration action is a collection of configuration instructions.
[0038] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, the configuration action is subjected to parallel grayscale testing, specifically including:
[0039] Parse the fields defined in the Kubernetes resources in each configuration action to identify hardware resource conflicts between configuration actions;
[0040] Obtain historical operation logs of the container instance; and identify policy-level dependencies between configuration actions based on the historical operation logs;
[0041] Select configuration actions to form test groups; there is no hardware resource conflict and policy-level dependency between any two configuration actions in any test group;
[0042] Parallel grayscale testing is performed on the test group, specifically including: selecting container instances through a label selector to form a grayscale instance group; deploying configuration actions in the test group to the grayscale instance group; routing production traffic to the grayscale instance group, and blocking the container instances in the grayscale instance group from writing to the production data storage; collecting performance indicators of the grayscale instance group, and judging whether the test group passes the test based on the performance indicators.
[0043] As an optimal solution for the one-click deployment and lifecycle management method of the Kubernetes cluster described in the present application, the policy-level dependency includes: the frequency of collaborative operation of any two configuration actions in the historical operation data exceeds a preset frequency threshold; the collaborative operation is the synchronous execution of configuration instructions of different configuration actions; the policy-level dependency also includes: different configuration actions have fault correlation in the historical operation data; the fault correlation includes the correlation of fault triggering and the correlation of fault recovery.
[0044] As a preferred solution of the one-click deployment and lifecycle management method of the Kubernetes cluster described in this application, the test results of the parallel grayscale test include pass test and fail test;
[0045] The full deployment or state rollback of the container instance specifically includes:
[0046] If the test group passes the test, the configuration actions in the test group are deployed to all container instances;
[0047] If the test group fails the test, a conflicting action is identified, where the conflicting action is a configuration action that causes the test group to fail the test, and the deployment state of the conflicting action in the grayscale instance group is rolled back.
[0048] In the second aspect, the present application provides a Kubernetes cluster one-click deployment and lifecycle management system, including a fingerprint generation module, a node allocation module, a resource configuration module, a grayscale testing module, and a deployment management module; wherein:
[0049] The fingerprint generation module is used to collect node parameters and metadata of container images, and generate a node hardware fingerprint of each node based on the node parameters, and generate a behavioral fingerprint of each container image based on the metadata;
[0050] The node allocation module allocates the global optimal node to each container instance based on the node hardware fingerprint and the behavior fingerprint of the container image;
[0051] The resource configuration module generates resource configuration policies based on the business characteristics of the container instance and splits the resource configuration policies into independent configuration actions;
[0052] The grayscale test module is used to perform parallel grayscale tests on configuration actions and feed back the results of the parallel grayscale tests to the deployment management module;
[0053] The deployment management module performs full deployment or status rollback of container instances based on the results of the parallel grayscale test.
[0054] Compared with the prior art, the beneficial effects achieved by this application are as follows:
[0055] This application comprehensively considers the dynamic behavior and implicit dependencies of static data and container instances during runtime, builds node hardware fingerprints and container image behavior fingerprints, and allocates the global optimal node to the container instance through fingerprint matching and verification, solving the problem that the application runs normally locally but crashes in the cluster.
[0056] By building a business semantics-driven resource configuration template, the business characteristics of the application are automatically mapped to the Kubernetes resource configuration policy, and complex cluster-level changes are decomposed into multiple independently verifiable atomic operation units, reducing the reliance on manual experience, improving the accuracy and efficiency of grayscale testing, reducing the business anomaly rate caused by configuration changes, and ensuring the stability and reliability of the deployment process. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:
[0058] Figure 1 A flowchart of the one-click deployment and lifecycle management method for the Kubernetes cluster provided for this application;
[0059] Figure 2 Schematic diagram of the structure of the Kubernetes cluster one-click deployment and lifecycle management system provided for this application. DETAILED DESCRIPTION
[0060] The technical solution of the present application is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.
[0061] Example 1
[0062] This example introduces a one-click deployment and lifecycle management method for a Kubernetes cluster. Figure 1 , the method comprises the following steps:
[0063] Collect node parameters; generate node hardware fingerprints based on the node parameters; parse metadata of the container image; generate behavior fingerprints of the container image based on the metadata;
[0064] The node parameters include hardware topology parameters, system environment parameters, runtime state parameters, and edge device parameters; the metadata include explicit hardware dependencies, system environment dependencies, runtime behavior constraints, and edge computing requirements; any node parameter corresponds to a metadata.
[0065] Some of the node parameters and metadata preferred in this embodiment are as follows:
[0066] The hardware topology parameters in the node parameters include GPU architecture, number of physical cores, and storage medium type; the explicit hardware dependencies in the metadata include GPU architecture requirements, physical core number requirements, and storage medium type requirements; the correspondence between node parameters and metadata includes: GPU architecture corresponds to GPU architecture requirements, physical core number corresponds to physical core number requirements, and storage medium type corresponds to storage medium type requirements. The system environment parameters in the node parameters include kernel version and system call set; the system environment dependencies in the metadata include kernel version requirements and system call whitelist; the correspondence between node parameters and metadata also includes: kernel version corresponds to kernel version requirements, and system call set corresponds to system call whitelist. The runtime state parameters in the node parameters include interrupt affinity configuration, clock synchronization protocol, and NUMA node architecture; the runtime behavior constraints in the metadata include interrupt request binding requirements, clock synchronization accuracy requirements, and NUMA architecture dependencies; the edge device parameters in the node parameters include hot-swappable interfaces and cryptographic acceleration hardware; the edge computing requirements in the metadata include hot-swappable device support and cryptographic acceleration requirements; the correspondence between node parameters and metadata also includes: hot-swappable interfaces correspond to hot-swappable device support; and cryptographic acceleration hardware corresponds to cryptographic acceleration requirements.
[0067] The node hardware fingerprint is structured data containing all node parameters; the behavior fingerprint of the container image is structured data containing all metadata; the method of generating the node hardware fingerprint and the behavior fingerprint of the container image is as follows: the node parameters and the metadata of the container image are structuredly encoded respectively using the same encoding method to obtain the node hardware fingerprint and the behavior fingerprint of the container image. This embodiment preferably uses JSON or Protobuf format as the structured encoding format of the node hardware fingerprint and the behavior fingerprint, which can retain the readability of the original data and support matching and compatibility verification between fingerprints. By constructing the behavior fingerprint of the node hardware fingerprint and the container image, the present application not only considers static data (such as explicit hardware dependencies), but also considers the dynamic behavior and implicit dependencies of the container instance during runtime (such as interrupt request binding requirements, clock synchronization accuracy requirements, etc.). By matching and verifying the container instance and the node through the node hardware fingerprint and the behavior fingerprint, the deployment failure rate of the container instance can be reduced, which is more suitable for the heterogeneous hardware environment of the Kubernetes cluster, and can solve problems such as the container running normally locally but crashing in the cluster.
[0068] Based on the node hardware fingerprint and the behavior fingerprint of the container image, a global optimal node is allocated to each container instance; specifically, the following steps are included:
[0069] Obtain the behavior fingerprint of the container image corresponding to each container instance as the behavior fingerprint of the corresponding container instance;
[0070] Based on the node hardware fingerprint and the behavior fingerprint of the container instance, the matching weight between any container instance and any node is calculated;
[0071] Construct a weight matrix; each row in the weight matrix corresponds to a container instance, each column corresponds to a node, and the element value of each element is the matching weight between the corresponding container instance and the node;
[0072] The weight matrix is processed by an optimization algorithm to obtain a final matching solution; the final matching solution includes a set of element values selected in the weight matrix; wherein each row in the weight matrix has one and only one element value selected, and any selected element value is not 0;
[0073] Selecting an element value from each row in the weight matrix will form a matching scheme; since only one element value in each row is selected in any matching scheme, and each row in the weight matrix corresponds to a container instance, any matching scheme specifies a node for each container instance. This embodiment preferably uses the KM algorithm as the optimization algorithm, and uses the KM algorithm to find a matching scheme that maximizes the cumulative weight (i.e., the cumulative value of a selected set of element values) as the final matching scheme, so that the node specified for each container is its global optimal node. In the process of the KM algorithm exploring the final matching scheme, the CPU utilization of the node is used as a constraint to prevent multiple container instances from exceeding the processing capacity of the node when they are assigned to one node.
[0074] A global optimal node is allocated to each container instance based on the final matching scheme. Each element value selected in the final matching scheme corresponds to a container instance and its corresponding global optimal node, and the global optimal node corresponding to each container instance can be determined through the corresponding relationship of each element value.
[0075] The calculation of the matching weight between any container instance and any node specifically includes:
[0076] Obtain a weight scoring model for the container instance and the node; input the node hardware fingerprint and the behavior fingerprint of the container instance into the weight scoring model, and the weight scoring model calculates and outputs the matching weight between the corresponding container instance and the node;
[0077] The weight scoring model has built-in weight scoring rules; the weight scoring model calculates a single matching degree based on a node parameter in the node hardware fingerprint and the corresponding metadata of the behavior fingerprint according to any weight scoring rule; the weight scoring model multiplies all single matching degrees calculated based on each weight scoring rule to obtain a matching weight between the corresponding container instance and the node.
[0078] The preferred partial weight scoring rules of this embodiment are as follows: If the GPU architecture and the GPU architecture requirement are completely matched, the corresponding single matching degree is 1, otherwise, the corresponding single matching degree is 0; for example, if the GPU architecture and the GPU architecture requirement are both Ampere architecture, then they are completely matched. If the kernel version meets the kernel version requirement, the corresponding single matching degree is 1, otherwise, the corresponding single matching degree is 0; for example, if the kernel version is 5.1.1 and the kernel version requirement is greater than or equal to 5.1.0, then the kernel version meets the kernel version requirement.
[0079] If the system call set covers all system calls in the system call whitelist, the corresponding single item matching degree is 1, otherwise, the corresponding single item matching degree is 0; if the number of physical cores does not meet the physical core number requirement, the corresponding single item matching degree is 0, otherwise, the corresponding single item matching degree is assigned a value based on the number of physical cores. The more physical cores there are, the higher the single item matching degree is, and the value range of the single item matching degree is ; For example, if the number of physical cores required is no less than 8 cores, set the upper limit of the number of physical cores to , if the number of physical cores is not less than , the single item matching degree is 1; if the number of physical cores is less than 8, the single item matching degree is 0; if the number of physical cores is not less than 8 and less than , the value range is Since the matching weight between a container instance and a node is the product of all the individual matching degrees, when any individual matching degree is 0, the matching weight is also 0, and the node with a matching weight of 0 will not be used as the global optimal node for the corresponding container instance, thereby filtering out unmatched nodes through hard conditions.
[0080] Execute the deployment of container instances on the corresponding global optimal node, including:
[0081] Match the business characteristics of the container instance with the predefined resource configuration template to obtain the resource configuration policy;
[0082] The business characteristics of the container instance include at least one of the hardware type required for the container instance, the runtime performance indicators of the container instance, and the SLA definition (i.e., service level agreement definition) of the container instance; wherein the hardware type required for the container instance is extracted based on the explicit hardware dependencies of the corresponding container image, including GPU architecture, storage medium type, encryption acceleration hardware requirements, etc.; the runtime performance indicators of the container instance are extracted based on the runtime behavior constraints of the corresponding container image, including NUMA node architecture dependencies, interrupt request binding requirements, clock synchronization accuracy requirements, etc.; the SLA definition of the container instance is extracted based on the system environment dependencies and edge computing requirements of the corresponding container image, including the minimum number of replicas, maximum latency threshold, etc.;
[0083] The resource configuration strategy is an executable Kubernetes object description file, including at least one of a resource allocation strategy, a scheduling constraint strategy, and an elastic scaling strategy.
[0084] The resource configuration template is a structured configuration file, including resource allocation rules, scheduling constraint rules, and elastic scaling rules. This embodiment preferably uses YAML or JSON as the specific format of the resource configuration template, and associates the resource configuration template with the metadata of the container image through a version identifier to extract the business characteristics of the container instance from the corresponding metadata.
[0085] The resource allocation rules include: generating a resource allocation strategy based on the hardware type required by the container instance; the resource allocation strategy includes a device plug-in binding declaration and a restriction field of Kubernetes resources;
[0086] The scheduling constraint rules include: generating a scheduling constraint strategy based on the runtime performance indicators of the container instance; the scheduling constraint strategy includes node affinity configuration, NUMA core binding strategy, interrupt affinity declaration, etc.;
[0087] The elastic scaling rules include: generating an elastic scaling strategy based on the SLA definition of the container instance; the elastic scaling strategy includes a horizontal scaling strategy, a replica number configuration, a resource utilization scaling threshold, etc.
[0088] Split the resource configuration strategy into independent configuration actions; specifically including:
[0089] Based on the declarative syntax of Kubernetes resource objects, the resource configuration policy is parsed into atomic operation units; the atomic operation units include resource allocation declarations, scheduling constraint declarations, runtime optimization declarations, and network policy declarations;
[0090] Perform stateless verification on each atomic operation unit;
[0091] Methods for stateless verification include explicit dependency detection, specifically, by statically analyzing resource references in atomic operation units, identifying explicit dependencies between atomic operation units, such as service selection dependencies, configuration mapping name dependencies, etc. Dependency detection can ensure that the execution of any atomic operation unit does not depend on the output of other operation units. Methods for stateless verification also include resource isolation verification, ensuring that each atomic operation unit only operates independent Kubernetes resources, and only the corresponding Kubernetes resource objects need to be deleted when the state is rolled back. If the atomic operation unit fails the stateless verification, the atomic operation unit is manually reconstructed, such as manually merging atomic operation units with explicit dependencies.
[0092] The atomic operation units that pass the stateless check are encapsulated as configuration actions; any configuration action is a set of configuration instructions that can be executed independently. The configuration instructions contained in the configuration action include Kubernetes native API call commands, CRD operation commands, custom controller coordination instructions, etc. In this embodiment, the configuration action is stored in the form of a YAML manifest file or a Helm Chart template, and each configuration action is associated with a unique identifier to support the status tracking of subsequent grayscale testing.
[0093] Performing a parallel grayscale test on the configuration action; and performing a full deployment or state rollback of the container instance on the corresponding global optimal node based on the result of the parallel grayscale test.
[0094] The configuration actions are subjected to parallel grayscale testing, specifically including:
[0095] Parse the fields defined in the Kubernetes resources in each configuration action to identify hardware resource conflicts between configuration actions; identify hardware resource conflicts based on the field association of the Kubernetes resource definition, which include mutually exclusive access declarations of different configuration actions to the same hardware device (such as GPU, storage volume), over-allocation of resource quotas (such as CPU, memory quota) of different configuration actions, and competition of different configuration actions for underlying physical resources (such as LLC cache, memory bandwidth).
[0096] Obtain the historical operation log of the container instance; identify the policy-level dependency between configuration actions based on the historical operation log; perform logical segmentation on the historical operation log of the deployed container instance to obtain the atomic operation unit consistent with the current configuration action, thereby extracting the historical operation data of the configuration action; perform statistical analysis on the collaborative changes between configuration actions based on the historical operation data of the configuration action to obtain the policy-level dependency between configuration actions; the policy-level dependency includes: the frequency of collaborative operations of any two configuration actions in the historical operation data exceeds a preset frequency threshold; the collaborative operation is the synchronous execution of configuration instructions of different configuration actions; the policy-level dependency also includes: the existence of fault correlation between different configuration actions in the historical operation data; the fault correlation includes the correlation of fault triggering, that is, different configuration actions have associated faults, for example, configuration action A is a database connection pool configuration, and configuration action B is a database service Endpoint declaration, then when configuration action B fails due to network errors or other reasons, configuration action A will trigger a fault due to the inability to connect to the database. The fault correlation also includes the correlation of fault recovery, that is, different configuration actions have logical dependencies in fault recovery. For example, configuration action C is storage volume mounting, and configuration action D is storage class declaration. During fault recovery, configuration action C must be executed after configuration action D is successfully rolled back. Otherwise, data corruption will occur because the storage volume cannot be unmounted.
[0097] Atomic operation units are selected to form test groups; there is no hardware resource conflict and policy-level dependency between any two atomic operation units in any test group;
[0098] Conduct parallel grayscale testing on the test groups;
[0099] The steps of the parallel grayscale test are as follows:
[0100] Use the label selector to select container instances to form a grayscale instance group. For example, select 10% of the container instances as the grayscale instance group and mark the value of the canary label of each selected container instance as true.
[0101] Deploy the configuration actions in the test group to the grayscale instance group;
[0102] Route a specified proportion of production traffic to the grayscale instance group, and block the container instances in the grayscale instance group from writing to the production data store. For example, route 10% of production traffic to the grayscale instance group through the service grid, and then block the grayscale instance group from writing to the production data store through database shadow tables, message queue mirror topics, cache replica isolation, etc., to isolate test data from production data.
[0103] Collect performance indicators of grayscale instance groups, such as collection latency, error rate, resource utilization, and other performance indicators;
[0104] The test group is judged whether it has passed the test based on the preset performance indicator threshold. For example, if the error rate is less than 0.1%, the test has passed.
[0105] Record the test results of parallel grayscale testing and split the test groups that have completed parallel grayscale testing into independent atomic operation units. Split the test groups into independent configuration actions to provide independent operation objects for subsequent full deployment or status rollback.
[0106] The embodiment of the present application first ensures that there are no explicit dependencies between configuration actions through stateless verification, and then identifies implicit dependencies between configuration actions (hardware resource conflicts, policy-level dependencies), thereby achieving more sophisticated dependency management and optimizing the accuracy and efficiency of parallel grayscale testing.
[0107] The test results of the parallel grayscale test include a pass test and a fail test;
[0108] Based on the results of the parallel grayscale test, the container instance is fully deployed or its status is rolled back on the corresponding global optimal node, including:
[0109] If the test group passes the test, the configuration actions in the test group are deployed to all container instances. For example, through progressive traffic switching, the configuration actions in the grayscale instance group are gradually synchronized to all container instances.
[0110] If the test group fails the test, a conflicting action is identified, where the conflicting action is a configuration action that causes the test group to fail the test, and the deployment state of the conflicting action in the grayscale instance group is rolled back.
[0111] Based on real-time grayscale test data, conflicting actions can be identified in the test group. For example, the monitoring logs of parallel grayscale tests are recorded in real time. Error events in the monitoring logs can be used to identify conflicting actions. If the test group fails the parallel grayscale test, the deployment status of the conflicting action is rolled back, including deleting the Kubernetes resource definition of the conflicting action (such as Deployment, ConfigMap, etc.), deleting temporary data written by the conflicting unit during the test (such as database shadow tables), etc. There is no need to roll back other actions in the test group except the conflicting action, and retain their deployment status in the grayscale instance group without full deployment. Subsequently, a new test group is formed with other configuration actions. When the new test group passes the parallel grayscale test, the configuration actions in the new test group are deployed to all container instances.
[0112] The resource configuration and update of traditional container instances mostly rely on static rules preset by operation and maintenance personnel (such as manually setting CPU request value thresholds), and lack the ability to automatically adapt to the dynamic characteristics of business scenarios. This application achieves intelligent optimization through the following technical paths: First, build a business semantics-driven resource configuration template to automatically map the business characteristics of the application to Kubernetes resource configuration policies, and then use declarative resource change splitting to decouple complex cluster-level changes into multiple independently verifiable atomic operation units; finally, through dynamic traffic coloring and isolation technology, implement parallel grayscale verification of multi-dimensional resource configuration strategies to achieve progressive deployment decisions. The technical path of this application achieves the optimization effect of improving resource utilization while reducing dependence on manual experience, and reduces the business anomaly rate caused by configuration changes.
[0113] The one-click deployment and lifecycle management method for Kubernetes clusters described in the embodiments of this application is not only applicable to the initial installation phase of container instances, but also to the update and upgrade phase of applications. In the application upgrade scenario, by dynamically adapting the version differences before and after the upgrade, smooth rolling upgrades and grayscale releases are achieved. Specifically, when upgrading an application, the following key steps need to be adaptively adjusted:
[0114] When generating the behavioral fingerprint of the container image, it is necessary to parse the metadata of the upgraded image and generate differentiated behavioral fingerprints based on the incremental changes of the version upgrade. For example, if the upgraded container image adds runtime behavioral constraints on the real-time process scheduling policy (such as requiring the CPU core isolation mechanism to be disabled), it is necessary to expand the corresponding scheduler configuration requirement field in the behavioral fingerprint and expand the kernel real-time patch status in the node hardware fingerprint to incrementally match the new runtime behavioral constraints. Preferably, version compatibility verification rules are introduced in the weighted scoring model. For example, when the intersection of the upgraded system call whitelist and the node's system call set covers the old version whitelist, it is judged as a compatibility match, thereby avoiding node mismatch problems caused by version upgrades.
[0115] Furthermore, when constructing the weight matrix, the matching weights are dynamically modified in combination with the rolling upgrade strategy. For example, the node allocation status of the old version container instance is used as a constraint to ensure that the new version container instance is preferentially scheduled to idle nodes or low-load nodes, thereby reducing resource contention during the upgrade process. In addition, when using the optimization algorithm to process the weight matrix, version affinity constraints are introduced, such as taking the uniformity of version distribution of multiple container instances of the same service as the optimization goal to prevent the risk of single point failure caused by centralized version deployment.
[0116] Furthermore, in the resource configuration strategy generation stage, a version transition strategy is generated based on the differences in business characteristics before and after the upgrade. For example, if the upgraded container instance needs to perform data migration, the data migration rules are extended in the resource configuration template to generate a composite resource configuration strategy that includes storage volume declaration conversion, database schema version management, and cache data preheating strategy. Preferably, in the stateless verification of the atomic operation unit, isolation verification of cross-version resource references is added, such as detecting name conflicts between the new and old version configuration mappings, to ensure that the state can be restored by deleting a single resource object when the version is rolled back.
[0117] Furthermore, in the parallel grayscale testing phase, the traffic coloring and isolation mechanism is optimized for upgrade scenarios. For example, the mixed deployment status of the old and new versions of container instances is tested in parallel in the grayscale instance group, and the production requests are routed to the old and new versions of container instances at the same time through the traffic mirroring function of the service grid, and the performance indicator differences between the two are compared to verify the upgrade compatibility. If the test fails, the behavior fingerprint and resource configuration template of the old version are quickly rolled back based on the version identifier, and the incrementally generated difference fields are retained as upgrade failure logs for subsequent version iteration analysis.
[0118] Through the above-mentioned adaptive adjustment, the embodiment of the present application can effectively identify the implicit dependency changes caused by version differences in the application upgrade scenario, realize the dynamic adaptation of resource allocation strategies and precise control of grayscale releases, and improve the success rate and efficiency of cluster upgrades while ensuring service continuity, thereby realizing the lifecycle management of the Kubernetes cluster.
[0119] Example 2
[0120] This embodiment is the second embodiment of the present application; it is based on the same inventive concept as embodiment 1, and Figure 2 This embodiment introduces a one-click deployment and lifecycle management system for a Kubernetes cluster, including a fingerprint generation module, a node allocation module, a resource configuration module, a grayscale testing module, and a deployment management module; wherein:
[0121] The fingerprint generation module is used to collect node parameters and metadata of container images, and generate a node hardware fingerprint of each node based on the node parameters, and generate a behavioral fingerprint of each container image based on the metadata;
[0122] The node allocation module allocates the global optimal node to each container instance based on the node hardware fingerprint and the behavioral fingerprint of the container image; based on the node hardware fingerprint and the behavioral fingerprint of the container instance, this module calculates the matching weight between the container instance and the node, constructs a weight matrix and uses an optimization algorithm to process the weight matrix to obtain the final matching solution, thereby allocating the global optimal node to each container instance.
[0123] The resource configuration module generates a resource configuration policy based on the business characteristics of the container instance and splits the resource configuration policy into independent configuration actions. The module is configured with a resource configuration template, maps the business characteristics of the container instance into a resource configuration policy based on the resource configuration template, and splits the resource configuration policy into independent configuration actions through stateless verification.
[0124] The grayscale testing module is used to perform parallel grayscale testing on configuration actions and feed back the results of parallel grayscale testing to the deployment management module. This module generates test groups by identifying and avoiding hardware resource conflicts and policy-level dependencies between configuration actions, and performs parallel grayscale testing on the test groups to improve testing efficiency and accuracy.
[0125] The deployment management module performs full deployment or status rollback of container instances based on the results of the parallel grayscale test. If the test group passes the test, the deployment management module deploys the configuration actions in the test group to all container instances; if the test group fails the test, the deployment management module rolls back the deployment status of the conflicting actions in the grayscale instance group, while retaining the deployment status of other actions in the test group except the conflicting actions in the grayscale instance group, so that they can be combined with other configuration actions to form a new test group later.
[0126] The specific functions of the above modules are implemented by referring to the relevant contents of the one-click deployment and lifecycle management method of the Kubernetes cluster described in Example 1, which will not be repeated here.
[0127] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0128] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose and scope of protection of the present application, all of which are within the protection of the present application.
Claims
1. A one-click deployment and lifecycle management method for Kubernetes clusters, characterized in that: The following steps are involved: Collecting node parameters; generating node hardware fingerprints based on the node parameters; Parsing metadata of the container image; generating a behavioral fingerprint of the container image based on the metadata; Based on the node hardware fingerprint and the behavior fingerprint of the container image, a global optimal node is allocated to each container instance; Execute the deployment of container instances on the corresponding global optimal node, including: Match the business characteristics of the container instance with the predefined resource configuration template to obtain the resource configuration policy; Splitting the resource configuration strategy into independent configuration actions; Performing a parallel grayscale test on the configuration action; and based on the result of the parallel grayscale test, performing full deployment or state rollback of the container instance on the corresponding global optimal node.
2. The one-click deployment and lifecycle management method for Kubernetes clusters according to claim 1, characterized in that: The node parameters include hardware topology parameters, system environment parameters, runtime state parameters, and edge device parameters; the metadata include explicit hardware dependencies, system environment dependencies, runtime behavior constraints, and edge computing requirements; any node parameter corresponds to a metadata; The node hardware fingerprint is structured data containing all node parameters; the behavior fingerprint of the container image is structured data containing all metadata; the method of generating the node hardware fingerprint and the behavior fingerprint of the container image is as follows: the node parameters and the metadata of the container image are structuredly encoded respectively using the same encoding method to obtain the node hardware fingerprint and the behavior fingerprint of the container image.
3. The one-click deployment and lifecycle management method for Kubernetes clusters as described in claim 2, characterized in that: Based on the node hardware fingerprint and the behavior fingerprint of the container image, a global optimal node is allocated to each container instance, specifically including: Obtain the behavior fingerprint of the container image corresponding to each container instance as the behavior fingerprint of the corresponding container instance; Based on the node hardware fingerprint and the behavior fingerprint of the container instance, the matching weight between any container instance and any node is calculated; Construct a weight matrix; each row in the weight matrix corresponds to a container instance, each column corresponds to a node, and the element value of each element is the matching weight between the corresponding container instance and the node; The weight matrix is processed by an optimization algorithm to obtain a final matching solution; the final matching solution includes a set of element values selected in the weight matrix; wherein each row in the weight matrix has one and only one element value selected, and any selected element value is not 0; A global optimal node is allocated to each container instance based on the final matching solution.
4. The one-click deployment and lifecycle management method for Kubernetes clusters as described in claim 3, characterized in that: The calculation of the matching weight between any container instance and any node specifically includes: Obtain a weight scoring model for the container instance and the node; input the node hardware fingerprint and the behavior fingerprint of the container instance into the weight scoring model, and the weight scoring model calculates and outputs the matching weight between the corresponding container instance and the node; The weight scoring model has built-in weight scoring rules; the weight scoring model calculates a single matching degree based on a node parameter in the node hardware fingerprint and the corresponding metadata of the behavior fingerprint according to any weight scoring rule; the weight scoring model multiplies all single matching degrees calculated based on each weight scoring rule to obtain a matching weight between the corresponding container instance and the node.
5. The one-click deployment and lifecycle management method for Kubernetes clusters as described in claim 4, characterized in that: The business characteristics of the container instance include at least one of the hardware type required for the container instance, the runtime performance indicators of the container instance, and the SLA definition of the container instance; wherein the hardware type required for the container instance is extracted based on the explicit hardware dependency of the corresponding container image; the runtime performance indicators of the container instance are extracted based on the runtime behavior constraints of the corresponding container image; the SLA definition of the container instance is extracted based on the system environment dependency and edge computing requirements of the corresponding container image; The resource configuration strategy is a Kubernetes object description file, including at least one of a resource allocation strategy, a scheduling constraint strategy, and an elastic scaling strategy.
6. The one-click deployment and lifecycle management method for Kubernetes clusters as described in claim 5, characterized in that: The resource configuration template is a structured configuration file, including resource allocation rules, scheduling constraint rules, and elastic expansion rules; The resource allocation rules include: generating a resource allocation policy based on the hardware type required by the container instance; The scheduling constraint rules include: generating a scheduling constraint strategy based on runtime performance indicators of the container instance; The elastic scaling rule includes: generating an elastic scaling strategy based on the SLA definition of the container instance.
7. The one-click deployment and lifecycle management method for Kubernetes clusters according to claim 6, characterized in that: The resource configuration strategy is split into independent configuration actions, specifically including: Based on the declarative syntax of Kubernetes resource objects, the resource configuration policy is parsed into atomic operation units; Performing a stateless check on each atomic operation unit; the stateless check includes explicit dependency detection; the explicit dependency detection includes: identifying explicit dependencies between atomic operation units by statically analyzing resource references in the atomic operation units; The atomic operation units that pass the stateless check are encapsulated as configuration actions; any configuration action is a collection of configuration instructions.
8. The one-click deployment and lifecycle management method for Kubernetes clusters according to claim 7, characterized in that: The configuration actions are subjected to parallel grayscale testing, specifically including: Parse the fields defined in the Kubernetes resources in each configuration action to identify hardware resource conflicts between configuration actions; Obtain historical operation logs of the container instance; and identify policy-level dependencies between configuration actions based on the historical operation logs; Select configuration actions to form test groups; there is no hardware resource conflict and policy-level dependency between any two configuration actions in any test group; Parallel grayscale testing is performed on the test group, specifically including: selecting container instances through a label selector to form a grayscale instance group; deploying configuration actions in the test group to the grayscale instance group; routing production traffic to the grayscale instance group, and blocking the container instances in the grayscale instance group from writing to the production data storage; collecting performance indicators of the grayscale instance group, and judging whether the test group passes the test based on the performance indicators.
9. The one-click deployment and lifecycle management method for Kubernetes clusters as claimed in claim 8, characterized in that: The policy-level dependency includes: the frequency of collaborative operation of any two configuration actions in historical operation data exceeds a preset frequency threshold; the collaborative operation is the synchronous execution of configuration instructions of different configuration actions; the policy-level dependency also includes: the existence of fault correlation between different configuration actions in historical operation data; the fault correlation includes the correlation of fault triggering and the correlation of fault recovery.
10. The one-click deployment and lifecycle management method for Kubernetes clusters according to claim 9, characterized in that: The test results of the parallel grayscale test include a pass test and a fail test; The full deployment or state rollback of the container instance specifically includes: If the test group passes the test, the configuration actions in the test group are deployed to all container instances; If the test group fails the test, a conflicting action is identified, where the conflicting action is a configuration action that causes the test group to fail the test, and the deployment state of the conflicting action in the grayscale instance group is rolled back.
11. A one-click deployment and lifecycle management system for Kubernetes clusters, which is used to implement the one-click deployment and lifecycle management method for Kubernetes clusters as described in any one of claims 1 to 10, characterized in that it includes a fingerprint generation module, a node allocation module, a resource configuration module, a grayscale test module, and a deployment management module; wherein: The fingerprint generation module is used to collect node parameters and metadata of container images, and generate a node hardware fingerprint of each node based on the node parameters, and generate a behavioral fingerprint of each container image based on the metadata; The node allocation module allocates the global optimal node to each container instance based on the node hardware fingerprint and the behavior fingerprint of the container image; The resource configuration module generates resource configuration policies based on the business characteristics of the container instance and splits the resource configuration policies into independent configuration actions; The grayscale test module is used to perform parallel grayscale tests on configuration actions and feed back the results of the parallel grayscale tests to the deployment management module; The deployment management module performs full deployment or status rollback of container instances based on the results of the parallel grayscale test.
Citation Information
Patent Citations
Deployment method of Kubernetes cluster
CN116860386A
Managing deployment of workloads
CN114610477A
Container dynamic scheduling method and device, computer equipment and storage medium
CN114880100A
Automatically telescopic non-intrusive gray release system
CN116755764A
Kalman filter-based k8s scheduling method and device
CN119814878A
Cited By
ERP non-core service migration method based on K8s dynamic resource scheduling
CN120639862A