Stateful service management method and device and storage medium

Through custom resource definitions and Webhook services, the change tasks of stateful services are automatically handled, which solves the problem of inability to automate execution and record the change status in the existing technology, realizes the operation and maintenance management of long-connected services and abnormal self-healing, and improves the usability of the system.

CN120335833APending Publication Date: 2025-07-18TP-LINK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397878.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art cannot automate the execution of change tasks of stateful services, and cannot record and display the status of the change process, resulting in frequent abnormalities and manual intervention.

Method used

Through the custom resource definition (CRD) expansion mechanism, KeepAliveSet resources and custom Webhook services are designed, the differences between expected fields and actual fields are automatically processed, and the operation and maintenance management of long-connected services are realized, and state conversion and abnormal self-healing is adopted to control and tuning.

Benefits of technology

It realizes automated change management of stateful services, avoids abnormal occurrence and reduces manual intervention, and improves the usability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335833A_ABST
    Figure CN120335833A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of application environments, and provides a stateful service management method, which comprises the following steps: acquiring an expected field of a user and an actual field of stateful service; under the condition that a resource change event occurs, processing the expected field and the actual field to obtain processing process information; modifying the actual field according to the processing process information to obtain a current field; and comparing the expected field with the current field, and completing stateful service conversion under the condition that the expected field is the same as the current field. The invention further provides computer equipment and a computer readable storage medium. According to the invention, automatic change in the stateful service updating process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of application environments, and particularly relates to a management method, device, and storage medium for stateful services. Background Art

[0002] The change management of stateful services is different from common stateless services. Changes require special processing of specific states. When designing service automation change management, two common problems usually need to be solved. The first is how to automatically execute change tasks to achieve the transformation from the current state to the desired target state. The second is how to record and display the state of the change process, and how to change the initial state and target state of the change object.

[0003] In the prior art, the change management of stateful services is mainly achieved by the operation and maintenance system calling imperative scripts. Among them, various script units are used to execute a series of operation tasks to complete service change operations and state transformation. The operation and maintenance system will orchestrate script operations, and the script returns the processing result. The operation and maintenance system then records the version set, service state, and change result in the system database. The prior art cannot automatically execute change tasks, nor record and display the state of the change process. It cannot meet the management requirements of stateful services. Summary of the Invention

[0004] The embodiments of this application provide a management method, device, and storage medium for stateful services, which can solve the problems of inability to automatically execute change tasks and record and display the state of the change process in the management of stateful services.

[0005] In a first aspect, the embodiments of this application provide a management method for stateful services. The method includes:

[0006] Obtain the expected fields of the user and the actual fields of the stateful service;

[0007] In the case of a resource change event, process the expected fields and the actual fields to obtain process information;

[0008] Modify the actual fields according to the process information to obtain the current fields;

[0009] Compare the expected fields and the current fields. If the expected fields and the current fields are the same, complete the transformation of the stateful service.

[0010] For the management method of claim 1, before obtaining the expected fields of the user and the actual fields of the stateful service, it further includes:

[0011] Store the expected state input by the user as the expected fields;

[0012] Record the current state of the stateful service and store it as the actual fields;

[0013] Submit the expected fields and the actual fields.

[0014] Optionally, store the expected status input by the user as the expected field, including:

[0015] Obtain the expectation input by the user;

[0016] Verify and modify the expectation input by the user to obtain the expected field;

[0017] Store the expected field.

[0018] Optionally, the management method further includes:

[0019] Regularly compare the expected field and the current field. When the expected field and the actual field are different, process the expected field and the actual field to obtain the process information.

[0020] Optionally, modify the actual field according to the process information to obtain the current field, including:

[0021] Modify the actual field to obtain the current field by creating a resource;

[0022] Modify the actual field to obtain the current field through an external resource.

[0023] Optionally, the management method further includes:

[0024] Annotate the interface for obtaining the expected field and the actual field in a preset manner to obtain the interface annotation;

[0025] Add the interface annotation to the Pod template.

[0026] Optionally, the management method further includes:

[0027] Set the declared field in the expected field.

[0028] In a second aspect, an embodiment of the present application provides a management system for a stateful service, the system includes:

[0029] An acquisition module: used to acquire the expected field of the user and the actual field of the stateful service;

[0030] A processing module: used to process the expected field and the actual field to obtain the process information when a resource change event occurs;

[0031] A modification module: used to modify the actual field according to the process information to obtain the current field;

[0032] A comparison module: used to compare the expected field and the current field, and complete the conversion management of the stateful service when the expected field and the current field are the same.

[0033] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above is implemented.

[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in any one of the above.

[0035] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: Avoiding uncontrollable exceptions and reducing manual intervention during the update process. It can effectively avoid exceptions introduced by network or service failures in a distributed system. By continuously controlling the loop to check the actual state and the expected state, anomaly self-healing can be achieved, improving system availability. Description of the Drawings

[0036] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 It is a flowchart of a method for managing a stateful service provided by an embodiment of the present application;

[0038] Figure 2 It is a flowchart of a method for managing a stateful service provided by another embodiment of the present application;

[0039] Figure 3 It is a flowchart of step S201 provided by an embodiment of the present application;

[0040] Figure 4 It is a flowchart of step S103 provided by an embodiment of the present application;

[0041] Figure 5 It is a flowchart of a method for managing a stateful service provided by another embodiment of the present application;

[0042] Figure 6 It is a schematic internal structure diagram of a computer device provided by an embodiment of the present application. Detailed Embodiments

[0043] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0044] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0045] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0046] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0047] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0048] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0049] This application is used to solve the state change of stateful services, among which the persistent connection service is a stateful service. The client and the service maintain a TCP persistent connection for a long time. For a service with a single instance carrying tens of thousands of TCP persistent connections, it is impossible to directly perform rolling updates like ordinary stateless services. The update process requires connection migration, adjustment of load balancing weights, etc. In this embodiment, the client is an IoT device.

[0050] For the core applications connected to IoT devices, a conservative strategy is usually adopted during upgrades and changes, and no failures are allowed. Therefore, strategies such as grayscale updates and batch changes are also adopted. After the service has been running in grayscale for a period of time, the grayscale range will be gradually expanded.

[0051] Since stateful services usually have unique operation and maintenance management methods, the built-in StatefulSet of K8s cannot meet the operation and maintenance management of most stateful services. Therefore, K8s provides a custom resource definition (CRD) extension mechanism, allowing users to expand the functions of K8s and perform secondary development on K8s built-in resources and controllers. This application proposes a management method for stateful services based on the CRD extension mechanism. Kubernetes, abbreviated as K8s, is an abbreviation that replaces the 8 characters "ubernete" in the middle of the name with 8. It is an open source containerized application for managing multiple hosts in a cloud platform. The goal of Kubernetes is to make the deployment of containerized applications simple and powerful. Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance.

[0052] This management method can declare the expected state of the service in the resource and view the current state of the service. In conjunction with the self-developed WebHook and Controller services, it provides a management system for long-connection stateful services. In this embodiment, custom resources can be created, updated, and deleted using clients such as Kubectl, just like using K8s native resources.

[0053] Please refer to Figure 1 , which is a flow chart of a method for managing a stateful service provided in an embodiment of the present application, the method specifically includes:

[0054] Step S101, obtaining the user's expected fields and the actual fields of the stateful service.

[0055] In this embodiment, the user first submits the KeepAliveSet resource for managing the long connection service to the K8s Apiserver. Then, the resource processing enters the custom Webhook service. After the processing is completed, the resource is written to the K8s Etcd service. Among them, the KeepAliveSet resource is divided into an expected field and an actual field. Among them, the expected field is also called Spec; the actual field is also called Status. Spec records the state expected by the user, and Status records the current actual state of the service.

[0056] The custom resource KeepAliveSet designed in this embodiment can describe the expected state Spec and the actual state Status of the long connection service operation and maintenance management.

[0057] Among them, the content included in KeepAliveSet.Spec is as follows:

[0058] (a) Basic version information: such as the current running main version and the number of instances, the gray version and the number of gray instances;

[0059] (b) Interface information for connection transfer: such as the interface information for starting connection transfer and querying the number of connections, the threshold of the number of connections for service offline, the connection transfer speed and switch declaration. These information can be used to perform migration operations on long connections during the update process;

[0060] (c) Declaration of various change methods: such as gray update, expansion, batch update, batch update batch, minimum online quantity, shrinkage and shrinkage object, eviction and eviction object.

[0061] Among them, the content included in KeepAliveSet.Status is as follows:

[0062] (a) The total number of current service instances and the number of ready instances;

[0063] (b) Whether the current state is consistent with the state expected by Spec;

[0064] (c) The specific steps of the current update and the status of the specific steps. For example, the status of each batch during batch update;

[0065] (d) The name of the Pod currently undergoing connection transfer and the connection transfer status. For example, the connection transfer speed and the result of connection transfer interface call, etc.

[0066] Step S102, in the case of a resource change event, process the expected field and the actual field to obtain the processing process information.

[0067] In this embodiment, the custom controller uses the control tuning method to register the Event change events of resources such as KeepAliveSet, Deployment, and Pod. The Event change event is also called the resource change event. After receiving the resource change event, the controller compares the Spec field of the service status expected by the KeepAliveSet resource user with the Status field of the actual service status to obtain the processing process information.

[0068] Step S103, modify the actual field according to the processing process information to obtain the current field.

[0069] In this embodiment, record the processing process information into the Status field of KeepAliveSet according to the processing process information to obtain the current field.

[0070] Step S104, compare the expected field and the current field. When the expected field and the current field are the same, the stateful service conversion is completed.

[0071] In this embodiment, compare the expected field and the current field. When the actual status Status field is consistent with the expected status Spec field, the service status conversion process is completed.

[0072] The above embodiment can automate the common operation and maintenance operations of long connections through the management method of stateful services. For example, for batch updates, the number of batches can be set to reduce the resources required for creating new version Pods. It supports common version gray-scale updates and can continue to create gray-scale versions when there are existing gray-scale versions. It supports the Pod eviction operation that must be used for Node replacement in K8s.

[0073] Please refer to Figure 2 , which is the flowchart of the management method of stateful services provided by another embodiment of this application. Before obtaining the expected field of the user and the actual field of the stateful service, the method further includes:

[0074] Step S201, store the expected status input by the user as the expected field.

[0075] In this embodiment, the expected status input by the user is a custom resource. When adding, deleting, querying, and modifying custom resources, the system configures through MutatingWebhookConfiguration that such resources need to be forwarded to the self-developed Webhook service for processing. After the resource is submitted to the K8s apiserver service through clients such as Kubectl, the apiserver will forward the request to the Webhook.

[0076] Step S202, record the current status of the stateful service and store it as the actual field.

[0077] In this embodiment, the Webhook checks whether there is an update in progress and rejects the changes submitted by the user when there is an ongoing update, unless the change type is manual synchronization. The Webhook also validates the update plan. For example, when evicting or scaling down, it is necessary to check that the declared Pods actually exist, and the number of Pods in each version after eviction must match the number of target Pods to be evicted.

[0078] Step S203: Submit the expected fields and the actual fields.

[0079] In this embodiment, the Webhook also generates an update plan for the controller based on the expected state declared by the user and the state of the current target service. The content includes:

[0080] (a) The number of Pods to be newly created and the number of Pods to be deleted for each processing.

[0081] (b) For batch updates, there are also the number of connection migrators and connection in - migrators in each batch.

[0082] (c) For scaling down and eviction, there is also a list of target Pods to be processed for scaling down.

[0083] Obtain the expected fields and the actual fields according to the above content, and submit the expected fields and the actual fields.

[0084] In the above - mentioned embodiment, through the custom resources input by the user and processed by the Webhook service, it is possible to avoid uncontrollable exceptions and reduce manual intervention during the update process.

[0085] Please refer to Figure 3 , which is the flowchart of step S201 provided by an embodiment of the present application. Step S201: Store the expected state input by the user as the expected field, specifically including:

[0086] Step S2011: Obtain the expectation input by the user.

[0087] In this embodiment, the expected data of the user for the current system is obtained through a hardware input device.

[0088] Step S2012: Validate and modify the expectation input by the user to obtain the expected field.

[0089] In this embodiment, Webhook first performs basic type checks on each value of the resource. For example, the version number defined in KeepAliveSet must conform to the version number regular match, and the update type must be an enumerated type specified by the system. After passing the check, WebHook will also reset and modify some values. For example, the total number of service instances is determined by the number of main version instances declared and the number of instances in each gray version. Therefore, the total number of instances will be automatically calculated and overwritten by Webhook.

[0090] Step S2013, store the expected fields.

[0091] In this embodiment, store the expected fields automatically calculated by Webhook, and completely encapsulate the update process of the stateful service into the declaration file, so as to realize the conversion of user expectations into controller language.

[0092] In the above embodiment, the management method of the stateful service is used for the operation and maintenance management of the long connection service. The update process of the stateful service can be completely encapsulated into the declaration file, and the changes can be audited by the DIFF change of the declaration file content before the update.

[0093] The management method of the stateful service provided by another embodiment of this application further includes: regularly comparing the expected fields and the current fields, and when the expected fields and the actual fields are different, processing the expected fields and the actual fields to obtain process information.

[0094] In this embodiment, the controller also sets a regularly triggered method to avoid missing or losing Event change events caused by the restart of the controller or the interruption of the network segment between the controller and the K8s Apiserver. The regularly triggered time is set to 10 seconds. The numbers mentioned in this embodiment are only examples and not limitations, and specific data shall be subject to actual applications.

[0095] The above embodiment can realize the self-healing of the long-state system exception. For example, it can be automatically restored in case of an exception, and can be retried continuously by means of control tuning, and push the current state to the target expected state. Such as Pod cannot be created, resource shortage, network exception, and LB exception, etc. The above exceptions can be self-healed and restored by means of the declarative regular comparison of K8s itself.

[0096] Please refer to Figure 4 , which is the flowchart of step S103 provided by an embodiment of this application. Step S103, modify the actual fields according to the process information to obtain the current fields, specifically including:

[0097] Step S1031, create a resource to modify the actual fields to obtain the current fields.

[0098] A Pod is the smallest scheduling unit in Kubernetes. A Pod encapsulates one or more containers that share resources such as storage and network. The containers in a Pod run on the same host, use the same network namespace, IP address, and ports, and can communicate via localhost. A container is a group of processes with restricted and isolated resources, instantiated from a container image that contains the program to be executed and all its dependencies, such as code, runtime, system libraries, etc.

[0099] In this embodiment, the controller is responsible for implementing the transformation of the service state. When resources are created, changed, or deleted, the controller will monitor the resource changes and target object changes, and perform corresponding processing by comparing the current declared state and the actual service state, such as adding Pods, deleting Pods, and evicting Pods. When a version upgrade requires migrating long connections from the old version to the new version, the target Pods to be processed in each batch are automatically calculated based on the minimum online number, batch times, etc., and the connection migration is sequentially executed on the Pods according to the connection migration speed declared by the user. When the number of connections of the Pod is lower than the offline threshold, the Pod is closed to complete the replacement of the old and new versions.

[0100] Among them, the loop from the outer layer to the inner layer is mainly divided into global tuning, update type tuning, CLB status modification tuning, and connection migration tuning.

[0101] By creating resources and modifying the actual fields, the current fields are obtained, including global tuning and update type tuning. Global Reconcile: Obtain the current state, compare it with the expected state, add an update lock to the Status.planLocked field before executing the update plan, and release the lock after the update plan is completed. Global tuning enters the lower-level tuning according to the update plan type.

[0102] UpdateType Reconcile: According to types such as batch update, gray scale, scale out / scale in, kick connections only, and eviction, handle the update, such as creating Deployment resources and scaling Deployment resources. When it is determined that connection transfer is required, enter the lower-level loop.

[0103] Step S1032, modify the actual fields through external resources to obtain the current fields.

[0104] In this embodiment, through external resources, the actual fields are modified to obtain the current fields, including CLB status modification tuning and connection migration tuning. CLB Status Modification Tuning (CLB Reconcile): Before performing connection transfer, by modifying the custom ReadinessGate of the Pod to Unready, after waiting for the LB controller to detect that the Pod is in the Unready state, synchronize the weight of the CLB to the target Pod to 0. Specifically, the weight here is the weight of the load balancer. The previous weight was the default value of 10, and the weights of each pod were the same.

[0105] Connection Migration Tuning (Kick Reconcile): Mark the connection migration roles of the Pods, including the connection releaser (kickee), the connection receiver (reciever), and the connection keeper (keeper). Verify the readiness status of each Pod before connection transfer to ensure that the connection is removed at the connection transfer speed declared by the user when kickee <= reciever.

[0106] The above embodiment can implement the viewing of the update process status and the connection transfer status. Among them, the viewing of the update process status can view the current steps of batch updates through Status, and the specific status of each batch of updates, such as preparing resources, marking Pods, connection transfer, and Pod going offline. The viewing of the connection transfer status can view the connection transfer process through Status in real time, such as the Pods performing connection transfer, the connection transfer speed, and the remaining connection transfer time.

[0107] Please refer to Figure 5 , which is the flowchart of the management method for stateful services provided by another embodiment of this application. Before obtaining the expected fields and the actual fields of the stateful service, this method further includes:

[0108] Step S501, annotate the interface for obtaining the expected fields and the actual fields in a preset manner to obtain interface annotations.

[0109] In this embodiment, before use, it is necessary to define the Pod template using the defined KeepAliveSet. The system requires the long connection service to provide an interface for processing connections and annotate it into the Pod template through annotation.

[0110] Step S502, add the interface annotations to the Pod template.

[0111] In this embodiment, the interfaces include: an interface for starting to remove connections at a specified speed, a pause connection removal interface, and an interface for obtaining the current remaining connection count through monitoring. There are no restrictions on other contents of the Pod template.

[0112] Flowchart of the management method for stateful services provided by another embodiment of the present application. Before obtaining the expected fields of the user and the actual fields of the stateful service, the method further includes: setting the declared fields in the expected fields.

[0113] In this embodiment, in addition to the Pod template, it is also necessary to set the declared fields in the following Spec of the KeepAliveSet. The declared fields include:

[0114] 1. type: Change type;

[0115] 2. versionSet.main.version: Main version number, which is the version with the largest number of replicas;

[0116] 3. versionSet.main.replicas: Number of replicas of the main version;

[0117] 4. versionSet.grayList: The field is a List, with a maximum length of 2. The content is the version number and the number of replicas of the gray scale;

[0118] 5. batchUpdateSteps: Number of batches during batch update. For example, when upgrading 50 old versions to 50 new versions, the number of batches is 5, and 10 Pods are updated each time;

[0119] 6. kickSwitch: true represents that the kick connection switch is turned on, and false represents that the kick connection switch is turned off. If the kick connection switch is turned off, the Pod will remain in the pending state before the kick connection;

[0120] 7. kickTps: Kick connection tps. When all change types are designed to kick the connection, it is performed on a per-Pod basis;

[0121] 8. targetPodList: Specify the Pods to be processed when type is scaleDown, drain, or kick. When the type is kick, it is also necessary to specify the duration of the kick connection, kickDurationSecondsUsedOnTypeKick;

[0122] 9. readyReplicasMin: Minimum number of available replicas of the service. To prevent failures during the change process from causing new long connections to overwhelm the backend service, the number of online services needs to be greater than this value.

[0123] 10. receiverReadyTimeOutSeconds: Timeout for waiting for the newly created Pod to enter the Ready state. If this time is exceeded, the creation of the Pod is abnormal. Usually, it can be set to 100s;

[0124] 11. offlineTolerationConnectionsMax: After performing the connection kicking task on a Pod, if the number of connections is less than this value, it can be considered that the connection kicking is completed. This is used in production environments where some connections cannot be kicked.

[0125] 12. kickDurationSecondsUsedOnTypeKick: When the type is kick, set the duration for kicking connections, which is used to kick redundant long connections from a certain Pod.

[0126] In some embodiments, the usage scenarios for managing long - connection services using KeepAliveSet are usually when the service is first launched and the ScaleUp method of KeepAliveSet is used to create the service. Or when the service is already managed by Deployment, KeepAliveSet's ManualSync is used to take over existing Pods.

[0127] In some embodiments, using system declarative management for long - connection services, the following operation and maintenance operations are performed:

[0128] 1. Batch update (resource change type = batchUpdate): Used to batch - update the current version to a new version or a version that has completed gray - scale verification.

[0129] 2. Gray - scale update (type = gray): Used to create a gray - scale version of the service. In the case of an existing gray - scale version, a new gray - scale version can be created again. The system limits a maximum of two gray - scale versions in addition to the main version.

[0130] 3. Scale - up (type = scaleUp): Used to scale up the main version or gray - scale version.

[0131] 4. Scale - down (type = scaleDown): Specify the Pod name for shutdown operation to reduce the number of services.

[0132] 5. Drain (type = drain): Since the underlying Node nodes managed by K8s will be maintained and replaced, in this case, this method can be used to drain the long - connection service to other K8s nodes.

[0133] 6. Only kick connections (type = kick): Due to the special nature of long - connection services, there may be a situation where the number of connections between service instances is unbalanced. For such scenarios, this method can be used to kick the connections of Pod nodes with excessive connection numbers to other Pods.

[0134] The control tuning technology adopted by this system can effectively avoid anomalies introduced by distributed system network or service failures. By continuously checking the actual state and the expected state through a control loop, anomaly self-healing can be achieved, improving system availability.

[0135] In the above embodiments, as the core access service, long connections have a large number of instances and a long update process span (in this embodiment, in units of days). The following operation methods are reserved in the system design:

[0136] (a) The KeepAliveSet resource can be gracefully launched. After the service has been deployed using Deployment, the method reserved in the system with the type of ManualSync can be used to take over the existing Pods without rebuilding the Pods.

[0137] (b) For exceptions occurring during the core service update process, the update action needs to be immediately paused. The reserved ManualSync method in the system can also be used to abandon the change, stop all connection kicking actions, and restore the load balancing weights.

[0138] (c) During the connection transfer process, the kickSwitch switch can be used to control in real time whether to pause or start kicking connections, which is used to pause connection kicking when there are important alarms during the core service update process.

[0139] For the core access layer service that manages tens of millions of connections in this system, any uncontrollable anomalies need to be avoided and manual intervention during the update process needs to be reduced. The control tuning technology adopted by this system can effectively avoid anomalies introduced by distributed system network or service failures. By continuously checking the actual state and the expected state through a control loop, anomaly self-healing can be achieved, improving system availability.

[0140] The embodiment of this application provides a management system for stateful services. The system includes: an acquisition module, a processing module, a modification module, and a comparison module. Among them:

[0141] The acquisition module: is used to acquire the expected fields of the user and the actual fields of the stateful service.

[0142] In this embodiment, the user first submits the KeepAliveSet resource for managing the long connection service to the K8s Apiserver, and then the resource processing enters the custom Webhook service. After the processing is completed, the resource is written to the K8s Etcd service. Among them, the KeepAliveSet resource is divided into expected fields and actual fields. Among them, the expected fields are also called Spec; the actual fields are also called Status. Spec records the state expected by the user, and Status records the current actual state of the service.

[0143] The custom resource KeepAliveSet designed in this embodiment can describe the expected state Spec and the actual state Status of the long - connection service operation and maintenance management.

[0144] Among them, the content included in KeepAliveSet.Spec is as follows:

[0145] (a) Basic version information: such as the main version currently running and the number of instances, the gray - scale version and the number of gray - scale instances;

[0146] (b) Interface information for connection transfer: such as the interface information for starting connection transfer and querying the number of connections, the threshold of the number of connections for service offline, the connection transfer speed and switch declaration, which can be used to migrate long - connections during the update process;

[0147] (c) Declaration of various change methods: such as gray - scale update, expansion, batch update, batch update batch, minimum online quantity, shrinkage and shrinkage object, eviction and eviction object.

[0148] Among them, the content included in KeepAliveSet.Status is as follows:

[0149] (a) The total number of instances of the current service and the number of ready instances;

[0150] (b) Whether the current state is consistent with the state expected by Spec;

[0151] (c) The specific steps of the current update and the status of the specific steps. For example, the status of each batch during batch update;

[0152] (d) The name of the Pod currently undergoing connection transfer and the connection transfer status. For example, the connection transfer speed, the result of connection transfer interface call, etc.

[0153] Processing module: used to process the expected fields and actual fields to obtain process information in the case of resource change events.

[0154] In this embodiment, the custom controller uses the control tuning method to register Event change events of resources such as KeepAliveSet, Deployment, and Pod. Event change events are also called resource change events. After receiving a resource change event, the controller will compare the service state Spec field expected by the KeepAliveSet resource user and the service actual state Status field to obtain process information.

[0155] Modification module: used to modify the actual fields according to the process information to obtain the current fields.

[0156] In this embodiment, the processing process information is recorded into the Status field of the KeepAliveSet according to the processing process information to obtain the current field.

[0157] Comparison module: used to compare the expected field and the current field. When the expected field and the current field are the same, the stateful service conversion is completed.

[0158] Through the management method of the stateful service in the above embodiment, the common operation and maintenance operations of the long connection can be automated. For example, for batch update, the number of batches can be set to reduce the additional resources required for creating a new version of the Pod. It supports the common version gray update and can continue to create a gray version when there is already a gray version. It supports the Pod eviction operation that must be used for Node replacement in K8s.

[0159] Please refer to Figure 6 , which is a schematic internal structure diagram of a computer device in an embodiment. The computer device 900 includes a memory 910 and a processor 920. The memory 910 stores a computer program. When the computer program is executed by the processor, the processor 920 executes the steps of any one of the above methods.

[0160] The computer device 900 further includes a processor 920, a memory 910 and a network interface 940 connected through a system bus 930. Among them, the memory 910 includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device 900 stores an operating system and can also store a computer program. When the computer program is executed by the processor 920, the processor 920 can implement the management method of the stateful service. The internal memory 910 can also store a computer program. When the computer program is executed by the processor, the processor can execute the management method of the stateful service.

[0161] Among them, the memory 910 includes at least one type of computer-readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 910 can be an internal storage unit of the computer device 900, such as the hard disk of the computer device 900. In other embodiments, the memory 910 can also be an external storage device of the computer device 900, such as a plug-in hard disk equipped on the computer device 900, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 910 can also include both an internal storage unit and an external storage device of the computer device 900. The memory 910 can be used not only to store application software installed on the computer device 900 and various types of data, such as computer programs for the management method of the status service, but also to temporarily store data that has been output or will be output, such as data generated by the execution of the management method of the status service. In some feasible embodiments, the processor 920 can be a Central Processing Unit (CPU), and can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0162] Specifically, the processor 920 executes the computer program of the management method of the status service to control the computer device 900 to implement the management method of the status service.

[0163] Further, the computer device 900 can also include a system bus 930, which can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0164] Specifically, the computer device 900 may further include a network interface 940. The network interface may optionally include a wired network interface and / or a wireless network interface (such as a WI-FI network interface, a Bluetooth network interface, etc.), and is generally used to establish a communication connection between the computer device 900 and other devices. For example, a communication connection between the computer device 900 and a waveform display device.

[0165] In some other feasible embodiments, the computer device 900 may further include a display component (not shown in the figure). The display component may be an LED (Light Emitting Diode) display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display component may also be appropriately referred to as a display device or a display unit, and is used to display the information processed in the computer device 900 and to display a visual user interface.

[0166] Figure 6 Only the computer device 900 with components 910-940 and the management method for implementing the status service is shown. Those skilled in the art can understand that Figure 6 The shown structure does not constitute a limitation on the computer device 900, and may include fewer or more components than shown, or combine certain components, or have different component arrangements. Since the computer device 900 adopts all the technical solutions of the above-mentioned all embodiments, it at least has all the beneficial effects brought by the technical solutions of the above-mentioned embodiments, which will not be elaborated here.

[0167] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of any one of the above methods. Specifically, the program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0168] It should be noted that for the content such as information interaction and execution process between the above devices / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0169] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details are not described herein again.

[0170] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0171] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0172] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included within the protection scope of this application.

Claims

1. A management method for stateful services, characterized in that The method includes: Obtaining the expected fields of the user and the actual fields of the stateful service; In the case of a resource change event, processing the expected fields and the actual fields to obtain processing process information; Modifying the actual fields according to the processing process information to obtain the current fields; Comparing the expected fields and the current fields, and completing the transformation management of the stateful service when the expected fields and the current fields are the same.

2. The management method according to claim 1, wherein Before obtaining the expected fields of the user and the actual fields of the stateful service, it further includes: Storing the expected state input by the user as the expected fields; Recording the current state of the stateful service and storing it as the actual fields; Submitting the expected fields and the actual fields.

3. The management method according to claim 2, characterized in that, The storing the expected state input by the user as the expected fields includes: Obtaining the expectation input by the user; Verifying and modifying the expectation input by the user to obtain the expected fields; Storing the expected fields.

4. The management method according to claim 1, characterized in that, The management method further includes: Regularly comparing the expected fields and the current fields, and in the case where the expected fields and the actual fields are different, processing the expected fields and the actual fields to obtain processing process information.

5. The management method according to claim 1, characterized in that, The modifying the actual fields according to the processing process information to obtain the current fields includes: Modifying the actual fields to obtain the current fields by creating resources; Modifying the actual fields to obtain the current fields through external resources.

6. The management method according to claim 1, wherein The management method further includes: Annotating the interface for obtaining the expected fields and the actual fields in a preset manner to obtain interface annotations; Adding the interface annotations to the Pod template.

7. The management method according to claim 1, characterized in that, The management method further includes: Setting the declared fields in the expected fields.

8. A management system for stateful services, characterized in that, The system includes: An obtaining module: used for obtaining the expected fields of the user and the actual fields of the stateful service; A processing module: used for processing the expected fields and the actual fields to obtain processing process information in the case of a resource change event; A modifying module: used for modifying the actual fields according to the processing process information to obtain the current fields; A comparing module: used for comparing the expected fields and the current fields, and completing the transformation management of the stateful service when the expected fields and the current fields are the same.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program implements the method according to any one of claims 1 to 7 when executed by the processor.