Multi-cloud service rollback control method and device, computer equipment and storage medium

By identifying the failed nodes in a multi-cloud environment and selecting appropriate rollback strategies, the system overload caused by traditional rollback technology during peak periods is solved, and efficient and flexible multi-cloud rollback control is achieved to ensure service stability and resource conservation.

CN120560705APending Publication Date: 2025-08-29BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510560964.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-29

Smart Images

  • Figure CN120560705A_ABST
    Figure CN120560705A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-cloud service rollback control method and device, computer equipment and a storage medium. The method comprises the following steps: detecting traffic and load of a plurality of cloud clusters providing services, and identifying whether a service fault occurs on each cloud cluster; in response to a service fault on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a fault node; detecting a real-time flow index of the fault node, and selecting a rollback strategy according to the real-time flow index; executing a rollback operation on the fault node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein the plurality of cloud clusters can synchronously execute respective corresponding rollback operations; and performing service verification and stability check on the target cloud cluster after rollback. According to the method and the device, the corresponding rollback operations are synchronously executed on different cloud clusters, so that the applicability of the rollback mode is improved, overload crash of the system cannot be caused, and the rollback time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a multi-cloud service rollback control method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of cloud computing technology, more and more enterprises are adopting a distributed architecture to deploy their applications across multiple cloud platforms. This facilitates elastic scalability and high availability. However, when applications experience failures or issues, rollbacks have become a common method for restoring services. Traditional rollback techniques typically roll back all nodes in the entire system, failing to fully consider the differences between cloud platforms and fluctuations in business traffic. During peak traffic periods, a unified full rollback across multiple cloud platforms can lead to system overload, causing service crashes or unavailability. Rollbacks often require manual selection of the target version, making the process cumbersome and time-consuming. Furthermore, rollbacks are typically performed in full, lacking the flexibility to roll back failed nodes, further increasing time costs. Traditional rollback methods ignore the load issues during peak traffic periods and directly roll back a large number of nodes, potentially causing the system to be unable to handle the peak traffic and resulting in service avalanches. Traditional rollback methods lack flexibility and cannot provide customized rollback strategies for different application types and cloud platform configurations. Summary of the Invention

[0003] Based on this, a multi-cloud service rollback control method, apparatus, computer equipment and storage medium are provided to solve the technical problems that during traffic peaks, using a unified rollback method for multiple cloud platforms may cause system overload, thereby causing service crashes or unavailability, and the rollback time is long and lacks flexibility.

[0004] In one aspect, a multi-cloud service rollback control method is provided, the method comprising:

[0005] Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster;

[0006] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster;

[0007] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index;

[0008] Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously execute their respective corresponding rollback operations;

[0009] Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

[0010] In one embodiment, in response to a service failure occurring on a target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and health check results of each node in the target cloud cluster includes:

[0011] In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster;

[0012] Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

[0013] In one embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0014] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0015] Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators;

[0016] If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

[0017] In one embodiment, obtaining a real-time traffic indicator of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic indicator includes:

[0018] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0019] The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

[0020] In one embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0021] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0022] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0023] When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy;

[0024] When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected;

[0025] When the service thread ratio is greater than a third ratio threshold, the blue-green rollback strategy is selected.

[0026] In one embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0027] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0028] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0029] Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications;

[0030] Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster;

[0031] Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters;

[0032] When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

[0033] In one embodiment, performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes:

[0034] When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down.

[0035] When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0036] When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0037] In another aspect, a multi-cloud service rollback control device is provided, the device comprising:

[0038] A fault monitoring module is used to detect the traffic and load of multiple cloud clusters providing services and identify whether a service failure occurs in each cloud cluster;

[0039] A rollback version and faulty node selection module is configured to determine a rollback version of the target cloud cluster in response to a service failure on the target cloud cluster, and identify the faulty node based on the version information and health check results of each node in the target cloud cluster;

[0040] A rollback strategy selection module is used to detect the real-time traffic index of the faulty node and select a rollback strategy according to the real-time traffic index;

[0041] A rollback operation execution module, configured to execute a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can execute their respective corresponding rollback operations simultaneously;

[0042] The rollback verification module is used to perform service verification and stability check on the target cloud cluster after the rollback to ensure that the target cloud cluster runs stably after the rollback.

[0043] In another aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:

[0044] Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster;

[0045] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster;

[0046] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index;

[0047] Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously execute their respective corresponding rollback operations;

[0048] Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

[0049] In another aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0050] Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster;

[0051] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster;

[0052] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index;

[0053] Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously execute their respective corresponding rollback operations;

[0054] Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

[0055] The above-mentioned multi-cloud service rollback control method, device, computer equipment and storage medium, by identifying faulty nodes for different cloud clusters and selecting rollback versions and rollback strategies, can simultaneously execute corresponding rollback operations on different cloud clusters, thereby improving the applicability of the rollback method, preventing system overload and crash, and shortening the rollback time. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1 This is an application environment diagram of a multi-cloud service rollback control method in one embodiment of the present application;

[0058] Figure 2 This is a flowchart of a multi-cloud service rollback control method in one embodiment of the present application;

[0059] Figure 3 A logical diagram of the steps for determining the rollback version of the target cloud cluster in one embodiment of the present application;

[0060] Figure 4 A logical diagram of the steps for identifying a faulty node based on the version information and health check results of each node in the target cloud cluster in one embodiment of the present application;

[0061] Figure 5 A logic diagram of the steps of performing a rollback operation on a failed node of the target cloud cluster according to a selected rollback version and rollback strategy in one embodiment of the present application;

[0062] Figure 6 This is a structural block diagram of a multi-cloud service rollback control device in one embodiment of the present application;

[0063] Figure 7 This is a diagram of the internal structure of a computer device in one embodiment of the present application. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0065] As described in the background, existing rollback technologies are widely used to recover application services in distributed and cloud environments. When an application fails, a rollback operation typically restores the service to a previously stable version. Existing rollback technologies have the following key features:

[0066] Full rollback: When a failure occurs, the rollback operation is performed on all nodes, restoring to the previous version. It is suitable for situations where there is a problem with the entire service.

[0067] Rollback process: Existing technologies require manual selection of the target version during rollback, increasing operational complexity. Most systems only support restoring to a specific version and do not provide a one-click rollback function to the previous stable version.

[0068] To solve the above problems, the present invention creatively proposes a multi-cloud service rollback control method to realize a unified rollback control engine mode in a multi-cloud heterogeneous environment, which is different from the existing technology that only supports the rollback mechanism of a single containerized cluster (such as K8S). Figure 1 As shown, the present invention pioneers a unified scheduling framework for multi-cloud heterogeneous resources, which can simultaneously support the mixed rollback of ECS virtual machines and containerized clusters such as Alibaba Cloud and Tencent Cloud, and realize dynamic adaptation of cross-cloud batch strategies. This can be achieved by setting up a multi-cloud adaptation layer: by encapsulating standardized API interfaces, shielding the differences in underlying cloud platforms, and providing a unified rollback control plane. Automatically match the optimal rollback rhythm based on the traffic and load differences of different cloud clusters. For high-load clusters, a blue-green rollback strategy is used to prioritize stability. For low-load clusters, a fast rollback strategy is enabled to quickly reclaim resources. The unified rollback control engine for multi-cloud heterogeneous environments adopts a hot-swappable architecture and has the ability to quickly access other cloud platforms.

[0069] In one embodiment, Figure 2 As shown, a multi-cloud service rollback control method is provided, including the following steps:

[0070] Step S1: Detect the traffic and load of multiple cloud clusters providing services and identify whether a service failure occurs in each cloud cluster;

[0071] Step S2: In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on the version information and health check results of each node in the target cloud cluster;

[0072] Step S3, detecting the real-time traffic index of the faulty node, and selecting a rollback strategy according to the real-time traffic index;

[0073] Step S4: performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously perform their respective corresponding rollback operations;

[0074] Step S5: performing service verification and stability check on the target cloud cluster after the rollback to ensure that the target cloud cluster after the rollback runs stably.

[0075] Rollback refers to the process of restoring an application to its previous normal version when an application fails or has problems.

[0076] This embodiment realizes the synchronous execution of corresponding rollback operations on different cloud clusters by identifying faulty nodes for different cloud clusters and selecting rollback versions and rollback strategies, thereby improving the applicability of the rollback method, preventing system overload and crash, and shortening the rollback time.

[0077] The specific effects of this embodiment are as follows:

[0078] (1) Improve the rollback efficiency in multi-cloud environments: By selecting an appropriate rollback strategy, unified and efficient rollback operations can be effectively implemented across multiple cloud platforms, reducing manual intervention and operational complexity.

[0079] (2) Reduce the risk of service interruption: The rollback strategy based on traffic fluctuations and node status can avoid system overload caused by large-scale rollback during business peaks and ensure high service availability.

[0080] (3) Saving resources and costs: By rolling back only the failed nodes instead of the entire system, the system can save a lot of computing resources and time, and reduce the costs and risks of rollbacks.

[0081] See also Figure 3 、 Figure 4 In this embodiment, in response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and health check results of each node in the target cloud cluster includes:

[0082] In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster;

[0083] Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

[0084] Among them, when you choose to roll back the version, the system will automatically roll back to the previous version of the application, reducing operation and communication costs and saving rollback time.

[0085] When selecting a rollback node, the system will identify the faulty nodes based on the version information and health check results of each node, and only roll back these problematic nodes, thereby speeding up system recovery and achieving rapid stop-loss.

[0086] See also Figure 5 In this embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index includes:

[0087] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0088] Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators;

[0089] If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

[0090] When selecting a rollback strategy, the system uses real-time traffic monitoring to determine whether traffic is currently peaking. If so, a blue-green rollback strategy or a standard rollback strategy is used to ensure that the rollback does not cause a service avalanche and affect service availability. During periods of low traffic, the system can select a rapid rollback strategy.

[0091] In this embodiment, obtaining the real-time traffic index of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic index includes:

[0092] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0093] The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

[0094] In this embodiment, a variety of rollback strategies are provided to cope with different business scenarios and requirements.

[0095] (1) Shielding cloud platform differences: Through unified interface and policy design, the architectural differences between different cloud platforms are shielded, so that rollback operations can be uniformly executed in a multi-cloud environment, improving rollback efficiency and consistency.

[0096] (2) Node status-based rollback: By recording the version information of each node, we can identify and roll back only those nodes with problems, rather than rolling back all nodes. This refined rollback method greatly improves the speed of system recovery.

[0097] (3) Flexible rollback strategy selection: Based on the current business traffic and the actual status of the application, the system provides three rollback strategies: blue-green rollback, fast rollback, and normal rollback. The system selects the most appropriate rollback strategy based on traffic conditions and application characteristics to avoid traffic overload during peak periods.

[0098] In other embodiments, detecting a real-time traffic indicator of the faulty node and selecting a rollback strategy based on the real-time traffic indicator includes:

[0099] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0100] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0101] When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy;

[0102] When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected;

[0103] When the service thread ratio is greater than a third ratio threshold, the blue-green rollback strategy is selected.

[0104] For example, the real-time traffic indicators of the cloud cluster where the faulty node is located are obtained, and based on the real-time resource water levels (CPU / DB connection pool / thread pool) of each cloud cluster, differentiated rollback strategies are dynamically recommended to achieve multi-cloud parallel rollback.

[0105] (1) CPU water level (i.e., processor occupancy): When it is higher than 60% (the first ratio threshold), the rapid rollback strategy is disabled and the blue-green rollback strategy is forced to be used;

[0106] (2) DB connection rate (i.e., database connection ratio): If it is higher than 65% (the second ratio threshold), the blue-green rollback strategy is disabled, the normal rollback strategy is used instead, and the connection pool expansion is triggered;

[0107] (3) Service thread pool (i.e., service thread ratio): When the number of active threads is greater than 60% (the third ratio threshold), the blue-green rollback strategy is forced to be used.

[0108] In other embodiments, detecting a real-time traffic indicator of the faulty node and selecting a rollback strategy based on the real-time traffic indicator includes:

[0109] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0110] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0111] Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications;

[0112] Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster;

[0113] Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters;

[0114] When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

[0115] Among them, processor utilization: CPU occupancy rate of application process (%);

[0116] Input / output wait time ratio: the ratio of the CPU time spent waiting for I / O operations (obtained through iostat or / proc / stat);

[0117] Database connection pool activity rate: number of active connections / maximum number of connections (%);

[0118] Thread pool blocking rate: the percentage of time a thread spends waiting for locks or IO.

[0119] Determining the application type of the cloud cluster according to the application feature data includes:

[0120] When the processor utilization rate is greater than 65% (average of a 5-minute sliding window), the input and output waiting time ratio is less than 20% (indicating that CPU computing is the main focus and IO waiting is minimal), and the thread pool blocking rate is less than 30% (threads are running most of the time), the cloud cluster application type is determined to be processor-intensive.

[0121] When the processor utilization is less than 40% (low computing demand), the input and output waiting time ratio is greater than 50% (a large amount of time waiting for disk / network IO), the database connection pool activity rate is greater than 60%, or the thread pool blocking rate is greater than 50%, the cloud cluster application type is determined to be an input and output intensive application;

[0122] Otherwise, it is determined that the application type of the cloud cluster is a hybrid application.

[0123] Among them, when the indicators of hybrid applications are between the above two categories (for example, CPU = 50%, IO wait = 35%), the type tendency is calculated according to the dynamic weight formula: type score = 0.6 × CPU utilization + 0.3 × IO wait time ratio + 0.1 × thread pool blocking rate; if the score is > 60 → CPU-intensive tendency; if the score is < 40 → IO-intensive tendency.

[0124] The values ​​of the dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio are set based on the application type of the cloud cluster, including:

[0125] When the application type of the cloud cluster is a processor-intensive application, setting the values ​​of the dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio to a first weight parameter set;

[0126] When the application type of the cloud cluster is an input / output intensive application, setting the dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio to a second weight parameter set;

[0127] When the application type of the cloud cluster is a hybrid application, the values ​​of the dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio are set to a third weight parameter set.

[0128] It is understandable that the method of setting the values ​​of the dynamic weight parameters of the processor occupancy rate, the database connection ratio and the service thread ratio based on the application type of the cloud cluster can also be replaced by: setting the values ​​of the dynamic weight parameters of the processor occupancy rate, the database connection ratio and the service thread ratio through real-time adjustment through a reinforcement learning algorithm based on Q-Learning.

[0129] The step of setting the values ​​of the dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio to be adjusted in real time by using a reinforcement learning algorithm based on Q-Learning includes:

[0130] The state space is defined as a triplet, which includes processor occupancy, database connection ratio, and service thread ratio;

[0131] Set the action space as the incremental adjustment of the values ​​of the dynamic weight parameters corresponding to the processor occupancy rate, the database connection ratio and the service thread ratio, and limit the step size of each time step;

[0132] The Q-value table is updated based on historical switching data at fixed intervals. The ε-greedy strategy is used for random exploration with probability ε, and the action with the largest current Q-value is selected with probability 1-ε to determine the values ​​of the dynamic weight parameters corresponding to the processor occupancy rate, database connection ratio, and service thread ratio.

[0133] Calculating the rollback policy weight by weighted summation based on the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters includes:

[0134] The formula for calculating the rollback policy weight is W = α·A+β·B+γ·C, where A is the processor occupancy rate, B is the database connection ratio, C is the service thread ratio, α is the value of the dynamic weight parameter of the processor occupancy rate, β is the value of the dynamic weight parameter of the database connection ratio, and γ is the value of the dynamic weight parameter of the service thread ratio.

[0135] See also Figure 5 In this embodiment, performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes:

[0136] When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down.

[0137] When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0138] When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0139] In this embodiment, during rollback execution and verification, the system rolls back nodes batch by batch according to the selected rollback strategy, provides a unified rollback executor, and shields the differences between different cloud platforms; after the rollback is completed, a health check is performed to ensure that the service is restored and stable after the rollback.

[0140] In the above-mentioned multi-cloud service rollback control method, by identifying the faulty nodes for different cloud clusters and selecting the rollback version and rollback strategy, the corresponding rollback operations are performed on different cloud clusters simultaneously, which improves the applicability of the rollback method, does not cause system overload and crash, and shortens the rollback time.

[0141] The above multi-cloud service rollback control method has the following specific effects:

[0142] (1) Improve the rollback efficiency in multi-cloud environments: By selecting an appropriate rollback strategy, unified and efficient rollback operations can be effectively implemented across multiple cloud platforms, reducing manual intervention and operational complexity.

[0143] (2) Reduce the risk of service interruption: The rollback strategy based on traffic fluctuations and node status can avoid system overload caused by large-scale rollback during business peaks and ensure high service availability.

[0144] (3) Saving resources and costs: By rolling back only the failed nodes instead of the entire system, the system can save a lot of computing resources and time, and reduce the costs and risks of rollbacks.

[0145] In one embodiment, Figure 6 As shown, a multi-cloud service rollback control device 10 is provided, including: a fault monitoring module 1, a rollback version and fault node selection module 2, a rollback strategy selection module 3, a rollback operation execution module 4, and a rollback verification module 5.

[0146] The fault monitoring module 1 is used to detect the traffic and load of multiple cloud clusters providing services, and identify whether a service fault occurs in each cloud cluster.

[0147] The rollback version and faulty node selection module 2 is used to determine the rollback version of the target cloud cluster in response to a service failure on the target cloud cluster, and identify the faulty node based on the version information and health check results of each node in the target cloud cluster.

[0148] The rollback strategy selection module 3 is used to detect the real-time traffic index of the faulty node and select a rollback strategy according to the real-time traffic index.

[0149] The rollback operation execution module 4 is used to execute a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can execute their corresponding rollback operations synchronously.

[0150] The rollback verification module 5 is used to perform service verification and stability check on the target cloud cluster after the rollback, to ensure that the target cloud cluster after the rollback runs stably.

[0151] In this embodiment, in response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and health check results of each node in the target cloud cluster includes:

[0152] In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster;

[0153] Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

[0154] In this embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0155] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0156] Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators;

[0157] If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

[0158] In this embodiment, obtaining the real-time traffic index of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic index includes:

[0159] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0160] The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

[0161] In one embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0162] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0163] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0164] When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy;

[0165] When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected;

[0166] When the service thread ratio is greater than a third ratio threshold, the blue-green rollback strategy is selected.

[0167] In one embodiment, detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0168] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0169] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0170] Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications;

[0171] Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster;

[0172] Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters;

[0173] When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

[0174] In this embodiment, performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes:

[0175] When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down.

[0176] When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0177] When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0178] In the above-mentioned multi-cloud service rollback control device, the faulty nodes are identified for different cloud clusters respectively and the rollback version and rollback strategy are selected to realize the synchronous execution of the corresponding rollback operations on different cloud clusters, thereby improving the applicability of the rollback method, preventing the system from overloading and crashing, and shortening the rollback time.

[0179] For the specific definition of the multi-cloud service rollback control device, please refer to the definition of the multi-cloud service rollback control method above, which will not be repeated here. Each module in the above-mentioned multi-cloud service rollback control device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0180] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0181] Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster;

[0182] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster;

[0183] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index;

[0184] Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously execute their respective corresponding rollback operations;

[0185] Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

[0186] In one embodiment, the computer program further performs the following steps when executed by a processor:

[0187] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and a health check result of each node in the target cloud cluster includes:

[0188] In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster;

[0189] Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

[0190] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0191] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0192] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0193] Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators;

[0194] If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

[0195] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0196] The obtaining of the real-time traffic index of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic index includes:

[0197] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0198] The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

[0199] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0200] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0201] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0202] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0203] When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy;

[0204] When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected;

[0205] When the service thread ratio is greater than a third ratio threshold, the blue-green rollback strategy is selected.

[0206] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0207] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0208] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0209] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0210] Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications;

[0211] Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster;

[0212] Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters;

[0213] When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

[0214] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0215] The performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes:

[0216] When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down.

[0217] When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0218] When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0219] For the specific limitations on the steps implemented when the computer program is executed by the processor, please refer to the limitations on the method for multi-cloud service rollback control above, which will not be repeated here.

[0220] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multi-cloud service rollback control data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a multi-cloud service rollback control method is implemented.

[0221] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0222] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0223] Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster;

[0224] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster;

[0225] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index;

[0226] Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously execute their respective corresponding rollback operations;

[0227] Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

[0228] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0229] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and a health check result of each node in the target cloud cluster includes:

[0230] In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster;

[0231] Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

[0232] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0233] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0234] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0235] Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators;

[0236] If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

[0237] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0238] The obtaining of the real-time traffic index of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic index includes:

[0239] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0240] The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

[0241] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0242] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0243] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0244] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0245] When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy;

[0246] When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected;

[0247] When the service thread ratio is greater than a third ratio threshold, the blue-green rollback strategy is selected.

[0248] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0249] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0250] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0251] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0252] Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications;

[0253] Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster;

[0254] Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters;

[0255] When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

[0256] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0257] The performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes:

[0258] When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down.

[0259] When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0260] When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0261] For the specific limitations on the steps implemented when the processor executes the computer program, please refer to the limitations on the method for multi-cloud service rollback control above, which will not be repeated here.

[0262] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0263] Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster;

[0264] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster;

[0265] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index;

[0266] Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can simultaneously execute their respective corresponding rollback operations;

[0267] Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

[0268] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0269] In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and a health check result of each node in the target cloud cluster includes:

[0270] In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster;

[0271] Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

[0272] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0273] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0274] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0275] Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators;

[0276] If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

[0277] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0278] The obtaining of the real-time traffic index of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic index includes:

[0279] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0280] The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

[0281] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0282] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0283] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0284] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0285] When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy;

[0286] When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected;

[0287] When the service thread ratio is greater than a third ratio threshold, the blue-green rollback strategy is selected.

[0288] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0289] Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes:

[0290] Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy;

[0291] Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio;

[0292] Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications;

[0293] Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster;

[0294] Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters;

[0295] When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

[0296] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0297] The performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes:

[0298] When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down.

[0299] When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0300] When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

[0301] For the specific limitations on the steps implemented when the computer program is executed by the processor, please refer to the limitations on the method for multi-cloud service rollback control above, which will not be repeated here.

[0302] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0303] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0304] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A multi-cloud service rollback control method, characterized in that: include: Detect the traffic and load of multiple cloud clusters providing services and identify whether service failures occur in each cloud cluster; In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying the failed node based on version information and health check results of each node in the target cloud cluster; Detecting the real-time traffic index of the faulty node and selecting a rollback strategy based on the real-time traffic index; Performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; Multiple cloud clusters can execute their corresponding rollback operations synchronously; Perform service verification and stability check on the target cloud cluster after rollback to ensure stable operation of the target cloud cluster after rollback.

2. The multi-cloud service rollback control method according to claim 1, characterized in that: In response to a service failure occurring on the target cloud cluster, determining a rollback version of the target cloud cluster, and identifying a failed node based on version information and a health check result of each node in the target cloud cluster includes: In response to a service failure occurring on a target cloud cluster, obtaining a system version running on each node of the target cloud cluster, determining a previous system version of the current system version, and using the previous system version as a rollback version of the target cloud cluster; Obtain the cloud computer room information of the node distribution of the target cloud cluster, perform functional operation health checks on the nodes in each cloud computer room based on the node distribution computer room information, obtain the version information and health check results of each node in each cloud computer room, and treat the nodes that cannot function as faulty nodes.

3. The multi-cloud service rollback control method according to claim 1, characterized in that: Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes: Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy; Obtain the real-time traffic indicators of the cloud cluster where the faulty node is located, and determine whether it is in a business peak period based on the real-time traffic indicators; If it is a business peak period, the blue-green rollback strategy or the common rollback strategy is selected; if it is not a business peak period, the fast rollback strategy is selected.

4. The multi-cloud service rollback control method according to claim 3, characterized in that: The obtaining of the real-time traffic index of the cloud cluster where the faulty node is located and determining whether the cloud cluster is in a peak period based on the real-time traffic index includes: Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio; The processor occupancy rate, the database connection ratio and the service thread ratio are weighted and summed to obtain a traffic peak evaluation value. When the traffic peak evaluation value is greater than the traffic peak threshold, it is determined to be in the business peak period; otherwise, it is determined to be in the business off-peak period.

5. The multi-cloud service rollback control method according to claim 1, characterized in that: Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes: Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy; Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio; When the processor occupancy rate is greater than a first ratio threshold, disabling the fast rollback strategy and selecting the blue-green rollback strategy; When the database connection ratio is greater than a second ratio threshold, the blue-green rollback strategy is disabled and the common rollback strategy is selected; When the service thread ratio is greater than a third ratio threshold, the fast rollback strategy is selected.

6. The multi-cloud service rollback control method according to claim 1, characterized in that: Detecting the real-time traffic index of the faulty node and selecting a rollback strategy according to the real-time traffic index includes: Setting the rollback strategy to include one of a blue-green rollback strategy, a normal rollback strategy, and a fast rollback strategy; Obtain real-time traffic indicators of the cloud cluster where the faulty node is located, wherein the real-time traffic indicators include processor occupancy rate, database connection ratio, and service thread ratio; Obtaining application characteristic data of the cloud cluster, the application characteristic data including processor utilization, input / output wait time ratio, database connection pool activity rate, and thread pool blocking rate; and determining an application type of the cloud cluster based on the application characteristic data, the application type of the cloud cluster including processor-intensive applications, input / output-intensive applications, and hybrid applications; Setting values ​​of dynamic weight parameters of the processor occupancy rate, the database connection ratio, and the service thread ratio based on the application type of the cloud cluster; Calculating a rollback policy weight by weighted summation according to the processor occupancy rate, the database connection ratio, the service thread ratio, and the values ​​of their corresponding dynamic weight parameters; When the rollback policy weight is greater than the first threshold, the blue-green rollback policy is selected; when the rollback policy weight is less than or equal to the first threshold and greater than the second threshold, the normal rollback policy is selected; when the rollback policy weight is less than or equal to the second threshold, the fast rollback policy is selected.

7. The multi-cloud service rollback control method according to any one of claims 3, 5 or 6, characterized in that: The performing a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy includes: When the selected rollback strategy is the blue-green rollback strategy, the non-faulty nodes of the target cloud cluster are controlled to continue execution, and the number of application nodes corresponding to the faulty nodes of the target cloud cluster is expanded. The selected rollback version is installed on the application nodes and then started. The services of the faulty nodes are transferred to the application nodes, and the faulty nodes are controlled to shut down. When the selected rollback strategy is a normal rollback strategy, a rollback batch is set according to the traffic of the faulty node of the target cloud cluster. The faulty nodes are controlled to stop and install the selected rollback version in batches by rolling back the batches one by one, and the faulty nodes are started to resume service after the selected rollback version is installed. When the selected rollback strategy is the fast rollback strategy, the faulty nodes of the target cloud cluster are set to roll back in two batches, with half of the faulty nodes rolled back in each batch. The faulty nodes are controlled to be shut down in batches to install the selected rollback version, and the faulty nodes are started to resume service after the selected rollback version is installed.

8. A multi-cloud service rollback control device, characterized in that: The device comprises: A fault monitoring module is used to detect the traffic and load of multiple cloud clusters providing services and identify whether a service failure occurs in each cloud cluster; A rollback version and faulty node selection module is configured to determine a rollback version of the target cloud cluster in response to a service failure on the target cloud cluster, and identify the faulty node based on the version information and health check results of each node in the target cloud cluster; A rollback strategy selection module is used to detect the real-time traffic index of the faulty node and select a rollback strategy according to the real-time traffic index; A rollback operation execution module, configured to execute a rollback operation on the failed node of the target cloud cluster according to the selected rollback version and rollback strategy; wherein multiple cloud clusters can execute their respective corresponding rollback operations simultaneously; The rollback verification module is used to perform service verification and stability check on the target cloud cluster after the rollback to ensure that the target cloud cluster runs stably after the rollback.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the multi-cloud service rollback control method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-cloud service rollback control method according to any one of claims 1 to 7 are implemented.