Dynamic disaster recovery method, system and device for cluster service shunting

By acquiring server node load information and calculating load weight values, alarm information and business diversion schemes are generated, solving the disaster recovery problem of existing intelligent diversion algorithms under high cluster load, realizing intelligent diversion and adaptive disaster recovery, and improving the system's stability and fault self-healing capability.

CN120980081APending Publication Date: 2025-11-18CHINA TELECOM INTELLIGENT NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511251732.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing intelligent traffic splitting algorithms lack disaster recovery methods for high overall cluster load in load balancing scenarios. They cannot perceive business scenarios, resulting in the need for manual intervention under overload conditions, and cannot achieve adaptive recovery.

Method used

By acquiring server node load information, calculating load weight values, generating alarm information and service diversion schemes, and dynamically adjusting traffic allocation, intelligent traffic diversion and adaptive disaster recovery are achieved, including a comprehensive assessment of CPU utilization, memory utilization, disk utilization, and network interface traffic.

Benefits of technology

It enables rapid response to changes in node load, avoids business interruptions caused by node overload, improves system stability and reliability, enhances fault self-healing capabilities, and reduces the frequency of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980081A_ABST
    Figure CN120980081A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic disaster recovery method, system and device for cluster service shunting. The method comprises: obtaining load information of each server node in a service cluster, the load information comprising at least one of a CPU occupancy rate, a memory utilization rate, a disk occupancy rate and network port traffic; determining a load weight value of each server node according to the load information; service shunting information of the service cluster is acquired in real time, wherein the service shunting information comprises the number of server nodes for processing each service and a server node list; and updating the load weight value of each server node according to the service shunting information to obtain an updated load weight value of each server node, and generating alarm information and a service shunting scheme based on the updated load weight value of each server node. The problem that an existing intelligent flow distribution algorithm can dynamically adjust the flow distribution condition according to the load condition of cluster nodes and lacks disaster recovery processing means when the overall load of a cluster is high is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of dynamic disaster recovery technology for cluster service offloading, and more specifically, to a dynamic disaster recovery method, a dynamic disaster recovery system, and a dynamic disaster recovery device for cluster service offloading. Background Technology

[0002] The rapid development of 5G converged private network services has made the need for interconnection between private networks and large networks increasingly urgent, but this has also brought about a series of problems such as network security, operator management and supervision.

[0003] To meet the needs of single or mixed business scenarios in SBI, N4 and AAAP modes, the C-IWF cluster traffic is complex.

[0004] In load balancing scenarios, traditional intelligent traffic distribution algorithms dynamically adjust traffic allocation based on the load of cluster nodes. They lack disaster recovery methods for high overall cluster load and cannot perceive business scenarios. They can only judge the load based on traffic volume, hardware load, etc. At the same time, they lack dynamic management methods. When overload occurs, manual intervention is required, and adaptive recovery cannot be achieved. Summary of the Invention

[0005] The main purpose of this application is to provide a dynamic disaster recovery method, dynamic disaster recovery system and dynamic disaster recovery device for cluster service traffic splitting, so as to at least solve the problem that existing intelligent traffic splitting algorithms dynamically adjust traffic allocation according to the load of cluster nodes and lack disaster recovery means when the cluster as a whole is under high load.

[0006] To achieve the above objectives, according to one aspect of this application, a dynamic disaster recovery method for cluster service traffic splitting is provided, comprising: acquiring load information of each server node in a service cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic; determining the load weight value of each server node based on the load information; acquiring service traffic splitting information of the service cluster in real time, wherein the service traffic splitting information includes the number of server nodes processing each service and a list of server nodes; updating the load weight value of each server node based on the service traffic splitting information to obtain an updated load weight value of each server node, and generating alarm information and a service traffic splitting scheme based on the updated load weight value of each server node.

[0007] Optionally, generating alarm information and a service diversion scheme based on the updated load weight values ​​of each of the server nodes includes: generating a first alarm information when all server nodes processing the target service are detected to be overloaded server nodes, and generating the service diversion scheme based on the first alarm information, wherein the first alarm information represents an overload alarm for the processing service resources of all server nodes processing the target service, and the overloaded server node represents the server node whose updated load weight value is less than a preset load weight value.

[0008] Optionally, the service diversion scheme is generated based on the first alarm information. The method further includes: obtaining the service message data of the target service, and distributing the service message data evenly to each of the overload server nodes processing the target service based on the first alarm information.

[0009] Optionally, after generating the service diversion scheme based on the first alarm information, the method further includes: detecting the ratio of the number of overloaded server nodes processing the target service to the total number of all server nodes processing the target service; and eliminating the first alarm information if the ratio is less than a preset ratio.

[0010] Optionally, determining the load weight value of each server node based on the load information includes: determining the load weight value of each server node using a threshold comparison, weighted average, or weight scaling algorithm based on the load information.

[0011] Optionally, generating alarm information and a service diversion scheme based on the updated load weight values ​​of each of the server nodes includes: generating a second alarm information when a preset number of server nodes processing the target service are detected as overloaded server nodes, and allocating the service packet data of the target service to the target server nodes according to the updated load weight values, wherein the preset number is less than the total number of server nodes processing the target service, the overloaded server node indicates that the updated load weight value of the server node is less than the preset load weight value, the second alarm information indicates that the processing service resources of the preset number of server nodes processing the target service are overloaded, and the target server node indicates that the updated load weight value of the server node is greater than or equal to the preset load weight value.

[0012] According to another aspect of this application, a dynamic disaster recovery system for cluster service traffic splitting is provided, comprising: multiple server nodes, wherein the server nodes are IP virtual servers deployed on a service cluster by real servers; a load balancing system, wherein the load balancing system includes an intelligent traffic splitting control module and a load balancing module, wherein the intelligent traffic splitting control module is used to execute any of the dynamic disaster recovery methods for cluster service traffic splitting described above; and the load balancing module is used to receive the load weight value output by the intelligent traffic splitting control module, determine service traffic splitting information according to the load weight value, and distribute service packets to the corresponding server nodes according to the service traffic splitting information.

[0013] Optionally, the system further includes: a network management platform for configuring and monitoring the dynamic disaster recovery system; and an operation status monitoring module for receiving alarm information output by the intelligent traffic splitting control module and the service splitting information from the load balancing module, and uploading the alarm information and the service splitting information to the network management platform.

[0014] Optionally, the server node includes a load information collection client and multiple business subsystems. The load information collection client is used to collect the load information of the server node and, based on ZeroMQ's in-band telemetry technology, multiplex the business links to transmit the load information to the intelligent traffic splitting control module.

[0015] According to another aspect of this application, a dynamic disaster recovery device for cluster service traffic splitting is provided, comprising: a first acquisition unit, configured to acquire load information of each server node in a service cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network port traffic; a determination unit, configured to determine the load weight value of each server node based on the load information; a second acquisition unit, configured to acquire service traffic splitting information of the service cluster in real time, wherein the service traffic splitting information includes the number of server nodes processing each service and a list of server nodes; and an update unit, configured to update the load weight value of each server node based on the service traffic splitting information to obtain an updated load weight value of each server node, and generate alarm information and a service traffic splitting scheme based on the updated load weight value of each server node.

[0016] This application's technical solution obtains the load information of each server node in a business cluster, including at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic. Based on the load information, the load weight value of each server node is determined. Real-time business traffic distribution information of the business cluster is acquired, including the number of server nodes processing each business and a list of server nodes. The load weight value of each server node is updated based on the business traffic distribution information, resulting in an updated load weight value for each server node. Alarm information and a business traffic distribution plan are generated based on the updated load weight values ​​of each server node. By dynamically detecting the load information of each server node and determining its load weight value, and based on the load weight value and business traffic distribution information of the business cluster, the solution ensures load balancing and dynamic adjustment of disaster recovery strategies for each server. This solution can quickly respond to changes in node load, achieving intelligent traffic distribution, dynamic weight adjustment, and adaptive disaster recovery, effectively avoiding business interruptions caused by node overload and improving system stability and reliability. It solves the problem that existing intelligent traffic distribution algorithms dynamically adjust traffic allocation based on the load of cluster nodes but lack disaster recovery measures for high overall cluster load. Attached Figure Description

[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 A hardware structure block diagram of a mobile terminal for performing a dynamic disaster recovery method for cluster service offloading according to an embodiment of this application is shown.

[0019] Figure 2 A flowchart illustrating a dynamic disaster recovery method for cluster service offloading provided according to an embodiment of this application is shown.

[0020] Figure 3 A structural diagram of a dynamic disaster recovery system for cluster service offloading provided according to an embodiment of this application is shown;

[0021] Figure 4 A flowchart illustrating a dynamic disaster recovery method for specific cluster service offloading provided according to an embodiment of this application is shown.

[0022] Figure 5 A structural block diagram of a dynamic disaster recovery device for cluster service offloading provided according to an embodiment of this application is shown. Detailed Implementation

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:

[0027] C-IWF (Customized-InterWorking Function): Signaling Interconnection Gateway;

[0028] IPVS (IP Virtual Server): An IP virtual server is a technology that provides load balancing functionality under LVS; it consists of virtual servers and physical servers.

[0029] WLC (Weighted Least-Connection): This is a weighted Least-Connection algorithm that distributes new traffic to real servers with fewer active connections based on their weights.

[0030] LB (Load Balance): Load balancing;

[0031] AMF (Access and Mobility Management Function): A core 5G network element responsible for registration, connectivity, reachability, mobility, security, access management, and service authorization;

[0032] UDM (User Data Management): A 5G network element responsible for managing and storing user-related data, providing authentication, authorization, and user configuration functions for the 5G network;

[0033] SMF (Session Management function): This function is responsible for tunnel maintenance, IP address allocation and management, UPF selection, policy enforcement and QoS control, billing data collection, roaming, etc.

[0034] UPF (User plane function): User plane functions include packet routing and forwarding, policy enforcement, traffic reporting, and QoS processing.

[0035] As described in the background section, existing intelligent traffic splitting algorithms dynamically adjust traffic allocation based on the load of cluster nodes, but lack disaster recovery measures when the cluster as a whole is under high load. To address the problem that existing intelligent traffic splitting algorithms dynamically adjust traffic allocation based on the load of cluster nodes but lack disaster recovery measures when the cluster as a whole is under high load, embodiments of this application provide a dynamic disaster recovery method, dynamic disaster recovery system, and dynamic disaster recovery device for cluster service splitting.

[0036] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0037] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for a dynamic disaster recovery method of cluster service offloading according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0038] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the dynamic disaster recovery method for cluster service offloading in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-described networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-described networks may include wireless networks provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0039] This embodiment provides a dynamic disaster recovery method for cluster service offloading running on mobile terminals, computer terminals or similar computing devices. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0040] Figure 2 This is a flowchart of a dynamic disaster recovery method for cluster service offloading according to an embodiment of this application. For example... Figure 2 As shown, the method includes the following steps:

[0041] Step S201: Obtain the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic.

[0042] Step S202: Determine the load weight value of each of the above server nodes based on the load information.

[0043] Step S203: Obtain the service routing information of the above-mentioned service cluster in real time. The service routing information includes the number of server nodes and the list of server nodes that process each service.

[0044] Step S204: Update the load weight value of each of the above-mentioned server nodes according to the above-mentioned service diversion information to obtain the updated load weight value of each of the above-mentioned server nodes, and generate alarm information and service diversion scheme based on the updated load weight value of each of the above-mentioned server nodes.

[0045] This embodiment, by applying steps S201, S202, S203, and S204, dynamically detects the load information of each server node to determine the load weight value of each server node. Based on the load weight value and the business traffic distribution information of the business cluster, it ensures the dynamic adjustment of load balancing and disaster recovery strategies for each server. This solution can quickly respond to changes in node load, realize intelligent traffic distribution, dynamic weight adjustment, and adaptive disaster recovery, effectively avoiding business interruptions caused by node overload, and improving system stability and reliability. It solves the problem that existing intelligent traffic distribution algorithms dynamically adjust traffic allocation based on the load of cluster nodes, but lack disaster recovery methods when the entire cluster is under high load.

[0046] In the specific implementation process, alarm information and service diversion scheme are generated based on the updated load weight values ​​of each of the aforementioned server nodes, including: when all the aforementioned server nodes processing the target service are detected to be overloaded server nodes, a first alarm information is generated, and the aforementioned service diversion scheme is generated based on the first alarm information. The first alarm information represents an alarm indicating that the processing service resources of all server nodes processing the target service are overloaded, and the overloaded server node represents that the updated load weight value of the aforementioned server node is less than a preset load weight value.

[0047] The preset load weight value can be set to a value close to 0, such as 0.1. The higher the load weight value of a server node, the lower the server load.

[0048] This method, through the collaborative operation of the runtime monitoring module and the intelligent traffic distribution control module, achieves accurate identification and alarm for overload scenarios across all nodes of a specific service. When the load weight values ​​of all nodes processing a specific service are detected to be below a preset threshold, it indicates that these nodes cannot effectively handle new service traffic, at which point the runtime monitoring module will generate the first alarm message. The technology in this embodiment can promptly detect and warn of overload situations across all service nodes, avoiding service processing failures due to resource overruns and enhancing the system's fault self-healing capabilities.

[0049] Specifically, the above-mentioned service diversion scheme is generated based on the first alarm information, including: obtaining the service message data of the target service, and distributing the service message data evenly to each of the above-mentioned over-limit server nodes that process the target service based on the first alarm information.

[0050] If the weight of all nodes in the target service is 0, report an alarm that all nodes of the service have exceeded their resource limits. Set the load weight of all nodes of the service to 1 (the server node that processes the service has a weight of 1, and the load is distributed evenly to ensure that the service is not interrupted as much as possible).

[0051] This method protects existing service traffic even when all nodes are overloaded, through the decision-making of the intelligent traffic distribution control module. In principle, when the system detects that all nodes processing a specific service are overloaded, the intelligent traffic distribution control module automatically adjusts the weights of these nodes to ensure uninterrupted connections, while simultaneously distributing service packet data evenly across all nodes to maintain service continuity as much as possible. In terms of effectiveness, the weight reset protection measure prevents service traffic interruption and enhances service assurance.

[0052] More specifically, after generating the service diversion scheme based on the first alarm information, the method further includes: detecting the ratio of the number of the overloaded server nodes processing the target service to the total number of all the server nodes processing the target service; and eliminating the first alarm information if the ratio is less than a preset ratio.

[0053] The preset ratio can be set to 50%.

[0054] This method achieves automatic monitoring and handling of the overload recovery status of all service nodes through dynamic weight adjustment and alarm cancellation mechanisms. In principle, when the number of nodes exceeding the limit for processing a specific service is detected to be lower than a preset proportion, it indicates that the node load for that service has been alleviated. At this point, the first alarm message is automatically cleared, and the weights are recalculated based on the current load situation to restore normal traffic allocation. In terms of effectiveness, the technology in this embodiment can automatically monitor the node load recovery status, promptly eliminate alarms, avoid manual intervention, and improve the intelligent operation and maintenance level of the system.

[0055] Further, determining the load weight value of each of the aforementioned server nodes based on the aforementioned load information includes: determining the aforementioned load weight value of each of the aforementioned server nodes using threshold comparison, weighted average, and weight scaling algorithms based on the aforementioned load information.

[0056] This method achieves dynamic adjustment of load weights based on multi-dimensional health assessment through a comprehensive weight calculation mechanism. In principle, it calculates the load weight value of a node based on hardware indicators such as CPU utilization, memory utilization, disk utilization, and network interface traffic, combined with network performance indicators and service processing indicators, using algorithms such as threshold comparison, weighted averaging, and weight scaling. In terms of effectiveness, the technology in this embodiment can more accurately assess the health status of each server node, dynamically adjust weights, optimize resource allocation, and improve the load balancing and resource utilization of the cluster.

[0057] Furthermore, generating alarm information and a service diversion scheme based on the updated load weight values ​​of each of the aforementioned server nodes includes: generating a second alarm information when a preset number of server nodes processing the target service are detected as overloaded server nodes, and allocating the service packet data of the target service to the target server nodes according to the updated load weight values. The preset number is less than the total number of server nodes processing the target service; the overloaded server nodes indicate that the updated load weight value of the server node is less than the preset load weight value; the second alarm information indicates that the processing service resources of the preset number of server nodes processing the target service are overloaded; and the target server nodes are those whose updated load weight value is greater than or equal to the preset load weight value.

[0058] This method employs a tiered circuit breaker mechanism to achieve rapid response and alerts for overloaded nodes. In principle, when the load weight values ​​of a preset number of nodes processing a specific service are detected to be lower than a preset threshold, the intelligent traffic control module generates a second alarm to indicate that some nodes have exceeded their resource limits and require disaster recovery measures. In terms of effectiveness, the technology in this embodiment can promptly detect and warn of overloaded nodes, triggering the circuit breaker mechanism and preventing service processing failures due to resource overload, thus enhancing the system's disaster recovery capabilities.

[0059] This embodiment also relates to a dynamic disaster recovery system for cluster service offloading, such as... Figure 3 As shown, it includes: multiple server nodes, which are real servers deployed on the business cluster as IP virtual servers; a load balancing system, which includes an intelligent traffic splitting control module and a load balancing module. The intelligent traffic splitting control module is used to execute any of the above-mentioned dynamic disaster recovery methods for cluster business traffic splitting; the load balancing module is used to receive the load weight value output by the intelligent traffic splitting control module, determine the business traffic splitting information according to the load weight value, and distribute the business packets to the corresponding server nodes according to the business traffic splitting information.

[0060] This embodiment achieves intelligent traffic distribution and dynamic disaster recovery based on business awareness by constructing a dynamic disaster recovery system that includes multiple server nodes and a load balancing system. The intelligent traffic distribution control module calculates weights based on the load information of the nodes, and the load balancing module dynamically adjusts the traffic allocation strategy based on these weight values ​​to ensure reasonable distribution of business traffic and avoid node overload. In terms of effectiveness, the technology in this embodiment can realize intelligent traffic distribution, dynamic weight adjustment, and adaptive disaster recovery, effectively improving system stability and resource utilization. It solves the problem that existing intelligent traffic distribution algorithms dynamically adjust traffic allocation based on the load of cluster nodes, but lack disaster recovery methods when the entire cluster is under high load.

[0061] Furthermore, the above system also includes: a network management platform for configuring and monitoring the above dynamic disaster recovery system; and an operation status monitoring module for receiving alarm information output by the above intelligent traffic control module and the above service traffic information from the above load balancing module, and uploading the alarm information and the above service traffic information to the above network management platform.

[0062] Among them, such as Figure 3 As shown, the load balancing configuration information and intelligent traffic splitting enable configuration are distributed through the network management platform, and the configuration is stored in the data storage module; the network function control module processes the data and distributes the IPVS configuration to the load balancing module; at the same time, it distributes the intelligent traffic splitting configuration.

[0063] The load balancing module is responsible for controlling traffic distribution and monitoring the health status of each node. If any abnormality is found, it will be reported to the operation status monitoring module.

[0064] The load information collection client is based on ZeroMQ and reuses traffic links to report load information to the intelligent traffic control module. It can determine whether the load information is updated in a timely manner by linking with health detection and business IPVS RS.

[0065] The intelligent traffic splitting control module calculates the IPVS RS weight based on the load information and obtains the traffic splitting information of the running services from the load balancing module. Based on the above two types of information, it makes the final IPVS RS weight setting decision.

[0066] Furthermore, the aforementioned server node includes a load information collection client and multiple business subsystems. The load information collection client is used to collect the load information of the aforementioned server node and, based on ZeroMQ's in-band telemetry technology, multiplex the business links to transmit the load information to the aforementioned intelligent traffic splitting control module.

[0067] This system utilizes ZeroMQ's in-band telemetry technology, reusing business links to transmit load data and ensuring 99.999% reachability of monitoring data. By combining the load information acquisition client with ZeroMQ technology, highly reliable load information transmission is achieved. The load information acquisition client can collect node load information in real time, while ZeroMQ provides an efficient and reliable communication channel, ensuring timely updates and transmission of load information.

[0068] To enable those skilled in the art to better understand the technical solution of this application, the implementation process of the dynamic disaster recovery method for cluster service offloading of this application will be described in detail below with reference to specific embodiments.

[0069] This embodiment relates to a specific dynamic disaster recovery method for cluster service offloading, such as... Figure 4 As shown, it includes the following steps:

[0070] Step S1: Collect load information of each server node in the cluster, including CPU utilization, memory utilization, disk utilization, network port traffic, etc.

[0071] Step S2: Calculate the IPVS RS weight based on the load information using threshold comparison, weighted average, weight scaling algorithm, etc. The higher the load, the smaller the weight.

[0072] Step S3: Obtain service distribution information from the load balancing module, including the number of nodes processing each service and a list of node IDs;

[0073] Step S4: Iterate through the cluster node weight list calculated in step S2. If the node weight is 0 (when the IPVS RS weight is 0, it does not affect the original business traffic, but it cannot accept new traffic connections), report a single node resource over-limit alarm; if it is not 0, and there is an alarm, it will be automatically eliminated.

[0074] Step S5: Loop through the business routing information obtained in step S3:

[0075] Step S5.1: If the weight of all nodes of a certain service is 0, report an alarm that the resource limit of all nodes of the service is exceeded, set the IPVS RS weight of all nodes of the service to 1 (the IPVS RS weight of the service is 1, and the regression average distribution is used to ensure that the service is not interrupted as much as possible), and set the automatic recovery flag of the service to TRUE.

[0076] Step S5.2: If the weights of all nodes in a certain business are not all 0:

[0077] Step S5.2.1: If the automatic recovery flag for this service is TRUE, and if the number of nodes with a weight of non-zero for this service is greater than or equal to 50% of the nodes for this service, then cancel the resource over-limit alarm for all nodes of this service, set the automatic recovery flag for this service to FALSE, and update the IPVS RS weight according to the results of steps S3 and S4.

[0078] Step S5.2.2: If the service automatically recovers to Flag FLASE, then update the IPVS RS weight according to the results of steps S3 and S4.

[0079] This embodiment includes the following technical points:

[0080] 1. Pioneering business-aware load balancing mechanism: By acquiring business traffic information in real time, it can identify overload scenarios of all nodes of specific businesses (such as SBI signaling storm) and realize business-level circuit breaker protection;

[0081] 2. Construct a multi-dimensional health assessment system: integrate hardware indicators (CPU / memory), network indicators (traffic / latency), and business indicators for comprehensive weight calculation;

[0082] 3. Intelligent elastic recovery technology: It adopts a progressive weight recovery algorithm, which automatically increases the weight in a gradient when the node load decreases, so as to avoid load oscillation;

[0083] 4. Highly reliable communication channel: Based on ZeroMQ's in-band telemetry technology, the service link is reused to transmit load data, ensuring 99.999% reachability of monitoring data;

[0084] 5. Propose a "business-resource" dual-dimensional perception model to break through the limitations of traditional single-dimensional load balancing;

[0085] 6. Invent a dynamic weight compensation algorithm that automatically triggers weight zeroing protection when a node is overloaded, while maintaining existing connections without interruption;

[0086] 7. Design a tiered circuit breaker mechanism to achieve a three-tiered disaster recovery protection system: single node → full business node → cluster level;

[0087] 8. We have innovatively proposed the "soft shutdown" technology, which enables nodes to exit seamlessly by dynamically resetting their weights to zero, reducing the switching time by 80% compared to traditional HA solutions.

[0088] This embodiment has the following technical effects:

[0089] 1. Intelligent elastic scaling up and down of 5G private network core network elements (UPF / AMF);

[0090] 2. Cross-AZ traffic scheduling for cloud-native NFV infrastructure;

[0091] 3. Ensuring the business priority of industrial internet edge computing nodes;

[0092] 4. Adaptive scheduling of burst traffic in video live streaming CDN networks;

[0093] 5. Business continuity assurance for financial transaction systems.

[0094] 6. Enhanced fault self-healing capability: Reduces single-node fault recovery time and service interruption time;

[0095] 7. Resource utilization optimization: Improve cluster load balancing, save on hardware investment, and enable intelligent operation and maintenance;

[0096] 8. Reduce the frequency of manual intervention and improve alarm accuracy;

[0097] 9. Enhanced Business Assurance: Improved SLA compliance rate for core businesses.

[0098] This application also provides a dynamic disaster recovery device for cluster service offloading. It should be noted that this dynamic disaster recovery device for cluster service offloading can be used to execute the dynamic disaster recovery method for cluster service offloading provided in this application. This device is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0099] The following describes the dynamic disaster recovery device for cluster service offloading provided in the embodiments of this application.

[0100] Figure 5 This is a schematic diagram of a dynamic disaster recovery device for cluster service offloading according to an embodiment of this application. Figure 5 As shown, the device includes:

[0101] The first acquisition unit 51 is used to acquire the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic.

[0102] The determining unit 52 is used to determine the load weight value of each of the above-mentioned server nodes based on the above-mentioned load information.

[0103] The second acquisition unit 53 is used to acquire the service diversion information of the above-mentioned service cluster in real time. The above-mentioned service diversion information includes the number of server nodes and the list of server nodes that process each service.

[0104] The update unit 54 is used to update the load weight value of each of the above-mentioned server nodes according to the above-mentioned service diversion information, to obtain the updated load weight value of each of the above-mentioned server nodes, and to generate alarm information and service diversion scheme based on the updated load weight value of each of the above-mentioned server nodes.

[0105] In this embodiment, the first acquisition unit is used to acquire the load information of each server node in the service cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic; the determination unit is used to determine the load weight value of each server node based on the load information; the second acquisition unit is used to acquire the service routing information of the service cluster in real time, wherein the service routing information includes the number of server nodes processing each service and a list of server nodes; the update unit is used to update the load weight value of each server node based on the service routing information, thereby obtaining the updated load weight value of each server node, and generating alarm information and a service routing scheme based on the updated load weight value of each server node. By dynamically detecting the load information of each server node and determining the load weight value of each server node, and based on the load weight value and the service routing information of the service cluster, the dynamic adjustment of the load balancing and disaster recovery strategy of each server is ensured. It can quickly respond to changes in node load, realize intelligent routing dynamic weight adjustment and adaptive disaster recovery, effectively avoid service interruption caused by node overload, and improve the stability and reliability of the system. This solves the problem that existing intelligent traffic distribution algorithms dynamically adjust traffic allocation based on the load of cluster nodes, but lack disaster recovery measures when the cluster as a whole is under high load.

[0106] As an optional solution, the update unit includes a first detection module, which is used to generate a first alarm message when all the server nodes processing the target service are detected as overloaded server nodes, and generate the service diversion scheme based on the first alarm message. The first alarm message indicates that the processing service resources of all server nodes processing the target service are overloaded, and the overloaded server node indicates that the update load weight value of the server node is less than a preset load weight value.

[0107] In an optional embodiment, the updating unit further includes an allocation module, which, upon detecting that each of the aforementioned server nodes processing the target service is an over-limit server node, generates a first alarm message, obtains the service message data of the target service, and distributes the service message data evenly to each of the aforementioned over-limit server nodes processing the target service based on the first alarm message.

[0108] In one optional embodiment, the updating unit further includes a second detection module and an elimination module; the second detection module is used to generate a first alarm message after detecting that each of the aforementioned server nodes processing the target service is an over-limit server node, and then detect the ratio of the number of the over-limit server nodes processing the target service to the total number of all the aforementioned server nodes processing the target service; the elimination module is used to eliminate the first alarm message if the ratio is less than a preset ratio.

[0109] In one optional scheme, the determining unit includes a determining module, which is used to determine the load weight value of each of the above-mentioned server nodes based on the above-mentioned load information using a threshold comparison, weighted average, and weight scaling algorithm.

[0110] In one optional scheme, the updating unit includes a third detection module, configured to generate a second alarm message when a preset number of server nodes processing the target service are detected as overloaded server nodes, and to allocate the service message data of the target service to the target server nodes according to the updated load weight value. The preset number is less than the total number of server nodes processing the target service; the overloaded server nodes are characterized by their updated load weight value being less than a preset load weight value; the second alarm message indicates that the processing service resources of the preset number of server nodes processing the target service have exceeded the limit; and the target server nodes are characterized by their updated load weight value being greater than or equal to the preset load weight value.

[0111] The aforementioned dynamic disaster recovery device for cluster service offloading includes a processor and a memory. The first acquisition unit, the determination unit, the second acquisition unit, and the update unit are all stored as program units in the memory. The processor executes these program units stored in the memory to achieve the corresponding functions. All of the above modules are located in the same processor; alternatively, the modules may be located in different processors in any combination.

[0112] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured. By adjusting kernel parameters, the problem that existing intelligent traffic distribution algorithms dynamically adjust traffic allocation based on the load of cluster nodes, lacking disaster recovery mechanisms for high overall cluster loads, can be addressed.

[0113] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0114] This invention provides a computer-readable storage medium that includes a stored program, wherein the program, when running, controls the device where the computer-readable storage medium is located to execute the dynamic disaster recovery method for cluster service offloading.

[0115] Specifically, dynamic disaster recovery methods for cluster service offloading include:

[0116] Step S201: Obtain the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic.

[0117] Step S202: Determine the load weight value of each of the above server nodes based on the load information.

[0118] Step S203: Obtain the service routing information of the above-mentioned service cluster in real time. The service routing information includes the number of server nodes and the list of server nodes that process each service.

[0119] Step S204: Update the load weight value of each of the above-mentioned server nodes according to the above-mentioned service diversion information to obtain the updated load weight value of each of the above-mentioned server nodes, and generate alarm information and service diversion scheme based on the updated load weight value of each of the above-mentioned server nodes.

[0120] This invention provides a processor for running a program, wherein the program executes the dynamic disaster recovery method for cluster service offloading during runtime.

[0121] Specifically, dynamic disaster recovery methods for cluster service offloading include:

[0122] Step S201: Obtain the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic.

[0123] Step S202: Determine the load weight value of each of the above server nodes based on the load information.

[0124] Step S203: Obtain the service routing information of the above-mentioned service cluster in real time. The service routing information includes the number of server nodes and the list of server nodes that process each service.

[0125] Step S204: Update the load weight value of each of the above-mentioned server nodes according to the above-mentioned service diversion information to obtain the updated load weight value of each of the above-mentioned server nodes, and generate alarm information and service diversion scheme based on the updated load weight value of each of the above-mentioned server nodes.

[0126] This invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs at least the following steps:

[0127] Step S201: Obtain the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic.

[0128] Step S202: Determine the load weight value of each of the above server nodes based on the load information.

[0129] Step S203: Obtain the service routing information of the above-mentioned service cluster in real time. The service routing information includes the number of server nodes and the list of server nodes that process each service.

[0130] Step S204: Update the load weight value of each of the above-mentioned server nodes according to the above-mentioned service diversion information to obtain the updated load weight value of each of the above-mentioned server nodes, and generate alarm information and service diversion scheme based on the updated load weight value of each of the above-mentioned server nodes.

[0131] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0132] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program having at least the following method steps:

[0133] Step S201: Obtain the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic.

[0134] Step S202: Determine the load weight value of each of the above server nodes based on the load information.

[0135] Step S203: Obtain the service routing information of the above-mentioned service cluster in real time. The service routing information includes the number of server nodes and the list of server nodes that process each service.

[0136] Step S204: Update the load weight value of each of the above-mentioned server nodes according to the above-mentioned service diversion information to obtain the updated load weight value of each of the above-mentioned server nodes, and generate alarm information and service diversion scheme based on the updated load weight value of each of the above-mentioned server nodes.

[0137] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0138] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0139] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0142] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0143] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0144] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0145] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0146] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0147] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A dynamic disaster recovery method for cluster service offloading, characterized in that, include: Obtain the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network interface traffic; The load weight value of each server node is determined based on the load information; The service flow information of the service cluster is obtained in real time, including the number of server nodes and the list of server nodes that process each service. The load weight value of each server node is updated according to the service diversion information to obtain the updated load weight value of each server node, and alarm information and service diversion scheme are generated based on the updated load weight value of each server node.

2. The method according to claim 1, characterized in that, Based on the updated load weight values ​​of each of the aforementioned server nodes, alarm information and service diversion schemes are generated, including: If all server nodes processing the target service are detected to be overloaded server nodes, a first alarm message is generated, and the service diversion scheme is generated based on the first alarm message. The first alarm message indicates that the processing service resources of all server nodes processing the target service are overloaded, and the overloaded server node indicates that the updated load weight value of the server node is less than the preset load weight value.

3. The method according to claim 2, characterized in that, The service diversion plan is generated based on the first alarm information, including: The service message data of the target service is obtained, and the service message data is evenly distributed to each of the over-limit server nodes that process the target service based on the first alarm information.

4. The method according to claim 2, characterized in that, After generating the service diversion scheme based on the first alarm information, the method further includes: The ratio of the number of server nodes exceeding the limit that process the target service to the total number of server nodes processing the target service; If the ratio is less than a preset ratio, the first alarm message is cleared.

5. The method according to claim 1, characterized in that, Determining the load weight value of each server node based on the load information includes: Based on the load information, the load weight value of each server node is determined using threshold comparison, weighted average, and weight scaling algorithms.

6. The method according to claim 1, characterized in that, Based on the updated load weight values ​​of each of the aforementioned server nodes, alarm information and service diversion schemes are generated, including: If a preset number of server nodes processing the target service are detected as overloaded server nodes, a second alarm message is generated, and the service message data of the target service is allocated to the target server nodes according to the updated load weight value. The preset number is less than the total number of server nodes processing the target service. An overloaded server node is characterized by its updated load weight value being less than the preset load weight value. The second alarm message indicates that the preset number of server nodes processing the target service have exceeded their processing resource limits. The target server node is characterized by its updated load weight value being greater than or equal to the preset load weight value.

7. A dynamic disaster recovery system for cluster service offloading, characterized in that, include: Multiple server nodes, wherein the server nodes are IP virtual servers that are real servers deployed on the business cluster; The load balancing system includes an intelligent traffic splitting control module and a load balancing module. The intelligent traffic splitting control module is used to execute the dynamic disaster recovery method for cluster service splitting as described in any one of claims 1 to 5; The load balancing module is used to receive the load weight value output by the intelligent traffic splitting control module, determine the service splitting information according to the load weight value, and distribute the service packets to the corresponding server nodes according to the service splitting information.

8. The dynamic disaster recovery system for cluster service offloading according to claim 7, characterized in that, The system also includes: The network management platform is used for configuring and monitoring the dynamic disaster recovery system. The operation status monitoring module is used to receive alarm information output by the intelligent traffic splitting control module and the service traffic splitting information of the load balancing module, and upload the alarm information and the service traffic splitting information to the network management platform.

9. The dynamic disaster recovery system for cluster service offloading according to claim 7, characterized in that, The server node includes a load information collection client and multiple business subsystems. The load information collection client is used to collect the load information of the server node and, based on ZeroMQ's in-band telemetry technology, multiplex the business links to transmit the load information to the intelligent traffic splitting control module.

10. A dynamic disaster recovery device for cluster service offloading, characterized in that, include: The first acquisition unit is used to acquire the load information of each server node in the business cluster, wherein the load information includes at least one of CPU utilization, memory utilization, disk utilization, and network port traffic. A determining unit is used to determine the load weight value of each of the server nodes based on the load information; The second acquisition unit is used to acquire the service diversion information of the service cluster in real time. The service diversion information includes the number of server nodes processing each service and a list of server nodes. The update unit is used to update the load weight value of each server node according to the service diversion information, obtain the updated load weight value of each server node, and generate alarm information and service diversion scheme based on the updated load weight value of each server node.