Disaster recovery processing method for application system and nonvolatile storage medium
Patent Information
- Application Number
- CN202610969270.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-18
AI Technical Summary
[0005]本发明实施例提供了一种应用系统的容灾处理方法以及非易失性存储介质,以至少解决传统容灾方法因仅依赖孤立节点心跳或粗粒度集群监控而无法区分局部异常与全局风险、导致切换范围过大或过小、误切漏切频发的技术问题
[0017]In this embodiment of the invention, a disaster recovery method for an application system is employed. This involves acquiring real-time operational data of the main cluster running in the target application system. The main cluster includes multiple initial services. Based on the real-time operational data, a target topology graph corresponding to the main cluster is constructed. This target topology graph includes multiple nodes, edges between nodes, the states of each node, and the states of the edges between nodes. Each node represents an initial service, and each edge represents a call relationship between two initial services. The method addresses situations where any node has an abnormal state and/or any edge between nodes has an abnormal state. Based on the target topology map, the disaster recovery switching method of the main cluster is determined. The disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching, as well as the switching scope. The switching scope includes the entire cluster or the target group. The target group includes multiple target services in multiple initial services. This achieves the goal of quickly and accurately performing disaster recovery switching in the face of sudden failures, thereby realizing the technical effect of accurate identification and hierarchical response to the scope of failure impact. This solves the technical problem of traditional disaster recovery methods that cannot distinguish between local anomalies and global risks due to relying only on isolated node heartbeats or coarse-grained cluster monitoring, resulting in the switching scope being too large or too small, and frequent erroneous or missed switching.
Smart Images

Figure CN122777375A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a disaster recovery method for an application system and a non-volatile storage medium. Background Technology
[0002] In today's era of rapid digital transformation and cloud technology development, the importance of data is increasingly prominent. Simultaneously, the number and complexity of business services based on microservice architectures are also growing daily. Given this reality, data loss or business interruption can lead to significant economic losses for enterprises, making the development of efficient data protection and business continuity strategies particularly urgent. Disaster recovery, as a key measure to ensure business stability and data security, focuses on achieving health status monitoring and functional switching between two or more systems. This ensures that regardless of any unforeseen events—whether software failures, hardware damage, cybersecurity threats, or natural disasters—a rapid switch to a backup system can be achieved, maintaining uninterrupted application service operation.
[0003] In current cloud-native application systems, service dependencies are complex and dynamically changing. Traditional disaster recovery methods rely solely on single heartbeat detection or node-level health status assessments, making it difficult to accurately identify the scope of a fault's impact. This often leads to incorrect or untimely failovers. Due to a lack of comprehensive awareness of service call chains, resource usage, and business relationships, the system cannot distinguish between localized faults and global risks. It often blindly triggers a full cluster failover when a single non-core service fails, or fails to respond promptly even when core business chains are already affected by a cascading impact, resulting in low recovery efficiency, wasted resources, and prolonged business interruption.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This invention provides a disaster recovery method for an application system and a non-volatile storage medium to at least solve the technical problems of traditional disaster recovery methods, which rely solely on isolated node heartbeats or coarse-grained cluster monitoring and cannot distinguish between local anomalies and global risks, resulting in an excessively large or small switching range and frequent erroneous or missed switching.
[0006] According to one aspect of the present invention, a disaster recovery method for an application system is provided, comprising: acquiring real-time running data of a main cluster running in a target application system, wherein the main cluster includes multiple initial services; constructing a target topology graph corresponding to the main cluster based on the real-time running data, wherein the target topology graph includes multiple nodes, edges between multiple nodes, the states of each of the multiple nodes, and the states of each edge between the multiple nodes, where a node represents an initial service and an edge represents a call relationship between two initial services; and determining a disaster recovery switching mode for the main cluster based on the target topology graph when any one of the multiple nodes has an abnormal state and / or any one of the edges between the multiple nodes has an abnormal state, wherein the disaster recovery switching mode includes not performing a disaster recovery switching or performing a disaster recovery switching and a switching scope, the switching scope including the entire cluster or a target group, and the target group including multiple target services among the multiple initial services.
[0007] Optionally, the real-time operational data includes the operational logs of each of the multiple initial services, the memory usage data of the servers where the multiple initial services reside, the individual memory usage data of each of the multiple initial services, the call metrics between the multiple initial services, and the business types of each of the multiple initial services. Based on the real-time operational data, a target topology graph corresponding to the main cluster is constructed, including: treating the multiple initial services as multiple nodes; connecting the multiple nodes through multiple edges based on the call relationship between any two initial services to obtain an initial topology graph; determining the state of each of the multiple nodes based on the operational logs of each of the multiple initial services, the memory usage data of the servers where the multiple initial services reside, and the individual memory usage data of each of the multiple initial services; determining the state of each edge between the multiple nodes based on the call metrics between the multiple initial services; dividing the multiple nodes and the edges between the multiple nodes based on the call metrics between the multiple initial services to obtain multiple initial combinations; determining the state of each of the multiple initial combinations based on the state of each of the multiple nodes and the state of each edge between the multiple nodes; and marking the state of each of the multiple nodes, the state of each edge between the multiple nodes, the multiple initial combinations, the state of each of the multiple initial combinations, and the business types of each of the multiple initial services on the initial topology graph to obtain the target topology graph.
[0008] Optionally, based on the call metrics between multiple initial services, multiple nodes and the edges between them are divided to obtain multiple initial combinations. This includes: calculating the weights of the edges between multiple nodes based on the call relationship strength, call frequency, data traffic size, and business relevance in the call metrics; calculating the system modularity based on the weights of the edges between multiple nodes; taking one node as a first combination to obtain multiple first combinations; transferring any node from the multiple nodes to a first combination adjacent to that node, and calculating the increase in system modularity after the transfer; if the increase in system modularity is greater than a first preset threshold, retaining the above transfer operation, and continuing to traverse each node except for the arbitrary node, repeating the above transfer, calculation, and judgment steps until no more nodes are transferred, obtaining multiple second combinations; taking one second combination as an integrated node to obtain multiple integrated nodes; repeating the above transfer, calculation, and judgment steps for multiple integrated nodes until the system modularity reaches a second preset threshold, obtaining the final multiple initial combinations, wherein one initial combination includes multiple nodes, and one integrated node corresponds to multiple nodes.
[0009] Optionally, if any one of the multiple nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology map, including: if the service type of the initial service corresponding to the abnormal node is the target service type, determining the initial combination where the abnormal node is located; obtaining the state of the initial combination where the abnormal node is located; if the state of the initial combination where the abnormal node is located is abnormal, determining the disaster recovery switching method to perform disaster recovery switching and the switching scope to the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0010] Optionally, when any one of the multiple nodes is in an abnormal state and any one of the edges between the multiple nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology graph, including: when the service type of the initial service corresponding to the node in an abnormal state is not the target service type, determining the proportion of edges in an abnormal state among the edges directly or indirectly connected to the node in an abnormal state; when the proportion of edges in an abnormal state exceeds a third preset threshold, determining the disaster recovery switching method as performing a disaster recovery switching and the switching scope as the switching target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the node in an abnormal state is located.
[0011] Optionally, if any one of the multiple nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology map, including: if the memory usage data of the target server where the initial business corresponding to the abnormal node is located exceeds a fourth preset threshold, the proportion of abnormal nodes among all nodes located on the target server is determined; if the proportion of abnormal nodes exceeds a fifth preset threshold, the disaster recovery switching method is determined to be to perform a disaster recovery switching and the switching scope is the entire cluster.
[0012] Optionally, if any one of the nodes is in an abnormal state and / or any one of the edges between multiple nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology graph, including: determining the initial combination where the node with the abnormal state is located; obtaining the state of each edge and the state of each node in the initial combination where the node with the abnormal state is located; determining the proportion of abnormal nodes and / or abnormal edges; if either the proportion of abnormal nodes or the proportion of abnormal edges exceeds a sixth preset threshold, and the state of the edges between the initial combination where the abnormal node or abnormal edge is located and the adjacent combination is normal, the disaster recovery switching method is determined to be to perform disaster recovery switching and the switching scope is the target group, wherein multiple target services correspond to multiple nodes in the initial combination where the node with the abnormal state is located.
[0013] Optionally, after determining the disaster recovery switching method, a decision basis report is generated, which includes the triggering rules of the disaster recovery switching method, abnormal data in real-time operating data, abnormal subgraphs in the target topology map, the number of initial services affected, the target predicted value of recovery time, and the target predicted value of recovery point.
[0014] According to another aspect of the present invention, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored program, wherein, when the program is running, the device where the non-volatile storage medium is located controls the execution of any of the above-described application system's disaster recovery processing method.
[0015] According to another aspect of the present invention, a computer device is also provided, the computer device including a processor, the processor being configured to run a program, wherein the program executes the disaster recovery processing method of any of the above-described application systems during runtime.
[0016] According to another aspect of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the disaster recovery processing method of any of the above-described application systems.
[0017] In this embodiment of the invention, a disaster recovery method for an application system is employed. This involves acquiring real-time operational data of the main cluster running in the target application system. The main cluster includes multiple initial services. Based on the real-time operational data, a target topology graph corresponding to the main cluster is constructed. This target topology graph includes multiple nodes, edges between nodes, the states of each node, and the states of the edges between nodes. Each node represents an initial service, and each edge represents a call relationship between two initial services. The method addresses situations where any node has an abnormal state and / or any edge between nodes has an abnormal state. Based on the target topology map, the disaster recovery switching method of the main cluster is determined. The disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching, as well as the switching scope. The switching scope includes the entire cluster or the target group. The target group includes multiple target services in multiple initial services. This achieves the goal of quickly and accurately performing disaster recovery switching in the face of sudden failures, thereby realizing the technical effect of accurate identification and hierarchical response to the scope of failure impact. This solves the technical problem of traditional disaster recovery methods that cannot distinguish between local anomalies and global risks due to relying only on isolated node heartbeats or coarse-grained cluster monitoring, resulting in the switching scope being too large or too small, and frequent erroneous or missed switching. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0019] Figure 1 A hardware structure block diagram of a computer terminal for implementing a disaster recovery method for an application system is shown.
[0020] Figure 2 This is a flowchart illustrating the disaster recovery method for an application system provided according to an embodiment of the present invention;
[0021] Figure 3 This is an architecture diagram of a disaster recovery platform provided according to an optional embodiment of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0025] RTO, or Recovery Time Objective, refers to the maximum time required for the entire system to return to normal after a disaster.
[0026] RPO, or Recovery Point Objective, refers to the maximum possible duration of data loss.
[0027] Kubernetes, also known as K8S, is an open-source system for automatically deploying, scaling, and managing containerized applications.
[0028] The Kubernetes Operator is an extension of Kubernetes software that utilizes custom resource management applications and their components.
[0029] CR, or Custom Resource, is an extension mechanism in Kubernetes that allows users to define and use their own resource types to meet specific needs.
[0030] Longhorn is a lightweight, open-source distributed block storage system designed for Kubernetes, providing highly available and persistent storage solutions.
[0031] Persistent Volume Claims (PVCs) are resource objects in Kubernetes. Developers declare storage requirements (such as capacity and access mode) using PVCs without needing to concern themselves with the underlying storage details. Kubernetes automatically binds PVCs to matching PVCs, which can be actual storage resources provided by Longhorn, NFS, cloud disks, etc.
[0032] Docker is an open-source containerization platform used for developing, packaging, and running applications.
[0033] According to an embodiment of the present invention, a disaster recovery method for an application system is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal for implementing a disaster recovery method for an application system is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors (shown as 102a, 102b, ..., 102n in the figure) (the processor may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0035] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the disaster recovery processing method of the application system in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the disaster recovery processing method of the application system described above. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0037] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.
[0038] Figure 2 This is a flowchart illustrating the disaster recovery method for an application system provided according to an embodiment of the present invention, as shown below. Figure 2 As shown, the method includes the following steps:
[0039] Step S201: Obtain real-time running data of the main cluster running in the target application system, wherein the main cluster includes multiple initial services.
[0040] In this step, the target application system refers to business systems in a cloud computing environment that require disaster recovery support. These systems typically consist of multiple microservices that dynamically call and communicate with each other in a cloud-native environment, forming a complex service network. After identifying the target application systems requiring disaster recovery support, the next task is to select the primary and backup clusters. The primary cluster refers to the cluster that carries the actual business traffic during daily operation. It contains all necessary application services and data and is the core of normal business operation. The backup cluster is an environment mirrored from the primary cluster. It includes the same application services and data states as closely as possible, but is usually in a standby or low-load operating state.
[0041] Step S202: Based on real-time running data, construct the target topology graph corresponding to the main cluster. The target topology graph includes multiple nodes, edges between multiple nodes, the state of each node, and the state of each edge between multiple nodes. A node represents an initial service, and an edge represents the calling relationship between two initial services.
[0042] In this step, the target topology graph constructed based on real-time operational data is a dynamic knowledge graph formed by treating each initial service in the main cluster as a node and the actual call relationships between services as edges. The node status is calculated comprehensively from multiple dimensions such as service logs, container resource utilization, and server load, representing the health level of the service (e.g., normal, sub-healthy, or faulty). The edge status is dynamically labeled based on KPIs such as call success rate, response latency, jitter, and packet loss rate (e.g., normal, degraded, or interrupted), thus truly reflecting the interaction quality and dependency strength between services. This topology graph not only depicts the static call structure between services but also integrates real-time operational status and business attributes, forming a dynamic topology model with semantic awareness. It surpasses the point-like perception of traditional heartbeat detection, achieving a global understanding of fault propagation paths, impact range, and business relevance.
[0043] Step S203: In the case that any one of the nodes is in an abnormal state and / or any one of the edges between the nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology graph. The disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching and the switching scope. The switching scope includes the entire cluster or the target group. The target group includes multiple target services in multiple initial services.
[0044] In this step, when any node and / or edge in the target topology graph experiences an anomaly, it doesn't simply trigger a full cluster switchover. Instead, it uses intelligent reasoning based on the semantic relationships of the graph structure: First, it identifies the "target community" composed of highly coupled services through graph community partitioning. Then, it comprehensively judges the scope of the fault's impact by combining the node's business level (such as the core business of PO), the propagation ratio of the abnormal edge, the anomaly density within the community, and its isolation from neighboring communities. If the anomaly is concentrated in a certain target community and has not spread to other communities, only that community is switched over; if the anomaly originates from the underlying infrastructure and affects multiple communities, or if the core business community is abnormal as a whole, a full cluster switchover is triggered; if the anomaly is an isolated event and has not caused a chain reaction, no switchover is performed. This mechanism achieves a leap from "point-based alarms" to "graph-level impact assessment," enabling disaster recovery decisions to accurately match the actual scope of the fault's impact. It avoids the resource waste and business disruption caused by the "one-size-fits-all" approach of traditional methods, significantly improving the intelligence, security, and efficiency of disaster recovery.
[0045] Through the above steps, the goal of rapid and accurate disaster recovery switching in the face of sudden failures is achieved, thereby realizing the technical effect of accurate identification and graded response to the scope of failure impact. This solves the technical problem that traditional disaster recovery methods cannot distinguish between local anomalies and global risks due to relying only on isolated node heartbeats or coarse-grained cluster monitoring, resulting in an excessively large or small switching range and frequent erroneous or missed switching.
[0046] As an optional implementation, the real-time runtime data includes the runtime logs of each of the multiple initial services, the memory usage data of the servers where the multiple initial services reside, the individual memory usage data of each of the multiple initial services, the call metrics between the multiple initial services, and the service type of each of the multiple initial services. Based on the real-time runtime data, a target topology graph corresponding to the main cluster is constructed, including: treating the multiple initial services as multiple nodes; connecting the multiple nodes through multiple edges based on the call relationship between any two initial services to obtain an initial topology graph; determining the state of each of the multiple nodes based on the runtime logs of each of the multiple initial services, the memory usage data of the servers where the multiple initial services reside, and the individual memory usage data of each of the multiple initial services; determining the state of each edge between the multiple nodes based on the call metrics between the multiple initial services; dividing the multiple nodes and the edges between the multiple nodes based on the call metrics between the multiple initial services to obtain multiple initial combinations; determining the state of each of the multiple initial combinations based on the state of each of the multiple nodes and the state of each edge between the multiple nodes; and marking the state of each of the multiple nodes, the state of each edge between the multiple nodes, the multiple initial combinations, the state of each of the multiple initial combinations, and the service type of each of the multiple initial services on the initial topology graph to obtain a target topology graph.
[0047] Optionally, the process of constructing the target topology graph based on real-time runtime data is a core step in abstracting the microservice system into a dynamic knowledge graph. First, each initial business is mapped to a node in the graph, and edges are established based on the call relationships between businesses to form the initial topology. Then, the node status (e.g., healthy, sub-healthy, faulty) is comprehensively evaluated by combining the runtime logs of each business (e.g., OOM exceptions, call errors), the memory usage of the server, and the memory consumption of the business itself. Simultaneously, the health status of edges (e.g., normal, deteriorated, interrupted) is determined based on call metrics (latency, error rate, success rate). On this basis, the graph is weighted by call intensity, frequency, traffic, and business logic relevance, and the Louvain algorithm is used to partition the graph into communities, forming multiple initial combinations (i.e., microservice clusters). Each combination is then assigned an aggregate health score based on the combined status of its internal nodes and edges. Finally, all node states, edge states, community division results, community states, and business priority labels (such as PO / P1) are uniformly labeled on the topology graph to form a target topology graph with rich semantics and dynamic perception capabilities. This provides a structured and reasonable decision-making basis for subsequent accurate identification of root causes, assessment of the scope of impact, and implementation of fine-grained disaster recovery switching.
[0048] As an optional embodiment, based on the call metrics between multiple initial services, multiple nodes and the edges between them are divided to obtain multiple initial combinations. This includes: calculating the weights of the edges between multiple nodes based on the call relationship strength, call frequency, data traffic size, and business relevance in the call metrics; calculating the system modularity based on the weights of the edges between multiple nodes; taking one node as a first combination to obtain multiple first combinations; transferring any node from the multiple nodes to the first combination adjacent to that node, and calculating the increase in system modularity after the transfer; if the increase in system modularity is greater than a first preset threshold, retaining the above transfer operation, and continuing to traverse each node except for the arbitrary node, repeating the above transfer, calculation, and judgment steps until no more nodes are transferred, to obtain multiple second combinations; taking one second combination as an integrated node to obtain multiple integrated nodes; repeating the above transfer, calculation, and judgment steps for multiple integrated nodes until the system modularity reaches a second preset threshold, to obtain the final multiple initial combinations, wherein one initial combination includes multiple nodes, and one integrated node corresponds to multiple nodes.
[0049] Optionally, this process does not simply cluster based on a single call frequency or connection count. Instead, it integrates four dimensions of metrics: call relationship strength, call frequency, data traffic volume, and business relevance, constructing a weighted graph structure with business semantic awareness. Call relationship strength reflects the tightness of inter-service dependencies; for example, interfaces with high-frequency and stable calls have higher weights. Call frequency reflects the activity of service interactions; services with frequent interactions are more likely to form logical loops. Data traffic volume characterizes the load pressure of inter-service communication from a resource consumption perspective; high-traffic links often carry critical business flows. Business relevance is derived from metadata such as service naming conventions, deployment namespaces, and business tags (e.g., PO / P1) to identify service groups with the same business domain or functional modules. By normalizing and weighting these four heterogeneous metrics, each edge receives a comprehensive weight, making the graph reflect not only the technical call topology but also the inherent connections in business logic. Based on this, system modularity, a classic metric for measuring the quality of community partitioning, is used to quantify the clustering effect of "tight connections within communities and sparse connections between communities." The algorithm initially treats each node as an independent combination. Then, using a greedy strategy, it attempts to migrate nodes one by one to adjacent combinations, calculating the increment in modularity after migration. Migration is only accepted if this increment exceeds a first preset threshold (e.g., 0.01), ensuring that each adjustment brings structural optimization. This process continues to traverse all nodes until no nodes can be migrated, forming a preliminary "second combination." Subsequently, the algorithm enters the second stage—hierarchical aggregation: each second combination is abstracted as an "integrated node." At this point, the nodes in the new graph are combinations, and the edges represent the call relationships between combinations. The weight calculation and modularity optimization process is repeated until the system modularity reaches a second preset threshold (e.g., 0.75). The resulting multiple final combinations are microservice community units with high cohesion and low coupling. This process is essentially an enhanced implementation of the Louvain algorithm on a business-aware graph. Its advantage lies not only in its technical topology but also in its integration of business semantics and operational load, giving the partitioning results practical operational significance and providing a precise structural basis for subsequent "community-based switching" rather than "full cluster switching."
[0050] As an optional implementation, in the event that any one of the multiple nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology map, including: if the service type of the initial service corresponding to the abnormal node is the target service type, determining the initial combination where the abnormal node is located; obtaining the state of the initial combination where the abnormal node is located; if the state of the initial combination where the abnormal node is located is abnormal, determining the disaster recovery switching method to perform disaster recovery switching and the switching scope to the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0051] Optionally, when an abnormal node status is detected, the system first determines whether its corresponding initial business is a core business (i.e., the target business type). If so, it immediately locates the initial group to which the node belongs—that is, the highly cohesive microservice community partitioned by the Louvain algorithm. Subsequently, it calculates the overall health by combining the status of all nodes (failure, sub-health) and the status of edges (interruption, degradation) within the group, and determines whether the group is abnormal as a whole. If a high proportion of abnormal nodes or critical links are interrupted within the group, indicating that the fault has spread to the entire business unit rather than being an isolated event, it triggers a precise disaster recovery switch for the "target group," switching only all related services within the group, rather than migrating the entire cluster.
[0052] As an optional embodiment, when any one of the multiple nodes has an abnormal state and any one of the edges between the multiple nodes has an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology graph, including: when the service type of the initial service corresponding to the abnormal node is not the target service type, determining the proportion of abnormal edges among the edges directly or indirectly connected to the abnormal node; when the proportion of abnormal edges exceeds a third preset threshold, determining the disaster recovery switching method as performing disaster recovery switching and the switching scope as the switching target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0053] Optionally, when an abnormal node corresponds to a non-core business (non-target business type), the abnormality is not simply ignored. Instead, the system traces its directly affected links back through the target topology graph, statistically analyzing the proportion of abnormal edges in a "degraded" or "interrupted" state among all edges directly or indirectly connected to it. This proportion reflects whether the abnormal node has triggered cascading fault propagation. If this proportion exceeds a third preset threshold (e.g., 30%), it indicates that even if the initial fault originated from a non-core service, its impact has penetrated to the critical business path, posing a potential risk of business link disruption. At this point, the system determines that the abnormality constitutes a "hidden core threat" and locates the initial combination to which the node belongs—that is, the highly cohesive service group partitioned by the Louvain algorithm. Due to the high coupling of services within this combination, the propagation of abnormal edges is essentially a collapse of the group's internal stability. Therefore, a precise disaster recovery switch for the "target group" is triggered, rather than a full cluster switch, thereby blocking fault propagation and ensuring the availability of core services without affecting other normal businesses.
[0054] As an optional implementation, in the event that any one of the multiple nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology map, including: if the memory usage data of the target server where the initial service corresponding to the abnormal node is located exceeds a fourth preset threshold, the proportion of abnormal nodes among all nodes located on the target server is determined; if the proportion of abnormal nodes exceeds a fifth preset threshold, the disaster recovery switching method is determined to be to perform a disaster recovery switching and the switching scope is the entire cluster.
[0055] Optionally, when the memory utilization of a target server hosting a node exceeds the fourth preset threshold (e.g., 95%), the focus shifts from solely addressing individual service anomalies to elevating the risk assessment to the infrastructure layer. This involves comprehensively analyzing the proportion of all microservice nodes deployed on that server that are in a "faulty" or "sub-healthy" state. If this proportion exceeds the fifth preset threshold (e.g., 40%), it indicates that the physical node has experienced systemic resource degradation, potentially triggering a large-scale service cascading crash. The risk exceeds the controllable range of a single business group. At this point, the server is considered a "high-risk fault domain," where multiple service combinations are at risk of overall failure. Partial failover cannot eliminate the persistent threat at the underlying hardware or host level. To ensure overall system stability, a "full cluster disaster recovery switch" is triggered, migrating all services from the primary cluster to the backup cluster, achieving risk isolation and system-level recovery.
[0056] As an optional embodiment, in the case where any one of the nodes is in an abnormal state and / or any one of the edges between the nodes is in an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology graph, including: determining the initial combination where the node with the abnormal state is located; obtaining the state of each edge and the state of each node in the initial combination where the node with the abnormal state is located; determining the proportion of abnormal nodes and / or abnormal edges; if either the proportion of abnormal nodes or the proportion of abnormal edges exceeds a sixth preset threshold, and the state of the edges between the initial combination where the abnormal node or abnormal edge is located and the adjacent combination is normal, the disaster recovery switching method is determined to be to perform disaster recovery switching and the switching scope is the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the node with the abnormal state is located.
[0057] Optionally, when an anomaly is detected in a node or edge, the initial combination to which it belongs (i.e., the microservice community divided by the Louvain algorithm) is first located, and the health status within the combination is comprehensively assessed: the proportion of all nodes in the combination that are in a "faulty" or "sub-healthy" state, and the proportion of all edges that are in a "degraded" or "interrupted" state are statistically analyzed. If either of these proportions exceeds the sixth preset value (e.g., 40%), it indicates that a systemic functional degradation has occurred within the combination, exhibiting obvious characteristics of local collapse. At this point, the status of the connection edges between this combination and adjacent combinations is further examined. If these cross-combination edges remain "normal," it indicates that the anomaly has not spread externally, the isolation is good, and the fault is effectively confined within the current community. This condition means that the combination is a "self-consistent fault unit," and its anomaly originates from internal logic or resource issues, rather than global interference. Under this premise, it is determined that a full cluster switch is not necessary; instead, a precise disaster recovery switch should be implemented only for this "isolated" anomaly group, migrating all related services within it to the backup cluster to achieve minimal disturbance to fault isolation and business recovery.
[0058] As an optional implementation, after determining the disaster recovery switching method, a decision basis report is generated. The decision basis report includes the triggering rules of the disaster recovery switching method, abnormal data in the real-time running data, abnormal subgraphs in the target topology map, the number of affected initial services, the target predicted value of recovery time, and the target predicted value of recovery point.
[0059] Optionally, after determining the disaster recovery switchover method, a structured decision-making basis report is automatically generated, achieving a closed loop from automated decision-making to auditable and interpretable data. This report fully traces the decision-making logic: First, it clarifies the triggering rules (e.g., "the anomaly rate of the core business combination exceeds 40%), and attaches the original anomaly data (e.g., OOM logs for service A, service B call error rate soaring to 48%), ensuring that each judgment is supported by data; then, it embeds an anomaly subgraph from the target topology map, visually displaying the degraded edges of faulty nodes and the affected initial combinations, allowing operations personnel to intuitively understand the impact path; simultaneously, it counts the number of initially affected services, quantifying the scope of business impact; finally, it combines historical recovery data with the current topology scale to predict RTO (Recovery Time Objective) and RPO (Recovery Point Objective), providing a quantitative basis for switchover timeliness and data integrity.
[0060] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0061] Through the above description of the embodiments, those skilled in the art can clearly understand that the disaster recovery processing method of the application system according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0062] According to an optional embodiment of the present invention, a disaster recovery platform is also provided. Figure 3 This is an architecture diagram of a disaster recovery platform provided according to an optional embodiment of the present invention, such as... Figure 3 As shown, the platform includes an intelligent analysis module, a cluster management module, and a data management module. The platform will be described below.
[0063] The intelligent analysis module applies any of the disaster recovery methods used in the above-mentioned application systems.
[0064] Optionally, the intelligent analysis module obtains basic data from the data management module and performs intelligent analysis to provide information for intelligent disaster recovery switching. Based on GraphRAG, the intelligent analysis module can perform intelligent microservice relationship analysis, using multi-source data to generate target topology graphs between services, create microservice communities, and provide recovery paths and scope for disaster recovery switching. The intelligent analysis module runs in Kubernetes Operator mode, and disaster recovery information is stored through custom resources (CRs). The module monitors resource status and performs management operations, using a control loop to continuously compare the expected state with the actual state, ensuring the application system is always in a healthy state.
[0065] The cluster management module is used to manage all clusters in the target application system.
[0066] Optionally, the cluster management module uses Kubernetes technology to manage service operation. For example... Figure 3 As shown, the cluster management module can include one primary cluster management module and one backup cluster management module. The primary cluster management module and the backup cluster management module have the ability to sense each other's heartbeats. Combined with the intelligent analysis module, it can achieve automated anomaly detection and disaster recovery switching.
[0067] The data management module is used to collect and store the operational data of all services in the cluster.
[0068] Optionally, the data management module can be divided into multiple modules in a real disaster recovery scenario, and the backup data management module can be configured to be multiple according to actual needs. The data management module is not only responsible for collecting and storing business operation data from all clusters, but also ensures the synchronization and consistency of data between the primary and backup clusters, and is a key information source for the intelligent analysis module to make disaster recovery decisions.
[0069] The platform described above automates the entire recovery process, from event detection to backup data, selecting services to restore, and executing recovery tasks, integrating GraphRAG, the cloud platform, and data backup tools. Due to the complexity of the recovery process, this automation prevents errors that may occur during manual operations and ensures consistent recovery, thereby improving the overall performance and stability of the system.
[0070] Embodiments of the present invention may provide a computer device. Optionally, in this embodiment, the computer device may be located in at least one of a plurality of network devices in a computer network. The computer device includes a memory and a processor.
[0071] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the disaster recovery processing method and apparatus of the application system in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the disaster recovery processing method of the application system described above. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0072] The processor can access information and applications stored in memory via a transmission device to perform the following steps: Obtain real-time operational data of the main cluster running in the target application system, wherein the main cluster includes multiple initial services; Based on the real-time operational data, construct a target topology graph corresponding to the main cluster, wherein the target topology graph includes multiple nodes, edges between multiple nodes, the states of each node, and the states of each edge between multiple nodes, where a node represents an initial service and an edge represents the calling relationship between two initial services; In the event that any node among the multiple nodes has an abnormal state and / or any edge among the multiple nodes has an abnormal state, determine the disaster recovery switching method of the main cluster based on the target topology graph, wherein the disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching, and the switching scope, where the switching scope includes the entire cluster or a target group, and the target group includes multiple target services among the multiple initial services.
[0073] Optionally, the processor may also execute program code for the following steps: Real-time runtime data includes the runtime logs of multiple initial services, memory usage data of the servers where the multiple initial services reside, memory usage data of each of the multiple initial services, call metrics between the multiple initial services, and the service type of each of the multiple initial services. Based on the real-time runtime data, a target topology graph corresponding to the main cluster is constructed, including: treating the multiple initial services as multiple nodes; connecting the multiple nodes through multiple edges based on the call relationship between any two initial services to obtain an initial topology graph; determining the state of each of the multiple nodes based on the runtime logs of the multiple initial services, the memory usage data of the servers where the multiple initial services reside, and the memory usage data of each of the multiple initial services; determining the state of each edge between the multiple nodes based on the call metrics between the multiple initial services; dividing the multiple nodes and the edges between the multiple nodes based on the call metrics between the multiple initial services to obtain multiple initial combinations; determining the state of each of the multiple initial combinations based on the state of each of the multiple nodes and the state of each edge between the multiple nodes; and marking the state of each of the multiple nodes, the state of each edge between the multiple nodes, the multiple initial combinations, the state of each of the multiple initial combinations, and the service type of each of the multiple initial services on the initial topology graph to obtain a target topology graph.
[0074] Optionally, the processor may also execute program code with the following steps: based on the call metrics between multiple initial services, divide multiple nodes and the edges between multiple nodes to obtain multiple initial combinations, including: calculating the weights of the edges between multiple nodes based on the call relationship strength, call frequency, data traffic size, and business relevance in the call metrics; calculating the system modularity based on the weights of the edges between multiple nodes; taking one node as a first combination to obtain multiple first combinations; transferring any node from the multiple nodes to the first combination adjacent to any node, and calculating the increase in system modularity after the transfer; if the increase in system modularity is greater than a first preset threshold, retaining the above transfer operation, continuing to traverse each node except for any node in the multiple nodes, repeating the above transfer, calculation, and judgment steps until the node no longer needs to be transferred, to obtain multiple second combinations; taking one second combination as an integrated node to obtain multiple integrated nodes; repeating the above transfer, calculation, and judgment steps for multiple integrated nodes until the system modularity reaches a second preset threshold, to obtain the final multiple initial combinations, wherein one initial combination includes multiple nodes, and one integrated node corresponds to multiple nodes.
[0075] Optionally, the processor may also execute program code with the following steps: In the event that any one of the multiple nodes is in an abnormal state, based on the target topology, determine the disaster recovery switching method of the main cluster, including: if the service type of the initial service corresponding to the abnormal node is the target service type, determine the initial combination where the abnormal node is located; obtain the state of the initial combination where the abnormal node is located; if the state of the initial combination where the abnormal node is located is abnormal, determine that the disaster recovery switching method is to perform disaster recovery switching and the switching scope is the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0076] Optionally, the processor may also execute program code with the following steps: In the case where any one of the multiple nodes is in an abnormal state and any one of the edges between the multiple nodes is in an abnormal state, based on the target topology graph, determine the disaster recovery switching method for the main cluster, including: if the service type of the initial service corresponding to the abnormal node is not the target service type, determine the proportion of abnormal edges among the edges directly or indirectly connected to the abnormal node; if the proportion of abnormal edges exceeds a third preset threshold, determine the disaster recovery switching method as performing a disaster recovery switching and the switching scope as the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0077] Optionally, the processor may also execute program code with the following steps: In the event that any one of the nodes is in an abnormal state, based on the target topology, determine the disaster recovery switching method for the main cluster, including: if the memory usage data of the target server where the initial service corresponding to the abnormal node is located exceeds a fourth preset threshold, determine the proportion of abnormal nodes among all nodes located on the target server; if the proportion of abnormal nodes exceeds a fifth preset threshold, determine that the disaster recovery switching method is to perform a disaster recovery switching and the switching scope is the entire cluster.
[0078] Optionally, the processor may also execute program code with the following steps: In the case where any one of the nodes is in an abnormal state and / or any one of the edges between the nodes is in an abnormal state, based on the target topology graph, determine the disaster recovery switching method for the main cluster, including: determining the initial combination where the node with the abnormal state is located; obtaining the state of each edge and each node in the initial combination where the node with the abnormal state is located; determining the proportion of abnormal nodes and / or abnormal edges; if either the proportion of abnormal nodes or the proportion of abnormal edges exceeds a sixth preset threshold, and the state of the edges between the initial combination where the abnormal node or abnormal edge is located and adjacent combinations is normal, determine the disaster recovery switching method as performing a disaster recovery switching, with the switching scope being the target group, wherein multiple target services correspond to multiple nodes in the initial combination where the node with the abnormal state is located.
[0079] Optionally, the processor may also execute program code for the following steps: after determining the disaster recovery switching method, generating a decision basis report, wherein the decision basis report includes the triggering rules of the disaster recovery switching method, abnormal data in real-time running data, abnormal subgraphs in the target topology map, the number of affected initial services, the target predicted value of recovery time, and the target predicted value of recovery point.
[0080] The present invention provides a disaster recovery method for an application system. By acquiring real-time operational data of the main cluster running in the target application system, which includes multiple initial services, a target topology graph corresponding to the main cluster is constructed based on the real-time operational data. This graph includes multiple nodes, edges between nodes, the individual states of nodes, and the states of edges between nodes. Each node represents an initial service, and each edge represents the call relationship between two initial services. In the event of any abnormal state of any node and / or any abnormal state of any edge between nodes, the disaster recovery switching method for the main cluster is determined based on the target topology graph. This method includes either no disaster recovery switching or a disaster recovery switching operation, as well as the switching scope. The switching scope includes the entire cluster or a target group, where the target group includes multiple target services among the initial services. This achieves the goal of rapid and accurate disaster recovery switching in the face of sudden failures, thus realizing the technical effect of accurate identification and graded response to the scope of fault impact. Furthermore, it solves the technical problems of traditional disaster recovery methods that rely solely on isolated node heartbeats or coarse-grained cluster monitoring, which cannot distinguish between local anomalies and global risks, leading to excessively large or small switching scopes and frequent erroneous or missed switching.
[0081] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a non-volatile storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0082] Embodiments of the present invention also provide a non-volatile storage medium. Optionally, in this embodiment, the aforementioned non-volatile storage medium can be used to store the program code executed by the disaster recovery processing method of the application system provided in the above embodiments.
[0083] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0084] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining real-time running data of the main cluster running in the target application system, wherein the main cluster includes multiple initial services; constructing a target topology graph corresponding to the main cluster based on the real-time running data, wherein the target topology graph includes multiple nodes, edges between multiple nodes, the states of each of the multiple nodes, and the states of each edge between multiple nodes, where a node represents an initial service and an edge represents the calling relationship between two initial services; in the event that any node among the multiple nodes has an abnormal state and / or any edge among the multiple nodes has an abnormal state, determining the disaster recovery switching method of the main cluster based on the target topology graph, wherein the disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching and the switching scope, the switching scope including the entire cluster or target group, and the target group including multiple target services among the multiple initial services.
[0085] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: real-time runtime data includes runtime logs of each of the multiple initial services, memory usage data of the servers where the multiple initial services reside, memory usage data of each of the multiple initial services, call metrics between the multiple initial services, and the service types of each of the multiple initial services. Based on the real-time runtime data, a target topology graph corresponding to the main cluster is constructed, including: treating the multiple initial services as multiple nodes; connecting the multiple nodes through multiple edges based on the call relationship between any two initial services to obtain an initial topology graph; and based on the runtime logs of each of the multiple initial services, memory usage data of the servers where the multiple initial services reside, call metrics between the multiple initial services, and the service types of each of the multiple initial services. The memory usage data of the server where the service is located and the memory usage data of each of the multiple initial services are used to determine the state of each of the multiple nodes; based on the call indicators between the multiple initial services, the state of each edge between the multiple nodes is determined; based on the call indicators between the multiple initial services, the multiple nodes and the edges between the multiple nodes are divided to obtain multiple initial combinations; based on the state of each of the multiple nodes and the state of each edge between the multiple nodes, the state of each of the multiple initial combinations, the state of each of the multiple initial combinations, and the business type of each of the multiple initial services are marked on the initial topology graph to obtain the target topology graph.
[0086] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: based on the call indicators between multiple initial services, multiple nodes and the edges between multiple nodes are divided to obtain multiple initial combinations, including: calculating the weights corresponding to the edges between multiple nodes based on the call relationship strength, call frequency, data traffic size, and business relevance in the call indicators; calculating the system modularity based on the weights corresponding to the edges between multiple nodes; taking one node as a first combination to obtain multiple first combinations; transferring any one of the multiple nodes to the first combination adjacent to any one node, and calculating the increase in system modularity after the transfer; if the increase in system modularity is greater than a first preset threshold, retaining the above transfer operation, continuing to traverse each node except for any one node in the multiple nodes, repeating the above transfer, calculation, and judgment steps until the node no longer needs to be transferred, to obtain multiple second combinations; taking one second combination as an integrated node to obtain multiple integrated nodes; repeating the above transfer, calculation, and judgment steps for multiple integrated nodes until the system modularity reaches a second preset threshold, to obtain the final multiple initial combinations, wherein one initial combination includes multiple nodes, and one integrated node corresponds to multiple nodes.
[0087] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: when any one of the multiple nodes is in an abnormal state, based on the target topology, determine the disaster recovery switching method of the main cluster, including: when the service type of the initial service corresponding to the abnormal node is the target service type, determine the initial combination where the abnormal node is located; obtain the state of the initial combination where the abnormal node is located; when the state of the initial combination where the abnormal node is located is abnormal, determine the disaster recovery switching method to perform disaster recovery switching and the switching scope is the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0088] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: In the case where any one of the multiple nodes is in an abnormal state and any one of the edges between the multiple nodes is in an abnormal state, based on the target topology graph, determine the disaster recovery switching method for the main cluster, including: if the service type of the initial service corresponding to the abnormal node is not the target service type, determine the proportion of abnormal edges among the edges directly or indirectly connected to the abnormal node; if the proportion of abnormal edges exceeds a third preset threshold, determine the disaster recovery switching method as performing a disaster recovery switching and the switching scope as the target group, wherein the multiple target services respectively correspond to multiple nodes in the initial combination where the abnormal node is located.
[0089] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: in the event that any one of the multiple nodes is in an abnormal state, based on the target topology, determine the disaster recovery switching method of the main cluster, including: if the memory usage data of the target server where the initial service corresponding to the abnormal node is located exceeds a fourth preset threshold, determine the proportion of abnormal nodes among all nodes located on the target server; if the proportion of abnormal nodes exceeds a fifth preset threshold, determine that the disaster recovery switching method is to perform a disaster recovery switching and the switching scope is the entire cluster.
[0090] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: In the case where any one of the multiple nodes is in an abnormal state and / or any one of the edges between multiple nodes is in an abnormal state, based on the target topology graph, determine the disaster recovery switching method for the main cluster, including: determining the initial combination where the abnormal node is located; obtaining the state of each edge and the state of each node in the initial combination where the abnormal node is located; determining the proportion of abnormal nodes and / or abnormal edges; if either the proportion of abnormal nodes or the proportion of abnormal edges exceeds a sixth preset threshold, and the state of the edges between the initial combination where the abnormal node or abnormal edge is located and adjacent combinations is normal, determine the disaster recovery switching method as performing a disaster recovery switching, and the switching scope is the target group, wherein multiple target services correspond to multiple nodes in the initial combination where the abnormal node is located.
[0091] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: after determining the disaster recovery switching method, generating a decision basis report, wherein the decision basis report includes the triggering rules of the disaster recovery switching method, abnormal data in the real-time running data, abnormal subgraphs in the target topology map, the number of affected initial services, the target predicted value of recovery time, and the target predicted value of recovery point.
[0092] Embodiments of the present invention also provide a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it can: acquire real-time running data of a main cluster running in a target application system, wherein the main cluster includes multiple initial services; construct a target topology graph corresponding to the main cluster based on the real-time running data, wherein the target topology graph includes multiple nodes, edges between multiple nodes, the states of each of the multiple nodes, and the states of each edge between the multiple nodes, where a node represents an initial service and an edge represents a call relationship between two initial services; and determine a disaster recovery switching method for the main cluster based on the target topology graph when any node among the multiple nodes has an abnormal state and / or any edge among the multiple nodes has an abnormal state, wherein the disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching and the switching scope, the switching scope including the entire cluster or a target group, and the target group including multiple target services among the multiple initial services.
[0093] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0094] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0095] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0099] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A disaster recovery method for an application system, characterized in that, include: Obtain real-time running data of the main cluster running in the target application system, wherein the main cluster includes multiple initial services; Based on the real-time running data, a target topology graph corresponding to the main cluster is constructed. The target topology graph includes multiple nodes, edges between the multiple nodes, the state of each of the multiple nodes, and the state of each edge between the multiple nodes. Each node represents one of the initial services, and each edge represents the calling relationship between two initial services. If any one of the multiple nodes has an abnormal state and / or any one of the edges between the multiple nodes has an abnormal state, the disaster recovery switching method of the main cluster is determined based on the target topology graph. The disaster recovery switching method includes not performing disaster recovery switching or performing disaster recovery switching and the switching scope. The switching scope includes the entire cluster or the target group. The target group includes multiple target services among the multiple initial services.
2. The method according to claim 1, characterized in that, The real-time operational data includes the operational logs of each of the multiple initial services, the memory usage data of the servers where the multiple initial services reside, the memory usage data of each of the multiple initial services, the call metrics between the multiple initial services, and the service type of each of the multiple initial services. The step of constructing the target topology map corresponding to the main cluster based on the real-time operational data includes: The multiple initial services are used as the multiple nodes; Based on the calling relationship between any two of the multiple initial services, the multiple nodes are connected through multiple edges to obtain an initial topology graph; Based on the operation logs of each of the multiple initial services, the memory usage data of the servers where the multiple initial services are located, and the memory usage data of each of the multiple initial services, the status of each of the multiple nodes is determined; Based on the call metrics among the multiple initial services, determine the state of each edge among the multiple nodes; Based on the call metrics between the multiple initial services, the multiple nodes and the edges between the multiple nodes are divided to obtain multiple initial combinations; Based on the individual states of the nodes and the states of the edges between the nodes, the states of the initial combinations are determined. The states of each of the multiple nodes, the states of the edges between the multiple nodes, the multiple initial combinations, the states of the multiple initial combinations, and the service types of the multiple initial services are marked on the initial topology graph to obtain the target topology graph.
3. The method according to claim 2, characterized in that, Based on the call metrics between the multiple initial services, the multiple nodes and the edges between them are divided to obtain multiple initial combinations, including: Based on the call relationship strength, call frequency, data traffic volume, and business relevance in the call metrics, calculate the weights of the edges between the multiple nodes. The system modularity is calculated based on the weights of the edges between the nodes. By taking one of the nodes as a first combination, multiple first combinations are obtained; Transfer any one of the plurality of nodes to the first combination adjacent to the arbitrary node, and calculate the increase in the modularity of the system after the transfer; If the increase in the system modularity exceeds the first preset threshold, the above transfer operation is retained, and the process continues to traverse each node except for any one of the nodes, repeating the above transfer, calculation and judgment steps until the node no longer undergoes transfer, resulting in multiple second combinations; By using one of the second combinations as an integration node, multiple integration nodes are obtained; The above-described transfer, calculation, and judgment steps are repeated for the plurality of integrated nodes until the system modularity reaches a second preset threshold, thereby obtaining the final plurality of initial combinations, wherein one initial combination includes a plurality of nodes and one integrated node corresponds to a plurality of nodes.
4. The method according to claim 2, characterized in that, In the event that any one of the multiple nodes is in an abnormal state, the disaster recovery switchover method for the primary cluster is determined based on the target topology, including: If the service type of the initial service corresponding to the node with the abnormal status is the target service type, determine the initial combination where the node with the abnormal status is located; Obtain the state of the initial combination where the node with the abnormal state is located; If the initial combination where the node with the abnormal status is located is in an abnormal state, the disaster recovery switching method is determined to be performing a disaster recovery switching and the switching scope is switching the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the node with the abnormal status is located.
5. The method according to claim 2, characterized in that, In the case where any one of the multiple nodes is in an abnormal state and any one of the edges between the multiple nodes is in an abnormal state, the disaster recovery switchover method of the main cluster is determined based on the target topology graph, including: If the business type of the initial business corresponding to the node with abnormal state is not the target business type, determine the proportion of edges with abnormal state among the edges that are directly or indirectly connected to the node with abnormal state. If the proportion of edges with abnormal states exceeds a third preset threshold, the disaster recovery switching method is determined to be performing a disaster recovery switching and the switching scope is switching the target group, wherein the multiple target services correspond to multiple nodes in the initial combination where the nodes with abnormal states are located.
6. The method according to claim 2, characterized in that, In the event that any one of the multiple nodes is in an abnormal state, the disaster recovery switchover method for the primary cluster is determined based on the target topology, including: If the memory usage data of the target server where the initial service corresponding to the node with abnormal status is located exceeds the fourth preset threshold, determine the proportion of nodes with abnormal status among all nodes located on the target server. If the proportion of nodes in abnormal state exceeds a fifth preset threshold, the disaster recovery switching method is determined to be performing a disaster recovery switch and the switching scope is the entire cluster.
7. The method according to claim 2, characterized in that, In the event that any one of the multiple nodes is in an abnormal state and / or any one of the edges between the multiple nodes is in an abnormal state, the disaster recovery switchover method of the main cluster is determined based on the target topology graph, including: Determine the initial combination of nodes in the abnormal state; Obtain the state of each edge and the state of each node in the initial combination where the node with the abnormal state is located; Determine the proportion of abnormal nodes and / or abnormal edges; If either the proportion of abnormal nodes or the proportion of abnormal edges exceeds a sixth preset threshold, and the state of the edge between the initial combination and the adjacent combination where the abnormal node or the abnormal edge is located is normal, the disaster recovery switching method is determined to be performing disaster recovery switching and the switching range is switching the target group, wherein the multiple target services respectively correspond to multiple nodes in the initial combination where the abnormal node is located.
8. The method according to claim 1, characterized in that, Also includes: After determining the disaster recovery switching method, a decision basis report is generated, which includes the triggering rules of the disaster recovery switching method, abnormal data in the real-time operating data, abnormal subgraphs in the target topology map, the number of affected initial services, the target predicted value of recovery time, and the target predicted value of recovery point.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein, when the program is running, it controls the device where the non-volatile storage medium is located to execute the disaster recovery processing method of the application system according to any one of claims 1 to 8.
10. A computer device, characterized in that, include: Memory and processor The memory stores computer programs; The processor is configured to execute a computer program stored in the memory, wherein when the computer program is executed, the processor performs the disaster recovery processing method of the application system according to any one of claims 1 to 8.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the disaster recovery processing method of the application system according to any one of claims 1 to 8.