Disaster recovery visualization display system and method

By generating real-time snapshot data at nodes and sending it to the virtual machine, the virtual machine generates a target copy and sends it to the cloud operation and maintenance center for visual display, it solves the problem that failure association relationships are difficult to quickly obtain in complex network topology, and achieves fast and efficient fault display and improves system security.

CN114398204BActive Publication Date: 2025-07-18CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111541494.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-07-18
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The prior art is difficult to quickly acquire and demonstrate the association between failures in complex network topology, which makes it take a long time to display the failure of the disaster recovery system.

Method used

Real-time snapshot data is generated through nodes and sent to the virtual machine. When the virtual machine meets the conditions, it generates a target copy and sends it to the cloud operation and maintenance center. The operation and maintenance center performs visual display to avoid analyzing the entire network topology and is based only on the information of the target failed node and its associated nodes.

Benefits of technology

It realizes rapid and efficient acquisition and display of the association relationship between failures, improves the safety and reliability of the disaster recovery system, and reduces the occurrence and impact of failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398204B_ABST
    Figure CN114398204B_ABST
Patent Text Reader

Abstract

The present invention provides a disaster recovery visualization display system and method. Among them, the system includes: a node, which is used to generate real-time snapshot data of the current failure based on the current failure of the node when a failure occurs, and send a backup request for the real-time snapshot data to the virtual machine; the virtual machine is used to receive the backup request for the real-time snapshot data, accept the backup request for the real-time snapshot data when its own backup data does not reach the backup upper limit, and when the real-time snapshot data meets a predetermined first target condition, generate a target copy according to the real-time snapshot data and send the target copy to the cloud operation and maintenance center; the cloud operation and maintenance center is used to obtain the associated failure data of all failed nodes based on the backup information in the target copy, and issue the backup information based on the request of the operation and maintenance node for the associated failure data, and the operation and maintenance node or the cloud operation and maintenance center performs visual display. The disaster recovery visualization display system and method provided by the present invention can display the association relationship of failures faster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer networks, and in particular, to a disaster recovery visualization display system and method. Background Art

[0002] With the continuous expansion of the scale of business systems, the requirements for the performance, stability, and security of IT devices supporting the normal operation of business systems are getting higher and higher, and the network topology is becoming more and more complex. Currently, for the deployment of medium and large-scale business systems, a working center and two disaster recovery centers (two-site three-center or three-site three-center) are usually set up to achieve a perfect disaster tolerance system support. The network topology of the above business systems usually adopts a two-dimensional presentation method. As the scale increases (such as the continuous increase of functions and nodes), the network topology will become very complex. After some nodes in the business system fail, the business system performs disaster recovery. For the complex network topology in the prior art, it is difficult to obtain the correlation relationship between faults and display it. Summary of the Invention

[0003] The present invention provides a disaster recovery visualization display system and method to solve the defect of long time consumption for obtaining and displaying the correlation relationship between faults in the prior art, and to achieve fast and efficient display of fault information.

[0004] The present invention provides a disaster recovery visualization display system, including: a plurality of nodes, a plurality of virtual machines, and a cloud operation and maintenance center;

[0005] The node is used to generate real-time snapshot data of the current fault based on the current fault of the node when a fault occurs, and send a backup request for the real-time snapshot data to any virtual machine;

[0006] The virtual machine is used to receive the backup request for the real-time snapshot data, accept the backup request for the real-time snapshot data when its own backup data does not reach the backup limit, and when the real-time snapshot data meets a predetermined first target condition, generate a target copy according to the real-time snapshot data and send the target copy to the cloud operation and maintenance center;

[0007] The cloud operation and maintenance center is used to obtain the associated fault data of all the nodes that have failed based on the backup information in the target copy, and issue the backup information based on the display request of the operation and maintenance node for the associated fault data, and the operation and maintenance node or the cloud operation and maintenance center performs visual display.

[0008] A disaster recovery visualization display system provided by the present invention, the virtual machine is further configured to discard the real-time snapshot data when the real-time snapshot data does not meet the first target condition or meets the first update period of the virtual machine, or when the second target condition is met, the virtual machine sends the real-time snapshot data to a second virtual machine.

[0009] A disaster recovery visualization display system provided by the present invention, the virtual machine is further configured to record a backup log, the backup log includes node information for sending the real-time snapshot data, information of the second virtual machine that receives the real-time snapshot data of the virtual machine, and backup log update time, and sends the backup log to the cloud operation and maintenance center according to a third period.

[0010] A disaster recovery visualization display system provided by the present invention, the backup information includes: node information of a faulty node, fault type, fault occurrence time, and node information of an associated faulty node of the faulty node.

[0011] A disaster recovery visualization display system provided by the present invention, when there is a faulty associated node for a faulty node, the faulty node is further configured to send association information to the faulty associated node;

[0012] The faulty associated node is configured to record the association information, and when the faulty associated node fails, snapshot the association information into its own real-time snapshot data.

[0013] The present invention also provides a disaster recovery visualization display method, including:

[0014] When a node fails, generate real-time snapshot data of the current fault based on the current fault of the node, and send a backup request for the real-time snapshot data to any virtual machine;

[0015] The virtual machine receives the backup request for the real-time snapshot data, accepts the backup request for the real-time snapshot data when its own backup data does not reach the backup upper limit, and when the real-time snapshot data meets a predetermined first target condition, generate a target copy according to the real-time snapshot data, and send the target copy to the cloud operation and maintenance center;

[0016] The cloud operation and maintenance center obtains associated fault data of all faulty nodes based on the backup information in the target copy, and issues the backup information based on a display request of an operation and maintenance node for the associated fault data, and the operation and maintenance node or the cloud operation and maintenance center performs visualization display.

[0017] A disaster recovery visualization display method provided by the present invention, further including:

[0018] When the real-time snapshot data does not meet the first target condition or meets the first update period of the virtual machine, the virtual machine discards the real-time snapshot data, or when the second target condition is met, the virtual machine sends the real-time snapshot data to a second virtual machine.

[0019] A disaster recovery visualization display method provided by the present invention further includes:

[0020] The virtual machine is further configured to record a backup log, the backup log includes node information for sending the real-time snapshot data, information of the second virtual machine receiving the real-time snapshot data of the virtual machine, and backup log update time, and sends the backup log to the cloud operation and maintenance center according to a third period.

[0021] A disaster recovery visualization display method provided by the present invention further includes:

[0022] When there is a fault-associated node for the node with a fault, the node with the fault is further configured to send association information to the fault-associated node;

[0023] The fault-associated node is configured to record the association information, and when the fault-associated node fails, snapshot the association information into its own real-time snapshot data.

[0024] A disaster recovery visualization display method provided by the present invention, the cloud operation and maintenance center obtains associated fault data of all the nodes with faults based on the backup information in the target copy, specifically including:

[0025] The cloud operation and maintenance center performs association analysis on the backup information in the target copy or the experience data of each operation and maintenance node based on deep learning to determine the association relationship between faults.

[0026] The disaster recovery visualization display system and method provided by the present invention are such that when a node fails, based on the current failure of the node, real-time snapshot data of the current failure is generated, and a backup request for the real-time snapshot data is sent to any virtual machine. The virtual machine accepts the backup request and temporarily stores the real-time snapshot data. When the real-time snapshot data meets a predetermined first target condition, a target copy is generated according to the real-time snapshot data, and the target copy is sent to the cloud operation and maintenance center. The cloud operation and maintenance center obtains the associated failure data of all failed nodes based on the backup information in the target copy, and issues the backup information based on the display request of the operation and maintenance node for the associated failure data. The operation and maintenance node or the cloud operation and maintenance center performs visual display. A virtual machine designed for simply backing up the failures of nodes is used. When the cloud operation and maintenance center obtains the associated failure data of the failed nodes, it does not need to be based on the network topology structure of the entire business system, avoiding the analysis of intricate network connection relationships, but only needs to be based on the information of the target failed node and its associated nodes, and can obtain the associated relationship between failures faster and perform display. In a complex network system, it can quickly determine whether a failure will cause the occurrence of the next possible failure, so as to more quickly and efficiently perform visual display of the failures of the disaster recovery system. Moreover, both the virtual machine and the cloud operation and maintenance center can be backed up, with dual backup, which can improve the security and reliability of the business system. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0028] Figure 1 is a schematic structural diagram of the disaster recovery visualization display system provided by the present invention;

[0029] Figure 2 is a schematic flowchart of the disaster recovery visualization display method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0031] In the description of the embodiments of the present invention, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be construed as indicating or implying relative importance, nor are they related to order.

[0032] In the description of the embodiments of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific situations.

[0033] The following combines Figures 1 to 2 to describe the disaster recovery visualization display system and method provided by the present invention.

[0034] Figure 1 is a schematic structural diagram of the disaster recovery visualization display system provided by the present invention. The following combines Figure 1 to describe the disaster recovery visualization display system of the embodiments of the present invention. As Figure 1 shown, the system includes: several nodes 101, several virtual machines 102, and a cloud operation and maintenance center 103.

[0035] Specifically, the disaster recovery visualization display system mainly includes several nodes 101, several virtual machines 102, and a cloud operation and maintenance center 103. Several means one or more. Preferably, the disaster recovery visualization display system can include multiple nodes 101, multiple virtual machines 102, and a cloud operation and maintenance center 103.

[0036] Any node 101 can be a device, application subsystem, or server that has a local or global impact on the entire business system. Node 101 is a network node, and a network node is a node device, application subsystem, or server that has a local or global impact on the entire system.

[0037] Node 101 can be connected to at least one other node 101, or it can also not be connected to other nodes 101.

[0038] Each node 101 can correspond to several virtual machines 102. Preferably, each node 101 can correspond to multiple virtual machines 102. Each node 101 is communicatively connected to its corresponding virtual machines 102.

[0039] Each node 101 is relatively independent. Each virtual machine 102 is equal to each other. The virtual machines 102 corresponding to each node 101 can be flexibly allocated: any virtual machine 102 corresponding to this node 101 can be deployed locally, that is, it can be in the same area as this node 101 (such as the same city or the same industrial park, etc.); it can also be deployed non-locally, that is, it can be not in the same area as this node 101.

[0040] The virtual machines 102 can be dispatched to each node 101, that is, there is no one-to-one correspondence between the node 101 and the virtual machine 102, so that the number of virtual machines 102 can be reduced, a lightweight disaster recovery recording and display system can be implemented, and it is convenient for subsequent overall fault display.

[0041] The cloud operation and maintenance center 103, which can be an operation and maintenance center located in the cloud, can perform operation and maintenance management of the business system based on cloud computing technology. Each virtual machine 102 is communicatively connected to the cloud operation and maintenance center 103.

[0042] The node 101 is used to generate real-time snapshot data of the current fault based on the current fault of the node 101 when a fault occurs, and send a backup request for the real-time snapshot data to any virtual machine 102.

[0043] Specifically, for any node 101, when this node 101 has a fault, this node 101 can generate real-time snapshot data of this fault based on the currently occurring fault (that is, the current fault).

[0044] After generating the real-time snapshot data, this node 101 can send a backup request for the real-time snapshot data to any virtual machine 102, requesting to send this real-time snapshot data to this virtual machine 102.

[0045] Optionally, after generating the real-time snapshot data, this node 101 can send a backup request for the real-time snapshot data to any virtual machine 102 corresponding to this node 101.

[0046] Optionally, the virtual machines 102 corresponding to the node 101 can be saved in a list based on a certain order. After the node 101 generates a target copy, it can send a backup request to the first virtual machine 102 in this list, and this backup request can carry information of this node 101; after the first virtual machine 102 receives this backup request, it queries its own backup identifier, and the backup identifier is used to indicate whether it can perform a backup action; if it can, then perform a backup action, record backup information, save this real-time snapshot data, and can also transfer this real-time snapshot data to other virtual machines 102 corresponding to this node 101; if it cannot, it can return a rejection message to this node 101, and this rejection message can be used to point to the next virtual machine 102 in this list, so that this node 101 can send a backup request to the next virtual machine 102.

[0047] The virtual machine 102 is configured to receive a backup request for real-time snapshot data. When its own backup data does not reach the backup limit, it accepts the backup request for real-time snapshot data, and when the real-time snapshot data meets a predetermined first target condition, it generates a target copy according to the real-time snapshot data and sends the target copy to the cloud operation and maintenance center 103.

[0048] Specifically, each virtual machine 102 that receives the backup request for the real-time snapshot data can determine whether the backup data stored by itself reaches the backup limit.

[0049] In the case where the backup limit is not reached, the virtual machine 102 can accept the backup request, perform a backup operation, and record backup information. Specifically, after the virtual machine 102 accepts the backup request, the node 101 sends the real-time snapshot data to the virtual machine 102; the virtual machine 102 can receive the real-time snapshot data and save it.

[0050] After the virtual machine 102 receives the real-time snapshot data, it can determine whether the real-time snapshot data meets a predetermined first target condition.

[0051] The first target condition may include that at least one of the fault type, fault level, frequency of fault occurrence, and association status with other nodes of the target fault node meets a preset condition.

[0052] The target fault node refers to the node that generates the real-time snapshot data.

[0053] Optionally, that the fault type of the target fault node meets the preset condition may mean that the fault type of the target fault node is a target type. The target type may include power failure and serious database error, etc.

[0054] Optionally, that the fault level of the target fault node meets the preset condition may mean that the fault level of the target fault node reaches a preset level. Exemplarily, in the case where the severity of the fault level ranges from low to high as 1 - 5, if the preset level is 3, then when the fault level of the target fault node is 3 - 5, the fault level of the target fault node meets the preset condition, and when the fault level of the target fault node is 1 - 2, the fault level of the target fault node does not meet the preset condition.

[0055] Optionally, that the frequency of fault occurrence of the target fault node meets the preset condition may mean that the frequency of fault occurrence of the target fault node is greater than or equal to a preset frequency threshold. The specific value of the frequency threshold is not specifically limited in the embodiments of the present invention.

[0056] Optionally, the association status between the target faulty node and other nodes may refer to whether there is an association between the target faulty node and other nodes and whether they are directly connected, etc.

[0057] When the real-time snapshot data meets the first target condition, the virtual machine 102 may generate a target copy according to the real-time snapshot data.

[0058] The target copy may carry backup information.

[0059] After generating the target copy, the virtual machine 102 may send the target copy to the cloud operation and maintenance center 103.

[0060] The cloud operation and maintenance center 103 is used to obtain the associated fault data of all faulty nodes 101 based on the backup information in the target copy, and issue the backup information based on the display request of the operation and maintenance node for the associated fault data, and the operation and maintenance node or the cloud operation and maintenance center 103 performs visual display.

[0061] Specifically, the cloud operation and maintenance center 103 may receive the target copies sent by each virtual machine 102.

[0062] The cloud operation and maintenance center 103 may obtain the associated fault data of each faulty node 101 based on the backup information in each target copy.

[0063] Specifically, the cloud operation and maintenance center 103 may obtain the associated fault data of each faulty node 101 through empirical data or an association analysis model obtained through self-learning (such as a neural network model or a knowledge graph, etc.). Obtaining the associated fault data of the faulty node 101 through empirical data or a model obtained through self-learning may be implemented based on a written script.

[0064] The operation and maintenance node may send a display request to the cloud operation and maintenance center 103. The operation and maintenance node is used to perform operation and maintenance management of the business system. The display request may carry information of several nodes.

[0065] Optionally, the global framework may be constructed by dragging, that is, the display request carries information of nodes, and the icons of the nodes to be displayed can be dragged to the corresponding positions by dragging. That is, generating a display request includes the operation of dragging the icons of the nodes to be displayed to the corresponding positions.

[0066] Optionally, generating a display request may include the operation of clicking on the icon of the node to be displayed.

[0067] After the cloud operation and maintenance center 103 receives the display request of the operation and maintenance node, it can send the backup information corresponding to the display request to the operation and maintenance node; the operation and maintenance node can visually display the fault conditions of the above-mentioned several nodes (including the fault data of the node itself and the associated fault data) based on the backup information, that is, perform local display or single-point display.

[0068] The cloud operation and maintenance center 103 can also visually display the fault conditions of the above-mentioned several nodes, that is, perform overall display or centralized display.

[0069] Exemplarily, the fault nodes and other nodes that have the above-mentioned associated faults can be highlighted in the electronic map to distinguish the fault nodes and other nodes that have the above-mentioned associated faults from the other nodes.

[0070] Exemplarily, it is also possible to only display the information of the fault nodes and other nodes that have the above-mentioned associated faults, and identify the root fault node (i.e., the target fault node) and the associated fault nodes (i.e., other nodes that have the above-mentioned associated faults).

[0071] The operation and maintenance nodes (each subsystem or single node) summarize the collected relevant fault data and can display it in real time. The overall fault data is summarized and fed back to the display desktop in real time by the cloud operation and maintenance center 103 in the cloud for each basic level processing process and result.

[0072] It can be understood that the virtual machine 102 stores the real-time snapshot data for a period of time (i.e., temporarily stores the real-time snapshot data) to facilitate the rollback of the real-time snapshot data, and sends the important real-time snapshot data to the cloud operation and maintenance center 103 for storage according to the importance of the real-time snapshot data, which is convenient for the security and durability of the real-time snapshot data storage.

[0073] Traditional disaster recovery centers and working centers are generally set in the same computer room. While avoiding the impact of local fault accidents on faults and data loss, the process of restoring data in the disaster recovery center will affect the normal operation of the working center to a certain extent, resulting in more frequent generation of faults, and the faults are still limited within a single system, which is not conducive to the global control of faults. In the embodiments of the present invention, data recovery is performed by mounting the implementation snapshot data, which will not affect the normal operation of the working center, thereby reducing the generation of faults and being conducive to the global control of faults.

[0074] In an embodiment of the present invention, when a node fails, real-time snapshot data of the current failure is generated based on the current failure of the node, and a backup request for the real-time snapshot data is sent to any virtual machine. The virtual machine accepts the backup request and temporarily stores the real-time snapshot data. When the real-time snapshot data meets a predetermined first target condition, a target copy is generated according to the real-time snapshot data, and the target copy is sent to the cloud operation and maintenance center. The cloud operation and maintenance center obtains associated failure data of all failed nodes based on the backup information in the target copy, and issues the backup information based on the display request of the operation and maintenance node for the associated failure data. The operation and maintenance node or the cloud operation and maintenance center performs visual display. A virtual machine that simply performs failure backup for node failures is designed. When the cloud operation and maintenance center obtains the associated failure data of the failed node, it does not need to be based on the network topology of the entire business system, avoiding the analysis of intricate network connection relationships, but only needs to be based on the information of the target failed node and its associated nodes, and can obtain the association relationship between failures faster and perform display. In a complex network system, it can quickly determine whether a failure will cause the occurrence of the next possible failure, so as to more quickly and efficiently perform visual display of the failures of the disaster recovery system. Moreover, both the virtual machine and the cloud operation and maintenance center can be backed up, with dual backup, which can improve the security and reliability of the business system.

[0075] Based on the content of any of the above embodiments, the virtual machine 102 is further configured to discard the real-time snapshot data when the real-time snapshot data does not meet the first target condition or meets the first update period of the virtual machine 102, or when the second target condition is met, the virtual machine 102 sends the real-time snapshot data to the second virtual machine.

[0076] Specifically, after receiving the real-time snapshot data, the virtual machine 102 can determine whether the real-time snapshot data meets a predetermined first target condition.

[0077] In the case where the real-time snapshot data does not meet the first target condition, the virtual machine 102 can discard the real-time snapshot data, that is, no longer store the real-time snapshot data. Not meeting the first target condition indicates that the failure (or the real-time snapshot data) is unimportant or there is no associated failed node, and it is meaningless to continue storing the real-time snapshot data.

[0078] The virtual machine 102 can also periodically discard the stored real-time snapshot data, that is, when the storage duration of the real-time snapshot data reaches the first update period, it indicates that the real-time snapshot data is too old, and the virtual machine 102 can discard the real-time snapshot data.

[0079] The first update period can be preset according to actual needs. The specific duration of the first update period is not specifically limited in the embodiments of the present invention.

[0080] When the second target condition is met, the virtual machine 102 can send the real-time snapshot data to the second virtual machine and delete the real-time snapshot data stored by itself, which is convenient for automatically mounting the previous real-time snapshot data (from the virtual machine 102 or the second virtual machine) to the node for dual protection.

[0081] The second virtual machine is another virtual machine different from the virtual machine 102. Therefore, the real-time snapshot data of a certain node 101 stored by the virtual machine 102 itself is a local snapshot, and the real-time snapshot data stored by the second virtual machine is an off-site snapshot.

[0082] It should be noted that both the local snapshot and the off-site snapshot can be sent to the cloud operation and maintenance center 103 and stored in the cloud operation and maintenance center 103, which is convenient for fault data rollback and automatic restart when a fault occurs.

[0083] The second target condition may include meeting the second update period of the virtual machine or the real-time snapshot data meeting the backup upper limit of the current virtual machine.

[0084] The second update period can be preset according to actual needs. The specific duration of the second update period is not specifically limited in the embodiments of the present invention.

[0085] In the embodiments of the present invention, by discarding the real-time snapshot data when it can be discarded, the virtual machine can be lightened, meaningless data occupying storage space can be avoided, the backup burden of the virtual machine can be reduced, and when the real-time snapshot data cannot be discarded, the real-time snapshot data is sent to other virtual machines, which is convenient for the data of the virtual machine to be rolled back and updated in time and avoid being lost without reason.

[0086] Based on the content of any of the above embodiments, the virtual machine 102 is further configured to record a backup log, where the backup log includes node information for sending the real-time snapshot data, information of the second virtual machine receiving the real-time snapshot data of the virtual machine, and the backup log update time, and send the backup log to the cloud operation and maintenance center 103 according to a third period.

[0087] Specifically, the virtual machine 102 can also record a backup log.

[0088] The backup log may include information such as the ID, IP address, and port number of the node sending the real-time snapshot data, information such as the ID of the second virtual machine, and the update time of the backup log.

[0089] For the backup log recorded by the virtual machine 102, the virtual machine can send it to the cloud operation and maintenance center 103 periodically, that is, when the duration since the last sending of the backup log reaches the third period, the backup log is sent to the cloud operation and maintenance center 103. The third period can be preset according to actual needs. The specific duration of the third period is not specifically limited in the embodiments of the present invention.

[0090] In the embodiments of the present invention, backup logs are recorded through virtual machines, making the nodes relatively independent. As a result, when the cloud operation and maintenance center determines the associated faults of a faulty node, it does not need to be based on the network topology of the entire business system, avoiding the analysis of intricate network connection relationships. Instead, it only needs to be based on the information of the faulty node and its associated nodes, and can obtain the associated relationships between faults faster and display them, enabling a faster and more efficient visual display of the faults in the disaster recovery system.

[0091] Based on the content of any of the above embodiments, the backup information includes: the node information of the faulty node, the fault type, the fault occurrence time, and the basic information of the node of the faulty node.

[0092] Specifically, the node information of a node may include the identification information (ID), IP address, port number, etc. of the node.

[0093] The associated nodes of a faulty node refer to other nodes that are associated with this node. Usually, the set composed of other nodes affected by this faulty node is a subset of the set composed of the associated nodes of the faulty node.

[0094] The backup information of the faulty node can be used as the input of the aforementioned association analysis model to obtain the associated fault data of the faulty node output by the association analysis model.

[0095] In the embodiments of the present invention, through the node information of the faulty node, the fault type, the fault occurrence time, and the node information of the associated nodes of the target faulty node, the associated relationships between faults can be obtained more accurately and quickly. Moreover, it is convenient for big data correlation analysis to obtain the relationships between faults and for subsequent disaster recovery display.

[0096] Based on the content of any of the above embodiments, when there are associated fault nodes for the faulty node 101 that has a fault, the faulty node 101 is also used to send association information to the associated fault nodes.

[0097] Specifically, when there are associated fault nodes for the faulty node 101 (i.e., the faulty node) (i.e., other faulty nodes, and the faults that occur to other faulty nodes are associated with the fault that occurs to this node 101), this node 101 can send association information to its own associated fault nodes.

[0098] The association information is used to describe the association relationship between the fault that occurs to the associated fault node and the fault that occurs to this node 101.

[0099] The associated fault nodes are used to record the association information and, when a fault occurs to the associated fault nodes, snapshot the association information into their own real-time snapshot data.

[0100] Specifically, the fault-associated node can receive the above-mentioned association information and record it. When a fault occurs in the fault-associated node, relevant information and association information about the current fault can be snapshotted into its own real-time snapshot data.

[0101] In the embodiment of the present invention, the node with a fault sends association information to the fault-associated node, the fault-associated node records the association information, and when a fault occurs in the fault-associated node, the association information is snapshotted into its own real-time snapshot data, which facilitates the timely recording of the association information and is conducive to the cloud operation and maintenance center and each operation and maintenance node to read the association information. Moreover, the fault information of a single node at the basic level is sent to each relevant subsystem or associated node through various channels, so that the cloud operation and maintenance center can calculate the relationship between node faults based on the fault information of each node, and can obtain the association relationship between faults faster and more conveniently.

[0102] Figure 2 It is a schematic flowchart of the disaster recovery visualization display method provided by the present invention. The following combines Figure 2 to describe the disaster recovery visualization display method of the embodiment of the present invention. As Figure 2 shown, the method includes: Step 201, Step 202, and Step 203.

[0103] Specifically, the disaster recovery visualization display method provided by the embodiment of the present invention can be implemented based on the disaster recovery visualization display system provided by any of the above-mentioned disaster recovery visualization display system embodiments. That is, the execution subject of the disaster recovery visualization display method provided by the embodiment of the present invention can be any of the above-mentioned disaster recovery visualization display systems.

[0104] Step 201, when a fault occurs in the node, generate real-time snapshot data of the current fault based on the current fault of the node, and send a backup request for the real-time snapshot data to any virtual machine.

[0105] Specifically, for any node 101, when a fault occurs in the node 101, the node 101 can generate real-time snapshot data of the fault based on the currently occurring fault (i.e., the current fault).

[0106] After generating the real-time snapshot data, the node 101 can send a backup request for the real-time snapshot data to any virtual machine 102, requesting to send the real-time snapshot data to the virtual machine 102.

[0107] Step 202, the virtual machine receives the backup request for the real-time snapshot data, when its own backup data does not reach the backup limit, accepts the backup request for the real-time snapshot data, and when the real-time snapshot data meets a predetermined first target condition, generates a target copy according to the real-time snapshot data, and sends the target copy to the cloud operation and maintenance center.

[0108] Specifically, each virtual machine 102 that receives a backup request for the real-time snapshot data can determine whether the backup data stored locally reaches the backup limit.

[0109] In the case where the backup limit is not reached, the virtual machine 102 can accept the backup request, perform a backup operation, and record backup information. Specifically, after the virtual machine 102 accepts the backup request, the node 101 sends the real-time snapshot data to the virtual machine 102; the virtual machine 102 can receive and save the real-time snapshot data.

[0110] After receiving the real-time snapshot data, the virtual machine 102 can determine whether the real-time snapshot data meets a predetermined first target condition.

[0111] The first target condition can include that at least one of the failure type, failure level, frequency of failure occurrence, and association status with other nodes of the target failure node meets a preset condition.

[0112] The target failure node refers to the node that generates the real-time snapshot data.

[0113] Optionally, that the failure type of the target failure node meets the preset condition can mean that the failure type of the target failure node is the target type. The target type can include power-off and serious database errors, etc.

[0114] Optionally, that the failure level of the target failure node meets the preset condition can mean that the failure level of the target failure node reaches a preset level. Exemplarily, if the severity of the failure level ranges from 1 to 5 in ascending order, and the preset level is 3, then when the failure level of the target failure node is 3 - 5, the failure level of the target failure node meets the preset condition, and when the failure level of the target failure node is 1 - 2, the failure level of the target failure node does not meet the preset condition.

[0115] Optionally, that the frequency of failure occurrence of the target failure node meets the preset condition can mean that the frequency of failure occurrence of the target failure node is greater than or equal to a preset frequency threshold. The specific value of the frequency threshold is not specifically limited in the embodiments of the present invention.

[0116] Optionally, the association status between the target failure node and other nodes can refer to whether there is an association between the target failure node and other nodes and whether they are directly connected, etc.

[0117] When the real-time snapshot data meets the first target condition, the virtual machine 102 can generate a target copy based on the real-time snapshot data.

[0118] The target copy can carry backup information.

[0119] After generating the target copy, the virtual machine 102 may send the target copy to the cloud operation and maintenance center 103.

[0120] Step 203: The cloud operation and maintenance center obtains the associated fault data of all the failed nodes based on the backup information in the target copy, and issues the backup information based on the display request of the operation and maintenance node for the associated fault data. The operation and maintenance node or the cloud operation and maintenance center performs visual display.

[0121] Specifically, the cloud operation and maintenance center 103 may receive the target copies sent by each virtual machine 102.

[0122] The cloud operation and maintenance center 103 may obtain the associated fault data of each failed node 101 based on the backup information in each target copy.

[0123] Specifically, the cloud operation and maintenance center 103 may obtain the associated fault data of each failed node 101 through empirical data or an association analysis model obtained through self-learning (such as a neural network model or a knowledge graph, etc.). Obtaining the associated fault data of the failed node 101 through empirical data or a model obtained through self-learning can be implemented based on a written script.

[0124] The operation and maintenance node may send a display request to the cloud operation and maintenance center 103. The operation and maintenance node is used for the operation and maintenance management of the business system. The display request may carry the information of several nodes.

[0125] After receiving the display request from the operation and maintenance node, the cloud operation and maintenance center 103 may issue the backup information corresponding to the display request to the operation and maintenance node; the operation and maintenance node may, based on the backup information, perform visual display on the fault conditions of the above-mentioned several nodes (including the fault data of the nodes themselves and the associated fault data), that is, perform local display or single-point display.

[0126] In the embodiment of the present invention, when a node fails, real-time snapshot data of the current failure is generated based on the current failure of the node, and a backup request for the real-time snapshot data is sent to any virtual machine. The virtual machine accepts the backup request and temporarily stores the real-time snapshot data. When the real-time snapshot data meets a predetermined first target condition, a target copy is generated according to the real-time snapshot data, and the target copy is sent to the cloud operation and maintenance center. The cloud operation and maintenance center obtains associated failure data of all failed nodes based on the backup information in the target copy, and issues the backup information based on the display request of the operation and maintenance node for the associated failure data. The operation and maintenance node or the cloud operation and maintenance center performs visual display. A virtual machine that simply performs failure backup for node failures is designed. When the cloud operation and maintenance center obtains the associated failure data of the failed node, it does not need to be based on the network topology of the entire business system, avoiding the analysis of intricate network connection relationships, but only needs to be based on the information of the target failed node and its associated nodes, and can obtain the association relationship between failures faster and perform display. In a complex network system, it can quickly determine whether a failure will cause the occurrence of the next possible failure, so that the failures of the disaster recovery system can be visually displayed more quickly and efficiently. Moreover, both the virtual machine and the cloud operation and maintenance center can be backed up, with dual backup, which can improve the security and reliability of the business system.

[0127] Based on the content of any of the above embodiments, the disaster recovery visualization display method further includes: when the real-time snapshot data does not meet the first target condition or meets the first update period of the virtual machine, the virtual machine discards the real-time snapshot data, or when the second target condition is met, the virtual machine sends the real-time snapshot data to the second virtual machine.

[0128] Specifically, after the virtual machine 102 receives the real-time snapshot data, it can determine whether the real-time snapshot data meets a predetermined first target condition.

[0129] In the case where the real-time snapshot data does not meet the first target condition, the virtual machine 102 can discard the real-time snapshot data, that is, no longer store the real-time snapshot data. Not meeting the first target condition indicates that the failure (or the real-time snapshot data) is unimportant or there is no associated failed node, and it is meaningless to continue storing the real-time snapshot data.

[0130] The virtual machine 102 can also periodically discard the stored real-time snapshot data, that is, when the storage duration of the real-time snapshot data reaches the first update period, it indicates that the real-time snapshot data is too old, and the virtual machine 102 can discard the real-time snapshot data.

[0131] The first update period can be preset according to actual needs. The specific duration of the first update period is not specifically limited in the embodiment of the present invention.

[0132] When the second target condition is met, virtual machine 102 can send the real-time snapshot data to the second virtual machine and delete the real-time snapshot data stored in itself, so as to automatically mount the previous real-time snapshot data (from virtual machine 102 or the second virtual machine) to the node for double protection.

[0133] The second virtual machine is another virtual machine different from the virtual machine 102. Therefore, the real-time snapshot data of a certain node 101 stored in the virtual machine 102 itself is a local snapshot, and the real-time snapshot data stored in the second virtual machine is a remote snapshot.

[0134] The second target condition may include satisfying a second update period of the virtual machine or the real-time snapshot data satisfying a backup upper limit of the current virtual machine.

[0135] The second update period can be preset according to actual needs. The specific duration of the second update period is not specifically limited in the embodiment of the present invention.

[0136] The embodiments of the present invention can lightweight the virtual machine, avoid meaningless data occupying storage space, and reduce the backup burden of the virtual machine by discarding the real-time snapshot data when the real-time snapshot data can be discarded, and by sending the real-time snapshot data to other virtual machines when the real-time snapshot data cannot be discarded, it is convenient for the virtual machine data to be rolled back and updated in a timely manner to avoid loss without reason.

[0137] Based on the content of any of the above embodiments, the disaster recovery visualization display method also includes: a virtual machine, which is also used to record a backup log, the backup log contains the node information that sends the real-time snapshot data, the information of the second virtual machine that receives the real-time snapshot data of the virtual machine, and the backup log update time, and sends the backup log to the cloud operation and maintenance center according to the third cycle.

[0138] Specifically, the virtual machine 102 may also record a backup log.

[0139] The backup log may include information such as the ID, IP address, and port number of the node that sends the real-time snapshot data, information such as the ID of the second virtual machine, and the update time of the backup log.

[0140] For the backup log recorded by the virtual machine 102, the virtual machine can periodically send it to the cloud operation and maintenance center 103, that is, when the time from the last sending of the backup log reaches the third period, the backup log is sent to the cloud operation and maintenance center 103. The third period can be preset according to actual needs. The specific duration of the third period is not specifically limited in the embodiment of the present invention.

[0141] In the embodiment of the present invention, the virtual machine records the backup log, making the nodes relatively independent. Therefore, when the cloud operation and maintenance center determines the associated faults of the faulty nodes, it does not need to be based on the network topology of the entire business system, avoiding the analysis of intricate network connection relationships. Instead, it only needs to be based on the information of the faulty nodes and their associated nodes, and can obtain the association relationships between faults faster and display them, enabling a faster and more efficient visual display of the faults in the disaster recovery system.

[0142] Based on the content of any of the above embodiments, the disaster recovery visualization display method further includes: when there are fault-associated nodes for the node with a fault, the node with the fault is also used to send association information to the fault-associated nodes.

[0143] Specifically, when the faulty node 101 (i.e., the node with a fault) has fault-associated nodes (i.e., other nodes with faults, and the faults that occur in the other nodes with faults are associated with the fault that occurs in this node 101), this node 101 can send association information to its own fault-associated nodes.

[0144] The association information is used to describe the association relationship between the fault that occurs in the fault-associated node and the fault that occurs in this node 101.

[0145] The fault-associated node is used to record the association information, and when the fault-associated node has a fault, snapshot the association information into its own real-time snapshot data.

[0146] Specifically, the fault-associated node can receive the above association information and record it. When this fault-associated node has a fault, it can snapshot the relevant information of the current fault and the association information into its own real-time snapshot data.

[0147] In the embodiment of the present invention, the node with a fault sends association information to the fault-associated nodes, the fault-associated nodes record the association information, and when the fault-associated node has a fault, snapshot the association information into its own real-time snapshot data, facilitating the timely recording of the association information and being beneficial for the cloud operation and maintenance center and each operation and maintenance node to read the association information. Moreover, by sending the fault information of a single node at the basic level to each related subsystem or associated node through various channels, the cloud operation and maintenance center can calculate the relationship between node faults based on the fault information of each node, and can obtain the association relationships between faults faster and more conveniently.

[0148] Based on the content of any of the above embodiments, the cloud operation and maintenance center obtains the associated fault data of all nodes with faults based on the backup information in the target copy, specifically including: the cloud operation and maintenance center performs association analysis on the backup information in the target copy or the experience data of each operation and maintenance node based on deep learning to determine the association relationships between faults.

[0149] Specifically, the cloud operation and maintenance center can pre - learn based on big data technology according to the sample backup information to obtain a trained correlation analysis model.

[0150] Optionally, methods such as deep learning can be used for self - learning.

[0151] Optionally, the backup information in each target copy can be input into the correlation analysis model, and the correlation analysis model can perform correlation analysis on the backup information in each target copy to obtain the correlation relationship between the faults output by the correlation analysis model.

[0152] Optionally, the experience data of each operation and maintenance node can be input into the correlation analysis model, and the correlation analysis model can perform correlation analysis on the experience data of each operation and maintenance node to obtain the correlation relationship between the faults output by the correlation analysis model.

[0153] By performing correlation analysis on the backup information in the target copy or the experience data of each operation and maintenance node based on deep learning, the embodiments of the present invention can obtain correlation fault data and can more quickly and accurately determine the correlation relationship between fault nodes.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A disaster recovery visualization display system, characterized in that, Including: A number of nodes, a number of virtual machines, and a cloud operation and maintenance center; The node is used to generate real-time snapshot data of the current fault based on the current fault of the node when a fault occurs, and send a backup request for the real-time snapshot data to any virtual machine; The virtual machine is used to receive the backup request for the real-time snapshot data. When its own backup data does not reach the backup limit, it accepts the backup request for the real-time snapshot data, and when the real-time snapshot data meets a predetermined first target condition, it generates a target copy according to the real-time snapshot data and sends the target copy to the cloud operation and maintenance center; The cloud operation and maintenance center is used to obtain associated fault data of all the nodes that have failed based on the backup information in the target copy, and issue the backup information based on the display request of the operation and maintenance node for the associated fault data, and the operation and maintenance node or the cloud operation and maintenance center performs visual display; When there is a fault-associated node for the node that has failed, the node that has failed is further used to send association information to the fault-associated node; the fault-associated node is used to record the association information, and when the fault-associated node fails, snapshot the association information into its own real-time snapshot data.

2. The disaster recovery visualization display system according to claim 1, wherein, The virtual machine is further used to discard the real-time snapshot data when the real-time snapshot data does not meet the first target condition or meets the first update period of the virtual machine, or when the second target condition is met, the virtual machine sends the real-time snapshot data to a second virtual machine.

3. The disaster recovery visualization display system according to claim 2, wherein The virtual machine is further used to record a backup log, which includes the node information of the node that sent the real-time snapshot data, the information of the second virtual machine that received the real-time snapshot data of the virtual machine, and the backup log update time, and send the backup log to the cloud operation and maintenance center according to a third period.

4. The disaster recovery visualization display system according to any one of claims 1 to 3, characterized in that The backup information includes: the node information of the faulty node, the fault type, the fault occurrence time, and the node information of the fault-associated node of the faulty node.

5. A disaster recovery visualization display method, characterized in that, Including: When a fault occurs, the node generates real-time snapshot data of the current fault based on the current fault of the node, and sends a backup request for the real-time snapshot data to any virtual machine; The virtual machine receives the backup request for the real-time snapshot data. When its own backup data does not reach the backup limit, it accepts the backup request for the real-time snapshot data, and when the real-time snapshot data meets a predetermined first target condition, it generates a target copy according to the real-time snapshot data and sends the target copy to the cloud operation and maintenance center; The cloud operation and maintenance center obtains associated fault data of all the nodes that have failed based on the backup information in the target copy, and issues the backup information based on the display request of the operation and maintenance node for the associated fault data, and the operation and maintenance node or the cloud operation and maintenance center performs visual display; When there is a fault-associated node for the node that has failed, the node that has failed is further used to send association information to the fault-associated node; The fault-associated node is used to record the association information, and when the fault-associated node fails, snapshot the association information into its own real-time snapshot data.

6. The disaster recovery visualization display method according to claim 5, wherein Also including: When the real-time snapshot data does not meet the first target condition or meets the first update period of the virtual machine, the virtual machine discards the real-time snapshot data, or when the second target condition is met, the virtual machine sends the real-time snapshot data to a second virtual machine.

7. The disaster recovery visualization display method according to claim 6, wherein, Further included: The virtual machine is further configured to record a backup log, where the backup log includes node information for sending the real-time snapshot data, information about the second virtual machine that receives the real-time snapshot data of the virtual machine, and the backup log update time, and sends the backup log to the cloud operation and maintenance center according to a third period.

8. The disaster recovery visualization display method according to any one of claims 5 to 7, characterized in that Based on the backup information in the target copy, the cloud operation and maintenance center obtains the associated failure data of all the nodes that have failed, specifically including: Based on deep learning, the cloud operation and maintenance center performs an association analysis on the backup information in the target copy or the experience data of each operation and maintenance node to determine the association relationship between the failures.

Citation Information

Patent Citations

  • Intelligent voice alarm system and method

    CN107045469A

  • A method for processing cloud disaster recovery data

    CN109408289A