A distributed monitoring system and method

By using a distributed monitoring system to achieve data synchronization and alarm notifications among nodes, the paralysis problem caused by the failure of the central node in a centralized monitoring system is solved, and the system's disaster recovery capability and scalability are improved.

CN114449227BActive Publication Date: 2025-12-19AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210232192.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-12-19
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

If the central node of a centralized monitoring system fails, the entire monitoring system will be paralyzed and unable to meet disaster recovery requirements.

Method used

A distributed monitoring system is adopted, in which data is synchronized and alarms are triggered between monitoring nodes, forming a decentralized monitoring architecture. The first node, intermediate nodes and the tail node back each other up, ensuring data flow between nodes and realizing disaster recovery function.

Benefits of technology

Even if a monitoring node fails, the system can still operate normally, demonstrating good disaster recovery and scalability, thus reducing the risk of monitoring system failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114449227B_ABST
    Figure CN114449227B_ABST
Patent Text Reader

Abstract

The application discloses a distributed monitoring system and method. The system comprises: a plurality of monitoring nodes; a console of monitoring software is arranged on each of the plurality of monitoring nodes; the plurality of monitoring nodes comprise a head node, a tail node and intermediate nodes; the head node is used for synchronizing monitoring data collected by the monitoring software of the head node to the console of the intermediate node when the console of the head node fails; the intermediate node is used for synchronizing monitoring data collected by the monitoring software of the intermediate node to the console of the tail node when the console of the intermediate node fails; the tail node is used for synchronizing monitoring data collected by the monitoring software of the tail node to the console of the head node when the console of the tail node fails; and any node of the plurality of monitoring nodes is further used for giving an alarm prompt when the console of the any node determines that the monitoring data is abnormal. It can be seen that the distributed monitoring mode can meet the disaster recovery requirement of the monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a distributed monitoring system and method. BACKGROUND

[0002] For business needs, enterprise computer rooms can be deployed in multiple cities such as Shenzhen, Beijing, Shanghai, etc. For these distributed computer rooms and data centers, unified monitoring and management is needed to make the operation more efficient and more refined.

[0003] The common monitoring system at present is centralized, that is, there is only one unified central node in the distributed enterprise computer rooms. The centralized multi-node monitoring is a common solution for monitoring multiple nodes. The main console of the monitoring software is installed at the headquarters (such as the central node). Then, the remote node of the monitoring software is installed on only one system at each remote node. The remote node calls the main console of the central node using a secure communication protocol. The remote node performs monitoring and only sends the results to the main console of the central node, which performs alarm.

[0004] The existing centralized monitoring solution monitors multiple nodes. Once the central node is damaged, the monitoring of the entire multi-center system will be in a paralyzed state, which is difficult to meet the disaster recovery requirements. SUMMARY

[0005] In order to solve the above technical problems, the present application provides a distributed monitoring system and method, which can meet the disaster recovery requirements of the monitoring system.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] In a first aspect, the present application provides a distributed monitoring system, comprising a plurality of monitoring nodes; each of the plurality of monitoring nodes is deployed with a console of monitoring software; the plurality of monitoring nodes comprise a head node, a tail node and intermediate nodes;

[0008] The head node is configured to synchronize monitoring data collected by the monitoring software of the head node to the console of the intermediate node when the console of the head node fails;

[0009] The intermediate node is configured to synchronize monitoring data collected by the monitoring software of the intermediate node to the console of the tail node when the console of the intermediate node fails;

[0010] The tail node is configured to synchronize monitoring data collected by the monitoring software of the tail node to the console of the head node when the console of the tail node fails;

[0011] Any node of the plurality of monitoring nodes is further configured to perform an alarm prompt when a console of the any node determines that the monitoring data is abnormal.

[0012] In some possible implementation manners, the first node is further configured to synchronize, to a console of the tail node, monitoring data collected by monitoring software of the first node when a console of the first node fails.

[0013] In some possible implementation manners, the intermediate node is configured to synchronize, to a console of the first node, monitoring data collected by monitoring software of the intermediate node when a console of the intermediate node fails.

[0014] In some possible implementation manners, the tail node is configured to synchronize, to a console of the intermediate node, monitoring data collected by monitoring software of the tail node when a console of the tail node fails.

[0015] In some possible implementation manners, any node of the plurality of monitoring nodes is further configured to present, by a console, monitoring data collected by monitoring software of at least one node.

[0016] In a second aspect, the present application provides a distributed monitoring method applied to a plurality of monitoring nodes, each of the plurality of monitoring nodes being deployed with a console of monitoring software, the plurality of monitoring nodes including a first node, a tail node and an intermediate node; the method comprising:

[0017] synchronizing, by the first node, to a console of the intermediate node, monitoring data collected by monitoring software of the first node when a console of the first node fails;

[0018] synchronizing, by the intermediate node, to a console of the tail node, monitoring data collected by monitoring software of the intermediate node when a console of the intermediate node fails;

[0019] synchronizing, by the tail node, to a console of the first node, monitoring data collected by monitoring software of the tail node when a console of the tail node fails.

[0020] performing an alarm prompt when a console of any node of the plurality of monitoring nodes determines that the monitoring data is abnormal.

[0021] In some possible implementation manners, the method further comprises:

[0022] synchronizing, by the first node, to a console of the tail node, monitoring data of monitoring software of the first node when a console of the first node fails.

[0023] In some possible implementation manners, the method further comprises:

[0024] When the console of the intermediate node fails, the intermediate node synchronizes monitoring data collected by the monitoring software of the intermediate node to the console of the head node.

[0025] In some possible implementation manners, the method further includes:

[0026] When the console of the tail node fails, the tail node synchronizes monitoring data collected by the monitoring software of the tail node to the console of the intermediate node.

[0027] In some possible implementation manners, the method further includes:

[0028] Any node of the plurality of monitoring nodes presents monitoring data collected by the monitoring software of at least one node through a console.

[0029] The technical scheme provided in the application has the following beneficial effects:

[0030] The application provides a distributed monitoring system, including a plurality of monitoring nodes; each node of the plurality of monitoring nodes is deployed with a console of monitoring software; the plurality of monitoring nodes include a head node, a tail node and an intermediate node; the head node is configured to synchronize monitoring data collected by the monitoring software of the head node to the console of the intermediate node when the console of the head node fails; the intermediate node is configured to synchronize monitoring data collected by the monitoring software of the intermediate node to the console of the tail node when the console of the intermediate node fails; the tail node is configured to synchronize monitoring data collected by the monitoring software of the tail node to the console of the head node when the console of the tail node fails; and any node of the plurality of monitoring nodes is further configured to perform alarm prompting when the console of the any node determines that the monitoring data is abnormal. Through the distributed monitoring manner, when one of the monitoring nodes fails, for example, the console fails, the next node of the monitoring node that fails monitors the part to be monitored by the monitoring node that fails, and it can be seen that the distributed monitoring system has good disaster recovery function. BRIEF DESCRIPTION OF DRAWINGS

[0031] In order to more clearly illustrate the technical scheme in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0032] Figure 1 It is a schematic view of the centralized monitoring system in the prior art;

[0033] Figure 2A schematic diagram of a distributed monitoring system provided by an embodiment of the present application;

[0034] Figure 3 A schematic diagram of another distributed monitoring system provided by an embodiment of the present application;

[0035] Figure 4 A flowchart of a distributed monitoring method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0036] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0037] At present, in order to meet the business needs, enterprises will establish their own data centers. For example, for banks, it is necessary to monitor a plurality of information of financial business, and the monitoring contents include but are not limited to the room environment of information system carrying financial transaction, network connection of information system and other systems, cluster, server, CPU, memory, middleware, database and state of application, etc.

[0038] When the monitored object (such as the above information system) fails, the operation, management and decision personnel need to repair, locate the professional field where the problem may occur, analyze the cause and quickly implement remedial measures through quickly obtaining the monitoring data, and the running of the monitored object at the first time.

[0039] The data center of an enterprise is developing towards multi-site and multi-center deployment. The core, important, general and other data centers are deployed in different geographical locations, and are jointly operated by teams located in different locations. As a result, the monitoring and operation teams also need to be deployed in multiple data centers.

[0040] In order to ensure data security, each data center needs to be monitored. The common monitoring system design at present is centralized, as shown in FIG. 1. Figure 1 The centralized monitoring system only deploys a control console of monitoring software in a center node, and then deploys monitoring software in each remote node. The remote node calls the control console of the center node through a secure communication protocol, so as to send the monitoring data obtained by the remote node to the control console of the center node. Based on the monitoring data, the control console of the center node determines the abnormality, and then performs alarm prompt.

[0041] However, once the central node is damaged, such as the console of the central node fails, the whole centralized monitoring system will be paralyzed and cannot meet the disaster recovery requirements.

[0042] Based on this, the embodiment of the present application provides a distributed monitoring system, as shown in the figure, which is a schematic diagram of the distributed monitoring system provided by the present application. The distributed monitoring system comprises a plurality of monitoring nodes, the consoles of the monitoring software are deployed in the plurality of monitoring nodes, and the plurality of monitoring nodes comprise a head node 210, an intermediate node 220 and a tail node 230. Among them, the intermediate node can be a plurality of nodes or one node. In order to facilitate understanding, the present application takes the intermediate node as one node as an example for introduction, and the principle is similar when the intermediate node is a plurality of nodes. Figure 2

[0043] Among them, the head node is used to synchronize the monitoring data collected by the monitoring software of the head node to the console of the intermediate node when the console of the head node fails. The intermediate node synchronizes the monitoring data collected by the monitoring software of the intermediate node to the console of the tail node when the console of the intermediate node fails. The console of the tail node synchronizes the monitoring data collected by the monitoring software of the tail node to the console of the head node when the console of the tail node fails. Any node in the plurality of monitoring nodes is also used to alarm when the console of the any node determines that the monitoring data is abnormal.

[0044] In this way, even if the console of any monitoring node in the plurality of monitoring nodes fails, the monitoring node whose console fails can synchronize the monitoring data to the next monitoring node, and the console of the next monitoring node processes the monitoring data. When the console of the next monitoring node determines that the monitoring data is abnormal, an alarm is given.

[0045] In some embodiments, the head node is also used to synchronize the monitoring data collected by the monitoring software of the head node to the console of the tail node when the console of the head node fails. Similarly, the intermediate node is also used to synchronize the monitoring data collected by the monitoring software of the intermediate node to the console of the head node when the console of the intermediate node fails; and the tail node is also used to synchronize the monitoring data collected by the monitoring software of the tail node to the console of the intermediate node when the console of the tail node fails.

[0046] ​In some embodiments, any one of the plurality of monitoring nodes is further configured to present monitoring data collected by the monitoring software of at least one monitoring node through a console. For example, the first node can present at least one of the monitoring data of the first node, the monitoring data of intermediate nodes, and the monitoring data of the tail node through its console. Similarly, intermediate nodes can also present at least one of the monitoring data of the first node, the monitoring data of intermediate nodes, and the monitoring data of the tail node through their consoles; and the tail node can also present at least one of the monitoring data of the first node, the monitoring data of intermediate nodes, and the monitoring data of the tail node through its console.

[0047] Based on the above description, in this embodiment of the application, a console is deployed as a central node at each monitoring node. Each monitoring node has equal status. Through a distributed monitoring method, the monitoring nodes synchronize monitoring data to achieve distributed monitoring. When a problem occurs at a certain node, other nodes can still monitor the entire system, thus meeting disaster recovery requirements.

[0048] In this distributed monitoring system, the monitoring nodes form a chain. Each monitoring node records its next monitoring node in the chain, and the next monitoring node after the last monitoring node (tail node) in the chain becomes the head node. Therefore, the distributed monitoring system provided in this embodiment achieves decentralization, reduces the risk of the entire monitoring system crashing due to the failure of a single monitoring node, improves disaster recovery capabilities, and has good scalability.

[0049] For ease of understanding, the following uses 6 nodes as an example to introduce the distributed monitoring system provided in this application embodiment. Each monitoring node collects monitoring data, and after each monitoring alarm information is collected, it is synchronized among the nodes, and each monitoring node saves a complete copy.

[0050] like Figure 3 As shown, the data on each node is stored in a chain-like data structure of monitoring data blocks, with each monitoring data block recording one monitoring alarm message. Each block consists of a control block portion and a data block portion. The control block portion records the previous data block identifier, the current data block identifier, and a timestamp. The data block portion records the fields of the specific monitoring alarm information, including but not limited to the following:

[0051] The monitoring alarm data block comprises an alarm name alert name, an alarm level alert level, an alarm object alert target, an alarm IP alert IP, an alarm host name alert hostname, an alarm processing state alert status, a first occurrence time first occur time, a latest occurrence time last occur time, a duration maintain time, an accumulated number of times accumulated times, an alarm description alert description, an application system application name, a system administrator system administrator, and an alarm source alert resource.

[0052] It can be seen that the system can realize decentralization of the monitoring system, and each monitoring node is completely equal, and there is no centralized node in the traditional monitoring, and therefore the system has better high availability and disaster recovery functions. When one monitoring site is down, other sites can continue to work normally, and each node is both a production node and a backup node. Moreover, the monitoring record generated by the monitoring method cannot be modified or deleted at will, and after a monitoring alarm information is generated, it will be immediately synchronized to other nodes, and an error alarm can be found by other sites. Meanwhile, the node where the alarm occurs can be tracked, and each monitoring alarm data block can be traced back.

[0053] The distributed monitoring method provided by the embodiment of the application will be introduced below. Referring to FIG. 1, Figure 4 which is a flowchart of a distributed monitoring method provided by the embodiment of the application. The method is applied to a plurality of monitoring nodes, each node of the plurality of monitoring nodes is deployed with a console of monitoring software, and the plurality of monitoring nodes comprise a head node, a tail node and intermediate nodes. The method comprises the following steps.

[0054] S301, when the console of the head node is faulty, the head node synchronizes monitoring data collected by the monitoring software of the head node to the console of the intermediate node.

[0055] S302, when the console of the intermediate node is faulty, the intermediate node synchronizes monitoring data collected by the monitoring software of the intermediate node to the console of the tail node.

[0056] S303, when the console of the tail node is faulty, the tail node synchronizes monitoring data collected by the monitoring software of the tail node to the console of the head node.

[0057] It should be noted that S301-S303 can be executed simultaneously or sequentially, and the present application does not limit the execution sequence of S301-S303, and a person skilled in the art can determine the execution sequence of S301-S303 according to actual needs.

[0058] S304, when the console of any node in the plurality of monitoring nodes determines that the monitoring data is abnormal, an alarm prompt is performed.

[0059] In some embodiments, the method further comprises: when the console of the head node fails, the head node synchronizes the monitoring data collected by the monitoring software of the head node to the console of the tail node.

[0060] In some embodiments, the method further comprises: when the console of the intermediate node fails, the intermediate node synchronizes the monitoring data collected by the monitoring software of the intermediate node to the console of the head node.

[0061] In some embodiments, the method further comprises: when the console of the tail node fails, the tail node synchronizes the monitoring data collected by the monitoring software of the tail node to the console of the intermediate node.

[0062] In some embodiments, the method further comprises: any node in the plurality of monitoring nodes presents the monitoring data collected by the monitoring software of at least one node through the console.

[0063] Based on the above description, in the present application, the console is deployed in the form of a central node in each monitoring node, and each monitoring node is equal in status. Through the distributed monitoring mode and the synchronization of monitoring data between the monitoring nodes, distributed monitoring is achieved. When a problem occurs in a node, other nodes can also monitor the entire system, meeting the disaster recovery requirements.

[0064] Each monitoring node forms a chain, and each monitoring node records the next monitoring node in the chain. The next monitoring node of the last monitoring node (tail node) in the chain is the head node. It can be seen that the distributed monitoring system provided by the present application realizes decentralization, reduces the risk of the entire monitoring system being paralyzed due to the failure of a monitoring node, improves the disaster recovery capability, and has good expansibility.

[0065] The various embodiments described in this specification are described in progressive order, with each embodiment building on the previous one. Each embodiment is described in detail, with reference to the same or similar components in each embodiment. Each embodiment is directed to the differences between the embodiments. The method embodiments are described in more detail in the specification, as the method embodiments are substantially similar to the system embodiments. The method embodiments described above are merely exemplary. One of ordinary skill in the art without any inventive effort, can understand and implement the method embodiments.

[0066] It should be understood that, in this specification, “at least one” means one or more, and “multiple” means two or more. “And / or” is used to describe the relationship between associated objects, which means that there can be three relationships, for example, “A and / or B” can mean that only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character “ / ” generally represents an “or” relationship between the associated objects. “At least one of the following” or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, “a and b”, “a and c”, “b and c”, or “a and b and c”, where a, b, and c can be singular or plural.

[0067] The above is only the preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with the preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make many possible changes and modifications to the technical solutions of the present application, or modify equivalent embodiments with the above disclosed methods and technical contents, without departing from the scope of the technical solutions of the present application. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application, without departing from the content of the technical solutions of the present application, still falls within the scope of protection of the technical solutions of the present application.

Claims

1. A distributed monitoring system, characterized in that, It includes multiple monitoring nodes; each of the multiple monitoring nodes has a console for monitoring software deployed; the multiple monitoring nodes include a head node, a tail node, and intermediate nodes; the multiple monitoring nodes form a chain, and each monitoring node records its next monitoring node in the chain, with the next monitoring node of the tail node being the head node; The first node is used to synchronize the monitoring data collected by the monitoring software of the first node only to the console of the intermediate node when the console of the first node fails. The intermediate node is used to synchronize the monitoring data collected by the monitoring software of the intermediate node only to the console of the tail node when the console of the intermediate node fails. The tail node is used to synchronize the monitoring data collected by the monitoring software of the tail node only to the console of the head node when the console of the tail node fails; any of the plurality of monitoring nodes is also used to issue an alarm when the console of any node determines that the monitoring data synchronized by the previous monitoring node is abnormal.

2. The system according to claim 1, characterized in that, Any of the plurality of monitoring nodes is also used to present the monitoring data collected by the monitoring software of at least one node through the console.

3. A distributed monitoring method, characterized in that, The method is applied to multiple monitoring nodes; each of the multiple monitoring nodes has a console for monitoring software deployed; the multiple monitoring nodes include a head node, a tail node, and intermediate nodes; the multiple monitoring nodes form a chain, and each monitoring node records its next monitoring node in the chain, with the next monitoring node after the tail node being the head node; the method includes: When the console of the first node fails, the first node only synchronizes the monitoring data collected by the monitoring software of the first node to the console of the intermediate node; When the console of the intermediate node fails, the intermediate node only synchronizes the monitoring data collected by the monitoring software of the intermediate node to the console of the tail node; When the console of the tail node fails, only the monitoring data collected by the monitoring software of the tail node is synchronized to the console of the head node; When the console of any of the multiple monitoring nodes determines that the monitoring data synchronized by the previous monitoring node is abnormal, an alarm is issued.

Citation Information

Patent Citations

  • Joint monitoring method and device based on block chain and computer storage medium

    CN110602222A

  • Dynamic monitoring method and device based on block chain

    CN111935289A