A method, device and medium for monitoring the running state of a distributed system

CN116701114BActive Publication Date: 2026-09-29INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310665539.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-05
Publication Date
2026-09-29
Estimated Expiration
2043-06-05

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种用于分布式系统的运行状态监控方法、设备及介质,用于解决当前分布式系统的运行状态监控过于耗费人力资源,且监控效率低,用户使用体验差的问题

Benefits of technology

[0044]通过上述技术方案,本申请通过待监控节点的配置接入、运行状态信息收集、运行日志收集,实现对待监控分布式系统中各个待监控节点的监控信息汇总,并进行展示监控信息,以及进行告警。能够及时地展示不同待监控节点上的软件系统的运行状况,实现软件的健康状况监控,便于用户快速定位软件系统在不同待监控节点上的问题,无需耗费大量人力,进行一一查看或根据经验进行排查各待监控节点问题。从而避免当前分布式系统的运行状态监控过于耗费人力资源,提高了监控效率及用户使用体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116701114B_ABST
    Figure CN116701114B_ABST
Patent Text Reader

Abstract

The application provides a running state monitoring method, device and medium for a distributed system. The method obtains configuration information of each to-be-monitored node in the to-be-monitored distributed system, and generates first running state monitoring information based on heartbeat detection information between the to-be-monitored node and the to-be-monitored node accessed according to the configuration information within a predetermined time interval; and obtains running log information of the to-be-monitored node and takes the running log information as second running state monitoring information. Based on node attributes of each to-be-monitored node and a preset node attribute classification table, each first running state monitoring information and each second running state monitoring information are classified and processed to obtain at least one monitoring information set. Based on each monitoring information set, a preset monitoring alarm strategy is matched to generate alarm information of a to-be-monitored distributed system and / or a to-be-monitored node whose running state reaches an alarm state, and the corresponding running state is displayed to a user terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed system technology, and in particular to a method, device and medium for monitoring the operating status of distributed systems. Background Technology

[0002] As business complexity increases, distributed systems and microservices have emerged. Traditional single projects are vertically split according to business rules, and distributed systems are used to implement distributed business logic. In distributed systems, critical service modules are typically deployed in clusters, and load balancing is used to call distributed system services, preventing single points of failure.

[0003] However, as the number of nodes in a clustered deployment increases, the amount of data related to the operational status of server nodes also grows. Users find it difficult to centrally monitor the operational status of each server node, and the scattered operational status data makes it difficult to establish a convenient monitoring process for viewing the status. Currently, monitoring the operational status of distributed systems is too manpower-intensive, inefficient, and results in a poor user experience. Summary of the Invention

[0004] This application provides a method, device, and medium for monitoring the operational status of distributed systems, which addresses the problems of current distributed system operational status monitoring being too manpower-intensive, having low monitoring efficiency, and providing a poor user experience.

[0005] On one hand, embodiments of this application provide a method for monitoring the operational status of a distributed system, the method comprising:

[0006] Obtain the configuration information of each entity in the distributed system to be monitored; wherein, the configuration information to be monitored includes at least the node access configuration information of the cluster server of the software to be monitored in the distributed system after load balancing.

[0007] Based on the corresponding heartbeat detection information between the monitored nodes and the monitored nodes accessed according to the monitored configuration information within a predetermined time interval, first operating status monitoring information is generated; and

[0008] Obtain the operation log information of the node to be monitored, and use the operation log information as the second operation status monitoring information;

[0009] Based on the node attributes of each node to be monitored and a preset node attribute classification table, the corresponding first operating status monitoring information and second operating status monitoring information are classified to obtain at least one set of monitoring information; wherein, the node attribute classification table includes at least a plurality of node attributes for tree classification; there is an attribute priority relationship among the plurality of node attributes in the tree classification.

[0010] Based on the monitoring information sets, corresponding preset monitoring and alarm strategies are matched to generate alarm information for the monitored distributed system and / or the monitored node whose operating status has reached the alarm state, and the corresponding operating status is displayed to the user terminal.

[0011] In one implementation of this application, the node access configuration information includes at least the server address, port number, and monitoring type of the node to be monitored; the monitoring type includes: system service running status, system server logs, and system component interaction logs.

[0012] In one implementation of this application, first operating status monitoring information is generated based on the corresponding heartbeat detection information between the monitored node and the monitored node accessed according to the monitored configuration information within a predetermined time interval, specifically including:

[0013] Heartbeat packets are sent to the monitored node with which a communication connection has been established at the predetermined time intervals.

[0014] If the monitored node responds to the heartbeat packet, the heartbeat detection information of the monitored node is determined to be in a live state, and the corresponding first running state monitoring information is generated.

[0015] Otherwise, if the heartbeat detection information of the node to be monitored is determined to be in an abnormal state, the corresponding first operating status monitoring information is generated.

[0016] In one implementation of this application, obtaining the operation log information of the node to be monitored specifically includes:

[0017] Based on the preset log types to be collected and the corresponding log directories, incremental collection is performed on the operation logs in the monitored nodes to obtain the operation log information; wherein, the operation logs include at least: the system service logs, middleware usage logs, and database usage logs of the monitored nodes; the incremental collection is performed through message middleware.

[0018] In one implementation of this application, based on the node attributes of each of the monitored nodes and a preset node attribute classification table, the corresponding first operating status monitoring information and second operating status monitoring information are classified to obtain at least one monitoring information set, specifically including:

[0019] Determine the node attributes corresponding to each node to be monitored and the corresponding first and second operating status monitoring information; the node attributes include at least system type, operating status, log level, and date;

[0020] Each node attribute is matched with the node attribute classification table to determine the tree-like classification node order corresponding to each node attribute. Then, according to the tree-like classification node order, each first running status monitoring information and each second running status monitoring information corresponding to the same tree branch are added to the same monitoring information set.

[0021] In one implementation of this application, before determining the node attributes corresponding to each of the nodes to be monitored and the corresponding first and second operating status monitoring information, the method further includes:

[0022] Determine the data display rules corresponding to the node to be monitored; wherein, the data display rules correspond to displaying data according to a predetermined attribute priority relationship;

[0023] Based on the data display rules, the corresponding node attribute classification table is determined.

[0024] In one implementation of this application, the method includes:

[0025] Upon obtaining the monitoring information set, the monitoring information set is stored in a preset relational database; and

[0026] Each monitoring information text in the monitoring information set is segmented into words to determine the keywords corresponding to each monitoring information text; wherein, the monitoring information text includes heartbeat detection information text and operation log information text; the keywords are used to query the corresponding monitoring information text;

[0027] A monitoring information display interface for the monitoring information set is generated and sent to the corresponding user terminal; the monitoring information display interface includes at least a keyword query box and monitoring information text of the monitoring information set displayed according to the attribute priority of the tree-like classification.

[0028] In one implementation of this application, a corresponding preset monitoring alarm strategy is matched based on each of the monitoring information sets, specifically including:

[0029] Load the pre-set rule engine and determine the rule trigger expression corresponding to the rule engine; the rule trigger expression corresponds to the rule trigger condition.

[0030] The alarm level corresponding to each monitoring information text in each monitoring information set is determined by the rule-triggered expression, so as to determine the monitoring alarm strategy based on the alarm level.

[0031] On the other hand, embodiments of this application also provide a device for monitoring the operational status of a distributed system, the device comprising:

[0032] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:

[0033] Obtain the configuration information of each entity in the distributed system to be monitored; wherein, the configuration information to be monitored includes at least the node access configuration information of the cluster server of the software to be monitored in the distributed system after load balancing.

[0034] Based on the corresponding heartbeat detection information between the monitored nodes and the monitored nodes accessed according to the monitored configuration information within a predetermined time interval, first operating status monitoring information is generated; and

[0035] Obtain the operation log information of the node to be monitored, and use the operation log information as the second operation status monitoring information;

[0036] Based on the node attributes of each node to be monitored and a preset node attribute classification table, the corresponding first operating status monitoring information and second operating status monitoring information are classified to obtain at least one set of monitoring information; wherein, the node attribute classification table includes at least a plurality of node attributes for tree classification; there is an attribute priority relationship among the plurality of node attributes in the tree classification.

[0037] Based on the monitoring information sets, corresponding preset monitoring and alarm strategies are matched to generate alarm information for the monitored distributed system and / or the monitored node whose operating status has reached the alarm state, and the corresponding operating status is displayed to the user terminal.

[0038] Furthermore, embodiments of this application also provide a non-volatile computer storage medium for monitoring the operational status of a distributed system, storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:

[0039] Obtain the configuration information of each entity in the distributed system to be monitored; wherein, the configuration information to be monitored includes at least the node access configuration information of the cluster server of the software to be monitored in the distributed system after load balancing.

[0040] Based on the corresponding heartbeat detection information between the monitored nodes and the monitored nodes accessed according to the monitored configuration information within a predetermined time interval, first operating status monitoring information is generated; and

[0041] Obtain the operation log information of the node to be monitored, and use the operation log information as the second operation status monitoring information;

[0042] Based on the node attributes of each node to be monitored and a preset node attribute classification table, the corresponding first operating status monitoring information and second operating status monitoring information are classified to obtain at least one set of monitoring information; wherein, the node attribute classification table includes at least a plurality of node attributes for tree classification; there is an attribute priority relationship among the plurality of node attributes in the tree classification.

[0043] Based on the monitoring information sets, corresponding preset monitoring and alarm strategies are matched to generate alarm information for the monitored distributed system and / or the monitored node whose operating status has reached the alarm state, and the corresponding operating status is displayed to the user terminal.

[0044] Through the above technical solution, this application achieves the aggregation and display of monitoring information for each node in the distributed system under monitoring by configuring and accessing the monitored nodes, collecting operational status information, and collecting operational logs. It also provides alarms. This allows for timely display of the operating status of the software system on different monitored nodes, enabling health monitoring of the software. This facilitates users in quickly locating problems in the software system on different monitored nodes, eliminating the need for extensive manpower to check each node individually or troubleshoot based on experience. This avoids the excessive manpower required for monitoring the operational status of current distributed systems, improving monitoring efficiency and user experience. Attached Figure Description

[0045] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0046] Figure 1 This is a flowchart illustrating a method for monitoring the operational status of a distributed system in an embodiment of this application.

[0047] Figure 2 This is another flowchart illustrating a method for monitoring the operational status of a distributed system according to an embodiment of this application.

[0048] Figure 3 This is a schematic diagram of the structure of a distributed system operation status monitoring device according to an embodiment of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] With increasing business complexity and the rise of distributed systems and microservices, traditional single projects are vertically split according to business rules. Furthermore, to prevent single points of failure, critical service modules are often deployed in clusters, using load balancing to call distributed system services. As the number of nodes increases, the operational status and logs of each distributed system are scattered across various servers. This presents a significant challenge in terms of workload and difficulty for those performing log analysis and monitoring system operation. There is an urgent need for a technical solution that can collect logs and system operational status data scattered across various server nodes, perform aggregation calculations, display the data, and provide alerts and push notifications based on error levels.

[0051] This application provides a method, device, and medium for monitoring the operational status of distributed systems, which addresses the problems of current distributed system operational status monitoring being too manpower-intensive, having low monitoring efficiency, and providing a poor user experience.

[0052] The various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0053] This application provides a method for monitoring the operational status of a distributed system, such as... Figure 1 As shown, the method may include steps S101-S104:

[0054] S101, the server obtains the configuration information of each entity to be monitored in the distributed system to be monitored.

[0055] The configuration information to be monitored includes at least the node access configuration information of the cluster server in the load-balanced distributed system where the software to be monitored is located.

[0056] It should be noted that the server, as the execution subject of the method for monitoring the running status of a distributed system, exists only as an example, and the execution subject is not limited to the server. This application does not make any specific limitation in this regard.

[0057] The node access configuration information mentioned above should include at least the server address, port number, and monitoring type of the node to be monitored. Monitoring types include: system service running status, system server logs, and system component interaction logs.

[0058] In this embodiment, the system software for monitoring the operational status of a distributed system may be deployed on different monitored nodes in different operating systems and clusters. The monitored nodes can be servers, and the server used for operational status monitoring can establish normal communication connections with each monitored node. The operational status of the system services monitored in this application includes, for example, normal, stopped, and abnormal. System server logs include log data from each system's own startup, shutdown, and operation processes. System component interaction logs include logs generated during operation from interactions with other components, such as gateways and Nginx. Load balancing refers to distributing tasks across multiple operational units for execution, such as web servers, FTP servers, enterprise critical application servers, and other critical task servers, thereby collaboratively completing the work tasks.

[0059] S102, the server generates first operational status monitoring information based on the corresponding heartbeat detection information between itself and the monitored nodes connected according to the monitored configuration information within a predetermined time interval. It also obtains the operational log information of the monitored nodes and uses the operational log information as second operational status monitoring information.

[0060] In this embodiment, first operational status monitoring information is generated based on the corresponding heartbeat detection information between the monitored node and the node accessed according to the monitored configuration information within a predetermined time interval, specifically including:

[0061] The server sends heartbeat packets to the monitored nodes with which a communication connection has been established at predetermined time intervals. If a response to the heartbeat packet is received from the monitored node, the server determines that the monitored node's heartbeat detection information indicates a live state and generates corresponding first operational status monitoring information. Otherwise, the server determines that the monitored node's heartbeat detection information indicates an abnormal state and generates corresponding first operational status monitoring information.

[0062] The predetermined time interval can be a user-defined time based on actual usage scenarios. During this predetermined time interval, heartbeat checks are performed on each monitored node. This application does not specify the exact time of the predetermined time interval. The server periodically sends heartbeat packets to each monitored node in the distributed system to complete the heartbeat check. If the server detects a heartbeat packet from a monitored node, it indicates that the service is running well. If no heartbeat is detected, i.e., no heartbeat response is received, it indicates that the monitored node has encountered a problem or has been interrupted or stopped. In this case, a message log indicating an abnormal running status of the monitored node needs to be generated and pushed to the corresponding message log collection component. The message logs indicating that the monitored node is running well or has encountered a problem are added to the first running status monitoring information of each monitored node.

[0063] In this embodiment of the application, the server can also obtain the operation log information of the node to be monitored, specifically including:

[0064] The server incrementally collects operational logs from the monitored nodes based on preset log types and corresponding log directories to obtain operational log information. These operational logs include at least: system service logs, middleware usage logs, and database usage logs from the monitored nodes. Incremental collection is performed through a message middleware.

[0065] The log types to be collected can be service log types pre-defined by the user. Users can also pre-define the directory of the logs to be collected, such as redo logs or rollback logs. The log directory can be the path to the logs to be collected. The server can collect logs through Apache Kafka software or a message broker based on the configured log types and directories. The entire collection process is real-time, automatically detecting changes in logs across various systems, and then incrementally collecting only the most recently generated logs before adding them to the runtime log information.

[0066] S103, the server classifies the corresponding first running status monitoring information and second running status monitoring information based on the node attributes of each node to be monitored and the preset node attribute classification table, so as to obtain at least one set of monitoring information.

[0067] The node attribute classification table includes at least several node attributes used for tree-like classification. There are attribute priority relationships among the multiple node attributes in the tree-like classification.

[0068] In this application embodiment, the node attribute classification table can be determined through the following examples:

[0069] The server determines the data display rules for the nodes to be monitored. These rules correspond to displaying data according to a predetermined attribute priority relationship. Based on the data display rules, a corresponding node attribute classification table is determined.

[0070] In practical use, different data display rules can be applied to different monitored nodes. These rules can be based on the order in which node attributes are displayed, such as system type, running status, log level, and date. For example, system type might precede running status, with system type attributes having higher priority than running status attributes. The server can obtain the data display rules for the monitored nodes based on the monitoring software and generate a node attribute classification table according to attribute priority.

[0071] In this embodiment, based on the node attributes of each node to be monitored and a preset node attribute classification table, the corresponding first operating state monitoring information and second operating state monitoring information are classified to obtain at least one monitoring information set, specifically including:

[0072] The server determines the node attributes corresponding to each node to be monitored and its corresponding first and second running status monitoring information. Node attributes include at least system type, running status, log level, and date. Each node attribute is matched against a node attribute classification table to determine the tree-structured classification node order. Following this tree-structured classification node order, each first and second running status monitoring piece of information corresponding to the same tree branch is added to the same monitoring information set.

[0073] In other words, the server can collect operational status and logs, and then aggregate the logs distributed across multiple monitored nodes. Since each monitored node service is independently developed and deployed, the data formats to be aggregated may vary greatly. Therefore, the collected data needs to be stored in a unified structure (e.g., by system, by operational status, by log level, by date, etc.), such as in a relational database or a distributed full-text search engine like ElasticSearch.

[0074] The tree-like classification node order described above can be understood as follows: if system type is the first-level node, and there are Type 1 and Type 2 systems, then under Type 1 there are child nodes for running status, which may also have node log levels or node dates, arranged sequentially. Similarly, under Type 2, the running status child nodes are stored, which may also have node log levels or node dates, arranged sequentially. This allows the server to sort the first and second running status monitoring information of each monitored node according to the tree structure. In actual use, for example, when a user clicks the control corresponding to the system type, they can view the collapsed running status child nodes below it. Clicking the running status control allows them to view the first and second running status monitoring information for each corresponding log level. Furthermore, the log levels can be sorted by date for the first and second running status monitoring information, and clicking on a log level can expand the collapsed first and second running status monitoring information corresponding to different dates.

[0075] S104, the server matches the corresponding preset monitoring alarm policies based on each monitoring information set to generate alarm information for the monitored distributed system and / or monitored node whose operating status has reached the alarm state, and displays the corresponding operating status to the user terminal.

[0076] In this embodiment, upon obtaining the monitoring information set, the server stores the monitoring information set in a preset relational database. The server then performs word segmentation on each monitoring information text in the monitoring information set to determine the keywords corresponding to each text. The monitoring information text includes heartbeat detection information text and operation log information text. Keywords are used to query the corresponding monitoring information text. A monitoring information display interface for the monitoring information set is generated and sent to the corresponding user terminal. The monitoring information display interface includes at least a keyword query box and monitoring information text of the monitoring information set displayed according to attribute priority in a tree-like classification.

[0077] The above embodiments can summarize system operation data, allowing users to clearly see the current status of each monitored node (whether the operation is normal, whether there are serious errors during operation, etc.). For monitored nodes with errors, users can drill down to see what type of error was reported on which day, enabling quick problem localization. Secondly, a comprehensive query function page is provided, allowing monitors to customize conditions for flexible multi-dimensional queries. Furthermore, word segmentation technology is used to achieve a keyword search function similar to Baidu.

[0078] In this embodiment of the application, based on each monitoring information set, a corresponding preset monitoring alarm strategy is matched, specifically including:

[0079] The server loads a pre-configured rule engine and determines the corresponding rule trigger expression. The rule trigger expression corresponds to the rule trigger condition. Using the rule trigger expression, the alarm level corresponding to each monitoring information text in each monitoring information set is determined, and the monitoring alarm strategy is determined based on the alarm level.

[0080] In this embodiment, users can set rule trigger conditions, i.e., rule trigger expressions, through the rule engine. For example, if the operating status of a monitored node becomes abnormal, reaching the first-level alarm level, or if a null value appears in the log of a monitored node, reaching the second-level alarm level. After generating the monitoring information set, the server needs to be able to notify the development and maintenance personnel of each system in real time regarding any issues, so that they can promptly identify and resolve problems and avoid unnecessary losses. Therefore, the server can use the rule engine to set rule trigger conditions, specifying what conditions must be met for each system to trigger an alarm. Once an alarm is triggered, the responsible user terminal (such as a mobile phone, computer, or other device) of the corresponding monitored node is promptly notified via SMS and email.

[0081] Figure 2 Here is another flowchart illustrating a method for monitoring the operational status of a distributed system, such as... Figure 2As shown, through the technical solution for monitoring the operation status of distributed systems in this application, user 210 can view each distributed system 220 (software system), complete data collection and summary calculation 230 through distributed system access configuration 231, distributed system data collection 232, and distributed system data classification and summarization 233, and complete data display and early warning 240 by displaying summary data 241, setting alarm rules 242, and sending alarm messages 243.

[0082] Through the above technical solution, this application achieves the aggregation and display of monitoring information for each node in the distributed system under monitoring by configuring and accessing the monitored nodes, collecting operational status information, and collecting operational logs. It also provides alarms. This allows for timely display of the operating status of the software system on different monitored nodes, enabling health monitoring of the software. This facilitates users in quickly locating problems in the software system on different monitored nodes, eliminating the need for extensive manpower to check each node individually or troubleshoot based on experience. This avoids the excessive manpower required for monitoring the operational status of current distributed systems, improving monitoring efficiency and user experience.

[0083] Figure 3 This application provides a schematic diagram of the structure of a device for monitoring the operational status of a distributed system, the device comprising:

[0084] At least one processor; and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0085] The system acquires the configuration information of each node in the distributed system to be monitored. This configuration information includes at least the node access configuration information of the load-balanced cluster server where the monitored software resides. Based on the heartbeat detection information between the monitored nodes and the nodes connected according to the configuration information within a predetermined time interval, first operational status monitoring information is generated. The system also acquires the operational log information of the monitored nodes and uses this log information as second operational status monitoring information. Based on the node attributes of each monitored node and a preset node attribute classification table, the corresponding first and second operational status monitoring information are classified to obtain at least one set of monitoring information. The node attribute classification table includes at least several node attributes for tree-like classification. There are attribute priority relationships among the multiple node attributes in the tree-like classification. Based on each monitoring information set, a corresponding preset monitoring alarm strategy is matched to generate alarm information for the monitored distributed system and / or monitored nodes whose operational status has reached an alarm state, and the corresponding operational status is displayed to the user terminal.

[0086] This application embodiment also provides a non-volatile computer storage medium for monitoring the operational status of a distributed system, storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:

[0087] The system acquires the configuration information of each node in the distributed system to be monitored. This configuration information includes at least the node access configuration information of the load-balanced cluster server where the monitored software resides. Based on the heartbeat detection information between the monitored nodes and the nodes connected according to the configuration information within a predetermined time interval, first operational status monitoring information is generated. The system also acquires the operational log information of the monitored nodes and uses this log information as second operational status monitoring information. Based on the node attributes of each monitored node and a preset node attribute classification table, the corresponding first and second operational status monitoring information are classified to obtain at least one set of monitoring information. The node attribute classification table includes at least several node attributes for tree-like classification. There are attribute priority relationships among the multiple node attributes in the tree-like classification. Based on each monitoring information set, a corresponding preset monitoring alarm strategy is matched to generate alarm information for the monitored distributed system and / or monitored nodes whose operational status has reached an alarm state, and the corresponding operational status is displayed to the user terminal.

[0088] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0089] The devices, media, and methods provided in this application are one-to-one correspondences. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0090] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0091] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for monitoring the operational status of a distributed system, characterized in that, The method includes: Obtain the configuration information of each entity in the distributed system to be monitored; wherein, the configuration information to be monitored includes at least the node access configuration information of the cluster server of the software to be monitored in the distributed system after load balancing. Based on the corresponding heartbeat detection information between the monitored nodes and the monitored nodes accessed according to the monitored configuration information within a predetermined time interval, first operating status monitoring information is generated; and Obtain the operation log information of the node to be monitored, and use the operation log information as the second operation status monitoring information; Based on the node attributes of each node to be monitored and a preset node attribute classification table, the corresponding first operating status monitoring information and second operating status monitoring information are classified to obtain at least one set of monitoring information; wherein, the node attribute classification table includes at least a plurality of node attributes for tree classification; there is an attribute priority relationship among the plurality of node attributes in the tree classification. Based on the monitoring information sets, corresponding preset monitoring and alarm strategies are matched to generate alarm information for the monitored distributed system and / or the monitored node whose operating status has reached the alarm state, and the corresponding operating status is displayed to the user terminal.

2. The method for monitoring the operational status of a distributed system according to claim 1, characterized in that, The node access configuration information includes at least the server address, port number, and monitoring type of the node to be monitored; the monitoring type includes: system service running status, system server logs, and system component interaction logs.

3. The method for monitoring the operational status of a distributed system according to claim 1, characterized in that, Based on the corresponding heartbeat detection information between the monitored node and the monitored node accessed according to the monitored configuration information within a predetermined time interval, first operating status monitoring information is generated, specifically including: Heartbeat packets are sent to the monitored node with which a communication connection has been established at the predetermined time intervals. If the monitored node responds to the heartbeat packet, the heartbeat detection information of the monitored node is determined to be in a live state, and the corresponding first running state monitoring information is generated. Otherwise, if the heartbeat detection information of the node to be monitored is determined to be in an abnormal state, the corresponding first operating status monitoring information is generated.

4. The method for monitoring the operational status of a distributed system according to claim 1, characterized in that, Obtaining the operation log information of the node to be monitored specifically includes: Based on the preset log types to be collected and the corresponding log directories, incremental collection is performed on the operation logs in the monitored nodes to obtain the operation log information; wherein, the operation logs include at least: the system service logs, middleware usage logs, and database usage logs of the monitored nodes; the incremental collection is performed through message middleware.

5. A method for monitoring the operational status of a distributed system according to claim 1, characterized in that, Based on the node attributes of each node to be monitored and a preset node attribute classification table, the corresponding first operating status monitoring information and second operating status monitoring information are classified to obtain at least one set of monitoring information, specifically including: Determine the node attributes corresponding to each node to be monitored and the corresponding first and second operating status monitoring information; the node attributes include at least system type, operating status, log level, and date; Each node attribute is matched with the node attribute classification table to determine the tree-like classification node order corresponding to each node attribute. Then, according to the tree-like classification node order, each first running status monitoring information and each second running status monitoring information corresponding to the same tree branch are added to the same monitoring information set.

6. The method for monitoring the operational status of a distributed system according to claim 5, characterized in that, Before determining the node attributes corresponding to each of the nodes to be monitored and the corresponding first and second operating status monitoring information, the method further includes: Determine the data display rules corresponding to the node to be monitored; wherein, the data display rules correspond to displaying data according to a predetermined attribute priority relationship; Based on the data display rules, the corresponding node attribute classification table is determined.

7. The method for monitoring the operational status of a distributed system according to claim 1, characterized in that, The method includes: Upon obtaining the monitoring information set, the monitoring information set is stored in a preset relational database; and Each monitoring information text in the monitoring information set is segmented into words to determine the keywords corresponding to each monitoring information text; wherein, the monitoring information text includes heartbeat detection information text and operation log information text; the keywords are used to query the corresponding monitoring information text; A monitoring information display interface for the monitoring information set is generated and sent to the corresponding user terminal; the monitoring information display interface includes at least a keyword query box and monitoring information text of the monitoring information set displayed according to the attribute priority of the tree-like classification.

8. A method for monitoring the operational status of a distributed system according to claim 1, characterized in that, Based on the aforementioned monitoring information sets, corresponding preset monitoring alarm strategies are matched, specifically including: Load a pre-set rule engine and determine the rule trigger expression corresponding to the rule engine; the rule trigger expression corresponds to the rule trigger condition. The alarm level corresponding to each monitoring information text in each monitoring information set is determined by the rule-triggered expression, so as to determine the monitoring alarm strategy based on the alarm level.

9. A device for monitoring the operational status of a distributed system, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform a method for monitoring the operational status of a distributed system as described in any one of claims 1-8.

10. A non-volatile computer storage medium for monitoring the operational status of a distributed system, storing computer-executable instructions, characterized in that, The computer-executable instructions are capable of executing a method for monitoring the operational status of a distributed system as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Mobile equipment secure data transmission method

    CN105516110A

  • Multistage dispatching distributed parallel computing-oriented monitoring system and monitoring method

    CN105703940A