Availability monitoring system and method for a distributed network system

By deploying monitoring nodes in a distributed network system and utilizing a central node for real-time data analysis and alarms, the problem of response latency in centralized monitoring solutions is solved, achieving efficient distributed network system monitoring and fault prediction, and improving system stability and user experience.

CN118612113BActive Publication Date: 2026-01-20SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410844978.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2026-01-20
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

Existing centralized monitoring solutions cannot effectively solve the response latency problem of distributed network systems, making it difficult to ensure their high availability.

Method used

In a distributed network system, monitoring nodes are deployed to collect system data in real time. Availability analysis and anomaly detection are performed through a central node, and alarm information is quickly triggered for operation and maintenance.

Benefits of technology

It improves the stability and reliability of distributed network systems, enables real-time monitoring and fault diagnosis, enhances network service quality and user experience, and reduces the risk of failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118612113B_ABST
    Figure CN118612113B_ABST
Patent Text Reader

Abstract

The application provides an availability monitoring system of a distributed network system, which comprises a center node and a plurality of monitoring nodes, the plurality of monitoring nodes are distributed in corresponding distributed networks of the distributed network system; wherein: the center node is configured to configure a monitoring task and distribute the monitoring task to the plurality of monitoring nodes; each monitoring node is configured to collect system data of the distributed network system in the region to which the monitoring node belongs according to the monitoring task and upload the system data to the center node; the center node is further configured to analyze the availability of the distributed network system according to the system data uploaded by each monitoring node, determine whether the distributed network system is abnormal, and push alarm information to an operation and maintenance team of the distributed network system when it is determined that the distributed network system is abnormal, so that the operation and maintenance team performs operation and maintenance processing according to the alarm information. The application can quickly trigger an alarm when a problem is found and quickly respond to a fault.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of system monitoring technology, and in particular to an availability monitoring system and method for a distributed network system. Background Technology

[0002] Distributed network systems play a vital role in modern society, but ensuring their high availability remains a challenge. Currently, centralized monitoring of distributed network systems is primarily handled by a monitoring center. However, this traditional centralized monitoring approach has several drawbacks, such as response latency. Therefore, it is necessary to provide a new monitoring solution. Summary of the Invention

[0003] To address at least one of the above-mentioned technical problems, embodiments of the present invention provide an availability monitoring system and method for a distributed network system.

[0004] In a first aspect, the availability monitoring system for a distributed network system provided in this embodiment of the invention includes a central node and multiple monitoring nodes, wherein the multiple monitoring nodes are distributedly deployed in a corresponding distributed network of the distributed network system; wherein:

[0005] The central node is used to configure monitoring tasks and distribute the monitoring tasks to the multiple monitoring nodes; wherein, the monitoring task includes the monitoring frequency, the entry address of the distributed network system in each region, and the address of each monitoring node;

[0006] Each monitoring node is used to: collect system data of the distributed network system in the area to which the monitoring node belongs, according to the monitoring task, and upload the system data to the central node;

[0007] The central node is also used to: analyze the availability of the distributed network system based on the system data uploaded by each monitoring node, determine whether the distributed network system is abnormal, and when it is determined that the distributed network system is abnormal, push alarm information to the operation and maintenance team of the distributed network system so that the operation and maintenance team can perform operation and maintenance processing based on the alarm information.

[0008] In one embodiment, the system data includes system performance metrics, system logs, and / or user access data; wherein the system performance metrics include at least one of response time, throughput, and error rate; and the user access data includes at least one of access frequency, page dwell time, and bounce rate.

[0009] In one embodiment, the central node is specifically used to: perform anomaly detection operations on the system logs using text mining methods or pattern matching methods; if anomalies are found in the system logs, determine that the distributed network system is abnormal, and generate an operational status assessment report of the distributed network system based on the level and degree of anomaly of the system logs.

[0010] In one embodiment, the central node is specifically used to: calculate a statistical measure of the system performance index based on the system performance index within a preset time period; determine the performance status of the distributed network system based on the statistical measure, and determine whether the distributed network system has an anomaly based on the performance status; wherein the statistical measure includes at least one of mean, variance, and maximum value.

[0011] In one embodiment, the central node is further configured to: analyze the user's experience using the distributed network system based on the user access data, and analyze potential problems encountered by the user during the use of the distributed network system based on the experience.

[0012] In one embodiment, the central node is further configured to: pre-configure an alarm policy, the alarm policy including an anomaly analysis algorithm, an alarm information push method, and an alarm information receiving address for the operations and maintenance team under the push method; wherein, the anomaly matching algorithm is used to analyze whether the distributed network system has an anomaly based on the system data; the alarm information push method includes at least one of SMS, email, WeChat, and telephone;

[0013] Correspondingly, the central node is specifically used to: analyze the system data using the anomaly analysis algorithm to obtain an analysis result indicating whether the distributed network system has an anomaly; if the analysis result indicates an anomaly, then push the alarm information to the operation and maintenance team according to the push method in the alarm strategy and the alarm information receiving address.

[0014] In one embodiment, the central node maintains metadata information for each monitoring node, including the region, server address, and node information of each monitoring node.

[0015] Correspondingly, each monitoring node is also used to: report heartbeat data to the central node at preset intervals; the central node is also used to: determine whether it receives heartbeat data from each monitoring node at preset intervals based on the metadata information of each monitoring node it maintains, and obtain a determination result; based on the determination result, determine whether each monitoring node is in working state, so as to send the monitoring task to the monitoring nodes in working state.

[0016] In one embodiment, each monitoring node is further configured to: acquire historical system data, preprocess the historical system data, input the preprocessed historical system data into a machine learning model to obtain a prediction result of whether the distributed network system will experience a failure or performance bottleneck in the future preset time period, and push the prediction result to the operation and maintenance team so that the operation and maintenance team can formulate maintenance measures to reduce the possibility of the distributed network system experiencing a failure or performance bottleneck in the future preset time period.

[0017] In one embodiment, the central node is further configured to: generate a visual chart reflecting the operating status of the distributed network system based on the system data reported by each monitoring node, and display the visual chart on a page; and / or, the central node is specifically configured to: analyze the system data uploaded by each monitoring node using regression analysis, classification algorithm or clustering analysis method to obtain the analysis results of the availability of the distributed network system.

[0018] According to a second aspect, the availability monitoring method for a distributed network system provided in the embodiments of the present invention is implemented based on the availability monitoring system provided in the first aspect, and the method includes:

[0019] The central node configures monitoring tasks and distributes these tasks to the multiple monitoring nodes. Each monitoring task includes a monitoring frequency, the entry address of the distributed network system in each region, and the address of each monitoring node.

[0020] Each monitoring node collects system data of the distributed network system in the area to which it belongs, according to the monitoring task, and uploads the system data to the central node.

[0021] The central node analyzes the availability of the distributed network system based on the system data uploaded by each monitoring node, determines whether the distributed network system is abnormal, and pushes alarm information to the operation and maintenance team of the distributed network system when an abnormality is determined, so that the operation and maintenance team can perform operation and maintenance based on the alarm information.

[0022] The availability monitoring system and method for distributed network systems provided in this invention deploy monitoring nodes in the distributed network to collect and upload system data in real time. Then, the central node performs availability analysis based on the system data uploaded by these monitoring nodes and quickly triggers alarms when problems are detected. This improves the stability and reliability of the distributed network system, enables real-time monitoring, data collection, analysis, and fault diagnosis of the distributed network system, helps improve the quality of network services, enhances user experience, and improves the overall utilization efficiency of network resources. It can ensure the efficient operation of large-scale network services and rapid response to faults. Attached Figure Description

[0023] Figure 1 This is a structural block diagram of an availability monitoring system for a distributed network system according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating the interaction between the user, monitoring node, and central node in one embodiment of the present invention.

[0025] Figure 3 This is a flowchart of an availability monitoring method for a distributed network system according to an embodiment of the present invention. Detailed Implementation

[0026] In a first aspect, embodiments of the present invention provide an availability monitoring system for a distributed network system, see [link to relevant documentation]. Figure 1 The system includes a central node and multiple monitoring nodes, which are distributed and deployed in the corresponding distributed network of the distributed network system; wherein:

[0027] The central node is used to configure monitoring tasks and distribute the monitoring tasks to the multiple monitoring nodes; wherein, the monitoring task includes the monitoring frequency, the entry address of the distributed network system in each region, and the address of each monitoring node;

[0028] Each monitoring node is used to: collect system data of the distributed network system in the area to which the monitoring node belongs, according to the monitoring task, and upload the system data to the central node;

[0029] The central node is also used to: analyze the availability of the distributed network system based on the system data uploaded by each monitoring node, determine whether the distributed network system is abnormal, and when it is determined that the distributed network system is abnormal, push alarm information to the operation and maintenance team of the distributed network system so that the operation and maintenance team can perform operation and maintenance processing based on the alarm information.

[0030] The availability monitoring system includes a central node and multiple monitoring nodes, all of which are communicatively connected to the central node.

[0031] Users can configure monitoring tasks on the central node, which then distributes these tasks to each monitoring node, which in turn collects data according to the assigned tasks. The monitoring tasks specify a monitoring frequency, which can also be understood as the data collection frequency. The distributed network system can be accessed via its entry address; therefore, monitoring nodes can access the system based on the entry address corresponding to their location. The address of each monitoring node represents its respective region.

[0032] Each monitoring node collects system data from the distributed network system within its assigned area based on the received monitoring task, and then uploads the system data to the central node.

[0033] The central node receives system data uploaded from various monitoring nodes, analyzes the availability of the distributed network system based on the data, determines whether the distributed network system is abnormal, and pushes alarm information to the operation and maintenance team of the distributed network system when an anomaly is determined. The operation and maintenance team then performs operation and maintenance procedures based on the alarm information.

[0034] In one embodiment, the system data may include system performance metrics, system logs, and / or user access data; wherein the system performance metrics may include at least one of response time, throughput, and error rate; and the user access data may include at least one of access frequency, page dwell time, and bounce rate.

[0035] Furthermore, the central node can be specifically used to: perform anomaly detection operations on the system logs using text mining methods or pattern matching methods; if anomalies are found in the system logs, it is determined that the distributed network system is abnormal, and an operational status assessment report of the distributed network system is generated based on the level and degree of anomaly of the system logs.

[0036] Text mining methods refer to methods for extracting anomalous text information from system data according to certain rules. Pattern matching methods refer to methods for matching anomalous information from system data according to rule-based matching methods.

[0037] In other words, the central node can use text mining or pattern matching methods to detect anomalies in the system logs. If anomalies are found, it indicates that an anomaly has occurred in the distributed network system. At this point, an operational status assessment report will be generated based on the level of the system log and the severity of the anomaly. For example, if the system log level is the highest and the anomaly severity is the most severe, the operational status in the operational status assessment report will be the worst.

[0038] Furthermore, the central node can be specifically used to: calculate the statistics of the system performance indicators based on the system performance indicators within a preset time period; determine the performance status of the distributed network system based on the statistics, and determine whether the distributed network system has an anomaly based on the performance status; wherein the statistics include at least one of the mean, variance, and maximum value.

[0039] In other words, the central node can calculate at least one of the statistical measures of mean, variance, and maximum value based on the system performance indicators over a period of time. It can then determine the performance state of the distributed network system based on these statistics, and further determine whether the distributed network system is experiencing anomalies based on the performance state. For example, if the performance state is below normal, the distributed network system is considered to be experiencing anomalies.

[0040] Furthermore, the central node can be specifically used to: analyze the user's experience using the distributed network system based on the user access data, and analyze potential problems encountered by the user in the process of using the distributed network system based on the experience.

[0041] In other words, the central node can determine the user experience in the distributed network system based on access frequency, page dwell time, and bounce rate. High access frequency, long page dwell time, and low bounce rate indicate a good user experience; otherwise, a poor user experience. This data then helps determine if the user experience in the distributed network system is unsatisfactory, highlighting potential issues with low user engagement.

[0042] In one embodiment, the central node can also be used to: pre-configure an alarm policy, the alarm policy including an anomaly analysis algorithm, an alarm information push method, and an alarm information receiving address for the operations and maintenance team under the push method; wherein, the anomaly matching algorithm is used to analyze whether the distributed network system has an anomaly based on the system data; the alarm information push method includes at least one of SMS, email, WeChat, and telephone;

[0043] Correspondingly, the central node is specifically used to: analyze the system data using the anomaly analysis algorithm to obtain an analysis result indicating whether the distributed network system has an anomaly; if the analysis result indicates an anomaly, then push the alarm information to the operation and maintenance team according to the push method in the alarm strategy and the alarm information receiving address.

[0044] In other words, users can configure alarm policies on the central node. These policies include anomaly analysis algorithms, alarm push methods, and the receiving addresses for those push methods. Anomaly analysis algorithms refer to algorithms that analyze whether anomalies have occurred in the distributed network system, such as the text mining and pattern matching methods mentioned above. Push methods can include email, SMS, and WeChat. For email, the receiving address is the email address of the operations and maintenance team members; for SMS, it's their mobile phone number; and for WeChat, it's their WeChat account.

[0045] Based on the above configuration, the central node can use anomaly analysis algorithms to analyze the system data and obtain an analysis result indicating whether the distributed network system has an anomaly. If the analysis result indicates an anomaly, the alarm information is pushed to the alarm information receiving address according to the push method in the alarm policy, so that the operation and maintenance team will receive the alarm information.

[0046] In one embodiment, the central node maintains metadata information for each monitoring node, including the region, server address, and node information of each monitoring node.

[0047] Correspondingly, each monitoring node can also be used to: report heartbeat data to the central node at preset intervals; the central node is also used to: determine whether it receives heartbeat data from each monitoring node at preset intervals based on the metadata information of each monitoring node it maintains, and obtain a determination result; based on the determination result, determine whether each monitoring node is in a working state, so as to send the monitoring task to the monitoring nodes in a working state.

[0048] The metadata information maintained by the central node for each monitoring node includes the region where each monitoring node is located, its server address (i.e., entry address), and node information, such as the monitoring node's name and identifier. Each monitoring node reports heartbeat data to the central node at preset intervals. If the central node receives heartbeat data from a monitoring node at preset intervals, it indicates that the monitoring node is in a working state. If the central node does not receive heartbeat data from a monitoring node, it indicates that the monitoring node is not in a working state. The central node sends monitoring tasks to monitoring nodes that are in a working state, and for monitoring nodes that are not in a working state, the central node will remind relevant personnel to investigate the problem.

[0049] In one embodiment, each monitoring node may also be used to: acquire historical system data, preprocess the historical system data, input the preprocessed historical system data into a machine learning model, obtain a prediction result of whether the distributed network system will experience a failure or performance bottleneck in the future preset time period, and push the prediction result to the operation and maintenance team so that the operation and maintenance team can formulate maintenance measures to reduce the possibility of the distributed network system experiencing a failure or performance bottleneck in the future preset time period.

[0050] In other words, each monitoring node acquires historical system data, preprocesses it (e.g., normalization, data format standardization), and inputs the preprocessed historical system data into a machine learning model. This model can predict whether the distributed network system will experience failures or performance bottlenecks within a preset future timeframe. The prediction results are then pushed to the operations and maintenance team, who will then implement maintenance measures to reduce the likelihood of failures or performance bottlenecks in the distributed network system within the preset future timeframe, thereby improving the reliability of the distributed network system.

[0051] In one embodiment, the central node can also be used to: generate a visual chart reflecting the operating status of the distributed network system based on the system data reported by each monitoring node, and display the visual chart on a page.

[0052] In other words, the central node will generate visualization charts from the system data reported by each monitoring node, such as bar charts or line charts of a certain monitoring indicator. These visualization charts can reflect the operating status of the distributed network system and be displayed on the page so that operation and maintenance personnel can understand the trend of the operating status of the distributed network system.

[0053] In one embodiment, the central node may be specifically used to: analyze the system data uploaded by each monitoring node using regression analysis, classification algorithm, or clustering analysis to obtain the analysis results of the availability of the distributed network system.

[0054] In other words, by processing system data through regression analysis, classification algorithms, or cluster analysis, the distribution or classification of the operating indicators of the distributed network system can be obtained. Based on the distribution or classification, the availability level of the distributed network system can be determined.

[0055] See Figure 2 , Figure 2 The service in the middle actually refers to the central node. Figure 2 The nodes in this context actually refer to monitoring nodes, and the monitoring tasks refer to monitoring tasks. Users configure monitoring tasks and alarm policies on the central node, send monitoring tasks to the monitoring nodes that report heartbeat data, receive system data reported by the monitoring nodes, clean and process the system data reported by multiple monitoring nodes, and then determine whether a fault has occurred. If a fault occurs, an alarm notification is pushed to the central node.

[0056] In practical scenarios, an efficient data transmission network is established between each monitoring node and the central node to ensure that the data collected by the monitoring nodes can be reliably and in real time transmitted to the central node. Specifically, the collected system data can be sent to the central node via HTTP / 2, gRPC, etc.

[0057] This invention involves "distributed network monitoring technology" and "application performance management technology." By deploying monitoring nodes in a distributed network, system data is collected and uploaded in real time. A central node then performs availability analysis based on this data, quickly triggering alarms when problems are detected. This improves the stability and reliability of the distributed network system, enabling real-time monitoring, data collection, analysis, and fault diagnosis. This helps improve network service quality, enhance user experience, and increase the overall utilization efficiency of network resources, ensuring efficient operation and rapid fault response for large-scale network services. Furthermore, the system provided in this invention uses machine learning models to predict potential faults and performance issues, taking proactive measures to reduce the risk of failures.

[0058] As can be seen, the availability monitoring system provided by the embodiments of the present invention can monitor the availability of distributed network systems in real time, promptly identify and address problems, improve the stability and reliability of distributed network systems, and reduce the risk of failures or bottlenecks by predicting whether failures or performance bottlenecks will occur in the future. The system provided by the embodiments of the present invention has a wide range of applications, including but not limited to cloud computing services, big data platforms, Internet of Things (IoT) applications, and network services, and is particularly suitable for applications in cloud computing, big data processing, and Internet services.

[0059] Secondly, embodiments of the present invention provide an availability monitoring method for a distributed network system, the method being implemented based on the availability monitoring system provided in the first aspect, see [link to relevant documentation]. Figure 3 The method includes the following steps S110 to S130:

[0060] S110. Configure monitoring tasks through the central node and distribute the monitoring tasks to the multiple monitoring nodes; wherein, the monitoring task includes the monitoring frequency, the entry address of the distributed network system in each region, and the address of each monitoring node;

[0061] S120. Each monitoring node collects system data of the distributed network system in the area to which the monitoring node belongs according to the monitoring task, and uploads the system data to the central node.

[0062] S130. The central node analyzes the availability of the distributed network system based on the system data uploaded by each monitoring node, determines whether the distributed network system is abnormal, and pushes alarm information to the operation and maintenance team of the distributed network system when it is determined that the distributed network system is abnormal, so that the operation and maintenance team can perform operation and maintenance processing based on the alarm information.

[0063] In one embodiment, the system data includes system performance metrics, system logs, and / or user access data; wherein the system performance metrics include at least one of response time, throughput, and error rate; and the user access data includes at least one of access frequency, page dwell time, and bounce rate.

[0064] In one embodiment, determining whether the distributed network system has malfunctioned includes:

[0065] The central node uses text mining or pattern matching methods to perform anomaly detection on the system logs; if anomalies are found in the system logs, it is determined that the distributed network system is abnormal, and an operational status assessment report of the distributed network system is generated based on the level and severity of the anomalies in the system logs.

[0066] In one embodiment, determining whether the distributed network system has malfunctioned includes:

[0067] The central node calculates statistics of the system performance indicators based on the system performance indicators within a preset time period; determines the performance status of the distributed network system based on the statistics, and determines whether the distributed network system has an anomaly based on the performance status; wherein the statistics include at least one of mean, variance, and maximum value.

[0068] In one embodiment, determining whether the distributed network system has malfunctioned includes:

[0069] The central node analyzes the user's experience using the distributed network system based on the user access data, and analyzes potential problems encountered by the user in using the distributed network system based on the experience.

[0070] In one embodiment, the method further includes:

[0071] The central node pre-configures an alarm policy, which includes an anomaly analysis algorithm, an alarm information push method, and an alarm information receiving address for the operations and maintenance team under the push method. The anomaly matching algorithm is used to analyze whether the distributed network system is experiencing anomalies based on system data. The alarm information push method includes at least one of SMS, email, WeChat, and telephone.

[0072] Correspondingly, determining whether the distributed network system is abnormal, and when it is determined that the distributed network system is abnormal, pushing alarm information to the operation and maintenance team of the distributed network system, includes: analyzing the system data through the central node using the anomaly analysis algorithm to obtain the analysis result of whether the distributed network system is abnormal; if the analysis result indicates that an anomaly has occurred, then pushing the alarm information to the operation and maintenance team according to the push method in the alarm strategy and the alarm information receiving address.

[0073] In one embodiment, the method further includes:

[0074] The central node maintains metadata information for each monitoring node, including the region, server address, and node information of each monitoring node.

[0075] Each monitoring node reports heartbeat data to the central node at preset intervals.

[0076] The central node determines whether it receives heartbeat data from each monitoring node at preset intervals based on the metadata information of each monitoring node it maintains, and obtains a judgment result. Based on the judgment result, it determines whether each monitoring node is in working state, and sends the monitoring task to the monitoring nodes in working state.

[0077] In one embodiment, the method further includes:

[0078] Historical system data is acquired through each monitoring node, preprocessed, and then input into a machine learning model to obtain a prediction of whether the distributed network system will experience a failure or performance bottleneck within a preset future time period. The prediction result is then pushed to the operations and maintenance team so that the team can formulate maintenance measures to reduce the likelihood of the distributed network system experiencing a failure or performance bottleneck within the preset future time period.

[0079] In one embodiment, the method further includes:

[0080] The central node generates a visual chart reflecting the operational status of the distributed network system based on the system data reported by each monitoring node, and displays the visual chart on the page.

[0081] In one embodiment, the method further includes:

[0082] By employing regression analysis, classification algorithms, or cluster analysis methods at the central node, the system data uploaded by each monitoring node is analyzed to obtain the availability analysis results of the distributed network system.

[0083] It is understood that explanations, specific implementation methods, beneficial effects, examples, etc. of the methods provided in the embodiments of the present invention can be found in the corresponding parts of the system provided in the first aspect, and will not be repeated here.

[0084] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0085] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using hardware, software, widgets, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.

[0086] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.

Claims

1. An availability monitoring system for a distributed network system, characterized in that The system comprises a center node and a plurality of monitoring nodes, which are distributed in corresponding distributed network systems; The center node is configured to configure a monitoring task and distribute the monitoring task to the plurality of monitoring nodes, wherein the monitoring task comprises a monitoring frequency, an entry address of the distributed network system in each region, and an address of each monitoring node; Each monitoring node is configured to collect system data of the distributed network system in a region to which the monitoring node belongs according to the monitoring task and upload the system data to the center node; The center node is further configured to analyze the availability of the distributed network system according to the system data uploaded by each monitoring node, determine whether the distributed network system is abnormal, and push alarm information to an operation and maintenance team of the distributed network system when it is determined that the distributed network system is abnormal, so that the operation and maintenance team performs operation and maintenance processing according to the alarm information; The system data comprises system performance indicators, system logs, and user access data, wherein the system performance indicators comprise at least one of a response time, a throughput, and an error rate, and the user access data comprises at least one of an access frequency, a page dwell time, and a bounce rate; The center node is specifically configured to perform an abnormality discovery operation on the system logs by using a text mining method or a pattern matching method, determine that the distributed network system is abnormal if it is found that the system logs are abnormal, and generate a running status evaluation report of the distributed network system according to a level and an abnormality degree of the system logs; The center node is specifically configured to calculate a statistical quantity of the system performance indicators according to the system performance indicators in a preset time period, determine a performance state of the distributed network system according to the statistical quantity, and determine whether the distributed network system is abnormal according to the performance state, wherein the statistical quantity comprises at least one of a mean value, a variance, and a maximum value; The center node is further specifically configured to analyze an experience of users using the distributed network system according to the user access data, and analyze potential problems of the users in the process of using the distributed network system according to the experience; Each monitoring node is further configured to obtain historical system data, pre-process the historical system data, input the pre-processed historical system data into a machine learning model, obtain a prediction result of whether the distributed network system will have a fault or a performance bottleneck in a future preset time period, and push the prediction result to the operation and maintenance team, so that the operation and maintenance team formulates maintenance measures to reduce the possibility of the distributed network system having a fault or a performance bottleneck in the future preset time period.

2. The system of claim 1, wherein, The center node is further configured to pre-configure an alarm strategy, wherein the alarm strategy comprises an abnormality analysis algorithm, a pushing manner of alarm information, and an alarm information receiving address of the operation and maintenance team under the pushing manner; wherein the abnormality analysis algorithm is used to analyze whether the distributed network system is abnormal according to the system data; and the pushing manner of alarm information comprises at least one of short message, email, WeChat and telephone. Correspondingly, the center node is specifically configured to analyze the system data by using the abnormality analysis algorithm to obtain an analysis result of whether the distributed network system is abnormal; and if the analysis result is abnormal, push the alarm information to the operation and maintenance team according to the pushing manner and the alarm information receiving address in the alarm strategy.

3. The system of claim 1, wherein, The center node maintains metadata information of each monitoring node, wherein the metadata information comprises a region, a server address and node information of each monitoring node; Correspondingly, each monitoring node is further configured to report heartbeat data to the center node every preset time length; and the center node is further configured to determine whether the heartbeat data from each monitoring node is received every preset time length according to the maintained metadata information of each monitoring node to obtain a determination result; and determine whether each monitoring node is in a working state according to the determination result to send the monitoring task to the monitoring node in the working state.

4. The system of claim 1, wherein, The center node is further configured to generate a visual chart reflecting the running state of the distributed network system according to the system data reported by each monitoring node, and display the visual chart on a page. And / or, the center node is specifically configured to analyze the system data uploaded by each monitoring node by using a regression analysis method, a classification algorithm or a clustering analysis method to obtain an analysis result of the availability of the distributed network system.

5. A method of availability monitoring of a distributed network system, characterized in that, The method is implemented based on the availability monitoring system of any one of claims 1-4, and the method comprises: configuring a monitoring task by the center node and distributing the monitoring task to the plurality of monitoring nodes; wherein the monitoring task comprises a monitoring frequency, an entry address of the distributed network system in each region, and an address of each monitoring node; collecting, by each monitoring node, system data of the distributed network system in the region to which the monitoring node belongs according to the monitoring task, and uploading the system data to the center node; analyzing, by the center node, the availability of the distributed network system according to the system data uploaded by each monitoring node, determining whether the distributed network system is abnormal, and pushing alarm information to an operation and maintenance team of the distributed network system when it is determined that the distributed network system is abnormal, so that the operation and maintenance team performs operation and maintenance processing according to the alarm information.

Citation Information

Patent Citations

  • Network anomaly monitoring method and device

    CN106254153A

  • Large-scale mixed heterogeneous storage system-oriented node fault prediction system and method

    CN108415789A