Operation and maintenance alarm method and device, nonvolatile storage medium and electronic equipment

By comparing the system monitoring data in the timing database and using expression tree analysis, the high computing power consumption and delay problems caused by the polling method in the prior art are solved, and more efficient operation and maintenance alarms are achieved.

CN120066906APending Publication Date: 2025-05-30CHINA TELECOM INTELLIGENT NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112273.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, operation and maintenance alarms lead to high consumption of computing power resources through polling, which is under great pressure to database query and has obvious delay problems.

Method used

By comparing the system monitoring data collected at the first moment and the system monitoring data collected at the second moment in the timing database, the data is analyzed using the expression tree to determine whether there are abnormal events, thereby updating the alarm status of the alarm.

Benefits of technology

Reduces the computing power resources required for operation and maintenance alarms, reduces the pressure on database query, and improves operation and maintenance performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066906A_ABST
    Figure CN120066906A_ABST
Patent Text Reader

Abstract

The invention discloses an operation and maintenance alarm method and device, a nonvolatile storage medium and electronic equipment. The method comprises the following steps: in the process of writing first system monitoring data acquired at a first moment into a time sequence database, comparing the first system monitoring data acquired at the first moment with second system monitoring data acquired at a second moment; under the condition that the first system monitoring data and the second system monitoring data are inconsistent, according to the data type of the first system monitoring data, determining an expression tree corresponding to the data type; analyzing the first system monitoring data by adopting the expression tree to obtain an analysis result; and updating the alarm state of the alarm according to the data information of the first system monitoring data and the judgment result of the root node under the condition of determining that the abnormal event exists. According to the method and the device, the technical problem of relatively low operation and maintenance performance caused by operation and maintenance alarm in a polling mode in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of operation and maintenance, and more particularly, to an operation and maintenance alarm method, apparatus, non-volatile storage medium, and electronic device. Background Art

[0002] In the related art, when performing operation and maintenance alarms on a system, a polling method is usually adopted. The alarm device regularly obtains relevant data at a preset period and determines whether to trigger an alarm event according to a preset alarm rule. The problem with this method is that it consumes a large amount of computing resources, has a large query pressure on the database, and there are obvious latency problems.

[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide an operation and maintenance alarm method, apparatus, non-volatile storage medium, and electronic device to at least solve the technical problem of low operation and maintenance performance caused by using the polling method for operation and maintenance alarms in the related art.

[0005] According to one aspect of the embodiments of this application, an operation and maintenance alarm method is provided, including: during the process of writing the first system monitoring data collected at the first moment into the time series database, comparing the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; in the case where the first system monitoring data and the second system monitoring data are inconsistent, determining an expression tree corresponding to the data type according to the data type of the first system monitoring data, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a preset calculation operation on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation result of the intermediate node; analyzing the first system monitoring data using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; in the case where it is determined that there is an abnormal event, updating the alarm status of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node.

[0006] Optionally, analyzing the first system monitoring data using the expression tree to obtain an analysis result includes: the execution levels between the intermediate nodes in the expression tree, where the execution levels are used to determine the input-output relationships between the intermediate nodes; calling each intermediate node according to the execution levels to perform a preset calculation operation on the first system monitoring data, and determining the analysis result according to the calculation result of the preset calculation operation.

[0007] Optionally, each intermediate node is called according to the execution level to perform a predefined calculation operation on the first system monitoring data, and the analysis result is determined according to the calculation result of the predefined calculation operation, including: determining a first target node, and determining whether there is a second target node, where the first target node is any intermediate node among the intermediate nodes, and the second target node is an intermediate node at the same execution level as the first target node; in the case where there is no second target node, determining whether the calculation result of the first target node is consistent with the corresponding historical calculation result of the first target node, and in the case of consistency, determining that the analysis result is that there is no abnormal event, and in the case of determining inconsistency, determining a third target node, and inputting the calculation result of the first target node into the third target node, where the input of the third target node is the output of the second target node, and the third target node is an intermediate node or a root node; in the case where there is a second target node, determining whether the calculation result of the first target node is consistent with the corresponding historical calculation result of the first target node, and determining whether the calculation result of the second target node is consistent with the corresponding historical calculation result of the second target node, and in the case where the judgment results corresponding to the first target node and the second target node are both consistent, determining that the analysis result is that there is no abnormal event, and in the case where there is inconsistency in the judgment results corresponding to the first target node and the second target node, determining a fourth target node, and inputting the calculation results of the first target node and the second target node into the fourth target node, where the input of the fourth target node is the output of the second target node, and the fourth target node is an intermediate node or a root node.

[0008] Optionally, in the process of writing the first system monitoring data collected at the first moment into the time series database, comparing the first system monitoring data collected at the first moment and the second system monitoring data collected at the second moment includes: obtaining the first system monitoring data through a collector, where the first system monitoring data includes native metric data and downsampled metric data summary; filtering the first system monitoring data successively by a first filter and a second filter, and writing the obtained first filtering result into the time series database, where the first filter includes a Bloom filter and the second filter includes a tag set filter; comparing the first filtering result and the second filtering result obtained by filtering the second system monitoring data.

[0009] Optionally, the operation and maintenance alarm method further includes: updating the expression tree, and updating the expression tree cache of the expression tree according to the update result of the expression tree; updating the first filter and the second filter according to the update result of the expression tree cache.

[0010] Optionally, the expression tree is determined by the following method: obtaining a predefined alarm rule; determining the alarm rule expression of the predefined alarm rule; splitting the alarm rule expression into multiple expression nodes, where the expression nodes are intermediate nodes.

[0011] Optionally, the operation and maintenance alarm method further includes: after the predefined alarm rule is updated, updating the expression tree corresponding to the predefined alarm rule.

[0012] According to another aspect of the embodiments of the present application, there is also provided an operation and maintenance alarm system, including: a time series database, a metric listener, an alarm rule parser, a rule execution engine, and an alarm device. Among them, the alarm rule parser is used to generate an expression tree according to the predefined alarm rule, where the expression tree includes leaf nodes, intermediate nodes, and root nodes. The leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform predefined calculation operations on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes; the metric listener is used to compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment during the process of writing the first system monitoring data collected at the first moment into the time series database, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; the rule execution engine is used to determine the expression tree corresponding to the data type according to the data type of the first system monitoring data in the case where the first system monitoring data and the second system monitoring data are inconsistent; analyzing the first system monitoring data by using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; the alarm device is used to update the alarm state of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node in the case where it is determined that there is an abnormal event.

[0013] According to another aspect of the embodiments of the present application, an operation and maintenance warning device is further provided, including: a first processing module, configured to compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment during the process of writing the first system monitoring data collected at the first moment into the time series database, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; a second processing module, configured to determine an expression tree corresponding to the data type according to the data type of the first system monitoring data in the case where the first system monitoring data and the second system monitoring data are inconsistent, where the expression tree includes leaf nodes, intermediate nodes, and a root node, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a pre-designed calculation operation on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation result of the intermediate nodes; a third processing module, configured to analyze the first system monitoring data by using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; a fourth processing module, configured to update the warning state of the warning device according to the data information of the first system monitoring data and the judgment result of the root node in the case where it is determined that there is an abnormal event.

[0014] According to another aspect of the embodiments of the present application, a non-volatile storage medium is further provided. A program is stored in the non-volatile storage medium. When the program runs, it controls the device where the non-volatile storage medium is located to execute the operation and maintenance warning method.

[0015] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory and a processor, where the processor is configured to run the program stored in the memory. When the program runs, it executes the operation and maintenance warning method.

[0016] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, where the computer program implements the operation and maintenance warning method when executed by a processor.

[0017] In the embodiment of the present application, during the process of writing the first system monitoring data collected at the first moment into the time series database, the first system monitoring data collected at the first moment is compared with the second system monitoring data collected at the second moment, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; in the case where the first system monitoring data and the second system monitoring data are inconsistent, an expression tree corresponding to the data type is determined according to the data type of the first system monitoring data, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a pre-designed calculation operation on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation result of the intermediate node; the expression tree is used to analyze the first system monitoring data to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; in the case where it is determined that there is an abnormal event, the alarm status of the alarm device is updated according to the data information of the first system monitoring data and the judgment result of the root node. By determining whether there is an alarm event and whether the alarm status needs to be updated through the expression tree when it is determined that there is a change in the system monitoring data, the purpose of reducing the computing resources required for operation and maintenance alarms and reducing the query pressure on the database is achieved, thereby achieving the technical effect of improving the operation and maintenance performance, and further solving the technical problem of low operation and maintenance performance caused by using the polling method for operation and maintenance alarms in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0019] Figure 1 is a schematic structural diagram of an operation and maintenance alarm system provided according to an embodiment of the present application;

[0020] Figure 2 is a schematic diagram of an alarm logic architecture provided according to an embodiment of the present application;

[0021] Figure 3 is a schematic flowchart of an operation and maintenance alarm method provided according to an embodiment of the present application;

[0022] Figure 4 is a schematic flowchart of an operation and maintenance alarm initialization process provided according to an embodiment of the present application;

[0023] Figure 5 is a schematic flowchart of an operation and maintenance alarm process provided according to an embodiment of the present application;

[0024] Figure 6It is a schematic structural diagram of an operation and maintenance warning device provided according to an embodiment of the present application;

[0025] Figure 7 It is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. Specific embodiments

[0026] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0029] Metrics: Refers to a certain measurement value of the system at a specific time point, such as CPU usage rate or the number of requests, etc. These metrics are distinguished by a tag set. In addition, the metric name is a special tag (__name__) in the tag set.

[0030] Alert: An abnormal event generated by the monitoring system after the system state lasts for a period of time, which can help users discover and respond to potential problems in the system in a timely manner. It mainly includes an alert name (Name), an alert description (Description), and an alert summary (Summary). From the perspective of the mainstream monitoring system, an alert is also a special metric (Metrics).

[0031] A time-series database (TSDB) is a database management system specifically designed to store, process, and analyze time-series data. This type of data is arranged in chronological order and indexed by the time dimension, usually consisting of multiple data series.

[0032] Metric Exporter: Its main function is to collect metrics from various data sources and format them into a format suitable for storage in a time-series database (TSDB), and provide data externally through an Http port.

[0033] The main function of the Metric Scraper is to periodically extract raw metric data from the Metric Exporter and, after performing basic data processing procedures (including but not limited to data format conversion, data filtering, and data aggregation), store the processed data in the time-series database.

[0034] A LabelSet is a collection of key / value pairs used to describe a time series. It consists of a series of labels, which give rich context information to the metrics.

[0035] Series: The organizational model of metric data. The data is arranged indexed by timestamps and each series is usually identified by a combination of a timestamp and a LabelSet.

[0036] DownSampling / Recording: By reducing the resolution and volume of metric data within a specific time period, it enables efficient data storage and optimization. In time-series databases and other analysis fields, downsampling techniques help improve system performance, reduce storage costs, and enhance processing efficiency. Usually, the downsampled data needs to be recorded and persisted in the time-series database (TSDB).

[0037] Alert Rule: Alert rules are the key elements in a monitoring system that determine when to issue an alert. It is based on specific conditions and threshold settings to monitor the health and performance of the system. Alert rules mainly include an Alert Name, a trigger condition or expression, and a duration / evaluation time. These rules must follow the standards of Prometheus alert rules.

[0038] Alert Name: The identifying name of an alert rule.

[0039] Trigger condition: A specific expression used to judge the status of the current data series (Series).

[0040] Duration / Evaluation time: The minimum time threshold required to trigger an alarm.

[0041] Alarm label set (Labelset): A set of labels used to enhance the context information of the alarm.

[0042] Alarm state machine: Used to control and track the transition of the alarm state.

[0043] Pending: When the evaluation result of the alarm rule expression is positive but the duration requirement is not yet met, the alarm state is in the pending state to avoid false alarms due to short-term fluctuations.

[0044] Firing: When the alarm state is continuously evaluated as positive and exceeds the set duration, the alarm state will change to the firing state and the alarm information will be passed to the downstream system.

[0045] Resolved: When the alarm state is evaluated as negative, the alarm state will change to the resolved state and the resolved alarm information will be passed to the downstream system.

[0046] Suppressed: Under certain specific conditions, such as within a specific time period, alarms for specific components may be suppressed, and in this case, the alarm information will not be passed to the downstream system.

[0047] Currently, the software infrastructure is undergoing profound changes, mainly manifested in the following aspects: First, the operating environment of all software has shifted to the cloud or edge cloud. Second, the trend of microservices is obvious, that is, the traditional monolithic software architecture is decomposed into multiple microservices. Third, container cloud technology has been widely promoted, especially container clouds based on technologies such as Kubernetes and Mesos Marathon. In addition, the speed of software deployment and release has increased significantly. These changes have put forward higher requirements for the overall observability of the system.

[0048] The alarm function plays a crucial role in the field of observability. Its main purpose is to timely identify and respond to abnormal conditions in the system. It is an essential part of ensuring system observability. Currently, all alarm mechanisms are based on continuous monitoring of the data in the time series database (TSDB), and judge whether to trigger an alarm by continuously evaluating whether the result of the alarm rule expression is positive. Once the evaluation result continuously shows a positive state for a period of time, the system will trigger an alarm and pass the relevant information to the personnel or system responsible for listening.

[0049] Generally, the construction of an alarm system requires the collaborative work of the following key components:

[0050] Data collector: Its main responsibility is to capture target metrics and transmit the data to the Time Series Database (TSDB).

[0051] Time Series Database (TSDB): This database is based on a time series data model and provides support for the storage and retrieval of metric data and alarm status.

[0052] Alerts: This component evaluates by querying the metric data in the TSDB according to the preset rule expressions. If the metric status is positive, it constructs a tag set and metadata based on this status, generates an alarm message, and pushes it to the corresponding notification channels.

[0053] Notification channels: Responsible for receiving and processing alarm messages, with multiple functions such as alarm grouping, suppression, routing, notification template customization, recording, and integration.

[0054] The above components together constitute the basic architecture of the alarm system. The data collector is responsible for collecting metric data and storing it in the time series database, and the alerts are responsible for evaluating the metric data and generating alarms, and finally passing the alarm messages to users through the notification channels. This process constitutes the core function of the alarm system and represents the mainstream architecture of current alarm generation.

[0055] In the current mode, the alerts will actively poll the TSDB database at the set polling interval to evaluate the results of the alarm rule expressions. If the expression result is positive, it will drive the state transition of the alarm state machine, thus triggering the generation of an alarm. The implementation of this design is relatively simple and can handle a certain number of alarm rules. However, this timed and passive processing method has obvious defects. Since it is necessary to periodically traverse the rules and poll the Time Series Database (TSDB), this will cause a large query pressure on the TSDB. As the number of rules increases, the query pressure grows exponentially, which is particularly evident in terms of CPU, memory, and disk I / O resource consumption. In addition, this approach will also lead to resource competition for the TSDB business query service. Due to the nature of timed polling, there is a delay in the alarm response time, and the theoretical maximum delay is twice the polling interval plus the execution queue processing time. In actual operation, if the execution queue becomes congested, the maximum delay will become quite significant. Therefore, in situations where the alarm response time requirement is extremely strict (such as second-level response) and in cases where system resources are limited, this mode cannot meet the system performance requirements.

[0056] In summary, the monitoring and alerting system is usually deployed synchronously with the business system. As a secondary system, its resource requirements are usually limited, that is, its system priority is relatively low. Given that the polling mechanism for alert triggering increases the resource requirements as the scale of alerts and the system expands, resulting in a gradually increasing alert delay, which cannot meet strict alert requirements, such as the ability to trigger alerts within seconds and strict budget limits on computing resource consumption. In addition, as the business pressure increases, the resources of low-priority tasks in the monitoring system may be snatched by other tasks, which may cause congestion when alert rules are executed, resulting in unforeseen delays and preventing the monitoring from operating ideally.

[0057] To solve the above problems, relevant solutions are provided in the embodiments of the present application, and a reactive-based alert mechanism is provided, which will be described in detail below.

[0058] According to the embodiments of the present application, a method embodiment of an operation and maintenance alert method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0059] The method embodiments provided by the embodiments of the present application can be executed in an operation and maintenance alert system as Figure 1 shown. It can be seen from Figure 1 that the operation and maintenance alert system includes: a time series database 10, a metric listener 12, an alert rule parser 14, a rule execution engine 16, and an alertor 18, where:

[0060] An alarm rule parser 14 is used to generate an expression tree according to predefined alarm rules. The expression tree includes leaf nodes, intermediate nodes, and a root node. The leaf nodes are used to indicate the data types corresponding to the expression tree. The intermediate nodes are used to perform predefined calculation operations on the first system monitoring data. The root node is used to determine whether there is an abnormal event based on the calculation results of the intermediate nodes. An index listener 12 is used to compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment during the process of writing the first system monitoring data collected at the first moment into the time series database 10. The second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same. A rule execution engine 16 is used to determine an expression tree corresponding to the data type according to the data type of the first system monitoring data in the case where the first system monitoring data is inconsistent with the second system monitoring data. The first system monitoring data is analyzed using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event. An alarm device 18 is used to update the alarm state of the alarm device 18 according to the data information of the first system monitoring data and the judgment result of the root node in the case where it is determined that there is an abnormal event.

[0061] In some embodiments of the present application, the alarm logic of the operation and maintenance alarm system is as Figure 2 shown. As can be seen from Figure 2 , the alarm rule parser 14 follows the syntax rules of PromQL and is responsible for parsing the alarm rules into a syntax tree composed of metrics, functions, and expressions. Each alarm rule is converted into an independent parse tree and stored in the memory cache pool. The index listener 12 is responsible for configuring the component for monitoring the changes of the time series database 10 (tsdb) metrics. This component obtains the leaf nodes of the parse tree from the alarm rule parser 14, and these leaf nodes include atomic metrics and downsampled metrics. When the corresponding atomic metrics and downsampled metrics are inserted into the tsdb, the index listener 12 will trigger the hierarchical calculation of the parse tree. In addition, the index listener 12 will perform matching according to the Labelset of the metrics to determine whether to trigger the execution of the corresponding rules.

[0062] After receiving the event notification sent by the index listener 12, the rule execution engine 16 will select the corresponding parse tree from the alarm parse tree pool and perform hierarchical parsing from bottom to top. During the parsing process, the engine will query the time series database (TSDB) as needed to obtain data and cache the query results. Once the rule execution engine 16 completes the calculation of the expression, it will send an event notification to the alarm device 18.

[0063] The alarm device is responsible for maintaining and periodically persisting the alarm status to the time series database (TSDB). It is also responsible for receiving event notifications from the rule execution engine 16 and evaluating whether the current alarm state machine has undergone a state transition due to the event notifications from the rule execution engine 16. State transitions may include changing from an inactive state to a pending state, from a pending state to a firing state, and from a firing state to a resolved state, etc.

[0064] Under the above operating environment, an embodiment of the present application provides an operation and maintenance alarm method, as Figure 3 shown, the method includes the following steps:

[0065] Step S302, during the process of writing the first system monitoring data collected at the first moment into the time series database, compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same;

[0066] In the technical solution provided in step S302, the step of comparing the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment during the process of writing the first system monitoring data collected at the first moment into the time series database includes: obtaining the first system monitoring data through a collector, where the first system monitoring data includes native metric data and downsampled metric data summaries; filtering the first system monitoring data sequentially using a first filter and a second filter, and writing the obtained first filtering result into the time series database, where the first filter includes a Bloom filter and the second filter includes a tag set filter; comparing the first filtering result with the second filtering result obtained by filtering the second system monitoring data. The above data types include data types such as CPU load rate and cache occupancy rate that can reflect the system operation status. The above second moment may be the nearest historical collection moment to the first moment.

[0067] In some embodiments of the present application, after the time series database (i.e., the above-mentioned time series database) is started, the built-in listening manager will be activated. According to the configuration of the TSDB, the listening manager initializes the gRPC metric listening interface and the notification interface. When the TSDB performs the write-ahead log (WAL) operation, it will monitor the metric data and its downsampling results. The specific operations include configuring two filtering mechanisms: First, the original metrics to be monitored and the downsampling summary are input into the Bloom filter to achieve fast screening; Second, a label set matching filter is set as an auxiliary filtering mechanism to reduce the possible misjudgments of the Bloom filter. After passing the matching checks of these two filters, the qualified metric data will trigger the operation of writing to the notification queue.

[0068] As an alternative implementation, the downsampling summary includes the number information of the downsampled data, where the number information is used to indicate the position of the data in the time series database.

[0069] In some embodiments of the present application, the metric collector can write the data metrics (compliant with the Prometheus Metrics standard) collected from the exporter into the TSDB through the standard prometheus remotewrite protocol. Then, during the process of processing the Prometheus RemoteWrite write request, the time series database (TSDB) is responsible for parsing and deserializing these requests, and converting them into metric series. Subsequently, these series will be recorded in the write-ahead log (WAL). At this stage, the TSDB has pre-set filters to ensure that the metric series go through the processing of two layers of filters. This process will trigger the detection of monitoring metric changes, and notify the change information to the execution engine through the gRPC interface to wake it up for execution.

[0070] Step S304, in the case where the first system monitoring data is inconsistent with the second system monitoring data, determine the expression tree corresponding to the data type according to the data type of the first system monitoring data, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, and the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform pre-designed calculation operations on the first system monitoring data, and the root nodes are used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes;

[0071] In the technical solution provided in step S304, the expression tree is determined in the following manner: obtain the predefined alarm rules; determine the alarm rule expression of the predefined alarm rules; split the alarm rule expression into multiple expression nodes, where the expression nodes are intermediate nodes.

[0072] As an alternative implementation, the operation and maintenance alerting method further includes: after the predefined alerting rules are updated, updating the expression tree corresponding to the predefined alerting rules.

[0073] Specifically, an alerting rule parser can be started. The alerting rule parser processes the predefined alerting rules (PromQL) through lexical analysis and syntactic analysis, and disassembles the alerting rule expressions into multiple expression nodes. These intermediate nodes cover structures such as function calls and range queries, while the leaf nodes correspond to downsampling (Recording) and atomic metrics. The parsed expression tree is bound to the alerting rules and stored in the cache. In addition, the alerting rule parser uses the monitoring file mechanism and interface notifications to Reload for reloading, continuously monitoring changes to the alerting configuration file. Once a change in the configuration file is detected, the parser will trigger a re-parsing process and synchronously update the cache content.

[0074] In some embodiments of the present application, the operation and maintenance alerting method further includes: updating the expression tree, and updating the expression tree cache of the expression tree according to the update result of the expression tree; updating the first filter and the second filter according to the update result of the expression tree cache.

[0075] As an alternative implementation, a metric listener can also be started to perform a listening setting operation based on the leaf nodes parsed from the alerting rules. Specifically, two filters are set at the TSDB write layer using the listening setting interface. The raw metrics / downsampled summaries to be listened to are sent to the Bloom filter (mainly for fast filtering), and at the same time, a Labelset matching filter is also set as a secondary filter to perform a matching check on the raw metrics and the downsampled label set (Labelset). Once the filters match, a write notification is triggered. In addition, the listener will establish an association with the alerting rule expression tree in the cache and synchronously update the expression tree cache. Once the expression tree cache changes, both layers of filters will be updated. When a write change occurs to the metric or its downsampled data being listened to, the system will locate the expression tree through the summary hash value of the metric or downsampled label set and trigger the execution engine to perform related operations.

[0076] Step S306, analyzing the first system monitoring data using the expression tree to obtain an analysis result, where the analysis result is used to determine whether an abnormal event exists;

[0077] In the technical solution provided in step S306, the steps of analyzing the first system monitoring data by using an expression tree to obtain an analysis result include: determining the execution levels among the intermediate nodes in the expression tree, where the execution levels are used to determine the input-output relationships among the intermediate nodes; calling each intermediate node according to the execution levels to perform a preset calculation operation on the first system monitoring data, and determining the analysis result according to the calculation result of the preset calculation operation.

[0078] As an optional implementation manner, the steps of calling each intermediate node according to the execution levels to perform a preset calculation operation on the first system monitoring data, and determining the analysis result according to the calculation result of the preset calculation operation include: determining a first target node and determining whether there is a second target node, where the first target node is any intermediate node among the intermediate nodes, and the second target node is an intermediate node at the same execution level as the first target node; in the case where there is no second target node, determining whether the calculation result of the first target node is consistent with the corresponding historical calculation result of the first target node, and in the case of consistency, determining that the analysis result is that there is no abnormal event, and in the case of determining inconsistency, determining a third target node and inputting the calculation result of the first target node into the third target node, where the input of the third target node is the output of the second target node, and the third target node is an intermediate node or a root node; in the case where there is a second target node, determining whether the calculation result of the first target node is consistent with the corresponding historical calculation result of the first target node, and determining whether the calculation result of the second target node is consistent with the corresponding historical calculation result of the second target node, and in the case where the judgment results corresponding to the first target node and the second target node are both consistent, determining that the analysis result is that there is no abnormal event, and in the case where there is inconsistency in the judgment results corresponding to the first target node and the second target node, determining a fourth target node and inputting the calculation results of the first target node and the second target node into the fourth target node, where the input of the fourth target node is the output of the second target node, and the fourth target node is an intermediate node or a root node.

[0079] Specifically, after the execution engine is awakened, it will execute the expression trees related to the changed metrics (raw metrics / downsampling). Each expression tree is bound to an execution cache and related alarm rules. By comparing the difference between the current calculation result of the node and the previous result, it can be determined whether to continue executing the expression calculation at the previous level. If the current calculation result is the same as the previous result (including the data comparison of LabelSet and Values), there is no need to continue executing upward, nor to update the alarm status; otherwise, continue to execute upward to the root node of the expression tree to complete the calculation of the entire expression tree. Finally, the sequences with positive execution results are used as the results and sent to the listening interface for alarm triggering to complete the notification process. The sequences with positive execution results include the data types processed by the expression tree, specific monitoring values, calculation results of each level of nodes, and the final judgment result of the root node. A positive execution result means that an abnormal event is finally determined.

[0080] Step S308, in the case of determining that there is an abnormal event, update the alarm status of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node.

[0081] In the technical solution provided in step S308, the alarm device will receive the positive result notification from the execution engine through the alarm status listening interface. This notification includes the positive sequence of the execution result and the timestamp information. The alarm device is responsible for coordinating the conversion of the alarm state machine, including adding an alarm (Pending), upgrading the alarm (Pending) to the active (Active) state, and adjusting the active alarm (Active) to the resolved (Resolved) state, etc.

[0082] According to the embodiments of the present application, there is also provided an initialization process as Figure 4 shown, including the following steps:

[0083] Step S402, initialize the alarm listening interface of the alarm device, the alarm status manager of the alarm device, the TSDB listening interface and the notification interface, and initialize the Bloom filter and the label set filter;

[0084] Step S404, the alarm rule parser parses the alarm expression into an expression tree, and in the case of determining that there is no parsing exception, initialize the execution engine and bind the expression tree and the node cache, and then jump to step S406; if there is a parsing exception, end the initialization process and prompt the parsing exception;

[0085] Step S406, the execution engine calls the TSDB interface to set up listening;

[0086] Step S408, the TSDB updates the Bloom filter and the label set filter;

[0087] Step S410, the execution engine polls the alarm rules for the first time to populate the cache and publish status transition notifications.

[0088] Specifically, after the execution engine is started, it will poll the alarm rules for the first time to populate the internal cache of the expression tree, and use the positive results in the evaluation results as evaluation result transition notifications to the alarm. Subsequently, the alarm will initialize the internal alarm manager, establish corresponding alarm sets and corresponding alarm state machines. After the alarm is started, it will initialize the alarm manager and the alarm state monitoring interface inside the alarm. The alarm manager is responsible for maintaining the state machines of each alarm and regularly synchronizing the alarm status to the time series database. The alarm state monitoring interface will monitor the notifications from the execution engine regarding the evaluation result transitions.

[0089] According to the embodiments of the present application, there is also provided an Figure 5 operation and maintenance alarm process as shown in

[0090] Step S502, the metric collector writes the collected data into the TSDB, and the TSDB filters the written data through two layers of filters;

[0091] Step S504, determine whether to trigger the monitored metric by comparing whether the currently collected data is consistent with the historical data collected most recently, and confirm to trigger the monitored metric when it is determined that they are inconsistent;

[0092] Step S506, after the monitored metric is triggered, wake up the stealth, select the corresponding expression tree, and execute the expressions corresponding to the leaf nodes, intermediate nodes, and root nodes at all levels in the expression tree step by step, and determine whether there is a difference between the calculated values of each level of nodes and the calculation results cached last time;

[0093] Step S508, when there is a difference, send the positive result of the expression to the alarm, and the alarm coordinates the alarm state machine.

[0094] By comparing the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment during the process of writing the first system monitoring data collected at the first moment into the time series database, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; in the case where the first system monitoring data and the second system monitoring data are inconsistent, determining an expression tree corresponding to the data type according to the data type of the first system monitoring data, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a pre-designed calculation operation on the first system monitoring data, and the root nodes are used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes; using the expression tree to analyze the first system monitoring data to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; in the case where it is determined that there is an abnormal event, updating the alarm status of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node, by determining whether there is an alarm event and whether the alarm status needs to be updated through the expression tree when it is determined that there is a change in the system monitoring data, the purpose of reducing the computing power resources required for operation and maintenance alarms and reducing the query pressure on the database is achieved, thereby achieving the technical effect of improving the operation and maintenance performance, and further solving the technical problem of low operation and maintenance performance caused by using the polling method for operation and maintenance alarms in the related art.

[0095] An embodiment of the present application provides an operation and maintenance alarm device. Figure 6 It is a schematic structural diagram of the device. As can be seen from Figure 6 it, the device includes: a first processing module 60, configured to compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment during the process of writing the first system monitoring data collected at the first moment into the time series database, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; a second processing module 62, configured to determine an expression tree corresponding to the data type according to the data type of the first system monitoring data in the case where the first system monitoring data and the second system monitoring data are inconsistent, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a pre-designed calculation operation on the first system monitoring data, and the root nodes are used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes; a third processing module 64, configured to analyze the first system monitoring data using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; a fourth processing module 66, configured to update the alarm status of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node in the case where it is determined that there is an abnormal event.

[0096] In some embodiments of the present application, in the process of writing the first system monitoring data collected at the first moment into the time series database, the steps of comparing the first system monitoring data collected at the first moment and the second system monitoring data collected at the second moment include: obtaining the first system monitoring data through a collector, where the first system monitoring data includes native metric data and downsampled metric data summary; filtering the first system monitoring data sequentially using a first filter and a second filter, and writing the obtained first filtering result into the time series database, where the first filter includes a Bloom filter and the second filter includes a tag set filter; comparing the first filtering result with the second filtering result obtained by filtering the second system monitoring data.

[0097] In some embodiments of the present application, the expression tree is determined in the following manner: obtaining a predefined alarm rule; determining the alarm rule expression of the predefined alarm rule; splitting the alarm rule expression into multiple expression nodes, where the expression nodes are intermediate nodes.

[0098] In some embodiments of the present application, the operation and maintenance alarm device is further configured to: update the expression tree corresponding to the predefined alarm rule after the predefined alarm rule is updated.

[0099] In some embodiments of the present application, the operation and maintenance alarm device is further configured to: update the expression tree, and update the expression tree cache of the expression tree according to the update result of the expression tree; update the first filter and the second filter according to the update result of the expression tree cache.

[0100] In some embodiments of the present application, the steps of the third processing module 64 analyzing the first system monitoring data using the expression tree to obtain an analysis result include: determining the execution levels between the intermediate nodes in the expression tree, where the execution levels are used to determine the input-output relationships between the intermediate nodes; calling each intermediate node to perform a predefined calculation operation on the first system monitoring data according to the execution levels, and determining the analysis result according to the calculation result of the predefined calculation operation.

[0101] In some embodiments of the present application, the steps for the third processing module 64 to call each intermediate node according to the execution level to perform a preset calculation operation on the first system monitoring data and determine the analysis result based on the calculation result of the preset calculation operation include: determining a first target node and determining whether there is a second target node, where the first target node is any intermediate node among the intermediate nodes, and the second target node is an intermediate node at the same execution level as the first target node; in the case where there is no second target node, determining whether the calculation result of the first target node is consistent with the corresponding historical calculation result of the first target node, and in the case of consistency, determining that the analysis result is that there is no abnormal event, and in the case of determining inconsistency, determining a third target node and inputting the calculation result of the first target node into the third target node, where the input of the third target node is the output of the second target node, and the third target node is an intermediate node or a root node; in the case where there is a second target node, determining whether the calculation result of the first target node is consistent with the corresponding historical calculation result of the first target node, and determining whether the calculation result of the second target node is consistent with the corresponding historical calculation result of the second target node, and in the case where the judgment results corresponding to the first target node and the second target node are both consistent, determining that the analysis result is that there is no abnormal event, and in the case where there is inconsistency in the judgment results corresponding to the first target node and the second target node, determining a fourth target node and inputting the calculation results of the first target node and the second target node into the fourth target node, where the input of the fourth target node is the output of the second target node, and the fourth target node is an intermediate node or a root node.

[0102] It should be noted that each module in the above operation and maintenance warning device can be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the presentation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.

[0103] Figure 7 shows a hardware structure block diagram of an electronic device for implementing the operation and maintenance warning method. As Figure 7 shown, the electronic device 70 may include one or more (shown as 702a, 702b,..., 702n in the figure) processors 702 (the processor 702 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 704 for storing data, and a transmission device 706 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand,Figure 7 The structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 70 may further include more or fewer components than those shown in Figure 7 or have a different configuration from that shown in Figure 7 .

[0104] It should be noted that the above one or more processors 702 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 70 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0105] The memory 704 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the operation and maintenance warning method in the embodiments of the present application. The processor 702 executes various functional applications and data processing by running the software programs and modules stored in the memory 704, that is, implements the above-mentioned operation and maintenance warning method. The memory 704 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 704 may further include a memory remotely provided with respect to the processor 702, and these remote memories can be connected to the electronic device 70 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0106] The transmission device 706 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the electronic device 70. In one instance, the transmission device 706 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 706 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0107] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the electronic device 70.

[0108] According to an embodiment of the present application, there is also provided a non-volatile storage medium storing a program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the following operation and maintenance warning method: During the process of writing the first system monitoring data collected at the first moment into the time series database, compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; in the case where the first system monitoring data and the second system monitoring data are inconsistent, determine an expression tree corresponding to the data type according to the data type of the first system monitoring data, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a pre-designed calculation operation on the first system monitoring data, and the root nodes are used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes; analyze the first system monitoring data using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; in the case where it is determined that there is an abnormal event, update the warning status of the warning device according to the data information of the first system monitoring data and the judgment result of the root node.

[0109] According to an embodiment of the present application, there is also provided a computer program product including a computer program, which when executed by a processor, implements the following operation and maintenance warning method: During the process of writing the first system monitoring data collected at the first moment into the time series database, compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment, where the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; in the case where the first system monitoring data and the second system monitoring data are inconsistent, determine an expression tree corresponding to the data type according to the data type of the first system monitoring data, where the expression tree includes leaf nodes, intermediate nodes, and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a pre-designed calculation operation on the first system monitoring data, and the root nodes are used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes; analyze the first system monitoring data using the expression tree to obtain an analysis result, where the analysis result is used to determine whether there is an abnormal event; in the case where it is determined that there is an abnormal event, update the warning status of the warning device according to the data information of the first system monitoring data and the judgment result of the root node.

[0110] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0111] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0112] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0113] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0114] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical disks and other various media that can store program codes.

[0115] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. An operation and maintenance alarm method, characterized in that: include: In the process of writing the first system monitoring data collected at the first moment into the time series database, comparing the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment, wherein the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; In the case that the first system monitoring data and the second system monitoring data are inconsistent, an expression tree corresponding to the data type is determined according to the data type of the first system monitoring data, wherein the expression tree includes leaf nodes, intermediate nodes and root nodes, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform preset calculation operations on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation results of the intermediate nodes; Using the expression tree to analyze the first system monitoring data to obtain an analysis result, wherein the analysis result is used to determine whether the abnormal event exists; When it is determined that the abnormal event exists, the alarm status of the alarm device is updated according to the data information of the first system monitoring data and the judgment result of the root node.

2. The operation and maintenance alarm method according to claim 1, characterized in that: The expression tree is used to analyze the first system monitoring data, and the analysis results obtained include: Determine an execution level between each of the intermediate nodes in the expression tree, wherein the execution level is used to determine an input-output relationship between each of the intermediate nodes; Each of the intermediate nodes is called according to the execution level to perform the preset calculation operation on the first system monitoring data, and the analysis result is determined according to the calculation result of the preset calculation operation.

3. The operation and maintenance alarm method according to claim 2, characterized in that: Calling each of the intermediate nodes to perform the preset calculation operation on the first system monitoring data according to the execution level, and determining the analysis result according to the calculation result of the preset calculation operation includes: Determine a first target node, and determine whether a second target node exists, wherein the first target node is any intermediate node among the intermediate nodes, and the second target node is an intermediate node at the same execution level as the first target node; In the case where the second target node does not exist, determining whether the calculation result of the first target node is consistent with the historical calculation result corresponding to the first target node, and in the case of consistency, determining that the analysis result is that the abnormal event does not exist, and in the case of inconsistency, determining a third target node, and inputting the calculation result of the first target node into the third target node, wherein the input of the third target node is the output of the second target node, and the third target node is an intermediate node or a root node; In the case where the second target node exists, determine whether the calculation result of the first target node is consistent with the historical calculation result corresponding to the first target node, and determine whether the calculation result of the second target node is consistent with the historical calculation result corresponding to the second target node, and when the judgment results corresponding to the first target node and the second target node are consistent, determine that the analysis result is that the abnormal event does not exist, and when there is an inconsistency in the judgment results corresponding to the first target node and the second target node, determine the fourth target node, and input the calculation results of the first target node and the second target node into the fourth target node, wherein the input of the fourth target node is the output of the second target node, and the fourth target node is an intermediate node or a root node.

4. The operation and maintenance alarm method according to claim 1, characterized in that: In the process of writing the first system monitoring data collected at the first moment into the time series database, comparing the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment includes: Acquire the first system monitoring data through a collector, wherein the first system monitoring data includes native indicator data and a summary of downsampled indicator data; Using a first filter and a second filter to filter the first system monitoring data in sequence, and writing a first filtering result obtained by filtering into the time series database, wherein the first filter includes a Bloom filter and the second filter includes a label set filter; Compare the first filtering result with a second filtering result obtained by filtering the second system monitoring data.

5. The operation and maintenance alarm method according to claim 4, characterized in that: The operation and maintenance alarm method also includes: Updating the expression tree, and updating the expression tree cache of the expression tree according to the update result of the expression tree; The first filter and the second filter are updated according to the update result of the expression tree cache.

6. The operation and maintenance alarm method according to claim 1, characterized in that: The expression tree is determined in the following way: Get predefined alarm rules; Determining an alarm rule expression of the predefined alarm rule; The alarm rule expression is split into multiple expression nodes, wherein the expression nodes are the intermediate nodes.

7. The operation and maintenance alarm method according to claim 6, characterized in that: The operation and maintenance alarm method also includes: After the predefined alarm rule is updated, the expression tree corresponding to the predefined alarm rule is updated.

8. An operation and maintenance alarm system, characterized in that: It includes time series database, indicator listener, alarm rule parser, rule execution engine, and alarm. The alarm rule parser is used to generate an expression tree according to a predefined alarm rule, wherein the expression tree includes leaf nodes, intermediate nodes and a root node, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a preset calculation operation on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation result of the intermediate node; The indicator listener is used to compare the first system monitoring data collected at the first moment with the second system monitoring data collected at the second moment in the process of writing the first system monitoring data collected at the first moment into the time series database, wherein the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; The rule execution engine is used to determine the expression tree corresponding to the data type according to the data type of the first system monitoring data when the first system monitoring data and the second system monitoring data are inconsistent; use the expression tree to analyze the first system monitoring data to obtain an analysis result, wherein the analysis result is used to determine whether the abnormal event exists; The alarm device is used to update the alarm status of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node when it is determined that the abnormal event exists.

9. An operation and maintenance alarm device, characterized in that: include: A first processing module, configured to compare the first system monitoring data collected at the first moment with the second system monitoring data collected at a second moment in a process of writing the first system monitoring data collected at the first moment into a time series database, wherein the second moment is a moment before the first moment, and the data types of the second system monitoring data and the first system monitoring data are the same; a second processing module, configured to determine, in the case where the first system monitoring data and the second system monitoring data are inconsistent, an expression tree corresponding to the data type according to the data type of the first system monitoring data, wherein the expression tree includes leaf nodes, intermediate nodes and a root node, the leaf nodes are used to indicate the data type corresponding to the expression tree, the intermediate nodes are used to perform a preset calculation operation on the first system monitoring data, and the root node is used to determine whether there is an abnormal event according to the calculation result of the intermediate node; A third processing module, configured to analyze the first system monitoring data using the expression tree to obtain an analysis result, wherein the analysis result is used to determine whether the abnormal event exists; The fourth processing module is used to update the alarm status of the alarm device according to the data information of the first system monitoring data and the judgment result of the root node when it is determined that the abnormal event exists.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the operation and maintenance alarm method according to any one of claims 1 to 7.

11. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the operation and maintenance alarm method described in any one of claims 1 to 7 is executed when the program is run.

12. A computer program product, characterized in that It comprises a computer program, which, when executed by a processor, implements the operation and maintenance alarm method according to any one of claims 1 to 7.