Intelligent operation and maintenance monitoring alarm method and system

Through multi-dimensional data fusion and dynamic adaptive models, combined with device node network diagrams for intelligent anomaly judgment, the problem of neglected interaction relationships and topological structures between devices in existing technologies is solved, and high-precision real-time monitoring and troubleshooting are achieved.

CN120639591AActive Publication Date: 2025-09-12NEWLIXON TECH CO LTD

Patent Information

Application Number
CN202511135596.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-09-12
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

Existing technologies have weak abnormality diagnosis capabilities in complex scenarios and ignore the interactive relationships between devices and network topology, resulting in low monitoring accuracy, untimely response and difficulty in troubleshooting.

Method used

It uses multi-dimensional data fusion, label association analysis and dynamic adaptive models to perform intelligent anomaly judgment, conducts status assessment based on the device node network diagram, generates adaptive alarm thresholds, and detects and locates anomalies through multi-dimensional dynamic labels and elastic rules.

Benefits of technology

It achieves accurate perception of equipment operating status and real-time anomaly identification, improves the intelligence level and adaptability of operation and maintenance monitoring, and enhances the fault tolerance and flexibility of the system in complex abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639591A_ABST
    Figure CN120639591A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent operation and maintenance monitoring alarm method and system, and belongs to the technical field of monitoring alarm, and the method specifically comprises the steps: collecting the real-time communication data and a multi-dimensional dynamic label of each equipment node, the multi-dimensional dynamic label comprises a dynamically generated equipment state label, automatically updating label attributes and life cycles according to real-time communication data, generating self-adaptive alarm thresholds based on label matching dynamic threshold templates and historical data, constructing an equipment node network diagram, performing state evaluation based on the equipment node network diagram, and comparing, analyzing and judging whether the real-time state of each node is abnormal or not. Generating alarm information for the detected abnormal behavior, wherein the alarm information comprises a specific abnormal node, an abnormal data index and an abnormal reason prompt; according to the method, the dynamic threshold template is matched by using the historical data and the dynamic label, the adaptive alarm threshold is automatically generated, the fault node and the affected link are quickly positioned, the abnormal level is judged through real-time comparison and analysis, and detailed alarm information is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of monitoring and alarm technology, and specifically to an intelligent operation and maintenance monitoring and alarm method and system. Background Art

[0002] With the rapid development of technologies such as the Internet, the Internet of Things, and cloud computing, the number of device nodes in modern industrial and information systems has exploded, creating a complex network environment. Traditional operations and maintenance monitoring technologies primarily rely on fixed thresholds or static rules to monitor single indicators, such as data transmission rate, latency, and packet loss rate. This approach suffers from the following drawbacks: 1. Static threshold limitations: Fixed thresholds cannot adapt to the dynamic changes in data fluctuations in network environments; 2. Insufficient single-dimensional monitoring: The ability to diagnose anomalies in complex scenarios is weak; 3. Lack of network topology analysis: Existing monitoring systems often ignore the interactions between devices and the network topology, making it difficult to achieve global anomaly detection; 4. Insufficient adaptability: The lack of an adaptive adjustment mechanism makes it difficult to form an effective real-time monitoring and early warning closed loop.

[0003] For example, a Chinese patent with authorization announcement number CN114328118B discloses an intelligent alarm method, device, equipment and medium for operation and maintenance monitoring data. The method includes: training and correcting a data prediction model based on historical operation and maintenance monitoring data; inputting the operation and maintenance monitoring data of all time periods within several time periods closest to the current time into the data prediction model, and obtaining the operation and maintenance monitoring prediction data of a certain time period within several time periods after the current time; comparing the operation and maintenance monitoring prediction data of a certain time period within several time periods after the current time with the alarm threshold corresponding to the operation and maintenance monitoring data in the time period; if the operation and maintenance monitoring prediction data of a certain time period within a certain time period after the current time is greater than the alarm threshold corresponding to the operation and maintenance monitoring data in the time period, an early warning is issued, thereby effectively improving the accuracy and practicality of the operation and maintenance monitoring data alarm.

[0004] For example, a Chinese patent with authorization announcement number CN111782487B discloses an alarm notification method and device, in which the method applied to the alarm intelligent aggregation and merging engine includes: obtaining a target alarm information set monitored by the alarm monitoring platform within a target time period; determining the target operation and maintenance position corresponding to each alarm information based on the system identifier corresponding to each alarm information in the target alarm information set; aggregating the alarm information corresponding to the same target operation and maintenance position in the target alarm information set to obtain an alarm information set corresponding to each target operation and maintenance position; generating alarm broadcast information corresponding to each target operation and maintenance position, and sending it to the voice call platform, so that the voice call platform notifies the corresponding operation and maintenance personnel of the alarm broadcast information corresponding to each target operation and maintenance position. This application can promptly and quickly notify the operation and maintenance personnel of the alarm broadcast information corresponding to each target operation and maintenance position, and the alarm notification efficiency is higher; and this application will not miss important alarm information.

[0005] The defects of the above technical solution are: weak ability to diagnose abnormalities in complex scenarios, and neglect of the interaction between devices and network topology.

[0006] Therefore, existing technologies have problems such as low monitoring accuracy, untimely response and difficult troubleshooting in large-scale, dynamically changing equipment network environments. There is an urgent need for a new monitoring and alarm method that can utilize multi-dimensional dynamic tags, adaptive thresholds and global network graph analysis. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention proposes an intelligent operation and maintenance monitoring and alarm method and system, which performs intelligent anomaly judgment based on multi-dimensional data fusion, label association analysis and dynamic adaptive models, and issues alarms based on adaptive thresholds.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] Intelligent operation and maintenance monitoring and alarm methods include:

[0010] Collecting real-time communication data and multi-dimensional dynamic tags of each device node, including dynamically generated device status tags, and automatically updating tag attributes and lifecycles based on real-time communication data, where the lifecycle is the lifecycle of the device status tag;

[0011] Generate adaptive alarm thresholds based on tag matching dynamic threshold templates and historical data;

[0012] Build a device node network diagram and evaluate the status of the device node network diagram to determine whether the real-time status of each node is abnormal. When a special anomaly is detected, match the elastic rules based on the multi-dimensional dynamic label combination of the abnormal node;

[0013] Generate alarm information for detected abnormal behaviors, including specific abnormal nodes, abnormal data indicators and abnormal cause prompts.

[0014] Specifically, the tag-based matching dynamic threshold template and the generation of an adaptive alarm threshold in combination with historical data include:

[0015] Creating a threshold template based on the dynamic tag classification of the device, the threshold template including: initial threshold, alarm level, indicator, and indicator weight;

[0016] Automatically match the most suitable threshold template through multi-dimensional dynamic labels. If the multi-dimensional dynamic labels conflict, a weighted priority algorithm is used to select the template.

[0017] Roll statistics on historical data according to the preset time window to calculate the indicator baseline value, that is, the indicator dynamic threshold. Compare and correct the initial threshold of the threshold template with the indicator baseline value to obtain the corrected indicator dynamic threshold.

[0018] Specifically, the device node network diagram is constructed, and based on the status evaluation of the device node network diagram, whether the real-time status of each node is abnormal is determined. When a special abnormality is detected, elastic rules are matched based on the multi-dimensional dynamic label combination of the abnormal node, including:

[0019] Each device is regarded as a node, and the real-time data transmission relationship between devices is regarded as an edge to build a device node network graph;

[0020] Compare the currently collected real-time data with the revised indicator dynamic threshold to detect whether the data exceeds the normal fluctuation range and calculate the degree of deviation;

[0021] Combined with the multi-dimensional dynamic tag information of device nodes, it confirms and locates abnormal situations, evaluates the operating status of nodes and links, and classifies different abnormality levels according to the degree of deviation and evaluation results. The abnormality levels are divided into: special abnormality, level 1 abnormality, level 2 abnormality, level 3 abnormality, and level 4 abnormality;

[0022] When a special anomaly is detected, the elasticity rule is matched based on the multi-dimensional dynamic label combination of the abnormal node, the threshold relaxation coefficient is dynamically calculated, and the elasticity range is limited by combining the business priority label. The above steps are repeated to perform anomaly detection.

[0023] Specifically, the flexible rule includes a flexible rule base, supports multi-label conflict arbitration, and when a multi-label conflict occurs, a final strategy is selected according to label priority weights.

[0024] Specifically, the generation of the multi-dimensional dynamic label includes:

[0025] Dynamically generate a device status tag based on real-time communication data, the device status tag including device load, temporary service type, and abnormal event mark;

[0026] Define a lifecycle for each tag and automatically remove the tag when the trigger condition fails.

[0027] Specifically, the calculation of the indicator baseline value excludes data from the historical alarm period and is dynamically updated through a sliding window algorithm.

[0028] Specifically, the intelligent operation and maintenance monitoring and alarm methods also include:

[0029] Perform differential encoding compression on multi-dimensional dynamic tag data and only transmit the tag changes;

[0030] Complete label classification and initial screening at the edge node, filter the normal label data and then upload it.

[0031] Specifically, the intelligent operation and maintenance monitoring and alarm system is used to implement the intelligent operation and maintenance monitoring and alarm method, including: an acquisition module, a threshold generation module, an abnormality judgment module and an alarm module;

[0032] The acquisition module is used to collect real-time communication data and multi-dimensional dynamic tags of each device node, the multi-dimensional dynamic tags including dynamically generated device status tags, and automatically update tag attributes and life cycle based on real-time communication data, the life cycle being the life cycle of the device status tag;

[0033] The threshold generation module generates an adaptive alarm threshold based on the tag matching dynamic threshold template and historical data;

[0034] The abnormality judgment module is used to construct a device node network diagram, and based on the status evaluation of the device node network diagram, judge whether the real-time status of each node is abnormal. When a special abnormality is detected, it matches the elastic rules according to the multi-dimensional dynamic label combination of the abnormal node;

[0035] The alarm module is used to generate alarm information for detected abnormal behaviors, including specific abnormal nodes, abnormal data indicators and possible fault causes.

[0036] Specifically, the anomaly judgment module includes: a graph construction unit, a deviation calculation unit and an anomaly judgment unit;

[0037] The graph construction unit is used to construct a device node network graph by treating each device as a node and the real-time data transmission relationship between devices as an edge;

[0038] The deviation calculation unit is used to compare the currently collected real-time data with the revised indicator dynamic threshold, detect whether the data exceeds the normal fluctuation range, and calculate the degree of deviation;

[0039] The abnormality determination unit is used to confirm and locate abnormal situations in combination with the multi-dimensional dynamic tag information of the device nodes, evaluate the operating status of the nodes and links, and classify different abnormality levels according to the degree of deviation and the evaluation results.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The present invention proposes an intelligent operation and maintenance monitoring and alarm method. By introducing multi-dimensional dynamic labels, dynamic threshold templates and device node network diagram modeling, it realizes accurate perception of the device operation status and real-time anomaly identification, and has the ability of multi-source information fusion, state adaptation and anomaly level judgment. The method can dynamically generate labels according to the actual operation characteristics of the equipment, and continuously correct the alarm threshold based on historical data, effectively solving the problems of fixed thresholds, frequent false alarms, and lack of upstream and downstream linkage judgment in traditional operation and maintenance systems. At the same time, by constructing a device node network diagram and a flexible alarm strategy, the system's fault tolerance and flexibility in the face of complex abnormal situations are improved, comprehensive monitoring is achieved, and the intelligence level and adaptability of operation and maintenance monitoring are significantly enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Flowchart of the intelligent operation and maintenance monitoring and alarm method provided by the present invention;

[0043] Figure 2 The operation and maintenance monitoring alarm flow chart provided by the present invention;

[0044] Figure 3 The elastic matching flow chart provided by the present invention;

[0045] Figure 4 This is the architecture diagram of the intelligent operation and maintenance monitoring and alarm system provided by the present invention. DETAILED DESCRIPTION

[0046] The present application is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that those skilled in the art may make several variations and improvements without departing from the scope of the present application. These all fall within the scope of protection of the present application.

[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0048] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. In addition, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.

[0049] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.

[0050] Example 1

[0051] See also Figure 1-Figure 3 The present invention provides an embodiment: an intelligent operation and maintenance monitoring and alarm method, comprising the following specific steps:

[0052] Step S1: Collect real-time communication data and multi-dimensional dynamic tags of each device node, the multi-dimensional dynamic tags include dynamically generated device status tags, and automatically update tag attributes and life cycle according to real-time communication data, the life cycle is the life cycle of the device status tag.

[0053] Collect real-time communication data of each device node, and collect data from each device node through embedded software or hardware monitoring modules. The main monitored indicators include but are not limited to transmission rate, delay, packet loss rate, data flow, connection status, etc.

[0054] Data not only comes from the network transmission process, but may also include the operating status of the device itself (such as temperature, load, etc.) and external environmental data (such as geographic location, network topology information). By integrating these multi-dimensional data, a comprehensive view of the device operating status can be constructed.

[0055] Traditional tags are static and cannot reflect real-time device status changes, such as service switching or temporary load surges. Dynamic tags automatically generate or update tag attributes based on real-time device communication data. For example, if a node's CPU usage is high for five consecutive minutes, a high-load temporary tag is automatically added. If the node's data transmission volume suddenly increases tenfold, a suspected DDoS attack event tag is added.

[0056] The dimensions of the multi-dimensional dynamic label include: device type, business level, geographic location, etc., and weights are assigned based on business scenarios. For example, basic attribute labels (such as device type, geographic location) → fixed weight (such as set to 30%), business attribute labels (such as business level, service SLA) → dynamic weight (such as set to 40%-60%), device status labels (such as fault history, real-time load) → temporary weight (such as set to 10%-30%). It should be noted that the weight setting is set by personnel in this field based on actual conditions or a large number of simulation experiments, and the weight values ​​here are only examples.

[0057] Tag lifecycle management: Define the validity period of device status tags. When the validity period expires, the status tags will immediately become invalid. For example, a high-load tag will automatically become invalid after the CPU returns to normal, preventing historical tags from interfering with current analysis.

[0058] Dynamic tags can maintain high resolution capabilities even when facing complex network topologies and massive device nodes.

[0059] The generation of multi-dimensional dynamic labels includes:

[0060] A device status tag is dynamically generated based on real-time communication data, wherein the device status tag includes device load, temporary service type, and abnormal event mark.

[0061] Define a lifecycle for each tag and automatically remove the tag when the trigger condition fails.

[0062] When transmitting multi-dimensional dynamic label data, differential coding compression is required first, and only the label changes are transmitted; label classification and initial screening are completed at the edge node, and normal label data is filtered before uploading.

[0063] Step S2: Generate an adaptive alarm threshold based on the tag matching dynamic threshold template and historical data.

[0064] The specific steps of step S2 are:

[0065] Step S201: creating a threshold template according to the dynamic tag classification of the device, wherein the threshold template includes: an initial threshold, an alarm level, an indicator, and an indicator weight.

[0066] In this embodiment, first, based on the data generated by the equipment during simulation operation, label classification is dynamically generated, that is, the state characteristics of the equipment under a specific training task are identified. Subsequently, the same type of equipment is clustered, and common behavior patterns are extracted. Combined with the equipment technical specifications and historical fault data, an initial threshold is established, and a multi-level alarm level is set according to the operational risk level to clearly trigger warnings and take response strategies. In terms of indicators, key monitoring parameters are selected based on the criticality of the equipment and the degree of impact on training safety, and expert knowledge or hierarchical analysis method is introduced to assign weights to each indicator.

[0067] Step S202: Automatically match the most suitable threshold template through multi-dimensional dynamic labels. If the multi-dimensional dynamic labels conflict, use a weighted priority algorithm to select a template.

[0068] In this embodiment, after collecting the real-time operation data of the equipment, multi-dimensional dynamic labels are automatically generated in combination with the historical operation mode, task type and environmental parameters. These labels simultaneously characterize the mechanical state, environmental adaptability and operational behavior characteristics of the equipment. The candidate template set associated with these labels is retrieved from the preset threshold template database. If the equipment has only a single label, the corresponding template is directly matched; however, in the case of label overlap or conflict, the weighted priority algorithm is activated to calculate the influence of each label in the current scenario. The algorithm refers to factors such as task urgency, risk level, and equipment sensitivity to assign a weight score to each label, and the template corresponding to the one with the highest weighted score is used as the final matching result, thereby improving the adaptability and accuracy of the matching template.

[0069] Step S203: Rolling statistics of historical data according to a preset time window to calculate the indicator baseline value, that is, the indicator dynamic threshold value. The initial threshold value of the threshold template is compared with the indicator baseline value and corrected to obtain the corrected indicator dynamic threshold value. The calculation formula of the dynamic threshold value is:

[0070] ;

[0071] in, Indicates the modified dynamic threshold of the indicator, represents the initial threshold in the threshold template based on label matching, Represents the template weight coefficient, which is used to control the fusion ratio of the initial threshold value in the threshold template and the indicator baseline value. Indicates the average value of historical data in the preset time window, reflecting the normal operation level of the equipment. It represents the standard deviation of historical data in the preset time window, reflecting the degree of data fluctuation. K represents the variance coefficient, which is dynamically adjusted according to business priorities.

[0072] In this embodiment, the template weight coefficient The range is 0~1. =1, it completely relies on the initial threshold in the threshold template, which is suitable for new devices or scenarios with insufficient historical data. =0, it completely relies on historical data and is suitable for mature and stable equipment. It can automatically transition from "template-driven" to "data-driven" to reduce operation and maintenance costs. The variance coefficient k adjusts the magnification of the standard deviation according to business priority to control the sensitivity of the threshold to fluctuations. The larger the k value, the wider the threshold range, which is suitable for high-tolerance scenarios such as log synchronization. The smaller the k value, the narrower the threshold range, which is suitable for sensitive businesses such as real-time transactions.

[0073] Through the dynamic threshold correction algorithm, the system achieves a balance between global stability and local flexibility.

[0074] The calculation of the indicator baseline value excludes the data of the historical alarm period and is dynamically updated through the sliding window algorithm.

[0075] Step S3: Construct a device node network diagram, and based on the status evaluation of the device node network diagram, perform comparative analysis to determine whether the real-time status of each node is abnormal.

[0076] The specific steps of step S3 are:

[0077] Step S301: Each device is regarded as a node, and the real-time data transmission relationship between devices is regarded as an edge to construct a device node network graph.

[0078] In this embodiment, nodes contain multi-dimensional dynamic labels of devices, and edges are attached with transmission data indicators (such as data volume, latency, packet loss rate, etc.). This information can reflect the performance and stability of the link; based on the defined nodes and edges, the entire network topology is constructed using a graph data structure. This structure can intuitively display the connection relationship between device nodes and the data transmission path.

[0079] Step S302: Compare the currently collected real-time data with the revised indicator dynamic threshold to detect whether the data exceeds the normal fluctuation range and calculate the degree of deviation.

[0080] In this embodiment, the currently collected real-time data is compared with the dynamic threshold of the corresponding indicator to determine whether the data exceeds the normal fluctuation range. If so, anomaly detection is triggered and the weighted Euclidean distance algorithm is used to calculate the comprehensive deviation of the multi-dimensional indicators.

[0081] The advantages are: dynamic adaptation to complex scenarios, breaking through the limitations of traditional static thresholds, and being able to automatically match differentiated thresholds based on device tags; adapting to business burst traffic scenarios; dynamically adjusting thresholds as equipment ages; and the dynamic threshold fluctuation range design ensures that normal business fluctuations do not trigger alarms, eliminating environmental interference.

[0082] Step S303: Combined with the multi-dimensional dynamic tag information of the device node, further confirm and locate the abnormal situation, evaluate the operating status of the node and link, and divide different abnormality levels according to the degree of deviation and the evaluation results. The abnormality levels are divided into: special abnormality, level 1 abnormality, level 2 abnormality, level 3 abnormality and level 4 abnormality.

[0083] In this embodiment, the abnormality levels are divided into: special abnormality, level one abnormality (minor abnormality), level two abnormality (moderate abnormality), level three abnormality (serious abnormality) and level four abnormality (critical / catastrophic abnormality).

[0084] Special anomalies include those caused by external factors or human factors, such as power outages on some devices, the need to shut down some devices, and other human factors. Such anomalies are generally recovered quickly.

[0085] Level 1 anomaly, classification basis: data indicators have only slight fluctuations, the deviation from the normal state is small, and it does not reach the set major threshold.

[0086] Level 2 abnormality: classification based on: significant fluctuations in certain key indicators, with a moderate degree of deviation from normal conditions.

[0087] Level 3 abnormality, based on: multiple indicators deviate from the normal state at the same time, or a single key indicator reaches a high abnormal level.

[0088] Level 4 anomaly, classification basis: The anomaly is serious and spreads to key nodes, the overall network performance is greatly reduced, and some core functions may be interrupted.

[0089] Step S304: When a special anomaly is detected, the elasticity rule is matched according to the multi-dimensional dynamic label combination of the abnormal node, the threshold relaxation coefficient is dynamically calculated, and the elasticity range is limited in combination with the business priority label, and the above steps are repeated to perform anomaly detection.

[0090] like Figure 3 As shown, the elastic rule generates a dynamic rule base by multi-dimensional dynamic labels and anomaly types, calculates the elasticity coefficient, and dynamically relaxes the threshold. When multi-label conflict occurs, the final strategy is selected according to the label priority weight.

[0091] In this embodiment, the elasticity rules include an elasticity rule base, for example: Rule 1: If one of the labels of a node is "core switch" and the exception type is "packet loss rate exceeds the threshold", the delay threshold of all downstream devices will be automatically relaxed (+30%); Rule 2: If more than 30% of the nodes under a certain geographical label (such as "Data Center B") simultaneously alarm, the bandwidth threshold of the nodes in the area will be relaxed (+20%); the device labels (function, service level, location) are used to judge its criticality in the network to decide whether to trigger the elasticity rules. It should be noted that the specific relaxed thresholds here need to be set by personnel in this field according to actual conditions.

[0092] Dynamic calculation of threshold relaxation coefficient: Based on the severity of the anomaly (such as the alarm level and the number of affected devices) and the service tag weight, the elasticity coefficient is dynamically generated. The threshold is then adjusted in real time based on the elasticity coefficient. The elasticity rule base is used to determine the fluctuating range of the threshold relaxation. The threshold relaxation value is specifically calculated based on the elasticity coefficient. The elasticity coefficient of high-priority service nodes increases slightly to ensure their monitoring sensitivity. While the anomaly persists, the elasticity coefficient is periodically increased until the anomaly is resolved.

[0093] Business priority tags are combined to limit the elasticity range. Through deep coupling of elasticity rules and tags, the elasticity policy priority of each tag combination is defined. For example: Tag combination (core devices + financial services) → elasticity policy priority = 1 (the highest level, only 10% relaxation allowed); Tag combination (edge ​​devices + log service) → elasticity policy priority = 3 (50% relaxation allowed). If a device matches multiple elasticity rules simultaneously, that is, it belongs to at least two tag combinations at the same time, and a multi-tag conflict occurs, such as belonging to both "core devices" and "disaster recovery nodes", a weighted voting method is used to select the final policy and determine the relaxation threshold. The elasticity range limit for business priority tags is calculated based on the elasticity coefficient to determine the threshold relaxation value. If it does not exceed the elasticity range of the business priority tags, the relaxation is performed. If it exceeds, the elasticity coefficient is regenerated and the threshold relaxation value is calculated until it falls within the elasticity range of the business priority tags. It should be noted that the specific relaxation threshold here needs to be set by personnel in this field according to actual conditions.

[0094] Through the deep coupling of elastic rules and labels, fine-grained control is achieved, avoiding policy conflicts in multi-rule scenarios. Temporary threshold relaxation is supported in burst traffic scenarios. The elasticity range is limited by the business priority label to prevent excessive relaxation from missed faults. Gradual rollback avoids frequent threshold changes, effectively improving the intelligence level and fault tolerance of the overall operation and maintenance system.

[0095] Step S4: Generate alarm information for the detected abnormal behavior, including specific abnormal nodes, abnormal data indicators and possible abnormal causes.

[0096] The specific steps of step S4 are:

[0097] Step S401: Abnormal node identification and labeling. Based on the abnormality assessment results of the nodes in the network diagram, the nodes judged to be abnormal are automatically extracted and labeled with abnormal identification labels, including metadata such as device number, label combination (such as type, location, business role), etc., to ensure accurate positioning.

[0098] Step S402: Extract and describe abnormal indicators. For each abnormal node, extract the core indicators that trigger the abnormality (such as delay, packet loss rate, traffic surge, etc.), record the current value, historical baseline value and deviation of the indicator, and generate a quantitative indicator description to reflect the severity and trend of the abnormality.

[0099] Step S403: Preliminary abnormality cause analysis and prompts. This process uses tag combinations, historical event libraries, and node upstream and downstream link status to analyze possible fault causes and generate preliminary abnormality cause prompts, such as link congestion, abnormal upstream node transmission, and hardware resource exhaustion, to provide guidance for manual or automated processing.

[0100] Step S404: Structural encapsulation of alarm information. The above contents are uniformly encapsulated into a structured alarm information package. The format includes but is not limited to: node ID, alarm level, abnormal indicators and values, abnormality type, possible causes and detection timestamp, etc., to facilitate system processing, manual viewing, interface transmission and log archiving.

[0101] Step S405: Alarm information distribution and push: Determine the alarm push strategy based on the abnormality level and business priority, support multi-channel real-time sending (such as SMS, email, pop-up window, API interface), and ensure that relevant personnel or system modules are notified as soon as possible.

[0102] In this embodiment, the alarm rules are matched with the push strategy, and preset alarm rules are constructed based on the device type, geographical location, business priority, historical abnormal experience, etc., to automatically match the most suitable push strategy; a dynamic adjustment mechanism supports adjusting the alarm push strategy during special events or peak periods, such as avoiding repeated sending of the same event or pushing in batches to prevent information flooding and personnel fatigue.

[0103] Alarm classification based on abnormality level and business priority ensures that abnormalities of different levels will not be pushed to all personnel at the same time, avoiding duplication and irrelevant information interference. Through preset rules and dynamic adjustment mechanisms, the alarm push strategy can adapt to changes in different network environments and business scenarios and can operate stably in various situations.

[0104] Example 2

[0105] See also Figure 4, another embodiment provided by the present invention: an intelligent operation and maintenance monitoring and alarm system, comprising: a collection module, a threshold generation module, an abnormality judgment module and an alarm module;

[0106] The acquisition module is used to collect real-time communication data and multi-dimensional dynamic tags of each device node, wherein the multi-dimensional dynamic tags include dynamically generated device status tags, and automatically update tag attributes and lifecycles according to the real-time communication data;

[0107] The threshold generation module generates an adaptive alarm threshold based on the tag matching dynamic threshold template and historical data;

[0108] The abnormality judgment module is used to construct a device node network diagram, and based on the status evaluation of the device node network diagram, compare and analyze to determine whether the real-time status of each node is abnormal;

[0109] The alarm module is used to generate alarm information for detected abnormal behaviors, including specific abnormal nodes, abnormal data indicators and possible fault causes.

[0110] The anomaly judgment module includes: a graph construction unit, a deviation calculation unit and an anomaly judgment unit;

[0111] The graph construction unit is used to construct a device node network graph by treating each device as a node and the real-time data transmission relationship between devices as an edge;

[0112] The deviation calculation unit is used to compare the currently collected real-time data with the revised indicator dynamic threshold, detect whether the data exceeds the normal fluctuation range, and calculate the degree of deviation;

[0113] The abnormality determination unit is used to further confirm and locate the abnormal situation in combination with the multi-dimensional dynamic tag information of the device node, evaluate the operating status of the node and link, and divide different abnormality levels according to the deviation degree and the evaluation results.

[0114] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive redundancy.

[0115] The above-described specific embodiments further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is merely a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. Intelligent operation and maintenance monitoring and alarm method, characterized in that: include: Collecting real-time communication data and multi-dimensional dynamic tags of each device node, including dynamically generated device status tags, and automatically updating tag attributes and lifecycles based on real-time communication data, where the lifecycle is the lifecycle of the device status tag; Generate adaptive alarm thresholds based on tag matching dynamic threshold templates and historical data; Build a device node network diagram and evaluate the status of the device node network diagram to determine whether the real-time status of each node is abnormal. When a special anomaly is detected, match the elastic rules based on the multi-dimensional dynamic label combination of the abnormal node; Generate alarm information for detected abnormal behaviors, including specific abnormal nodes, abnormal data indicators and abnormal cause prompts.

2. The intelligent operation and maintenance monitoring and alarm method according to claim 1, characterized in that: The tag-based matching dynamic threshold template and the combination of historical data to generate an adaptive alarm threshold include: Creating a threshold template based on the dynamic tag classification of the device, the threshold template including: initial threshold, alarm level, indicator, and indicator weight; Automatically match the most suitable threshold template through multi-dimensional dynamic labels. If the multi-dimensional dynamic labels conflict, a weighted priority algorithm is used to select the template. Roll statistics on historical data according to the preset time window to calculate the indicator baseline value, that is, the indicator dynamic threshold. Compare and correct the initial threshold of the threshold template with the indicator baseline value to obtain the corrected indicator dynamic threshold.

3. The intelligent operation and maintenance monitoring and alarm method according to claim 1, characterized in that: The device node network diagram is constructed, and based on the status evaluation of the device node network diagram, the real-time status of each node is judged to be abnormal. When a special abnormality is detected, elastic rules are matched based on the multi-dimensional dynamic label combination of the abnormal node, including: Each device is regarded as a node, and the real-time data transmission relationship between devices is regarded as an edge to build a device node network graph; Compare the currently collected real-time data with the revised indicator dynamic threshold to detect whether the data exceeds the normal fluctuation range and calculate the degree of deviation; Combined with the multi-dimensional dynamic tag information of device nodes, it confirms and locates abnormal situations, evaluates the operating status of nodes and links, and classifies different abnormality levels according to the degree of deviation and evaluation results. The abnormality levels are divided into: special abnormality, level 1 abnormality, level 2 abnormality, level 3 abnormality, and level 4 abnormality; When a special anomaly is detected, the elasticity rule is matched based on the multi-dimensional dynamic label combination of the abnormal node, the threshold relaxation coefficient is dynamically calculated, and the elasticity range is limited by combining the business priority label. The above steps are repeated to perform anomaly detection.

4. The intelligent operation and maintenance monitoring and alarm method according to claim 3, characterized in that: The elastic rule includes an elastic rule base. When a multi-label conflict occurs, a final strategy is selected according to the label priority weight.

5. The intelligent operation and maintenance monitoring and alarm method according to claim 1, characterized in that: The generation of the multi-dimensional dynamic label includes: Dynamically generate a device status tag based on real-time communication data, the device status tag including device load, temporary service type, and abnormal event mark; Define a lifecycle for each tag and automatically remove the tag when the trigger condition fails.

6. The intelligent operation and maintenance monitoring and alarm method according to claim 2, characterized in that: The calculation of the indicator baseline value excludes data from the historical alarm period and is dynamically updated through a sliding window algorithm.

7. The intelligent operation and maintenance monitoring and alarm method according to claim 1, characterized in that: The method further comprises: Perform differential encoding compression on multi-dimensional dynamic tag data and only transmit the tag changes; Complete label classification and initial screening at the edge node, filter the normal label data and then upload it.

8. An intelligent operation and maintenance monitoring and alarm system, used to implement the intelligent operation and maintenance monitoring and alarm method according to any one of claims 1 to 7, characterized in that: include: Acquisition module, threshold generation module, abnormality judgment module and alarm module; The acquisition module is used to collect real-time communication data and multi-dimensional dynamic tags of each device node, the multi-dimensional dynamic tags including dynamically generated device status tags, and automatically update tag attributes and life cycle based on real-time communication data, the life cycle being the life cycle of the device status tag; The threshold generation module generates an adaptive alarm threshold based on the tag matching dynamic threshold template and historical data; The abnormality judgment module is used to construct a device node network diagram, and based on the status evaluation of the device node network diagram, judge whether the real-time status of each node is abnormal. When a special abnormality is detected, it matches the elastic rules according to the multi-dimensional dynamic label combination of the abnormal node; The alarm module is used to generate alarm information for detected abnormal behaviors, including specific abnormal nodes, abnormal data indicators and abnormal cause prompts.

9. The intelligent operation and maintenance monitoring and alarm system according to claim 8, characterized in that: The abnormality judgment module includes: a graph construction unit, a deviation calculation unit and an abnormality judgment unit; The graph construction unit is used to construct a device node network graph by treating each device as a node and the real-time data transmission relationship between devices as an edge; The deviation calculation unit is used to compare the currently collected real-time data with the revised indicator dynamic threshold, detect whether the data exceeds the normal fluctuation range, and calculate the degree of deviation; The abnormality determination unit is used to confirm and locate abnormal situations in combination with the multi-dimensional dynamic tag information of the device nodes, evaluate the operating status of the nodes and links, and classify different abnormality levels according to the degree of deviation and the evaluation results.

Citation Information

Patent Citations

  • A method and device for warning notification

    CN111782487B

  • A method, device, equipment and medium for intelligent alarm of operation and maintenance monitoring data

    CN114328118B

  • Telecommunication bearer network abnormal node positioning method based on terminal data

    CN108521346A

  • Multi-index intelligent dynamic threshold monitoring method and system

    CN115794532A

  • Security alarm information data interaction system and method

    CN118486152A

Cited By

  • Server data supervision method and electronic equipment

    CN120849222A

  • Communication network operation and maintenance system based on knowledge graph

    CN121619223A

  • Information system and equipment service work order closed-loop management system and method

    CN122155335A