Intelligent operation and maintenance monitoring and alarming method and system
By integrating multi-dimensional data and dynamic adaptive models, and combining them with device node network diagrams for intelligent operation and maintenance monitoring, the problem of neglecting the interaction relationships between devices and network topology in existing technologies has been solved. This enables accurate anomaly detection and real-time response, improving the intelligence and fault tolerance of the operation and maintenance system.
Patent Information
- Application Number
- CN202511135596.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing technologies have weak anomaly diagnosis capabilities in complex scenarios, neglecting the interaction between devices and network topology, resulting in low monitoring accuracy, untimely response, and difficulty in troubleshooting.
Intelligent anomaly judgment is achieved by employing multi-dimensional data fusion, tag association analysis, and dynamic adaptive models. Status assessment is performed based on the device node network diagram, adaptive alarm thresholds are generated, and anomaly detection and location are achieved through multi-dimensional dynamic tags and elastic rules.
It enables accurate perception of equipment operating status and real-time anomaly identification, improves the intelligence level and adaptability of operation and maintenance monitoring, and enhances the system's fault tolerance and flexibility under complex abnormal situations.
Smart Images

Figure CN120639591B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of monitoring and alarming, in particular to an intelligent operation and maintenance monitoring and alarming method and system. BACKGROUND
[0002] With the rapid development of technologies such as the Internet, the Internet of Things, and cloud computing, the number of device nodes in modern industrial and information systems has increased explosively, forming a complex network environment. Traditional operation and maintenance monitoring technologies mainly rely on fixed thresholds or static rules to monitor single indicators such as data transmission rate, delay, and packet loss rate. This method has the following defects: 1. Static threshold limitations: fixed thresholds cannot adapt to the dynamic changes of data fluctuations in the network environment; 2. Single-dimensional monitoring deficiency: the diagnostic ability of abnormalities in complex scenarios is weak; 3. Lack of network topology analysis: existing monitoring systems usually ignore the interaction between devices and the network topology structure, making it difficult to achieve global anomaly detection; 4. Insufficient adaptive ability: lacking adaptive adjustment mechanisms, it is difficult to form an effective real-time monitoring and early warning closed loop.
[0003] A kind of operation and maintenance monitoring data intelligent alarming method, device, equipment and medium are disclosed in Chinese patent with authorized announcement No.CN114328118B, which includes: according to operation and maintenance monitoring historical data, the training and correction of data prediction model are carried out;The operation and maintenance monitoring data of all time periods in several time periods from the current time is input into data prediction model, and the operation and maintenance monitoring prediction data of a certain time period in several time periods after the current time are obtained;According to the operation and maintenance monitoring prediction data of a certain time period in several time periods after the current time and the alarm threshold corresponding to the operation and maintenance monitoring data in the time period, comparison is carried out;If the operation and maintenance monitoring prediction data of a certain time period in a certain time period after the current time is greater than the alarm threshold corresponding to the operation and maintenance monitoring data in the time period, early warning is carried out, which effectively improves the accuracy and practicality of operation and maintenance monitoring data alarm.
[0004] A kind of alarm notification method and device are disclosed in Chinese patent with authorization announcement No.CN111782487B, wherein the method applied to alarm intelligent aggregation merging engine includes: obtaining target alarm information set monitored by alarm monitoring platform in target period;According to the system identifier corresponding to each alarm information in target alarm information set, determine the target operation and maintenance post corresponding to each alarm information;Alarm information corresponding to the same target operation and maintenance post in target alarm information set is aggregated, to obtain the alarm information set corresponding to each target operation and maintenance post;Each target operation and maintenance post corresponding alarm broadcast information is generated, and it is sent to voice calling platform, so that voice calling platform informs each target operation and maintenance post corresponding alarm broadcast information to corresponding operation and maintenance personnel.The application can quickly inform each target operation and maintenance post corresponding alarm broadcast information to operation and maintenance personnel in time, and the alarm notification efficiency is higher;And, the application will not appear the situation of missing important alarm information.
[0005] The defects of the above technical solutions are: the diagnosis ability of abnormality in complex scene is weak, and the interaction relationship between devices and network topology structure are ignored.
[0006] Therefore, the prior art has problems such as low monitoring accuracy, slow response and difficult fault troubleshooting in large-scale, dynamic device network environment, and there is an urgent need for a new monitoring and alarm method that can use multi-dimensional dynamic labels, adaptive threshold and global network graph analysis. SUMMARY
[0007] In view of the deficiencies of the prior art, the present application provides an intelligent operation and maintenance monitoring and alarm method and system, which performs intelligent abnormality judgment based on multi-dimensional data fusion, label correlation analysis and dynamic adaptive model, and performs alarm according to adaptive threshold.
[0008] To achieve the above purpose, the present application provides the following technical solutions:
[0009] The intelligent operation and maintenance monitoring and alarm method comprises:
[0010] Collect real-time communication data and multi-dimensional dynamic labels of each device node, the multi-dimensional dynamic labels include dynamically generated device state labels, and the label attributes and life cycle are automatically updated according to the real-time communication data, the life cycle is the life cycle of the device state label;
[0011] Based on label matching dynamic threshold template, generate adaptive alarm threshold combined with historical data;
[0012] Construct device node network graph, based on the state evaluation of device node network graph, judge whether the real-time state of each node is abnormal, when detecting special abnormality, match flexible rules according to multi-dimensional dynamic label combination of abnormal node;
[0013] The alarm information is generated for the detected abnormal behavior, including specific abnormal nodes, abnormal data indicators and abnormal reason prompts.
[0014] Specifically, the dynamic threshold template based on label matching generates an adaptive alarm threshold in combination with historical data, including:
[0015] According to the dynamic label classification of the device, a threshold template is created, including an initial threshold, an alarm level, an indicator and an indicator weight.
[0016] The most suitable threshold template is automatically matched through multi-dimensional dynamic labels. If the multi-dimensional dynamic labels conflict, a weight priority algorithm is used to select the template.
[0017] The historical data is statistically rolled up according to a preset time window, the baseline value of the indicator, i.e. the dynamic threshold of the indicator, is calculated, the initial threshold of the threshold template is compared with the baseline value of the indicator, and the corrected dynamic threshold of the indicator is obtained.
[0018] Specifically, the device node network graph is constructed, the real-time state of each node is judged based on the state evaluation of the device node network graph, when a special abnormality is detected, the elastic rule is matched according to the multi-dimensional dynamic label combination of the abnormal node, including:
[0019] Each device is taken as a node, and the real-time communication data transmission relationship between devices is taken as an edge, so as to construct a device node network graph.
[0020] The current collected real-time communication data is compared with the corrected dynamic threshold of the indicator, whether the data exceeds the normal fluctuation range is detected, and the deviation degree is calculated.
[0021] In combination with the multi-dimensional dynamic label information of the device node, the abnormal situation is confirmed and located, the running state of the node and the link is evaluated, different abnormal levels are divided according to the deviation degree and the evaluation result, and the abnormal levels include special abnormality, first-level abnormality, second-level abnormality, third-level abnormality and fourth-level abnormality.
[0022] When a special abnormality is detected, the elastic rule is matched according to the multi-dimensional dynamic label combination of the abnormal node, the threshold relaxation coefficient is dynamically calculated, the elastic range is limited in combination with the business priority label, and the above steps are repeated for abnormality detection.
[0023] Specifically, the elastic rule includes an elastic rule library, supports multi-label conflict arbitration, and selects the final strategy according to the label priority weight when multi-label conflict occurs.
[0024] Specifically, the generation of the multi-dimensional dynamic label includes:
[0025] The device state label is dynamically generated based on real-time communication data, and the device state label includes device load, temporary service type and abnormal event mark;
[0026] A life cycle is defined for each label, and the label is automatically removed when a trigger condition is invalid.
[0027] Specifically, the calculation of the index baseline value excludes data of a historical alarm period, and is dynamically updated by a sliding window algorithm.
[0028] Specifically, the intelligent operation and maintenance monitoring alarm method further comprises:
[0029] The multi-dimensional dynamic label data is differentially encoded and compressed, and only the changed part of the label is transmitted.
[0030] Label classification and preliminary screening are completed at the edge node, and normal label data is filtered and uploaded.
[0031] Specifically, the intelligent operation and maintenance monitoring alarm system is used to implement the intelligent operation and maintenance monitoring alarm method, and comprises a collection module, a threshold generation module, an abnormality judgment module and an alarm module.
[0032] The collection module is used to collect real-time communication data and multi-dimensional dynamic labels of each device node, the multi-dimensional dynamic labels include dynamically generated device state labels, and the label attributes and life cycle are automatically updated according to the real-time communication data, and the life cycle is the life cycle of the device state label.
[0033] The threshold generation module generates adaptive alarm thresholds based on label matching dynamic threshold templates and historical data.
[0034] The abnormality judgment module is used to construct a device node network graph, judge whether the real-time state of each node is abnormal based on state evaluation of the device node network graph, and when a special abnormality is detected, match elastic rules according to multi-dimensional dynamic labels of the abnormal node.
[0035] The alarm module is used to generate alarm information for the detected abnormal behavior, including specific abnormal nodes, abnormal data indicators and possible fault cause prompts.
[0036] Specifically, the abnormality judgment module comprises a graph construction unit, a deviation calculation unit and an abnormality determination unit.
[0037] The graph construction unit is used to construct a device node network graph by taking each device as a node and taking real-time communication data transmission relationship between devices as an edge.
[0038] The deviation calculation unit is used for comparing the real-time communication data collected currently with the modified dynamic threshold of the index, detecting whether the data exceeds the normal fluctuation range, and calculating the deviation degree;
[0039] The abnormality determination unit is used for confirming and locating the abnormal situation in combination with the multi-dimensional dynamic label information of the device node, evaluating the running state of the node and the link, and dividing different abnormality levels according to the deviation degree and the evaluation result.
[0040] Compared with the prior art, the beneficial effects of the present application are:
[0041] The intelligent operation and maintenance monitoring alarm method is proposed, the multi-dimensional dynamic label, the dynamic threshold template and the device node network graph modeling are introduced, the accurate perception of the device running state and the real-time abnormality identification are realized, the multi-source information fusion, the state self-adaptation and the abnormality level determination capability are possessed, the method can dynamically generate the label according to the actual running characteristics of the device, and the alarm threshold is continuously modified in combination with the historical data, the problems of fixed threshold, frequent false alarm and lack of upstream and downstream linkage judgment in the traditional operation and maintenance system are effectively solved, meanwhile, the device node network graph and the elastic alarm strategy are constructed, the fault tolerance and the flexibility of the system in the face of complex abnormal situations are improved, the comprehensive monitoring is realized, and the intelligent level and the adaptability of the operation and maintenance monitoring are significantly enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 The intelligent operation and maintenance monitoring alarm method flowchart is provided for the present application;
[0043] Figure 2 The operation and maintenance monitoring alarm flowchart is provided for the present application;
[0044] Figure 3 The elastic matching flowchart is provided for the present application;
[0045] Figure 4 The intelligent operation and maintenance monitoring alarm system architecture diagram is provided for the present application. DETAILED DESCRIPTION
[0046] The present application will be described in detail below in combination with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0047] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present application, and do not limit the present application.
[0048] It should be noted that various features of the embodiments of the present application can be combined with each other, and are within the protection scope of the present application, if there is no conflict. In addition, although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. In addition, the "first", "second", "third" and the like used in the present application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and effect.
[0049] Unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The term "and / or" used in the present application includes any and all combinations of one or more related listed items.
[0050] Embodiment 1
[0051] Please refer to Figures 1-3 The present application provides an embodiment: an intelligent operation and maintenance monitoring and alarm method, comprising the following specific steps:
[0052] Step S1: collecting real-time communication data of each device node and multi-dimensional dynamic labels, the multi-dimensional dynamic labels including dynamically generated device state labels, and automatically updating label attributes and life cycle according to real-time communication data, the life cycle being the life cycle of the device state label.
[0053] Collecting real-time communication data of each device node, collecting data of each device node through an embedded software or hardware monitoring module, and the main monitored indicators including but not limited to transmission rate, delay, packet loss rate, data flow, connection state, etc.
[0054] The data not only comes from the network transmission process, but also can include the running state of the device itself (such as temperature, load, etc.) and external environment data (such as geographic location, network topology information), and by integrating these multi-dimensional data, a comprehensive device running state view can be constructed.
[0055] The traditional label is statically configured and cannot reflect the real-time state change of the device, such as service switching and temporary load surge, and the dynamic label automatically generates or updates label attributes according to real-time communication data of the device, for example: if the CPU usage rate of a node is high for 5 minutes, a temporary high-load label is automatically added; if the transmission data volume of the node increases by 10 times, a suspected DDoS attack event label is marked, etc.
[0056] The dimensions of the multi-dimensional dynamic label include device type, service level, geographic location, etc., and weights are assigned based on service scenarios, for example, basic attribute label (such as device type, geographic location) → fixed weight (such as set to 30%), service attribute label (such as service level, service SLA) → dynamic weight (such as set to 40%-60%), device state label (such as fault history, real-time load) → temporary weight (such as set to 10%-30%), it should be noted that the weight setting is set by personnel in the art according to actual conditions or a large number of simulation experiments, and the weight values here are only examples.
[0057] Label life cycle management: define the validity period of the device state label, and when the validity period is exceeded, the state label is immediately invalidated, such as the high load label automatically invalidating after the CPU returns to normal, avoiding historical label interference with current analysis.
[0058] The dynamic label can maintain high resolution capability when facing complex network topology and massive device nodes.
[0059] The generation of the multi-dimensional dynamic label includes:
[0060] Based on real-time communication data, device state labels are dynamically generated, including device load, temporary service type, and abnormal event markers.
[0061] A life cycle is defined for each label, and the label is automatically removed when the trigger condition is invalidated.
[0062] When the multi-dimensional dynamic label data is transmitted, it needs to be first compressed by differential encoding, and only the label change part is transmitted; the edge node completes label classification and preliminary screening, and filters normal label data for uploading.
[0063] Step S2: based on label matching dynamic threshold template, combined with historical data to generate adaptive alarm threshold.
[0064] The specific steps of step S2 are:
[0065] Step S201: according to the dynamic label classification of the device, create a threshold template, the threshold template includes: initial threshold, alarm level, index and index weight.
[0066] In this embodiment, first, according to the data generated by the device in the simulation run, the label classification is dynamically generated, that is, the state characteristics of the device under a specific training task are identified, then the same type of devices are clustered, the common behavior patterns are extracted, the initial threshold is established combined with the device technical specification and historical fault data, the multi-level alarm level is set according to the operation risk level, the warning is triggered and the coping strategy is clear; in terms of indicators, the key monitoring parameters are selected according to the criticality of the device and the influence degree on the training safety, and the expert knowledge or analytic hierarchy process is introduced to assign weights to each indicator.
[0067] Step S202: The most suitable threshold template is automatically matched through multi-dimensional dynamic labels, and if the multi-dimensional dynamic labels conflict, a weight priority algorithm is used to select the template.
[0068] In this embodiment, after collecting the real-time running data of the device, the multi-dimensional dynamic labels are automatically generated combined with the historical running mode, task type and environmental parameters, these labels represent the mechanical state, environmental adaptability and operation behavior characteristics of the device at the same time, the candidate template set associated with these labels is retrieved from the preset threshold template database, if the device has only a single label, the corresponding template is directly matched; but in the case of label overlap or conflict, the weight priority algorithm is started, the influence degree of each label in the current scene is calculated, the algorithm refers to the task urgency, risk level, device sensitivity and other factors, and assigns a weight score to each label, and the template corresponding to the highest weighted score is taken as the final matching result, which improves the adaptability and accuracy of the matched template.
[0069] Step S203: The historical data is statistically calculated in a preset time window, the index baseline value, that is, the index dynamic threshold, is calculated, the initial threshold of the threshold template is compared and corrected with the index baseline value, and the corrected index dynamic threshold is obtained, and the calculation formula of the dynamic threshold is:
[0070] ;
[0071] Among them, represents the corrected index dynamic threshold, represents the initial threshold in the threshold template based on label matching, represents the template weight coefficient, which is used to control the fusion proportion of the initial threshold in the threshold template and the index baseline value, represents the mean value of the historical data in the preset time window, which reflects the normal operation level of the device, represents the standard deviation of the historical data in the preset time window, which reflects the data fluctuation degree, and k represents the variance coefficient, which is dynamically adjusted according to the business priority.
[0072] In this embodiment, the template weight coefficient ranges from 0 to 1, and when = 1, completely rely on the initial threshold in the threshold template, suitable for new devices or scenarios with insufficient historical data, when = 0, completely rely on historical data, suitable for mature devices with stable operation, which can automatically transition from "template dominant" to "data dominant", reducing operation and maintenance costs; the variance coefficient k adjusts the amplification multiple of the standard deviation according to the business priority, controls the sensitivity of the threshold to fluctuations, and the larger the value of k, the wider the threshold range, which is suitable for high tolerance scenarios such as log synchronization, and the smaller the value of k, the narrower the threshold range, which is suitable for sensitive businesses such as real-time transactions.
[0073] Through the dynamic threshold correction algorithm, the system realizes the balance between global stability and local flexibility.
[0074] The calculation of the baseline value of the index excludes the data of the historical alarm period and dynamically updates through the sliding window algorithm.
[0075] Step S3: Construct a device node network graph, and based on the state evaluation of the device node network graph, compare and analyze whether the real-time state of each node is abnormal.
[0076] The specific steps of step S3 are:
[0077] Step S301: Take each device as a node, and the real-time communication data transmission relationship between devices as an edge, to construct a device node network graph.
[0078] In this embodiment, the node contains the multi-dimensional dynamic label of the device, and the transmission data indicators (such as data volume, delay, and packet loss rate) are attached to the edge. These information can reflect the performance and stability of the link; according to the defined nodes and edges, the entire network topology is constructed using the graph data structure, which can intuitively display the connection relationship between device nodes and the data transmission path.
[0079] Step S302: Compare the currently collected real-time communication data with the corrected dynamic threshold of the index, detect whether the data exceeds the normal fluctuation range, and calculate the deviation degree.
[0080] In this embodiment, the currently collected real-time communication data is compared with the dynamic threshold of the corresponding index to determine whether the data exceeds the normal fluctuation range. If it does, an abnormality detection is triggered, and a weighted Euclidean distance algorithm is used to calculate the comprehensive deviation degree of the multi-dimensional index.
[0081] The advantages are: dynamically adapt to complex scenarios, break through the limitations of traditional static thresholds, and automatically match differentiated thresholds according to device labels; adapt to business burst traffic scenarios; dynamically adjust the threshold as the device ages; the dynamic threshold fluctuation range design does not trigger an alarm for normal business fluctuations, eliminating environmental interference.
[0082] Step S303: Combine the multi-dimensional dynamic tag information of the device nodes to further confirm and locate the abnormal situation, evaluate the operating status of the nodes and links, and classify different abnormal levels according to the degree of deviation and evaluation results. The abnormal levels are divided into: special abnormality, level 1 abnormality, level 2 abnormality, level 3 abnormality and level 4 abnormality.
[0083] In this embodiment, the anomaly levels are divided into: special anomaly, level 1 anomaly (minor anomaly), level 2 anomaly (moderate anomaly), level 3 anomaly (serious anomaly), and level 4 anomaly (critical / catastrophic anomaly).
[0084] Special anomalies include those caused by external or human factors, such as power outages in some equipment or the need to shut down some equipment. These anomalies are usually resolved quickly.
[0085] Level 1 anomaly, defined as follows: data indicators fluctuate only slightly, deviate from the normal state by a small margin, and do not reach the set main threshold.
[0086] Level 2 abnormality, defined as: significant fluctuations in certain key indicators, with a moderate degree of deviation from the normal state.
[0087] Level 3 abnormality is defined as follows: multiple indicators deviate from the normal state at the same time, or a single key indicator reaches a high abnormal level.
[0088] Level 4 anomaly is defined as follows: the anomaly is severe and has spread to critical nodes, resulting in a significant decrease in overall network performance and potential interruption of some core functions.
[0089] Step S304: When a special anomaly is detected, the threshold relaxation coefficient is dynamically calculated based on the multi-dimensional dynamic label combination matching elastic rules of the anomaly node, and the elastic range is limited by the business priority label. Then, the above steps are repeated to detect the anomaly.
[0090] like Figure 3 As shown, the elastic rules are generated from multi-dimensional dynamic labels and anomaly types to form a dynamic rule base, and the elasticity coefficient is calculated to dynamically relax the threshold. When multiple label conflicts occur, the final strategy is selected according to the label priority weight.
[0091] In the embodiment, the elastic rule includes an elastic rule library, for example: rule 1: if one of the node labels is "core switch" and the abnormal type is "packet loss rate exceeds threshold", then automatically relax the delay threshold of all devices downstream (+30%); rule 2: if more than 30% of the nodes in a certain geographic label (such as "data center B") are simultaneously alarmed, then relax the bandwidth threshold of the nodes in the region (+20%); the key of a device in the network is determined by the device label (function, service level, location) to determine whether to trigger the elastic rule. It should be noted that the relaxed threshold value needs to be set by a person skilled in the art according to the actual situation.
[0092] Dynamic calculation of threshold relaxation coefficient: according to the abnormal severity (such as alarm level, number of affected devices) and service label weight, an elastic coefficient is dynamically generated, and the threshold is real-time corrected based on the elastic coefficient. The elastic rule library is used to determine the fluctuating threshold relaxation range, and the value of the threshold relaxation is calculated according to the elastic coefficient. The elastic coefficient of the high-priority service node has a small increase, which ensures its monitoring sensitivity. During the abnormal period, the elastic coefficient is periodically increased until the abnormality is resolved.
[0093] The elastic amplitude is limited by the service priority label, and the elastic rule is deeply coupled with the label to define the elastic strategy priority of the label combination, for example: label combination (core device + financial service) → elastic strategy priority = 1 (highest level, only allowed to relax 10%), label combination (edge device + log service) → elastic strategy priority = 3 (allowed to relax 50%). If a device matches multiple elastic rules at the same time, that is, it belongs to not less than two label combinations at the same time, a multi-label conflict occurs, such as belonging to "core device" and "backup node" at the same time. The final strategy is determined by using a weighted voting method to select the relaxed threshold. The elastic amplitude limit combined with the service priority label is used to calculate the value of the threshold relaxation according to the elastic coefficient. If it does not exceed the elastic amplitude limit combined with the service priority label, it is relaxed. If it exceeds, the elastic coefficient is regenerated, the value of the threshold relaxation is calculated, and it is within the elastic amplitude limit combined with the service priority label. It should be noted that the relaxed threshold value needs to be set by a person skilled in the art according to the actual situation.
[0094] Through the deep coupling of the elastic rule and the label, fine-grained control is realized, the strategy contradiction in the multi-rule scene is avoided, the temporary threshold relaxation in the burst traffic scene is supported, the elastic amplitude is limited by the service priority label, the fault is prevented from being missed due to excessive relaxation, the gradual rollback avoids frequent changes of the threshold, and the intelligent level and fault tolerance of the overall operation and maintenance system are effectively improved.
[0095] Step S4: generating alarm information for the detected abnormal behavior, including specific abnormal nodes, abnormal data indicators, and possible abnormal reason prompts.
[0096] The specific steps of step S4 are:
[0097] Step S401: Abnormal node identification and labeling. Based on the abnormal evaluation results of the nodes in the network graph, the nodes determined to be abnormal are automatically extracted, and the abnormal nodes are labeled with abnormal labels, including device number, label combination (such as type, location, business role), and other meta information, to ensure accurate positioning.
[0098] Step S402: Abnormal index extraction and description. For each abnormal node, the core index (such as delay, packet loss rate, traffic surge, etc.) that triggers the abnormality is extracted, the current value, historical baseline value, and deviation amplitude of the index are recorded, and a quantitative index description is generated to reflect the severity and trend of the abnormality.
[0099] Step S403: Preliminary abnormal reason analysis and prompt. Using label combination, historical event library, node upstream and downstream link state, etc., the possible fault causes are analyzed, and a preliminary abnormal reason prompt is generated, such as link congestion, upstream node abnormal conduction, hardware resource depletion, etc., to provide guidance for manual or automated processing;
[0100] Step S404: Structured packaging of alarm information. The above contents are uniformly packaged into a structured alarm information package, the format of which includes but is not limited to: node ID, alarm level, abnormal index and value, abnormal type, possible reason, and detection timestamp, etc., to facilitate system processing, manual viewing, interface transmission, and log archiving.
[0101] Step S405: Alarm information distribution and push. According to the abnormal level and business priority, the alarm push strategy is determined, supporting multi-channel real-time sending (such as SMS, email, pop-up window, API interface), to ensure that relevant personnel or system modules are notified in the first time.
[0102] In this embodiment, the alarm rules and push strategy are matched. The preset alarm rules are constructed according to device type, geographical location, business priority, historical abnormal experience, etc., and the most suitable push strategy is automatically matched. The dynamic adjustment mechanism supports adjusting the alarm push strategy during special events or peak periods, such as avoiding repeated sending or batch pushing for the same event, to prevent information flooding and personnel fatigue.
[0103] The alarm classification based on abnormal level and business priority ensures that different levels of abnormalities are not pushed to all personnel at the same time, avoiding repeated and irrelevant information interference. Through the preset rules and dynamic adjustment mechanism, the alarm push strategy can adapt to changes in different network environments and business scenarios, and can operate stably in various situations.
[0104] Embodiment 2
[0105] Please refer to Figure 4The application provides another embodiment: an intelligent operation and maintenance monitoring alarm system, comprising a collection module, a threshold generation module, an abnormality judgment module and an alarm module.
[0106] The collection module is used for collecting real-time communication data and multi-dimensional dynamic labels of each device node, wherein the multi-dimensional dynamic labels comprise dynamically generated device state labels, and the label attributes and life cycle are automatically updated according to the real-time communication data.
[0107] The threshold generation module generates self-adaptive alarm thresholds based on label matching dynamic threshold templates and in combination with historical data.
[0108] The abnormality judgment module is used for constructing a device node network graph, and comparing and analyzing whether the real-time state of each node is abnormal based on state evaluation of the device node network graph.
[0109] The alarm module is used for generating alarm information for the detected abnormal behaviors, including specific abnormal nodes, abnormal data indicators and possible fault cause prompts.
[0110] The abnormality judgment module comprises a graph construction unit, a deviation calculation unit and an abnormality judgment unit.
[0111] The graph construction unit is used for constructing a device node network graph by taking each device as a node and taking real-time communication data transmission relationships between devices as edges.
[0112] The deviation calculation unit is used for comparing the currently collected real-time communication data with the corrected index dynamic threshold, detecting whether the data exceeds a normal fluctuation range, and calculating a deviation degree.
[0113] The abnormality judgment unit is used for further confirming and positioning abnormal conditions in combination with multi-dimensional dynamic label information of the device node, evaluating the running state of nodes and links, and dividing different abnormality levels according to the deviation degree and the evaluation result.
[0114] In addition, the parts of the above technical solutions in the embodiments of the application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail, so as to avoid excessive repetition.
[0115] The specific embodiments described above further specifically describe the purposes, technical solutions and beneficial effects of the application. It should be understood that the above description is only a specific embodiment of the application, and is not used to limit the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application should be included in the protection scope of the application.
Claims
1. An intelligent operation and maintenance monitoring and alarm method, characterized in that, The method comprises the following steps: Collecting real-time communication data of each device and multi-dimensional dynamic tags, wherein the multi-dimensional dynamic tags comprise dynamically generated device state tags, and the tag attributes and life cycle are automatically updated according to the real-time communication data, and the life cycle is the life cycle of the device state tag; Adaptive alarm thresholds are generated based on the matching of dynamic threshold templates and historical data; A device node network graph is constructed, and based on the state evaluation of the device node network graph, it is determined whether the real-time state of each node is abnormal; when a special abnormality is detected, the elastic rules are matched according to the multi-dimensional dynamic tags of the abnormal node; Alarm information is generated for the detected abnormality, including specific abnormal nodes, abnormal data indicators and abnormal reason prompts; The adaptive alarm thresholds are generated based on the matching of dynamic threshold templates and historical data, comprising: According to the dynamic tag classification of the device, a threshold template is created, and the threshold template comprises an initial threshold, an alarm level, an abnormal data indicator and an abnormal data indicator weight; The multi-dimensional dynamic tags are automatically matched with the adaptive threshold template; The historical data is statistically rolled up according to a preset time window, the index dynamic threshold is calculated, the initial threshold of the threshold template is compared with the index dynamic threshold, and the modified index dynamic threshold is obtained; The device node network graph is constructed, and based on the state evaluation of the device node network graph, it is determined whether the real-time state of each node is abnormal; when a special abnormality is detected, the elastic rules are matched according to the multi-dimensional dynamic tags of the abnormal node, comprising: Each device is taken as a node, and the real-time communication data transmission relationship between devices is taken as an edge, so as to construct a device node network graph; The current collected real-time communication data is compared with the modified index dynamic threshold, it is detected whether the data exceeds the normal fluctuation range, if yes, the abnormality detection is triggered, and the deviation degree is calculated; The abnormality is confirmed and located in combination with the multi-dimensional dynamic tag information of the device, the running state of the node and the link is evaluated, different abnormality levels are divided according to the deviation degree and the evaluation result, and the abnormality levels comprise a special abnormality, a first-level abnormality, a second-level abnormality, a third-level abnormality and a fourth-level abnormality; When a special abnormality is detected, the index dynamic threshold relaxation range is determined according to the matching of the multi-dimensional dynamic tags of the abnormal node and the elastic rules, the index dynamic threshold elastic range is limited in combination with the business priority, and the above steps are repeated for abnormality detection; wherein each rule in the elastic rules comprises multi-dimensional dynamic tags, an abnormality type and an index dynamic threshold relaxation range. 2.The intelligent operation and maintenance monitoring alarm method of claim 1, wherein, The generation of the multi-dimensional dynamic tags comprises: The device state tags are dynamically generated based on the real-time communication data, and the device state tags comprise device load, temporary business type and abnormal event mark; The life cycle of each tag is defined, and the tag is automatically removed when the trigger condition is invalid. 3.The intelligent operation and maintenance monitoring alarm method of claim 2, wherein, The calculation of the index baseline value excludes the data of the historical alarm period, and is dynamically updated through a sliding window algorithm. 4.The intelligent operation and maintenance monitoring alarm method of claim 1, wherein, The method further comprises: The multi-dimensional dynamic tag data is differentially encoded and compressed, and only the tag change part is transmitted; The label classification and preliminary screening are completed at the edge node, and the normal tag data is filtered and uploaded.
5. The intelligent operation and maintenance monitoring and alarm system for implementing the intelligent operation and maintenance monitoring and alarm method of any one of claims 1-4, characterized in that, The method comprises the following steps: The collection module, the threshold generation module, the abnormality judgment module and the alarm module; The collection module is used for collecting real-time communication data of each device and multi-dimensional dynamic labels, the multi-dimensional dynamic labels include dynamically generated device state labels, and label attributes and a life cycle are automatically updated according to the real-time communication data, the life cycle is a life cycle of the device state label; The threshold generation module generates self-adapting alarm thresholds based on label matching dynamic threshold templates and in combination with historical data; The abnormality judgment module is used for constructing a device node network graph, judging whether the real-time state of each node is abnormal based on state evaluation of the device node network graph, and when a special abnormality is detected, matching elastic rules according to multi-dimensional dynamic label combinations of the abnormal node; The alarm module is used for generating alarm information for the detected abnormal behavior, including a specific abnormal node, abnormal data indicators and abnormal reason prompts. 6.The intelligent operation and maintenance monitoring alarm system of claim 5, wherein, The abnormality judgment module includes a graph construction unit, a deviation calculation unit and an abnormality judgment unit; The graph construction unit is used for constructing a device node network graph by taking each device as a node and taking real-time communication data transmission relationships between devices as edges; The deviation calculation unit is used for comparing the currently collected real-time communication data with corrected index dynamic thresholds, detecting whether the data exceeds a normal fluctuation range, triggering abnormality detection if the data exceeds the normal fluctuation range, and calculating a deviation degree; The abnormality judgment unit is used for confirming and positioning abnormal conditions in combination with multi-dimensional dynamic label information of the device, evaluating running states of nodes and links, and dividing different abnormality levels according to the deviation degree and the evaluation result.
Citation Information
Patent Citations
A method and device for warning notification
CN111782487B
A method, device, equipment and medium for intelligent alarm of operation and maintenance monitoring data
CN114328118B
Service network flow analysis system and method
CN120166027A
Index alarm method, system and equipment based on dynamic baseline and medium
CN120431691A