Alarm data processing system and analysis processing method for communication distribution network management

Through data collection, rules engine and intelligent analysis modules, invalid alarms are automatically filtered, duplicate alarms are identified and fault root causes are generated, which solves the problems of low screening efficiency and mismatch in the massive alarm data, and achieves efficient fault location and rapid processing.

CN120416005APending Publication Date: 2025-08-01HANGZHOU MOTANNI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422764.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Faced with massive alarm data, manual screening is inefficient, which affects fault location and troubleshooting timeliness; even if alarm filtering is completed, the number of effective alarms is still large, and some equipment repeatedly reports similar alarms frequently occur after abnormalities, resulting in difficulty in judging faults; the alarm level in professional network management does not match the actual fault level on site, and the topology diagram and the circuit diagram are difficult to intuitively present the fault correlation, and it is impossible to quickly determine the fault location.

Method used

The data acquisition module is used to collect and format the alarm data in real time, and automatically filter invalid alarms through the rule engine and intelligent analysis module, identify duplicate alarms and generate root causes of failures. Combined with the alarm display and processing module, provide automated response and manual intervention, feedback and optimization module optimize filtering rules, integrate edge computing module for preliminary filtering and preprocessing, and use a distributed architecture to ensure system stability.

Benefits of technology

It realizes automatic identification and filtering of invalid alarms, improves fault response efficiency, merges similar alarms to avoid interference from redundant information, automatically generates accurate alarm levels and equipment status, automatically locates fault locations, improves troubleshooting efficiency, and solves the problem of grade mismatch and topology diagrams in professional network management that are difficult to visually present.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to the technical field of network communication, in particular to an alarm data processing system and analysis processing method for communication distribution network management, which comprises a data acquisition module, an alarm filtering engine, an alarm storage module, an alarm display and processing module and a feedback and optimization module. The innovative characteristics of a self-adaptive filtering mechanism, a multi-level intelligent cooperation system, cross-domain technology fusion (such as edge calculation and deep reinforcement learning), high-throughput and low-delay data processing capacity, a self-feedback intelligent evolution mechanism and the like are combined. According to the innovation points, the performance and the intelligent level of the alarm filtering system are improved, the stability and the expandability of the system are remarkably improved, and the requirement for efficiently managing alarm data in a large-scale network environment is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network communication technologies, and particularly to an alarm data processing system and an analysis and processing method for communication network distribution management. Background Art

[0002] With the continuous development of network communication technologies, the scale of network systems has become increasingly large, the number of network devices has increased sharply, and the total amount of events and alarm data collected and reported in real time by various network management systems has also increased significantly, generating tens of thousands to hundreds of thousands of alarm entries per day. Faced with such a huge amount of alarm data, manual screening is inefficient, seriously affecting the timeliness of fault location and troubleshooting. Even after alarm filtering, the number of effective alarms is still large, and some devices will repeatedly report the same type of alarm after an abnormality, and coupled with the frequent occurrence of alarms, it brings great interference to fault judgment.

[0003] On the other hand, the alarm levels in professional network management are pre-customized by the system, but in the on-site operation and maintenance environment, operation and maintenance personnel need to comprehensively judge the fault level based on the line type and the impact of the alarm on the device service, which results in the alarm levels on professional network management often not matching the actual on-site fault levels, making it difficult to accurately judge faults and handle them in a timely manner. In addition, the sorting rule of devices in professional network management is based on the order of network access time. Under this rule, it is difficult to visually present the fault correlation in the topology diagram and the line diagram, and it also cannot help operation and maintenance personnel quickly judge the initial fault point.

[0004] Regarding the solution to the operation and maintenance status of simultaneously monitoring multiple heterogeneous network management systems, the massive alarm data leads to low efficiency of manual screening, seriously affecting the timeliness of fault location and troubleshooting; even after alarm filtering, the number of effective alarms is still large; there are problems such as some devices repeatedly reporting the same type of alarm after an abnormality and the frequent occurrence of alarms; however, in the on-site operation and maintenance environment, operation and maintenance personnel need to comprehensively judge the fault level based on the line type and the impact of the alarm on the device service, which results in the alarm levels on professional network management often not matching the actual on-site fault levels. Currently, there have been some invention patents. For example: Patent No. CN107196804A discloses a power system terminal communication access network alarm centralized monitoring system and method, which realizes the collection, processing, and monitoring of alarms through modules such as alarm collection, alarm preprocessing, alarm monitoring, fault diagnosis, and fault handling scheduling. However, this patent still has the problem of further expanding the functions and performance of the alarm collection module to meet the requirements of real-time, comprehensive, and accurate alarm collection in different device network management systems.

[0005] Patent No. CN111585782A discloses an integrated centralized alarm automatic processing system and method. This patent realizes the collection and processing of alarm information through functional modules such as alarm collection, alarm centralized monitoring, alarm processing, alarm query and statistics, alarm association, alarm forwarding, and alarm rule management. However, there are still user experience and feedback issues with this patent. It can process and associate the user's business information and related information, realize new automated processing and association, and issue and associate new alarms. It can handle the user's business information and management information, improve the user's business and efficiency, but there are problems with reducing the efficiency and accuracy of the user experience and management.

[0006] Therefore, there is an urgent need for an alarm data analysis and processing method that can automatically identify and filter out invalid alarms, merge similar alarms into one piece of data to avoid interference from redundant information, automatically generate accurately matched alarm levels and device statuses, and automatically and accurately generate the root cause of the fault, providing strong guidance for maintenance personnel to quickly locate the fault location, thereby improving the fault response efficiency of maintenance, greatly enhancing the fault troubleshooting efficiency, and effectively reducing the fault handling duration. Summary of the Invention

[0007] (I) Technical Problems to be Solved 1. Facing a large amount of alarm data, the efficiency of manual screening is low, seriously affecting the timeliness of fault location and troubleshooting.

[0008] 2. Even after alarm filtering, the number of effective alarms is still large. After some devices are abnormal, they will repeatedly report similar alarms. Coupled with the frequent occurrence of alarms, it brings great interference to fault judgment.

[0009] 3. The alarm levels in the professional network management are pre-customized by the system. However, in the on-site maintenance environment, maintenance personnel need to comprehensively judge the fault level based on the line type and the impact of the alarm on the device business. This results in the alarm levels on the professional network management often not matching the actual on-site fault levels, making it difficult to accurately judge and handle faults in a timely manner.

[0010] 4. The sorting rule of devices in the professional network management is based on the order of network access time. Under this rule, it is difficult to intuitively present the fault correlation in the topology diagram and the line diagram, and it is impossible to directly locate the position where the fault occurs and the root cause of the fault, nor can it help maintenance personnel quickly judge and handle faults.

[0011] 5. The existing technologies cannot meet the requirements of real-time, unified, comprehensive, and accurate alarm collection in different device network management systems.

[0012] (II) Technical Solutions To achieve the above object, the present invention provides the following technical solution: An alarm data processing system for communication distribution network management, including: The data collection module is used to collect alarm and event data from multiple network management platforms in real time, and send the collected data to the alarm filtering engine after unified formatting processing; The alarm filtering engine, as the core module of the system, includes a rule engine and an intelligent analysis module. The rule engine is used to automatically analyze and filter out invalid alarms according to preset rules for information such as the severity, source, and time series of alarms. The intelligent analysis module, based on machine learning and deep learning technologies, automatically identifies and filters out potential invalid alarms such as duplicate alarms and temporary faults, and can analyze various types of information of related devices to accurately locate the location and root cause of the fault; The alarm storage module is used to store the filtered valid alarm data for subsequent historical query and statistical analysis; The alarm display and processing module provides an interface for operation and maintenance personnel to view and process alarms, supports automated response and manual intervention functions, and can automatically suggest operation steps for operation and maintenance personnel, such as restarting the device or checking the configuration; The feedback and optimization module is responsible for collecting feedback information from operation and maintenance personnel, and automatically optimizing the alarm filtering rules by analyzing the feedback data to improve the filtering accuracy and the overall performance of the system.

[0013] Preferably, the rule engine further includes a time window rule, a threshold rule, a dependency rule, a device / system status rule, and an event correlation rule, where: The time window rule is used to determine whether the same type of alarm is a duplicate alarm within a set time window and decide whether to ignore it; The threshold rule automatically ignores low-frequency alarms according to whether the alarm frequency of the device or system is lower than a preset threshold; The dependency rule is used to identify related derivative alarms caused by a certain root problem and perform merging or ignoring processing on subsequent related alarms; The device / system status rule decides whether to ignore the corresponding alarm according to the maintenance status or known downtime status of the device or system; The event correlation rule determines whether multiple alarms belong to the same event through information such as alarm timestamps and device IDs, and merges them into a single alarm for processing.

[0014] Preferably, the intelligent analysis module includes an anomaly detection unit, a classification and clustering unit, and a comprehensive disposal unit, where: The anomaly detection unit adopts a model based on time series analysis, such as a long short-term memory network or an autoregressive integrated moving average model, trains the model through historical alarm data, and automatically detects and filters out duplicate, false, or unnecessary alarms; The classification and clustering unit uses classification models such as support vector machines and random forests to classify alarms into different categories according to characteristics such as the device type, alarm level, and trigger conditions of the alarms, so as to facilitate screening and processing. At the same time, a clustering algorithm is used to classify similar alarms.

[0015] The comprehensive disposal unit constructs an intelligent fault diagnosis system based on deep learning algorithms. By integrating real-time device status monitoring data and the historical operation and maintenance experience knowledge base, it realizes the correlation analysis of multi-source heterogeneous data of alarm events, completes fault location and root cause tracing, effectively reduces the frequency of manual inspections, and improves the efficiency of fault disposal.

[0016] Preferably, the alarm display and processing module further includes: An automated response function that can automatically execute certain processing operations based on preset operation strategies, such as device restart, configuration check, etc. (when it is not an extreme or major problem, the platform will not execute operations according to the strategy, but only provide the root cause of the fault and the location where the fault occurred to assist users in fault handling); An artificial intervention interface that allows operation and maintenance personnel to perform manual processing when necessary and provides operation suggestions to assist operation and maintenance decision-making; The system interface supports real-time alarm push and provides multiple view modes, such as list view, graph view, and dashboard view, to facilitate operation and maintenance personnel to quickly locate and analyze alarm information.

[0017] Preferably, the feedback and optimization module includes: An artificial feedback mechanism that allows operation and maintenance personnel to manually mark false alarms or irrelevant alarms through interface functions such as "mark as invalid", and use this feedback information to reverse adjust the rule engine; A self-learning mechanism. Based on the collected historical feedback data and alarm patterns, the intelligent analysis module regularly self-optimizes and dynamically adjusts the filtering strategy to improve the filtering accuracy and the overall performance of the system; A self-feedback intelligent evolution mechanism that not only relies on manual marking but also includes correction suggestions automatically generated by the system to help operation and maintenance personnel adjust and optimize alarm filtering rules more quickly and effectively.

[0018] Preferably, an edge computing module is also integrated in the system architecture, which is used to process some alarm data in real time at the source of data generation (such as the device side), perform preliminary filtering and preprocessing, thereby reducing the transmission burden of alarm data and improving the response speed of the system, especially suitable for alarm processing in a large-scale distributed network environment.

[0019] Preferably, the intelligent analysis module further includes a deep reinforcement learning unit, which adopts a deep reinforcement learning algorithm. By continuously interacting with the environment, it optimizes the alarm filtering decision-making strategy, enabling the system to automatically learn and adjust the filtering strategy according to historical feedback and environmental changes, thereby improving the accuracy of alarms and reducing false alarms and missed alarms.

[0020] Preferably, the alarm storage module adopts a distributed database architecture, which supports high-throughput and low-latency data storage and retrieval. Through data sharding and replication mechanisms, it ensures that the system can still perform historical queries and statistical analysis efficiently and stably in the case of large-scale alarm data.

[0021] Preferably, the system architecture is designed as a distributed self-optimizing architecture, which can dynamically perform horizontal expansion according to alarm traffic and network status, ensuring that the system can still operate stably in the case of massive data; moreover, the architecture has an adaptive high-availability design, which can dynamically adjust the fault tolerance mechanism according to traffic changes and node status. For example, it automatically adds processing nodes under high load or quickly switches to standby nodes in case of node failures, ensuring the high availability and continuous stable operation of the system.

[0022] A method for alarm data analysis and processing in communication distribution network management, the steps are as follows Step 1: According to whether the alarm affects the device status and services, the system automatically analyzes an alarm table. Step 2: When receiving the alarm table, the system automatically analyzes and queries, and classifies the information in the alarm table as the status to be stored in the database. The whole process of this table system can be intervened manually, and useless alarms are filtered out. Step 3: Judge whether the new alarm is stored in the database according to the set alarm filtering rules (this process can be automatically processed by the system or intervened manually). If so, execute Step 4; if not, return to Step 2. Step 4: Compare the alarm data in the database. If there is an uneliminated alarm with the same name in the database, only keep the alarm information that has been stored in the database before, record the update time of the alarm, and complete the aggregation processing of the alarm. Step 5: Judge whether there are alarm data that occur repeatedly and are eliminated. If so, execute Step 6; if not, return to Step 3. Step 6: For alarm data that occur repeatedly and are eliminated, during the process of the system processing the alarm, give it a judgment time. Within this judgment period, no matter how many times the alarm of this device occurs and is eliminated, it is defaulted that there is only one alarm information. When the elimination entry of this alarm is received for the last time, start the alarm buffer time. If no more alarms of this kind are received within the buffer time, the system determines that this alarm is truly eliminated after the buffer time ends, and eliminates this alarm; otherwise, it is considered that the alarm always exists and belongs to the uneliminated state during the jitter process. Step 7: Collect on-site alarm information and redefine the basic alarm level according to the impact of the alarm on the service. Step 8: Distinguish specific situations such as single-end break, double-end break, disconnection, and power failure according to the on-site link type. When an alarm is received, query the device status, and match the impact of the alarm on the device status according to the device status. If there are alarms at both ends and it affects the service, the alarm level will be upgraded. Step 9: According to the system link fault continuity and relevance judgment rules, which are the system processing logics sorted out by our side in combination with information such as the alarm situation and distance sorting of on-site devices, automatically and accurately generate the root cause of the fault.

[0023] (III) Beneficial effects Compared with the prior art, the present invention provides an alarm data processing system and an analysis and processing method for communication network distribution management, which have the following beneficial effects: 1. Automatically identify and filter invalid alarms through a rule engine, effectively improving the fault response efficiency of operation and maintenance, and solving the problem of low efficiency of manual screening in the face of a large amount of alarm data, which seriously affects the timeliness of fault location and troubleshooting. 2. Merge similar alarms into one piece of data to avoid redundant information interfering with fault judgment, and solve the problems that even after alarm filtering, the number of valid alarms is still large, some devices report similar alarms repeatedly after anomalies, and alarms occur frequently. 3. Automatically generate accurately matched alarm levels and device statuses, which highly conform to business logics and on-site operation and maintenance methods, and can more effectively help operation and maintenance personnel accurately judge faults and handle them in a timely manner, solving the problem that the alarm levels on professional network management are often not matched with the actual fault levels on site. 4. Automatically and accurately generate the root cause of the fault, providing strong guidance for operation and maintenance personnel to quickly locate the fault location, greatly improving the fault troubleshooting efficiency, effectively reducing the fault handling time, and solving the problem that the device sorting rules in professional network management are executed based on the sequence of network access time, and it is difficult for the topology map and line map to intuitively present the fault relevance and cannot help operation and maintenance personnel quickly judge the initial fault point. Specific implementation manners

[0024] In order to better understand the purpose, structure and function of the present invention, the present invention will be further described in detail below for an alarm data processing system and an analysis and processing method for communication network distribution management.

[0025] An alarm data processing system and an analysis and processing method for communication network distribution management, wherein the system can be divided into the following modules: Data acquisition module: responsible for collecting alarm and event data from each network management platform in real time, and sending them to the alarm filtering engine after unified formatting.

[0026] Alarm filtering engine: This is the core module that automatically analyzes and filters invalid alarms according to preset rules. This module will consist of two parts: Rule Engine: Defines the rules for alarm filtering and determines whether an alarm is valid based on information such as the severity, source, and time series of the alarm.

[0027] Intelligent Analysis Module: Based on machine learning / deep learning technologies, automatically identifies potential invalid alarms (e.g., duplicate alarms, temporary failures, etc.), and analyzes the fault location and root cause.

[0028] Alarm Storage Module: Stores the filtered alarm data for historical query and statistical analysis. The key point is to store the valid alarms in the database for subsequent query and statistics.

[0029] Alarm Display and Handling Module: Provides an interface for operation and maintenance personnel to view and handle alarms, pushes valid alarms to the operation and maintenance personnel's interface, supports quick response and handling, supports automated response and manual intervention, and automatically provides suggested operations for operation and maintenance personnel (such as restarting the device, checking the configuration, etc.).

[0030] Feedback and Optimization Module: Collects the feedback from operation and maintenance personnel and automatically optimizes the alarm filtering rules to improve the filtering accuracy.

[0031] Among them According to the actual business requirements, the following types of rules can be designed for the alarm filtering rules: Time Window Rule: Alarms of the same type that are repeatedly triggered within a short period of time are judged as duplicate alarms and can be ignored.

[0032] Threshold Rule: If the alarm frequency of a certain device or system is lower than the preset threshold, these low-frequency alarms can be automatically ignored.

[0033] Dependency Rule: Some alarms may be caused by a fundamental problem. Subsequent related alarms can be regarded as "derivative alarms" and no longer processed separately. For example, if an alarm of the main switch is triggered, alarms of all downstream devices dependent on this device can be temporarily ignored.

[0034] Device / System Status Rule: If a device or system is in a maintenance state or a known downtime state, all alarms can be ignored to avoid false alarms.

[0035] Event Correlation Rule: Based on information such as the alarm timestamp and device ID, determine whether multiple alarms triggered by the same event belong to the same problem, and merge them into one alarm for processing.

[0036] Furthermore, an adaptive alarm filtering engine can be adopted: The current alarm filtering engine relies on a rule engine and an intelligent analysis module for filtering, but its innovation can be enhanced by introducing an adaptive alarm filtering mechanism. The adaptive mechanism can automatically adjust the filtering rules according to the changes in real-time data. For example, based on factors such as real-time load and alarm frequency, dynamically adjust the filtering strategy and model to avoid the limitations of fixed rules in certain scenarios.

[0037] A multi-level intelligent collaborative filtering system can also be adopted: During the alarm filtering process, in addition to the existing rule engine and machine learning module, add a multi-level intelligent collaborative system. This system will combine multiple data sources (such as device status, historical alarm data, feedback from operation and maintenance personnel, etc.) for multi-dimensional analysis to filter alarms in a more intelligent and accurate manner. For example, by combining big data analysis and artificial intelligence, automatically identify undefined abnormal patterns or potential faults in the system and adjust the filtering rules in a timely manner.

[0038] The integration of cross-domain technologies can also be carried out For example, the combination of edge computing and alarm filtering: Consider introducing edge computing into the system architecture. Edge computing can process part of the alarm data in real time at the source of data generation (such as the device side), perform preliminary filtering and preprocessing, reduce the transmission burden of alarm data, and improve the response speed. This method is particularly effective for alarm processing in a large-scale distributed network environment.

[0039] Optimizing alarm decision-making with deep reinforcement learning: In existing machine learning and deep learning models, deep reinforcement learning can be introduced to optimize the decision-making process of alarm filtering. Reinforcement learning optimizes the decision-making strategy by continuously interacting with the environment. This method can enable the alarm system to automatically learn and adjust the filtering strategy according to historical feedback and environmental changes, thereby improving the accuracy of alarms and reducing false alarms.

[0040] Considering the timeliness of alarm data, the system needs to support high-throughput and low-latency data processing capabilities to achieve fast alarm filtering.

[0041] With the popularization of 5G networks, the introduction of real-time alarm stream processing based on 5G networks can be considered in the system architecture. The low-latency feature of 5G can improve the response speed of the system and reduce the alarm processing time, which is particularly important for network environments with high real-time requirements.

[0042] To ensure that the system can still work properly under a large amount of alarm traffic, a highly available architecture needs to be designed to ensure the stable operation of the system. With an adaptive high-availability design, the system can dynamically adjust its fault tolerance mechanism according to factors such as traffic and node status. For example, it automatically adds processing nodes under high load, or quickly switches to standby nodes for alarm data processing when some nodes fail, ensuring the high availability of the system.

[0043] The system adopts a distributed self-optimizing architecture that can automatically expand. This architecture can horizontally expand dynamically according to alarm traffic, network status, etc., ensuring that the system can still operate stably in the case of massive data. In addition, this architecture can automatically adjust the system configuration according to real-time feedback data for performance optimization, ensuring that the system can provide the best performance in different environments.

[0044] The system can provide a "mark as invalid" function, allowing operation and maintenance personnel to manually mark false alarms or irrelevant alarms. These feedbacks will reverse-adjust the rule engine to improve the filtering accuracy.

[0045] By collecting historical feedback data and alarm patterns, the intelligent analysis module can regularly self-optimize and adjust the filtering strategy.

[0046] Based on the existing manual feedback mechanism and self-learning mechanism, enhance the self-feedback intelligent evolution mechanism. Through alarm data and the feedback of operation and maintenance personnel, the system can continuously self-learn and optimize to improve the accuracy of alarm filtering. At the same time, the feedback is not limited to manual marking, but also includes correction suggestions automatically generated by the system to help operation and maintenance personnel adjust the alarm filtering rules more quickly and effectively.

[0047] A method for analyzing and processing alarm data in communication distribution network management, the steps are as follows Step 1: According to whether the alarm affects the device status and business, the system automatically analyzes an alarm table; Step 2: When receiving the alarm table, the system automatically analyzes and queries, and classifies the information in the alarm table as the status to be stored in the database. The whole process of this table's automatic processing can be intervened by humans, and useless alarms are filtered out; Step 3: Judge whether the new alarm is stored in the database according to the set alarm filtering rules (this process can be automatically processed by the system or intervened by humans). If so, execute Step 4. If not, return to Step 2; Step 4: Compare the alarm data in the database. If there is an uneliminated alarm with the same name in the database, only keep the alarm information that has been stored in the database before, record the update time of the alarm, and complete the aggregation processing of the alarm; Step 5: Judge whether there are alarm data with repeated occurrences and eliminations. If so, execute Step 6. If not, return to Step 3; Step 6: For the alarm data that recurs and is eliminated, during the process of the system handling the alarm, give it a judgment time. Within this judgment period, regardless of how many times the alarm of the device occurs and is eliminated, it is defaulted that there is only one alarm message. When the last defect elimination entry of this alarm is received, start the alarm buffer time. If the alarm is not received again within the buffer time, the system determines that this alarm is truly eliminated and eliminates the alarm only after the buffer time ends. Otherwise, it is considered that the alarm always exists and is not eliminated during the jitter process; Step 7: Collect on-site alarm information and redefine the basic alarm level according to the impact degree of the alarm on the service; Step 8: Distinguish specific situations such as single-ended break, double-ended break, disconnection, power failure, etc. according to the on-site link type. When an alarm is received, query the device status, and match the impact of the alarm on the device status according to the device status. If there are alarms at both ends and it affects the service, the alarm level will be upgraded; Step 9: According to the system link fault continuity and relevance judgment rules, which are the system processing logics sorted out by our side by combining information such as the alarm situation of on-site devices and distance sorting, automatically and accurately generate the root cause of the fault.

[0048] The steps in Step 1 include: Step 101: Divide alarms into different levels according to the impact degree of on-site device alarms on the service; Step 102: Determine the impact degree of the alarm on the device status according to the device type and link type; Step 103: Associate information such as the alarm level, device type, link type, and service impact degree to form a corresponding table of alarms and services.

[0049] The steps in Step 6 include: Step 601: Set the judgment time period for alarm anti-jitter, such as 5 minutes; Step 602: Within the judgment time period, if the occurrence and defect elimination information of this alarm is received, it is considered that the alarm always exists and no processing is performed; Step 603: After the judgment time period ends, if the last received information for this alarm is defect elimination information, start the alarm buffer time, such as 2 minutes; Step 604: Within the buffer time, if the occurrence information of this alarm is not received again, it is determined that this alarm is truly eliminated and alarm elimination processing is performed; if the occurrence information of this alarm is received, return to Step 602 and continue to judge.

[0050] The steps in Step 9 include: Step 901: Collect on-site device alarm information, device location information, line information, device-to-device distance information, and device relevance information; Step 902: Analyze the alarm information, device location, and line information to determine the initial point where the fault occurred. Step 903: Combine the device - to - device distance information and device correlation information to analyze the impact scope and propagation path of the fault. Step 904: Automatically generate a root - cause analysis report of the fault based on the above analysis results.

[0051] Embodiment 1: A method for analyzing and processing alarm data in communication distribution network management according to the present invention includes the following steps: Step 1: Based on whether the on - site alarm affects the device status and services, sort out a corresponding table of alarms and services.

[0052] Step 101: Classify the alarms into four levels: severe level, important level, minor level, and prompt level according to the impact degree of on - site device alarms on services. Among them, severe - level alarms indicate that the device has a serious fault and the service is interrupted; important - level alarms indicate that the device has an important fault and the service is greatly affected; minor - level alarms indicate that the device has a general fault and the service is somewhat affected; prompt - level alarms indicate that the device has a minor fault and the service is basically normal.

[0053] Step 102: Determine the impact degree of the alarm on the device status according to the device type and link type. For example, for optical fiber communication devices, a single - port fault belongs to a minor - level impact; a dual - port fault belongs to an important - level impact; a main control board fault belongs to a severe - level impact. For wireless communication devices, a base - station power - off belongs to a severe - level impact; a base - station radio - frequency module fault belongs to an important - level impact; an antenna fault belongs to a minor - level impact.

[0054] Step 103: Correlate information such as alarm level, device type, link type, and service impact degree to form a corresponding table of alarms and services.

[0055] Step 2: When receiving an alarm, query the corresponding table, filter out the useless alarm information, and store the useful alarm information in the database.

[0056] Step 3: Determine whether the new alarm is stored in the database. If so, execute Step 4; if not, return to Step 2.

[0057] Step 4: Compare the alarm data in the database. If there is an un - eliminated alarm with the same name in the database, only retain the alarm information that has been stored in the database before, record the update time of the alarm, and complete the aggregation processing of the alarm.

[0058] Step 5: Determine whether there is alarm data that occurs repeatedly and has been eliminated. If so, execute Step 6; if not, return to Step 3.

[0059] Step 6: For the alarm data that recurs and is eliminated, during the process of the system handling the alarm, give it a judgment time period, which is set to 20 minutes.

[0060] Step 601: Set the judgment time period for anti-jitter of the alarm to 20 minutes.

[0061] Step 602: Within the 20-minute judgment time period, if the occurrence and elimination information of the alarm is received, it is considered that the alarm always exists and no processing is performed.

[0062] Step 603: After the judgment time period ends, if the last received information for this alarm is elimination information, start the alarm buffer time, which is set to 20 minutes.

[0063] Step 604: Within the 20-minute buffer time, if the occurrence information of this alarm is not received again, it is determined that the alarm is truly eliminated and alarm elimination processing is performed; if the occurrence information of this alarm is received, return to Step 602 and continue to judge.

[0064] Step 7: Collect on-site alarm information and re-define the basic alarm level according to the impact degree of the alarm on the service.

[0065] Step 8: Distinguish specific situations such as single-end break, double-end break, disconnection, and power failure according to the on-site link type. When receiving an alarm, query the device status, and match the impact of the alarm on the device status according to the device status. If there are alarms at both ends and it affects the service, the alarm level will be upgraded.

[0066] Step 9: According to the system link fault continuity and relevance judgment rules, which are the system processing logics sorted out by our side by combining information such as the alarm situation of on-site devices and distance sorting, automatically and accurately generate the root cause of the fault.

[0067] Step 901: Collect on-site device alarm information, device location information, line information, device distance information between devices, and device relevance information.

[0068] Step 902: Analyze the alarm information, device location, and line information to judge the initial point where the fault occurs.

[0069] Step 903: Combine the device distance information between devices and the device relevance information to analyze the impact scope and propagation path of the fault.

[0070] Step 904: According to the above analysis results, automatically generate a root cause analysis report of the fault.

[0071] Embodiment 2: A method for alarm data analysis and processing in communication distribution network management of the present invention includes the following steps: Step 1: Based on the impact of on-site alarms on the device status and services, compile a corresponding table of alarms and services.

[0072] Step 101: According to the degree of impact of on-site device alarms on services, classify the alarms into three levels: severe level, important level, and general level. Among them, severe-level alarms indicate that the device has a serious fault and the service is completely interrupted; important-level alarms indicate that the device has an important fault and the service is greatly affected; general-level alarms indicate that the device has a general fault and the device performance is affected to a certain extent.

[0073] Step 102: Determine the degree of impact of the alarm on the device status according to the device type and link type. For example, for optical fiber communication devices, a single-port fault belongs to a general-level impact; a double-port fault belongs to a severe-level impact; important-level alarms, main control board faults, or power supply faults in multiple devices on the same link belong to severe-level impacts. For wireless communication devices, base station power failure or main device failure belongs to a severe-level impact; base station radio frequency module failure belongs to an important-level impact; antenna failure or auxiliary device failure belongs to a general-level impact.

[0074] Step 103: Correlate information such as alarm level, device type, link type, and service impact degree to form a corresponding table of alarms and services.

[0075] The remaining steps are the same as those in Embodiment 1.

[0076] Embodiment 3: An alarm data analysis and processing method for communication distribution network management according to the present invention includes the following steps: Step 1: Based on the impact of on-site alarms on the device status and services, compile a corresponding table of alarms and services.

[0077] Step 101: According to the degree of impact of on-site device alarms on services, classify the alarms into three levels: severe level, important level, and general level. Severe-level alarms correspond to complete interruption of services; important-level alarms correspond to a large impact on services and a performance drop of more than 70%; general-level alarms correspond to basically normal services and a performance drop within 30%.

[0078] Step 102: Determine the degree of impact of the alarm on the device status according to the device type and link type. For example, for optical fiber communication devices, a single-port fault belongs to a general-level impact; a double-port fault belongs to a severe-level impact; important-level alarms, main control board faults, or power supply faults in multiple devices on the same link belong to severe-level impacts. For wireless communication devices, base station power failure or main device failure belongs to a severe-level impact; base station radio frequency module failure belongs to an important-level impact; antenna failure or auxiliary device failure belongs to a general-level impact.

[0079] Step 103: Associate information such as the alarm level, device type, link type, and degree of business impact to form a corresponding table of alarms and services.

[0080] The remaining steps are the same as those in the first embodiment.

[0081] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. An alarm data processing system for communication distribution network management, characterized in that, Including: A data collection module, which is used to collect alarm and event data from multiple network management platforms in real time, and send the collected data to the alarm filtering engine after unified formatting; The alarm filtering engine, as the core module of the system, includes a rule engine and an intelligent analysis module. The rule engine is used to automatically analyze and filter out invalid alarms according to preset rules for information such as the severity, source, and time series of alarms. The intelligent analysis module, based on machine learning and deep learning technologies, automatically identifies and filters out potential invalid alarms such as duplicate alarms and temporary faults, and can analyze multiple types of information of related devices to accurately locate the position and root cause of the fault; An alarm storage module, which is used to store the filtered valid alarm data for subsequent historical query and statistical analysis; An alarm display and processing module, which provides an interface for operation and maintenance personnel to view and process alarms, supports automated response and manual intervention functions, and can automatically suggest operation steps for operation and maintenance personnel, such as restarting the device or checking the configuration; A feedback and optimization module, which is responsible for collecting feedback information from operation and maintenance personnel, and automatically optimizing the alarm filtering rules by analyzing the feedback data to improve the filtering accuracy and the overall performance of the system.

2. The alarm data processing system for communication network distribution management according to claim 1, characterized in that The rule engine further includes a time window rule, a threshold rule, a dependency rule, a device / system status rule, and an event association rule, where: The time window rule is used to determine whether the same type of alarm is a duplicate alarm within a set time window and decide whether to ignore it; The threshold rule automatically ignores low-frequency alarms according to whether the alarm frequency of the device or system is lower than a preset threshold; The dependency rule is used to identify related derivative alarms caused by a certain root problem, and perform merging or ignoring processing on subsequent related alarms; The device / system status rule decides whether to ignore the corresponding alarm according to the maintenance status or known downtime status of the device or system; The event association rule judges whether multiple alarms belong to the same event through information such as alarm timestamps and device IDs, and merges them into a single alarm for processing.

3. The alarm data processing system for communication network distribution management according to claim 1, wherein The intelligent analysis module includes an anomaly detection unit, a classification and clustering unit, and a comprehensive disposal unit, where: The anomaly detection unit adopts a model based on time series analysis, such as a long short-term memory network or an autoregressive integrated moving average model, trains the model through historical alarm data, and automatically detects and filters out duplicate, false, or unnecessary alarms; The classification and clustering unit uses classification models such as support vector machines and random forests to classify alarms into different categories according to characteristics such as device types, alarm levels, and triggering conditions for easy screening and processing. At the same time, clustering algorithms are used to classify similar alarms, thereby reducing the manual processing burden of operation and maintenance personnel; The comprehensive disposal unit constructs an intelligent fault diagnosis system based on deep learning algorithms, realizes the correlation analysis of multi-source heterogeneous data of alarm events by integrating device real-time status monitoring data and historical operation and maintenance experience knowledge base, completes fault location and root cause tracing, effectively reduces the frequency of manual inspections, and improves the fault disposal efficiency.

4. The alarm data processing system for communication network distribution management according to claim 1, characterized in that The alarm display and processing module further includes: An automated response function that can automatically execute certain processing operations based on preset operation strategies, such as device restart, configuration check, etc.; A manual intervention interface that allows operation and maintenance personnel to perform manual processing when necessary and provides operation suggestions to assist operation and maintenance decision-making; The system interface supports real-time alarm push and provides multiple view modes, such as list view, graphical view, and dashboard view, to facilitate operation and maintenance personnel to quickly locate and analyze alarm information.

5. The alarm data processing system for communication network distribution management according to claim 1, characterized in that The feedback and optimization module includes: A manual feedback mechanism that allows operation and maintenance personnel to manually mark false alarms or irrelevant alarms through interface functions such as "mark as invalid" and use this feedback information to reverse-adjust the rule engine; A self-learning mechanism that, based on the collected historical feedback data and alarm patterns, the intelligent analysis module periodically self-optimizes and dynamically adjusts the filtering strategy to improve the accuracy of filtering and the overall performance of the system; A self-feedback intelligent evolution mechanism that not only relies on manual marking but also includes correction suggestions automatically generated by the system to help operation and maintenance personnel adjust and optimize alarm filtering rules more quickly and effectively.

6. The alarm data processing system for communication network distribution management according to claim 1, characterized in that, An edge computing module is also integrated in the system architecture to process some alarm data in real time at the source of data generation (such as the device side), perform preliminary filtering and preprocessing, thereby reducing the transmission burden of alarm data and improving the response speed of the system, especially suitable for alarm processing in a large-scale distributed network environment.

7. The alarm data processing system for communication network distribution management according to claim 1, wherein The intelligent analysis module further includes a deep reinforcement learning unit that uses deep reinforcement learning algorithms to optimize the alarm filtering decision-making strategy through continuous interaction with the environment, enabling the system to automatically learn and adjust the filtering strategy according to historical feedback and environmental changes, thereby improving the accuracy of alarms and reducing false alarms and missed alarms.

8. The alarm data processing system for communication network distribution management according to claim 1, wherein The alarm storage module adopts a distributed database architecture, supports high-throughput and low-latency data storage and retrieval, and ensures that the system can still perform historical queries and statistical analysis efficiently and stably in the case of large-scale alarm data through data sharding and replica mechanisms.

9. The alarm data processing system for communication network distribution management according to claim 1, characterized in that The system architecture is designed as a distributed self-optimizing architecture that can dynamically scale horizontally according to alarm traffic and network status to ensure stable operation of the system in the case of massive data; moreover, the architecture has an adaptive high-availability design that can dynamically adjust the fault tolerance mechanism according to traffic changes and node status, such as automatically adding processing nodes under high load or quickly switching to standby nodes in case of node failure, ensuring the high availability and continuous stable operation of the system.

10. A method for analyzing and processing alarm data in communication distribution network management, characterized in that, The steps are as follows: Step 1: The system automatically analyzes an alarm table according to whether the alarm affects the device status and business; Step 2: When receiving the alarm table, the system automatically analyzes and queries, classifies the information in the alarm table as the to-be-stored state, and the whole process of this table's automatic processing can be intervened by humans, and useless alarms are filtered out; Step 3: Judge whether the new alarm is stored in the database according to the set alarm filtering rules (this process can be automatically processed by the system or intervened by humans). If so, execute Step 4; if not, return to Step 2; Step 4: Compare the alarm data in the database. If there is an uneliminated alarm with the same name in the database, only retain the alarm information that has been stored in the database previously, record the update time of the alarm, and complete the aggregation process of the alarm; Step 5: Determine whether there is alarm data that recurs and is eliminated. If so, execute Step 6; if not, return to Step 3; Step 6: For alarm data that recurs and is eliminated, during the process of the system handling the alarm, give it a judgment time. Within this judgment period, regardless of how many times the alarm of this device occurs and is eliminated, it is defaulted that there is only one alarm information. When the elimination entry of this alarm is received for the last time, start the alarm buffer time. If no more alarms of this type are received within the buffer time, the system determines that this alarm is truly eliminated after the buffer time ends and eliminates this alarm; otherwise, it is considered that the alarm always exists and is uneliminated during the jitter process; Step 7: Collect on-site alarm information and redefine the basic alarm level according to the impact of the alarm on the service; Step 8: Distinguish specific situations such as single-end break, double-end break, disconnection, and power failure according to the on-site link type. When an alarm is received, query the device status and match the impact of the alarm on the device status according to the device status. If there are alarms at both ends and it affects the service, the alarm level will be upgraded; Step 9: According to the system link fault continuity and relevance judgment rules, which are the system processing logics sorted out by our side in combination with information such as the alarm situation and distance of on-site devices, automatically and accurately generate the root cause of the fault.

Citation Information

Patent Citations

  • Centralized alarm monitoring system and method of power system terminal communication access network

    CN107196804A

  • Integrated centralized alarm automatic processing system and method

    CN111585782A