Alarm management system, method and device and storage medium
By collecting and cleaning multi-source alarm records, applying machine learning and automated workflows, establishing a historical fault case library, solving the problem that alarm information in the existing technology fails to effectively identify mutually influencing relationships, and achieving efficient alarm management and rapid business recovery.
Patent Information
- Application Number
- CN202510167358.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing operation and maintenance monitoring system cannot effectively identify the mutual influence relationship between multi-source alarms, resulting in low analysis and troubleshooting efficiency, lack of automation integration capabilities, unable to achieve real-time automatic recovery of business availability, long response time, and alarm information does not fully utilize historical faults and expert experience value, and low information utilization rate.
By collecting alarm records from different monitoring tools and logs, performing data cleaning and deduplication strategies, aggregating low-level alarm records, applying machine learning model optimization strategies, establishing a library of historical failure cases, and implementing automated workflow modes to achieve automated integration and real-time response of alarm management systems.
It improves the efficiency of analysis and troubleshooting, reduces the average recovery time of faults, realizes automation integration capabilities, ensures immediate response to alarm events, accelerates business recovery progress, fully integrates historical fault cases and expert operation and maintenance knowledge, and improves operation and maintenance skills inheritance and team growth.
Smart Images

Figure CN120066897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of alarm management, and specifically to an alarm management system, method, device, and storage medium. Background Art
[0002] With the development of information technology, the complexity of IT systems has been continuously increasing, resulting in a huge number of alarm messages generated during the operation and maintenance process and diverse sources. Traditional manual troubleshooting and handling are inefficient, prone to causing delays in critical alarms, and affecting business continuity and user experience.
[0003] Existing operation and maintenance monitoring systems usually rely on single rule matching for alarm classification and filtering, lacking in-depth correlation analysis and intelligent processing mechanisms for alarm messages. Defects of the existing technology: This traditional processing mode cannot effectively identify the mutual influence relationships among multi-source alarms, reducing the accuracy of alarms and the efficiency of analysis and troubleshooting, and it is difficult to meet the high-efficiency operation and maintenance requirements of large-scale distributed systems.
[0004] This solution proposes an alarm management system, method, device, and storage medium to solve the problems that in a large amount of scattered alarm data, the mutual influence relationships cannot be effectively identified, resulting in low analysis and troubleshooting efficiency; lacking automation integration capabilities, unable to achieve real-time automatic recovery of business availability, with a long response time; and the alarm messages do not fully utilize the value of historical faults and expert experience, with low information utilization rate. Summary of the Invention
[0005] The present invention provides an alarm management system, method, device, and storage medium, which contribute to solving the problems mentioned in the above background art.
[0006] In a first aspect, the present application provides an alarm management method, adopting the following technical solution: An alarm management method includes: Step 1: Collect alarm records from different monitoring tools and logs; By collecting alarm records from different monitoring tools and logs, the system can obtain multi-dimensional fault information, ensuring coverage of potential problems at each technical level. The multi-channel data collection method improves the system's perception ability, avoids information blind spots, and enhances the early warning function. This not only makes fault detection more accurate but also helps the system identify and take actions in a timely manner before problems occur, reducing the losses caused by faults.
[0007] Step 2: Execute a preset data cleaning strategy for each alarm record to remove duplicate alarm records and eliminate invalid alarm records based on a preset rule for invalid alarm fields; Step 3: Classify the alarm records into high-level alarm records and low-level alarm records according to the alarm level, where high-level alarm records need to be processed immediately; For low-level alarm records, execute the aggregation and classification strategy to merge low-level alarm records that can be processed simultaneously into one alarm record; Step 4: For high-level alarm records and the alarm records merged from low-level alarm records, execute the machine learning model optimization strategy and apply the optimization results of the model to the operation of the alarm management system; Step 5: Execute the historical fault case library reserve strategy, collect historical faults and solutions, and form a fault case library; Step 6: Execute the automated workflow mode, connect to various automated repair tools, and repair faults.
[0008] Preferably, for each alarm record, execute the preset data cleaning strategy to remove duplicate alarm records, and based on the preset list of invalid alarm features, eliminate invalid alarm records, including: Duplicate removal processing: Set specific fields for judging whether alarm records are equal to obtain field set A; ; Perform hash processing on each alarm record, and the calculation method is as follows: , and the result is recorded as the hash value of the alarm record; Obtain the hash values of all alarm records, identify alarm records with the same hash value as the same alarm records, and remove duplicate alarm records; Record each record after removing duplicate alarm records as the first alarm record; Invalid alarm filtering: Set the rules for invalid alarm fields; For each first alarm record, judge whether there are fields in the first alarm record that conform to the rules for invalid alarm fields; If so, determine that the first alarm record is invalid and delete the invalid first alarm record; Invalid alarm filtering: Set the rules for invalid alarm fields; For each first alarm record, judge whether there are fields in the first alarm record that conform to the rules for invalid alarm fields; If so, determine that the first alarm record is invalid and delete the invalid first alarm record.
[0009] By eliminating duplicate and invalid alarms through data cleaning and duplicate removal strategies, the system response is ensured to be more accurate and efficient. The duplicate removal operation avoids the processing of redundant alarms through the hash algorithm and reduces false alarms. The filtering of invalid alarms clears alarms without practical significance through rules, ensuring that operation and maintenance personnel can focus on the problems that really need to be processed, thereby improving the processing efficiency and accuracy of the system.
[0010] Preferably, for low-level alarm records, an aggregation and classification strategy is executed to merge low-level alarm records that can be processed simultaneously into one type of alarm record, including: Classify all low-level alarm records according to the alarm category, specifically: , where is the category of low-level alarm records, is the alarm category, C is the number of low-level alarm records, D is the number of alarm categories, and δ is an indicator function used to judge and are equal. If they are equal, return 1; if not, return 0; For any one alarm category, execute a time-window-based aggregation determination strategy to determine whether all low-level alarm records in the same alarm category can be processed as one alarm record, specifically: Set the duration size f of the time window; Calculation formula , where is the alarm category, is a time interval representing the start time of the time period and the end time , is the aggregation result at time t. Each low-level alarm category in the aggregation result is and the time stamp of each low-level alarm category is within the time interval; Recognize all low-level alarm records in the aggregation result as one alarm record.
[0011] The low-level alarm aggregation strategy merges similar or duplicate alarms, reducing the number of alarms and improving the efficiency of alarm classification and processing. By combining the time window and alarm category, the system can automatically determine which alarms can be merged for processing, thus reducing redundant operations. In this way, the operation and maintenance personnel can focus on more important high-priority alarms, improving the overall efficiency and response speed of alarm management.
[0012] Preferably, for high-level alarm records and alarm records merged from low-level alarm records, execute a machine learning model optimization strategy and apply the optimization result of the model to the operation of the alarm management system, including: Data preparation: For any one alarm record among high-level alarm records and alarm records merged from low-level alarm records, extract the features of the alarm record. The features include: Alarm level: ; Alarm category: ; Timestamp: T; Identification of the associated system: ; Keywords of the alarm content: M; Occurrence frequency: ; Establish a feature set for alarm records: Map the alarm level to a numerical value: , where 1 represents a high-level alarm record and 2 represents a low-level alarm record; One-hot encode the alarm category: ; Process the time interval as a sequence feature: ; Alarm keyword embedding: Generate keyword vectors using the word embedding algorithm: ; Statistical feature extraction: ; Convert the alarm record into a feature set: , and input it into the machine learning model.
[0013] Preferably, for high-level alarm records and alarm records merged from low-level alarm records, execute the machine learning model optimization strategy, and apply the optimization result of the model to the work of the alarm management system, including: Alarm association rule mining: Obtain all alarm records to form an alarm record set ; Arbitrarily obtain two elements in the alarm record set and ; Calculate ; Execute the following formula to calculate the confidence level, and the confidence level is When it occurs, Occurrence probability; Set the confidence level threshold; Compare the confidence levels of the alarm records and with the confidence level threshold; If the confidence levels of the alarm records and are greater than the confidence level threshold, then when the alarm record occurs, is given priority for processing; Priority prediction: , where u is the number of alarm records, K is the number of priority categories, indicates whether the i-th alarm record belongs to alarm record k, is the probability that the model predicts the i-th alarm record belongs to the alarm category k; Obtain the priority of each predicted alarm record, sort the alarm records from largest to smallest according to the priority, and process the alarm records in the order of sorting.
[0014] Enable the alarm management system to have self-adaptive capabilities through machine learning optimization strategies, and improve the accuracy of alarm prediction by analyzing historical data. The machine learning model can accurately identify fault patterns, reduce misjudgments, and automatically optimize alarm classification and response strategies, thereby improving the system's decision-making ability and response speed. As the system continues to learn, the accuracy of the model will gradually increase, ensuring more intelligent and precise alarm handling.
[0015] Preferably, implement the historical fault case library reserve strategy, collect historical faults and solutions, and form a fault case library, including: Collect historical fault cases: Represent each historical fault case as a feature vector , where represents the characteristics of the fault case; Obtain all historical fault cases to form a set , m is the total number of historical fault cases, is the feature vector of the fault case; Refinement of the expert operation and maintenance manual: Map expert experience to a rule set , where each rule defines a processing solution for a specific fault situation; Record the current fault as ; Execute the following formula to calculate the similarity metric function , where is the value of the k-th feature in the fault case , is the value of the k-th feature in the current fault , is the feature 's weight, indicating the importance of this feature for fault diagnosis; Calculate the optimal fault solution: Optimal solution = , where is used to measure the similarity between the expert rule and the historical fault case .
[0016] The establishment of the historical fault case library enables the alarm management system to quickly diagnose current faults using existing experience. By calculating similarities, the system can automatically match solutions to similar problems in history, accelerating the fault handling process. This approach reduces manual intervention, improves the efficiency of fault location and resolution, ensures that faults can be repaired quickly and accurately, and avoids the occurrence of repeated errors.
[0017] Preferably, the execution of the automated workflow mode interfaces with various automated repair tools to repair faults, including: The automated workflow mode is specifically as follows: Set a standardized message queue protocol; Establish connections between the alarm management system and external automated scripts and third-party service providers through the message queue protocol; When an alarm record occurs, the alarm management system triggers the adaptive workflow engine, and at the same time obtains the alarm type and impact scope of the alarm record, and selects an operation instruction set to solve the current alarm record.
[0018] The automated workflow interfaces with external automated tools through the message queue protocol, realizing the automation of fault handling. When an alarm is triggered, the system can automatically select an appropriate repair instruction set according to the alarm type, reducing the time and error rate of manual operations. The automated workflow makes the fault response faster and the processing process more precise, reducing the risks brought by human factors and improving the overall efficiency of fault repair.
[0019] In a second aspect, the present application provides an alarm management system, adopting the following technical solution: An alarm management system includes: An alarm data collection and management system that processes the collection, cleaning, deduplication, and filtering of alarms; An alarm classification and aggregation module that processes alarms according to the alarm level and aggregation strategy; A machine learning optimization module that optimizes the alarm management system according to a machine learning model; A historical fault case library management module that provides a fault case library and expert experience support; An automated workflow engine and execution module that realizes an automated repair process; An alarm priority prediction and association rule mining module that performs priority prediction and association analysis on alarms; A user interface and monitoring panel that provides a user interface for operation and maintenance personnel to manage and monitor.
[0020] In a third aspect, the present application provides a device, adopting the following technical solution: A device; includes: A data collection device: servers, network devices, IoT devices, etc.; Data processing and storage devices: computing servers, databases, distributed storage systems, etc.; Network and message queue devices: message queues, API gateways, etc.; Automated repair devices: automated tools, remote control devices, etc.; Machine learning and optimization devices: training servers, cluster computing devices, etc.; User interface devices: workstations, monitors, notification devices, etc.; Security devices: firewalls, IDS / IPS systems, etc.; Alarm history case library devices: document management systems and knowledge bases.
[0021] Fourthly, the present application provides a computer-readable storage medium, adopting the following technical solution: A computer-readable storage medium, comprising: executable program instructions stored thereon, characterized in that when the executable program instructions are executed by a processor, the alarm management method according to any one of claims 1 to 7 is implemented.
[0022] The present invention has the following beneficial effects: 1. The alarm management method adopts deep mining and correlation analysis techniques for alarm information, achieving the effects of improving the analysis and troubleshooting efficiency and reducing the mean time to recover from faults. It realizes the automated integration ability, achieving the goal of instant response to alarm events and accelerating the business recovery progress. It fully integrates historical fault cases and expert operation and maintenance knowledge, improves the value conversion of alarm information, and promotes the inheritance of operation and maintenance skills and the growth of the team.
[0023] 2. The alarm management method collects alarm records from different monitoring tools and logs, ensuring the comprehensiveness and diversity of alarm data, and can cover fault information at all levels, including multiple fields such as networks, servers, applications, databases, etc. This multi-dimensional data collection provides a global perspective on faults, ensuring that the system can perceive in real time regardless of where the fault occurs. Moreover, by combining data from multiple sources, it can avoid the blind spots that may be brought by a single data source, improving the alarm accuracy and comprehensiveness of the system.
[0024] 3. The alarm management method performs data cleaning and deduplication, which can greatly reduce redundant data and improve the response efficiency of the alarm system. The deduplication operation judges duplicate alarms through a hash algorithm, avoiding the interference of duplicate alarm information, enabling operation and maintenance personnel to focus on the alarms that really need to be processed. By eliminating invalid alarm records, unnecessary alarm processing and resource consumption can be reduced, thereby improving the overall working efficiency of the system. In addition, data cleaning can also reduce the workload of operation and maintenance personnel and avoid judgment errors caused by redundant data.
[0025] 4. For this alarm management method, the aggregation and classification of low-level alarm records can significantly improve the alarm handling efficiency. By merging similar types of alarms into a unified record, it is possible to avoid repeatedly responding to multiple similar alarms, reducing the system burden. This strategy is particularly applicable to high-frequency and less impactful alarms. Through time-window aggregation judgment, it can reduce redundant operations and improve the effective utilization of resources without missing key information. In addition, this aggregation method enables operation and maintenance personnel to focus on core issues, thereby enhancing the response speed and resolution efficiency.
[0026] 5. For this alarm management method, the application of machine learning optimization strategies enables the alarm management system to continuously self-learn and optimize, improving the accuracy of alarm recognition. By training machine learning models, the alarm system can analyze the priority, type, and handling methods of alarms based on historical data, providing intelligent predictions for future alarms. Machine learning can reduce human errors and subjective judgments, making alarm handling more efficient and intelligent. At the same time, it can also adjust alarm handling strategies in real time to ensure the adaptability and flexibility of the handling process, further enhancing the reliability and response speed of the system.
[0027] 6. For this alarm management method, the historical fault case library reserve strategy can effectively accumulate experience and improve the efficiency of fault diagnosis and handling. By converting historical fault cases into feature vectors and establishing a fault case library, the alarm management system can quickly find matching solutions when encountering similar faults. This not only speeds up the response speed of fault handling but also avoids repetitive errors and improves the system's self-healing ability for faults. In addition, the extraction of expert experience enables the system to have the ability of self-improvement and fault prediction, further enhancing the accuracy and efficiency of fault handling.
[0028] 7. For this alarm management method, the implementation of an automated workflow can significantly improve the efficiency of fault response. Through a standardized message queue protocol, the alarm management system can seamlessly interface with external automated scripts and third-party repair tools to achieve automatic triggering and handling of alarm information. This strategy not only reduces the interference of human operations and the risk of human errors but also automatically selects the most appropriate repair operation according to the type and level of the alarm, making fault handling faster and more accurate. Through the automated workflow, fault repair can be quickly launched, enhancing the overall efficiency and response ability of the operation and maintenance team. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the method of the present invention.
[0030] Figure 2 It is a schematic diagram of the system of the present invention.
[0031] Figure 3 It is a schematic diagram of the process of establishing a fault case library of the present invention.
[0032] Figure 4 This is a schematic diagram of the process for the automated workflow mode implemented by the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] Embodiment 1, referring to Figure 1 , an alarm management method, step 1: collect alarm records from different monitoring tools and logs; Step 2: Execute a preset data cleaning strategy for each alarm record, remove duplicate alarm records, and based on the preset rules for invalid alarm fields, eliminate invalid alarm records; Step 3: Classify the alarm records into high-level alarm records and low-level alarm records according to the alarm level. Among them, high-level alarm records need to be processed immediately; For low-level alarm records, execute an aggregation and classification strategy to merge low-level alarm records that can be processed simultaneously into one alarm record; Step 4: For high-level alarm records and alarm records merged from low-level alarm records, execute a machine learning model optimization strategy and apply the optimization results of the model to the work of the alarm management system; Step 5: Execute a historical fault case library storage strategy, collect historical faults and solutions, and form a fault case library; Step 6: Execute an automated workflow mode to connect to various automated repair tools to repair faults.
[0035] In this embodiment, in a data center operation and maintenance scenario, the alarm management system first collects raw alarm information from multiple channels such as network monitoring and server health detection. Through data cleaning strategies, such as deduplication within a time window, false alarms or duplicate alarms are excluded to ensure the validity of alarm data. Subsequently, primary aggregation and screening are performed using the grading criteria of alarm levels and predefined time series analysis. Then, a pre-trained deep neural network model is used to perform advanced correlation analysis on the alarm information, extract potential influencing factors and trend changes, and automatically adjust the severity level of the alarms. In addition, an expert knowledge base is introduced. In cases where the machine learning model cannot cover, through a human-machine collaboration method, the experience and wisdom of senior operation and maintenance personnel are utilized to diagnose complex abnormal situations. Finally, an automated operation and maintenance workflow is triggered, such as restarting services and adjusting load balancing, to quickly respond to alarm events and ensure the stable operation of the business.
[0036] Executing a preset data cleaning strategy for each alarm record, removing duplicate alarm records, and eliminating invalid alarm records based on a preset list of invalid alarm features, including: Deduplication processing: Set specific fields for determining whether alarm records are equal to obtain a field set A; ; Perform hash processing on each alarm record, and the calculation method is as follows: , and the result is recorded as the hash value of the alarm record; Obtain the hash values of all alarm records, identify alarm records with the same hash value as the same alarm records, and remove duplicate alarm records; Record each record after removing duplicate alarm records as the first alarm record; Invalid alarm filtering: Set rules for invalid alarm fields; For each first alarm record, determine whether there are fields in the first alarm record that conform to the rules for invalid alarm fields; If so, determine that the first alarm record is invalid and delete the invalid first alarm record.
[0037] Performing data cleaning and deduplication can greatly reduce redundant data and improve the response efficiency of the alarm system. The deduplication operation uses a hash algorithm to judge duplicate alarms, avoiding the interference of duplicate alarm information, enabling operation and maintenance personnel to focus on the alarms that really need to be processed. By eliminating invalid alarm records, unnecessary alarm processing and resource consumption can be reduced, thereby improving the overall working efficiency of the system. In addition, data cleaning can also reduce the workload of operation and maintenance personnel and avoid judgment errors caused by redundant data.
[0038] For the low - level alarm records, execute the aggregation and classification strategy, and merge the low - level alarm records that can be processed simultaneously into one type of alarm record, including: Classify all the low - level alarm records according to the alarm category. Specifically: , where is the category of the low - level alarm record, is the alarm category, C is the number of low - level alarm records, D is the number of alarm categories, and δ is an indicator function used to judge and are equal. If they are equal, return 1; if not, return 0; For any one alarm category, execute the time - window - based aggregation determination strategy to determine whether all the low - level alarm records in the same alarm category can be processed as one alarm record. Specifically: Set the duration size f of the time window; Calculation formula , where is the alarm category, is a time interval representing the start time of the time period and the end time , is the aggregation result at time t. Each low - level alarm category in the aggregation result is and the time stamp of each low - level alarm category is within the time interval; Recognize all the low - level alarm records in the aggregation result as one alarm record.
[0039] The aggregation and classification of low - level alarm records can significantly improve the alarm processing efficiency. By merging similar alarms into a unified record, it is possible to avoid repeated responses to multiple similar alarms and reduce the system burden. This strategy is particularly suitable for high - frequency and less - impactful alarms. Through time - window aggregation judgment, redundant operations can be reduced and the effective utilization of resources can be improved without missing key information. In addition, this aggregation method enables operation and maintenance personnel to focus on core issues, thereby improving the response speed and resolution efficiency.
[0040] For the high - level alarm records and the alarm records merged from low - level alarm records, execute the machine - learning - model optimization strategy and apply the optimization result of the model to the operation of the alarm management system, including: Data preparation: For any one alarm record among the high - level alarm records and the alarm records merged from low - level alarm records, extract the features of the alarm record. The features include: Alarm level: ; Alarm category: ; Timestamp: T; Identification of associated system: ; Keywords of alarm content: M; Occurrence frequency: ; Feature set for establishing alarm records: Alarm level mapped to a numerical value: , where 1 represents a high-level alarm record and 2 represents a low-level alarm record; One-hot encoding of alarm categories: ; Time interval processed as a sequence feature: ; Alarm keyword embedding: Generating keyword vectors using the word embedding algorithm: ; Statistical feature extraction: ; Converting alarm records into a feature set: , and inputting it into a machine learning model.
[0041] For the high-level alarm records and the alarm records merged from low-level alarm records, execute the machine learning model optimization strategy, and apply the optimization results of the model to the work of the alarm management system, including: Alarm association rule mining: Obtain all alarm records to form an alarm record set ; Randomly obtain two elements in the alarm record set and ; Calculate ; Execute the following formula to calculate the confidence level, where the confidence level is When it occurs, Occurrence probability; Set the confidence level threshold; Compare the confidence levels of the alarm records and with the confidence level threshold; If the confidence levels of the alarm records and are greater than the confidence level threshold, then when the alarm record occurs, give priority to processing ; Priority prediction: , where u is the number of alarm records, K is the number of priority categories, indicates whether the i-th alarm record belongs to the alarm record k, is the probability that the model predicts the i-th alarm record belongs to the alarm category k; Obtain the priority of each predicted alarm record, sort the alarm records from largest to smallest according to the priority, and process the alarm records in the order of sorting.
[0042] Execute the historical fault case library reserve strategy, collect historical faults and solutions, and form a fault case library, including: In this embodiment, refer to Figure 3 ; Collect historical fault cases: Represent each historical fault case as a feature vector , where represents the features of the fault case; Obtain all historical fault cases to form a set , m is the total number of historical fault cases, is the feature vector of the fault case; Refinement of the expert operation and maintenance manual: Map expert experience to a rule set , where each rule defines a processing solution for a specific fault situation; Record the current fault as ; Execute the following formula to calculate the similarity metric function , where is the value of the k-th feature in the fault case , is the value of the k-th feature in the current fault , is the weight of the feature , indicating the importance of the feature for fault diagnosis; Calculate the optimal fault solution: Optimal solution = , where is used to measure the similarity between the expert rule and the historical fault case .
[0043] In this embodiment, past operation and maintenance records are collected, especially those events that have caused large-scale failures. Common warning indicators and technical countermeasures are summarized to form a local or cloud-based historical case database. At the same time, regular training meetings are organized to invite front-line operation and maintenance personnel to share practical experience and best practices, which are sorted into written documents and incorporated into the online operation and maintenance manual. Based on this, a cognitive computing framework with natural language understanding and reasoning capabilities is developed. It can quickly retrieve successful solutions in similar scenarios while receiving new alarms, provide them for front-line engineers for reference, or directly incorporate them into the automated decision-making process to accelerate the troubleshooting process.
[0044] The reserve strategy of the historical fault case database can effectively accumulate experience and improve the efficiency of fault diagnosis and handling. By converting historical fault cases into feature vectors and establishing a fault case database, the alarm management system can quickly find matching solutions when encountering similar faults. This not only speeds up the response speed of fault handling but also avoids repetitive errors and improves the system's self-healing ability for faults. In addition, the refinement of expert experience enables the system to have the ability of self-improvement and fault prediction, further enhancing the accuracy and efficiency of fault handling.
[0045] Execute the automated workflow mode to connect to various automated repair tools to repair faults, including: The automated workflow mode is specifically as follows: Set a standardized message queue protocol; Establish connections between the alarm management system and external automated scripts and third-party service providers through the message queue protocol; When an alarm record occurs, the alarm management system triggers the adaptive workflow engine, and at the same time obtains the alarm type and impact scope of the alarm record, and selects an operation instruction set to solve the current alarm record.
[0046] In this embodiment, refer to Figure 4 , execute the schematic diagram of the automated workflow mode process.
[0047] In this embodiment, a set of standardized message queue protocols are established, allowing the alarm management system to seamlessly connect with external automated scripts and third-party service providers. When an alarm occurs, the system immediately triggers the adaptive workflow engine and automatically selects the most suitable operation instruction set according to the alarm type and impact scope. For example, for an alarm about a decrease in database performance, an index optimization script can be started; in the face of a high-incidence period of network latency, call the traffic scheduling service of the cloud service provider to flexibly allocate network resources. The entire process not only reduces the need for human intervention but also can continuously evaluate the effect of automated processing through a real-time feedback mechanism and continuously iterate and optimize the strategy.
[0048] The implementation of an automated workflow can significantly improve the efficiency of fault response. Through a standardized message queue protocol, the alarm management system can seamlessly interface with external automated scripts and third-party repair tools to achieve the automatic triggering and processing of alarm information. This strategy not only reduces the interference of human operations and the risk of human errors but also automatically selects the most appropriate repair operation based on the type and level of the alarm, making fault handling faster and more accurate. Through the automated workflow, fault repair can be quickly launched, enhancing the overall efficiency and response ability of the operation and maintenance team.
[0049] Embodiment 2, referring to Figure 2 , an alarm management system, characterized by comprising: An alarm data collection and management system, which processes the collection, cleaning, deduplication, and filtering of alarms; An alarm classification and aggregation module, which processes alarms according to the alarm level and aggregation strategy; A machine learning optimization module, which optimizes the alarm management system according to a machine learning model; A historical fault case library management module, which provides a fault case library and expert experience support; An automated workflow engine and execution module, which implements an automated repair process; An alarm priority prediction and association rule mining module, which performs priority prediction and association analysis on alarms; A user interface and monitoring panel, which provides a user interface for operation and maintenance personnel to manage and monitor.
[0050] Through the design of an automated alarm processing mechanism and its workflow engine, the fault response speed and repair success rate have been significantly improved. The alarm intelligent analysis architecture that integrates a knowledge graph and an expert system enhances the anomaly recognition ability and fault location accuracy. An efficient alarm data cleaning and preliminary clustering algorithm reduces the interference of invalid and redundant information and improves the quality of subsequent analysis.
[0051] Embodiment 3, referring to Figure 3 , a device, characterized by comprising: Data collection devices: servers, network devices, IoT devices, etc.; Data processing and storage devices: computing servers, databases, distributed storage systems, etc.; Network and message queue devices: message queues, API gateways, etc.; Automated repair devices: automated tools, remote control devices, etc.; Machine learning and optimization devices: training servers, cluster computing devices, etc.; User interface devices: workstations, monitors, notification devices, etc.; Security devices: firewalls, IDS / IPS systems, etc.; Alarm historical case library device: document management system and knowledge base.
[0052] Embodiment 4. A computer-readable storage medium, characterized in that executable program instructions are stored thereon, and the executable program instructions, when executed by a processor, implement the alarm management method according to any one of claims 1 to 7.
[0053] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0054] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An alarm management method, characterized in that: include: Step 1: Collect alarm records from different monitoring tools and logs; Step 2: Execute the preset data cleaning strategy for each alarm record, remove duplicate alarm records, and eliminate invalid alarm records based on the preset invalid alarm field rules; Step 3: Classify the alarm records into high-level alarm records and low-level alarm records according to the alarm level, among which the high-level alarm records need to be processed immediately; For low-level alarm records, an aggregation classification strategy is implemented to merge low-level alarm records that can be processed simultaneously into one alarm record; Step 4: Execute the machine learning model optimization strategy for high-level alarm records and alarm records merged from low-level alarm records, and apply the optimization results of the model to the alarm management system; Step 5: Execute the historical fault case library reserve strategy, collect historical faults and solutions, and form a fault case library; Step 6: Execute the automated workflow mode, connect to various automated repair tools, and repair faults.
2. The alarm management method according to claim 1, characterized in that: The method of executing a preset data cleaning strategy on each alarm record, removing duplicate alarm records, and eliminating invalid alarm records based on a preset invalid alarm feature list includes: Deduplication processing: Set a specific field used to determine whether the alarm records are equal, and obtain a field set A; ; Hash processing is performed on each alarm record, and the calculation method is as follows: , the result is recorded as the hash value of the alarm record; Obtain the hash values of all alarm records, identify alarm records with the same hash value as the same alarm record, and remove duplicate alarm records; Recording each record from which duplicate alarm records are removed as a first alarm record; Invalid alarm filtering: Set invalid alarm field rules; For each first alarm record, determining whether there is a field in the first alarm record that complies with an invalid alarm field rule; If it exists, the first alarm record is deemed invalid and the invalid first alarm record is deleted.
3. The alarm management method according to claim 1, characterized in that: The aggregation classification strategy is executed for the low-level alarm records, and the low-level alarm records that can be processed simultaneously are merged into one type of alarm records, including: All low-level alarm records are classified according to the alarm category, specifically: ,in, The category of low-level alarm records. is the alarm category, C is the number of low-level alarm records, D is the number of alarm categories, and δ is the indicator function used to judge and Are they equal? If they are equal, it returns 1; if they are not equal, it returns 0; For any alarm category, execute the time window aggregation decision strategy to determine whether all low-level alarm records in the same alarm category can be processed as one alarm record. Specifically: Set the duration f of the time window; Calculation formula ,in, is the alarm category, A time interval, indicating the start time of the time period and end time , is the aggregation result at time t, and each low-level alarm category in the aggregation result is And the timestamp of each low-level alarm category All in time interval middle; All low-level alarm records in the aggregation result are considered as one alarm record.
4. The alarm management method according to claim 1, characterized in that: The method of executing a machine learning model optimization strategy for high-level alarm records and alarm records formed by merging low-level alarm records, and applying the optimization results of the model to the alarm management system includes: Data preparation: For any one of the high-level alarm records and the alarm records formed by merging the low-level alarm records, extract features of the alarm record, the features including: Alarm level: ; Alarm category: ; Timestamp: T; Identification of the associated system: ; Alarm content keywords: M; Occurrence frequency: ; Create a feature set for alarm records: Alarm levels are mapped to numeric values: , where 1 represents a high-level alarm record and 2 represents a low-level alarm record; Alarm categories are one-hot encoded: ; Time intervals are treated as sequence features: ; Alarm keyword embedding: Generate keyword vectors using word embedding algorithm: ; Statistical feature extraction: ; Convert alarm records into feature sets: , which is input into the machine learning model.
5. The alarm management method according to claim 1, characterized in that: The method of executing a machine learning model optimization strategy for high-level alarm records and alarm records formed by merging low-level alarm records, and applying the optimization results of the model to the alarm management system includes: Alarm association rule mining: Get all alarm records and form an alarm record set ; Get any two elements from the alarm record set and ; calculate ; Execute the following formula to calculate the confidence level, which is When it happens, Probability of occurrence; Set confidence threshold; Record the alarm and The confidence level is compared with the confidence threshold; If the alarm record and If the confidence level of the alarm record is greater than the confidence threshold, When it happens, give priority to ; Priority prediction: , where u is the number of alarm records, K is the number of priority categories, is whether the i-th alarm record belongs to alarm record k, The model predicts the probability that the i-th alarm record belongs to alarm category k; Obtain the priority of each predicted alarm record, sort the alarm records from high to low according to the priority, and process the alarm records in the order of sorting.
6. The alarm management method according to claim 1, characterized in that: The execution of the historical fault case library reserve strategy collects historical faults and solutions to form a fault case library, including: Collect historical fault cases: Represent each failure case in history as a feature vector ,in Characterize the failure case; Get all the historical fault cases to form a collection , m is the total number of historical failure cases, is the feature vector of the fault case; Expert operation and maintenance manual refinement: Mapping expert experience into rule sets , where each rule Defines a solution for a specific fault scenario; Record the current fault as ; Execute the following formula to calculate the similarity measurement function ,in, For failure cases The value of the kth feature in , Is the current fault The value of the kth feature in , It is a feature The weight of represents the importance of this feature to fault diagnosis; Calculate the optimal fault solution: Optimal solution = ,in, Used to measure expert rules and historical failure cases similarity.
7. The alarm management method according to claim 1, characterized in that: The automated workflow mode is executed to connect to various automated repair tools to repair faults, including: The automated workflow mode is specifically: Set up a standardized message queue protocol; Connect the alarm management system with external automation scripts and third-party service providers through the message queue protocol; When an alarm record occurs, the alarm management system triggers the adaptive workflow engine, obtains the alarm type and impact scope of the alarm record, and selects an operation instruction set to resolve the current alarm record.
8. An alarm management system, characterized in that: include: Alarm data collection and management system, handling alarm collection, cleaning, deduplication and filtering; The alarm classification and aggregation module processes alarms according to the alarm level and aggregation strategy; Machine learning optimization module, which optimizes the alarm management system based on the machine learning model; Historical fault case library management module, providing fault case library and expert experience support; Automated workflow engine and execution module to realize automated repair process; Alarm priority prediction and association rule mining module, which performs priority prediction and association analysis on alarms; The user interface and monitoring panel provide a management and monitoring user interface for operation and maintenance personnel.
9. A device, characterized in that: include: Data collection devices: servers, network equipment, IoT devices, etc. Data processing and storage devices: computing servers, databases, distributed storage systems, etc.; Network and message queue devices: message queues, API gateways, etc. Automated repair devices: automated tools, remote control equipment, etc.; Machine learning and optimization devices: training servers, cluster computing equipment, etc. User interface devices: workstations, displays, notification devices, etc.; Security devices: firewall, IDS / IPS system, etc.; Alarm history case library device: document management system and knowledge base.
10. A computer-readable storage medium, characterized in that: Executable program instructions are stored thereon, and it is characterized in that when the executable program instructions are executed by a processor, the alarm management method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Alarm processing method and device and readable storage medium
CN107908525A
Fault self-recovery system and method based on MySQL database
CN117632651A
Alarm aggregation method and device
CN118316782A
Alarm message processing method and device, electronic equipment and storage medium
CN118779184A
Fault diagnosis and self-healing method and system based on workflow automatic arrangement
CN118860724A
Cited By
Monitoring stability self-adaptive monitoring system based on automatic configuration
CN120750795A