Alarm monitoring method and device based on interference feedback and electronic equipment
By obtaining alarm data and filtering and processing it using preset quantitative rules, calculating the interference duration, and generating a monitoring and tuning strategy, the problem of the failure to effectively quantify monitoring alarm interference in existing technologies is solved, thereby improving operation and maintenance efficiency and system stability.
Patent Information
- Application Number
- CN202511112364.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-14
AI Technical Summary
Existing technologies fail to effectively quantify and evaluate the interference level of monitoring alarms from the operation and maintenance load side, which causes great pressure on operation and maintenance personnel and affects operation and maintenance efficiency and system stability.
By obtaining alarm data, filtering and processing it using preset quantitative rules, calculating the interference duration, and generating a monitoring and tuning strategy when the interference duration exceeds the preset tolerance threshold, the application monitoring platform is guided to perform parameter tuning.
Accurately quantify the impact of alarm interference on operation and maintenance efficiency, reduce invalid alarms, improve operation and maintenance responsiveness and system stability, and reduce the burden of operation and maintenance work.
Smart Images

Figure CN120785718A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, in particular to an alarm monitoring method and device based on interference feedback, and an electronic device. BACKGROUND
[0002] In the rapidly developing digital era, computer information systems are becoming increasingly complex, and the demand for application operation and maintenance and monitoring alarms is becoming more urgent. In the field of application operation and maintenance, monitoring systems have become an indispensable tool for maintaining system stability and security. However, despite the significant progress made in monitoring technology and alarm systems in terms of monitoring capabilities and data collection speeds, the problem of high alarm noise and false alarm rates has not been effectively solved, especially the degree of interference of monitoring alarms from the load end of the operation and maintenance personnel has not been quantitatively interpreted.
[0003] Under the existing technical framework, the analysis of alarm events usually focuses on the number of alarms and the content of alarm information, trying to reduce alarm noise by filtering repeated events and non-critical information. Although this improves the accuracy of the alarm to some extent, it ignores the actual work pressure and interference experienced by the operation and maintenance personnel, making it difficult to fully measure the real effect of the monitoring system. For example, even if the system successfully filters most low-level alarms, if high-level alarms are still frequent and distributed unreasonably, the operation and maintenance personnel will still face continuous high-intensity work pressure, which not only affects the operation and maintenance efficiency, but also may cause human errors and pose a potential threat to the stability of the system.
[0004] The related art fails to effectively quantify the degree of interference of monitoring alarms from the load end of the operation and maintenance, thereby failing to focus on the actual pressure of the operation and maintenance personnel, resulting in a lack of pertinence in alarm management, and thus affecting the operation and maintenance efficiency and the overall management level of the system.
[0005] In view of the above problems, no effective solution has been proposed so far. SUMMARY
[0006] The embodiments of the present application provide an alarm monitoring method and device based on interference feedback, and an electronic device, to at least solve the technical problem that the degree of interference of monitoring alarms cannot be effectively quantified from the load end of the operation and maintenance in the related art, thereby failing to focus on the operation and maintenance pressure, resulting in low operation and maintenance efficiency.
[0007] According to one aspect of an embodiment of the present invention, an alarm monitoring method based on interference feedback is provided, which includes: obtaining alarm data from an application monitoring platform, wherein the alarm data includes at least: alarm period, alarm level, alarm source application and alarm details; filtering and processing the alarm data using preset quantization rules to obtain filtered data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events and exclude plan change interference; calculating interference duration based on the filtered data, wherein the interference duration is used to reflect the workload of the operation and maintenance end; when the interference duration is greater than a preset tolerance threshold, generating a monitoring tuning strategy based on the filtered data, wherein the monitoring tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0008] Furthermore, the alarm data is filtered and processed using preset quantification rules to obtain the steps of filtering data, including: obtaining the alarm time period in the preset quantification rules, wherein the alarm time period is a preset minimum working activity period of the operation and maintenance end; filtering out all data falling within the alarm time period from the alarm data to obtain first pre-selected data; obtaining the alarm level boundary in the preset quantification rules, wherein the alarm level boundary refers to the preset minimum alarm level standard that the operation and maintenance end pays attention to; filtering out all data higher than or equal to the alarm level boundary from the first pre-selected data to obtain second pre-selected data.
[0009] Furthermore, the step of filtering and processing the alarm data using preset quantization rules to obtain filtered data also includes: obtaining an aggregation window in the preset quantization rule, wherein the aggregation window refers to a pre-set time window for performing similar aggregation of alarm events; segmenting the second pre-selected data based on the aggregation window to obtain window data; for each of the window data, analyzing the event types of the N alarm events recorded in the window data, and merging the alarm events of the same type to obtain merged window data containing M of the alarm events, wherein N is a positive integer and M is less than or equal to N; aggregating the merged window data corresponding to all the window data to obtain third pre-selected data.
[0010] Furthermore, the step of filtering the alarm data using preset quantitative rules to obtain filtered data also includes: obtaining the plan change window recorded in the preset quantitative rules, wherein the plan change window refers to the time window occupied by the system plan change activities pre-filed at the operation and maintenance end; for each of the alarm events in the third pre-selected data, checking whether the alarm time period of the alarm event coincides with the plan change window to obtain an inspection result; if the inspection result indicates no overlap, determining the third pre-selected data as the filtered data; or, if the inspection result indicates overlap, excluding the data corresponding to the overlapping time period from the third pre-selected data, and obtaining the filtered data after excluding all data corresponding to the overlapping time periods.
[0011] Furthermore, the step of calculating the interference duration based on the filtered data includes: setting the initial value of the interference duration to 0; traversing all alarm events in the filtered data; during the traversal process, adding the duration of the alarm period triggered by each alarm event to the interference duration to update the interference duration; after the traversal is completed, determining the interference duration.
[0012] Furthermore, the step of generating a monitoring and tuning strategy based on the filtered data includes: analyzing the filtered data within the interference duration to obtain the first alarm number of each of the alarm source applications; when the first alarm number exceeds a preset threshold, judging whether the alarm is a false alarm noise based on the inspection result of the operation log of the alarm source application to obtain a first judgment result; generating the monitoring and tuning strategy based on the first judgment result, wherein the monitoring and tuning strategy includes: adjusting the alarm threshold.
[0013] Furthermore, the step of generating a monitoring and tuning strategy based on the filtered data also includes: analyzing the filtered data within the interference duration to obtain a second alarm number for each type of alarm event; when the second alarm number exceeds a preset threshold, determining whether the event type of the alarm event is an invalid alarm by checking historical alarm records to obtain a second judgment result; generating the monitoring and tuning strategy based on the second judgment result, wherein the monitoring and tuning strategy includes: setting event type filtering conditions.
[0014] According to another aspect of an embodiment of the present invention, an alarm monitoring device based on interference feedback is also provided, which includes: an acquisition unit for acquiring alarm data from an application monitoring platform, wherein the alarm data includes at least: an alarm period, an alarm level, an alarm source application, and alarm details; a screening unit for screening and processing the alarm data using preset quantization rules to obtain screened data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events, and exclude plan change interference; a calculation unit for calculating the interference duration based on the screened data, wherein the interference duration is used to reflect the workload of the operation and maintenance end; a generation unit for generating a monitoring tuning strategy based on the screened data when the interference duration is greater than a preset tolerance threshold, wherein the monitoring tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0015] Furthermore, the screening unit includes: a first acquisition module, used to obtain the alarm time period in the preset quantification rule, wherein the alarm time period is a preset minimum working activity period of the operation and maintenance end; a first screening module, used to screen out all data falling within the alarm time period from the alarm data to obtain first pre-selected data; a second acquisition module, used to obtain the alarm level boundary in the preset quantification rule, wherein the alarm level boundary refers to the preset minimum alarm level standard that the operation and maintenance end pays attention to; a second screening module, used to screen out all data higher than or equal to the alarm level boundary from the first pre-selected data to obtain second pre-selected data.
[0016] Furthermore, the screening unit also includes: a third acquisition module, used to obtain the aggregation window in the preset quantization rule, wherein the aggregation window refers to a pre-set time window for performing similar aggregation of alarm events; a segmentation module, used to segment the second pre-selected data based on the aggregation window to obtain window data; a merging module, used to analyze the event types of the N alarm events recorded in the window data for each of the window data, and merge the alarm events of the same type to obtain merged window data containing M alarm events, wherein N is a positive integer and M is less than or equal to N; an aggregation module, used to aggregate the merged window data corresponding to all the window data to obtain third pre-selected data.
[0017] Furthermore, the screening unit also includes: a fourth acquisition module, used to obtain the plan change window recorded in the preset quantification rule, wherein the plan change window refers to the time window occupied by the system plan change activity pre-filed at the operation and maintenance end; an inspection module, used to check whether the alarm time period of each alarm event in the third pre-selected data coincides with the plan change window, and obtain an inspection result; a first determination module, used to determine the third pre-selected data as the screening data if the inspection result indicates no overlap; and an elimination module, used to eliminate the data corresponding to the overlapping time period from the third pre-selected data if the inspection result indicates overlap, and obtain the screening data after eliminating all data corresponding to the overlapping time periods.
[0018] Furthermore, the calculation unit includes: a setting module for setting the initial value of the interference duration to 0; a traversal module for traversing all alarm events in the filtered data; an accumulation module for adding the duration of the alarm period triggered by each alarm event to the interference duration during the traversal process to update the interference duration; and a second determination module for determining the interference duration after the traversal is completed.
[0019] Furthermore, the generation unit includes: a first analysis module, used to analyze the screening data within the interference duration to obtain the first alarm number of each alarm source application; a first judgment module, used to judge whether the alarm is a false alarm noise based on the inspection result of the operation log of the alarm source application when the first alarm number exceeds a preset threshold, and obtain a first judgment result; a first generation module, used to generate the monitoring and tuning strategy based on the first judgment result, wherein the monitoring and tuning strategy includes: adjusting the alarm threshold.
[0020] Furthermore, the generation unit also includes: a second analysis module, used to analyze the screening data within the interference duration to obtain the second alarm number of each type of alarm event; a second judgment module, used to determine whether the event type of the alarm event is an invalid alarm by checking historical alarm records when the second alarm number exceeds a preset threshold, and obtain a second judgment result; a second generation module, used to generate the monitoring and tuning strategy based on the second judgment result, wherein the monitoring and tuning strategy includes: setting event type filtering conditions.
[0021] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any of the above-mentioned alarm monitoring methods based on interference feedback.
[0022] According to another aspect of an embodiment of the present invention, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-mentioned alarm monitoring methods based on interference feedback.
[0023] In the present invention, an alarm monitoring method based on interference feedback is proposed, which first obtains alarm data from an application monitoring platform, wherein the alarm data at least includes: alarm period, alarm level, alarm source application and alarm details, and then uses preset quantization rules to filter and process the alarm data to obtain filtered data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events and eliminate interference from plan changes, and then calculate the interference duration based on the filtered data, wherein the interference duration is used to reflect the workload of the operation and maintenance end, and finally, when the interference duration is greater than the preset tolerance threshold, a monitoring tuning strategy is generated based on the filtered data, wherein the monitoring tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0024] In the present invention, by constructing a refined interference feedback model, the purpose of accurately quantifying the impact of alarm interference on operation and maintenance efficiency is achieved, thereby achieving the technical effect of improving operation and maintenance response capabilities and overall system stability. Specifically, first, preset quantification rules are used to filter valuable information flows from complex monitoring alarms. For example, alarms during non-critical periods are filtered, repeated events are aggregated to avoid excessive interference, and alarm fluctuations caused by planned maintenance are eliminated, thereby ensuring focus on real and urgent system problems. Subsequently, a data analysis algorithm is used to calculate the actual interference duration of the alarm event. This indicator can directly reflect the load and efficiency of the operation and maintenance work. When it is identified that the interference duration exceeds the preset tolerance threshold, that is, the degree of interference exceeds the tolerance range of normal operation, a targeted monitoring tuning strategy is generated based on the filtered data to guide the application monitoring platform to optimize parameters to reduce invalid alarms, improve the quality and efficiency of alarm event processing, and reduce the burden of operation and maintenance work, thereby significantly enhancing the system's adaptability and the collaborative efficiency of the operation and maintenance team. This solves the technical problem in related technologies that fail to effectively quantify the interference degree of monitoring alarms from the operation and maintenance load end, thereby failing to pay attention to operation and maintenance pressure and resulting in low operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0026] Figure 1 is a schematic diagram of modules of an optional operation and maintenance management system according to an embodiment of the present invention;
[0027] Figure 2 is a flow chart of an optional alarm monitoring method based on interference feedback according to an embodiment of the present invention;
[0028] Figure 3 is a flowchart of an optional method for evaluating an alarm interference degree according to an embodiment of the present invention;
[0029] Figure 4 is a flowchart of an example of an optional method for calculating interference time according to an embodiment of the present invention;
[0030] Figure 5 is a schematic diagram of an optional alarm monitoring device based on interference feedback according to an embodiment of the present invention;
[0031] Figure 6 The figure is a structural block diagram of an electronic device for executing an alarm monitoring method based on interference feedback according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] The following embodiments of the present invention can be applied to various systems / applications / devices that require quantitative assessment of alarm interference and monitoring system tuning, and can implement monitoring alarm management functions based on the actual workload of operation and maintenance personnel. The present invention uses preset quantitative rules to filter and process the collected alarm data, and then calculates the interference duration based on the filtered data to reflect the actual working time of the operation and maintenance personnel affected by handling the alarms; by analyzing the relationship between the interference duration and the preset tolerance threshold, it can automatically or semi-automatically generate a monitoring and tuning strategy when the interference duration exceeds a reasonable range, guiding the application monitoring platform to perform parameter tuning, thereby reducing ineffective interference to the operation and maintenance personnel and improving operation and maintenance efficiency and system stability.
[0035] Specifically, the core of the present invention is to build an operation and maintenance-oriented alarm management mechanism through in-depth analysis and efficient processing of alarm data. This mechanism first ensures that only alarm events that have a substantial impact on operation and maintenance work are included in the analysis scope, avoiding interference caused by non-abnormal alarms such as planned changes; then, by calculating the interference duration, the degree of disruption of the alarm event to the work rhythm of operation and maintenance personnel is intuitively displayed, providing data support for the tuning strategy; in the stage of generating monitoring and tuning strategies, the focus is on identifying and solving the root causes of high-interference alarms. Whether it is optimizing alarm rules, adjusting notification frequency or improving alarm aggregation algorithms, it is all about reducing unnecessary alarms without affecting critical fault responses, thereby achieving the dual goals of improving operation and maintenance efficiency and reducing system false alarm rates.
[0036] The present invention will be described in detail below with reference to various embodiments.
[0037] Example 1
[0038] According to an embodiment of the present invention, an embodiment of an alarm monitoring method based on interference feedback is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0039] The present invention is implemented as follows Figure 2The interference feedback-based alarm monitoring method shown in the figure is implemented by the operation and maintenance management system. It combines data analysis and quantitative evaluation technology and is used in application operation and maintenance and monitoring alarm scenarios, especially for the problem of the impact of excessive alarm interference on operation and maintenance efficiency. By constructing an interference feedback model and generating a monitoring and tuning strategy, it specifically includes the following steps: first, obtaining the original alarm data including the alarm period, alarm level, alarm source application and alarm details from the application monitoring platform; second, using preset quantitative rules to filter and process the alarm data to obtain more accurate filtered data; third, calculating the interference duration based on the filtered data to reflect the workload of the operation and maintenance personnel; finally, when the interference duration exceeds the preset tolerance threshold, generating a monitoring and tuning strategy based on the filtered data to guide the application monitoring platform to perform parameter tuning, so as to achieve the purpose of reducing the workload of the operation and maintenance personnel, improving operation and maintenance efficiency and system stability.
[0040] Figure 1 is a schematic diagram of a module of an optional operation and maintenance management system according to an embodiment of the present invention, such as Figure 1 As shown, the operation and maintenance management system includes: a data acquisition module, a rule management module, a data analysis module and a result output module. The operation and maintenance management system is connected to the application monitoring system, which is described in detail as follows:
[0041] The application monitoring system is used as the alarm data source. By monitoring factors such as log output, resource usage, and performance indicators generated by the computer information system, alarm notifications and alarm data are generated in the event of system abnormalities.
[0042] The data collection module connects to the application monitoring system to obtain alarm data, including but not limited to the application name, alarm title, trigger time, alarm level, and alarm details;
[0043] The rule management module supports setting calculation rules such as time period, filtering conditions, aggregation algorithms, and exception scenarios. It can also set "tolerable time" thresholds based on different scenarios for the data analysis module to call rules;
[0044] The data analysis module is used to analyze the alarm data transmitted by the data collection module, calculate the interruption time of the operation and maintenance end in combination with the rules called from the rule management module, and evaluate whether the interruption time is less than the tolerable time;
[0045] The result output module is used to output the results of the data analysis module. For example, it can be output as a file report for the operation and maintenance side to tune the monitoring parameters of the application monitoring system.
[0046] Figure 2 FIG. 1 is a flow chart of an optional alarm monitoring method based on interference feedback according to an embodiment of the present invention. Figure 2As shown, the method includes the following steps:
[0047] Step S201: Acquire alarm data from the application monitoring platform, wherein the alarm data at least includes: alarm period, alarm level, alarm source application and alarm details.
[0048] It should be noted that the application monitoring platform is used to monitor key indicators and events in information computer systems around the clock, such as server load, network latency, system logs and performance bottlenecks, collect, analyze and present system data from multiple sources, and immediately trigger an alarm once an abnormal situation that deviates from the normal operating state is detected to remind operation and maintenance personnel to quickly intervene to deal with potential problems and prevent minor faults from turning into major accidents.
[0049] Alarm data is the core output of the application monitoring platform, including real-time notifications of system anomalies. The alarm data in the embodiments of the present invention goes beyond superficial descriptions of events and deeply integrates temporal characteristics (alarm period), severity indicators (alarm level), source information (the application from which the alarm originated), and detailed background information (alarm details). This integration of information enables operations and maintenance personnel to quickly locate issues, determine their priority, and take appropriate measures, greatly improving the efficiency and accuracy of problem resolution.
[0050] Specifically, the alarm period indicates the specific time window in which the alarm event occurs; the alarm level is used to describe the severity of the alarm event, which is generally divided into multiple levels, such as minor, serious, warning, etc.; the alarm source application can be used to identify the specific application or service module that triggers the alarm; the alarm details are used to provide detailed background and possible causes of the alarm event, including error codes, log fragments, and the status of related resources.
[0051] Step S202 , screening the alarm data using preset quantization rules to obtain screened data, wherein the preset quantization rules are used to filter alarm time periods, screen alarm levels, aggregate alarm events, and eliminate plan change interference.
[0052] It's important to note that pre-set quantitative rules are logical standards used to process raw alarm data. They aim to filter and optimize the flow of alarm information through well-defined parameters and conditions. These rules include, but are not limited to: filtering specific time periods to ensure that alarms are analyzed only within the time period of interest; filtering alarm levels, typically ignoring low-level alarms due to their minimal impact on system stability; aggregating alarm events to combine alarms within the same time period or of the same type to avoid duplicate alerts; and eliminating scheduled change interference to ensure that relevant alarms during known maintenance or update periods are treated as normal and not counted towards interference metrics. By applying these rules, the number of invalid alarms can be effectively reduced, improving the accuracy and utility of the alarm system.
[0053] Filtered data is an alarm data set that has been processed using preset quantitative rules. Compared with raw alarm data, filtered data focuses more on alarm events that are truly meaningful to operation and maintenance personnel and require immediate response. This helps operation and maintenance personnel quickly identify key issues and reduce the workload and interference caused by irrelevant information.
[0054] The technical problem addressed by the above steps is to optimize the alarm management process, reduce the workload of operations and maintenance personnel caused by alarm noise, and improve problem diagnosis and response speed. In an unfiltered alarm environment, operations and maintenance personnel may face information overload, making it difficult to distinguish between true emergencies and routine fluctuations. This not only consumes valuable attention resources but also increases the risk of misjudgment, posing a threat to system stability and business continuity. Using pre-set quantitative rules to filter data overcomes this problem, ensuring the effectiveness of the alarm system while reducing the burden on operations and maintenance personnel.
[0055] Alternatively, in an embodiment of the present invention, a machine learning algorithm may be used for intelligent alarm classification. Specifically, a model may be trained to identify the importance of alarms, and filtering and aggregation rules may be automatically generated based on historical data without manual presetting, thereby achieving more automated and accurate alarm management.
[0056] Alternatively, an embodiment of the present invention may also introduce a user feedback mechanism to optimize alarm rules, allowing operation and maintenance personnel to provide feedback on received alarms, for example, marking them as false alarms or non-urgent. The system then adjusts the alarm thresholds and rules accordingly, gradually reducing unnecessary alarms.
[0057] In order to reduce invalid warnings and achieve more accurate information filtering, the alarm data is further filtered and processed using preset quantitative rules to obtain the steps of filtering data, including: obtaining the alarm time period in the preset quantitative rules, wherein the alarm time period is a preset minimum working activity period of the operation and maintenance end; filtering out all data falling within the alarm time period from the alarm data to obtain first pre-selected data; obtaining the alarm level boundary in the preset quantitative rules, wherein the alarm level boundary refers to the preset minimum alarm level standard that the operation and maintenance end pays attention to; filtering out all data higher than or equal to the alarm level boundary from the first pre-selected data to obtain second pre-selected data.
[0058] It should be noted that the operations and maintenance department can set a minimum work activity period based on its own working hours and system characteristics as the basis for the alarm time period, ensuring that important alarms issued by the system during these periods are not missed. The alarm data will be compared with the set time period, and only those alarms generated during the focus period will be retained, reducing invalid information in the time dimension. A minimum alarm level standard can also be set to reflect the minimum threshold of alarms that the operations and maintenance department believes should be paid attention to. This further narrows the scope to only those important alarms that meet or exceed the set level, eliminating interference from low-level events and ensuring that operations and maintenance resources are rationally allocated to the most critical issues.
[0059] For example, if the operation and maintenance personnel's shift cycle is 24 hours, then the alarm time period is set to a smaller time period, such as 8 hours, as the time window of the alarm event, focusing attention on this time period to eliminate alarm noise in the past or future time periods, reducing the amount of analysis data, so as to facilitate a more detailed analysis of the workload of the operation and maintenance personnel in each working period, ensuring that the alarm analysis is neither too macro nor too micro, and accurately reflects the actual work rhythm and load.
[0060] In another embodiment, if operations personnel determine that only alerts above level 3 require immediate response, the alert level boundary can be set to level 3, ensuring that the system only processes alerts that could potentially impact system stability and business continuity. This level screening ensures that subsequent analysis and processing efforts focus on the most urgent and critical issues, avoiding distractions to operations personnel due to lower-level alerts.
[0061] In order to reduce the interference of repeated alarm events on operation and maintenance personnel, optimize the alarm processing process, and improve operation and maintenance efficiency, the alarm data is further filtered and processed using preset quantization rules to obtain the filtered data step, which also includes: obtaining the aggregation window in the preset quantization rule, wherein the aggregation window refers to a pre-set time window for performing similar aggregation of alarm events; segmenting the second pre-selected data based on the aggregation window to obtain window data; for each window data, analyzing the event types of the N alarm events recorded in the window data, and merging the alarm events of the same type to obtain merged window data containing M alarm events, wherein N is a positive integer and M is less than or equal to N; aggregating the merged window data corresponding to all window data to obtain third pre-selected data.
[0062] It should be noted that the aggregation window is a time range defined in the preset quantification rules, which is used to merge the same type of alarm events occurring in the same period. It aims to reduce the frequency of operation and maintenance personnel handling similar alarms and avoid repeatedly receiving alarm notifications of the same nature in a short period of time, thereby reducing the work pressure of operation and maintenance personnel and increasing their attention to key events.
[0063] Based on the size of the aggregation window, the second pre-selected data is divided into a series of independent time periods according to the time series. The alarm events in each time period constitute a set of window data, which can better manage and analyze the distribution of alarm events in different time periods and prepare for subsequent event aggregation.
[0064] Event type refers to the specific classification of alarm events, such as system errors, hardware failures, software anomalies, etc. When processing window data, the event type of each alarm event is identified and analyzed to facilitate merging alarm events of the same type, reducing redundant information and improving the quality of alarm data.
[0065] In an optional embodiment, if the aggregation window is set to 1 hour, then the same type of alarm events occurring within 1 hour will be treated as one event and processed; after obtaining the second pre-selected data, the data will be segmented in time according to the size of the aggregation window to form multiple window data, each window data contains alarm event information that occurred within a specific time window.
[0066] Next, the alarm events in each window data are analyzed to identify which events belong to the same category. For example, multiple alarms indicating excessive CPU utilization are considered to be of the same type. Once identified, these events are merged, with multiple similar alarms counted as one, forming a merged window containing M alarm events. If M is less than or equal to N, the number of alarm events has been effectively reduced through the aggregation window processing.
[0067] Finally, the merged window data from all window data is further aggregated to eliminate overlap and redundancy between time windows, forming the final third pre-selected data set. This is the final set of alarm information after dual filtering by time window and event type. This third pre-selected data is not only compressed in quantity but also optimized in content, ensuring that operations personnel receive comprehensive and accurate alarm information.
[0068] In order to reduce the impact of false positive alarms on the operation and maintenance end, and reduce the time for operation and maintenance personnel to handle non-emergency alarm events, especially alarms generated during system maintenance or changes, the alarm data is further filtered and processed using preset quantitative rules to obtain the filtered data step, which also includes: obtaining the plan change window recorded in the preset quantitative rules, wherein the plan change window refers to the time window occupied by the system plan change activities pre-filed at the operation and maintenance end; for each alarm event in the third pre-selected data, checking whether the alarm time period of the alarm event coincides with the plan change window to obtain the inspection result; if the inspection result indicates no overlap, the third pre-selected data is determined as the filtered data; or, if the inspection result indicates overlap, the data corresponding to the overlapping time period is eliminated from the third pre-selected data, and after eliminating all data corresponding to the overlapping time periods, the filtered data is obtained.
[0069] It's important to note that a planned change window is a predefined timeframe within operations management that covers planned maintenance, upgrades, or configuration changes to computer systems. During these activities, a series of expected alerts are generated. While these alerts reflect changes in system status, they are typically a normal byproduct of planned operations and don't require immediate intervention from operations personnel. By defining a planned change window, it's possible to identify alerts generated during planned maintenance, preventing these alerts from being treated as anomalies.
[0070] When processing alert data, the trigger time (alert period) of each alert event is examined to determine whether the event occurred within any pre-set plan change windows. If the alert period coincides with a plan change window, the alert is considered to be the result of a planned activity and can be ignored or handled differently. Conversely, if the alert period does not coincide with a plan change window, the alert is retained as an event requiring further evaluation and response.
[0071] Alternatively, embodiments of the present invention can replace the passive checking of planned change windows by building a proactive, integrated change management system. This change management system seamlessly interfaces with the application monitoring system, automatically eliminating the impact of planned change events during the alert data generation phase. By inputting relevant information into the monitoring system before the change activity begins, planned alerts can be eliminated in real time, eliminating the need for subsequent screening steps, streamlining the process, and increasing the immediate availability of alert data.
[0072] Alternatively, embodiments of the present invention can incorporate event correlation analysis technology to not only check whether the alarm period coincides with the plan change window, but also analyze the correlation between the alarm event and historical plan change events. For example, if the system detects a pattern of alarms that has occurred during multiple past system upgrades, even without clear time window records, these patterns can be intelligently identified and eliminated, further reducing false alarms and interference.
[0073] Step S203: Calculate the interference duration based on the filtered data, where the interference duration is used to reflect the workload of the operation and maintenance end.
[0074] It should be noted that the interference duration is a quantitative indicator used to measure the impact on the operation and maintenance end due to processing system alarms. It aims to reflect the actual time that the operation and maintenance end is occupied by invalid or low-priority alarms within a specific time period.
[0075] The technical problem solved in the above steps is mainly the alarm noise phenomenon that is prevalent in existing monitoring systems. That is, a large number of non-urgent, repetitive or planned events are mixed in with the alarm data, making it difficult for operation and maintenance personnel to effectively identify and respond to system problems that really need attention when faced with massive alarms; and calculating the interference duration based on the filtered alarm data deeply explores the direct connection between alarm events and operation and maintenance work, which helps to more accurately evaluate the performance of the alarm system and the burden level of operation and maintenance personnel, and then facilitates the targeted optimization of monitoring rules, reduces the occurrence of invalid alarms, and improves operation and maintenance efficiency.
[0076] In order to quantify the load pressure on the operation and maintenance end, further, the steps of calculating the interference duration based on the filtered data include: setting the initial value of the interference duration to 0; traversing all alarm events in the filtered data; during the traversal process, adding the duration of the alarm period triggered by each alarm event to the interference duration to update the interference duration; after the traversal is completed, determining the interference duration.
[0077] It's important to note that the cumulative value of this metric is initialized to zero before calculating the duration of interference. This ensures a clear starting point for the calculation, avoids any residual influence from historical data, and ensures that each calculation is based on the latest filtered dataset. Traversing each alarm event in the filtered dataset means individually examining and processing all alarm information in the dataset. This traversal process is the basis for calculating the duration of interference and ensures that all required alarm events are systematically analyzed and calculated.
[0078] For each traversed alarm event, the duration of the triggered alarm period is calculated and added to the current interference duration. The alarm period here refers to the entire time range from the alarm event triggering to its end. By accumulating the duration of each alarm period, the total length of time the operation and maintenance side was affected by invalid or duplicate alarms within a certain time range can be accurately calculated. After traversing all alarm events in the filtered dataset, the accumulated interference duration is the final result.
[0079] In another alternative embodiment, the degree of interference can be assessed by calculating the density of alarm events per unit time, rather than directly accumulating the duration of the alarm period. For example, a threshold for the number of alarm events per minute or hour can be set; once this threshold is exceeded, the operator is considered to be experiencing interference. This approach takes into account the impact of event frequency on operators and may be more sensitive to the impact of a large number of alarm events in a short period of time.
[0080] Furthermore, time series analysis can be used to not only calculate the duration of interference but also predict possible future periods of interference. Specifically, a time series model can be built based on historical alarm data to predict periods of time when alarm frequency may be abnormally high. This allows for early warning, allowing the operations team to prepare or adjust monitoring rules to avoid excessive alarms during these periods.
[0081] Furthermore, when calculating interference duration, you can factor in event priority and urgency weights. Not all alarm events are equally important. By assigning higher weights to high-priority, urgent alarms and lower weights to lower-priority alarms, we can calculate weighted interference duration, making the calculation more tailored to actual O&M needs.
[0082] In addition to considering the alarm event itself, user behavior analysis can also be used to assess whether the alarm actually affects the work of the operation and maintenance personnel. For example, the response time of the operation and maintenance personnel after receiving the alarm, the handling actions, and whether the escalation of the fault was successfully avoided can be recorded. This method can more accurately measure the extent to which the alarm affects actual work, but its implementation complexity and the required data are more abundant, and it may require the integration of more sensors and user behavior data. This application does not go into details here.
[0083] Step S204 : When the interference duration is greater than a preset tolerance threshold, a monitoring and tuning strategy is generated based on the filtered data, wherein the monitoring and tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0084] It's important to note that in the high-density alert environment of existing technologies, operations and maintenance teams face difficulty distinguishing signal from noise. Many invalid alerts not only waste time and resources but also delay responses to truly critical situations. The lack of unified evaluation standards and tuning guidance makes the process of optimizing application monitoring platforms unclear and difficult to quantify.
[0085] The quantification mechanism proposed in this embodiment of the present invention, namely a preset tolerance threshold, is used to set the maximum interference duration that the operation and maintenance end can reasonably tolerate. When the actual interference duration exceeds this threshold, the next step is triggered: generating a monitoring and tuning strategy based on the filtered data. This strategy contains specific recommended measures, such as adjusting alarm thresholds, optimizing alarm rules, and improving alarm aggregation logic. This strategy guides the application monitoring platform to perform necessary parameter tuning, reduce interference duration, improve monitoring efficiency, and enhance the job satisfaction of operation and maintenance personnel.
[0086] In order to reduce the alarm noise in the monitoring system, further, the step of generating a monitoring tuning strategy based on the filtered data includes: analyzing the filtered data within the interference period to obtain the first alarm number of each alarm source application; when the first alarm number exceeds the preset threshold, judging whether the alarm is a false alarm noise based on the inspection result of the operation log of the alarm source application, and obtaining a first judgment result; generating a monitoring tuning strategy based on the first judgment result, wherein the monitoring tuning strategy includes: adjusting the alarm threshold.
[0087] The alarm source application refers to the specific application or service that triggers the monitoring system's alarm. Runtime logs contain detailed information about the application's runtime, including but not limited to performance metrics, operational activities, and abnormal conditions. By examining run logs, you can verify the authenticity and validity of alarm events and determine whether the alarm accurately reflects the application's operational status or is merely a false alarm or noise.
[0088] False alarms refer to alarms that are mistakenly triggered by the monitoring system or have no real value. These may be caused by system oversensitivity, improper monitoring rule settings, or normal changes in the application itself, rather than actual failures or abnormal behavior. These false alarms increase the burden on operations and maintenance personnel, leading to wasted resources and distraction.
[0089] Based on analysis of interference duration and alarm sources, the system then examines the operation logs to determine which alarm events represent false alarms. Based on this determination, a monitoring optimization strategy is generated. Common strategies include adjusting monitoring thresholds to reduce the frequency of false alarms or irrelevant alarms. This process aims to streamline and optimize alarm rules, reduce disruptions for operations personnel, and improve monitoring system accuracy and operational efficiency.
[0090] It should be noted that the preset threshold is a key parameter used to distinguish normal alarms from abnormally high-frequency alarms. The determination method is usually based on the following points: statistically analyzing previous alarm records to find the average trigger frequency of various alarms under normal operating conditions and the peak number of alarms in special events (such as system upgrades and network fluctuations), so as to set a reasonable threshold; consulting operation and maintenance personnel about which types of alarm events most often cause work interference and at what frequency these events begin to become troublesome, and then designing a reasonable threshold; comprehensively considering the potential impact of various alarm events on business continuity and service quality, maintaining a lower threshold for events with greater business risks, and setting a higher threshold for events that generally do not affect the business; after setting the initial threshold, based on the feedback from actual monitoring and operation and maintenance operations, gradually adjust the threshold until the optimal balance point is found that can effectively filter out invalid alarms while not missing key issues.
[0091] Once an application's alarm is determined to be a false alarm, that is, the first judgment result is "false alarm", a monitoring optimization strategy is generated based on these results. The main optimization strategy is to adjust the alarm threshold. Specific methods may include:
[0092] Raise the alarm trigger threshold: For applications with frequent false alarms, you can raise the threshold of the monitoring indicator to reduce the alarm trigger frequency and prevent alarms from being triggered by small fluctuations.
[0093] Introducing more filtering rules: You can add alarm filtering rules for specific applications or types, such as anomaly detection rules based on time windows, specific system states, or historical data, to more accurately identify false alarms.
[0094] Optimize the alarm convergence algorithm: By improving the alarm convergence algorithm, duplicate alarms from the same source can be automatically aggregated, and the summarized alarm information is only pushed once to the operation and maintenance personnel, thereby reducing interference.
[0095] In an optional specific implementation scenario, assume that in the monitoring system of a certain data center, application A triggered a large number of alarms within a week, causing frequent interference to the operation and maintenance personnel. After analyzing the filtered data during the interference period, it was found that the number of first alarms of application A exceeded the preset threshold. By further checking the operation log of application A, it was determined that these alarms were mainly caused by resource bottlenecks within the application, and these bottlenecks were considered abnormal under the current monitoring rules. Based on these judgments, the following monitoring and tuning strategies can be generated: increase the resource utilization monitoring threshold of application A to adapt to higher resource consumption requirements; formulate a set of filtering rules for the resource bottlenecks of application A, such as automatically filtering out real resource leakage alarms during high-load periods, while ignoring regular resource usage fluctuations; improve the alarm convergence algorithm, and when repeated alarms related to resource usage are detected, only push summary information to the operation and maintenance personnel once to avoid multiple interruptions.
[0096] In another optional embodiment, a machine learning algorithm can be used to analyze historical alarm data and operation logs, automatically learn and adjust monitoring thresholds, and dynamically optimize the alarm triggering conditions based on the daily performance of the application and the response mode of the operation and maintenance end, thereby reducing false alarms and capturing real abnormal situations in a timely manner.
[0097] In order to reduce the interference of invalid alarms generated by the monitoring system on the operation and maintenance end, the step of generating a monitoring and tuning strategy based on the filtered data further includes: analyzing the filtered data within the interference period to obtain the second alarm number of each type of alarm event; when the second alarm number exceeds the preset threshold, judging whether the event type of the alarm event is an invalid alarm by checking the historical alarm records, and obtaining a second judgment result; generating a monitoring and tuning strategy based on the second judgment result, wherein the monitoring and tuning strategy includes: setting event type filtering conditions.
[0098] It's important to note that when processing and filtering data, we analyze the triggering frequency of various alarm events during the interference period and record the number of times each type of alarm event is triggered, which is the second alarm count. By counting the second alarm count, we can identify which types of alarm events occur frequently, which may be false alarms or invalid alarms from the monitoring system.
[0099] Historical alarm records contain all alarm events generated by the monitoring system over a period of time, including information such as event type, trigger time, and processing results. They serve as an important data source for analyzing alarm effectiveness and formulating tuning strategies. By analyzing historical data, typical invalid alarm patterns can be learned and identified.
[0100] Invalid alarms are those that don't actually require an immediate response from operations personnel. They are caused by overly sensitive monitoring rules, inherent system volatility, or other non-critical factors. Although triggered, these alarms often don't cause system failures or performance degradation, but instead unnecessarily increase the workload and distractions for operations personnel.
[0101] In one optional implementation scenario, the operations and maintenance department of a large network service company was frequently plagued by a specific type of memory usage alarm, which frequently triggered during off-peak hours at night but generally did not cause system performance degradation or business interruptions. Analysis of filtered data during the interference period revealed that the number of secondary alarms for this type of memory usage alarm event far exceeded the preset threshold, and an inspection of historical alarm records revealed that most of these events were marked as invalid. Based on this analysis, the following monitoring and tuning strategy was generated:
[0102] Set event type filtering conditions to increase the trigger threshold for memory usage alarms during non-peak hours at night, for example, from 75% to 90%. At the same time, introduce time window filtering to trigger the alarm only when the threshold is reached for more than 30 consecutive minutes to avoid false alarms caused by short-term memory fluctuations.
[0103] Automatically aggregate similar events and similar memory usage alarms within the same time period. If the alarm level is deemed insufficient to attract attention after aggregation, it will be automatically filtered to reduce the burden on operation and maintenance personnel.
[0104] Through the above steps S201 to S204, alarm data can be first obtained from the application monitoring platform, wherein the alarm data at least includes: alarm period, alarm level, alarm source application and alarm details, and then the alarm data is filtered and processed using preset quantization rules to obtain filtered data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events and eliminate interference from plan changes, and then the interference duration is calculated based on the filtered data, wherein the interference duration is used to reflect the workload of the operation and maintenance end, and finally, when the interference duration is greater than the preset tolerance threshold, a monitoring and tuning strategy is generated based on the filtered data, wherein the monitoring and tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0105] In an embodiment of the present invention, by constructing a refined interference feedback model, the purpose of accurately quantifying the impact of alarm interference on operation and maintenance efficiency is achieved, thereby achieving the technical effect of improving operation and maintenance response capabilities and overall system stability. Specifically, first, preset quantification rules are used to filter valuable information streams from complex monitoring alarms. For example, alarms during non-critical periods are filtered, repeated events are aggregated to avoid excessive interference, and alarm fluctuations caused by planned maintenance are eliminated, thereby ensuring that the focus is on real and urgent system problems. Subsequently, a data analysis algorithm is used to calculate the actual interference duration of the alarm event. This indicator can directly reflect the load and efficiency of the operation and maintenance work. When it is identified that the interference duration exceeds the preset tolerance threshold, that is, the degree of interference exceeds the tolerance range of normal operation, a targeted monitoring tuning strategy is generated based on the filtered data to guide the application monitoring platform to optimize parameters to reduce invalid alarms, improve the quality and efficiency of alarm event processing, and reduce the burden of operation and maintenance work. This significantly enhances the system's adaptability and the collaborative efficiency of the operation and maintenance team, thereby solving the technical problem in related technologies that fail to effectively quantify the interference degree of monitoring alarms from the operation and maintenance load end, thereby failing to pay attention to operation and maintenance pressure and resulting in low operation and maintenance efficiency.
[0106] The present invention will be described below in conjunction with another specific embodiment.
[0107] Figure 3 is a flow chart of an optional method for evaluating the degree of alarm interference according to an embodiment of the present invention. Figure 3As shown, the method includes the following process:
[0108] S1. Alarm data generation: the application monitoring system generates alarm data.
[0109] S2. Alarm data collection: call the application monitoring system interface to obtain alarm data.
[0110] S3. Set the interference time calculation rules, including time period, filtering conditions, aggregation algorithm, exception scenarios and other calculation rules.
[0111] In an optional embodiment, exemplary rules are as follows: 1) The time period is set to be consistent with the duty period, from 8:00 am every day to 8:00 am the next day; 2) The filtering conditions include alarms generated by applications named IC / ASI / IR / SCH / RCT / NTF, and alarms with alarm levels > level 2; 3) The aggregation algorithm is aggregated by hour, and multiple alarms occurring within 1 hour are counted as 1; 4) For exceptional scenarios, alarms during application plan changes should be excluded.
[0112] S4. Set a tolerable time threshold. This can be determined based on different dimensions, such as importance. High-level systems can have a higher tolerance time to avoid missed alerts, while low-level systems can have a lower tolerance time to reduce false alerts. More specifically, for example, setting the flight management duty officer's tolerance time for a single on-call period to 2 means a maximum of two hours of police interference can be tolerated per day.
[0113] S5. Calculate the interference time based on the alarm data and the rule data.
[0114] Figure 4 FIG. 1 is a flow chart of an example of an optional method for calculating the interference time according to an embodiment of the present invention. Figure 4 As shown, the method includes the following process steps:
[0115] S5.1. Filter alarm data by time to select alarms that occurred during the last duty cycle.
[0116] S5.2. Filter alarm data by application. Continue to select alarm data with an alarm level greater than Level 2 and an application name of IC, ASI, IR, SCH, RCT, or NTF. If an alarm exists, proceed to S5.3; otherwise, proceed to S5.6.
[0117] S5.3. Obtain application plan change data. If there is a change that overlaps with the alarm time, proceed to S5.4; otherwise, proceed to S5.5.
[0118] S5.4. Remove alarms that overlap with the change time;
[0119] S5.5. Obtain the interference hour list by hourly aggregation. For example, if there are 2 alarm data at 10 o'clock and 1 alarm data at 13 o'clock, then the interference hour list is [10, 13].
[0120] S5.6. Count the number of interference hours to obtain the interference time. For example, if the interference hour list is [10, 13], the interference time is 2 hours.
[0121] S5.7. Compare the tolerable times and note the results.
[0122] S6. Output the results to a file.
[0123] S7: Is the interference time less than or equal to the tolerable time? If so, the process ends; otherwise, the process proceeds to S8.
[0124] S8. Monitor and optimize, analyze the alarms when the interference time exceeds the tolerable time, and perform adjustments until the interference time is less than or equal to the tolerable time.
[0125] Prior to the embodiments of the present invention, existing methods for evaluating the degree of alarm interference all started from alarm events and analyzed, filtered, and aggregated the number of alarms and alarm information. However, this method had high implementation costs and poor accuracy, and could not truly reflect the feelings of operation and maintenance personnel regarding receiving alarms.
[0126] The embodiments of the present invention, however, aim to eliminate alarm interference, focusing on the feelings of operations and maintenance personnel. These embodiments have the following advantages: The interference duration truly reflects the degree to which operations and maintenance personnel are disturbed by alarms; fuzzy expressions are converted into quantitative expressions, such as "I feel too many alarms are interfering with my work" is converted into "The interference duration is 4, which exceeds the tolerable time"; quantitative standards are provided for monitoring system tuning and alarm data management; implementation costs are low; and in addition to computer information system application monitoring, other fields involving monitoring and alarms can also refer to the methods of the embodiments of the present invention.
[0127] The present invention is described below in conjunction with another optional embodiment.
[0128] Example 2
[0129] An alarm monitoring device based on interference feedback provided in this embodiment includes multiple implementation units, each implementation unit corresponding to each implementation step in the above-mentioned embodiment 1.
[0130] Figure 5 is a schematic diagram of an optional alarm monitoring device based on interference feedback according to an embodiment of the present invention, such as Figure 5 As shown, the device may include: an acquisition unit 51, a screening unit 52, a calculation unit 53, and a generation unit 54.
[0131] The acquisition unit 51 is configured to acquire alarm data from the application monitoring platform, wherein the alarm data at least includes: alarm period, alarm level, alarm source application, and alarm details.
[0132] The screening unit 52 is used to screen the alarm data using preset quantitative rules to obtain screened data, wherein the preset quantitative rules are used to filter alarm time periods, screen alarm levels, aggregate alarm events, and eliminate plan change interference.
[0133] The calculation unit 53 is used to calculate the interference duration based on the filtered data, wherein the interference duration is used to reflect the workload of the operation and maintenance end.
[0134] The generating unit 54 is configured to generate a monitoring and tuning strategy based on the screening data when the interference duration is greater than a preset tolerance threshold, wherein the monitoring and tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0135] The above-mentioned alarm monitoring device based on interference feedback can first obtain alarm data from the application monitoring platform through the acquisition unit 51, wherein the alarm data at least includes: alarm period, alarm level, alarm source application and alarm details, and then filter and process the alarm data using preset quantization rules through the screening unit 52 to obtain filtered data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events and eliminate plan change interference, and then calculate the interference duration based on the filtered data through the calculation unit 53, wherein the interference duration is used to reflect the workload of the operation and maintenance end, and finally generate a monitoring tuning strategy based on the filtered data through the generation unit 54 when the interference duration is greater than the preset tolerance threshold, wherein the monitoring tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0136] In an embodiment of the present invention, by constructing a refined interference feedback model, the purpose of accurately quantifying the impact of alarm interference on operation and maintenance efficiency is achieved, thereby achieving the technical effect of improving operation and maintenance responsiveness and overall system stability. Specifically, preset quantification rules are first used to filter valuable information streams from complex monitoring alarms. For example, alarms during non-critical periods are filtered, repeated events are aggregated to avoid excessive interference, and alarm fluctuations caused by planned maintenance are eliminated, thereby ensuring focus on real and urgent system problems. Subsequently, a data analysis algorithm is used to calculate the actual interference duration of the alarm event. This indicator can directly reflect the load and efficiency of the operation and maintenance work. When it is identified that the interference duration exceeds a preset tolerance threshold, that is, the degree of interference exceeds the tolerance range of normal operation, a targeted monitoring tuning strategy is generated based on the filtered data to guide the application monitoring platform to optimize parameters to reduce invalid alarms, improve the quality and efficiency of alarm event processing, and reduce the burden of operation and maintenance work, thereby significantly enhancing the system's adaptability and the collaborative efficiency of the operation and maintenance team. This solves the technical problem in related technologies that fail to effectively quantify the interference degree of monitoring alarms from the operation and maintenance load end, thereby failing to pay attention to operation and maintenance pressure and resulting in low operation and maintenance efficiency.
[0137] Furthermore, the screening unit includes: a first acquisition module, used to obtain the alarm time period in the preset quantitative rules, wherein the alarm time period is a preset minimum working activity period of the operation and maintenance end; a first screening module, used to screen out all data falling within the alarm time period from the alarm data to obtain first pre-selected data; a second acquisition module, used to obtain the alarm level boundary in the preset quantitative rules, wherein the alarm level boundary refers to the preset minimum alarm level standard that the operation and maintenance end pays attention to; a second screening module, used to screen out all data higher than or equal to the alarm level boundary from the first pre-selected data to obtain second pre-selected data.
[0138] Furthermore, the screening unit also includes: a third acquisition module, used to obtain the aggregation window in the preset quantization rule, wherein the aggregation window refers to a pre-set time window for performing similar aggregation of alarm events; a segmentation module, used to segment the second pre-selected data based on the aggregation window to obtain window data; a merging module, used to analyze the event types of N alarm events recorded in the window data for each window data, and merge alarm events of the same type to obtain merged window data containing M alarm events, wherein N is a positive integer and M is less than or equal to N; an aggregation module, used to aggregate the merged window data corresponding to all window data to obtain third pre-selected data.
[0139] Furthermore, the screening unit also includes: a fourth acquisition module, used to obtain the plan change window recorded in the preset quantitative rules, wherein the plan change window refers to the time window occupied by the system plan change activities pre-filed at the operation and maintenance end; an inspection module, used to check whether the alarm period of each alarm event in the third pre-selected data coincides with the plan change window, and obtain an inspection result; a first determination module, used to determine the third pre-selected data as screening data when the inspection result indicates no overlap; and an elimination module, used to eliminate the data corresponding to the overlapping period from the third pre-selected data when the inspection result indicates overlap, and obtain the screening data after eliminating all data corresponding to the overlapping period.
[0140] Furthermore, the calculation unit includes: a setting module for setting the initial value of the interference duration to 0; a traversal module for traversing all alarm events in the filtered data; an accumulation module for adding the duration of the alarm period triggered by each alarm event to the interference duration during the traversal process to update the interference duration; and a second determination module for determining the interference duration after the traversal is completed.
[0141] Furthermore, the generation unit includes: a first analysis module, which is used to analyze the screening data within the interference duration to obtain the first alarm number of each alarm source application; a first judgment module, which is used to judge whether the alarm is a false alarm noise based on the inspection result of the operation log of the alarm source application when the first alarm number exceeds the preset threshold, and obtain a first judgment result; a first generation module, which is used to generate a monitoring and tuning strategy based on the first judgment result, wherein the monitoring and tuning strategy includes: adjusting the alarm threshold.
[0142] Furthermore, the generation unit also includes: a second analysis module, which is used to analyze the screening data within the interference duration to obtain the second alarm number of each type of alarm event; a second judgment module, which is used to determine whether the event type of the alarm event is an invalid alarm by checking the historical alarm records when the second alarm number exceeds the preset threshold, and obtain a second judgment result; a second generation module, which is used to generate a monitoring and tuning strategy based on the second judgment result, wherein the monitoring and tuning strategy includes: setting event type filtering conditions.
[0143] The above-mentioned interference feedback-based alarm monitoring device can also include a processor and a memory. The above-mentioned acquisition unit 51, screening unit 52, calculation unit 53, generation unit 54, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.
[0144] The processor includes a kernel that retrieves the corresponding program unit from memory. One or more kernels can be configured. By adjusting kernel parameters, the kernel uses preset quantization rules to filter and process alarm data, generating filtered data. The interference duration is calculated based on the filtered data. If the interference duration exceeds a preset tolerance threshold, a monitoring and tuning strategy is generated based on the filtered data to guide parameter tuning for the application monitoring platform.
[0145] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0146] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program initialized with the following method steps: obtaining alarm data from an application monitoring platform, wherein the alarm data includes at least: alarm period, alarm level, alarm source application and alarm details; filtering and processing the alarm data using preset quantization rules to obtain filtered data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events and eliminate interference from plan changes; calculating interference duration based on the filtered data, wherein the interference duration is used to reflect the workload of the operation and maintenance end; when the interference duration is greater than a preset tolerance threshold, generating a monitoring and tuning strategy based on the filtered data, wherein the monitoring and tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
[0147] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the alarm monitoring method based on interference feedback of any one of the above-mentioned embodiments.
[0148] According to another aspect of an embodiment of the present invention, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the alarm monitoring method based on interference feedback of any one of the above-mentioned embodiments.
[0149] Figure 6 is a structural block diagram of an electronic device for executing an alarm monitoring method based on interference feedback according to an embodiment of the present invention, such as Figure 6 As shown, the electronic device may include: one or more ( Figure 6 Only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0150] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the alarm monitoring method and device based on interference feedback in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned alarm monitoring method based on interference feedback. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0151] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), or a PAD. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.
[0152] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0153] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0154] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0155] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0156] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0157] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0158] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0159] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An alarm monitoring method based on interference feedback, characterized in that: include: Acquire alarm data from the application monitoring platform, wherein the alarm data includes at least: alarm period, alarm level, alarm source application, and alarm details; Filtering the alarm data using preset quantization rules to obtain filtered data, wherein the preset quantization rules are used to filter alarm time periods, filter alarm levels, aggregate alarm events, and eliminate interference from plan changes; Calculating the interference duration based on the filtered data, wherein the interference duration is used to reflect the workload of the operation and maintenance end; When the interference duration is greater than a preset tolerance threshold, a monitoring and tuning strategy is generated based on the screening data, wherein the monitoring and tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
2. The alarm monitoring method according to claim 1, characterized in that: The step of screening the alarm data using a preset quantization rule to obtain the screened data includes: Obtaining an alarm time period in the preset quantification rule, wherein the alarm time period is a preset minimum working activity period of the operation and maintenance end; Filtering all data falling within the alarm time period from the alarm data to obtain first pre-selected data; Obtaining an alarm level boundary in the preset quantification rule, wherein the alarm level boundary refers to a preset minimum alarm level standard that the operation and maintenance end pays attention to; All data higher than or equal to the alarm level boundary are filtered out from the first pre-selected data to obtain second pre-selected data.
3. The alarm monitoring method according to claim 2, characterized in that: The step of screening the alarm data using a preset quantization rule to obtain screened data further includes: Obtaining an aggregation window in the preset quantification rule, wherein the aggregation window refers to a preset time window for aggregating similar alarm events; Segmenting the second pre-selected data based on the aggregation window to obtain window data; For each window data, analyzing the event types of N alarm events recorded in the window data, and merging the alarm events of the same type to obtain merged window data containing M alarm events, where N is a positive integer and M is less than or equal to N; Aggregate the merged window data corresponding to all the window data to obtain third pre-selected data.
4. The alarm monitoring method according to claim 3, characterized in that: The step of screening the alarm data using a preset quantization rule to obtain screened data further includes: Obtaining a plan change window recorded in the preset quantitative rule, wherein the plan change window refers to a time window occupied by a system plan change activity pre-filed at the operation and maintenance end; For each of the alarm events in the third pre-selected data, checking whether an alarm period of the alarm event coincides with the plan change window, and obtaining a checking result; If the inspection result indicates that there is no overlap, determining the third pre-selected data as the screening data; or, In the case that the inspection result indicates overlap, the data corresponding to the overlap period is removed from the third pre-selected data, and the screening data is obtained after removing the data corresponding to all overlap periods.
5. The alarm monitoring method according to claim 1, characterized in that: The steps for calculating the duration of interference based on the filtered data include: Setting the initial value of the interference duration to 0; Traverse all alarm events in the screening data; During the traversal process, the duration of the alarm period triggered by each of the alarm events is added to the interference duration to update the interference duration; After the traversal is completed, the interference duration is determined.
6. The alarm monitoring method according to claim 1, characterized in that: The step of generating a monitoring and tuning strategy based on the screened data includes: Analyzing the screening data within the interference duration to obtain a first alarm number applied by each of the alarm sources; When the first alarm number exceeds a preset threshold, determining whether the alarm is a false alarm noise based on a result of checking the operation log of the alarm source application, and obtaining a first judgment result; The monitoring and tuning strategy is generated based on the first judgment result, wherein the monitoring and tuning strategy includes: adjusting an alarm threshold.
7. The alarm monitoring method according to claim 1, characterized in that: The step of generating a monitoring and tuning strategy based on the filtered data further includes: Analyzing the screening data within the interference duration to obtain a second alarm number for each type of alarm event; When the second alarm number exceeds a preset threshold, determining whether the event type of the alarm event is an invalid alarm by checking historical alarm records, and obtaining a second judgment result; The monitoring and tuning strategy is generated based on the second judgment result, wherein the monitoring and tuning strategy includes: setting an event type filtering condition.
8. An alarm monitoring device based on interference feedback, characterized in that: include: An acquisition unit is configured to acquire alarm data from an application monitoring platform, wherein the alarm data includes at least: an alarm period, an alarm level, an alarm source application, and alarm details; a screening unit, configured to screen the alarm data using preset quantitative rules to obtain screened data, wherein the preset quantitative rules are used to filter alarm time periods, screen alarm levels, aggregate alarm events, and eliminate interference from plan changes; a calculation unit, configured to calculate an interference duration based on the filtered data, wherein the interference duration is used to reflect a workload of the operation and maintenance end; A generating unit is used to generate a monitoring and tuning strategy based on the screening data when the interference duration is greater than a preset tolerance threshold, wherein the monitoring and tuning strategy is used to guide the application monitoring platform to perform parameter tuning.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the alarm monitoring method based on interference feedback according to any one of claims 1 to 7.
10. An electronic device, characterized in that: It includes one or more processors and a memory, the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the alarm monitoring method based on interference feedback as described in any one of claims 1 to 7.