Monitoring stability self-adaptive monitoring system based on automatic configuration

Through automated configuration and intelligent processing, the problem of existing monitoring systems relying on manual configuration and insufficient self-adaptation capabilities has been solved, and efficient and accurate monitoring strategy adaptation has been achieved, reducing operation and maintenance costs, and improving system stability and fault handling efficiency.

CN120750795APending Publication Date: 2025-10-03QIMING INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510944874.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing monitoring systems rely on manual configuration and lack intelligence and self-adaptation capabilities, resulting in configuration errors, frequent alarm interference, long time consumption, difficulty in adapting to complex business scenarios, and inability to respond to system changes in a timely manner.

Method used

An automatically configured monitoring system is used, including data source collection management, threshold rule management, message notification and alarm modules. Reinforcement learning algorithms are used to generate optimal monitoring strategies. Event fingerprints and time window deduplication and correlation analysis are used to optimize alarm information, achieving self-adaptation and intelligent decision-making.

Benefits of technology

It improves the efficiency and accuracy of monitoring configuration, reduces manual intervention, lowers operation and maintenance costs, improves system stability and fault handling efficiency, and ensures that the monitoring system responds to complex environmental changes in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750795A_ABST
    Figure CN120750795A_ABST
Patent Text Reader

Abstract

The invention discloses a monitoring stability self-adaptive monitoring system based on automatic configuration, which comprises a data source acquisition management module, a threshold rule management module, a message notification module, an alarm module and an exception handling module, and is characterized in that the threshold rule management module is used for configuring related monitoring rules and setting an alarm threshold and an alarm level; the alarm module is used for managing historical alarm information and analyzing and processing the historical alarm information; and the exception handling module is used for generating and executing an optimal monitoring strategy based on the historical alarm information. The problems that an existing monitoring system depends on manual configuration, is large in alarm interference, long in adjustment time consumption and lack of self-adaption capacity and the like can be solved, display and processing of alarm information can be further optimized through the message center and the alarm convergence function, automation, intelligence and high efficiency of the monitoring system are achieved, and the monitoring efficiency is improved. The monitoring accuracy and efficiency are improved, the operation and maintenance cost is reduced, and stable operation of the system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent operation and maintenance technology, and in particular to a monitoring stability self-adaptive monitoring system based on automatic configuration. Background Art

[0002] Currently, some technologies and solutions already exist in the field of monitoring systems. For example, some monitoring systems support basic threshold alarm functions, triggering alarms by setting fixed thresholds. Some monitoring systems also integrate simple automated configuration functions, but these functions are often limited to specific scenarios or indicators and lack versatility and flexibility. In terms of intelligence, some advanced monitoring systems have begun to try to introduce machine learning algorithms for anomaly detection and alarm prediction, but these algorithms usually require a large amount of historical data for training, and model updating and optimization are relatively difficult. In addition, existing monitoring systems often perform poorly when dealing with complex business scenarios and multi-dimensional data, making it difficult to accurately locate problems and predict potential risks: 1. Dependence on manual configuration: Most existing monitoring systems rely on operations and maintenance personnel to manually configure monitoring rules. This not only consumes a lot of time and manpower, but is also prone to configuration errors due to human negligence, resulting in false positives and missed positives.

[0003] 2. Extensive interference: Due to the lack of intelligent screening and filtering mechanisms, existing monitoring systems are easily affected by various interference factors, generating a large number of meaningless alarms, making it difficult for operation and maintenance personnel to screen out the truly important issues.

[0004] 3. Time-consuming: When faced with complex business scenarios and system architectures, the configuration and adjustment process of existing monitoring systems is often very time-consuming and unable to adapt to changes in business and environment in a timely manner, resulting in a significant reduction in the timeliness of monitoring.

[0005] 4. Lack of self-adaptation: Existing monitoring systems generally lack the ability to self-learn and self-evolve, and are unable to automatically adjust monitoring strategies based on actual monitoring data and business needs, resulting in difficulty in continuously improving monitoring results.

[0006] In summary, existing monitoring systems have obvious deficiencies in terms of automated configuration, intelligent decision-making, and self-adaptation, and are unable to meet the monitoring needs of today's complex business environments. Therefore, developing a self-adaptive monitoring method and system for monitoring platform stability based on automated configuration is of great practical significance. Summary of the Invention

[0007] In order to solve the above problems, the present invention provides a monitoring stability self-adaptive monitoring system based on automatic configuration, including a data source acquisition management module, a threshold rule management module, a message notification module, an alarm module, and an exception handling module; The data source acquisition management module is used to collect, preprocess and manage relevant data, and configure monitoring indicators; the threshold rule management module is used to configure relevant monitoring rules and set alarm thresholds and alarm levels; the message notification module is used to transmit alarm information to staff and confirm whether the message has been received; the alarm module is used to manage historical alarm information and perform analysis and processing; the exception handling module is used to generate and execute the optimal monitoring strategy based on historical alarm information.

[0008] Furthermore, the data source collection and management module collects relevant data specifically including: infrastructure data, business processing data, and historical data; the infrastructure data includes: node load, network jitter, and storage IOPS; the business processing data includes: transaction volume peak, user geographical distribution, and API call chain delay; the historical data includes: root causes of similar failures and policy adjustment effect records.

[0009] Furthermore, the relevant rules configured in the threshold rule management module specifically include: monitoring basic rules, compound rules, and context rules; the alarm levels include: urgent, important, and general.

[0010] Furthermore, the alarm module includes: an alarm information management submodule and an alarm convergence submodule; the alarm information management submodule is used to search, statistically analyze and generate multi-dimensional alarm views of alarm information; the alarm convergence submodule is used to group, deduplicate and perform correlation analysis on alarm information.

[0011] Furthermore, the grouping of alarm information includes: hardware failure, software anomaly, performance problem; Deduplication of alarm information specifically includes: Deduplication based on event fingerprints: generating a unique hash value for key fields of alarm information (error code, exception stack, timestamp), retaining only one identical hash value within a certain period of time; and dynamic aggregation based on time windows: aggregating high-frequency alarm information according to time windows, counting the number of times the same type of alarm information is triggered, and retaining the statistical information. The event fingerprint is calculated as MD5 (error code + exception stack + affected instance ID).

[0012] Furthermore, the alarm compression ratio is calculated after deduplication of the alarm information; wherein, the calculation formula of the alarm compression ratio is: alarm compression ratio = amount of alarms after deduplication / amount of original alarms.

[0013] Furthermore, an optimization module is included for real-time optimization of the optimal monitoring strategy generated by the exception handling module. Specifically, it calculates reward values ​​based on system status and strategy execution actions. Positive rewards are given when alarms are accurate and timely. Negative rewards are given when false alarms or missed alarms occur. Positive rewards are given when system overhead is low, and negative rewards are given when system overhead is excessive.

[0014] The present invention provides a monitoring stability self-adaptive monitoring system based on automatic configuration, which has the following beneficial effects: 1. Improve monitoring configuration efficiency and accuracy: The automated configuration metadata model abstracts parameters such as monitoring rules into programmable metadata, greatly reducing the workload of manual configuration and the risk of human error. This significantly improves monitoring configuration efficiency and accuracy, enabling the monitoring system to adapt to the needs of different business scenarios more quickly and accurately.

[0015] 2. Realize dynamic self-adaptation of monitoring strategies: The dynamic self-adaptation mechanism enables the monitoring system to automatically adjust the monitoring strategy according to business and environmental changes, respond to dynamic changes in the system in a timely manner, ensure the timeliness and accuracy of monitoring, and enhance the system's adaptability to complex environments.

[0016] 3. Optimize alarm information display and processing: The message notification module serves as a unified display platform for alarm information, with classification, hierarchical display, query, filtering, and export functions, allowing operation and maintenance personnel to quickly locate and track alarms. The alarm convergence function effectively reduces the number of alarms through grouping, deduplication, and correlation analysis, avoiding alarm storms, reducing the workload of operation and maintenance personnel, and improving fault handling efficiency.

[0017] 4. Reduce operation and maintenance costs and improve system stability: Combining the above functions, the present invention reduces manual intervention, improves the level of monitoring automation, and reduces operation and maintenance manpower costs; at the same time, more accurate and timely monitoring and alarm mechanisms help to quickly discover and solve system problems, improve the overall stability of the system, and ensure business continuity. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0019] Figure 1 This is a schematic diagram of the system structure provided by the present invention. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0021] The following describes the implementation method of the present invention in detail with reference to the accompanying drawings. The description only includes some embodiments, not all embodiments. For the purpose of clarity, representations and descriptions that are not related to the present invention are omitted in the drawings and descriptions.

[0022] In order to have a clearer understanding of the technical features, purposes and beneficial effects of the present invention, the technical solutions of the present invention are now described in detail below. Obviously, the implementation cases described are part of the embodiments of the present invention, not all of them, and should not be understood as limiting the scope of the implementation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0023] like Figure 1 As shown, the present invention provides a self-adaptive monitoring system for monitoring stability based on automatic configuration, including a data source acquisition management module, a threshold rule management module, a message notification module, an alarm module, and an exception handling module.

[0024] The data source collection and management module is used to collect, preprocess, and manage relevant data, and configure monitoring indicators. The relevant data collected specifically includes infrastructure data, business processing data, and historical data. Infrastructure data includes node load, network jitter, and storage IOPS. Business processing data includes peak transaction volume, user geographical distribution, and API call chain latency. Historical data includes the root causes of similar failures and records of the effectiveness of policy adjustments.

[0025] Build a metadata model for monitoring configuration, abstracting parameters such as monitoring rules, alarm thresholds, and sampling frequencies into programmable metadata. This approach automates and standardizes monitoring configuration, reduces manual intervention, lowers the risk of configuration errors, and improves configuration efficiency.

[0026] Multi-dimensional environmental variables are collected, including data from the infrastructure layer (such as node load, network jitter, and storage IOPS), the business layer (such as transaction volume peaks, user regional distribution, and API call chain latency), and the historical layer (such as the root causes of similar failures and the effectiveness of policy adjustments). This multi-dimensional data is first preprocessed, including data cleaning and normalization, to ensure data quality and consistency. Key features are extracted from this preprocessed data and used as input to represent the current state of the system. Using reinforcement learning algorithms (such as Q-Learning), the optimal monitoring strategy is derived in real time, optimizing alert effectiveness and system overhead.

[0027] The threshold rule management module is used to configure relevant monitoring rules and set alarm thresholds and alarm levels. Configuring relevant rules specifically includes: monitoring basic rules, compound rules, and context rules; the alarm levels include: urgent, important, and general.

[0028] The message notification module transmits alarm information to personnel and verifies receipt. As a unified display platform for alarm information, it boasts powerful data processing and display capabilities. It receives and integrates alarm information from various monitoring sources, classifying alarms by severity (urgent, important, general) and type (hardware failure, software anomaly, performance issue, etc.). It updates alarm status in real time based on the output of the decision engine, ensuring that operations and maintenance personnel receive the latest and most accurate alarm information.

[0029] In the alarm module, setting reasonable deduplication rules is the key to reducing alarm noise and improving operation and maintenance efficiency. It is used to manage historical alarm information and perform analysis and processing.

[0030] The alarm module includes an alarm information management submodule and an alarm convergence submodule. The alarm information management submodule searches for alarm information, performs statistical analysis, and generates multi-dimensional alarm views. The alarm convergence submodule groups alarm information, removes duplicates, and performs correlation analysis. Alarm information is grouped into categories such as hardware failure, software anomalies, and performance issues.

[0031] Deduplication of alarm information specifically includes: Deduplication based on event fingerprints: A unique hash value is generated for key fields in the alert information (error code, exception stack, and timestamp). Only one identical hash value is retained within a certain period of time. For example, an event fingerprint = MD5 (error code + exception stack + affected instance ID). If a fingerprint is repeated within three consecutive minutes, subsequent alerts are merged into a "duplicate event count + 1" value, and only the first alert triggers a notification.

[0032] Dynamic aggregation based on time windows: High-frequency alarms are aggregated according to the time window, and the number of times the same alarm type is triggered is counted and retained. The event fingerprint is calculated as MD5 (error code + exception stack + affected instance ID). Alarm compression ratio is calculated after deduplication; the formula for calculating the alarm compression ratio is: Alarm compression ratio = number of deduplicated alarms / number of original alarms. The aggregation window is dynamically adjusted based on business peak / off-peak periods, such as shortening it to 3 minutes during peak periods and extending it to 10 minutes during off-peak periods. It can also be set to a fixed value. Aggregation is triggered if the same alarm occurs three or more times within the same window. A whitelist mechanism is enabled for critical businesses to disable deduplication to ensure that no issues are missed.

[0033] Root cause deduplication and correlation analysis based on topological relationships: This function mines the potential relationships between alarms and combines multiple alarms caused by the same root cause into a comprehensive alarm, providing more comprehensive fault information.

[0034] Alarm convergence effectively reduces alarm storms, with an alarm compression rate greater than 80% (alarm compression rate = number of deduplicated alarms / number of original alarms), reducing the workload of operation and maintenance personnel and improving fault handling efficiency.

[0035] The exception handling module is used to generate and execute the optimal monitoring strategy based on historical alarm information.

[0036] The system also includes an optimization module, which performs real-time optimization of the optimal monitoring strategy generated by the exception handling module. Specifically, it calculates reward values ​​based on system status and strategy execution actions. Positive rewards are given when alarms are accurate and timely. Negative rewards are given when false alarms or missed alarms occur. Positive rewards are given when system overhead is low, and negative rewards are given when system overhead is excessive.

[0037] During the startup phase, expert rules are used to generate an initial policy (such as an alarm when the CPU exceeds 85%). Transfer learning is then used to reuse historical data to train the Q network. If the Q network fails to converge, a weighted hybrid strategy is used to allocate 50% of traffic to the reinforcement learning strategy and 50% to the traditional strategy. This strategy is monitored for seven days. When the false alarm rate, false negative rate, and system overhead are all better than those of the traditional strategy, the entire strategy can be switched to the reinforcement learning strategy.

[0038] The present invention can not only solve the problems of existing monitoring systems such as reliance on manual configuration, frequent alarm interference, time-consuming adjustment and lack of self-adaptation capabilities, but also further optimize the display and processing of alarm information through the message center and alarm convergence functions, realize the automation, intelligence and efficiency of the monitoring system, improve the accuracy and efficiency of monitoring, reduce operation and maintenance costs, and ensure the stable operation of the system.

[0039] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.

Claims

1. A monitoring stability self-adaptive monitoring system based on automatic configuration, characterized in that: It includes data source collection management module, threshold rule management module, message notification module, alarm module, and exception handling module; The data source acquisition and management module is used to collect, pre-process and manage relevant data, and configure monitoring indicators; The threshold rule management module is used to configure relevant monitoring rules and set alarm thresholds and alarm levels; The message notification module is used to transmit the alarm information to the staff and confirm whether the message has been received; The alarm module is used to manage historical alarm information and perform analysis and processing; The exception handling module is used to generate and execute the optimal monitoring strategy based on historical alarm information.

2. The monitoring stability self-adaptive monitoring system based on automatic configuration according to claim 1 is characterized in that: The data source collection and management module collects relevant data, including infrastructure data, business processing data, and historical data; the infrastructure data includes node load, network jitter, and storage IOPS; the business processing data includes transaction volume peaks, user geographical distribution, and API call chain delays; the historical data includes root causes of similar failures and records of policy adjustment effects.

3. The monitoring stability self-adaptive monitoring system based on automatic configuration according to claim 1 is characterized in that: The relevant rules configured in the threshold rule management module specifically include: monitoring basic rules, compound rules, and context rules; the alarm levels include: urgent, important, and general.

4. The monitoring stability self-adaptive monitoring system based on automatic configuration according to claim 1 is characterized in that: The alarm module includes: an alarm information management submodule and an alarm convergence submodule; the alarm information management submodule is used to search, statistically analyze and generate multi-dimensional alarm views of alarm information; the alarm convergence submodule is used to group, deduplicate and perform correlation analysis on alarm information.

5. The monitoring stability self-adaptive monitoring system based on automatic configuration according to claim 4 is characterized in that: The grouping of alarm information includes: hardware failure, software anomaly, and performance problem; Deduplication of the alarm information specifically includes: deduplication based on event fingerprints: generating a unique hash value for the key fields of the alarm information, and retaining only one identical hash value within a period of time; Dynamic aggregation based on time windows: Aggregate high-frequency alarm information according to time windows, count the number of times the same type of alarm information is triggered, and retain the statistical information.

6. The monitoring stability self-adaptive monitoring system based on automatic configuration according to claim 5 is characterized in that: The alarm compression ratio is calculated after deduplication of the alarm information; wherein, the calculation formula of the alarm compression ratio is: alarm compression ratio = amount of alarms after deduplication / amount of original alarms.

7. The monitoring stability self-adaptive monitoring system based on automatic configuration according to claim 1 is characterized in that: It also includes an optimization module for performing real-time optimization processing on the optimal monitoring strategy generated in the exception handling module, specifically calculating the reward value through the system status and strategy execution action.