Dynamic monitoring alarm system of distributed storage system
The dynamic monitoring and alarm system solves the problems of high false alarm rate, low resource utilization efficiency and poor scalability of monitoring and alarm systems in distributed storage systems. It achieves efficient fault identification and adaptive resource management, improves system reliability and fault self-recovery capability, reduces the burden on operation and maintenance personnel and improves system reliability and fault self-recovery capability.
Patent Information
- Application Number
- CN202511253023.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-28
AI Technical Summary
Existing monitoring and alarm systems for distributed storage systems suffer from numerous false alarms or missed alarms, low resource utilization efficiency, high module coupling, poor scalability, and insufficient fault self-recovery capabilities. In particular, they struggle to adapt to dynamic load changes and identify complex fault modes in large-scale clusters.
A dynamic monitoring and alarm system is adopted. The data acquisition layer adjusts the acquisition frequency and granularity differently, stores monitoring data in layers, generates alarm thresholds that adapt to load changes and builds an alarm knowledge graph in the intelligent analysis layer, and realizes automatic repair in the alarm self-healing layer. The modular design improves scalability and resource utilization efficiency.
It significantly reduced false alarm rate and resource consumption, improved fault identification accuracy and self-healing capability, reduced the burden on maintenance personnel, enhanced system scalability and resource utilization, and achieved efficient fault self-recovery.
Smart Images

Figure CN121029094A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent operation and maintenance monitoring, in particular to a dynamic monitoring and alarming system for a distributed storage system. BACKGROUND
[0002] With the continuous expansion of system scale and the continuous improvement of business complexity, the number of nodes of large enterprise-level distributed storage clusters has expanded from dozens in the early days to thousands or even tens of thousands now, and the storage capacity has also jumped from TB level to PB level. In such a super large scale environment, hardware failure, network jitter, performance bottleneck and other problems have become the norm.
[0003] The traditional distributed storage monitoring and alarming system mainly relies on a static threshold strategy, that is, the administrator pre-sets fixed thresholds for each indicator (such as CPU usage exceeding 90%, disk remaining space being less than 10%, etc.), and triggers an alarm when the monitoring data exceeds these thresholds. For example, Chinese patent CN119892600A "Cloud cluster monitoring response method, device, equipment, storage medium and product" generates an alarm event according to the size relationship between the acquired monitoring indicator data and the pre-set alarm threshold. Although this static threshold strategy is simple and direct, it has exposed many shortcomings in actual application:
[0004] (1) The static threshold is difficult to adapt to the dynamic changes of different business scenarios and workloads, resulting in a large number of false positives or false negatives;
[0005] (2) The fixed threshold cannot reflect the timing characteristics and correlation of system behavior, making it difficult to identify complex failure patterns. In particular, in large-scale clusters, the node heterogeneity and workload diversity make the "one-size-fits-all" static threshold strategy ineffective.
[0006] However, the architectural complexity of modern distributed storage systems also brings new challenges to monitoring and alarming. Taking a typical distributed storage system as an example, it usually contains control nodes, data nodes, metadata nodes, cache nodes and other roles, each of which focuses on different monitoring indicators and health standards. At the same time, the inherent state dispersion of distributed systems makes it difficult to diagnose problems from a global perspective, and performance problems of a node may be indirectly caused by abnormal behavior of other nodes. In addition, the widespread application of virtualization and containerization technologies makes the deployment topology of storage services more dynamic and variable, and the traditional monitoring method based on fixed IP and port has been unable to meet the needs. Therefore, using the traditional static threshold strategy will cause the following shortcomings:
[0007] (1) Alarm storm is the primary challenge for large-scale distributed storage systems. When network failures or shared resource contention occur, a large number of similar alarms are often triggered in a short period of time. For example, network jitter in the machine room can cause all nodes in the cluster to report network communication exception alarms almost simultaneously, forming an alarm "tsunami". Although existing systems provide fine-grained alarm granularity (such as accurate to a single disk or service), they lack effective alarm aggregation and correlation analysis mechanisms, causing operations personnel to be overwhelmed by redundant information and making it difficult to identify the true root cause.
[0008] (2) Static threshold is not adaptive enough, which is another common shortcoming. Most existing systems use a fixed threshold strategy, such as triggering an alarm when the disk usage exceeds 85%. This approach ignores the dynamic characteristics and time correlation of workloads. For example, a CPU usage of 100% during batch processing may be normal, while the same value during an idle period may indicate an anomaly. Static thresholds cannot adapt to business cycle fluctuations and hardware performance degradation, resulting in either a high sensitivity that generates a large number of ineffective alarms or a threshold that is too loose and misses early failure signals. In particular, in distributed storage systems, the roles of different nodes (such as hot and cold data nodes) and the business pressure at different times (such as backup time windows) differ significantly, making it difficult for a uniform static threshold to meet actual needs.
[0009] (3) Low resource utilization efficiency is reflected in data collection and analysis. Existing solutions typically use periodic full-scan methods to obtain monitoring data, maintaining a fixed sampling frequency regardless of cluster load, resulting in unnecessary performance overhead. When monitoring data reaches 20,000 per second, the CPU usage of traditional monitoring services can become a bottleneck, causing data loss and alarm delays. In addition, raw monitoring data is often filtered simply and used directly for alarm judgment, lacking depth analysis and long-term trend mining, and failing to fully realize the value of data.
[0010] (4) High module coupling limits the scalability and flexibility of the system. Traditional solutions typically integrate data collection, monitoring, analysis, and alarm functions closely, forming a monolithic architecture. This design results in redundancy between parts, requiring multiple code changes to add new monitoring indicators or adjust alarm strategies, resulting in high maintenance costs and the potential for errors.
[0011] (5) Lack of fault self-recovery capability is a significant shortcoming of existing technologies. Most alarm systems focus only on problem detection and notification, lacking an automatic repair mechanism. When disk failures, node outages, and other common problems occur, manual intervention is still required, prolonging service unavailability. SUMMARY
[0012] Based on this, and in response to the aforementioned technical problems, a dynamic monitoring and alarm system for distributed storage systems is provided to solve the problems of existing distributed storage alarm systems, such as triggering a large number of similar alarms in a short period of time, insufficient adaptability of static thresholds leading to inaccurate alarms, and low resource utilization efficiency.
[0013] In a first aspect, a dynamic monitoring and alarm system for a distributed storage system, the system comprising:
[0014] Data acquisition layer: Used to acquire monitoring time-series data of each node in the distributed storage system through different monitoring acquisition devices, and adjust the acquisition frequency and granularity according to the primary characteristics of the monitoring time-series data;
[0015] Data transmission and storage layer: used to acquire the monitoring time-series data obtained by the monitoring acquisition device, and determine the corresponding separate storage database according to the IP information of the monitoring acquisition device; the location of the storage database is set based on the hierarchical storage strategy, and its specific hierarchical location is determined according to the second characteristic of the monitoring time-series data acquired by the monitoring acquisition device;
[0016] The intelligent analysis layer is used to preprocess and calculate metrics for the monitoring time-series data in the storage database. The preprocessed time-series data and corresponding metrics are input into a dynamic threshold prediction model to predict thresholds and generate alarm thresholds that adapt to workload changes. Initial alarm information is generated by comparing the monitoring time-series data in the storage databases of all nodes with the corresponding alarm thresholds. Granger causality tests are performed on all initial alarm information and corresponding monitoring time-series data to generate a monitoring data causal relationship graph. When any received monitoring time-series data exceeds the corresponding alarm threshold, it is recorded as the root monitoring time-series data. Sub-monitoring time-series data following the root monitoring time-series data in the relationship graph are then determined. Alarm event information is generated for the root monitoring time-series data, and alarm events are suppressed from being generated by the nodes corresponding to the sub-monitoring time-series data. A composite primary key is generated from the alarm event information to determine if a duplicate alarm event already exists. If not, an alarm event is issued. The alarm event information includes the generation time, alarm type, and alarm subject.
[0017] Alarm self-healing layer: Used to query the repair strategy library for corresponding repair action instructions based on the issued alarm event. If there is a corresponding repair action instruction, the repair action instruction is executed; otherwise, a prompt message is output.
[0018] Optionally, in the above scheme, the system further includes a visualization and interaction layer, which is used to display the issued alarm events, prompts, and the association graph; and has an operation entry point for user interaction.
[0019] Optionally, in the above scheme, the monitoring and acquisition device uses the native interface supported by each node in the distributed storage system for acquisition.
[0020] In the above scheme, optionally, the first characteristic includes: importance, frequency of change, and business requirements; the second characteristic includes: access frequency, data type, and retention period.
[0021] Optionally, in the above scheme, determining the corresponding individual storage database based on the IP information of the monitoring and acquisition device includes:
[0022] A consistent hash calculation is performed on the IP information of the monitoring and acquisition device to determine the corresponding separate storage database.
[0023] Optionally, in the above scheme, preprocessing the monitoring time-series data in the storage database includes: standardizing and enhancing the monitoring time-series data; the standardization and enhancement processing includes unit unification processing, missing value imputation processing, and feature engineering processing.
[0024] Optionally, in the above scheme, the calculation of the indicator includes: calculating the moving average and standard deviation.
[0025] Optionally, in the above scheme, the training of the dynamic threshold prediction model specifically includes:
[0026] The monitoring time-series data with occurrence timestamps from each node in the historical distributed storage system are obtained, and anomaly removal, missing interpolation, unit normalization, and time alignment are performed to obtain a stationary sequence of monitoring time-series data from each node with a unified timestamp.
[0027] Seasonal decomposition is performed on the stationary sequence of time series data monitored at each node with a unified timestamp to obtain the baseline component representing long-term changes and the fluctuation component representing random fluctuations.
[0028] A first model for predicting future baselines is trained using the baseline components, and a second model for estimating future fluctuation ranges is trained using the fluctuation components.
[0029] The baseline prediction value output by the first model and the fluctuation estimate value output by the second model are combined to form a dynamic threshold.
[0030] Optionally, in the above scheme, the intelligent analysis layer is further used for:
[0031] Based on the monitoring time-series data, indicators, and corresponding alarm thresholds in the storage database, trend prediction is performed, and trend prediction results are generated and output.
[0032] This application has at least the following beneficial effects:
[0033] This application's data acquisition layer acquires monitoring time-series data from each node in the distributed storage system using different monitoring acquisition devices, adjusting the acquisition frequency and granularity based on the primary characteristics of the monitoring time-series data. This differentiated strategy significantly reduces system overhead. Simultaneously, monitoring data for the same node in the data transmission and storage layers is always routed to the same database, with the database location determined by a hierarchical processing approach based on data properties, resulting in improved query performance and reduced storage costs. Furthermore, the intelligent analysis layer generates alarm thresholds that adapt to workload changes, effectively reducing false alarm rates during different business periods. It also constructs an alarm knowledge graph, using topological relationships and dependency chains to analyze alarm propagation paths, identify root causes of faults, issue alarms only to the root causes, retain only critical alarms, and suppress derived alarms. Additionally, it uses composite primary keys to determine if the same alarm information has been issued in a short period, thus suppressing alarm storms. The alarm self-healing layer can query the repair strategy library for corresponding repair action instructions based on alarm events. If found, it executes the repair action instructions, enabling fault self-healing.
[0034] Meanwhile, the monitoring and alarm process is divided into four independent modules: data collection, analysis and storage, display, and anomaly alarm. Each module communicates through a standard interface. This architecture allows each component to be expanded and upgraded independently. For example, when storage capacity is insufficient, the analysis and storage module can be expanded separately without affecting other services. Attached Figure Description
[0035] Figure 1 This is a structural diagram of a dynamic monitoring and alarm system for a distributed storage system provided in one embodiment of this application. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0037] In the description of this application: unless otherwise stated, "a plurality of" means two or more. The terms "first," "second," "third," etc., in this application are intended to distinguish the objects referred to and do not have any special meaning in terms of technical connotation (e.g., they should not be construed as an emphasis on importance or order). Expressions such as "comprising," "including," and "having" also mean "not limited to" (certain units, components, materials, steps, etc.).
[0038] In one embodiment, such as Figure 1 As shown, a dynamic monitoring and alarm system for a distributed storage system is provided, the system comprising:
[0039] Data acquisition layer: Used to acquire monitoring time-series data of each node in the distributed storage system through different monitoring acquisition devices, and adjust the acquisition frequency and granularity according to the primary characteristics of the monitoring time-series data.
[0040] In the data acquisition layer, monitoring and acquisition devices use the native interfaces supported by each node in the distributed storage system for data collection. Different data interfaces are used for different types of data sources, such as obtaining network device status via the SNMP protocol, extracting detailed storage array information using the CLI or API specific to each storage vendor, calling Linux tools (such as iostat and smartctl) to collect host-level metrics, and obtaining application-level performance data through the Instrumentation framework.
[0041] Specifically, the data acquisition layer, as the foundation of the entire monitoring system, directly impacts the quality of monitoring data and system overhead. This application employs an asynchronous, timed acquisition framework based on the Python Celery library, which supports two data acquisition modes: the data acquisition function built into the distributed storage system and user-defined acquisition tasks. In actual deployment, the system dynamically adjusts the acquisition frequency and granularity according to the characteristics of the monitoring target—using high-frequency acquisition at the second level for key performance indicators (such as latency and IOPS), and using minute-level acquisition for slowly changing capacity indicators. This differentiated strategy significantly reduces system overhead. The first characteristics in this application include: importance, frequency of change, and business requirements.
[0042] Data transmission and storage layer: used to acquire the monitoring time-series data obtained by the monitoring acquisition device, and determine the corresponding separate storage database according to the IP information of the monitoring acquisition device; the location of the storage database is set based on the hierarchical storage strategy, and its specific hierarchical location is determined according to the second characteristic of the monitoring time-series data acquired by the monitoring acquisition device.
[0043] Specifically, the second characteristic includes: access frequency, data type, and retention period.
[0044] Meanwhile, the data transmission and storage layer is responsible for the efficient flow and persistence of monitoring data. To address the throughput issue of monitoring data in large-scale clusters, this application proposes an innovative distribution mechanism: the distribution device (such as a server running HAProxy) receives time-series data sent by all monitoring acquisition devices, then performs consistent hash calculations on the IP information of each acquisition device to determine the corresponding storage database, achieving database sharding and load balancing of monitoring data. This design ensures that monitoring data from the same node is always routed to the same database, facilitating subsequent analysis and querying, while avoiding the performance bottleneck of a single database. Data storage adopts a hierarchical strategy: after the raw data is cleaned, temporary data (such as CPU utilization in the last 5 minutes) is stored in Redis for fast access; time-series indicators (such as disk read / write history) are stored in OpenTSDB to support time-range queries; and configuration and alarm logs that need to be retained long-term are persisted to MySQL. Implementation in a financial system shows that this hierarchical storage scheme improves query performance by 3 times and reduces storage costs by 40%.
[0045] The intelligent analysis layer is used to preprocess and calculate metrics for the monitoring time-series data in the storage database. The preprocessed time-series data and corresponding metrics are input into a dynamic threshold prediction model to predict thresholds and generate alarm thresholds that adapt to workload changes. Initial alarm information is generated by comparing the monitoring time-series data in the storage databases of all nodes with the corresponding alarm thresholds. Granger causality tests are performed on all initial alarm information and corresponding monitoring time-series data to generate a monitoring data causal relationship graph. When any received monitoring time-series data exceeds the corresponding alarm threshold, it is recorded as the root monitoring time-series data. Sub-monitoring time-series data following the root monitoring time-series data in the relationship graph are determined, and alarm event information is generated for the root monitoring time-series data. The generation of alarm events by the nodes corresponding to the sub-monitoring time-series data is suppressed. A composite primary key is generated from the alarm event information to determine if a duplicate alarm event already exists. If not, an alarm event is issued. The alarm event information includes the generation time, alarm type, and alarm subject.
[0046] Specifically, the intelligent analytics layer is the core of the dynamic monitoring strategy, realizing the transformation from raw data to operational insights. The advanced analytics pipeline comprises multiple processing stages: First, the raw data is standardized and enhanced, including unit unification, missing value imputation, and feature engineering; then, stream processing technologies (such as Apache Flink) are applied for real-time statistical analysis, calculating derived indicators such as moving averages and standard deviations; next, the data is fed into a dynamic threshold engine, which combines seasonal decomposition and machine learning prediction technologies (such as Facebook Prophet) to generate alarm thresholds that adapt to workload changes; finally, correlation analysis is performed between multiple monitoring data, using methods such as Granger causality tests to identify lead-lag relationships between monitoring data and construct a correlation graph between monitoring data. A joint primary key mechanism (generation time + alarm type + alarm subject) enables accurate alarm deduplication, preventing the same anomaly from being reported multiple times. The results output by the analytics layer include real-time alarms, trend predictions, capacity planning suggestions, and other intelligent operational products.
[0047] Meanwhile, this application's dynamic threshold adjustment algorithm, unlike fixed thresholds, automatically calculates a reasonable range based on historical data and context. This application proposes a threshold modeling method based on extreme value theory. This technique first collects and processes the raw data of database indicators, converting it into an input format suitable for threshold model matching. Then, it calibrates the thresholds for each node. Finally, it achieves accurate alarms by comparing real-time monitoring indicators with the dynamically calculated upper and lower limits of the thresholds. A more advanced implementation combines time series forecasting techniques, such as using ARIMA or LSTM models to predict the normal fluctuation range of the threshold. An alarm is triggered when the actual value significantly deviates from the predicted range. This method is particularly suitable for workloads with obvious periodic characteristics (such as scenarios where daytime online trading alternates with nighttime batch processing), effectively reducing the false alarm rate during different business periods.
[0048] Furthermore, the multi-dimensional alarm correlation analysis technology is dedicated to solving the alarm storm problem. Its core idea is to aggregate and correlate original alarms from multiple dimensions such as time, space, and semantics to identify the root cause. Specifically, it first determines the original alarms based on dynamic thresholds and initial data, then analyzes the correlation between multiple alarms to construct an alarm knowledge graph, and uses topological relationships and dependency chains to analyze alarm propagation paths. For example, when a power failure in a server rack causes multiple storage nodes to go offline simultaneously, the system can identify the power failure as the root cause, retaining only critical alarms and suppressing derived alarms. Simultaneously, it merges alarms with identical content within a deduplication configuration-defined time window; then, it filters out transient anomalies (such as instantaneous peak CPU usage) through debouncing configuration; finally, it distributes the refined alarms to the relevant responsible parties according to permission configurations.
[0049] Alarm self-healing layer: Used to query the repair strategy library for corresponding repair action instructions based on the issued alarm event. If there is a corresponding repair action instruction, the repair action instruction is executed; otherwise, a prompt message is output.
[0050] The alarm self-healing layer upgrades traditional passive alarms to an active repair system. This application describes a batch task monitoring framework based on database distributed locks. This framework designs multiple status tables, such as a monitoring execution table, a delay monitoring table, and an alarm execution table, to decouple the monitoring and alarm processes. When an anomaly is detected, the system first queries a predefined repair strategy library. If an applicable strategy is found, the system automatically executes the repair action (such as restarting the service or switching replicas). If no matching strategy is found, the system proceeds to manual processing, while recording the resolution process to enrich the strategy library. More advanced systems will introduce a reinforcement learning mechanism to continuously try different repair actions and observe their effects, gradually optimizing the self-healing strategy and forming a complete closed loop of "monitoring-analysis-repair-learning".
[0051] The fault self-healing mechanism proposed in this application represents an advanced stage of development for alarm systems. The proposed solution not only detects anomalies but also executes repair actions based on customized monitoring items: for known fault modes (such as disk full or process crashes), the system automatically attempts recovery (e.g., cleaning temporary files or restarting services); when automatic repair fails or encounters unknown anomalies, detailed diagnostic information and manual troubleshooting guidelines are provided; simultaneously, a machine learning mechanism is introduced to train the self-healing model using feedback from manual repair experience, forming a closed-loop learning system. A vendor's practice shows that this self-healing mechanism can handle approximately 60% of common storage faults, reducing the mean time to repair (MTTR) by more than 70%. More complex implementations will incorporate chaos engineering principles, proactively injecting fault tolerance capabilities into the fault testing system and identifying monitoring blind spots in advance.
[0052] In the dynamic monitoring and alarm system of the aforementioned distributed storage system, the data acquisition layer obtains monitoring time-series data of each node in the distributed storage system through different monitoring acquisition devices, and adjusts the acquisition frequency and granularity based on the primary characteristics of the monitoring time-series data. This differentiated strategy significantly reduces system overhead. Simultaneously, monitoring data for the same node in the data transmission and storage layers is always routed to the same database, and the database location is determined through hierarchical processing based on the nature of the data, resulting in improved query performance and reduced storage costs. Furthermore, the intelligent analysis layer generates alarm thresholds that adapt to workload changes, effectively reducing false alarm rates during different business periods; it also constructs an alarm knowledge graph, using topological relationships and dependency chains to analyze alarm propagation paths, identify root causes of faults, issue alarms only for the root causes, retain only critical alarms, and suppress derived alarms; and it uses composite primary keys to determine whether the same alarm information has been issued in a short period of time, thus suppressing alarm storms. The alarm self-healing layer can query the repair strategy library for corresponding repair action instructions based on alarm events; if found, it executes the repair action instructions, enabling fault self-healing.
[0053] Meanwhile, a microservice-based monitoring architecture is adopted to improve scalability by decoupling system components. This application proposes dividing the monitoring and alarm process into four independent modules: data acquisition, analysis and storage, display, and anomaly alarms. These modules communicate through standard interfaces. The data acquisition layer is developed based on an asynchronous timing framework using Python Celery, supporting flexible combinations of built-in acquisition functions of distributed storage systems and user-defined tasks. The analysis and storage layer selects different storage engines such as Redis, OpenTSDB, or MySQL based on data types (temporary, time-series, long-term storage, etc.). The display layer supports default views and custom dashboards. The anomaly alarm layer integrates a rule engine with multiple notification channels. This architecture allows each component to be independently expanded and upgraded; for example, the analysis and storage module can be expanded separately when storage capacity is insufficient without affecting other services.
[0054] In one embodiment, the system further includes a visualization and interaction layer, which is used to display alarm events, prompts, and the association graph; and has an operation entry point for user interaction.
[0055] In this embodiment, the visualization and interaction layer provides users with an intuitive understanding of system status and operational access. Modern monitoring systems support two display modes: the default display presets key indicators and views based on common needs of distributed storage clusters, such as topology diagrams, performance heatmaps, and capacity level tables; the customized display allows users to freely choose data sources and visualization formats (line charts, bar charts, scatter plots, pie charts, dashboards, etc.) and save them as personal dashboards. Visualization design should adhere to the principle of minimizing cognitive load, using color coding (e.g., green / yellow / red for normal / warning / abnormal), animation transitions, drill-down analysis, and other techniques to help users quickly locate problems without being overwhelmed by complex data. The interaction layer also provides collaborative tools such as alarm confirmation, fault marking, and root cause analysis, supporting multi-person teams to efficiently handle large-scale cluster operation and maintenance events.
[0056] In one embodiment, training the dynamic threshold prediction model specifically includes:
[0057] The monitoring time-series data with occurrence timestamps from each node in the historical distributed storage system are obtained, and anomaly removal, missing interpolation, unit normalization, and time alignment are performed to obtain a stationary sequence of monitoring time-series data from each node with a unified timestamp.
[0058] Seasonal decomposition is performed on the stationary sequence of time series data monitored at each node with a unified timestamp to obtain the baseline component representing long-term changes and the fluctuation component representing random fluctuations.
[0059] A first model for predicting future baselines is trained using the baseline components, and a second model for estimating future fluctuation ranges is trained using the fluctuation components.
[0060] The baseline prediction value output by the first model and the fluctuation estimate value output by the second model are combined to form a dynamic threshold.
[0061] In one embodiment, modern distributed storage alerting systems typically employ a hybrid topology in their actual deployment architecture: lightweight collectors are deployed on each storage node to perform edge computing; regional aggregation nodes are responsible for intermediate data aggregation and preliminary analysis; and central nodes run complex correlation analysis and machine learning algorithms. This layered processing architecture can support large-scale clusters with over 100,000 monitoring items per second while maintaining alert latency in the sub-second range. All functional components are designed as containerized microservices, supporting elastic scaling under Kubernetes orchestration to ensure maximum resource utilization efficiency.
[0062] In this embodiment, during actual deployment, data acquisition and storage are achieved by deploying data collectors at storage nodes and incorporating the data acquisition layer and data transmission and storage layer functions of the system described in this application. Data calculation and analysis are performed at regional aggregation nodes, which incorporate some functions of the intelligent analysis layer, such as preprocessing monitoring time-series data in the storage database and calculating indicators. The central node is equipped with a dynamic threshold prediction engine for dynamic threshold prediction, performs Granger causality testing, generates indicator correlation graphs, makes final alarm decisions, and incorporates an alarm self-healing layer for alarm self-healing.
[0063] In this embodiment, edge intelligent analytics optimizes resource utilization efficiency. Traditional centralized analytics requires transmitting all monitoring data to a central node, putting pressure on network bandwidth and computing resources. This paper proposes performing monitoring metric collection and preliminary analysis on child nodes of a distributed file system, only encapsulating abnormal data exceeding thresholds into alarm information and sending it to the master node. This edge computing model significantly reduces data transmission volume, making it particularly suitable for large-scale distributed storage environments. Further research deploys lightweight machine learning models to storage nodes to achieve local real-time anomaly detection, uploading only the detection results and feature vectors, rather than the raw data, to the central node, reducing system overhead while maintaining analytical accuracy.
[0064] The priority of this application is as follows:
[0065] (1) Reduce duplicate alarms and associated alarm explosions to alleviate the fatigue of operation and maintenance personnel and prevent real problems from being covered up;
[0066] (2) Use dynamic thresholds to adapt to dynamic load and reduce false alarm rate or false alarm rate;
[0067] (3) Decouple the data acquisition, analysis, and alarm modules to improve system scalability;
[0068] (4) Provide self-healing capabilities to reduce human intervention and improve service availability;
[0069] (5) Make full use of resources, edge sampling, data analysis, and reduce single-machine overhead.
[0070] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A dynamic monitoring and alarm system for a distributed storage system, characterized in that, The system includes: Data acquisition layer: Used to acquire monitoring time-series data of each node in the distributed storage system through different monitoring acquisition devices, and adjust the acquisition frequency and granularity according to the primary characteristics of the monitoring time-series data; Data transmission and storage layer: used to acquire the monitoring time-series data obtained by the monitoring acquisition device, and determine the corresponding separate storage database according to the IP information of the monitoring acquisition device; the location of the storage database is set based on the hierarchical storage strategy, and its specific hierarchical location is determined according to the second characteristic of the monitoring time-series data acquired by the monitoring acquisition device; The intelligent analysis layer is used to preprocess and calculate metrics for the monitoring time-series data in the storage database. The preprocessed time-series data and corresponding metrics are input into a dynamic threshold prediction model to predict thresholds and generate alarm thresholds that adapt to workload changes. Initial alarm information is generated by comparing the monitoring time-series data in the storage databases of all nodes with the corresponding alarm thresholds. Granger causality tests are performed on all initial alarm information and corresponding monitoring time-series data to generate a monitoring data causal relationship graph. When any received monitoring time-series data exceeds the corresponding alarm threshold, it is recorded as the root monitoring time-series data. Sub-monitoring time-series data following the root monitoring time-series data in the relationship graph are then determined. Alarm event information is generated for the root monitoring time-series data, and alarm events are suppressed from being generated by the nodes corresponding to the sub-monitoring time-series data. A composite primary key is generated from the alarm event information to determine if a duplicate alarm event already exists. If not, an alarm event is issued. The alarm event information includes the generation time, alarm type, and alarm subject. Alarm self-healing layer: Used to query the repair strategy library for corresponding repair action instructions based on the issued alarm event. If there is a corresponding repair action instruction, the repair action instruction is executed; otherwise, a prompt message is output.
2. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The system also includes a visualization and interaction layer, which is used to display alarm events, prompts, and the associated graph; and has an operation entry point for user interaction.
3. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The monitoring and acquisition device uses the native interfaces supported by each node in the distributed storage system for data acquisition.
4. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The first characteristic includes: importance, frequency of change, and business requirements; the second characteristic includes: access frequency, data type, and retention period.
5. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, Determining the corresponding individual storage database based on the IP information of the monitoring and acquisition device includes: A consistent hash calculation is performed on the IP information of the monitoring and acquisition device to determine the corresponding separate storage database.
6. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, Preprocessing the monitoring time-series data in the storage database includes: standardizing and enhancing the monitoring time-series data; the standardization and enhancement processes include unit unification, missing value imputation, and feature engineering.
7. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The calculation of the indicators includes: calculating the moving average and standard deviation.
8. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The training of the dynamic threshold prediction model specifically includes: The monitoring time-series data with occurrence timestamps from each node in the historical distributed storage system are obtained, and anomaly removal, missing interpolation, unit normalization, and time alignment are performed to obtain a stationary sequence of monitoring time-series data from each node with a unified timestamp. Seasonal decomposition is performed on the stationary sequence of time series data monitored at each node with a unified timestamp to obtain the baseline component representing long-term changes and the fluctuation component representing random fluctuations. A first model for predicting future baselines is trained using the baseline components, and a second model for estimating future fluctuation ranges is trained using the fluctuation components. The baseline prediction value output by the first model and the fluctuation estimate value output by the second model are combined to form a dynamic threshold.
9. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The intelligent analysis layer is also used for: Based on the monitoring time-series data, indicators, and corresponding alarm thresholds in the storage database, trend prediction is performed, and trend prediction results are generated and output.
10. The dynamic monitoring and alarm system for a distributed storage system according to claim 1, characterized in that, The alarm self-healing layer is used to record all manually input repair execution instructions corresponding to the issued alarm events, and add them to the repair strategy library.
Citation Information
Patent Citations
Cloud cluster monitoring response method and device, equipment, storage medium and product
CN119892600A
Cited By
Monitoring and automatic alarm method and system
CN121455777A