Alarm management method and device, electronic equipment and medium
By grading and managing the alarm information of the distributed storage system, the problem of excessive alarm information is solved, the identification and effective prevention of potential risks is achieved, and the stability and operation and maintenance efficiency of the system are improved.
Patent Information
- Application Number
- CN202510667517.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-15
AI Technical Summary
The alarm management method of distributed storage systems in the prior art is prone to generate a large amount of unnecessary alarm information, resulting in the neglect of emergencies and lack of ability to identify and prevent potential risks.
By obtaining the alarm information in the distributed storage system, performing hierarchical processing according to preset rules, determining the level of alarm information and setting a time window, managing the alarm status to achieve cancellation or continuous status, reducing unnecessary alarm information and identifying potential problems.
Effectively filter unnecessary alarm information to ensure that key alarm information is not ignored, and improve the operation and maintenance efficiency and system stability of distributed storage systems.
Smart Images

Figure CN120492981A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of alarm management technology, and in particular to alarm management methods, devices, electronic equipment and media. Background Art
[0002] Alarm management solutions typically employ a simple threshold-based trigger mechanism, generating an immediate alarm when a system metric exceeds a preset threshold. This approach can easily generate a large amount of redundant alarm information, leading to information overload and alarm fatigue for operations and maintenance personnel, and potentially causing truly urgent issues to go unnoticed. Furthermore, this approach lacks predictive capabilities and intelligent processing mechanisms for alarm information, making it difficult to proactively identify potential risks and implement effective preventative measures. Summary of the Invention
[0003] The present application provides an alarm management method, device, electronic device and medium to at least solve the problem in the related art that a large amount of unnecessary alarm information is easily generated, resulting in real emergencies being ignored, and it is difficult to identify potential risks in advance and take effective preventive measures.
[0004] This application provides an alarm management method, including:
[0005] Obtaining first alarm information in the distributed storage system;
[0006] Performing hierarchical processing on the first alarm information according to a first preset rule to obtain a level corresponding to the first alarm information;
[0007] Determining a time window corresponding to the first alarm information based on the level corresponding to the first alarm information;
[0008] The first alarm information is managed based on the time window and the alarm status of the first alarm information, where the alarm status includes at least one of a canceled state and a persistent state.
[0009] This application also provides an alarm management device, comprising:
[0010] An acquiring unit, configured to acquire first alarm information in a distributed storage system;
[0011] a grading unit, configured to perform grading processing on the first alarm information according to a first preset rule to obtain a level corresponding to the first alarm information;
[0012] a determining unit, configured to determine a time window corresponding to the first alarm information based on a level corresponding to the first alarm information;
[0013] The alarm management unit is used to manage the first alarm information based on the time window and the alarm status of the first alarm information, where the alarm status includes at least one of a cancel state and a continuous state.
[0014] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned alarm management methods when executing the computer program.
[0015] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned alarm management methods are implemented.
[0016] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned alarm management methods when executed by a processor.
[0017] Through the present application, the first alarm information in the distributed storage system is obtained; the first alarm information is graded according to the first preset rule to obtain the level corresponding to the first alarm information; based on the level corresponding to the first alarm information, the time window corresponding to the first alarm information is determined; based on the time window and the alarm status of the first alarm information, the first alarm information is managed, and the alarm status includes at least one of a cancellation state and a continuous state. This solves the technical problems in related schemes that a large amount of unnecessary alarm information is easily generated, resulting in real emergency situations being ignored, and it is difficult to identify potential alarm information, and achieves the technical effect of filtering unnecessary alarm information and determining potential alarm information at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of an alarm management method provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of the structure of an alarm management module provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the association relationship between first alarm information provided in an embodiment of the present application;
[0022] Figure 4 A flowchart of a delayed alarm method provided in an embodiment of the present application;
[0023] Figure 5 A schematic diagram of the structure of an alarm management device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0026] In order to facilitate those skilled in the art to better understand the technical solutions described in the embodiments of the present disclosure, the technical terms in the embodiments of the present disclosure are explained as follows before introducing the embodiments of the present disclosure.
[0027] Object Storage Service (OSS): OSS is a data storage solution that provides high availability, scalability, and high performance. It allows users to store and retrieve any type of data in the form of objects. Each object contains the data itself, its related metadata, and a globally unique identifier. It is suitable for storing large amounts of unstructured data, such as images, videos, backup files, etc., and supports access through standard APIs.
[0028] Cluster Control Service (CCS): CCS is one of the core components of a distributed storage system, responsible for managing and coordinating the operations of each node in the cluster. Its main responsibilities include, but are not limited to, monitoring cluster status, managing resource allocation and scheduling, ensuring data consistency, executing fault detection and recovery mechanisms, and processing alarm information. Through these functions, CCS can ensure the efficient operation and stability of the entire distributed storage system.
[0029] Distributed Metadata Server (DMS): DMS is a service for managing metadata in distributed storage environments. Metadata refers to data that describes data, such as the name, size, creation time, and location of a file or object. DMS efficiently stores, queries, and manages this metadata, supporting rapid location and access requirements in large-scale data storage systems. Through its distributed architecture, DMS provides enhanced reliability, scalability, and performance to meet the needs of complex application scenarios.
[0030] Object Storage Device (OSD): An OSD cluster is a network consisting of a series of interconnected object storage devices. Each device can independently store data objects and their metadata. This architecture allows data and computing power to be distributed throughout the cluster rather than concentrated on a single node, thereby improving system reliability and performance.
[0031] Central Processing Unit (CPU): The CPU is the core component of the computer, responsible for executing instructions in the instruction set to process data. It is the brain of the computer, performing arithmetic and logical operations, controlling data flow, and executing instructions of the operating system.
[0032] Distributed storage systems, as data storage solutions that offer high availability, scalability, and performance, have been widely adopted across various fields. However, as systems grow in size and complexity, the management and maintenance of distributed storage systems face increasing challenges. Alarm management is a crucial component of system operations and maintenance, directly impacting system reliability and stability.
[0033] Traditional alarm management methods typically rely on simple threshold settings, triggering an alarm when a system metric exceeds a threshold. However, this approach has two major problems: first, it easily generates a large number of unnecessary alarm messages, leading to manager fatigue and potentially overlooking true emergencies; second, it lacks the ability to predict and intelligently process alarm information, making it difficult to effectively prevent potential problems.
[0034] Through the present application, the first alarm information in the distributed storage system is obtained; the first alarm information is graded according to the first preset rule to obtain the level corresponding to the first alarm information; based on the level corresponding to the first alarm information, the time window corresponding to the first alarm information is determined; based on the time window and the alarm status of the first alarm information, the first alarm information is managed, and the alarm status includes at least one of a cancellation state and a continuous state. This solves the technical problems in related schemes that a large amount of unnecessary alarm information is easily generated, resulting in real emergency situations being ignored, and it is difficult to identify potential alarm information, and achieves the technical effect of filtering unnecessary alarm information and determining potential alarm information at the same time.
[0035] The present disclosure provides an alarm management method applicable to scenarios such as data center operations and maintenance management and cloud computing service platforms that rely on distributed storage systems to support their business operations or service provision. This method can be executed by an alarm management module or system, specifically a dedicated component or service within the distributed storage system architecture responsible for monitoring, collecting, processing, and managing alarm information.
[0036] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0037] Figure 1 A flowchart of an alarm management method provided by an embodiment of the present disclosure.
[0038] like Figure 1 As shown, the method comprises the following steps:
[0039] Step 101: Obtain first alarm information in a distributed storage system;
[0040] In some embodiments, the alarm management method provided by the present application is implemented based on a distributed storage system. The distributed storage system includes core storage components such as CCS, OSS, and DMS, as well as protocol connection components such as objects, blocks, and files. Each of the aforementioned modules has a corresponding alarm interface. At the same time, each node in the distributed storage system has an inspection script for monitoring the hardware status and generating an alarm message when the hardware status is abnormal. Specifically, once any abnormal situation is detected, such as disk failure, network slowdown, insufficient memory, or an error in a service, the monitoring tool will automatically generate a corresponding first alarm message.
[0041] In some embodiments, data that may indicate abnormal conditions can be collected from various components and nodes of the distributed storage system to determine a first alarm message, which includes details about the abnormality, such as the time of occurrence, abnormality type, severity, and specific error description.
[0042] In some embodiments, as Figure 2 As shown, Figure 2 A structural diagram of an alarm management module provided in an embodiment of the present application. After obtaining the first alarm information in the distributed storage system, it is necessary to standardize the structure of the obtained first alarm information and concentrate all the first alarm information in an alarm management module to facilitate the management of the alarm information. Specifically, the alarm management module includes at least an object protocol module, a file protocol module, a block protocol module, a CSS cluster, and an OSD cluster.
[0043] Step 102: performing classification processing on the first alarm information according to a first preset rule to obtain a level corresponding to the first alarm information;
[0044] In some embodiments, the first alarm information can be divided into the following levels: minor exceptions, general exceptions, serious exceptions and emergency exceptions. Among them, minor exceptions refer to exceptions that can be recovered in a short time, general exceptions refer to exceptions that do not affect the operation of the distributed storage system, serious exceptions refer to situations that cause performance degradation and insufficient capacity of the distributed storage system, and emergency exceptions refer to exceptions that make the distributed sequential system unusable.
[0045] In some embodiments, the present application does not limit the level corresponding to the first alarm information, which can be the aforementioned four levels, or the first alarm information can be divided into more and finer levels according to actual needs.
[0046] In some embodiments, the first preset rule refers to a preset specific grading rule for dividing the first alarm information. The first preset rule is usually determined based on the abnormality type of the first alarm information, the scope of impact of the abnormality, and the possibility and duration of abnormal recovery. Specifically, the abnormality type may include hardware failure, such as disk failure, network slowdown, component failure, such as single component abnormality, storage pool abnormality, such as insufficient capacity, data inconsistency; the scope of impact of the abnormality includes whether it is a local problem or a global impact and whether it will affect the key functions of the system; the possibility and duration of abnormal recovery include whether the abnormality can be automatically recovered in a short time and whether manual intervention is required.
[0047] In some embodiments, each first alarm message is evaluated according to the above rules and assigned an appropriate level label through an automated script or algorithm.
[0048] Step 103: determining a time window corresponding to the first alarm information based on the level corresponding to the first alarm information;
[0049] In some embodiments, it is necessary to set corresponding processing time windows for alarms of different levels. These time windows are determined based on a comprehensive consideration of factors such as the potential impact of the alarm, the time required to resolve it, and the time the system tolerates abnormalities. For example, in the aforementioned embodiment, the four levels of minor abnormalities, general abnormalities, severe abnormalities, and emergency abnormalities can set time windows to 10 minutes, 5 minutes, 2 minutes, and 30 seconds, respectively. This application does not limit the specific value of the time window, and it can be adjusted according to needs.
[0050] In some embodiments, for minor anomalies, such as brief input / output delays that are expected to recover quickly, a 10-minute time window can be set. For severe anomalies, such as persistently high input / output error rates that affect some data read and write operations, a 30-minute time window might be set. For urgent anomalies, such as complete disk damage that causes the entire node to fail, a smaller time window should be set, or no time window should be set at all.
[0051] In some embodiments, by determining the time window corresponding to the first alarm information based on the level corresponding to the first alarm information, the alarm processing process in the distributed storage system can be effectively managed and optimized, which not only reduces unnecessary alarm notifications, but also ensures that key alarm information will not be ignored, thereby improving the operation and maintenance efficiency of the distributed storage system.
[0052] Step 104: Manage the first alarm information based on the time window and the alarm status of the first alarm information. The alarm status includes at least one of a canceled state and a persistent state.
[0053] In some embodiments, the canceled state indicates that the problem indicated by the alarm has been resolved or returned to normal, and the persistent state indicates that the problem indicated by the alarm still exists and has not been resolved.
[0054] In some embodiments, for the first alarm message of a minor anomaly, if the alarm status is canceled within the set time window, the alarm message will not be reported to avoid excessive false alarms; if the alarm is still in a "persistent state" at the end of the time window, it will be decided whether to report or continue to delay reporting based on the specific situation. For the first alarm message of a general anomaly, if the alarm status changes to a canceled state within this time period, it will not be reported. If the alarm status is a persistent state exceeding the time window, it should be reported to the relevant personnel, but may be handled with a lower priority because such problems have less impact on system operation; for the first alarm message of a serious anomaly, if the alarm status changes to a canceled state within the specified time, no further action is required; otherwise, it needs to be reported immediately and corresponding measures must be taken to prevent the situation from worsening; for the first alarm message of an emergency anomaly, once detected, it will be reported immediately and the highest level response mechanism will be activated to ensure that quick action can be taken to restore service.
[0055] Through the present application, the first alarm information in the distributed storage system is obtained; the first alarm information is graded according to the first preset rule to obtain the level corresponding to the first alarm information; based on the level corresponding to the first alarm information, the time window corresponding to the first alarm information is determined; based on the time window and the alarm status of the first alarm information, the first alarm information is managed, and the alarm status includes at least one of a cancellation state and a continuous state. This solves the technical problems in related schemes that a large amount of unnecessary alarm information is easily generated, resulting in real emergency situations being ignored, and it is difficult to identify potential alarm information, and achieves the technical effect of filtering unnecessary alarm information and determining potential alarm information at the same time.
[0056] In some embodiments, obtaining first alarm information in the distributed storage system includes:
[0057] Get the hardware status of the distributed storage system;
[0058] In some embodiments, the hardware status of the distributed storage system may be acquired using a monitoring tool or inspection script deployed on each node (server) in the distributed storage system.
[0059] In some embodiments, hardware status refers to the status of hardware components, including but not limited to disk, memory, network interface, power supply, etc.
[0060] In some embodiments, the monitoring tool may periodically collect key performance indicators from various hardware components, such as CPU usage, memory utilization, disk read and write speeds, network latency, and bandwidth utilization.
[0061] In some embodiments, the hardware status data collected from each node can be transmitted to a central management system or a dedicated alarm management module for preprocessing, including but not limited to data cleaning and format conversion of the collected data.
[0062] Based on the hardware status, first alarm information corresponding to the hardware status is determined.
[0063] In some embodiments, the first alarm information corresponding to the hardware status may be determined based on historical data analysis results. For example, if the CPU utilization rate exceeds 80%, it is considered that an alarm is required.
[0064] In some embodiments, by determining the first alarm information corresponding to the hardware status based on the hardware status, the hardware health of the distributed storage system can be effectively monitored, potential problems can be discovered in time and corresponding measures can be taken, thereby ensuring the stability and reliability of the system.
[0065] In some embodiments, managing the first alarm information based on the time window and the alarm status of the first alarm information includes:
[0066] In response to the alarm state of the first alarm information being a canceled state within the time window, storing the first alarm information and stopping reporting the first alarm information;
[0067] In some embodiments, the distributed storage system continuously monitors the status change of the first alarm information. If the alarm status changes from a persistent state to a canceled state within a specified time window, such as 5 minutes, 30 minutes, etc., depending on the level of the first alarm information, it indicates that the anomaly has been automatically resolved or restored to normal.
[0068] In some embodiments, even though the alarm has been canceled, the distributed storage system still needs to record this event for subsequent analysis. Therefore, this first alarm information needs to be stored in a dedicated log or database. The stored information should include but is not limited to the timestamp of the alarm, the alarm type, the severity of the alarm, and the recovery time of the alarm.
[0069] In some embodiments, stopping reporting the first alarm information means that the alarm has been canceled and has no further impact. The distributed storage system immediately stops all reporting processes for this alarm, indicating that no notification will be sent to the operation and maintenance personnel, nor will it be displayed on the real-time alarm panel, thereby reducing unnecessary alarm interference and avoiding alarm fatigue.
[0070] In some embodiments, the alarm management module should update its internal status and mark the alarm as processed to ensure that future queries can accurately reflect the true life cycle of the alarm.
[0071] In response to the alarm status of the first alarm information being a persistent state within the time window, second alarm information is determined based on the level corresponding to the first alarm information and reported.
[0072] In some embodiments, a low-level alarm may cause a high-level alarm. If the low-level alarm is not canceled within the time window, the low-level alarm and the high-level alarm need to be reported at the same time to ensure that the alarm information is not missed. Specifically, the low-level alarm may cause a high-level alarm, including but not limited to hardware abnormalities may cause software component abnormalities, a single component abnormality may cause similar component cluster abnormalities or data security risks, component cluster abnormalities may cause cluster service abnormalities and thus lead to customer business interruption, etc.
[0073] In some embodiments, by managing the first alarm information based on the time window and the alarm status of the first alarm information, automatic upgrade and accurate reporting of the alarm can be achieved, thereby improving system reliability.
[0074] In some embodiments, based on the level corresponding to the first alarm information, determining the second alarm information and reporting the second alarm information includes:
[0075] Establishing an association relationship between the first alarm information based on the levels corresponding to the first alarm information;
[0076] In some embodiments, the association relationship between the first alarm information can be expressed in the form of a directed graph, a table, or a tree structure.
[0077] In some embodiments, it can be determined based on historical data analysis, historical experience, or system models which levels of alarms may trigger other levels of alarms. For example, a disk failure in a hardware failure may cause a component failure, such as a service being unable to start, which in turn causes a storage pool abnormality. If minor abnormalities are not handled in a timely manner, they may escalate to more serious abnormalities.
[0078] Based on the association relationship, the second alarm information is determined and reported.
[0079] In some embodiments, when a first alarm message is received, all other alarms that may be directly or indirectly related to this alarm are identified based on the aforementioned association relationships, such as the alarm relationship directed graph. For example, if an increase in disk input / output delay is detected (minor abnormality), it should be considered that it may cause a longer service response time (general abnormality) or failure of other related hardware or software components.
[0080] In some embodiments, the second alarm information refers to other alarm information that may be directly or indirectly caused by the first alarm information, and is usually alarm information with a higher level than the first alarm information.
[0081] In some embodiments, by determining the second alarm information based on the association relationship and reporting the second alarm information, an association relationship between the alarm information can be effectively established, which not only improves the intelligence level of alarm processing, but also enhances the stability and reliability of the system.
[0082] In some embodiments, establishing an association relationship between the first alarm information based on the levels corresponding to the first alarm information includes:
[0083] Classify the first alarm information according to a second preset rule to obtain a category corresponding to the first alarm information;
[0084] In some embodiments, the categories corresponding to the first alarm information may be hardware failure, component failure, and storage pool abnormality. Specifically, hardware failure includes disk failure, network slowdown, insufficient memory, node abnormality, etc.; component failure refers to a single abnormality of a certain type of component; storage pool abnormality includes insufficient capacity, inconsistent data, abnormal cluster status, etc.
[0085] In some embodiments, the present application does not limit the categories corresponding to the first alarm information. It can be the three categories mentioned above, or the first alarm information can be divided into more and more detailed categories according to actual needs.
[0086] In some embodiments, the second preset rule refers to a preset standard or condition for classifying alarm information. The second preset rule can be determined based on system requirements and operation and maintenance experience. Specifically, the second preset rule can include, but is not limited to, the source of the alarm (hardware, software), the specific manifestation of the anomaly (disk failure, network slowdown, etc.), and the affected business module (storage pool, data consistency, etc.).
[0087] An association relationship between the first alarm information is established based on the levels corresponding to the first alarm information and the categories corresponding to the first alarm information.
[0088] In some embodiments, as Figure 3 As shown, Figure 3 This diagram illustrates the relationships between first alarm messages provided in an embodiment of the present application. Each node in the diagram represents a first alarm message, and edges represent the relationships between these first alarm messages. Based on the hierarchical structure and mutual influence of alarms, a directed graph of alarm relationships is constructed. The lowest level is typically hardware status monitoring, with software components, clusters of similar components, and finally the customer service level.
[0089] In some embodiments, when a new first alarm message is received, all upstream and downstream alarms that may be affected are searched in the alarm relationship directed graph according to its level and type. If it is found that the related alarms have been triggered, the correlation of these alarms is strengthened; if they are not triggered, an early warning mechanism is set up to pay attention to possible problems in advance.
[0090] In some embodiments, before managing the first alarm information based on the time window and the alarm status of the first alarm information, the alarm management method further includes:
[0091] In response to the presence of low-level alarm information in the distributed storage system, first alarm information is stored, where the low-level alarm information includes at least one of a minor abnormality, a general abnormality, and a serious abnormality.
[0092] In some embodiments, minor exceptions are used to indicate occasional exceptions that can be recovered in a short time, such as network freezes, abnormal restart of a single component, etc.; general exceptions are used to indicate exceptions that do not affect system use, such as abnormal inability to restart a single component, etc.; serious exceptions are used to indicate exceptions that the cluster is close to being unusable, such as insufficient capacity, performance degradation, etc.; emergency exceptions are used to indicate that the system is unusable, such as complete downtime of cluster components, system inability to read and write, etc.
[0093] In some embodiments, as Figure 4 As shown, Figure 4 A flowchart of a delayed alarm method provided in an embodiment of the present application is provided. When a distributed storage system receives an alarm, it first determines whether it is a high-level alarm message. If it is a low-level alarm message, the alarm message is stored. If it is a high-level alarm message, it continues to determine whether there is a low-level delayed alarm message. If so, an alarm is triggered. If not, the alarm message is stored. It continues to determine whether the alarm has not been canceled due to timeout. If so, an alarm is triggered. If not, no alarm is triggered. Specifically, when the alarm is canceled within the time window, the alarm is no longer reported, reducing unnecessary alarm messages. For abnormalities with lower risks or alarm messages that may be false alarms, the alarm management module temporarily stores the alarm information and does not report it. A time window is used for filtering. If the alarm is canceled within a short period of time, the alarm management module no longer reports it. It is worth noting that low-level alarms may cause high-level alarms in the alarm relationship directed graph. Therefore, when ensuring that the low-level alarm is not reported, the high-level alarm generated by it must be blocked at the same time. If the low-level alarm is not canceled within the time window, the low-level alarm and the high-level alarm are reported simultaneously to ensure that no alarm information is missed.
[0094] In some embodiments, it should be noted that the low-level alarm information in this application refers to relatively low-level alarm information, rather than the lowest-level alarm information.
[0095] In some embodiments, the alarm management method further comprises:
[0096] Obtain target IDs, which are used to indicate target personnel of different roles;
[0097] In some embodiments, the target identifier can be any form of identity or communication identifier, such as a role name, user identifier, email address, mobile phone number, etc.
[0098] Based on the target identifier, target alarm information corresponding to the target identifier is determined.
[0099] In some embodiments, the target alarm information corresponding to the target identifier is determined based on the role responsibilities represented by the target identifier. For technical experts, more technical details and potential complexity analysis can be provided; for management or non-technical personnel, the technical details can be simplified, emphasizing the business impact and expected resolution time. Specifically, for on-site operations and maintenance personnel, they must always pay attention to the system operation status and require all alarm information; while for account managers, they are not sensitive to short-term alarms and alarms that do not affect usage. To reduce alarm notifications, only serious alarms can be reported to prevent customers from receiving too many low-level alarm information.
[0100] In some embodiments, by determining target alarm information corresponding to the target identifier based on the target identifier, the alarm information can be customized according to people with different roles, thereby reducing alarm notifications.
[0101] Through the present application, the first alarm information in the distributed storage system is obtained; the first alarm information is graded according to the first preset rule to obtain the level corresponding to the first alarm information; based on the level corresponding to the first alarm information, the time window corresponding to the first alarm information is determined; based on the time window and the alarm status of the first alarm information, the first alarm information is managed, and the alarm status includes at least one of a cancellation state and a continuous state. This solves the technical problems in related schemes that a large amount of unnecessary alarm information is easily generated, resulting in real emergency situations being ignored, and it is difficult to identify potential alarm information, and achieves the technical effect of filtering unnecessary alarm information and determining potential alarm information at the same time.
[0102] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0103] The embodiment of the present application further provides an alarm management device 500, Figure 5 A schematic diagram of the structure of an alarm management device provided by an embodiment of the present disclosure is shown in FIG. Figure 5 As shown, including:
[0104] An acquiring unit 501 is configured to acquire first alarm information in a distributed storage system;
[0105] The grading unit 502 is configured to perform grading processing on the first alarm information according to a first preset rule to obtain a level corresponding to the first alarm information;
[0106] A determining unit 503 is configured to determine a time window corresponding to the first alarm information based on the level corresponding to the first alarm information;
[0107] The alarm management unit 504 is configured to manage the first alarm information based on the time window and the alarm status of the first alarm information, where the alarm status includes at least one of a canceled state and a persistent state.
[0108] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 501 is configured to:
[0109] Get the hardware status of the distributed storage system;
[0110] Based on the hardware status, first alarm information corresponding to the hardware status is determined.
[0111] Furthermore, in a possible implementation of the embodiment of the present disclosure, the alarm management unit 504 is configured to:
[0112] In response to the alarm state of the first alarm information being a canceled state within the time window, storing the first alarm information and stopping reporting the first alarm information;
[0113] In response to the alarm status of the first alarm information being a persistent state within the time window, second alarm information is determined based on the level corresponding to the first alarm information and reported.
[0114] Furthermore, in a possible implementation of the embodiment of the present disclosure, the alarm management unit 504 is configured to:
[0115] Establishing an association relationship between the first alarm information based on the levels corresponding to the first alarm information;
[0116] Based on the association relationship, the second alarm information is determined and reported.
[0117] Furthermore, in a possible implementation of the embodiment of the present disclosure, the alarm management unit 504 is configured to:
[0118] Classify the first alarm information according to a second preset rule to obtain a category corresponding to the first alarm information;
[0119] An association relationship between the first alarm information is established based on the levels corresponding to the first alarm information and the categories corresponding to the first alarm information.
[0120] Furthermore, in a possible implementation of the embodiment of the present disclosure, the alarm management device 500 further includes a storage unit, which is configured to:
[0121] In response to the presence of low-level alarm information in the distributed storage system, first alarm information is stored, where the low-level alarm information includes at least one of a minor abnormality, a general abnormality, and a serious abnormality.
[0122] Furthermore, in a possible implementation of the embodiment of the present disclosure, the alarm management device 500 further includes a target alarm information determination unit, which is configured to:
[0123] Obtain target IDs, which are used to indicate target personnel of different roles;
[0124] Based on the target identifier, target alarm information corresponding to the target identifier is determined.
[0125] Through the present application, the first alarm information in the distributed storage system is obtained; the first alarm information is graded according to the first preset rule to obtain the level corresponding to the first alarm information; based on the level corresponding to the first alarm information, the time window corresponding to the first alarm information is determined; based on the time window and the alarm status of the first alarm information, the first alarm information is managed, and the alarm status includes at least one of a cancellation state and a continuous state. This solves the technical problems in related schemes that a large amount of unnecessary alarm information is easily generated, resulting in real emergency situations being ignored, and it is difficult to identify potential alarm information, and achieves the technical effect of filtering unnecessary alarm information and determining potential alarm information at the same time.
[0126] For the description of the features in the embodiment corresponding to the alarm management device, please refer to the relevant description of the embodiment corresponding to the alarm management method, and no further details will be given here.
[0127] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned alarm management method embodiments.
[0128] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned alarm management method embodiments when run.
[0129] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0130] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned alarm management method embodiments are implemented.
[0131] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned alarm management method embodiments.
[0132] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0133] The above is a detailed introduction to the alarm management method, device, electronic device and medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. An alarm management method, characterized in that: include: Obtaining first alarm information in the distributed storage system; Performing hierarchical processing on the first alarm information according to a first preset rule to obtain a level corresponding to the first alarm information; determining, based on the level corresponding to the first alarm information, a time window corresponding to the first alarm information; The first alarm information is managed based on the time window and an alarm state of the first alarm information, where the alarm state includes at least one of a cancel state and a persistent state.
2. The alarm management method according to claim 1, characterized in that: The obtaining of first alarm information in the distributed storage system includes: Get the hardware status of the distributed storage system; Based on the hardware status, first alarm information corresponding to the hardware status is determined.
3. The alarm management method according to claim 1, characterized in that: The managing the first alarm information based on the time window and the alarm status of the first alarm information includes: In response to the alarm state of the first alarm information being a canceled state within the time window, storing the first alarm information and stopping reporting the first alarm information; In response to the alarm status of the first alarm information being a persistent state within the time window, second alarm information is determined based on the level corresponding to the first alarm information and reported.
4. The alarm management method according to claim 3, characterized in that: The determining, based on the level corresponding to the first alarm information, second alarm information and reporting the second alarm information includes: establishing an association relationship between the first alarm information based on the levels corresponding to the first alarm information; Based on the association relationship, second alarm information is determined and reported.
5. The alarm management method according to claim 4, characterized in that: The establishing of the association relationship between the first alarm information based on the levels corresponding to the first alarm information includes: Classify the first alarm information according to a second preset rule to obtain a category corresponding to the first alarm information; An association relationship between the first alarm information is established based on the level corresponding to the first alarm information and the category corresponding to the first alarm information.
6. The alarm management method according to claim 1, characterized in that: Before managing the first alarm information based on the time window and the alarm status of the first alarm information, the method further includes: In response to the presence of low-level alarm information in the distributed storage system, the first alarm information is stored, where the low-level alarm information includes at least one of a minor abnormality, a general abnormality, and a serious abnormality.
7. The alarm management method according to claim 1, characterized in that: The method further comprises: Obtaining a target identifier, wherein the target identifier is used to indicate target personnel of different roles; Based on the target identifier, target alarm information corresponding to the target identifier is determined.
8. An alarm management device, characterized in that: include: An acquiring unit, configured to acquire first alarm information in a distributed storage system; a grading unit, configured to perform grading processing on the first alarm information according to a first preset rule to obtain a level corresponding to the first alarm information; a determining unit, configured to determine a time window corresponding to the first alarm information based on a level corresponding to the first alarm information; An alarm management unit is used to manage the first alarm information based on the time window and the alarm status of the first alarm information, where the alarm status includes at least one of a cancel state and a continuous state.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the alarm management method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the alarm management method according to any one of claims 1 to 7.