Cross-level storage device alarm method and device, equipment and storage medium
By introducing a master-slave node architecture into the storage array, dynamic threshold adjustment is carried out in combination with performance data and impact weights, the cross-level alarm system false alarm system errors and omissions of storage array devices are solved, and efficient and accurate alarm and fault management are achieved.
Patent Information
- Application Number
- CN202510637695.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, the threshold alarm system of the storage array device lacks cross-level correlation analysis, resulting in false alarms, missed alarms and inability to adapt to dynamic changes.
By introducing the master and slave architecture into the storage array, the master node synchronizes performance data and alarm thresholds, and dynamically adjusts the slave nodes, generates new alarm thresholds in combination with performance data and impact weights, and performs distributed collaborative work to improve alarm accuracy and system reliability.
It realizes timely and accurate alarms of storage array equipment, reduces misjudgment and misjudgment, improves the system's processing efficiency and fault tolerance, and enhances the convenience of fault location and management.
Smart Images

Figure CN120452163A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a cross-level storage device alarm method, apparatus, device, and storage medium. Background Art
[0002] Driven by the convergence of next-generation information technologies such as mobile internet, the Internet of Things, cloud computing, big data, and artificial intelligence, global data generation is experiencing exponential growth. As data centers expand, the number of servers, storage arrays, switches, and other equipment is also increasing exponentially. Traditional manual inspection methods are unable to meet real-time monitoring needs. The diverse equipment in enterprise data centers, encompassing heterogeneous hardware and hybrid multi-cloud architectures, places even higher demands on device management. Furthermore, the stability and high availability of IT systems are the lifeblood of enterprise services. Downtime or inaccessibility of core systems can result in short-term losses of millions.
[0003] Therefore, the emergence of threshold alarm technology coincided with the rapid development of information technology and the simultaneous evolution of operations and maintenance requirements. Related threshold alarm technologies include static threshold alarms and dynamic threshold alarms. However, static threshold alarms have a high learning and usage threshold, lack flexibility, and are unable to adapt to dynamic changes in storage array device performance and usage, leading to problems such as false alarms and missed alarms. Dynamic threshold alarms require a large amount of historical data for training and learning. In the early stages of device use or when data volumes are insufficient, it may not be possible to accurately establish a dynamic threshold model. Summary of the Invention
[0004] The present application provides a cross-level storage device alarm method, apparatus, device, and storage medium to at least solve the problem of high false alarm rate in threshold alarm technology in related technologies.
[0005] The present application provides a cross-layer storage device alarm method, which is applied to a storage array. The storage array includes a master node and at least one slave node. The method is executed by any slave node and includes:
[0006] Acquire performance data synchronously acquired by the master node from each of the plurality of sampling devices, and an alarm threshold corresponding to the performance data, wherein the performance data is associated data corresponding to an alarm factor of the storage array;
[0007] Determining, based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor, an impact weight of the performance data on the storage array alarm factor;
[0008] Determine a new alarm threshold corresponding to the storage array alarm factor based on the pre-acquired original alarm threshold corresponding to the storage array alarm factor, the indicator data, the impact weight, and the alarm threshold corresponding to the performance data;
[0009] When it is determined that an alarm operation is to be generated for the storage array alarm factor based on the indicator data and the new alarm threshold, alarm information is generated and sent to the master node so that the master node can execute the alarm operation.
[0010] The present application provides a cross-layer storage device alarm method, which is applied to a storage array. The storage array includes a master node and at least one slave node. The method is executed by the master node and includes:
[0011] When it is determined that the field value of the cross-level threshold policy indication field is a preset threshold value, performance data and an alarm threshold value corresponding to the performance data are obtained from each of the plurality of sampling devices through a pre-built data acquisition model, where the performance data is associated data corresponding to an alarm factor of the storage array;
[0012] Synchronizing performance data and alarm thresholds corresponding to the performance data to each slave node, so that each slave node determines whether to generate an alarm message based on the performance data and the alarm thresholds, as well as storage performance data on the slave node side, wherein the storage performance data includes indicator data corresponding to storage array alarm factors and the original alarm thresholds;
[0013] When receiving alarm information fed back by any slave node, the alarm operation is performed according to the alarm information.
[0014] The present application also provides a cross-level storage device alarm device, including:
[0015] an acquisition module, configured to acquire performance data acquired by the master node from each of the plurality of sampling devices, and an alarm threshold corresponding to the performance data, wherein the performance data is associated data corresponding to an alarm factor of the storage array;
[0016] a processing module configured to determine, based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor, an impact weight of the performance data on the storage array alarm factor; and determine, based on a pre-acquired original alarm threshold corresponding to the storage array alarm factor, the indicator data, the impact weight, and the alarm threshold corresponding to the performance data, a new alarm threshold corresponding to the storage array alarm factor;
[0017] An alarm generation module, configured to generate alarm information when an alarm operation is determined to be generated for an alarm factor of the storage array based on the indicator data and the new alarm threshold;
[0018] The sending module is used to send alarm information to the master node so that the master node can perform alarm operations.
[0019] The present application also provides a cross-level storage device alarm device, including:
[0020] A determination module, configured to determine whether a field value of a cross-level threshold policy indication field is a preset threshold;
[0021] an acquisition module configured to, when the determination module determines that the field value of the cross-level threshold policy indication field is a preset threshold, acquire performance data and an alarm threshold corresponding to the performance data from each of the plurality of sampling devices using a pre-built data acquisition model, wherein the performance data is associated data corresponding to a storage array alarm factor;
[0022] a synchronization module, configured to synchronize performance data and alarm thresholds corresponding to the performance data to each slave node, so that each slave node determines whether to generate an alarm message based on the performance data and the alarm thresholds, as well as storage performance data on the slave node side, wherein the storage performance data includes indicator data corresponding to storage array alarm factors and the original alarm thresholds;
[0023] A receiving module is used to receive alarm information fed back by any slave node;
[0024] The alarm execution module is used to execute alarm operations according to the alarm information.
[0025] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned cross-level storage device alarm methods are implemented.
[0026] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned cross-level storage device alarm methods when executed by a processor.
[0027] Through this application, the original alarm threshold is dynamically adjusted by combining performance data, indicator data and impact weights to generate new alarm thresholds. This method can more accurately reflect the current actual status of the storage array, avoiding misjudgments or missed judgments caused by the inability of fixed thresholds to adapt to different working conditions. Reduce unnecessary alarm interference. Moreover, different performance data may have different degrees of influence on the alarm factor. By calculating the impact weight, the actual significance of each performance data can be more accurately evaluated, thereby improving the accuracy of the alarm. The slave node can obtain the performance data and alarm thresholds synchronized by the master node in real time, and dynamically adjust them based on these data and the indicator data stored in itself. This real-time response mechanism enables the storage array to quickly adapt to various environmental changes and changes in business needs. Ensure the timeliness and accuracy of alarms.
[0028] In a storage array, multiple slave nodes can simultaneously assess and process alarms. This distributed collaborative approach improves the system's overall processing capacity and reliability. Each slave node independently assesses alarms based on local data and data synchronized with the master node, feeding the results back to the master node. The master node then integrates information from all slave nodes to make the final alarm decision. This approach not only improves system processing efficiency but also enhances fault tolerance.
[0029] Furthermore, the alarm information generated by this method is based on a comprehensive consideration of multiple factors and contains richer information. When an alarm occurs, administrators can locate and analyze the fault based on the alarm information and related performance and indicator data, quickly identifying the root cause and taking appropriate remedial measures. This helps improve system maintainability and reduce troubleshooting time. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0031] Figure 1 A schematic diagram of a cross-level storage device alarm method flow provided in an embodiment of the present application;
[0032] Figure 2 A simplified schematic diagram illustrating the corresponding relationship between each storage array alarm factor and the corresponding performance data, as well as the impact weight corresponding to each performance data, provided in an embodiment of the present application;
[0033] Figure 3 A flowchart of another cross-level storage device alarm method provided in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of the structure of a cross-level storage device alarm device provided in an embodiment of the present application;
[0035] Figure 5 A schematic diagram of the structure of another cross-level storage device alarm device provided in an embodiment of the present application;
[0036] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0039] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0040] Driven by the convergence of next-generation information technologies such as mobile internet, the Internet of Things, cloud computing, big data, and artificial intelligence, global data generation is experiencing exponential growth. As data centers expand, the number of servers, storage arrays, switches, and other equipment is also increasing exponentially. Traditional manual inspection methods are unable to meet real-time monitoring needs. The diverse equipment in enterprise data centers, encompassing heterogeneous hardware and hybrid multi-cloud architectures, places even higher demands on device management. Furthermore, the stability and high availability of IT systems are the lifeblood of enterprise services. Downtime or inaccessibility of core systems can result in short-term losses of millions.
[0041] Therefore, the birth of threshold alarm technology is accompanied by the rapid development of information technology and the synchronous evolution of operation and maintenance needs.
[0042] Related threshold alarm technologies are primarily based on static rules, dynamic adjustments, and intelligent analysis, aiming to improve the accuracy and real-time nature of device alarms while reducing false positives and missed alerts. In storage arrays, threshold alarm technology is primarily used to monitor hardware health, performance metrics, capacity utilization, and data reliability to ensure stable storage system operation. Similar to the threshold alarms used in data centers, storage arrays also employ the following types of threshold alarm technologies to proactively identify problems and risks:
[0043] (1) Static threshold alarm: By presetting a fixed value as a threshold, an alarm is triggered when a certain monitoring indicator of the storage array device reaches or exceeds this threshold. For example, the threshold of storage capacity utilization is set to 80%. When the actual storage capacity usage reaches or exceeds 80%, the system will issue an alarm to the administrator. Static threshold alarm technology is simple and direct, easy to set up, and is more suitable for indicators with relatively stable performance and small fluctuations. It also does not require complex algorithms and large amounts of historical data for analysis, and consumes less system resources.
[0044] (2) Percentage-based threshold alarm: A percentage-based threshold is used to trigger an alarm when the percentage of a certain indicator reaches or exceeds the set percentage. For example, if the percentage threshold for disk read / write speed drop is set to 20%, an alarm will be issued when the actual disk read / write speed drops by more than 20% compared to the normal level.
[0045] (3) Multi-level threshold alarms: By setting different thresholds, each threshold corresponds to a different alarm level and processing method, thus targeting a wider range of usage scenarios. For example, for storage capacity utilization, three threshold levels are set: a warning threshold of 70%, a major alarm threshold of 85%, and an emergency alarm threshold of 95%. When the utilization reaches different threshold levels, the system will issue an alarm of the corresponding level.
[0046] (4) Dynamic threshold alarm: By analyzing and learning the historical monitoring data of storage array devices, a dynamic threshold model is established. The model automatically adjusts the threshold based on factors such as device usage and time. For example, based on the storage capacity growth trend and usage pattern of the storage device in the past week, the capacity change in the future is predicted and the capacity utilization alarm threshold is dynamically adjusted. Dynamic threshold alarms can better adapt to the dynamic changes of storage array devices, improve the accuracy and timeliness of alarms, reduce false alarms and missed alarms caused by improper static threshold settings, and reduce the administrator's operation and maintenance costs.
[0047] Generally speaking, the first three types are static threshold alarms, mainly because the threshold alarm setting details and alarm conditions are different; dynamic alarms, on the other hand, require collecting and saving historical monitoring data, and dynamically adjusting the threshold settings by learning the changing trends of historical monitoring data.
[0048] However, the above threshold alarm technology still has the following problems:
[0049] (1) Static threshold alarms: For storage array administrators, they need to understand the capabilities of the storage array and the range of various indicator values, which has a high learning and usage threshold. At the same time, static threshold alarms lack flexibility and cannot adapt to dynamic changes in storage array device performance and usage, which can easily lead to problems such as false alarms and missed alarms.
[0050] (2) Dynamic threshold alarms require a large amount of historical data for training and learning. In the early stages of equipment use or when the amount of data is insufficient, it may not be possible to accurately establish a dynamic threshold model.
[0051] More importantly, the above solution only focuses on storage array devices. However, in many cases, certain indicators reaching thresholds may also be caused by other situations such as servers or network devices. However, the threshold alarm systems of servers, network devices, and storage arrays in related technologies usually operate independently and lack cross-layer correlation analysis (for example, storage IO delays may be related to network congestion or server load), making it impossible to perform early risk identification and alarms from a macro perspective.
[0052] To solve the above problems, the embodiment of the present application provides a cross-level storage device alarm method, see Figure 1 As shown, the method is applied to a storage array, the storage array includes a master node and at least one slave node, the method is executed by any slave node, and the method includes the following method steps:
[0053] Step S101 : Acquire performance data synchronously acquired by a master node from each of a plurality of sampling devices, and an alarm threshold value corresponding to the performance data.
[0054] Specifically, the sampling devices include servers and network devices such as switches. The servers and switches establish communication connections with storage array devices. The storage array devices, i.e., each node, are used to store some storage performance data from the sampling devices, including but not limited to the storage array alarm factors mentioned in this application. In an optional example, the storage array alarm factors include, for example, IO latency, total input / output operations per second (IOPS), and other data.
[0055] The associated data corresponding to the storage array alarm factor, that is, performance data, includes, for example, server CPU utilization, server memory usage, and server IO queue depth corresponding to IO latency, and server CPU utilization, switch port throughput, and server disk IO corresponding to total storage IOPS.
[0056] In an optional example, the master node may synchronize data to the slave node in the form of the following fields, as detailed below.
[0057] Table 1
[0058] Field Name Field Introduction Event_id Unique event identifier Event_type Event Type Per_stats Performance data Timestamp Timestamp version Version number (for conflict detection)
[0059] Step S102 : determining the influence weight of the performance data on the storage array alarm factor based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor.
[0060] Specifically, because multiple sampling devices can be used, multiple sets of performance indicator data can be collected for the indicator data corresponding to the storage array alarm factor. Even when multiple sets of performance indicator data are included, as previously described, the weight of each performance indicator's impact on the storage array alarm factor can be determined through fitting.
[0061] Step S103 : determining a new alarm threshold corresponding to the storage array alarm factor based on the pre-acquired original alarm threshold corresponding to the storage array alarm factor, the indicator data, the impact weight, and the alarm threshold corresponding to the performance data.
[0062] Specifically, first, based on the alarm threshold corresponding to the performance data and the performance data, it is determined whether the performance data is abnormal, and a determination result is obtained.
[0063] Specifically, when the value of the performance data is greater than the alarm threshold corresponding to the performance data, the determination result is 1; or when the value of the performance data is less than or equal to the alarm threshold corresponding to the performance data, the determination result is 0.
[0064] Then, a new alarm threshold is determined based on the determination result, the impact weight, and the original alarm threshold.
[0065] In an optional example, a new alarm threshold is determined according to the determination result, the impact weight, and the original alarm threshold, which is specifically expressed by the following expression:
[0066]
[0067] Among them, T i_new is the new alarm threshold corresponding to the i-th storage array alarm factor, T i_old is the original alarm threshold corresponding to the i-th storage array alarm factor, β a is the impact weight of the ath performance indicator data on the ith storage array alarm factor, P a To determine the result, when the value of the performance data is greater than the alarm threshold corresponding to the performance data, P aWhen the value of the performance data is less than or equal to the alarm threshold corresponding to the performance data, P a is 0.
[0068] Step S104 : When it is determined that an alarm operation is to be generated for the storage array alarm factor based on the indicator data and the new alarm threshold, alarm information is generated and sent to the master node so that the master node can execute the alarm operation.
[0069] Among them, when the indicator data is greater than the new alarm threshold, an alarm operation is generated;
[0070] or,
[0071] When the indicator data is less than or equal to the new alarm threshold, the alarm operation is prohibited.
[0072] An embodiment of the present application provides a cross-level storage device alarm method, which dynamically adjusts the original alarm threshold by combining performance data, indicator data and impact weights to generate a new alarm threshold. This method can more accurately reflect the current actual status of the storage array, avoiding misjudgments or missed judgments caused by the inability of fixed thresholds to adapt to different working conditions. Reduce unnecessary alarm interference. Moreover, different performance data may have different degrees of influence on the alarm factor. By calculating the impact weight, the actual significance of each performance data can be more accurately evaluated, thereby improving the accuracy of the alarm. The slave node can obtain the performance data and alarm threshold synchronized by the master node in real time, and dynamically adjust according to these data and the indicator data stored by itself. This real-time response mechanism enables the storage array to quickly adapt to various environmental changes and changes in business needs. Ensure the timeliness and accuracy of the alarm.
[0073] In a storage array, multiple slave nodes can simultaneously assess and process alarms. This distributed collaborative approach improves the system's overall processing capacity and reliability. Each slave node independently assesses alarms based on local data and data synchronized with the master node, feeding the results back to the master node. The master node then integrates information from all slave nodes to make the final alarm decision. This approach not only improves system processing efficiency but also enhances fault tolerance.
[0074] Furthermore, the alarm information generated by this method is based on a comprehensive consideration of multiple factors and contains richer information. When an alarm occurs, administrators can locate and analyze the fault based on the alarm information and related performance and indicator data, quickly identifying the root cause and taking appropriate remedial measures. This helps improve system maintainability and reduce troubleshooting time.
[0075] In an optional example, based on the above embodiment, when the performance data includes multiple types, the multiple types of performance data and the indicator data corresponding to the storage array alarm factor are fitted to determine the impact weight of each type of performance data on the storage array alarm factor, which is expressed by the following expression:
[0076] Y i =β 1i X 1i +β 2i X 2i +…+β ki X ki (Formula 2)
[0077] Among them, Y i is the indicator data corresponding to the i-th storage array alarm factor, β ki is the influence weight of the kth performance data on the ith storage array alarm factor, X ki is the kth performance data associated with the i-th storage array alarm factor, where β 1i +β 2i +…+β ki =1, i is a positive integer.
[0078] Taking the aforementioned storage array alarm factor and its corresponding performance data as an example, the influence weight of each performance data finally obtained on the storage array alarm factor is shown in FIG. Figure 2 As shown, Figure 2 The meaning of each field has been explained in the previous article, so I will not repeat it here.
[0079] In an optional example, based on the foregoing embodiment, before determining the weight of the impact of the performance data on the storage array alarm factor based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor, the method may further include the following method steps:
[0080] A threshold alarm enable signal corresponding to the storage array alarm factor synchronized with the master node is obtained, wherein the threshold alarm enable signal is used to indicate that the slave node has permission to monitor the storage array alarm factor.
[0081] Specifically, assuming that there are multiple storage array alarm factors, and the master node sends the threshold alarm enable signal corresponding to each storage array alarm factor to the slave node, the slave node determines that it has permission to monitor a certain storage array alarm factor and can then determine whether to enable that permission based on actual circumstances. That is, only after determining that permission is enabled can the method steps 102 to 104 above be executed for a specific storage array alarm factor.
[0082] It should be noted that when a new cluster node (newly connected device) is added, the new device can reuse the performance data of existing devices of the same type, the correlation (causal relationship) between storage array alarm factors and performance data, and the aforementioned method steps, improving the accuracy of threshold alarms for the newly added device during the promotion phase. Furthermore, this method does not require training and learning from a large amount of historical data; the dynamic threshold adjustment model (equation 1) can be directly constructed using the aforementioned method. This method is simpler and more convenient, and the determination of dynamic thresholds is more reasonable and accurate.
[0083] The embodiment of the present invention also provides a cross-level storage device alarm method, see Figure 3 As shown, the method includes the following steps:
[0084] Step S301: When it is determined that the field value of the cross-level threshold policy indication field is a preset threshold, performance data and an alarm threshold corresponding to the performance data are obtained from each of the multiple sampling devices through a pre-built data acquisition model.
[0085] Specifically, staff can use the storage array management interface to set a cross-level threshold policy indication switch. This switch is associated with a cross-level threshold policy indication field, such as "SpanThreshold." When the staff turns on the switch, the value of the cross-level threshold policy indication field "SpanThreshold" changes from false to true. This indicates that the master node can use a pre-built data acquisition model to obtain performance data from each of multiple sampling devices, as well as the alarm thresholds corresponding to the performance data. The performance data is associated with the storage array alarm factors.
[0086] Step S302: Synchronize the performance data and the alarm threshold corresponding to the performance data to each slave node.
[0087] Specifically, the performance data has been described in detail in the above embodiments. A storage array alarm factor may correspond to one or more performance data. Different performance data may be obtained from different sampling devices, such as servers or switches.
[0088] Each type of performance data can be assigned an alarm threshold. After synchronizing the performance data and alarm thresholds to each slave node, the slave node can determine whether to generate an alarm based on the performance data and alarm thresholds, as well as the storage performance data on the slave node. Storage performance data includes metrics corresponding to storage array alarm factors and the original alarm thresholds. The specific process for determining whether to generate an alarm has been detailed previously, so I won't elaborate on it here.
[0089] Step S303: When receiving alarm information fed back by any slave node, executing an alarm operation according to the alarm information.
[0090] Further optionally, while the master node synchronizes the performance data and the alarm threshold corresponding to the performance data to each slave node, the method may also include synchronizing the alarm threshold configuration switch to the slave node at the same time. Similar to the cross-level threshold policy switch, the alarm threshold configuration switch can also be implemented in the form of a configuration field. For example, the configuration enable signal is used to indicate whether the alarm threshold configuration switch is turned on. It should be noted that for each storage array alarm factor, a corresponding alarm threshold configuration switch is configured, which is used for the slave node to control whether the switch is turned on according to its own situation, that is, to control whether to monitor the relevant performance data corresponding to the storage array alarm factor, and to evaluate whether to generate alarm information according to the aforementioned method steps.
[0091] In an optional embodiment, the pre-built data acquisition model is specifically described below. Based on the existing model of the open network management standard, taking the performance data corresponding to each storage array alarm factor mentioned above as an example, the following server and network device interface model extensions can be added:
[0092] For example, a) in an openconfig-system (such as a general server), a performance data acquisition model such as CPU utilization, memory usage, and disk I / O is added.
[0093] The model example is shown below:
[0094]
[0095]
[0096] b) In openconfig-telemetry (for example, general switches), add performance data acquisition models such as port throughput, packet loss rate, queue depth, and link status. The network port model is similar to the above and will not be repeated here.
[0097] An embodiment of the present application provides a cross-level storage device alarm method, in which a master node obtains performance data and corresponding alarm thresholds from multiple sampling devices through a pre-built data acquisition model, and synchronizes them to the slave nodes. This avoids the duplication of work and waste of resources that may be caused by the independent acquisition of data by each slave node, reduces the network bandwidth usage and data transmission delay. At the same time, the pre-built data acquisition model has been optimized and can efficiently collect the required data, thereby improving the data processing efficiency of the entire storage array. After the master node synchronizes the performance data and alarm thresholds to the slave node, the slave node can make independent alarm judgments based on its own storage performance data. This distributed processing method makes full use of the computing resources of the slave node, reduces the processing burden of the master node, enables the system to respond and process large amounts of performance data more quickly, and improves the overall processing capability of the system.
[0098] Furthermore, slave nodes determine in real time whether to generate alarms based on synchronized data and their own storage performance data, and promptly report these to the master node. Upon receiving the alarm, the master node promptly executes the alarm action, allowing administrators to promptly understand storage array anomalies and take appropriate action to address them, preventing further escalation of the problem and ensuring the security and availability of stored data.
[0099] Furthermore, the cross-level threshold policy indicator field enables the system to employ a unified policy to monitor storage array performance. The master node determines whether to obtain performance data and alarm thresholds based on the value of this policy indicator field and then synchronizes these data with the slave nodes. This approach simplifies system management, allowing administrators to monitor and manage the entire storage array through unified policies, improving management convenience and efficiency.
[0100] In an optional example, as described above, the master node receives alarm information sent by each slave node. However, in some cases, the alarm information fed back by each slave node may be similar. Moreover, similar alarm information may refer to the same thing. In order to avoid repeatedly executing alarm operations corresponding to similar alarm information in a short period of time, the method may further include the following method steps:
[0101] Step a1: perform semantic clustering on multiple alarm messages received continuously to obtain similarities between the multiple alarm messages;
[0102] Step a2: merging the alarm information with similarity higher than a preset threshold and generating a summary report;
[0103] Step a3: performing an alarm operation based on the summary report.
[0104] Through the aforementioned method, clustering technology can group similar alarms into a single category, reducing processing time and computing resource consumption. Grouping similar alarms into a single category can simplify subsequent information processing, such as when performing the same alarm operation multiple times. Furthermore, merging multiple alarms into summaries can reduce alarm redundancy, making the information clearer and easier to understand. Summary reports can quickly convey key information, helping operators quickly identify problems and take action. Merging alarm information reduces the number of alarms operators need to handle, reducing operational difficulty and workload. Furthermore, summary reports can reduce the risk of misoperation or overlooking important alarms due to information overload. They can also more efficiently allocate resources, improving overall system stability and reliability.
[0105] Further optionally, the above-mentioned method of clustering multiple alarms can also be performed on the slave node side. For example, after the slave node generates the first alarm information and reports it to the master node, it may continue to generate new similar alarm information before the master node executes the alarm operation or during the execution of the alarm operation. Then, at this time, the slave node can compare the currently received alarm information with the alarm that has been sent. If the similarity is higher than a preset threshold, such as more than 95%, it means that the feedback is still a similar problem. In this case, a silent period can be set, for example, no similar alarm information will be sent to the master node within 5 minutes, and the system will switch to background monitoring. This avoids the frequent sending of repeated alarm information to the master node, reduces resource usage, and avoids resource waste.
[0106] Of course, the above functions can be configured with corresponding virtual controls or fields on each node to indicate whether the above functions are turned on, that is, the staff can determine whether to execute the above functions according to the actual situation.
[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0108] The embodiment of the present application also provides a cross-level storage device alarm device, which is applied to a storage array. The storage array includes a master node and at least one slave node. The device corresponds to any slave node. Figure 4 As shown, the device includes: an acquisition module 401, a processing module 402, an alarm generation module 403, and a sending module 404.
[0109] An acquisition module 401 is configured to acquire performance data synchronously acquired by a master node from each of a plurality of sampling devices, and an alarm threshold corresponding to the performance data, wherein the performance data is associated data corresponding to an alarm factor of a storage array;
[0110] Processing module 402 is configured to determine, based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor, an impact weight of the performance data on the storage array alarm factor; and determine, based on the pre-acquired original alarm threshold corresponding to the storage array alarm factor, the indicator data, the impact weight, and the alarm threshold corresponding to the performance data, a new alarm threshold corresponding to the storage array alarm factor.
[0111] The alarm generation module 403 is configured to generate alarm information when determining, based on the indicator data and the new alarm threshold, to generate an alarm operation for the storage array alarm factor;
[0112] The sending module 404 is used to send the alarm information to the master node so that the master node can perform an alarm operation.
[0113] In an optional example, when the performance data includes multiple types, the processing module 402 is specifically used to fit the multiple types of performance data and indicator data corresponding to the storage array alarm factor to determine the impact weight of each performance data on the storage array alarm factor.
[0114] In an optional example, the processing module 402 fits the indicator data corresponding to the various types of performance data and the storage array alarm factors to determine the impact weight of each type of performance data on the storage array alarm factor, which is expressed as the following expression:
[0115] Y i =β 1i X 1i +β 2i X 2i +…+β ki X ki
[0116] Among them, Y i is the indicator data corresponding to the i-th storage array alarm factor, β ki is the influence weight of the kth performance data on the ith storage array alarm factor, X ki is the kth performance data associated with the i-th storage array alarm factor, where β 1i +β 2i +…+β ki =1, i is a positive integer.
[0117] In an optional example, the processing module 402 is specifically configured to determine whether the performance data is abnormal based on the alarm threshold corresponding to the performance data and the performance data, and obtain a determination result;
[0118] A new alarm threshold is determined based on the determination result, the impact weight, and the original alarm threshold.
[0119] In an optional example, the processing module 402 determines a new alarm threshold according to the determination result, the impact weight, and the original alarm threshold, which is specifically expressed by the following expression:
[0120]
[0121] Among them, T i_new is the new alarm threshold corresponding to the i-th storage array alarm factor, T i_old is the original alarm threshold corresponding to the i-th storage array alarm factor, β a is the impact weight of the ath performance indicator data on the ith storage array alarm factor, P a To determine the result, when the value of the performance data is greater than the alarm threshold corresponding to the performance data, P a When the value of the performance data is less than or equal to the alarm threshold corresponding to the performance data, P a is 0.
[0122] In an optional example, the alarm generation module 403 is specifically configured to generate an alarm operation when the indicator data is greater than a new alarm threshold;
[0123] Alternatively, when the indicator data is less than or equal to the new alarm threshold, no alarm action is generated.
[0124] In an optional example, the acquisition module 401 is further configured to acquire a threshold alarm enable signal corresponding to the storage array alarm factor synchronized by the master node, wherein the threshold alarm enable signal is configured to indicate that the slave node has permission to monitor the storage array alarm factor.
[0125] The description of the features of the embodiment corresponding to the cross-level storage device alarm device provided in the embodiment of the present application can be found in Figure 1 The related descriptions of the illustrated embodiment and related embodiments will not be repeated here one by one.
[0126] An embodiment of the present application provides a cross-level storage device alarm device, which dynamically adjusts the original alarm threshold by combining performance data, indicator data and impact weights to generate a new alarm threshold. This method can more accurately reflect the current actual status of the storage array, avoiding misjudgments or missed judgments caused by the inability of fixed thresholds to adapt to different working conditions. Reduce unnecessary alarm interference. Moreover, different performance data may have different degrees of influence on the alarm factor. By calculating the impact weight, the actual significance of each performance data can be more accurately evaluated, thereby improving the accuracy of the alarm. The slave node can obtain the performance data and alarm threshold synchronized by the master node in real time, and dynamically adjust according to these data and the indicator data stored by itself. This real-time response mechanism enables the storage array to quickly adapt to various environmental changes and changes in business needs. Ensure the timeliness and accuracy of the alarm.
[0127] In a storage array, multiple slave nodes can simultaneously assess and process alarms. This distributed collaborative approach improves the system's overall processing capacity and reliability. Each slave node independently assesses alarms based on local data and data synchronized with the master node, feeding the results back to the master node. The master node then integrates information from all slave nodes to make the final alarm decision. This approach not only improves system processing efficiency but also enhances fault tolerance.
[0128] Furthermore, the alarm information generated by this method is based on a comprehensive consideration of multiple factors and contains richer information. When an alarm occurs, administrators can locate and analyze the fault based on the alarm information and related performance and indicator data, quickly identifying the root cause and taking appropriate remedial measures. This helps improve system maintainability and reduce troubleshooting time.
[0129] The embodiment of the present application also provides a cross-level storage device alarm device, which is applied to a storage array. The storage array includes a master node and at least one slave node. The device corresponds to the master node. Figure 5 As shown, the device includes: a determination module 501, an acquisition module 502, a synchronization module 503, a receiving module 504, and an alarm execution module 505.
[0130] Determining module 501, used to determine whether the field value of the cross-level threshold policy indication field is a preset threshold;
[0131] an acquisition module 502 configured to, when the determination module 501 determines that the field value of the cross-level threshold policy indication field is a preset threshold, acquire performance data and an alarm threshold corresponding to the performance data from each of the plurality of sampling devices using a pre-built data acquisition model, where the performance data is associated data corresponding to a storage array alarm factor;
[0132] Synchronization module 503, configured to synchronize performance data and alarm thresholds corresponding to the performance data to each slave node, so that each slave node determines whether to generate an alarm based on the performance data and alarm thresholds, as well as storage performance data on the slave node side, wherein the storage performance data includes indicator data corresponding to storage array alarm factors and the original alarm thresholds;
[0133] Receiving module 504, configured to receive alarm information fed back by any slave node;
[0134] The alarm execution module 505 is used to execute an alarm operation according to the alarm information.
[0135] The description of the features of the embodiment corresponding to the cross-level storage device alarm device provided in the embodiment of the present application can be found in Figure 2 The relevant descriptions of the corresponding embodiments will not be repeated here one by one.
[0136] An embodiment of the present application provides a cross-level storage device alarm device, in which the master node obtains performance data and corresponding alarm thresholds from multiple sampling devices through a pre-built data acquisition model, and synchronizes them to the slave nodes. This avoids the duplication of work and waste of resources that may be caused by the independent acquisition of data by each slave node, reduces the network bandwidth usage and data transmission delay. At the same time, the pre-built data acquisition model has been optimized and can efficiently collect the required data, thereby improving the data processing efficiency of the entire storage array. After the master node synchronizes the performance data and alarm thresholds to the slave node, the slave node can make independent alarm judgments based on its own storage performance data. This distributed processing method makes full use of the computing resources of the slave node, reduces the processing burden of the master node, enables the system to respond and process large amounts of performance data more quickly, and improves the overall processing capability of the system.
[0137] Furthermore, slave nodes determine in real time whether to generate alarms based on synchronized data and their own storage performance data, and promptly report these to the master node. Upon receiving the alarm, the master node promptly executes the alarm action, allowing administrators to promptly understand storage array anomalies and take appropriate action to address them, preventing further escalation of the problem and ensuring the security and availability of stored data.
[0138] Furthermore, the cross-level threshold policy indicator field enables the system to employ a unified policy to monitor storage array performance. The master node determines whether to obtain performance data and alarm thresholds based on the value of this policy indicator field and then synchronizes these data with the slave nodes. This approach simplifies system management, allowing administrators to monitor and manage the entire storage array through unified policies, improving management convenience and efficiency.
[0139] The embodiment of the present application also provides an electronic device, such as Figure 6 As shown, it includes a memory 10 and a processor 20, wherein the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above-mentioned cross-level storage device alarm method embodiments.
[0140] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned cross-level storage device alarm method embodiments, or execute the steps of any of the above-mentioned data reading method embodiments when running.
[0141] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0142] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned cross-level storage device alarm method embodiments.
[0143] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data reading method embodiments are implemented.
[0144] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0145] The above is a detailed introduction to a cross-level storage device alarm method, apparatus, equipment and storage medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A cross-level storage device alarm method, characterized in that: The method is applied to a storage array, the storage array including a master node and at least one slave node. The method is executed by any slave node, and the method includes: Acquire performance data synchronously acquired by the master node from each of the plurality of sampling devices, and an alarm threshold value corresponding to the performance data, wherein the performance data is associated data corresponding to an alarm factor of the storage array; Determining, based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor, an impact weight of the performance data on the storage array alarm factor; Determining a new alarm threshold corresponding to the storage array alarm factor based on the pre-acquired original alarm threshold corresponding to the storage array alarm factor, the indicator data, the impact weight, and the alarm threshold corresponding to the performance data; When it is determined based on the indicator data and the new alarm threshold that an alarm operation is to be generated for the storage array alarm factor, alarm information is generated and sent to the master node so that the master node performs the alarm operation.
2. The method according to claim 1, characterized in that When the performance data includes multiple types, determining the influence weight of the performance data on the storage array alarm factor based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor specifically includes: Fitting is performed on multiple types of performance data and indicator data corresponding to the storage array alarm factor to determine the influence weight of each type of performance data on the storage array alarm factor.
3. The method according to claim 2, characterized in that The multiple types of performance data and the indicator data corresponding to the storage array alarm factor are fitted to determine the influence weight of each type of performance data on the storage array alarm factor, which is expressed by the following expression: Y i =b 1i X 1i +b 2i X 2i +…+b ki X ki Among them, Y i is the index data corresponding to the i-th storage array alarm factor, β ki is the influence weight of the kth performance data on the ith storage array alarm factor, X ki is the kth performance data associated with the i-th storage array alarm factor, wherein β 1i +β 2i +…+β ki =1, i is a positive integer.
4. The method according to any one of claims 1 to 3, characterized in that Determining a new alarm threshold corresponding to the storage array alarm factor based on the pre-acquired original alarm threshold corresponding to the storage array alarm factor, the indicator data, the impact weight, and the alarm threshold corresponding to the performance data specifically includes: Determining whether the performance data is abnormal based on an alarm threshold corresponding to the performance data and the performance data, and obtaining a determination result; The new alarm threshold is determined according to the determination result, the impact weight, and the original alarm threshold.
5. The method according to claim 4, characterized in that The new alarm threshold is determined according to the determination result, the impact weight, and the original alarm threshold, which is specifically expressed by the following expression: Among them, T i_new is the new alarm threshold corresponding to the i-th storage array alarm factor, T i_old is the original alarm threshold corresponding to the i-th storage array alarm factor, β a is the influence weight of the ath performance indicator data on the ith storage array alarm factor, P a is the determination result, wherein, when the value of the performance data is greater than the alarm threshold corresponding to the performance data, the P a is 1, or when the value of the performance data is less than or equal to the alarm threshold corresponding to the performance data, the P a is 0.
6. The method according to claim 5, characterized in that Determining whether to generate an alarm operation for the storage array alarm factor based on the indicator data and the new alarm threshold specifically includes: When the indicator data is greater than the new alarm threshold, generating the alarm operation; Alternatively, when the indicator data is less than or equal to the new alarm threshold, the alarm operation is not generated.
7. The method according to any one of claims 1 to 3, characterized in that Before determining the influence weight of the performance data on the storage array alarm factor based on the performance data and pre-stored indicator data corresponding to the storage array alarm factor, the method further includes: A threshold alarm enable signal corresponding to the storage array alarm factor and synchronized with the master node is obtained, wherein the threshold alarm enable signal is used to indicate that the slave node has permission to monitor the storage array alarm factor.
8. A cross-level storage device alarm method, characterized in that: The method is applied to a storage array, the storage array including a master node and at least one slave node, and the method is executed by the master node, and the method includes: When it is determined that the field value of the cross-level threshold policy indication field is a preset threshold, performance data and an alarm threshold corresponding to the performance data are acquired from each of the plurality of sampling devices using a pre-built data acquisition model, the performance data being associated data corresponding to a storage array alarm factor; Synchronizing the performance data and an alarm threshold corresponding to the performance data to each of the slave nodes, so that each of the slave nodes determines whether to generate an alarm message based on the performance data and the alarm threshold, as well as storage performance data on the slave node side, wherein the storage performance data includes indicator data corresponding to the storage array alarm factor and the original alarm threshold; When receiving the alarm information fed back by any slave node, an alarm operation is performed according to the alarm information.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is configured to implement the steps of the cross-level storage device alarm method according to any one of claims 1 to 7 when executing the computer program, or implement the steps of the cross-level storage device alarm method according to claim 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the cross-level storage device alarm method as described in any one of claims 1 to 7, or when the computer program is executed by a processor, the computer program implements the steps of the cross-level storage device alarm method as described in claim 8.