Storage system anomaly detection methods, electronic devices, storage media, and software products

By collecting and correlating credit occupancy, release, and cache pressure data in the Fibre Channel storage area network, global detection of storage system anomalies is achieved, solving the problems of low detection reliability and correlation in existing technologies, and enabling precise location of faults.

CN120723519BActive Publication Date: 2025-11-14INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511188805.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-14
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

In existing technologies, anomaly detection in Fibre Channel storage area networks is usually performed in a single device, resulting in poor reliability and low correlation of anomaly detection.

Method used

Anomaly detection of the storage system is performed from a global perspective of the communication link. By collecting credit occupancy, credit release, and cache pressure data from host devices, switch devices, and storage devices, time alignment and link correlation are performed to detect related anomalies and locate the fault location.

Benefits of technology

It improves the reliability and correlation of anomaly detection in storage systems, enabling precise location of faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723519B_ABST
    Figure CN120723519B_ABST
Patent Text Reader

Abstract

This invention provides a storage system anomaly detection method, electronic device, storage medium, and program product, relating to the field of storage systems. It can collect credit occupancy data from host devices, credit release data from switch devices, and cache pressure data from storage devices, enabling monitoring of credit occupancy and release status from the host device, switch device, and storage device respectively. Subsequently, this invention can time-align the credit occupancy data, credit release data, and cache pressure data according to the collection time, and link-associate them according to the communication links formed by the host device, switch device, and storage device to obtain associated operational data. Anomaly detection is then performed on the associated operational data, and based on the anomaly detection results, the fault location in the host device, switch device, and storage device can be determined. This allows for the detection of anomalies in the storage system from a global link perspective, thereby improving the reliability and correlation of storage system anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage systems, and more particularly to methods for detecting anomalies in storage systems, electronic devices, storage media, and software products. Background Technology

[0002] Fibre Channel Storage Area Network (FC SAN) is a common form of storage system. In related technologies, anomaly detection for FC SANs is typically performed on a single device, resulting in poor reliability and low correlation in anomaly detection. Summary of the Invention

[0003] This invention provides a storage system anomaly detection method, electronic device, storage medium, and program product, which can detect storage system anomalies from a global perspective of the communication link and can pinpoint the fault location in detail, thereby improving the reliability and correlation of storage system anomaly detection.

[0004] To address the aforementioned technical problems, this invention provides a storage system anomaly detection method, comprising:

[0005] Credit occupancy data is collected from host devices, credit release data is collected from switch devices, and cache pressure data is collected from storage devices. Specifically, host devices occupy credit provided by storage devices by issuing read and write requests to storage devices, and storage devices release the credit occupied by host devices to host devices through switch devices when they complete the read and write request processing.

[0006] Credit occupancy data, credit release data, and cache pressure data are time-aligned according to the collection time, and linked according to the communication links formed by host devices, switch devices, and storage devices to obtain associated operational data.

[0007] Perform correlation anomaly detection on associated operational data, and determine the fault location in host devices, switch devices, and storage devices based on the anomaly detection results.

[0008] The present invention also provides an electronic device, comprising:

[0009] Memory, used to store computer programs;

[0010] The processor is used to implement the aforementioned storage system anomaly detection method when executing computer programs.

[0011] The present invention also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the above-described storage system anomaly detection method.

[0012] The present invention also provides a non-volatile computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the above-described storage system anomaly detection method.

[0013] The beneficial effects of this invention are as follows: First, it can collect credit occupancy data from host devices, credit release data from switch devices, and cache pressure data from storage devices. Specifically, the host device occupies credit provided by the storage device by issuing read / write requests, and the storage device releases the credit occupied by the host device to the host device via the switch device after completing the read / write request processing. That is, this invention can monitor the credit occupancy and release status of the host device, switch device, and storage device respectively. Subsequently, this invention can time-align the credit occupancy data, credit release data, and cache pressure data according to the collection time, and link them according to the communication links formed by the host device, switch device, and storage device to obtain associated operational data, thereby associating the working status of the host device, switch device, and storage device. Finally, this invention can perform associated anomaly detection on the associated operational data, and determine the fault location in the host device, switch device, and storage device based on the anomaly detection results. This allows for the detection of anomalies in the storage system from a global perspective of the communication links, and enables fine-grained fault location, thereby improving the reliability and correlation of anomaly detection in the storage system.

[0014] The present invention also provides an electronic device, a non-volatile computer-readable storage medium, and a computer program product, which have the above-mentioned beneficial effects. Attached Figure Description

[0015] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart of a storage system anomaly detection method provided in an embodiment of the present invention;

[0017] Figure 2 A structural block diagram of a storage system provided in an embodiment of the present invention;

[0018] Figure 3 A schematic diagram illustrating the workflow of a storage system provided in an embodiment of the present invention;

[0019] Figure 4 A schematic diagram illustrating the workflow of a host agent program provided in an embodiment of the present invention;

[0020] Figure 5 A schematic diagram illustrating the workflow of a switch probe program according to an embodiment of the present invention;

[0021] Figure 6 A schematic diagram illustrating the workflow of a storage agent program according to an embodiment of the present invention;

[0022] Figure 7 A schematic diagram of a communication link association provided in an embodiment of the present invention;

[0023] Figure 8 A structural block diagram of the data analysis system provided in an embodiment of the present invention;

[0024] Figure 9 This is a structural block diagram of a storage system anomaly detection device provided in an embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0026] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0027] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Fiber Channel Storage Area Networks (FC SANs) are a common form of storage system. In related technologies, anomaly detection in FC SANs is typically performed on a single device, resulting in poor reliability and low correlation. Therefore, this invention provides a storage system anomaly detection method that can correlate data generated by various devices in the storage system, detecting anomalies from a global communication link perspective, thereby improving the reliability and correlation of storage system anomaly detection.

[0029] It should be noted that this method can be executed by a standalone data analysis device, which is independent of the storage system and can establish communication connections with host devices, switch devices, and storage devices within the storage system. This embodiment does not limit the specific form of the data analysis device; it can be configured according to actual application requirements, such as a server, embedded device, or personal computer.

[0030] For easier understanding, please refer to Figure 1 , Figure 1 A flowchart illustrating a storage system anomaly detection method provided in an embodiment of the present invention. This method may include:

[0031] S101. Collect credit occupancy data from the host device, credit release data from the switch device, and cache pressure data from the storage device; wherein, the host device occupies the credit provided by the storage device by issuing read and write requests to the storage device, and the storage device releases the credit occupied by the host device to the host device through the switch device when it completes the read and write request processing.

[0032] To facilitate understanding, the operating modes of the storage systems applicable to the embodiments of the present invention will first be introduced. Please refer to... Figure 2 , Figure 2 This is a structural block diagram of a storage system provided in an embodiment of the present invention. The storage system may include a host device, a switch device, and a storage device. The host device is directly connected to the switch device, and the switch device is directly connected to the storage device. These three devices can communicate with each other via Fibre Channel (FC). The host device issues read / write requests to the storage device. The storage device is responsible for processing the read / write requests and feeding back the processing results to the host device. The switch device is responsible for transmitting data frames between the host device and the storage device. Furthermore, to facilitate the orderly processing of read / write requests by the storage device, a buffer queue can be set up in the storage device to cache read / write requests.

[0033] In this system architecture, to prevent the host device from sending excessive read / write requests to the storage device, a credit value mechanism can be established between the host device and the storage device. In this mechanism, the storage device can pre-set multiple credit values ​​and inform the host device of the number of credit values. Each time the host device sends a read / write request to the storage device, it consumes one credit value provided by the storage device; conversely, each time the storage device completes a read / write request, it releases one credit value back to the host device. Furthermore, when the host device determines that its credit values ​​are exhausted, it will no longer be able to send read / write requests to the storage device.

[0034] For a deeper understanding, please refer to Figure 3 , Figure 3 This invention provides a workflow for a storage system. First, a host device sends a data frame (FC frame) containing a read / write request to a switch device. The switch device forwards this data frame to a storage device to complete the allocation of one credit value. After processing the read / write request, the storage device releases the credit value and sends a credit value release frame (REC_F) to the switch device, which forwards the credit value release frame to the host device. Upon receiving the credit value release frame, the host device determines that the credit value has been released, thereby restoring its credit value quota. In short, the host device can allocate credit values, the storage device can release the credit values ​​allocated by the host device, and the switch is responsible for transmitting the credit value release information.

[0035] Clearly, the flow of credit points among host devices, switches, and storage devices affects the performance of the storage system. Furthermore, malfunctions in any of these devices can impact credit point flow. For example, a host device malfunction might send excessive read / write requests to the storage device, leading to rapid credit point depletion. A switch malfunction might decrease the forwarding rate of read / write requests or credit point release frames, thus affecting credit point flow efficiency. Similarly, a storage device malfunction might reduce the processing efficiency of read / write requests, resulting in reduced credit point release efficiency. Moreover, malfunctions in these devices can occur in combination. Therefore, to effectively locate storage system faults, this embodiment collects credit point usage data from host devices, credit point release data from switches, and cache pressure data from storage devices. This data is then linked by collection time and communication link to form associated operational data. Anomaly detection of the storage system is performed based on this associated operational data. This comprehensive approach to anomaly detection, considering the performance of each device, improves the reliability and relevance of the detection.

[0036] The following sections will introduce the specific content and collection methods of credit occupancy data, credit release data, and cache pressure data.

[0037] In one implementation, collecting credit occupancy data from the host device may include:

[0038] Step 11: At the current acquisition time, access the host bus adapter of the host device and read the credit usage and credit remaining in the register of the host bus adapter;

[0039] Step 12: When it is determined that the credit usage at the current collection time has changed compared to the historical credit usage at the previous collection time, determine the change in credit usage at the current collection time based on the credit usage and the historical credit usage.

[0040] Step 13: Set the credit occupancy change and credit remaining amount as credit occupancy data, and mark the first timestamp of the current collection time and the host port number of the host bus adapter for the credit occupancy data.

[0041] In this embodiment, credit usage data includes changes in credit usage and remaining credit. Changes in credit usage are collected when the credit usage changes; credit usage refers to the amount of credit currently occupied by the host device. Remaining credit refers to the amount of unused credit on the storage device. To determine changes in credit usage, this embodiment can determine whether the credit usage at the current collection time has changed compared to the historical credit usage at the previous collection time. If it has changed, the change in credit usage between the two can be determined. Subsequently, the changes in credit usage and remaining credit can be set as the aforementioned credit usage data. At this point, the credit usage situation can be understood from the host device side.

[0042] Furthermore, the credit occupancy and credit remaining amount can be stored in the PCIe registers (Peripheral Component Interconnect express, high-speed serial computer expansion bus standard) of the host bus adapter (HBA, Fibre Channel Host Bus Adapter). Specifically, the credit occupancy is stored in the Credit Request Register, and the credit remaining amount is stored in the RX Credit Value register. This embodiment can periodically access the host bus adapter and read the credit occupancy and credit remaining amount from the registers of the host bus adapter.

[0043] Furthermore, after completing the credit occupancy data collection, the credit occupancy data can be marked with the first timestamp of the current collection time and the host port number (WWN, World Wide Number, unique identifier) ​​of the host bus adapter, so that the credit occupancy data can be time-aligned and associated with other data and devices based on the first timestamp and the host port number.

[0044] Furthermore, to improve the efficiency of credit occupancy data collection, this embodiment can collect data from the registers of the host bus adapter in the kernel mode of the host device. Subsequently, the credit occupancy and remaining credit can be stored in a shared memory space between the kernel mode and user mode. Then, the credit occupancy and remaining credit can be read from the memory space in the user mode of the host device for subsequent processing.

[0045] In one implementation, accessing the host bus adapter of the host device and reading the credit usage and credit remaining amount from the registers of the host bus adapter may include:

[0046] Step 21: In the kernel mode of the host device, access the host bus adapter of the host device, read the credit usage and credit remaining in the register of the host bus adapter, and send the credit usage and credit remaining to the memory space shared by the kernel mode and user mode.

[0047] Step 22: In the user space of the host device, read the credit usage and credit remaining from the memory space.

[0048] In addition to collecting credit occupancy data from the host device, credit release latency can also be checked and collected for subsequent analysis. Credit release latency refers to the time from when the host device sends a data frame containing a read / write request to when the host device receives the credit value release frame. Furthermore, other registers on the host bus adapter can be accessed to obtain other information related to the communication link and data transmission. For example, the link status register of the host bus adapter can be accessed, and the connection / disconnection status of the physical link can be determined based on the value of that register.

[0049] Furthermore, steps 11 to 13 above can be executed by the agent program in the host device, and the host agent program needs to upload the credit occupancy data to the data analysis device. At this point, to improve the efficiency and security of data transmission, the agent program can encode, compress, encrypt, etc., the credit occupancy data, and then send the processed credit occupancy data to the data analysis device. For a better understanding of the host agent program's workflow, please refer to [link to documentation / reference]. Figure 4 , Figure 4 This is a schematic diagram illustrating the workflow of a host agent program provided in an embodiment of the present invention.

[0050] In one implementation, collecting credit release data from the switching device may include:

[0051] Step 31: Mirror the data traffic passing through the switching device to obtain mirrored data traffic, and match the preset frame header information with the data frame headers of each data frame in the mirrored data traffic to obtain the credit value release frame in the mirrored data traffic.

[0052] Step 32: Extract the source port number, destination port number, and credit release amount from the credit release frame.

[0053] Step 33: Obtain the acquisition time of the current credit value release frame and the historical acquisition time of the previous credit value release frame, and use the acquisition time and historical acquisition time to determine the credit release interval.

[0054] Step 34: Set the credit release interval and credit release amount to the credit release data, and mark the second timestamp of the collection time, source port number, and destination port number for the credit release data.

[0055] In this embodiment, the credit release data may include a credit release interval and a credit release amount. The credit release interval refers to the time interval between each credit release by the storage device, and the credit release amount refers to the amount of credit released by the storage device each time. This embodiment obtains the credit release data by collecting the aforementioned credit value release frames in the switching device.

[0056] Specifically, in this embodiment, the data traffic passing through the switching device is first mirrored to obtain mirrored data traffic. Then, the preset frame header information is matched with the data frame headers of each data frame in the mirrored data traffic to obtain the credit value release frame in the mirrored data traffic. The identification control field value of the credit value release frame header is 0x21, so the credit release frame can be extracted based on this field in the frame header.

[0057] Upon extracting a credit value release frame, its data frame can be parsed, and the source port number, destination port number, and credit release amount can be extracted. Since the credit value release frame is sent from the storage device to the host device, the source port number corresponds to the storage port of the storage device, and the destination port number corresponds to the host port of the host device. The credit release amount indicates how much credit value was released by the storage device this time. Furthermore, other information in the credit value release frame, such as the frame sequence number, can be parsed to perform time-series association of several credit value release frames based on the frame sequence number. Subsequently, to determine the credit value release interval, this embodiment can obtain the acquisition time of the current credit value release frame and the historical acquisition time of the previous credit value release frame, and use the acquisition time and historical acquisition time to determine the credit release interval. Then, the credit release interval and credit release amount can be set as credit release data, allowing the credit release status to be understood from the switch device side.

[0058] Furthermore, to facilitate time alignment and link association, a second timestamp, source port number, and destination port number can be used to mark the collection time of the credit release data.

[0059] Furthermore, steps 31-34 above can be executed by the probe program in the switch device, which needs to upload the credit release data to the data analysis device. At this point, to improve data transmission efficiency and security, the probe program can encode, compress, encrypt, etc., the credit occupancy data, and then send the processed credit occupancy data to the data analysis device. For a better understanding of the host agent program's workflow, please refer to [link to documentation / reference]. Figure 5 , Figure 5 This is a schematic diagram illustrating the workflow of a switch probe program provided in an embodiment of the present invention.

[0060] In one implementation, collecting cache pressure data from the storage device may include:

[0061] Step 41: Determine the corresponding access interface based on the device model of the storage device.

[0062] Step 42: At the current acquisition time, obtain the cache queue release amount, cache queue occupancy amount, cache queue release delay time, and cache allocation rate of the cache queues corresponding to each storage port of the storage device through the access interface.

[0063] Step 43: Set the cache queue release amount, cache queue occupancy amount, cache queue release delay time, and cache allocation rate as cache pressure data, and label the cache pressure data with the third timestamp of the current collection time, the storage port number of the storage port, and the logical unit number corresponding to the logical unit in the storage device.

[0064] In this embodiment, cache pressure data may include cache queue release amount, cache queue occupancy, cache queue release latency, and cache allocation rate. Cache queue release amount refers to the number of read / write requests completed from the cache queue in a single operation, i.e., the number of credits released from the cache queue in a single operation. Cache queue occupancy refers to the number of unprocessed read / write requests in the current cache queue, i.e., the number of credits still being occupied. Cache queue release latency refers to the time from when a read / write request is written to the cache queue to when it is removed from the cache queue. Cache allocation rate is the proportion of memory allocated to the cache queue. Cache queue release amount and cache queue occupancy reflect the current amount of credits occupied and released, cache queue release latency reflects the current credit release speed, and cache allocation rate reflects the memory utilization of the storage device. Therefore, these four values ​​reflect the overall cache pressure of the storage device.

[0065] Furthermore, to facilitate time alignment and link association, this embodiment sets the cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate as cache pressure data, and then labels the cache pressure data with the third timestamp of the current collection time, the storage port number of the storage port, and the logical unit number (LUN) corresponding to the logical unit number in the storage device. The logical unit number can further improve the location of the fault.

[0066] Furthermore, since different storage devices come from different manufacturers and are of different types, and each manufacturer and type of storage device has a different access interface (API, Application Programming Interface), this embodiment needs to determine the corresponding access interface based on the storage device model in order to obtain the aforementioned cache pressure data through the corresponding access interface.

[0067] Furthermore, regarding cache pressure detection, considering the varying importance of cache queue release amount, cache queue occupancy, cache queue release latency, and cache allocation rate, these factors can be weighted to obtain a cache pressure metric, which is then set as the cache pressure data. In this case, the cache pressure data better reflects the cache pressure.

[0068] In one implementation, the cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate are set as cache pressure data, including:

[0069] Step 51: Weight the cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate to obtain the cache pressure index;

[0070] Step 52: Set the cache pressure metric to the cache pressure data.

[0071] For example, one weighting method is:

[0072] BPI = 0.5 × Q d / Q max +0.3×B a / 100+0.2×500min(L at ,500) / 500;

[0073] Here, BPI represents the cache pressure index. Q d Q represents the current cache queue occupancy. max B represents the maximum queue capacity of the port. a L represents the cache allocation rate. at This indicates a delay in credit release.

[0074] Furthermore, steps 41-43 above can be executed by the agent program in the storage device, and the storage agent program needs to upload the cache pressure data to the data analysis device. At this point, to improve the efficiency and security of data transmission, the storage agent program can encode, compress, encrypt, etc., the credit occupancy data, and then send the processed credit occupancy data to the data analysis device. For a better understanding of the host agent program's workflow, please refer to [link to documentation / reference]. Figure 6 , Figure 6 This is a schematic diagram illustrating the workflow of a storage agent program provided in an embodiment of the present invention.

[0075] Furthermore, in addition to collecting credit usage data, credit release data, and cache pressure data, to facilitate further fault analysis of host devices, switch devices, and storage devices, this embodiment can also collect operating system operation data from host devices, physical layer operation data and protocol layer operation data from switch devices, and disk operation data from storage devices. For example, operating system operation data can be read / write request scheduling data and storage software operation data. Physical layer operation data can be optical attenuation data from the data transmission channel (TX) and data reception channel (RX) of the switch device. Protocol layer operation data can be the number of busy frames and rejected frames received by the switch device. Disk operation data can include RAID group latency (Redundant Arrays of Independent Disks), disk error rate, etc.

[0076] In one implementation, it may further include:

[0077] Step 61: Collect operating system running data from the host device, physical layer working data and protocol layer working data from the switch device, and disk working data from the storage device.

[0078] S102. Time-align the credit occupancy data, credit release data, and cache pressure data according to the collection time, and link them according to the communication links formed by the host device, switch device, and storage device to obtain the associated operation data.

[0079] In this embodiment, since timestamps and port numbers have been set in the credit occupancy data, credit release data, and cache pressure data, the credit occupancy data, credit release data, and cache pressure data can be time-aligned according to the collection time and linked according to the communication links formed by the host device, switch device, and storage device to obtain associated operation data.

[0080] In one implementation, credit occupancy data, credit release data, and cache pressure data are time-aligned according to their collection time and linked together according to the communication links formed by the host device, switch device, and storage device to obtain associated operational data, including:

[0081] Step 71: Extract the host port number and first timestamp from the credit occupancy data; extract the source port number, destination port number, and second timestamp from the credit release data; and extract the storage port number, logical unit number, and third timestamp from the cache pressure data.

[0082] Step 72: Align the credit usage data, credit release data, and cache pressure data according to the first, second, and third timestamps.

[0083] Step 73: Match the host port number, source port number, destination port number, and storage port number, and establish an association between the credit occupancy data, credit release data, and cache pressure data and the communication links formed by the host port, switch port, storage port, and logical units in the storage device based on the matching results and logical unit numbers to obtain the associated operation data.

[0084] Please refer to Figure 7 , Figure 7 This is a schematic diagram illustrating a communication link association provided in an embodiment of the present invention. Credit occupancy data, credit release data, and cache pressure data aligned according to timestamps can be associated with a communication link formed by host ports, switch ports, storage ports, and logical units in storage devices, thereby enabling monitoring of the operational status of each device node in the communication link according to time.

[0085] S103. Perform correlation anomaly detection on the associated running data, and determine the fault location in the host device, switch device, and storage device based on the anomaly detection results.

[0086] In this embodiment, associated anomaly detection can be directly performed on the associated operational data, and the fault location can be determined in the host device, switch device, and storage device based on the anomaly detection results. Specifically, multiple preset fault modes can be pre-set, each fault mode corresponding to a fault location, and each fault mode corresponding to a data distribution pattern and data change pattern of the associated operational data. Furthermore, in actual associated anomaly detection, this embodiment can acquire multiple sets of associated operational data within a preset time period, and determine the first data change data of credit occupancy data, the second data change data of credit release data, and the third data change data of cache pressure data within the preset time period based on the multiple sets of associated operational data. The data change data can include numerical change ranges, numerical increase / decrease amounts, numerical increase / decrease rates, and durations of numerical increase / decrease maintenance. Subsequently, this embodiment can match the first data change data, the second data change data, and the third data change data with the preset fault modes to determine the successfully matched target fault mode, thereby determining the fault location corresponding to the target fault mode.

[0087] In one implementation, associated anomaly detection is performed on the associated operational data, and the fault location is determined in the host device, switch device, and storage device based on the anomaly detection results, including:

[0088] Step 81: Obtain multiple sets of related operational data within a preset time period, and determine the first data change data of credit occupancy data, the second data change data of credit release data, and the third data change data of cache pressure data within the preset time period based on the multiple sets of related operational data;

[0089] Step 82: Among at least two preset fault modes, determine the target fault mode matched by the combination of the first data change data, the second data change data, and the third data change data, and determine the fault location corresponding to the target fault mode.

[0090] For example, when the credit usage of the host device increases significantly, the credit release detected by the switch device is normal, and the cache pressure in the storage device is high, the fault mode can be determined to be storage port overload, and the fault location can be determined to be the storage device.

[0091] Of course, in addition to using the aforementioned credit occupancy data, credit release data, and cache pressure data, fault mode matching can also be performed by combining other data from host devices, switch devices, and storage devices. For example, a preset fault mode matching method is as follows:

[0092] Table 1. Correlation Anomaly Detection

[0093]

[0094] Among them, ELS frame refers to Extended Link Service frame, CRC field is used to verify the integrity of data during transmission, CRC error indicates that ELS frame is incomplete.

[0095] Furthermore, to improve the targeting of detection by performing correlation anomaly detection only when the storage system is in poor condition, this embodiment can also preliminarily determine the health of the storage system based on credit flow, and only perform correlation anomaly detection on the storage system when the storage system health is poor. The method for determining the health of credit flow is described below.

[0096] In one embodiment, the method may further include:

[0097] Step 91: In the associated running data, extract the change in credit usage from the credit usage data, extract the credit release amount from the credit release data, and extract the cache queue release amount and cache queue usage from the cache pressure data.

[0098] Step 92: Determine the cache queue release rate based on the cache queue release amount and cache queue occupancy amount. Determine the initial expected credit release amount based on the credit occupancy change and cache queue release rate. Take the minimum value between the initial expected credit release amount and the maximum credit release amount of the switch device as the expected credit release amount.

[0099] Step 93: Determine the health of the credit flow using the credit release amount and the expected credit release amount, and determine whether the health of the credit flow is less than the preset threshold.

[0100] Step 94: If the credit flow health is less than the preset threshold, proceed to the step of detecting correlation anomalies in the associated running data.

[0101] In this embodiment, the expected credit release amount refers to the anticipated value determined based on the amount of credit occupied by the host device and the proportion of credit released by the storage device, and the expected credit release amount should not exceed the maximum credit release amount supported by the switch device. The actual credit release amount is the credit release amount detected in the switch device. Under normal circumstances, the actual credit release amount should be close to or the same as the expected credit release amount. However, due to factors such as switch port abnormalities or storage device processing abnormalities, the actual credit release amount may be less than the expected credit release amount. Therefore, this embodiment determines the difference between the actual and expected credit release amounts by collecting the actual credit release amount and calculating the expected credit release amount, and by determining the ratio of the actual to the expected credit release amount, thereby assessing the health of the storage system.

[0102] Specifically, the formula for calculating the expected credit release is as follows:

[0103] ;

[0104] Among them, F e This indicates the expected amount of credit to be released. R h This indicates the change in the host's credit usage. (C) r This represents the cache queue release rate, which is the number of cache queues released divided by the total number of cache queues occupied. R h ×C r According to the supply and demand relationship, based on the current cache queue release rate, the amount of credit occupied by the host can be released by the storage device in a timely manner. max This indicates the upper limit of the credit release frames that the physical link can carry.

[0105] The formula for calculating credit flow health is:

[0106] HealthScore = Actual number of credit releases (Fs) / Expected number of credit releases (Fe) × 100%;

[0107] Since a higher credit flow health value indicates that the actual credit release is closer to the expected release, and the health value is higher, and vice versa, a lower value indicates a lower health value, when the credit flow health value is determined to be less than the preset threshold, the storage system can be associated with anomaly detection.

[0108] Furthermore, credit flow health can also be used to adjust the frequency of data collection. The following are the actions that data analytics devices can trigger when credit flow health is at different thresholds:

[0109] Table 2 Data Analysis Equipment Operations

[0110]

[0111] Among them, when the secondary monitoring is triggered, the frequency of collecting credit occupancy data, credit release data, and cache pressure data can be increased.

[0112] Furthermore, after determining the location of the fault, root cause analysis can be performed by combining operating system operation data collected from the host device, physical layer and protocol layer operation data collected from the switch device, and disk operation data collected from the storage device. In this embodiment, the focus will be on detecting anomalies in the host device and storage device, which will be considered the primary causes of the fault; while network transmission faults in the switch device can be considered secondary causes.

[0113] In one implementation, after determining the fault location corresponding to the target fault mode, the method may further include:

[0114] Step 1001: When the fault location is determined to be the host device, the host device is used to perform fault detection using system operation data and credit usage data to obtain the main fault cause;

[0115] Step 1002: When the fault location is determined to be the storage device, the storage device is used to detect the fault using cache pressure data and disk working data to obtain the main cause of the fault;

[0116] Step 1003: Use physical layer working data and protocol layer working data to perform fault detection on the switch device and obtain the causes of minor faults;

[0117] Step 1004: Generate fault detection results using the primary fault cause and secondary fault cause.

[0118] Finally, please refer to Figure 8 , Figure 8 This is a structural block diagram of a data analysis system provided in an embodiment of the present invention.

[0119] Based on the above embodiments, the present invention first collects credit occupancy data from the host device, credit release data from the switch device, and cache pressure data from the storage device. Specifically, the host device occupies credit provided by the storage device by issuing read / write requests, and the storage device releases the credit occupied by the host device to the host device via the switch device upon completion of the read / write request processing. That is, the present invention can monitor the credit occupancy and release status of the host device, switch device, and storage device respectively. Subsequently, the present invention can time-align the credit occupancy data, credit release data, and cache pressure data according to the collection time, and link them according to the communication links formed by the host device, switch device, and storage device to obtain associated operational data, thereby associating the working status of the host device, switch device, and storage device. Finally, the present invention can perform associated anomaly detection on the associated operational data, and determine the fault location in the host device, switch device, and storage device based on the anomaly detection results. This enables the detection of anomalies in the storage system from a global perspective of the communication links and allows for fine-grained fault location, thereby improving the reliability and correlation of anomaly detection in the storage system.

[0120] Based on the above embodiments, in addition to collecting credit usage data and cache pressure data, this embodiment can also use credit usage data and cache pressure data for early warning, so as to further improve the reliability of anomaly detection results.

[0121] In one implementation, after determining that the credit usage at the current collection time has changed compared to the historical credit usage at the previous collection time, the method may further include:

[0122] S201. Determine whether the remaining credit balance is zero at least at two data collection times.

[0123] S202. If so, generate a host warning message indicating that the remaining credit has been exhausted.

[0124] In this embodiment, it is possible to detect in advance whether the host device has encountered a situation where the remaining credit value has been exhausted. If so, a host warning message can be generated and output in advance.

[0125] In one embodiment, the method may further include:

[0126] S301. Among at least two preset exception combinations, determine the matching target exception combination based on the values ​​of cache queue release amount, cache queue occupancy amount, cache queue release delay time, cache allocation rate, and cache pressure index.

[0127] S302. Generate storage device anomaly warning information based on the target anomaly combination.

[0128] In this embodiment, cache queue release amount, cache queue occupancy, cache queue release latency, cache allocation rate, and cache pressure indicators can also be used to predict potential storage device failures, thereby improving detection effectiveness. It should be noted that this embodiment is not limited to specific combinations of anomalies; for example, it can include:

[0129] Table 3 Predicted Root Causes

[0130]

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0132] Please refer to Figure 9 , Figure 9 This is a structural block diagram of a storage system anomaly detection device provided in an embodiment of the present invention. The device may include:

[0133] The data acquisition module 901 is used to collect credit occupancy data from the host device, credit release data from the switch device, and cache pressure data from the storage device. The host device occupies the credit provided by the storage device by issuing read and write requests to the storage device, and the storage device releases the credit occupied by the host device to the host device through the switch device when it completes the read and write request processing.

[0134] The data association module 902 is used to time-align credit occupancy data, credit release data, and cache pressure data according to the collection time, and to link-associate them according to the communication links formed by the host device, switch device, and storage device to obtain associated operation data.

[0135] The association detection module 903 is used to perform association anomaly detection on the associated running data and determine the fault location in the host device, switch device, and storage device based on the anomaly detection results.

[0136] Optionally, the device may further include:

[0137] The data extraction module is used to extract the change in credit usage from credit usage data, the amount of credit release from credit release data, and the amount of cache queue release and cache queue usage from cache pressure data in the associated running data.

[0138] The calculation module is used to determine the cache queue release rate based on the cache queue release amount and cache queue occupancy, and to determine the initial expected credit release amount based on the credit occupancy change and cache queue release rate. The minimum value between the initial expected credit release amount and the maximum credit release amount of the switch device is used as the expected credit release amount.

[0139] The judgment module is used to determine the health of the credit flow using the credit release amount and the expected credit release amount, and to determine whether the health of the credit flow is less than a preset threshold.

[0140] The association detection module 903 can also be used to perform association anomaly detection on the associated running data if the credit flow health is less than a preset threshold.

[0141] Optionally, the data association module 902 includes:

[0142] The data extraction submodule is used to extract the host port number and first timestamp from the credit occupancy data, the source port number, destination port number, and second timestamp from the credit release data, and the storage port number, logical unit number, and third timestamp from the cache pressure data.

[0143] The time-related submodule is used to align the credit occupancy data, credit release data, and cache pressure data according to the first, second, and third timestamps.

[0144] The link association submodule is used to match the host port number, source port number, destination port number, and storage port number, and establish an association between the credit occupancy data, credit release data, and cache pressure data and the communication links formed by the host port, switch port, storage port, and logical units in the storage device based on the matching results and logical unit numbers, thereby obtaining the associated operation data.

[0145] Optionally, the data acquisition module 901 includes:

[0146] The host acquisition submodule is used to access the host bus adapter of the host device at the current acquisition time and read the credit occupancy and credit remaining amount from the register of the host bus adapter. When it is determined that the credit occupancy at the current acquisition time has changed compared with the historical credit occupancy at the previous acquisition time, the change in credit occupancy at the current acquisition time is determined based on the credit occupancy and the historical credit occupancy. The change in credit occupancy and the credit remaining amount are set as credit occupancy data, and the credit occupancy data is marked with the first timestamp of the current acquisition time and the host port number of the host bus adapter.

[0147] Optionally, the host acquisition submodule can be used for:

[0148] In the kernel mode of the host device, the host bus adapter of the host device is accessed, the credit occupancy and credit remaining amount are read from the register of the host bus adapter, and the credit occupancy and credit remaining amount are sent to the memory space shared by the kernel mode and the user mode.

[0149] In the user space of the host device, the credit usage and credit remaining amount are read from the memory space.

[0150] Optionally, the data acquisition module 901 further includes:

[0151] The host warning module is used to determine whether the remaining credit balance is zero at least at two data collection times; if so, it generates a host warning message indicating that the remaining credit balance has been exhausted.

[0152] Optionally, the data acquisition module 901 further includes:

[0153] The switch acquisition submodule is used to mirror the data traffic passing through the switch device to obtain mirrored data traffic. It matches the preset frame header information with the data frame headers of each data frame in the mirrored data traffic to obtain the credit value release frame in the mirrored data traffic. It extracts the source port number, destination port number, and credit release amount from the credit value release frame. It obtains the acquisition time of the current credit value release frame and the historical acquisition time of the previous credit value release frame, and uses the acquisition time and historical acquisition time to determine the credit release interval. It sets the credit release interval and credit release amount as credit release data, and marks the credit release data with the second timestamp of the acquisition time, the source port number, and the destination port number.

[0154] Optionally, the data acquisition module 901 further includes:

[0155] The storage acquisition submodule is used to determine the corresponding access interface based on the device model of the storage device. At the current acquisition time, it obtains the cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate of the cache queues corresponding to each storage port of the storage device through the access interface. The cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate are set as cache pressure data, and the cache pressure data is labeled with the third timestamp of the current acquisition time, the storage port number of the storage port, and the logical unit number corresponding to the logical unit in the storage device.

[0156] Optionally, the storage acquisition submodule can be used for:

[0157] The cache pressure metric is obtained by weighting the cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate; the cache pressure metric is then set as the cache pressure data.

[0158] Optionally, the data acquisition module 901 also includes:

[0159] The storage alert submodule is used to determine the target abnormal combination from at least two preset abnormal combinations based on the values ​​of cache queue release amount, cache queue occupancy amount, cache queue release delay time, cache allocation rate, and cache pressure index; and to generate storage device abnormal alert information based on the target abnormal combination.

[0160] Optionally, the association detection module 903 is used for:

[0161] Acquire multiple sets of related operational data within a preset time period, and determine the first data change data of credit occupancy data, the second data change data of credit release data, and the third data change data of cache pressure data within the preset time period based on the multiple sets of related operational data;

[0162] In at least two preset fault modes, determine the target fault mode matched by the combination of the first data change data, the second data change data, and the third data change data, and determine the fault location corresponding to the target fault mode.

[0163] Optionally, the data acquisition module 901 can also be used to acquire operating system running data from the host device, physical layer working data and protocol layer working data from the switch device, and disk working data from the storage device.

[0164] The device may also include:

[0165] The fault cause detection module is used to detect faults in host devices by using system operation data and credit usage data when the fault location is determined to be a host device, and to obtain the primary fault cause; when the fault location is determined to be a storage device, it uses cache pressure data and disk operation data to detect faults in the storage device, and to obtain the primary fault cause; it uses physical layer operation data and protocol layer operation data to detect faults in switch devices, and to obtain the secondary fault causes; and it generates fault detection results using the primary and secondary fault causes.

[0166] For a description of the features in the embodiment corresponding to the storage system anomaly detection device, please refer to the relevant description of the embodiment corresponding to the storage system anomaly detection method, which will not be repeated here.

[0167] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the storage system anomaly detection method.

[0168] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the storage system anomaly detection method.

[0169] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0170] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the storage system anomaly detection method.

[0171] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described storage system anomaly detection method embodiments.

[0172] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0173] The above provides a detailed description of the storage system anomaly detection method, electronic device, storage medium, and program product provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of this invention.

Claims

1. A method for detecting anomalies in a storage system, characterized in that, include: The system collects credit usage data and operating system operation data from the host device, credit release data, physical layer operation data, and protocol layer operation data from the switch device, and cache pressure data and disk operation data from the storage device. Specifically, the host device uses the credit provided by the storage device by issuing read / write requests to the storage device, and the storage device releases the credit used by the host device to the host device through the switch device upon completion of the read / write request processing. The credit occupancy data, credit release data, and cache pressure data are time-aligned according to the collection time, and linked together according to the communication links formed by the host device, the switch device, and the storage device to obtain associated operation data; Acquire multiple sets of associated operational data within a preset time period, and determine the first data change data of the credit occupancy data, the second data change data of the credit release data, and the third data change data of the cache pressure data within the preset time period based on the multiple sets of associated operational data; In at least two preset fault modes, determine the target fault mode matched by the combination of the first data change data, the second data change data, and the third data change data, and determine the fault location corresponding to the target fault mode; When the fault location is determined to be the host device, the fault detection of the host device is performed using the system operation data and the credit occupancy data to obtain the main fault cause. When the fault location is determined to be a storage device, the storage device is used to perform fault detection using the cache pressure data and the disk working data to obtain the main fault cause. The physical layer working data and protocol layer working data are used to perform fault detection on the switch device to obtain the causes of minor faults. Fault detection results are generated using the primary fault cause and the secondary fault cause.

2. The storage system anomaly detection method according to claim 1, characterized in that, Also includes: In the associated operational data, the change in credit usage is extracted from the credit usage data, the credit release amount is extracted from the credit release data, and the cache queue release amount and cache queue usage amount are extracted from the cache pressure data; The cache queue release rate is determined based on the cache queue release amount and the cache queue occupancy amount. The initial expected credit release amount is determined based on the credit occupancy change amount and the cache queue release rate. The minimum value between the initial expected credit release amount and the maximum credit release amount of the switch device is taken as the expected credit release amount. The credit flow health is determined using the credit release amount and the expected credit release amount, and it is determined whether the credit flow health is less than a preset threshold. If the credit flow health is less than the preset threshold, then proceed to the step of detecting association anomalies in the associated running data.

3. The storage system anomaly detection method according to claim 1, characterized in that, The credit occupancy data, credit release data, and cache pressure data are time-aligned according to their collection time, and linked together according to the communication links formed by the host device, switch device, and storage device to obtain associated operational data, including: Extract the host port number and first timestamp from the credit occupancy data; extract the source port number, destination port number, and second timestamp from the credit release data; and extract the storage port number, logical unit number, and third timestamp from the cache pressure data. The credit occupancy data, credit release data, and cache pressure data are time-aligned based on the first timestamp, the second timestamp, and the third timestamp. The host port number, the source port number, the destination port number, and the storage port number are matched, and based on the matching result and the logical unit number, the credit occupancy data, the credit release data, and the cache pressure data are associated with the communication links formed by the host port, the switch port, the storage port, and the logical units in the storage device to obtain the associated operation data.

4. The storage system anomaly detection method according to claim 3, characterized in that, Collect credit usage data from the host device, including: At the current acquisition time, access the host bus adapter of the host device and read the credit usage and credit remaining amount from the register of the host bus adapter; When it is determined that the credit usage at the current collection time has changed compared to the historical credit usage at the previous collection time, the change in credit usage at the current collection time is determined based on the credit usage and the historical credit usage. The credit occupancy change and the remaining credit amount are set as the credit occupancy data, and the credit occupancy data is marked with the first timestamp of the current collection time and the host port number of the host bus adapter.

5. The storage system anomaly detection method according to claim 4, characterized in that, Access the host bus adapter of the host device and read the credit usage and credit remaining amount from the registers of the host bus adapter, including: In the kernel mode of the host device, the host bus adapter of the host device is accessed, the credit occupancy and credit remaining amount are read from the register of the host bus adapter, and the credit occupancy and credit remaining amount are sent to the memory space shared by the kernel mode and the user mode. In the user mode of the host device, the credit usage and credit remaining amount are read from the memory space.

6. The storage system anomaly detection method according to claim 4, characterized in that, After determining that the credit usage at the current collection time has changed compared to the historical credit usage at the previous collection time, the following is also included: Determine whether the remaining credit amount is zero at at least two data collection times; If so, a host warning message indicating that the remaining credit balance has been exhausted will be generated.

7. The storage system anomaly detection method according to claim 3, characterized in that, Collect credit release data from the switching device, including: The data traffic passing through the switch device is mirrored to obtain mirrored data traffic, and the preset frame header information is matched with the data frame header of each data frame in the mirrored data traffic to obtain the credit value release frame in the mirrored data traffic. Extract the source port number, the destination port number, and the credit release amount from the credit value release frame; Obtain the acquisition time of the current credit value release frame and the historical acquisition time of the previous credit value release frame, and use the acquisition time and the historical acquisition time to determine the credit release interval; Set the credit release interval and the credit release amount to the credit release data, and mark the credit release data with the second timestamp of the collection time, the source port number, and the destination port number.

8. The storage system anomaly detection method according to claim 3, characterized in that, Collect cache stress data from storage devices, including: The corresponding access interface is determined based on the device model of the storage device; At the current acquisition time, the cache queue release amount, cache queue occupancy amount, cache queue release delay time, and cache allocation rate of the cache queues corresponding to each storage port of the storage device are obtained through the access interface; The cache queue release amount, cache queue occupancy amount, cache queue release delay time, and cache allocation rate are set as the cache pressure data, and the cache pressure data is labeled with the third timestamp of the current collection time, the storage port number of the storage port, and the logical unit number corresponding to the logical unit in the storage device.

9. The storage system anomaly detection method according to claim 8, characterized in that, The cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate are set as the cache pressure data, including: The cache pressure index is obtained by weighting the cache queue release amount, cache queue occupancy, cache queue release delay time, and cache allocation rate. Set the cache pressure metric to the cache pressure data.

10. The storage system anomaly detection method according to claim 9, characterized in that, Also includes: Among at least two preset abnormal combinations, a matching target abnormal combination is determined based on the values ​​of the cache queue release amount, cache queue occupancy amount, cache queue release delay time, cache allocation rate, and the cache pressure index. Storage device anomaly warning information is generated based on the target anomaly combination.

11. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the storage system anomaly detection method as described in any one of claims 1 to 10.

12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the storage system anomaly detection method as described in any one of claims 1 to 10.

13. A non-volatile computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when loaded and executed by a processor, implement the storage system anomaly detection method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Shared cache release method and device based on credit back pressure

    CN118101597A

  • Fault positioning method, device, equipment and computer program product

    CN118802491A