Fault diagnosis methods, devices, electronic equipment and storage media for storage systems
By parsing the message data of the FC SAN storage system layer by layer, especially the performance analysis of the SCSI layer, the problems of delay and low accuracy in fault location in the existing technology have been solved, and fast and accurate fault diagnosis and location have been achieved.
Patent Information
- Application Number
- CN202511225690.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing technologies struggle to quickly and accurately locate I/O read/write faults in FC SAN storage systems, especially in multi-tiered architectures where the lack of in-depth analysis of the SCSI layer leads to delays and low accuracy in fault location.
By acquiring message data from the storage system, parsing protocol information layer by layer based on the multi-level structure, focusing on analyzing the performance data of the SCSI layer, and combining read and write operations, we can achieve full-level fault location and accurate diagnosis.
It enables rapid and accurate fault diagnosis of FC SAN storage systems, accurately identifies fault types and locates them at the device component level, thereby improving operation and maintenance efficiency and system reliability.
Smart Images

Figure CN120743716B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault diagnosis technology, and in particular to fault diagnosis methods, devices, electronic equipment and storage media for storage systems. Background Technology
[0002] Information centers handle various internal and external business processes, making their performance and stability crucial. Fiber Channel Storage Area Networks (FC SANs), with their high bandwidth and low latency, have become the preferred storage system configuration for mission-critical scenarios. However, as business grows, the scale of storage systems also expands, making it difficult to quickly detect and accurately pinpoint faults when anomalies occur. Therefore, accurate read / write fault diagnosis for FC SAN storage systems has become a pressing issue. Summary of the Invention
[0003] This application provides a method, apparatus, electronic device, and storage medium for fault diagnosis of storage systems, in order to at least solve the problem of how to accurately diagnose read and write faults in FC SAN storage systems in related technologies.
[0004] This application provides a fault diagnosis method for a storage system, including: acquiring message data of the storage system;
[0005] Based on the multi-level structure of the storage system, the message data is parsed layer by layer to obtain the protocol information corresponding to each layer. The multi-level structure includes at least the system interface layer, which is the highest layer in the multi-level structure of the storage system.
[0006] Based on the protocol information corresponding to each layer and the read and write operations of the system interface layer, determine the target performance data of the system interface layer;
[0007] Based on the target performance data, fault diagnosis is performed on the storage system to obtain the fault diagnosis results.
[0008] This application also provides a fault diagnosis device for a storage system, including: an acquisition module for acquiring message data of the storage system;
[0009] The processing module is used to parse the message data layer by layer based on the multi-level structure of the storage system to obtain the protocol information corresponding to each layer. The multi-level structure includes at least the system interface layer, which is the highest layer in the multi-level structure of the storage system.
[0010] The processing module is also used to determine the target performance data of the system interface layer based on the protocol information corresponding to each layer and the read and write operations of the system interface layer.
[0011] The processing module is also used to diagnose faults in the storage system based on target performance data and obtain fault diagnosis results.
[0012] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the fault diagnosis method of any of the above-described memory systems when executing the computer program.
[0013] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the fault diagnosis method of any of the above-described storage systems.
[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described fault diagnosis methods for storage systems.
[0015] This application obtains message data from a storage system; based on the multi-level structure of the storage system, the message data is parsed layer by layer to obtain the protocol information corresponding to each layer. The multi-level structure includes at least a system interface layer, which is the highest layer in the multi-level structure of the storage system; based on the protocol information corresponding to each layer and the read / write operations of the system interface layer, the target performance data of the system interface layer is determined; based on the target performance data, fault diagnosis is performed on the storage system to obtain the fault diagnosis results. In this scheme, data parsing can be performed on each layer of the storage system to achieve full-level fault location, and the performance data of the system interface layer, i.e., the SCSI layer, is analyzed in detail. This allows for accurate analysis of fault conditions in conjunction with read / write operations, enabling accurate diagnosis, precise location, and type determination of read / write faults in the storage system. Attached Figure Description
[0016] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating a fault diagnosis method for a storage system provided in this application embodiment. Figure 1 ;
[0018] Figure 2 A flowchart illustrating a fault diagnosis method for a storage system provided in this application embodiment. Figure 2 ;
[0019] Figure 3A schematic diagram of the structure of a fault diagnosis device for a storage system provided in an embodiment of this application;
[0020] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0022] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0023] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0024] In today's highly information-driven era, business systems across all industries heavily rely on the stable operation of information centers. Information centers handle various internal and external business processes, making their performance and stability paramount. Business systems running within information centers, such as databases and cloud computing platforms, typically require extensive data interaction with storage systems. FC SAN, with its high bandwidth and low latency, has become the preferred storage network solution for mission-critical scenarios.
[0025] FC SAN is a high-speed dedicated storage network built on Fibre Channel technology, used to connect servers and storage devices for efficient data transmission and centralized management. Based on the Fibre Channel (FC) protocol, FC SAN interconnects storage arrays (such as disk arrays and tape libraries) with server hosts through Fibre Channel switches (FC switches), forming an independent dedicated storage network. An FC SAN can include: Fibre Channel switches, responsible for connecting storage devices and servers, enabling high-speed data transmission and supporting multi-path redundancy (such as dual FC switches or dual HBA cards); storage devices, supporting Fibre Channel disk arrays, NAS, etc., connected to the switches via fiber optic interfaces; and servers, communicating with storage devices via the FC protocol to achieve data transmission and storage, supporting load balancing and fault tolerance mechanisms.
[0026] As FC SANs continue to expand in scale, their complexity also increases. When business systems experience performance issues, traditional Network Performance Monitoring (NPM) and Application Performance Management (APM) can help analyze the network and business application systems, quickly pinpointing whether the problem originates at the network layer or the host application layer. However, when the problem involves FC SAN storage connected to the database, fault location becomes extremely difficult. At this point, it becomes challenging to determine whether the problem originates within the database itself or within the backend FC SAN storage, such as the storage controller, disk array, or FC network links.
[0027] The development of FC SAN in China has made significant progress in recent years, but with the expansion of scale, related problems have also gradually increased. I / O read / write issues in FC SAN storage networks have become increasingly complex, mainly manifested in the following aspects:
[0028] Multi-level fault location is difficult: FC SAN involves multiple layers, including the FC physical layer, link layer, network layer, transport layer, and the upper Small Computer System Interface (SCSI) protocol layer. Faults at each layer can cause input / output (I / O) read / write anomalies, but existing commercially available tools are difficult to perform comprehensive and in-depth analysis of each layer.
[0029] High-speed data processing challenges: FC network transmission rates are constantly increasing, from the early 2Gbps and 4Gbps to the current 16Gbps and even higher 32Gbps and 64Gbps. Traditional software-based analysis tools struggle to perform real-time line-rate analysis of high-speed FC data streams, often resulting in the loss of critical fault information.
[0030] Missing SCSI Layer Performance Parameters: The root causes of I / O read / write failures are often closely related to SCSI layer operations, such as SCSI command execution time, retries, and error codes. However, existing solutions lack comprehensive statistics and analysis of SCSI layer performance parameters, failing to provide accurate basis for fault localization.
[0031] Statistics show that in large data centers, approximately 30% of storage performance issues originate within the FC SAN storage network. Due to a lack of effective diagnostic tools, the average time to locate these problems exceeds four hours, severely impacting business continuity and efficiency. Therefore, there is an urgent need for a precise diagnostic method for FC SAN storage network I / O read / write failures to quickly locate faults and improve system reliability and operational efficiency.
[0032] Current related technologies include a monitoring method for FC SAN storage networks based on Simple Network Management Protocol (SNMP). This method collects performance metrics of storage devices, such as bandwidth utilization and error counts, by deploying an SNMP agent and then constructs a network topology map using topology discovery technology. However, SNMP is a polling-based monitoring method with a low sampling frequency, making it unable to capture sudden I / O failures in real time. Furthermore, this method only obtains macroscopic performance metrics at the device level and cannot delve into the FC protocol and SCSI layers for detailed analysis, making it difficult to pinpoint specific fault locations.
[0033] In addition, some related technologies have proposed a method of deploying software probes on the host side to analyze I / O performance parameters by intercepting I / O requests between the host and the FC HBA (Host Bus Adapter). However, software probes need to run on the host side, consuming host CPU and memory resources and impacting the performance of the business system. Furthermore, this method can only monitor I / O requests on the host side and cannot obtain raw data packet information transmitted in the FC network, making it difficult to analyze network link and storage device internal faults.
[0034] In summary, the relevant technologies have many shortcomings in diagnosing I / O read / write faults in FC SAN storage networks, mainly in the following aspects:
[0035] Insufficient real-time performance: Traditional monitoring methods such as SNMP polling and software probes cannot respond to sudden failures in FC SAN in real time, resulting in delays in fault location.
[0036] Insufficient depth of analysis: Most existing tools can only perform performance monitoring at the device level or host side, lacking in-depth analysis of each layer of the FC protocol and the SCSI layer, and are unable to obtain key performance parameters and fault characteristics.
[0037] Insufficient high-speed processing capability: As FC network speeds continue to increase, software-based analysis tools struggle to meet the demands of high-speed data processing, resulting in packet loss and processing delays.
[0038] Low fault location accuracy: Due to the lack of comprehensive protocol analysis and performance parameter statistics, it is impossible to accurately distinguish whether the fault occurs inside the database, the FC network link, or the storage device, causing maintenance personnel to spend a lot of time troubleshooting.
[0039] To address all or part of the aforementioned technical problems, embodiments of this application provide a method, apparatus, electronic device, and storage medium for diagnosing faults in a storage system. To enable those skilled in the art to better understand the solutions provided in this application, the following detailed description is provided in conjunction with the accompanying drawings and specific embodiments.
[0040] like Figure 1 As shown, Figure 1 A flowchart of a fault diagnosis method for a storage system provided in this application embodiment is included, which may include the following steps:
[0041] 101. Obtain message data from the storage system.
[0042] In this embodiment, the storage system may be an FC SAN storage network, and the message data may be a Fibre Channel packet (FC PACKET) obtained by line-speed capture of the data stream in the FC network.
[0043] In some embodiments, the FC PACKET is the basic unit for data transmission in the FC protocol. It is designed to achieve high-speed, reliable, and low-latency block-level data transmission, and is a core support for the performance of FC SAN networks. The FC PACKET uses frames as its basic unit. Each frame contains: a Frame Header (24 bytes) containing addressing information (such as the World Wide Name (WWN) of the source / destination ports) and transmission control flags (such as priority and flow control information); a Payload (maximum 2112 bytes) carrying the actual data (such as SCSI commands and disk block data); and a Frame Trailer (4 bytes, Cyclic Redundancy Check (CRC) used to detect transmission errors).
[0044] 102. Based on the multi-level structure of the storage system, the message data is parsed layer by layer to obtain the corresponding protocol information for each layer.
[0045] In this embodiment of the application, the storage system can be configured as a multi-level structure, which can include at least a system interface layer, namely the SCSI layer. Since the root cause of I / O read and write failures is often closely related to the operation of the SCSI layer, the protocol information corresponding to the SCSI layer should be considered when parsing message data.
[0046] In some embodiments, the multi-level architecture of the storage system may include: a physical layer, a data link layer, a network layer, a transport layer, and a SCSI layer, where the SCSI layer may be the highest layer in the multi-level architecture of the storage system. The captured FC PACKET can be parsed layer by layer from FC-0 to FC-4 to extract key protocol information.
[0047] The FC SAN is comprised of several layers: the Physical Layer (FC-0) defines the transmission medium, interface specifications, and signaling mechanisms, forming the physical foundation of the FC SAN; the Link Layer (FC-1) is responsible for frame encoding / decoding, error detection, and link control, ensuring reliable data transmission; the Network Layer (FC-2) defines the frame structure, flow control, and Quality of Service (QoS), serving as the core transport layer of the FC SAN; for the Transport Layer (FC-3), the FC protocol integrates some transport layer functions (such as end-to-end flow control and error recovery) into the Link Layer or Network Layer (FC-1 / FC-2), achieving reliable transmission through the BB_Credit mechanism and CRC checksum, eliminating the need for a separate transport layer protocol; and the SCSI Layer (FC-4) encapsulates SCSI commands into FC frames, enabling block-level storage access, and is crucial for the interaction between the FC SAN and storage devices.
[0048] In some embodiments, the message data can be parsed layer by layer according to the five levels described above, which may include the following implementation methods:
[0049] The physical layer performs signal detection, encoding / decoding, and error detection. Specifically, it analyzes the signal characteristics of the FC physical layer, such as signal strength and noise level, to determine the connection quality. It decodes the encoded data from the FC physical layer to recover the original bitstream. For example, for 8B / 10B encoding, it decodes 10 bits of encoded data into 8 bits of original data. It calculates the CRC checksum of the physical layer and compares it with the received CRC code to detect transmission errors. When an error is detected, it records the error type and location and reports it to the upper-layer protocol.
[0050] The link layer can perform frame boundary identification, frame control information extraction, and link status monitoring. Specifically, it identifies the Start of Frame (SOF) and End of Frame (EOF) using a frame delimitation state machine, detects the SOF and EOF markers, identifies the boundaries of FC frames, and ensures the correct parsing of each FC frame, completing the boundary identification of one frame within four clock cycles. It parses the control fields of the FC frame to obtain information such as frame type (e.g., data frame, command frame, response frame), priority, and destination. It tracks FC link state changes, such as link initialization, link login, and link disconnection, recording the time and reason for link state transitions.
[0051] The network layer can perform functions such as address resolution, routing information extraction, and network performance statistics. Specifically, it parses the source and destination node addresses in FC frames to obtain information such as the node's WWN and port name. For frames transmitted through FC switches, it parses the switching path information and constructs the FC network topology. It also collects network layer performance parameters, such as frame forwarding latency and switch buffer utilization, to evaluate the performance of the FC network.
[0052] The transport layer performs connection management, flow control processing, and error recovery mechanisms. Specifically, it tracks the connection establishment and release process of the FC transport layer, recording the connection status and parameters. It parses the transport layer's flow control information, such as credit management and the sending and receiving of flow control frames, ensuring smooth data transmission. It handles transport layer error recovery operations, such as frame retransmission and connection reset, guaranteeing the reliability of data transmission.
[0053] The SCSI layer can perform protocol data extraction, command type identification, and key parameter extraction. Specifically, it extracts SCSI protocol data from the data fields of the FC frame, including the SCSI Command Description Block (CDB), data buffer, and status response. It parses the opcodes in the CDB to identify the type of SCSI command, such as READ, WRITE, and INQUIRY. It extracts key parameters from the SCSI commands, such as LUN number, Logical Block Address (LBA), transfer length, and command priority, providing a basis for subsequent performance statistics and fault diagnosis.
[0054] 103. Determine the target performance data of the system interface layer based on the protocol information corresponding to each layer and the read / write operations of the system interface layer.
[0055] In this embodiment of the application, since storage system failures may exist in every layer of the multi-level structure, but the failures are basically related to the read and write operations of the SCSI layer, after parsing the protocol information of each layer, we can focus on determining the various read / write performance parameters of the SCSI layer for the read and write operations of the SCSI layer, and conduct a full-stack analysis from the physical layer (signal integrity) → transport layer (FC frame error) → SCSI layer (command timeout) to obtain the target performance data of the system interface layer.
[0056] 104. Based on the target performance data, perform fault diagnosis on the storage system and obtain the fault diagnosis results.
[0057] In this embodiment of the application, after determining the target performance data, the target performance data can be analyzed to diagnose the storage system and obtain the fault diagnosis results.
[0058] In this embodiment, data parsing can be performed at each level of the storage system to achieve full-level fault location. The performance data of the system interface layer, namely the SCSI layer, is analyzed in detail. This allows for accurate analysis of fault conditions in conjunction with read and write operations, enabling accurate diagnosis, precise location, and type determination of read and write faults in the storage system.
[0059] like Figure 2 As shown, Figure 2 Another flowchart of a fault diagnosis method for a storage system provided in an embodiment of this application, the method may include the following steps:
[0060] 201. Collect data streams through the target processing chip.
[0061] In this embodiment, the target processing chip can be a high-speed processing chip, such as an ASIC or FPGA. The target processing chip can be a core component configured in the analysis device, integrating a high-speed data receiving engine and a cache management unit to ensure no packet loss under high load. After the data stream is introduced into the analysis device, the target processing chip can receive the data stream through the FC physical layer interface. This target processing chip can be a chip with the same network configuration as the storage system and in full-traffic capture mode.
[0062] In some embodiments, hardware deployment can be performed before data collection. Analysis devices can be installed in suitable locations within the FC SAN network, such as by connecting a splitter to the mirror port of an FC switch or in the link between the storage array and the switch. This allows FC data streams to be introduced into the FC interface of the analysis device. This access method does not affect the normal operation of the original FC network, ensuring the stability of the business system. Then, the high-speed processing chip (such as an FPGA) can be configured, setting parameters such as the FC interface transmission rate (e.g., 16Gbps) and encoding method (e.g., 8B / 10B or 128B / 130B) to ensure consistency with the FC network configuration. The target processing chip can then be set to full-traffic capture mode to ensure the capture of all FC PACKET packets without missing any data. A sufficiently large high-speed cache (e.g., 8MB) can be configured inside the chip for temporary storage of captured FC packets, preventing packet loss due to processing delays. The FC data stream is introduced into the analysis device through the splitter or the switch mirror port in the FC network, and then the target processing chip receives the data stream from the FC physical layer interface.
[0063] It's important to note that the hardware architecture of the analysis device uses a dedicated high-speed processing chip (such as an ASIC or FPGA) as its core processing unit, along with components like a high-speed FC interface chip, to achieve line-speed capture of the FC data stream. Configuring the high-speed processing chip (such as the FPGA) is a configuration operation targeting this core component. By setting parameters such as the target processing chip's transmission rate and encoding method, it ensures that the analysis device's configuration matches the FC network, thereby achieving accurate data stream capture. This is fundamental to the normal operation of the analysis device. In summary, the analysis device is a hardware device that includes the target processing chip (such as an FPGA), and the target processing chip is the core processing unit of the analysis device; the two are related as a whole and its core component.
[0064] 202. Perform data splitting processing on the data stream to obtain message data.
[0065] In this embodiment of the application, after the data stream is acquired, the data stream can be split into different processing queues. Specifically, different data streams can be assigned to different processing queues based on information such as the source address, destination address, and logical unit number (LUN) of the FC frame. That is, data streams belonging to the same LUN can be assigned to the same queue for subsequent protocol parsing and performance statistics.
[0066] In some embodiments, a dedicated processing chip (such as an ASIC or FPGA) is used to achieve line-rate analysis of data streams at speeds of 16Gbps and above, ensuring no packet loss and no latency, thus meeting the fault diagnosis requirements in a high-speed FC SAN environment.
[0067] In some embodiments, after acquiring the data stream, the data stream can be preprocessed by signal amplification, equalization, and clock recovery to eliminate attenuation and distortion during signal transmission and ensure the accuracy of the received data.
[0068] During transmission, the signal amplitude may decrease due to dielectric loss, distance attenuation, or noise interference, potentially falling below the sensitivity threshold of the receiving device. Signal amplification can be achieved by using electronic circuits (such as operational amplifiers) or optical devices (such as photomultiplier tubes) to proportionally enhance the voltage or current of the weak signal, restoring it to a manageable range. This can be understood as improving signal strength and signal-to-noise ratio (SNR) to counteract transmission loss and ensure that the signal amplitude is higher than the noise floor of the receiving device, thus avoiding data loss or misjudgment due to a weak signal.
[0069] Among them, the transmission medium (such as cable, optical fiber) has different attenuation degrees for signals of different frequencies, resulting in greater loss of high-frequency components and causing signal waveform distortion (such as inter-symbol interference). Signal equalization can compensate for the frequency domain characteristics through filters or digital algorithms, so that the attenuation of each frequency component tends to be consistent. By compensating for frequency domain distortion, inter-symbol interference is eliminated, the original waveform of the signal is restored, and the receiver can correctly decode the data.
[0070] In high-speed serial transmission, the clock signal is usually embedded in the data stream through data encoding (such as 8b / 10b). The receiving end needs to extract the clock component from the random data for synchronous sampling. Clock recovery can be achieved by tracking the data phase through a phase-locked loop (PLL) or a delay-locked loop (DLL) to generate a sampling clock synchronized with the transmitting end. The synchronous clock is extracted from the data stream to achieve accurate sampling, ensuring that the receiving end samples the data at the correct time and avoiding bit errors caused by clock offset.
[0071] In some embodiments, for data streams, the cache management unit inside the target processing chip can allocate multiple data cache queues for each FC port and schedule the queues according to the priority and type of the FC frame to avoid delays in high-priority data.
[0072] In FC (Focus-Frame) networks, the basic data unit for transmission has a priority attribute assigned to the data during transmission (a characteristic inherent to the acquired data itself), used to identify the importance or processing priority of the data. By scheduling queues according to the FC frame's own priority, it can be ensured that high-priority data (such as critical I / O request frames) will not be delayed due to the processing of low-priority data, thus guaranteeing the real-time performance and processing efficiency of core data. This is a crucial mechanism for cache management in high-speed data acquisition modules.
[0073] In some embodiments, when acquiring data streams, data flow can be monitored in real time. When the input data rate exceeds the chip's processing capacity, the flow control mechanism temporarily stops data transmission by sending a PAUSE frame to the FC network to prevent data loss.
[0074] 203. Input the message data into the first data layer of the multi-level structure of the storage system so that the first data layer can parse the message data and obtain the protocol information corresponding to the first data layer.
[0075] In this embodiment, the storage system has a multi-level structure, and data can be transmitted between adjacent layers. When parsing message data through the storage system, the message data can be parsed through each layer. That is, the input of each layer includes message data. In addition, since adjacent layers are interconnected, the input of the upper layer can also include the output of the lower layer. In other words, for the first data layer in the multi-level structure, since the first data layer is the bottom layer of the entire storage system, the input of the first data layer only includes message data. That is, the first data layer only parses the message data to obtain the protocol information corresponding to the first data layer.
[0076] 204. Input the message data and the protocol information corresponding to the (n-1)th data layer into the nth data layer of the multi-level structure of the storage system, so that the nth data layer can parse the message data and the protocol information corresponding to the (n-1)th data layer to obtain the protocol information corresponding to the nth data layer.
[0077] In the embodiments of this application, in the hierarchical structure of the storage system, in addition to the bottom layer, namely the first data layer, the input of other data layers may include, in addition to message data, the protocol information parsed by the previous data layer. That is to say, message data and the protocol information corresponding to the previous data layer can be parsed together.
[0078] It should be noted that in the context of a multi-layered data structure, where n is an integer greater than 1, this nth data layer can be the second data layer or a higher data layer. It can be understood as a data layer with a lower layer (i.e., the (n-1)th data layer). After the (n-1)th data layer completes parsing and obtains the corresponding protocol information, it can send the message data and the corresponding protocol information together to the nth data layer. The nth data layer can then parse the message data and the corresponding protocol information to obtain the protocol information for the nth data layer. Furthermore, if an (n+1)th data layer exists, the nth data layer can send the message data and the corresponding protocol information together to the (n+1)th data layer.
[0079] In some embodiments, in the storage system disclosed in this application, the physical layer is the lowest layer. Therefore, the physical layer can parse the message data to obtain the protocol information corresponding to the physical layer; then the link layer can parse the message data and the protocol information corresponding to the physical layer to obtain the protocol information corresponding to the link layer; then the network layer can parse the message data and the protocol information corresponding to the link layer to obtain the protocol information corresponding to the network layer; then the transport layer can parse the message data and the protocol information corresponding to the network layer to obtain the protocol information corresponding to the transport layer; and then the SCSI layer can parse the message data and the protocol information corresponding to the transport layer to obtain the protocol information corresponding to the SCSI layer.
[0080] In some embodiments, detailed information of each FC layer is analyzed in depth, and comprehensive performance parameters are statistically analyzed. This enables accurate differentiation of faults within the database, FC network links, and storage devices, with positioning accuracy reaching the device component level. Data at each layer is clearly distinguished during the parsing process, and the information extracted from each layer is targeted, providing a data foundation for the parsing of the upper layer layer by layer, ensuring the independence and accuracy of protocol information at each layer.
[0081] 205. Based on the protocol information corresponding to each layer, determine the classification information of at least one target command.
[0082] In this embodiment of the application, after parsing the protocol information corresponding to each layer, the SCSI operations can be classified and statistically analyzed to understand the distribution of different types of commands. Specifically, the SCSI commands parsed in real time can be classified according to the SCSI command opcode, such as read commands, write commands, control commands, etc. Then, the number of each type of command is counted according to the time interval (such as per second, per minute) to generate a command type distribution histogram in order to understand the composition of I / O load and obtain the classification information of at least one target command.
[0083] 206. Based on the classification information of at least one target command, determine the read and write operations corresponding to each target command, and perform performance analysis on each target command and its corresponding read and write operations to obtain the target performance data of the system interface layer.
[0084] In this embodiment of the application, different target commands may correspond to different read and write operations. Therefore, the target commands and read and write operations can be mapped together to determine the read and write operations corresponding to at least one target command. Then, performance analysis is performed on the target commands and their corresponding read and write operations to obtain the target performance data of the system interface layer. The target performance data may include at least command response time, data transmission volume, number of retries, and error code data.
[0085] In some embodiments, the target performance data can be calculated separately. Specifically, performance analysis is performed on at least one target command and its corresponding read / write operations to obtain the target performance data of the system interface layer. This can include: determining the command response time based on the sending and response times of the target command; determining the data transmission volume based on the data transmitted in the read / write operations; resending the target command and updating the retry count when no response command is received within a preset time period or when an error code exists in the response command; extracting error codes from the fields of the target command and classifying and statistically analyzing them according to error type to obtain error code data.
[0086] Each SCSI command can maintain a timestamp record. When a SCSI command is sent, the sending time is recorded. When the corresponding response frame is received, the response time is recorded. The difference between the two is the command response time. The execution latency of the command is calculated to determine the response performance of the storage device. For commands and responses transmitted across multiple FC frames, they are associated through frame sequence numbers.
[0087] This allows for the counting and summing of the data transferred in each SCSI read / write operation to obtain the total data transfer volume, thus analyzing the I / O load.
[0088] Specifically, if a SCSI command is sent but no response is received within a specified time, or if the received response contains an error code, the number of retries for that command is recorded. The processing of the retry command is tracked until the command is successfully executed or ultimately fails. An excessive number of retries may indicate problems such as network link instability or storage device failure.
[0089] This involves parsing the status fields and Sense data in the SCSI response, extracting error codes, and classifying and statistically analyzing them by error type, such as media errors, device errors, and parameter errors, and calculating the frequency of each type of error.
[0090] In some embodiments, based on in-depth analysis of the protocol information at each FC layer, further in-depth statistical analysis of the SCSI layer read / write performance parameters can achieve rapid and accurate location of I / O read / write faults in the FC SAN storage system, enabling full-process analysis of each layer structure and effectively improving the accuracy of fault diagnosis and location.
[0091] 207. Store the target performance data of the system interface layer into the database according to the time series, and create field indexes for the target performance data.
[0092] In this embodiment of the application, after determining the target performance data of the system interface layer, the target performance parameters can be stored in the database in a time series for subsequent historical data analysis and fault diagnosis.
[0093] The stored time series format can be: timestamp, LUN, command type, command response time, data transfer volume, number of retries, and error code data.
[0094] In some embodiments, the database may be a database suitable for storing time-series data, such as InfluxDB.
[0095] It should be noted that after storing the data in the database, field indexes can be created on the target performance data, which can include: timestamp indexes, LUN indexes, command type indexes, etc.
[0096] In some embodiments, storing performance data in a database according to a time series and creating field indexes can facilitate subsequent data queries and retrieval, thereby improving the efficiency of data query and analysis.
[0097] 208. When the storage system is in normal operation, continuously collect historical message data at multiple times corresponding to the storage system.
[0098] In this embodiment of the application, a performance baseline can be established before the performance data is detected. The performance baseline can be understood as the normal range of the performance indicators. If the performance exceeds the performance baseline, there may be an abnormal situation. Therefore, the performance baseline can be established based on the performance data during the historical normal operation process. Then, it is necessary to continuously collect historical message data corresponding to multiple moments of the storage system when the storage system is in normal operation. The historical message data can be continuous data over a period of time, such as one week or one month.
[0099] 209. Perform statistical analysis on historical message data at multiple times to obtain preset performance baselines corresponding to multiple performance indicators.
[0100] In this embodiment, the process of statistically analyzing historical message data from multiple time points is similar to the message data processing process. This involves inputting the historical message data into the multi-layered structure of the storage system for layer-by-layer parsing, and then statistically obtaining historical performance data for the system interface layer. Since historical message data from multiple time points is collected, corresponding historical performance data from multiple time points will also be obtained. Furthermore, the historical performance data may include multiple performance indicators. Therefore, for each performance indicator, statistical measures such as the average, standard deviation, and percentile of the performance data corresponding to multiple time points can be calculated to establish preset performance baselines for each performance indicator. For example, the average and 95th percentile of the SCSI read command response time can be calculated as a performance reference under normal conditions.
[0101] 210. When at least one performance indicator in the target performance data is detected to exceed the corresponding warning threshold, a warning event is triggered to perform fault diagnosis on the storage system based on the target performance data and obtain the fault diagnosis result.
[0102] In this embodiment of the application, the preset performance baseline can be pre-calculated based on historical message data. After obtaining the target performance data in the current scenario, the warning threshold corresponding to each performance indicator can be determined based on the preset performance baseline. Then, each performance indicator included in the target performance data is compared with the corresponding warning threshold. If at least one performance indicator exceeds the corresponding warning threshold, it indicates that there may be an abnormal situation. Therefore, a warning event can be triggered. The warning event can indicate that fault diagnosis of the storage system is required. Therefore, fault diagnosis of the storage system can be performed based on the target performance data to obtain the fault diagnosis result.
[0103] In some embodiments, by setting a warning threshold determined based on a performance baseline, a preliminary fault can be identified. Only when the performance index exceeds the warning threshold will subsequent precise fault diagnosis be performed. This avoids performing precise diagnosis on all data, reduces workload, and improves fault diagnosis efficiency.
[0104] In some embodiments, the aforementioned warning threshold is determined based on a preset performance baseline for the corresponding performance indicator. Furthermore, the warning threshold can also be determined based on the preset performance baseline corresponding to the performance indicator and current business requirements. That is, the specific value of the preset performance baseline can be adjusted based on current business requirements to obtain the warning threshold.
[0105] In some embodiments, when comparing performance indicators with corresponding warning thresholds, if the performance indicator exceeds the corresponding warning threshold, the warning level can be determined based on the degree of deviation and the duration. The degree of deviation is the difference between the performance indicator and the corresponding warning threshold, and the duration is the statistical duration of the performance indicator exceeding the corresponding warning threshold. Warning levels can include: warning, severe warning, fault, etc. Specific fault diagnosis strategies can be determined for different warning levels. For example, when the warning level is warning, fault diagnosis can be delayed, and changes in performance indicators can be observed. When the warning level is severe warning, further fault diagnosis can be performed. When the warning level is fault, fault reminder information can be directly output to staff, etc.
[0106] In some embodiments, when diagnosing faults in a storage system, in addition to the target performance data, fault diagnosis can also be performed based on the protocol information corresponding to each layer obtained by parsing the message data layer by layer. In other words, the fault can be accurately located by analyzing the correlation between the target performance data and the protocol information corresponding to each layer through the whole-level information.
[0107] In some embodiments, when diagnosing faults in a storage system, a combination of machine learning algorithms (such as learning models) and rule engines can be used to analyze target performance parameters. That is, based on target performance data, fault diagnosis of the storage system is performed to obtain fault diagnosis results. Specifically, this may include: diagnosing faults in the storage system based on target performance data, a pre-stored fault rule base, and a pre-trained fault diagnosis model to obtain fault diagnosis results.
[0108] It should be noted that the fault rule base can be a fault location rule base built on a rule engine, which can store fault phenomena, diagnostic rules, optimization suggestions, etc.; the fault diagnosis model can be a machine learning model obtained after pre-training, and performance data can be input into the fault diagnosis model to obtain diagnostic results.
[0109] Furthermore, based on the target performance data, the pre-stored fault rule base, and the pre-trained fault diagnosis model, the storage system is subjected to fault diagnosis to obtain fault diagnosis results. Specifically, this may include: extracting features from the target performance data and early warning events to obtain fault features; determining a first diagnosis result based on the fault features and the fault diagnosis model; determining a second diagnosis result based on the fault features and the fault rule base; and fusing the first and second diagnosis results to determine the final fault diagnosis result.
[0110] In this embodiment, fault features can be extracted from performance parameters and warning events, such as a sudden increase in response time, a sharp rise in the number of retries, and frequent occurrence of specific error codes. Then, the fault features are input into the fault diagnosis model to obtain the first diagnostic result output by the fault diagnosis model. At the same time, the fault features can be matched with data on various fault phenomena stored in the fault rule base to determine the data that matches the fault features as the second diagnostic result. Both the first and second diagnostic results are diagnostic results about the fault features. Finally, the first and second diagnostic results can be fused to obtain the final fault diagnosis result.
[0111] Furthermore, the process of fusing diagnostic results can be achieved through weighted summation. That is, the first diagnostic result and the second diagnostic result are fused to determine the fault diagnosis result. Specifically, this can include: weighted summation of the abnormal data deviation values indicated in the first diagnostic result and the abnormal data deviation values indicated in the second diagnostic result to obtain a fault probability score; and determining the fault diagnosis result based on the preset score range corresponding to the fault probability score.
[0112] It should be noted that the diagnostic results may indicate abnormal performance indicators and specific abnormalities, namely, the abnormal data deviation values of the abnormal indicators. These abnormal data deviation values can be understood as the difference between the abnormal indicator values and the normal values. The abnormal indicators indicated in the first and second diagnostic results may be the same or different. The corresponding weights can be determined for the indicators indicated in the first and second diagnostic results, and then the abnormal data deviation values of each indicator are weighted and summed according to their corresponding weights to obtain a fault probability score. Different score ranges can be pre-defined to correspond to different fault degrees, thus determining the score range in which the fault probability score falls, thereby obtaining the corresponding fault diagnosis result.
[0113] In some embodiments, both the first diagnostic result and the second diagnostic result may include specific fault results and fault probabilities. It is understood that the fault diagnosis model and the fault rule base are based on performance data for fault diagnosis, rather than detecting faults in actual scenarios. Therefore, the fault diagnosis results cannot be guaranteed to be absolutely correct. In addition, sometimes the fault phenomenon may correspond to multiple fault causes and fault types. Therefore, the first diagnostic result and the second diagnostic result may include multiple fault types and their corresponding probabilities.
[0114] Specifically, based on a fault rule base (such as rules indicating a large number of media error codes and increased response time corresponding to disk media failures), real-time extracted fault features (such as error code types and response time changes) are matched with fault patterns in the rule base. Combined with the frequency of occurrence of this pattern in historical cases, a preliminary probability range is determined. For the fault diagnosis model, after pre-training, it can automatically output the matching probability of each fault type based on the input fault features (such as the number of retries and error code distribution). This probability is directly calculated and generated by the algorithm model.
[0115] This allows us to sort by probability and determine the most likely type of failure.
[0116] In some embodiments, the process of acquiring the fault rule base and the fault diagnosis model may include: acquiring a first sample dataset when the storage system is in normal operation, and a second sample dataset and corresponding fault phenomena when the storage system is in fault operation; performing feature parsing on the first sample dataset and the second sample dataset to obtain a normal feature set and a fault feature set; training a preset model based on the normal feature set, the fault feature set, and the fault phenomena to obtain a fault diagnosis model; and storing the fault feature set and the fault phenomena accordingly to obtain a fault rule base.
[0117] It should be noted that the fault diagnosis model needs to be pre-trained. It can be understood that the fault diagnosis model performs fault diagnosis by analyzing the correlation between features and phenomena. Therefore, it is possible to obtain the first sample dataset when the storage system is fault-free, and the second sample dataset and corresponding fault phenomena when the storage system is faulty. After feature parsing, normal feature set and fault feature set are obtained. Then, the model learns from the normal feature set, fault feature set and fault phenomena to generate the fault diagnosis model.
[0118] It should be noted that since the fault rule base directly matches the corresponding fault type based on fault features, it can be understood that the fault rule base stores a large amount of data such as fault phenomena, fault causes, fault types, and fault features. Therefore, it is possible to obtain the second sample dataset and the corresponding fault phenomena when the storage system fails, perform feature parsing on the second sample dataset to obtain the fault feature set, and store the fault feature set and the fault phenomena accordingly to generate the fault rule base.
[0119] In some embodiments, a fault diagnosis system that integrates a rule engine and a machine learning model through both a fault rule base and a fault diagnosis model performs fault diagnosis simultaneously and outputs fault diagnosis results. This enables automatic identification and precise location of I / O read / write faults in FC SAN storage systems, reduces reliance on the experience of maintenance personnel, improves the accuracy and consistency of fault diagnosis, and shortens the diagnosis time.
[0120] 211. Based on the fault diagnosis results, output the corresponding fault warning information and fault handling strategy.
[0121] In this embodiment of the application, after obtaining the fault diagnosis result, the fault diagnosis result can be output to the staff. Specifically, fault warning information and fault handling strategy can be output. The fault warning information is used to indicate that there may be a fault in the current storage system, and the fault handling strategy is the handling strategy corresponding to the fault that may exist.
[0122] In some embodiments, operations and maintenance personnel may be allowed to configure personalized early warning rules according to business needs, including early warning thresholds, early warning levels, and early warning notification methods.
[0123] In some embodiments, when an early warning event occurs, relevant operations and maintenance personnel can be notified via various means such as email, SMS, and instant messaging tools.
[0124] In some embodiments, corresponding handling suggestions can be generated based on the fault diagnosis results and historical fault handling experience, such as checking FC link connections, restarting the storage controller, and replacing the faulty disk, to provide decision support for operation and maintenance personnel.
[0125] In some embodiments, visualization tools such as Grafana can be used to display the performance metrics and failure status of the FC SAN storage network in the form of charts.
[0126] In some embodiments, key performance indicators of the FC SAN storage network, such as SCSI command response time, data transfer rate, and error code statistics, can be displayed in real time in the form of dashboards, graphs, etc.
[0127] Specifically, an intuitive dashboard interface can be designed to display key performance indicators of the FC SAN storage network. Each indicator can be displayed in the form of charts, digital cards, etc., to help maintenance personnel quickly understand the system status.
[0128] Specifically, based on the topology information obtained from the FC network layer, a network topology map of the FC SAN can be generated, which can intuitively display the connection relationships and status of storage devices, switches, hosts and other devices.
[0129] Specifically, when the fault diagnosis result indicates an early warning, the alarm information, including the alarm level, fault location, and cause, can be displayed on the monitoring interface in a prominent manner (such as a red alarm box, sound prompts, etc.).
[0130] In some embodiments, historical data query and analysis functions can be provided, supporting filtering and statistics by time, device, command type and other dimensions to help operation and maintenance personnel discover potential performance trends and failure modes.
[0131] Specifically, it can provide flexible query condition settings, allowing operation and maintenance personnel to query historical performance data by time range, device type, LUN number, command type and other dimensions.
[0132] Specifically, the retrieved historical data can be displayed in the form of line charts, bar charts, pie charts, etc., to facilitate the analysis of performance trends and failure modes. For example, it can show the trend of SCSI read command response time over the past 24 hours.
[0133] Specifically, it can support comparative analysis of performance data from different time periods and different devices, helping maintenance personnel to discover performance differences and potential problems.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0135] like Figure 3 As shown, embodiments of this application also provide a fault diagnosis device for a storage system, which may include:
[0136] The acquisition module 301 is used to acquire message data from the storage system;
[0137] Processing module 302 is used to parse message data layer by layer based on the multi-level structure of the storage system to obtain the protocol information corresponding to each layer. The multi-level structure includes at least a system interface layer, which is the highest layer in the multi-level structure of the storage system.
[0138] The processing module 302 is also used to determine the target performance data of the system interface layer based on the protocol information corresponding to each layer and the read and write operations of the system interface layer;
[0139] The processing module 302 is also used to perform fault diagnosis on the storage system based on the target performance data and obtain fault diagnosis results.
[0140] In some embodiments, the acquisition module 301 is specifically used to acquire data streams through a target processing chip, wherein the target processing chip is a chip that has the same network configuration as the storage system and is in a full-traffic capture state;
[0141] The processing module 302 is specifically used to perform data splitting processing on the data stream to obtain message data.
[0142] In some embodiments, the processing module 302 is specifically used to input message data into the first data layer in the multi-level structure of the storage system, so that the first data layer parses the message data and obtains the protocol information corresponding to the first data layer.
[0143] The processing module 302 is specifically used to input the message data and the protocol information corresponding to the (n-1)th data layer into the nth data layer in the multi-level structure of the storage system, so that the nth data layer can parse the message data and the protocol information corresponding to the (n-1)th data layer to obtain the protocol information corresponding to the nth data layer; n is an integer greater than 1.
[0144] In some embodiments, the processing module 302 is specifically used to determine the classification information of at least one target command based on the protocol information corresponding to each layer.
[0145] The processing module 302 is specifically used to determine the read and write operations corresponding to at least one target command based on the classification information of at least one target command, and to perform performance analysis on at least one target command and its corresponding read and write operations to obtain target performance data of the system interface layer. The target performance data includes at least command response time, data transmission volume, number of retries and error code data.
[0146] In some embodiments, the processing module 302 is specifically used to determine the command response time based on the sending time and response time of the target command;
[0147] Processing module 302 is specifically used to determine the amount of data transmitted based on the data transmitted during the read / write operation;
[0148] The processing module 302 is specifically used to resend the target command and update the retry count when it detects that no response command has been received within a preset time period, or when there is an error code in the response command.
[0149] The processing module 302 is specifically used to extract error codes from the fields of the target command, classify and statistically analyze them according to error type, and obtain error code data.
[0150] In some embodiments, the processing module 302 is further configured to store the target performance data of the system interface layer into the database according to the time series, and to establish a field index for the target performance data.
[0151] In some embodiments, the processing module 302 is specifically used to trigger an early warning event when it is detected that at least one performance indicator in the target performance data exceeds the corresponding early warning threshold, so as to perform fault diagnosis on the storage system based on the target performance data and obtain the fault diagnosis result.
[0152] The warning threshold is determined based on the preset performance baseline of the corresponding performance index.
[0153] In some embodiments, the acquisition module 301 is further configured to continuously collect historical message data at multiple times corresponding to the storage system when the storage system is in normal operation.
[0154] The processing module 302 is also used to perform statistical analysis on historical message data at multiple times to obtain preset performance baselines corresponding to multiple performance indicators.
[0155] In some embodiments, the processing module 302 is specifically used to perform fault diagnosis on the storage system based on target performance data, a pre-stored fault rule base and a pre-trained fault diagnosis model, and obtain fault diagnosis results.
[0156] In some embodiments, the processing module 302 is specifically used to extract features from the target performance data and early warning events to obtain fault features;
[0157] The processing module 302 is specifically used to determine the first diagnostic result based on the fault characteristics and the fault diagnosis model;
[0158] The processing module 302 is specifically used to determine the second diagnostic result based on the fault characteristics and the fault rule base;
[0159] The processing module 302 is specifically used to integrate the first diagnostic result and the second diagnostic result to determine the fault diagnosis result.
[0160] In some embodiments, the acquisition module 301 is further configured to acquire a first sample dataset when the storage system is in normal operation, and a second sample dataset and corresponding fault phenomena when the storage system is in fault operation.
[0161] The processing module 302 is also used to perform feature parsing on the first sample dataset and the second sample dataset respectively to obtain the normal feature set and the fault feature set;
[0162] The processing module 302 is also used to train the preset model based on the normal feature set, the fault feature set and the fault phenomenon to obtain the fault diagnosis model;
[0163] The processing module 302 is also used to store the fault feature set and the fault phenomenon in correspondence to obtain the fault rule base.
[0164] In some embodiments, the processing module 302 is specifically used to perform a weighted summation of the abnormal data deviation value indicated in the first diagnostic result and the abnormal data deviation value indicated in the second diagnostic result to obtain a fault probability score.
[0165] The processing module 302 is specifically used to determine the fault diagnosis result based on the preset scoring range corresponding to the fault probability score.
[0166] In the embodiments of this application, the description of the features corresponding to the fault diagnosis device of the storage system can be found in the relevant description of the embodiment corresponding to the fault diagnosis method of the storage system, and will not be repeated here.
[0167] like Figure 4As shown, embodiments of this application also provide an electronic device, including a memory 401 and a processor 402. The memory 401 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above-described embodiments of the fault diagnosis method for a storage system.
[0168] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the fault diagnosis method for storage systems when it is run.
[0169] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0170] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the embodiments of any of the above-described storage system fault diagnosis methods.
[0171] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the embodiments of the fault diagnosis method for any of the above-described storage systems.
[0172] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0173] The foregoing has provided a detailed description of a fault diagnosis method, apparatus, electronic device, and storage medium for a storage system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A fault diagnosis method for a storage system, characterized in that, include: Obtain message data from the storage system; The message data is input into the first data layer in the multi-level structure of the storage system, so that the first data layer can parse the message data to obtain the protocol information corresponding to the first data layer. The message data and the protocol information corresponding to the (n-1)th data layer are input into the nth data layer of the multi-level structure of the storage system, so that the nth data layer parses the message data and the protocol information corresponding to the (n-1)th data layer to obtain the protocol information corresponding to the nth data layer; n is an integer greater than 1, and the multi-level structure includes at least the physical layer, link layer, network layer, transport layer and system interface layer, and the system interface layer is the highest layer in the multi-level structure of the storage system; Based on the protocol information corresponding to each layer and the read / write operations of the system interface layer, the target performance data of the system interface layer is determined. Based on the target performance data, fault diagnosis is performed on the storage system to obtain fault diagnosis results.
2. The method according to claim 1, characterized in that, The step of obtaining the message data of the storage system includes: The data stream is acquired through a target processing chip, which is a chip with the same network configuration as the storage system and is in a full-traffic capture state; The data stream is split into multiple streams to obtain the message data.
3. The method according to claim 1, characterized in that, The step of determining the target performance data of the system interface layer based on the protocol information corresponding to each layer and the read / write operations of the system interface layer includes: Based on the protocol information corresponding to each layer, at least one classification information for a target command is determined; Based on the classification information of the at least one target command, determine the read and write operations corresponding to the at least one target command respectively, and perform performance analysis on the at least one target command and the corresponding read and write operations to obtain the target performance data of the system interface layer. The target performance data includes at least command response time, data transmission volume, number of retries and error code data.
4. The method according to claim 3, characterized in that, The performance analysis of the at least one target command and its corresponding read / write operations to obtain the target performance data of the system interface layer includes: The command response time is determined based on the sending and response times of the target command. The amount of data transmitted is determined based on the data transmitted during the read / write operation; If no response command is received for the target command within a preset time period, or if an error code is found in the response command, the target command is resent and the number of retries is updated. Error codes are extracted from the fields of the target command and categorized and statistically analyzed according to error type to obtain the error code data.
5. The method according to claim 3, characterized in that, After determining the read / write operations corresponding to the at least one target command based on the classification information of the at least one target command, and performing performance analysis on the at least one target command and its corresponding read / write operations to obtain the target performance data of the system interface layer, the method further includes: The target performance data of the system interface layer is stored in the database according to the time series, and a field index is created for the target performance data.
6. The method according to claim 1, characterized in that, The step of performing fault diagnosis on the storage system based on the target performance data to obtain fault diagnosis results includes: When it is detected that at least one performance indicator in the target performance data exceeds the corresponding warning threshold, a warning event is triggered to perform fault diagnosis on the storage system based on the target performance data and obtain the fault diagnosis result. The warning threshold is determined based on a preset performance baseline of the corresponding performance index.
7. The method according to claim 6, characterized in that, The method further includes: When the storage system is in normal operation, historical message data at multiple times corresponding to the storage system are continuously collected; Statistical analysis is performed on the historical message data at the multiple time points to obtain preset performance baselines corresponding to multiple performance indicators.
8. The method according to claim 6, characterized in that, The step of performing fault diagnosis on the storage system based on the target performance data to obtain the fault diagnosis result includes: Based on the target performance data, the pre-stored fault rule base, and the pre-trained fault diagnosis model, the storage system is subjected to fault diagnosis to obtain the fault diagnosis results.
9. The method according to claim 8, characterized in that, The step of performing fault diagnosis on the storage system based on the target performance data, a pre-stored fault rule base, and a pre-trained fault diagnosis model, and obtaining the fault diagnosis result, includes: Feature extraction is performed on the target performance data and the early warning event to obtain fault features; Based on the fault characteristics and the fault diagnosis model, a first diagnostic result is determined; Based on the fault characteristics and the fault rule base, a second diagnostic result is determined; The fault diagnosis result is determined by combining the first diagnostic result and the second diagnostic result.
10. The method according to claim 9, characterized in that, The method further includes: A first sample dataset and a second sample dataset, along with the corresponding fault phenomena, are obtained when the storage system is in normal operating condition and when the storage system is in fault operating condition. Feature parsing is performed on the first sample dataset and the second sample dataset respectively to obtain a normal feature set and a fault feature set; The fault diagnosis model is obtained by training the preset model based on the normal feature set, the fault feature set, and the fault phenomenon. The fault feature set and the fault phenomenon are stored in correspondence to obtain the fault rule base.
11. The method according to claim 9, characterized in that, The process of fusing the first diagnostic result and the second diagnostic result to determine the fault diagnosis result includes: The abnormal data deviation values indicated in the first diagnostic result and the abnormal data deviation values indicated in the second diagnostic result are weighted and summed to obtain the fault probability score; The fault diagnosis result is determined based on the preset scoring range corresponding to the fault probability score.
12. A fault diagnosis device for a storage system, characterized in that, include: The acquisition module is used to acquire message data from the storage system; The processing module is used to input the message data into the first data layer in the multi-level structure of the storage system, so that the first data layer can parse the message data and obtain the protocol information corresponding to the first data layer. The processing module is further configured to input the message data and the protocol information corresponding to the (n-1)th data layer into the nth data layer in the multi-level structure of the storage system, so that the nth data layer parses the message data and the protocol information corresponding to the (n-1)th data layer to obtain the protocol information corresponding to the nth data layer. n is an integer greater than 1, and the multi-level structure includes at least a system interface layer, which is the highest layer in the multi-level structure of the storage system. The processing module is also used to determine the target performance data of the system interface layer based on the protocol information corresponding to each layer and the read and write operations of the system interface layer. The processing module is also used to perform fault diagnosis on the storage system based on the target performance data, and obtain fault diagnosis results.
13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the fault diagnosis method for the storage system as described in any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the fault diagnosis method for the storage system as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Data monitoring method and device, electronic equipment and computer readable storage medium
CN113934593A